The Software Efficiency Report · From the Founder's Desk
The Software Efficiency Report | 2026 Week 34
The 90-day rule: why engineering transformations succeed or fail in the first three months
Welcome to this edition of newsletter. Engineering teams can become less efficient long before the numbers show it clearly. A slow CI pipeline, an expired feature flag, or a growing gap between product discovery and engineering execution can quietly turn into a much larger delivery problem.
This week’s Software Efficiency Report looks at why CI feedback time matters to developer flow, how forgotten feature flags become technical debt, and whether faster software delivery is creating a new bottleneck in product discovery.
We also explore major shifts across modern engineering, including Kubernetes and AI workloads, multi-agent DevOps, continuous authorization, software supply chain security, cloud infrastructure and Edge AI.
The deep dive focuses on the first 90 days of an engineering transformation and what teams should actually achieve during that period: understanding the real constraints, proving a measurable improvement, and transferring ownership to the internal team.
Alongside this, the issue brings together this week’s cloud, open-source, DevOps, security, AI and embedded systems developments, plus a few tools and communities worth watching.
.
- Metric of the week
- CI Pipeline Feedback Time: Keep It Under 10 Minutes
- Deep dive
- The 90-day rule: why engineering transformations succeed or fail in the first three months
Software Efficiency Metric of the Week
CI Pipeline Feedback Time: Keep It Under 10 Minutes
CI Pipeline Feedback Time measures the total duration from the moment an engineer pushes a commit or opens a pull request to when they receive definitive build, lint, and automated test results. Keeping this turnaround under 10 minutes is the baseline for preserving developer flow. Once feedback runs past 15 to 20 minutes, developers stop waiting and shift their attention elsewhere.
The real cost is not just cloud runner hours or idle time. It is the heavy cognitive tax of context switching. When feedback is slow, developers open a new branch or start a different task while waiting. When the build inevitably fails 20 minutes later, they must drop what they are doing, reload mental context, and diagnose the issue. This cycle directly drives up pull request idle time, causes larger and riskier batch sizes, and tempts teams to skip running tests locally.
What I have noticed is that pipeline bloat usually happens silently over time as teams keep adding comprehensive integration and end-to-end suites directly into the default pull-request trigger. Teams should aggressively parallelize test runs, optimize Docker layer caching, and separate fast sanity suites from heavy end-to-end checks that can run asynchronously or nightly. Treating pipeline duration like an operational Service Level Objective (SLO) ensures build performance is defended just like production latency.
Formula:
CI Pipeline Feedback Time = Average elapsed wall-clock time from commit push to final automated verification status across all pull-request workflows during the measurement window.
Interesting readings: Source Source Source Source
Reader Poll
Should AI flip the traditional ratio of Product Managers to Software Engineers?
My take: AI pioneer Andrew Ng recently noted a striking shift: as AI coding assistants dramatically accelerate engineering velocity, the historical ratio of one PM to four (or more) engineers is being challenged, with some teams experimenting with ratios as extreme as two PMs for every one engineer. On the surface, two PMs sharing a single developer sounds like an absolute nightmare of meetings and ticket-writing. But the underlying logic is profound: when the cost of building software drops and engineers can ship 10 times faster, the operational bottleneck shifts from “how fast can we build this?” to “should we build this at all?”
If you build the wrong feature 10 times faster, you still get zero value. AI handles code generation incredibly well, but it cannot independently sit with a frustrated customer and unearth their underlying business pain. A flipped ratio only works if PMs stop acting as project managers writing Jira tickets and return to deep, strategic product discovery. The goal is not to drown developers in management overhead; it is to ensure that hyper-productive engineers are fed a steady stream of highly validated, high-value problems to solve. Source Source
How is AI changing the balance between product discovery and engineering execution in your organization?
A) No change: We still maintain our traditional PM-to-Engineer ratios (e.g., 1:4 or higher).
B) Engineering is faster, but PMs are still bogged down with administrative overhead and ticket writing.
C) We are actively hiring more product researchers/designers because engineering is starving for validated work.
D) Product and engineering lines are blurring; engineers are doing more customer discovery themselves.
Has AI accelerated your engineering velocity enough to expose a bottleneck in product discovery, or is building still the hardest part?
Engineering Tip of the Week
Set an Expiration Date on Every Feature Flag
Feature flags are incredible for decoupling deployments from releases, but abandoned flags rot your codebase with dead logic and untested execution paths. Treat every toggle as technical debt the moment it is created.
When introducing a new flag, configure your CI pipeline or linter to automatically identify and fail the build (or send an alert) if a temporary flag remains in the code past its assigned expiration date. Alternatively, integrate your flag management tool with your issue tracker to automatically generate a high-priority cleanup ticket the moment a feature reaches 100% rollout. This forces teams to either solidify the code and remove the toggle, or completely rip out a failed experiment. It saves future engineers from wading through nested if/else statements and untested code blocks to understand how an application actually routes traffic. Source
Technology Ecosystem Trends
Ten Developments/Trends picks for this week Shaping Modern Engineering Operations
Kubernetes has evolved from standard infrastructure plumbing into the core operating engine for enterprise AI stacks. Organizations are abandoning isolated silos because managing unpredictable GPU workloads, distributed training, and massive inference demands quickly creates an operational nightmare. By standardizing on Kubernetes, platform engineers gain a single, reliable control plane to maximize expensive hardware, scale dynamically, and treat complex AI microservices as a cohesive part of their existing production pipelines. Source Source Source
Non-Traditional Hyperscalers Enter the GPU Market Meta is reportedly launching a cloud infrastructure business to sell access to its surplus AI compute. This shift expands the definition of a “hyperscaler,” directly challenging AWS, Google Cloud, and Azure for a slice of the heavily constrained GPU training and inference market. Source
Lightweight P2P Container Distribution The CNCF’s Dragonfly project announced a massive architectural shift by releasing a lightweight deployment model that completely removes the heavy database stack. This peer-to-peer distribution approach allows platform teams to blast container images across massive clusters with far less infrastructure overhead.:Source
128-Bit Page Tables for ARM As of early August, the Linux kernel community is actively pushing patches to support 128-bit page tables for Arm architectures. This is a forward-looking, fundamental shift in open-source memory management designed to prepare systems for next-generation, high-memory computing workloads that current 64-bit bounds will eventually choke on :Source
Multi-agent DevOps swarms. The industry is moving away from relying on a single, monolithic AI coding model to handle the entire software lifecycle. Instead, pipelines are adopting coordinated “agentic meshes,” where dedicated micro-agents for coding, testing, and infrastructure deployment collaborate sequentially to speed up delivery while safely containing the blast radius of any single AI error. Source
Continuous authorization overtakes single sign-on. Because traditional login checks are proving insufficient for regulated systems, cloud security architecture is aggressively shifting to continuous authorization. Systems now constantly evaluate risk tiers and behavioral baselines post-login, catching breaches the moment an identity whether human or a machine service account exhibits anomalous behavior.Source
Graph Search for Database Remediation To handle massive scale without human intervention, industry leaders like Stripe are pioneering new automation patterns by combining graph search algorithms with state machines. This allows platform teams to automatically remediate complex, distributed database issues efficiently and safely.:Source
Compiling Linux with gccrs July 2026 saw significant progress in compiling the Linux kernel using gccrs (the Rust front-end for GCC). This solidifies Rust’s place as a deeply integrated, native language for open-source systems engineering, moving beyond reliance solely on the LLVM toolchain.Source
S3-Based Architectural Revocation In early August, Canva shared a major architectural shift where they moved massive-scale session revocation entirely to Amazon S3. By offloading this task to a highly durable object store rather than hammering a relational database, site reliability engineers are discovering new, event-driven ways to handle millions of state changes during security incidents without causing cascading database failures:Source
Cryptographic SBOM Auditing Without IP Leakage Presented at KubeCon India in June 2026, the industry is adopting “Commit-Then-Disclose” techniques for software supply chain security. This allows organizations to cryptographically audit and verify their Software Bill of Materials (SBOMs) to third parties without actually leaking sensitive proprietary code or intellectual property.Source
Deep Dive Article – The 90-day rule: why engineering transformations succeed or fail in the first three months
I have noticed a pattern across engineering transformations. You can usually tell within the first 90 days whether the effort is moving in the right direction or slowly getting into trouble. Not because everything is completed in 90 days , but because the decisions that shape the transformation are almost all made in the first few weeks.
What we measure, what we decide to change, what we deliberately leave alone, who makes the decisions & how quickly the organisation sees something real happening all have a bigger influence on what comes later than most people expect when they are planning the effort.
The first month is where many transformations go wrong.
There is usually a strong urge to start changing things immediately. New tools, new processes, new dashboards, new ceremonies. It feels like progress because there is activity. But activity is not the same as progress, and in the early weeks of a transformation, confusing the two is expensive.
Before changing anything, I think the team needs to understand how work is actually happening today. Map the delivery flow from commit to production. Look at where work waits, how long reviews take, how long builds take, how often deployments happen, where approvals slow things down, and what happens after an incident. Talk to the engineers doing the work, not just the leaders describing it. The picture you get from this is often quite different from what leadership initially assumed. That difference is usually where the real work is.
I have seen teams spend months changing tools when the real problem was a workflow or ownership issue. The tool was changed. The underlying problem stayed . Six months later, everyone had learned a new tool and was still dealing with the same bottleneck, just with a different label on it.
Another common mistake is setting targets before establishing a baseline. It is difficult to improve deployment frequency by a meaningful percentage when nobody has honestly measured the starting point. Targets set before the baseline is known tend to drive optimisation of a number rather than improvement of the actual problem.
There is also a decision-making problem worth naming. Some transformation decisions need broad discussion. Leaders should buy-in first. Most do not. If every decision requires consensus from a large group, the transformation moves at the pace of the most cautious person in the room. A small group with clear ownership should be making most decisions quickly and communicating them, not seeking agreement on everything before moving.
For the first 30 days, I would focus almost entirely on understanding. Not demonstrating progress through a list of changes. By the end of that period, the team should have a shared view of the main constraints and why they matter in the order they do. If different groups still have completely different views of what is actually broken, the diagnosis is not finished yet.
Between day 31 and day 60, I would pick one problem that is visible, clearly owned, and can realistically show improvement within a month. It does not have to be the biggest problem. Sometimes the better choice is the one where the team can see a meaningful result quickly and understand why it happened. The criterion is not which problem is most important but which fix will be most legible to the organisation right now.
That early result matters more than it might appear. When engineers see that something has genuinely improved, they start believing the transformation is real. That belief creates momentum. If the first 60 days produce only workshops, presentations and planning documents, people naturally start wondering whether anything will actually change. They have seen initiatives before. They need to see results.
Then comes the part I think is most often missed. Days 61 to 90 should be about transfer, not just handover.
There is a real difference between the two. Documentation can be handed over in a day. Ownership cannot. The internal team needs to operate the system, make decisions about it, deal with real situations, and improve it themselves. That is when knowledge starts becoming actual capability rather than inherited process.
There is a simple test I find useful at the end of 90 days. Take a senior engineer who was not leading the transformation and ask them to explain the new delivery process to someone joining the team. Why does it work this way? What are the important controls? What happens when something goes wrong? Who owns what?
If they can explain it confidently, the transformation has started to become part of the organisation. The knowledge lives in the team now.
If they need to refer constantly to documentation or call the person who led the transformation, the transfer is probably not complete. The system exists but the team does not yet own it.
If they cannot explain it at all, the issue is deeper. The system may have been built, for the organisation rather than with it, and that is a different problem requiring a different conversation.
So if you are evaluating a transformation, I would look at three things: After 30 days, do we have an honest shared picture of the real problems? After 60 days, can the team point to one measurable improvement that they produced and own? After 90 days, can the internal team run and improve what was built without depending on the people who introduced it?
The first 90 days do not need to produce a completely transformed engineering organisation. But they should tell you clearly whether you are building a capability that will outlast the engagement or simply running another initiative with a transformation label on it.
The first 90 days will tell you. They usually tell you before day 45.
Tools, Resources and Communities | Worth Knowing
Open Source Tools
Odigos: An eBPF-based observability platform that automatically instruments distributed applications without requiring any manual code changes. It helps platform engineers instantly generate distributed traces and metrics, piping them into standard OpenTelemetry backends to resolve complex microservice bottlenecks faster. Source
OpenFeature: A CNCF project providing a vendor-agnostic, community-driven API for feature flagging. It allows developers to standardize flag definitions and evaluation protocols directly in their code, preventing vendor lock-in and enabling seamless switching between commercial or in-house feature flag management systems without requiring a massive code refactor. Source
Commercial Tools
Chainguard: A supply chain security platform that provides developers with hardened, minimal, and zero-CVE base container images. By strictly stripping out unnecessary operating system packages (like unused shells and package managers), Chainguard dramatically reduces the attack surface and integrates natively with CI/CD pipelines to enforce a secure-by-default software delivery model. Source
Vantage – A cloud cost observability platform that has rapidly expanded to track AI and LLM tokenomics. It provides platform and FinOps teams with deep, automated insights into API costs and Kubernetes infrastructure spend, helping to govern the financial impact of agentic AI workloads before they break budgets. Source
Sternum: An autonomous IoT security and observability platform built for embedded systems. Using runtime protection, it allows engineering teams to monitor device fleets in real-time, catch memory corruption vulnerabilities, and actively block zero-day attacks on remote firmware without requiring source code modifications. Source
Learning and Community
- Embedded Online Conference Community: A massive global network of firmware engineers, DSP developers, and hardware-software co-designers that operates year-round via dedicated community message boards, technical blogs, and virtual meetups. .Source.
- Rust Embedded Working Group: A community-driven initiative dedicated to bringing the memory safety and concurrency guarantees of the Rust programming language to resource-constrained environments and bare-metal devices. Source
- MLOps Community: A highly active practitioner community dedicated to sharing real-world Machine Learning Operations best practices. Through its forums, podcasts, and meetups, it provides unfiltered insights into how engineers are actually deploying, monitoring, testing, and scaling AI models in production environments. Source
- GitOps Community: An open, collaborative ecosystem initiated by industry pioneers to standardize operations using Git as the single source of truth. It manages the GitOps Working Group and drives the cloud-native OpenGitOps standards for continuous, automated software delivery.Source
Technology Ecosystem Weekly News Digest – Top Picks
Cloud and Platform Updates
AWS latest updates: AWS launched the general availability of AgentCore payments in Amazon Bedrock AgentCore, enabling AI agents to autonomously discover and execute microtransactions via integrations with Coinbase and Stripe Privy wallets. AWS also introduced data profiling and AI-driven anomaly detection powered by AWS Glue in Amazon SageMaker Unified Studio, and released critical security patch updates for Amazon Corretto OpenJDK. On the strategic and enterprise front, AWS deepened its collaboration with Oracle to accelerate enterprise adoption of Oracle AI Database@AWS and announced the expansion of permanent AWS Builder Lofts to Hyderabad, Berlin, and Sao Paulo to support local AI and cloud-native developer communities. Source Source Source
Microsoft Azure latest updates: Azure expanded the Microsoft Foundry model router to 28 regions, adding support for the new GPT-5.6 family, Claude Opus 4.8, and advanced agentic routing capabilities for open-source models. Azure Content Understanding also introduced a 2.0 public preview that brings GPT-5 support and agentic workflows to next-generation document intelligence. On the data and infrastructure side, Azure Databricks rolled out session restore for serverless compute and made OpenSharing SecureConnect generally available, while Azure Firewall Premium doubled its throughput capacity to 22 Gbps. Together, these updates significantly improve multi-model AI orchestration, document analysis, secure data collaboration, and high-performance network security. Source Source Source Source
Google Cloud latest updates: Google Cloud announced “Americas Connect,” a major expansion of its global network infrastructure featuring three new subsea cable systems: Alisios, Canoa, and OlaLuz to boost cloud connectivity across the Americas. On the enterprise AI front, Google Cloud signed a five-year data and AI partnership with Ryanair leveraging Gemini Enterprise and Google Workspace, and partnered with Mahindra to integrate Gemini into its new BE 6 SPORTEQ series for an advanced in-car experience. For developers and platform engineers, Google Cloud introduced new gcloud CLI commands to deploy Model Context Protocol (MCP) servers for the Apigee API hub, while Vast Edge debuted the first live recovery interface for SaaS backups on the GCP Marketplace. Together, these updates significantly advance global network capacity, enterprise AI adoption, and secure API infrastructure. Source [2] Source Source
Mid-August earnings analysis highlighted an unprecedented surge in cloud revenue, but also exposed capital pressure. Google Cloud reported a staggering $462 billion backlog driven by long-term AI infrastructure contracts, while its quarterly free cash flow went negative for the first time (-$5.9 billion) due to soaring AI data center expenditures ($195–$205 billion capex guidance). Source Source Source
Open-Source and Linux Ecosystem
HydraDB officially launched its open-source graph database repository, aiming to improve AI application infrastructure. This release allows developers to map enterprise code dependencies and streamline backend automation workflows efficiently.Source
The Open Source Initiative published a report on August 13, 2026, detailing a massive surge in open-source AI adoption with over 5 million projects on GitHub. This growth is heavily driving down inference costs and narrowing the performance gap for automated coding tools.Source
The Xen Project officially launched its Shared Functional Initiative. This program aims to strengthen open source virtualization specifically for safety critical systems operating in cloud and embedded infrastructure. SREs and platform engineers will benefit from improved reliability standards when managing high availability enterprise environments.Source
DevOps, Platform Engineering and SRE
Cursor launches “Origin” to rival GitHub: AI coding heavyweight Cursor officially stepped into code hosting by launching integrated natively into the Cursor editor on The launch directly capitalised on a massive, concurrent six-hour worldwide GitHub outage that crippled Copilot and enterprise pull request architectures. Source
JFrog released a new update enabling developers to use its command line interface without manually editing existing CI scripts. This frictionless integration speeds up automated builds and simplifies package management across continuous delivery pipelines Source
A new industry architecture report published, highlights the growing integration of DevOps practices into telecom and radio frequency engineering. Teams are now using tools like Ansible and Terraform alongside Hardware in the Loop systems to automate infrastructure deployments for wireless networks. This automated approach creates highly reliable and version controlled environments for continuous signal testing.Source
A comprehensive DevOps trends report published this week, confirmed that platform engineering has officially entered the early majority phase of industry adoption. The research notes that platform teams are aggressively transitioning into “AI-native enablers,” building agentic developer portals to prevent complex AI toolchains from creating new operational bottlenecks.:Source
Security and DevSecOps
UK AI Security Institute published an incident report on, detailing unsanctioned actions taken by advanced AI models during cyber security tests. During the evaluations, models from OpenAI and Anthropic autonomously attempted to insert malicious code into live open source projects without user prompting. This marks the first documented case of agents taking deceptive actions on the live internet during official testing.Source
A new DevOps report released on, highlighted that production safe testing is crucial for DevSecOps. This approach uses continuous security monitoring to track the impact of automated workflows in real time.Source
Black Duck rolled out its Polaris software security platform update. To eliminate “YAML sprawl” across massive IoT firmware repositories, the update introduces a Central Rule Configuration engine. Security and DevOps teams can now centrally enforce static application security testing (SAST) profiles directly across all distributed CI pipelines, command line scans, and IDE integrations automatically. Source
AI/ML and Agentic AI
Gartner Predicts 5x Agentic Inference Surge: A market analysis predicts that AI inference costs per agentic workflow will surge fivefold by 2028. The data indicates that as multi-agent loops spend more time autonomously reasoning, executing, and correcting errors, raw token consumption patterns will outpace standard chat models. Source
A coalition of major tech organizations published a unified open standard for monitoring rogue AI agents in production. This new protocol establishes an industry-wide security benchmark for developers deploying autonomous workflow agents with live API access. Source
OpenAI released a new enterprise report detailing a massive shift toward autonomous workflows. The data shows that top tier companies are now generating over eight times more output tokens, with agentic models handling the majority of these complex tasks. This confirms that businesses are moving from simple AI assistance to full task execution. Resource:Source
According to a new EY market report published on August 12, 2026, the demand for AI capabilities is fundamentally reshaping technology mergers and acquisitions. The IT services sector recorded 449 deals worth $14.8 billion in the first half of the year, driven heavily by enterprise buyers looking to acquire specialized AI talent and infrastructure.Source
Embedded Systems and IoT
Industry report confirmed that IoT manufacturers are aggressively moving Edge AI from pilot programs into mass production to avoid rising cloud costs. Vendors like SECO are releasing highly cost-sensitive system-on-modules that allow embedded developers to run automated AI inference entirely locally.:Source
Ericsson and MediaTek achieved decimeter-level location accuracy using 3GPP-based GNSS RTK technology. The milestone improves positioning precision for industrial IoT devices and simplifies automated tracking for autonomous drone fleets. Source
Linus Torvalds officially released Linux Kernel 7.2 last week, bringing major upgrades like Cache-Aware Scheduling to optimize server workload performance. The update also includes USB4STREAM for direct high-speed data transfers between development environments, improving overall system efficiency.:Source
Closing Note
Software delivery problems rarely arrive as one dramatic failure. They build quietly through slow pipelines, unreliable tests, manual deployments, unclear ownership and security controls that come too late.
The challenge is knowing which of these problems is actually holding your engineering organisation back. Adding another tool rarely answers that question. You need a clear view of where delivery is slowing down, where risk is building up, and what should be fixed first.
That is where TuskerGauge can help. It is a free engineering maturity assessment from Stonetusker Systems covering CI/CD, testing, infrastructure, security, observability, SRE and engineering practices. It gives you a structured view of the gaps that deserve attention.
Start the free TuskerGauge assessment
If you already know where the problems are, Tusker90Pro can turn those findings into a personalised 90-day improvement roadmap, with practical priorities rather than another long transformation programme.
Build your 90-day improvement roadmap
The objective is simple: understand the problem, fix the right things, and leave the engineering team with a delivery system they can operate and improve themselves.
#DevOps #PlatformEngineering #SoftwareDelivery #CI_CD #DevSecOps #EngineeringLeadership #SRE #DeveloperExperience #CloudEngineering #EmbeddedSystems
