The Software Efficiency Report · From the Founder's Desk

The Software Efficiency Report | 2026 Week 41

DevSecOps for AI Model Deployment in Healthcare: Building Pipelines That Survive an Audit

Software delivery continues to change quickly, and engineering teams are having to adapt across cloud platforms, Kubernetes, security, developer tooling, and embedded systems. This week's developments show how quickly these areas are moving, with more automation entering cloud operations, stronger software supply-chain controls, and new approaches to cost, reliability, and security.

The security landscape is changing just as quickly. New vulnerabilities, attacks on exposed infrastructure, AI-assisted development, and the growing use of autonomous agents are creating new operational risks that engineering teams need to account for in their delivery systems.

This week's deep dive looks at DevSecOps for AI model deployment in healthcare, where the pipeline has to protect more than source code and containers. Data lineage, model artifacts, validation, clinical approval, security controls, and audit evidence all need to be part of the delivery process before a model reaches production.

Metric of the week
Customer-Reported Defect Ratio (CRDR): Target <10%
Deep dive
DevSecOps for AI Model Deployment in Healthcare: Building Pipelines That Survive an Audit

Software Efficiency Metric of the Week

Customer-Reported Defect Ratio (CRDR): Target <10%

This measures the percentage of confirmed production and field defects first discovered through customer reports or support tickets, rather than through engineering telemetry, automated monitoring, diagnostics, or internal investigation.

What it is: A detection-source KPI. While Defect Escape Rate (DER) measures how many bugs slip past pre-release controls into production, CRDR answers a different operational question: of the failures that reached production, how often did your customers find them first?

Why it matters: When customers serve as your earliest alarm system, external failure costs escalate quickly. In cloud services, customer-reported bugs drive emergency hotfixes, high support ticket volume, and churn. In connected hardware and embedded devices, they lead to warranty claims, truck rolls, and product returns. ASQ classifies these as external failure costs, including warranty claims, repairs, complaints, returns, and field service.

The Fix:

  • Automate post-deployment health checks: Run continuous or scheduled synthetic user journeys on critical cloud applications, with frequency calibrated to your SLO and detection requirements, and build automated self-test diagnostics on connected devices to detect failures before users do.
  • Stream automated crash telemetry: Pipe application error logs, traces, and crash reports, along with embedded crash dumps, into a central observability platform so on-call engineers get alerted immediately.
  • Tag the detection source on every ticket: Add a mandatory "Found By" field (APM, Synthetics, Fleet Telemetry, Internal Dogfooding, Customer Report) to incident records to track your detection performance over time.
  • Run detection retrospectives: For any defect reported by a customer, require a postmortem action item to build the specific monitor, alert, or hardware-in-the-loop test that will catch it automatically next time.

Formula: CRDR (%) = (Customer and Support Discovered Defects / Total Confirmed Production and Field Defects) * 100

Quick Example:

If 100 confirmed production and field defects occur in a quarter:

  • 65 caught by automated monitoring and telemetry
  • 20 caught by internal engineering inspection
  • 10 caught by automated field diagnostics
  • 5 reported by customers or support tickets

CRDR = 5 / 100 = 5%

Further Reading & References: Source Source Source Source

Reader Poll

Do you know how many of your paid AI seats actually get used each week, or are you paying for expensive digital shelfware?

My take: Most companies rolled out AI by purchasing blanket enterprise seats for every single developer on the roster. Fast forward a few months: a quarter of the team uses the tool daily, half the team opens it once a week, and the rest quietly reverted to their old workflows. Meanwhile, your power users are expensing personal frontier model API keys on corporate cards. If you aren't auditing seat activity every two weeks, you're just paying a recurring tax to model vendors. Source Source Source Source

How does your org manage AI coding licenses and API spend?

  • A) Active pruning: We audit seat usage every 14 days and reclaim inactive licenses automatically.
  • B) Blanket enterprise buy: Everyone gets a seat whether they open the tool or not.
  • C) Expense report sprawl: Devs expense personal ChatGPT or Claude subscriptions with zero central tracking.
  • D) The lockdown: Corporate security blocked external AI tools, so devs use nothing official.

Are AI tools driving real leverage on your team, or just sitting as unmanaged subscriptions on the company card?

Engineering Tip of the Week

Use git worktree instead of stashing when an urgent interruption strikes

When an urgent hotfix or code review interrupts your active work, stashing changes and checking out another branch forces your IDE to re-index files and invalidates local build caches. Run git worktree add ../hotfix main instead. This creates a separate working directory linked to your existing local repository, letting you test, patch, and push the fix in a completely isolated folder without touching your dirty working tree or rebuilding dependencies.

Top Developments/Trends picks for this week Shaping Modern Engineering Efficiency and Operations

CNCF is using AI agents to speed up its own governance: executive director Jonathan Bryce says agents now assist with technical due diligence at each stage of project review, pushing projects through graduation faster than at any point in the last 18 months. Source

Cloud CLIs are being rebuilt for agents, not humans: Cloudflare's new "cf" tool covers all 3,000-plus Cloudflare API operations after agent traffic hit 48% of CLI use, and within days Google Cloud shipped a remote MCP server for gcloud and AWS launched its own Well-Architected Agent preview. Source Source Source

Rust keeps eating core Linux tooling, one component at a time: the Mold linker shipped version 3.0 as a full rewrite from C++ to Rust, and Canonical reaffirmed Ubuntu 27.04 will default to the Rust-based ntpd-rs for time sync, extending the memory-safety migration Google ran on giflib last week. Source Source

Bug bounty programs are drowning in AI-written submissions: Google paused its Open Source Software Vulnerability Rewards Program after "a significant rise in automated submissions, the vast majority of which are not valid," the same complaint that made curl kill its own bounty program in January. Source Source

Gartner tells CISOs to treat AI agents as insider risks, not tools: its top five actions for the rest of 2026 include governing agents by what they're allowed to do rather than trusting them, replacing face and voice checks with layered defenses against deepfakes, and starting post-quantum migration now, echoing Cloudflare's quantum-safe certificate push from last week. Source

Shift-Left FinOps and Cost-as-Code in CI/CD: Engineering teams are integrating cloud cost estimation tools directly into pull request checks so developers immediately see the financial footprint of their architectural changes, preventing surprise month-end cloud bills and establishing true developer unit-economics ownership. Source

Predictive Mainline Queues for Legacy Monorepos: Teams managing high-volume monorepos are adopting machine-learning-driven speculative trees to pre-calculate merge conflicts and predict build outcomes, cutting CI compute consumption by over 50% while preventing broken PRs from blocking team releases. Source

Zero Standing Privileges (ZSP) via CI/CD Identity Federation: Security leaders are eliminating static service-account keys and hardcoded cloud secrets across build runners by enforcing short-lived OpenID Connect (OIDC) tokens and dynamic role assumption across all deployment workflows. Source

Database DevOps & Schema-as-Code Eradicate Release Blockers: Engineering orgs are treating database changes like first-class application code by embedding schema linters, backward-compatibility validation, and GitOps-driven DDL rollouts directly into CI/CD pipelines to eliminate the traditional "last-mile" migration bottleneck. Source

Outcome-Driven DevEx Replaces Vanity DORA Metrics: Technology executives are moving beyond raw deployment velocity metrics in favor of composite Developer Experience (DevEx) indicators—measuring CI wait times, review latency, and cognitive load—to stop teams from measuring productivity solely by raw, AI-generated code volume. Source

Tools, Resources and Communities | Worth Knowing

Open Source Tools

  • probe-rs: A modern, fast toolkit for flashing, debugging, and inspecting ARM Cortex-M and RISC-V microcontrollers. Written in Rust, it communicates directly with debug adapters like CMSIS-DAP, J-Link, and ST-Link, allowing developers to flash targets and stream real-time RTT logs from the command line or headless CI runners without clunky vendor IDEs. Source
  • Renovate: An automated multi-language dependency update tool. It continuously scans package manifests, groups updates logically, displays release notes inside pull requests, and automates minor patch merges so engineering teams avoid the maintenance burden of stale libraries. Source

Commercial Platforms

  • Golioth: A specialized cloud platform built for custom connected hardware fleets. It provides pre-built, secure OTA firmware deployment pipelines, device configuration state management, and field telemetry routing so embedded teams can manage IoT devices without building custom cloud infrastructure from scratch. Source
  • Balena: A fleet management and deployment platform for connected embedded Linux devices like Raspberry Pi, BeagleBone, and Nvidia Jetson boards. It packages application runtimes into containerized workloads, enabling safe over-the-air updates, remote container rollbacks, and host OS diagnostics across thousands of remote edge units. Source
  • Swarmia: An engineering productivity platform that connects Git repositories, issue trackers, and Slack. It surfaces pull request review delays, work-in-progress (WIP) limits, and engineering investment allocation without resorting to invasive activity tracking or vanity metrics. Source

Learning Resources – Interesting articles

  • Bug-Killing Coding Standard Rules for Embedded C: A practical set of coding rules designed specifically to prevent catastrophic embedded firmware bugs. Barr details defensive bracket hygiene, pointer typing, fixed-width integer types (stdint.h), and proper use of the volatile keyword around hardware registers. Source
  • Don't Overflow the Stack: A crisp analysis from safety-critical systems expert Phil Koopman on the danger of stack overflow in embedded microcontrollers. He outlines stack growth mechanics, memory corruption risks, stack sentinels, and why static analysis alone cannot guarantee safety. Source
  • How to Do a Code Review: Google's internal playbook for reviewing code changes. It emphasizes maintaining review speed, balancing technical perfection against forward progress, and structuring feedback so pull requests improve the codebase without stalling developer momentum. Source

Deep Dive Article : DevSecOps for AI Model Deployment in Healthcare: Building Pipelines That Survive an Audit

A model is not a normal build artifact and a hospital network is not a normal production environment. Here's how to build the pipeline for both.

Picture a sepsis-risk model (an AI early-warning system that flags patients at risk of a life-threatening infection) that starts behaving differently after a re-train. The code didn't change. The container digest looks the same. A clinical safety reviewer asks what happened, and "the code didn't change" is not an answer anyone in that room will accept.

This article covers what changes when you deploy AI models in healthcare and the pipeline design that holds up when compliance, security & clinical teams all start asking questions.

The short version

  • DevSecOps for AI model deployment in healthcare means building security controls, clinical validation, and audit evidence into the pipeline that trains, packages, and releases models, so every production version can be traced, verified, and rolled back.
  • A model's behavior depends on its training data, so the pipeline has to control data lineage and PHI exposure, not only source code and containers.
  • The serialized model file is an attack surface. Formats like Python pickle can run arbitrary code when loaded.
  • Promotion to production should depend on recorded validation evidence and a named approver, enforced by policy as code.
  • Auditors want proof, not intent. The best pipeline produces its own evidence as a side effect of running.

Why a model breaks the standard DevSecOps playbook

Most DevSecOps programs assume a simple flow: code goes in, a binary comes out, and you scan, sign, and deploy the binary. A model is the product of code, data, hyperparameters, and a training run that may not be perfectly repeatable.

Three things are genuinely new.

The training data is part of the build input. It carries PHI, it has provenance, and it can quietly change between runs.

The artifact itself can be hostile. Model files pulled from public hubs or passed between teams have been used to carry code that executes on load. Safer formats like safetensors exist for exactly this reason, yet plenty of internal pipelines still unpickle whatever lands in the bucket.

The failure mode is statistical. A model can pass every unit test and still be wrong for a patient population it was never validated on. No scanner catches that. You need a validation gate that a clinician or data scientist actually signs.

Test it: Pick a model running in production today. Can your team name the exact dataset snapshot, code commit, and container digest that produced it, without asking the person who trained it?

The pipeline, stage by stage

Each stage needs a control, and each control needs to leave evidence behind. If a stage produces no evidence, an auditor will treat the control as if it doesn't exist.

Data ingest. De-identification checks, role-scoped access, dataset versioning. Evidence: dataset hash, access log, de-identification report.

Training. Isolated compute, pinned dependencies, no outbound internet by default. Evidence: training config, environment lockfile, run ID.

Packaging. Safe serialization format, SBOM generation, image signing. Evidence: signed model, signed image, SBOM.

Scanning. Dependency and container scans, plus a scan of the model file for unsafe deserialization. Evidence: reports tied to the artifact digest.

Validation. Performance on held-out clinical data, subgroup checks, bias review. Evidence: a model card and a signed validation report.

Promotion. Policy as code, a named approver, a change record. Evidence: approval record and promotion log.

Runtime. Signature verification at admission, network policy, drift monitoring. Evidence: admission logs, drift dashboards, alert history.

Notice that the signature gets verified at runtime, not just created at build time. A signature nobody checks is decoration. The Sigstore ecosystem and the SLSA framework give you a recognized vocabulary for this, and both are worth reading before you design your own.

Test it: If someone swapped the model file in your registry tonight, would your cluster refuse to start it?

PHI boundaries and data lineage

This is where healthcare pipelines diverge hardest from everyone else's. A typical SaaS team can snapshot production data into a training bucket without much ceremony. You can't.

Decide early where PHI is allowed to exist in the pipeline, and put that boundary on a diagram that engineers and your privacy officer can both read. Some teams train in a controlled enclave on de-identified or limited datasets. Others train on identified data inside the covered entity's boundary and promote only the resulting artifact outward, which makes the question of what the model might memorize very real.

Treat the model artifact as potentially sensitive until a privacy review says otherwise. Large models trained on small clinical datasets are the risky combination.

Lineage has to be queryable. Every model version links to a dataset version, and every dataset version links to its source systems and the de-identification method applied. Keep that as structured metadata in your model registry, not in a spreadsheet someone updates when they remember. HIPAA's technical safeguards include audit controls, so access to training data and to the registry should be logged somewhere tamper-resistant.

Test it: A privacy officer asks which models were trained on data from one specific clinic. How long does it take you to answer, and how sure are you?

An eight step deployment playbook

This is the sequence we use at Stonetusker when helping a team stand up a regulated model pipeline from scratch. You can adopt it in stages, but the order matters because later steps depend on evidence from earlier ones.

  1. Classify the model and its intended use. Decide whether the software could be a regulated medical device, what clinical decisions it informs, and who owns the risk. Get the answer in writing. Tools like Greenlight Guru or Vanta help to log regulatory classifications and track risk registers in one place.
  2. Fix the data lineage and PHI boundary. Version datasets, record where identified data may live, and log all access before training begins. Tools like DVC help to snapshot exact dataset versions, while tools like Microsoft Presidio detect and redact PHI.
  3. Make training reproducible and isolated. Pin dependencies, use controlled compute, and block outbound network access unless approved. Tools like Poetry or Docker lock your execution environment, while tools like MLflow record seeds and hyperparameters for every run.
  4. Sign everything and generate an SBOM. Sign the model file, container image, and training configuration so vulnerability disclosures map directly to live code. Tools like Cosign help to cryptographically sign artifacts, while tools like Syft generate an automated Software Bill of Materials (SBOM).
  5. Scan the code, the dependencies, and the model file. Catch vulnerabilities across source code, third-party libraries, and serialized binaries. Tools like SonarQube or Snyk scan source code and dependencies, while tools like ModelScan detect unsafe code hidden inside model files.
  6. Gate promotion on validation evidence. Require a validation report, subgroup performance checks, and a named approver before promoting a model. Tools like Great Expectations run automated quality checks, while model registries like AWS SageMaker Registry store approval signatures.
  7. Deploy with policy as code and staged rollout. Verify signatures at admission, restrict network paths, and start in shadow mode. Tools like Kyverno or Open Policy Agent (OPA) block unsigned artifacts at admission, while tools like Argo Rollouts manage canary deployments.
  8. Monitor drift and rehearse rollback. Track live performance against validated baselines and practice emergency rollbacks. Tools like Arize AI or Evidently AI monitor input data drift, while tools like Prometheus and Grafana track real-time system performance.

Anti-patterns we keep finding

None of these are exotic. Most exist because a team under deadline took the fast path once and never went back.

  • Notebook to production by hand. No reproducibility, no approval trail. Promote only from the registry, through the pipeline.
  • Loading pickled models from shared buckets. Arbitrary code execution on load. Use safer formats, scanning, and signed artifacts.
  • Validation done once, at launch. Retrains inherit approval they never earned. Re-run and re-sign validation for every version.
  • One shared service account for the registry. The audit log can't say who did what. Use individual identities and short-lived credentials.
  • Auto-retrain straight to production. Skips clinical review entirely.
  • Evidence collected the week before the audit. Gaps, guesses, stress.

The auto-retrain one deserves a second look. Continuous learning sounds great on a slide. In a regulated setting, a model that changes itself without a recorded review is a compliance problem that happens to run on a schedule. Auto-retrain to staging if you like, but keep a human gate on promotion.

Test it: Who can push a new model version to production right now, and where's the record of why they were allowed to?

What happens when compliance breaks down? : A failed audit is an immediate operational crisis. Unapproved model updates risk FDA recalls, while data leaks trigger heavy penalties under HIPAA, GDPR, or local privacy laws. The same risks apply globally under frameworks like the EU AI Act. If you cannot prove data lineage, hospital risk teams will shut down your production model on the spot and block enterprise sales.

The compliance angle: HIPAA, FDA, and NIST

Three frameworks come up in almost every engagement. None hands you a pipeline design, but each shapes what your pipeline has to prove.

The HIPAA Security Rule governs how electronic PHI is protected, which covers training data, intermediate datasets, and potentially the model itself. Access control, audit controls, and integrity protections map directly onto registry permissions, logging, and signing.

For software that qualifies as a medical device, the FDA's work on AI-enabled device software includes the idea of a predetermined change control plan, which lets a manufacturer pre-specify certain model modifications. Your validation and promotion gates are what make such a plan executable. Whether your product is a device at all depends on intended use, so involve regulatory counsel rather than guessing.

On the engineering side, NIST SP 800-218 (the Secure Software Development Framework) and the NIST AI Risk Management Framework give you shared language for secure build practices and AI risk. Auditors tend to respond well when your controls map to named, public frameworks.

Quick answers to questions I get asked

Does HIPAA apply to a model trained on patient data? It applies to the PHI used to train it and to the systems handling that data, and a model can memorize or leak training records in some cases. Treat the artifact as potentially sensitive until your privacy review says otherwise, and talk to compliance counsel about your specific case.

Do we need FDA clearance to deploy a model? Only if the software meets the definition of a medical device, which depends on intended use. A scheduling or billing model usually doesn't. One that informs diagnosis or treatment often does.

What should we sign and scan? Sign the model file, the container image, and the training configuration, and attach an SBOM. Scan dependencies and container layers as usual, and scan the serialized model too.

How should drift be handled? Compare live input and output distributions to the validated baseline, with thresholds that trigger review. A drift alert should open a documented investigation and, if needed, a controlled rollback, not an automatic retrain that skips validation.

Can a small team do this without a platform group? Yes. Start with the controls that produce the most audit evidence: signed artifacts, a model registry with approvals, and logged promotions. Add the rest in stages.

Where to start

Pick one production model and try to answer the three "Test it" questions above. Wherever you can't answer quickly, that's your first gap, and it's usually a cheaper fix than people expect.

References for Further Reading

Technology Ecosystem Weekly News Digest – Top Picks

Cloud and Platform News

AWS this week.

  • AWS Continuum for Penetration Testing now runs continuous, CI/CD-integrated penetration tests instead of the periodic manual engagements most teams schedule once or twice a year. Source
  • Security Hub picked up guided remediation plans that prioritize fixes across findings, and GuardDuty Runtime Monitoring is now bundled into the existing Security Hub Threat Analytics plan at no extra line item. Source Source
  • AWS Config now tracks 77 more resource types, widening drift and compliance detection for anyone running infrastructure as code at scale. Source
  • Amazon ECS added native blue/green, linear, and canary deployments through VPC Lattice, so teams get progressive delivery without standing up a separate service mesh. Source
  • Amazon EKS and EKS Distro both now support Kubernetes 1.37. Source

Azure this week.

  • SQL Server on Azure Virtual Machines reached general availability inside Azure Bleu, the sovereign cloud Microsoft runs with Orange and Capgemini in France, giving regulated workloads a path to in-country data residency. Source
  • AKS on bare metal now runs on Ubuntu in public preview, widening node OS choice for container platform teams. Source
  • Azure HorizonDB, Microsoft's newer distributed Postgres-compatible database, expanded its preview to additional regions. Source

Google Cloud this week.

  • Google Cloud Modernize bundles Migration Center, VMware Engine, Mainframe Modernization, and a new preview tool for moving Kubernetes workloads off AWS EKS onto GKE under one "Modernization Hub," with an AI estimator that turns a VMware inventory export into a cost projection. Google cites NetEase Games cutting a scaling task from hours to five minutes using the tooling. Source Source
  • Security Command Center's Artifact Guard, the CI/CD-integrated artifact scanner, is being deprecated and shuts down October 30, so teams relying on it for pipeline scanning need a replacement lined up. Source
  • Cloud SDK 588.0.0 shipped a breaking change that removes the managed-flink-client component from the gcloud CLI, worth checking against any pipeline that scripts gcloud directly. Source
  • Google Cloud VMware Engine's bring-your-own-license management reached general availability, letting customers register and manage their own VCF license keys across projects as Broadcom's VMware licensing changes keep pushing people to reconsider the economics. Source

Open-Source and Linux Ecosystem News

The Linux Foundation's fintech arm launched OSERA with Deutsche Bank, Goldman Sachs, Morgan Stanley, NatWest, and RBC as founding backers, and in under a month the group has already published a patching and attestation standard and shipped compliant patches for more than 50 Spring and Java projects. The goal is to kill the "fork tax" where every bank quietly patches the same vulnerability on its own, and the roadmap calls for 80-plus patches a month once the full platform ships in November. Source

Red Hat and IBM's AI-assisted vulnerability hunt moves from research to a product. Their Lightwell initiative, whose AI agents found more than 400 previously unknown Java vulnerabilities (reported here last week), just took its "Clearinghouse" to general availability, so any enterprise can now submit its own dependency list for prioritized, AI-assisted remediation delivered through its existing scanners and pipelines rather than a tool swap. Source Source

Google is handing off the documentation theme that half of cloud native runs on. Docsy, the Hugo-based docs framework used by roughly 2,200 projects including Kubernetes, OpenTelemetry, and gRPC, is moving from Google's stewardship to the Linux Foundation. The project is also adding features aimed at AI coding agents rather than just human readers, including auto-generated per-page Markdown and a planned score for how agent-friendly a given doc set actually is. Source

CNCF and Kubernetes this week.

  • Node swap support in Kubernetes reached general availability in v1.34, and new project benchmarks show up to 3x higher pod density when swap is backed by fast NVMe storage, tested on bursty, idle-heavy workloads like AI agent sandboxes and CI jobs. Source
  • Kubernetes flagged cgroup v1 management as being in maintenance mode since v1.31, with cgroup v2 as the stable path going forward, worth checking before your next node image refresh. Source
  • Istio is moving its container images and Helm charts off Google Cloud and onto AWS and Docker Hub, citing changes to its funding model, and will run an "outage test" on October 13 that briefly disables the old endpoints before they retire for good in December. Source

DevOps, Platform Engineering and SRE News

GitHub this week.

  • Copilot CLI and the Copilot desktop app picked up a public preview of "computer use," letting the agent operate desktop applications directly rather than staying confined to a terminal or editor. Source
  • GitHub deprecated four Copilot models across every surface, Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code, and Claude Opus 4.7, pushing everyone onto their newer replacements, and code review can now be triggered through the REST and GraphQL APIs. Source

Dynatrace completed its acquisition of Arize, folding the AI-evaluation platform's trace analysis and auto-remediation tools into Dynatrace so teams can diagnose both "is the infrastructure healthy" and "is the agent reasoning correctly" from one place. Arize's existing "Signal" feature already reviews failure traces and opens pull requests to fix recurring problems, with roughly a 65 to 70 percent acceptance rate by its own count. Source Source

AWS open-sourced a sandbox built to stop autonomous agents from doing real damage. Strands Box adds a deterministic policy engine on top of normal OS-level isolation, so a team can cap how many times an agent is allowed to post to Slack or push to a git repo inside a given time window rather than trusting it to self-regulate. It targets teams running agents in so-called YOLO mode with actions auto-approved, and AWS plans to extend deployment to its own AgentCore service as well as ECS and Kubernetes. Source

Security and DevSecOps News

A stolen company login exposed 80% of Denmark's population register: attackers used a private firm's legitimate CPR access to run automated lookups for about ten days in September, exposing names, addresses and ID numbers for roughly 8.8 million people, a reminder that third-party access is now the weak link in national data systems. Source

Attackers hijacked three country-code domain registries just to get valid Google certificates. By compromising the registries behind Ghana, Sierra Leone, and American Samoa's top-level domains, attackers altered DNS records enough to pass routine domain-validation checks and obtain legitimate HTTPS certificates for google.com.gh, google.sl, and google.as, a dozen certificates across seven domains found so far. Google's own systems were never touched, and it says other organizations were likely hit the same way, though it hasn't named them. The episode is a sharp reminder that the trust chain behind TLS certificates is only as strong as the weakest registry it depends on. Source Source

Exposed AI servers are now a cryptomining target, not just a theoretical risk. A campaign researchers call PoeLLM has already infected more than 3,400 exposed AI and LLM servers, using them both to mine cryptocurrency and as scanning platforms to find the next vulnerable target. Anyone standing up self-hosted inference infrastructure without locking down network exposure is a direct match for what this campaign is scanning for. Source Source

A critical flaw in a popular LLM-serving tool has no patch yet. JFrog found that LMCache, the open-source caching layer that speeds up vLLM and similar serving stacks, runs an unauthenticated ZeroMQ socket that unpacks messages with Python's pickle, so one crafted message can execute commands as the process user, root on the official container images. The risk only shows up if an operator binds the socket to a routable address instead of localhost, but with no fix available yet, that configuration check is the only real mitigation right now. Source

AI/ML and Agentic AI News

An open-source agent framework just became a $1.5 billion company. Nous Research closed a $90 million Series B led by Robot Ventures, with Nvidia, Union Square Ventures, and Samsung among the backers, putting its Hermes agent framework at a $1.5 billion valuation. The company says Hermes has been cloned more than 24 million times and now drives roughly 2.5 percent of global AI token usage, and it's launching a paid enterprise version that lets companies deploy customized agents while keeping their data private, a serious signal that open-weight agent tooling can scale into real enterprise revenue. Source

Mistral's trillion-parameter model finally went public, access-gated. Large 4, nicknamed Le Chonk internally, launched behind a guardrail-gated endpoint while Mistral finishes safety testing, with full open-weight release planned for October 27. Mistral says it trained the model on only around 4,000 Nvidia GPUs, a fraction of what rivals reportedly use at this scale, and is explicitly targeting cybersecurity, finance, and chip-design customers. This is the same model that reportedly tried to break out of its sandbox during an earlier cybersecurity evaluation, a detail that didn't come up in either launch announcement. Source Source

Anthropic cut Haiku prices by as much as 90 percent. For requests under 100,000 tokens, which Anthropic says covers about 90 percent of Haiku traffic, input pricing dropped from $1.00 to $0.10 per million tokens and output from $5.00 to $0.50, landing exactly on OpenAI's GPT-6 Luna pricing up to Luna's token threshold. Anthropic says real-world workload costs fall by roughly 75 percent net of a slightly less efficient new tokenizer, a direct number to plug into the budget for any team running high-volume classification or routing on Haiku. Source

Cohere built an agent platform around the two things CIOs keep complaining about. North 2 adds per-user and per-department token and request budget caps with alerts, persistent memory across sessions, human-approval checkpoints in the orchestration flow, and prompt-injection guardrails, deployable in the cloud, on-prem, or fully air-gapped. It is aimed squarely at the two problems that keep coming up in agentic rollouts, runaway spend and no real audit trail, though Cohere hasn't given a general-availability date or public pricing yet. Source

AI tools are showing up cheap on underground marketplaces, short of full automation. A Halcyon Ransomware Research Center analysis of roughly 4,000 posts across Telegram channels and dark web forums found ads for AI criminal tooling jumping from under 50 a month in late 2025 to more than 1,400 a month by February, including stolen ChatGPT Plus accounts reselling for as little as 10 cents and jailbreak prompts going for under $10. Halcyon's research lead told Axios the one reassuring finding is that no fully autonomous end-to-end attack agent turned up for sale, AI is helping with discrete tasks like phishing kits and malware infrastructure rather than running an entire attack on its own, at least for now. Source

Embedded Systems and IoT News

Medical devices are years behind on crypto that doesn't yet exist in attacks. Forescout analyzed roughly 2.5 million devices across more than 50 healthcare organizations and found about half of general IT devices can support post-quantum cryptography, but that drops to 16 percent for operational technology and just 6 percent for infusion pumps, imaging systems, and other connected medical equipment. Only 31 percent of internet-reachable systems holding medical records even run TLS 1.3, the protocol current post-quantum standards build on, which matters because health data stolen today can still be decrypted once quantum computing catches up. Source Source

Silicon Labs is wiring AI coding assistants directly into chip-specific tooling. At its Works With summit, the company opened a beta SDK that gives tools like GitHub Copilot and OpenAI Codex structured access to its BLE documentation and build tools, open-sourced its Bluetooth Low Energy software stack for community contributions, and partnered with Databricks to move edge-device telemetry into cloud-based model training. It's a genuine step toward treating embedded firmware development like modern software, not just a feature bump. Source

The Operational Technology Cybersecurity Coalition is pushing CISA to issue a binding directive requiring federal civilian agencies to meet baseline OT security controls, asset visibility, segmentation, and verified backups among them, after a GAO report found only 7 of 22 reviewed agencies had actually completed a 2024 deadline to inventory their own networked OT and IoT devices. Source

Closing Note

This week's issue comes back to a simple idea: as engineering systems become more complex, the delivery system behind them needs to keep up.

Cloud cost, security, AI-assisted development, embedded testing, observability and release automation may look like separate concerns. In practice, they all affect the same thing: how reliably your team can move software from development to production.

If you want to see where your engineering delivery system has gaps, start with TuskerGauge. It is a free assessment covering CI/CD, testing, infrastructure, security, observability, SRE, deployment safety and engineering practices.

Find the gaps in your engineering delivery system

If you already know where the bottleneck is, Tusker90Pro turns those findings into a practical 90-day improvement roadmap focused on delivery flow, reliability, release management and engineering operations.

Build your 90-day improvement roadmap

The goal is simple: find the constraint, fix what is slowing the team down, and make software delivery more predictable.

If you want to discuss a specific engineering delivery problem, contact us at contact@stonetusker.com. No pitch deck. Just an engineering discussion around the problem you are dealing with.