Notes from the work, not the hype.
Practical thinking from the engineers and operators replacing manual work with systems that actually ship in production.

The harness we use to break an agent before production does: adversarial inputs, failure injection, tail latency, cost per resolved task, and injection probes.
Read more
An anonymised case study: where an agent's tokens actually went, the changes that cut spend ~60%, and the measurement discipline that proved quality held.
Read more
How to detect AI agent failure in production before users notice — drift detection, online monitoring, confidence-based fallbacks, and the telemetry that makes it work.
Read more
Most enterprise tasks handed to us as 'agent' projects ship better as constrained workflows. Concrete criteria for deciding when to give a model agency — and when to take it away.
Read more
Production agents balloon their tool-call counts with redundant retrievals and second-guessing loops. How to instrument, budget, and detect thrashing.
Read more
LLM self-reported confidence and logprobs are badly calibrated. Here's how that breaks human-in-the-loop gating in production agents — and the cheap fixes that work.
Read more
Golden eval sets silently rot within months. How we source real edge cases, version evals alongside prompts, and refresh without invalidating history.
Read more
An anonymised case study of an AI agent that worked in pilot but blew the budget at scale. Where the cost went, the fixes, and the before/after numbers.
Read more
AI-generated code ships exploitable flaws with total confidence. Why security review is the one step you can't delegate to the tool that caused the problem.
Read more
DeepMind's AI Control Roadmap signals the shift from hoping models stay aligned to securing agents like insider threats. Here's the executive mental model.
Read more
For agentic systems, the eval harness is what lets you ship with confidence and iterate safely. Here's what goes into a real one — and why teams pay for skipping it.
Read more
The specific failure modes that kill RAG systems after the demo — retrieval drift, chunk boundaries, stale indexes — and how to catch them with evals first.
Read more
A decision framework for CTOs and Heads of AI in 2026. Four paths — Build, Buy, Partner, Wait — with honest costs, timelines, and failure modes.
Read more
The month-four death has remarkably consistent causes. None of them are technical. Four organisational failure modes that kill AI pilots.
Read more
Case study: 100+ vehicles, 2 years, 40% downtime reduction. What we built, why it worked, and the patterns worth stealing.
Read more
There is a specific shape that enterprise AI projects take when they actually ship. Here's the week-by-week breakdown.
Read more
The question is almost never framed correctly. A framework for deciding how to stand up AI capability — honest about where each option breaks.
Read more
Most enterprise AI projects fail. Here's what separates the ones that ship from the ones that stall — based on what we've seen across 6 industries.
Read more
You need senior AI leadership but aren't ready for a full-time hire. Here's how the fractional model works and when it makes sense.
Read more