September 2026 • 7 min read
Choosing the right model and effort for the job
On one model, going from medium to max effort bought seven benchmark points for four and a half times the cost per task. The price sheet doesn't show that. Your invoice does.
Read article
September 2026 • 5 min read
Thirty-nine cents bought the draft, not the judgment
One agent run returned eighteen of eighteen sections of a go-to-market plan, then scored its own work at 0.82. That second number isn't evidence.
Read article
September 2026 • 6 min read
The tree you measured is not the tree that's running
Two working checkouts of one repository, both on main, one commit apart. Everything passed. Some of it was answering about the wrong tree.
Read article
September 2026 • 6 min read
Documented is not operational
An architecture record marked Accepted means we decided, not it runs. On 30 August I was about to write up three capabilities as working. None of them were.
Read article
September 2026 • 7 min read
The pilot isn't a trial run, it's the design
Region one sets the naming, the hierarchy and the rules every later region inherits. By region four, changing any of it means a migration.
Read article
September 2026 • 7 min read
One test sorts a group standard from a local choice
Does it roll up in group reporting, or does it move between countries? Run the governance list through that and most of it comes back local.
Read article
August 2026 • 5 min read
A green checkmark is not evidence
Three tools I use every day reported success this year. All three were wrong the same way: a check that never exercised the thing it claims to cover.
Read article
August 2026 • 6 min read
A pass is only as big as what it looked at
A link checker validated two links in two documents holding twelve external URLs, then raised 217 warnings about a virtualenv. The pass was real. The scope wasn't.
Read article
July 2026 • 7 min read
Build the shape once, fill it in locally
Compliance rules differ by country. The structure holding them doesn't have to. Four audit regimes over one control set taught me the same move.
Read article
July 2026 • 7 min read
Why I won't tell you which platform to buy in the first meeting
The platform is rarely why a rollout struggles. Governance and data usually are, and reading those honestly takes longer than one meeting.
Read article
April 2026 • 7 min read
Every model you depend on has a sunset date
Claude Sonnet 4 and Opus 4 retire June 15. GPT-5.5 just landed. Three model rotations in eighteen months. The teams with portable evals barely feel it.
Read article
April 2026 • 7 min read
Every cloud now sells an agent platform
Google Cloud Next, OpenAI Workspace Agents, Microsoft Agent Framework v1.0. Three platform launches in three weeks. The work they don't do is the work you still own.
Read article
April 2026 • 7 min read
When AI finds more bugs than your team can read
Why the March 2026 OpenSSF and Alpha-Omega funding matters for supply chain noise, triage, and how you run AppSec around AI-assisted scanning.
Read article
April 2026 • 7 min read
Three vendors launched agent stacks in March. You still own the hard parts.
GTC, Fusion, and SI headlines don't replace evals, messy integrations, or ownership. A field note on what changed and what didn't.
Read article
April 2026 • 6 min read
Identity before intelligence: what agent IAM forces you to decide
Three questions every platform team should answer before wiring another tool: what exists, what can connect, what's allowed.
Read article
March 2026 • 7 min read
Three reasons your AI pilot is stuck and what actually fixes them
Three structural reasons a pilot stalls: missing evals, underestimated integration complexity, and unclear ownership.
Read article