Skip to content

Insights

Notes on decisions, processes, and systems.

Short, practical writing on how organizations actually improve: framing the real problem, using evidence well, and building changes that last. We use data and AI where they help and say so plainly where they do not.

August 1, 2026

The word 'agent' stopped meaning much this year

Most products marketed as AI agents this year turned out to be automation with a new label.

A word can do a lot of damage once everyone starts using it for something slightly different. That happened to 'agent' this year. Industry analysts estimate that of the thousands of products now marketed as AI agents, only a small fraction, something like one in twenty, actually pursue a goal across multiple steps using real tools without a human approving each one. The rest are chatbots, scripted workflows, or ordinary automation wearing a new label.

The label matters because the buying decision changed underneath it. A year ago the question was whether a vendor had any AI at all. Now the question is whether what they built can finish a multi-step task on its own, and a rebranded script cannot do that, no matter how the demo is staged. The cost of getting that distinction wrong lands on whoever bought the thing.

The test practitioners converged on is simple enough to ask in a vendor meeting: give it a real, multi-step goal, hand it the actual tools it would use in production, and see how far it gets before a human has to step in. Anything that needs a person to approve every individual step is automation with a new name. Most of that automation is genuinely useful. The ask is just to call it what it is.

This is also the harder half of judging your own pilots, beyond a vendor demo. A model version that scores well on average can still be quietly failing the one case that matters, the same way a scripted tool can look agentic in a slide and fold the moment the task branches. Judging either one honestly takes the same discipline: specific cases, specific stakes, checked one at a time.

See this in the Agent version comparison

July 14, 2026

Agents reached production before governance did

Enterprise agent deployment jumped this year. The rules for supervising it did not keep pace.

Enterprise software is embedding agents at a pace that outran expectations this year. One mid-2026 estimate put embedded agents in roughly 40 percent of enterprise applications, up from under 5 percent at the start of last year. That is a real structural shift in how software gets built: agents that used to sit in a pilot are now running in production, taking actions, and touching real systems.

The oversight has not moved at the same speed. Analysts have flagged well over $200 billion in enterprise software spending as exposed to risk from agent deployments running ahead of governance, and infrastructure regulators have started asking for the basics, an override path, a record of what an agent decided and why, that a lot of production agents still do not have. The technology got easier to deploy faster than the practice of supervising it matured.

The fix is being specific about two different questions an agent has to answer before it acts: can I resolve this with confidence, and am I even the one allowed to decide this. The first is a judgment call the agent can improve at over time. The second is a rule that does not move no matter how confident the agent gets. Treating both as the same kind of caution produces exactly the failure mode teams are trying to avoid: a pilot that approves everything, or a pilot that asks a human to approve everything.

None of this requires a big governance program to start. It requires writing down, for every agent already running, which of its decisions are judgment calls and which are rules, and who signs off on the second kind. That list is usually short enough to make in an afternoon, and it is the difference between an agent that scales safely and one that becomes next year's line item in a risk report.

See this in the Agent pilot walkthrough

June 12, 2026

Shadow AI is already in your organization

By the time leaders debate whether to allow AI, their teams are already using it.

Most leaders we talk to are still deciding whether to allow AI. Their staff settled the question months ago. People are already pasting documents into chatbots, drafting donor emails, summarizing meetings, and cleaning up spreadsheets with whatever tool is open in another tab. The decision got made on the ground, by the people doing the work.

The reason is ordinary. A useful tool sits one click away, the official guidance is silence, and people reach for what helps them get through the day. The hard part is that nobody knows where it is happening, on what data, or with what review.

Start with visibility before policy. Ask one question across a few teams: where is AI already showing up in how we work? The answers tend to sort into three buckets. Work that is genuinely safe to keep doing. Work that is fine once you add a data rule or a review step. And a small number of workflows that should not be automated yet.

Those three buckets make the governance conversation concrete. You are writing guardrails around what is actually happening, with real examples in front of you, which is far easier than drafting rules for a hypothetical.

See this in the Shadow AI risk checklist

May 19, 2026

Most AI pilots fail for operating reasons

The model usually works. The handoff, the ownership, and the review are what break.

When an AI pilot stalls, the post-mortem usually reaches for a technical explanation. The model was not accurate enough. The data was messy. The tool was a poor fit. Sometimes that holds up. Often the technology did its job and the pilot still went nowhere.

The failures we see most are operating failures. No clear owner once the excitement faded. No agreement on what good looked like, so nobody could say whether it worked. No destination for the output, so it sat in a document nobody opened. A review step that everyone assumed someone else was doing.

These are ordinary problems, the same ones that decide whether any new process succeeds. AI raises the stakes because its output looks finished and confident even when it is wrong, so a missing review step does more damage here than it would almost anywhere else.

A pilot worth running has a named owner, a definition of success you can check, a real destination for the output, and a human review point someone is accountable for. Get those four right and the technology is usually the easy part.

See this in the Agent pilot walkthrough

April 22, 2026

Pick the workflow before you pick the tool

Start with one workflow you already run, and the tool question gets small and specific.

The most common way teams waste money on AI is to start with the tool. A vendor demo impresses someone, a license gets purchased, and then the real work begins: finding something for it to do. The tool ends up setting the agenda, and the spending starts before anyone has defined the problem.

Start with one workflow. Pick something repetitive that you do often enough for a small improvement to compound. Map it plainly. What is the job, what data does it touch, where does a human need to decide, and what happens when it goes wrong. One page is enough.

That page tells you almost everything. Whether AI even fits. Where the review has to live. What data must stay out. What better would actually mean. By the time you reach the tool question, it is small and specific, because you know exactly what you need the tool to do.

Teams that work this way spend less and trust the result more, because they understood the work before they automated any of it.

See this in the Initiative prioritization model

Get new insights by email

A few a year, each about one specific problem in decisions, processes, or systems.

Bring us the decision, process, or system that is not working.

We will help you understand the problem, determine what the evidence supports, and build a better way forward.