5 min read

Agentic maturity models focus on the wrong thing

Every AI maturity model measures autonomy. YOLO mode, parallel agents, building your own orchestrator. But that's only part of the picture and an answer to the wrong question. The better question is - are you ready?
Agentic maturity models focus on the wrong thing

Agentic AI maturity models keep popping up and they all measure the same thing: how much autonomy you've given the agent. Permissions off. YOLO mode. 10 parallel agents. Building your own orchestrator etc. The assumption is clear: more hands-off equals more mature equals better!

And for speed / adoption of AI, that might be right. But I think they're missing a critical axis entirely, readiness.

The autonomy ladder

The orchestration framework Gastown introduced one of the many maturity models being talked about: source

Stage 1: Zero or Near-Zero AI: maybe code completions, sometimes ask Chat questions
Stage 2: Coding agent in IDE, permissions turned on. A narrow coding agent in a sidebar asks your permission to run tools.
Stage 3: Agent in IDE, YOLO mode: Trust goes up. You turn off permissions, agent gets wider.
Stage 4: In IDE, wide agent: Your agent gradually grows to fill the screen. Code is just for diffs.
Stage 5: CLI, single agent. YOLO. Diffs scroll by. You may or may not look at them.
Stage 6: CLI, multi-agent, YOLO. You regularly use 3 to 5 parallel instances. You are very fast.
Stage 7: 10+ agents, hand-managed. You are starting to push the limits of hand-management.
Stage 8: Building your own orchestrator. You are on the frontier, automating your workflow.

Others are less charitable, but have a similar staged approach: Vibe Psychosis

And I quite like this Reddit thread

These models all tell you the same story: progress means giving up more control. But that's only part of the picture.

The missing axis

The agent you use and how you use it matters. But most of the repeatable value comes from the processes, pipelines and safeguards you have in place that let you move safely (well, a bit more safely) up the levels of agentic automation.

You need to be ready for agentic workflows and volume of output they can produce, as if your not, you run the risk of speed running failure, instead of success.

Think about it this way (straw man, but gets the point across). Two teams, both at "Stage 6" on the Gastown model, running multiple agents in YOLO mode:

  • Team A: No test suite, no linting, no pipeline checks, no code review. Agents committing to main and pushing to production.
  • Team B: Strict linting, comprehensive test coverage, TDD, CI/CD gates, architectural review, feature flagged implementations, QC/QA/Review before release to customers, clear metrics and instrumentation etc

Same "maturity level". Wildly different outcomes.

Team A is playing Russian roulette with their codebase, likely racking up debt. Team B is actually mature, delivering measurable and repeatable value.

The intern test (again)

I keep coming back to the intern analogy. But it's not my original idea: Punya Mishra put it brilliantly:

I have come to realize that working with generative AI is like having, at your beck and call, a really smart, but (occasionally) drunk, intern.

If you hired a junior developer, you wouldn't measure your software development maturity by how quickly you stopped reviewing their code. You'd measure it by how good the processes and systems are at giving them the freedom they need to do their job with support/guidance to deliver value, while minimizing business risk.

The more you take your hands off the wheel, the more maturity you need in the systems and tooling to let you do so. You have to be able to catch issues and provide feedback. That maturity needs to be built into the system, not assumed. We don't let developers edit live on prod (or we shouldn't), no matter how competent we think they are. We have systems, processes and checks to make sure we don't mess up.

And seeming competent is something AIs are very good at. Our path to better and better AI is littered with examples of these seemingly smart agents being unable to count the number of R's in Strawberry, or recommending you walk to the carwash because it's only 10 mins away, to wash your car. But, at the same time, those models have been able to pass the bar since 2023.

We need deterministic checks and balances to keep everyone, agents and humans alike, from making silly (or not so silly) mistakes. We look to our peers (again, agentic or human) to help make sure we haven't missed something obvious, or gone down a wrong path.

What agentic readiness looks like

How do you do more with AI? You do more without it, in most cases.

  • You have strict linting and code standards that catch problems before they compound
  • You have strict unit, integration and e2e testing standards that prove things work (and stay working)
  • You have feedback loops for quality, cost, value, correctness and architecture
  • You have vision of the why & who you are building for, not just the what, how and where

Most of the time, that means deterministic pipelines and quantitative measurement/gatekeepers for the first two, and qualitative review (perhaps AI-assisted) for the last two.

This isn't new. This is how you scale "traditional" software development teams. The difference is that AI has massively increased your effective team size overnight. You might still have the same number of humans, but you're producing code at the rate of a much, much larger team.

And with that increased output, you need the systems and sophistication to match. The processes that worked for a team of five won't hold up when those five are producing the output of twenty five+.

Without guardrails and controls, you won't know if the team (human or agentic) is speeding you towards success or catastrophic failure.

The Cloudflare team that rewrote NextJS support in a week didn't succeed because they turned everything to YOLO. They succeeded because they had a battle-tested test suite, clear architectural boundaries, and a senior engineer at the wheel. The agents were the engine. The humans and systems were the steering and brakes. They were ready.

Orchestration frameworks like Gastown and others that have sprung up are trying to address this by building review and feedback loops into the orchestration layer. That does help, just like it does with humans. But as we've seen, these AIs still able to make stupid mistakes, especially when checking their own work (just like us).

Using a 2nd foundation model to check the results of the primary model is a good step, but I find human in the loop still provides significant value, especially for more complex ideas/tasks, especially at the start and end of the execution.

AIs are not original, nor do they get inspired. They still need humans for the ideas and judgement calls.

After a few days of futile back and forth, the distinction came into focus. Humans are for ideas, AI is for execution.

So, before you start climbing the autonomy ladder, and scaling your execution too fast, ask yourself: are you ready? Do you have guardrails and support systems that are able to scale with you?

Because if not ready, you're not maturing, you're just getting better at making more mistakes, faster! (see https://www.acmconsulting.eu/post/tequila-vs-agentic-ai/)

Looking for more advice / guidance / support / mentorship ?

Please take a look at my Technology Consulting service, I might be able to help.