Blog
The Tideline

The tideline: where AI belongs in the enterprise

All work runs through the same six stages, from setting the goal to measuring the impact. Two lines cross them: the shoreline, where your agents stand at each stage today, which every model release pushes up, and the tideline, the mark showing how high agentic work can reach in your enterprise, which moves only when your organization moves it. Everything above the mark stays with a human at the helm.

Tags
Enterprise AIAgent AutonomyGovernance
Updated
July 22, 2026
Reading Time
9 min

Models raise the shoreline; only your organization raises the tideline

On 16 July, Replit's founders published a field report on becoming what they call a self-driving company: every employee now has a manager agent that can spawn further agents, and lines of code contributed rose 5.8× between January and June. If you run an AI program inside a normal enterprise, that post probably landed next to a strange pair of facts from your own shop: one agent already carries whole tickets from spec to merged change, while a workflow everyone agrees is automatable has sat untouched for a year.

Here is the picture I draw when I have to explain that split. Any piece of work runs through six stages, from setting the goal to measuring the impact; lay the stages out as a stretch of coast and let agent capability be the water. The shoreline is where your agents actually stand at each stage today. You rent that line: every model release pushes it up, on the labs' schedule, whether you do anything or not. Above it sits the tideline, the brown mark on the sand: the highest point agentic work can reach in your enterprise as it is built today. Models raise the shoreline; only your organization raises the tideline, with what it has written down, made checkable, and instrumented. And everything above the mark stays with a human at the helm.

Nathaniel Whittemore calls the economy-wide version of this gap the capability overhang: models can already do more than organizations deploy. The six stages are where the overhang becomes actionable, because the distance between the two lines is different at every stage, and so is the work that closes it. The route from here: the six stages, the two lines drawn across them, who stays at the helm, and the budget split that raises the tideline.

All work runs through the same six stages

StageWhat it isWho decides today
01 · Set goalsThe objective function. Objectives, key results, the definition of better, informed by AI reading 06's measured impact against internal and external data and decided by a human.Human
02 · Define how the work is doneThe method. Standards, guardrails, the codified process.Shared, and the bottleneck
03 · Plan the work to be doneDecomposition. The plan, the specs, the backlog.Shared, drifting agentward
04 · Do the workExecution.Agent
05 · Verify the work is doing the workConformance. Does the artifact meet its specification?Contested
06 · Measure impactValidity. Did the metric move, and was it us?Human

The six stages repeat at every size of work. A quarterly objective runs through all of them; so do an epic, a ticket, and a single agent turn, each with its own small version of setting the goal, doing the work, and checking it landed. That is why "our AI maturity is level three" tells you so little: on the same Tuesday, an organization can run near-total autonomy at stage 04 inside a ticket while the quarter's objective is still set by a human alone. Stage by stage, altitude by altitude, is the only resolution at which "how much of this can agents do" has an answer. The stages are also the same in finance and in engineering, which is what lets the capability compound across departments instead of staying trapped in one. (I first worked the six stages out in an earlier essay on why agents keep looking inward.)

Two features of the stages carry the rest of the essay. Stages 01 and 06 are one act performed twice: declaring what counts, before the work and again after it. Most companies weld 01 into a strategy deck and then find at 06 there is nothing falsifiable to audit impact against. And stages 05 and 06 are different tests: 05 asks whether the artifact meets its specification, 06 asks whether the specification was worth meeting. Keep them separate; the price of confusing them shows up two sections from now.

Two lines cross the stages, and you control one of them

Now draw the two lines across the stages. The levels in the chart are illustrative; the shape is the argument. When we benchmarked our own agents, autonomy behaved exactly this way: near-total where the task had a machine-checkable definition of done, collapsing where it did not (the measured version is here).

Chart of agent autonomy across six stages of work. A solid shoreline shows where agents stand today; a dashed brown tideline above it marks how high agentic work can reach in the enterprise; the zone above the mark stays with a human at the helm. Both lines peak at 'do the work' and thin toward 'set goals' and 'measure impact'.

The tideline (illustrative). Solid: the shoreline, where your agents stand at each stage today; models raise it. Dashed brown: the tideline, the mark your organization's readiness sets; only you raise it. Hatched: the gap the water fills as models improve and you deploy them. Above the mark: a human at the helm. Shape, not scale.

The solid line is the shoreline. It runs deepest at stage 04 because execution is where machine-checkable definitions of done live: "the tests pass" is a fact rather than an opinion, and wherever a task has that property (reconciliation has it, claims adjudication mostly has it), the water runs deep. Coding agents pulled stage 04 up in about a year at most companies I sit with, and that is what a rising shoreline looks like on the ground: nobody reorganized, the models improved, the line moved.

The dashed brown line is the tideline. It marks how high agentic work can reach in your enterprise as currently built, and its height is set by things no model release touches: whether the method at stage 02 is written down or lives in senior heads, whether stage 05 has a check independent of the agent that did the work, whether the objective at 01 was ever stated in a form 06 could falsify. The hatched gap between the lines closes on its own as models improve and you deploy them. The tideline's height changes only when the organization changes.

This is where the overhang bites. At the stages where your tideline sits low, capability is rarely the missing piece: models can already draft decompositions and check conformance far above most enterprises' tidelines. Nearly every enterprise I spoke to at the AI Summit in London this June had the constraint inverted, waiting on a model while sitting on a decade of tacit process nobody had written down. Raise the tideline and the shoreline catches up almost immediately, because the capability was already there.

Raising the tideline is the actual content of Replit's report, read with this chart in mind. The manager-agent layer, the escalation policy, and a year of practice are tideline work; none of it waited on a model. A self-driving company is one whose tideline has been pushed close to the top of the chart across the middle stages, so each model release pays them the day it ships. AI-native companies start with a head start there, carrying less legacy method and a habit of writing process down as software. But the direction of travel is open to everyone: the tideline is yours to move, at whatever speed you choose to move it.

The diagnostic for any candidate workflow is one question. What is the assertion that would fail? A workflow where someone can state it ("the ledger reconciles", "the claim matches the policy terms") is one where the tideline can rise this quarter. A workflow where nobody can, as in most brand strategy, is one where the tideline holds no matter which model you buy, and agent skill has nothing to do with it.

Before funding an agentic workflow, make someone state the assertion that would fail. If nobody can, spend the money writing the method down instead.

Above the tideline, a human stays at the helm

Everything above the tideline belongs to a human at the helm, and the striking thing in the Replit report is how plainly they keep it there at the very moment their middle stages run close to full autonomy:

A self-driving company is not one without people. People still choose the destination. They decide which problems matter, make difficult tradeoffs, exercise taste, and take responsibility for the outcome.

The reason stages 01 and 06 keep a deep human zone has little to do with capability. Frontier models are already decent at proposing an objective set, red-teaming it, and running the attribution afterwards. But an organization that delegates the authorship of its own objectives cannot treat hitting them as evidence of anything; the loop closes on itself. Someone has to be answerable for the goal and the verdict, and answerability can be spread across more names but it cannot be handed to a system. London's summit ran a headline session this June called "Humans at the Helm", and the phrase has the geometry right: the human's job migrates from standing in the middle of the work to holding the goal and the verdict.

I would not claim that boundary is fixed; it has already moved once in plain sight. Two years ago, human in the loop meant reading every diff; at Replit today it means taking escalations from a fleet. The helm will keep thinning as models absorb more of the drafting and checking beneath it, and a company comfortable with more delegation at 01 and 06 than I have described may well exist by the time you read this. What the record shows so far is narrower: every organization I can point at, including the most autonomous one on record, still keeps a named human choosing the problem and owning the outcome. When that changes, it will change on the calendar of boards, customers, and regulators, which is a slower calendar than the labs'.

Stage 05 is the trap: it looks like more execution

One stage deserves its own warning, because the chart makes it look like an invitation. Stage 05 sits next to the deep water at 04 and shares its texture: mechanical, testable, cheap to automate. Every instinct trained at execution says delegate it the same way. The gap at 05 is real, but closing it by handing verification to the agent that did the work, or to its clone, leaves the work effectively unchecked: a judge that shares the generator's priors shares its blind spots, which is the mechanism behind reward hacking and behind model-as-judge scores that flatter models of the same family.

The rule is short enough for a policy document. Verification must be independent along at least one axis. A different model family, a different author, a different objective function, or best of all a signal from the world: a real transaction, a real user, a real ledger. Set that against VivaTech's 2026 Trust Barometer, where 89 percent of executives said they trust AI to guide their company's decisions, while the discipline for checking that guidance is still being invented. The distance between those two facts is what the next two years of enterprise AI get spent closing, or paying for.

Fund the tideline: 60/30/10

The budget follows from the chart. Models will keep raising the shoreline whether you fund them or not, and agentic work stops at your tideline. A program that spends everything on execution is buying speed under a ceiling it never touches.

ShareWhere it goesWhy
60%Let the shoreline riseExecution, in domains where a machine-checkable definition of done exists or can be written this quarter. Payback is fast and uncontroversial and funds the rest. Nothing here needs to be clever; run it industrially.
30%Raise the tidelineStages 02 and 05: write the method down and make verification independent. Every guardrail, SOP, eval suite, and golden dataset permanently lifts what can be delegated at 03 through 05. The models are rented from a vendor; the written-down method stays yours when you switch.
10%Instrument the helmStages 01 and 06: objectives stated in falsifiable form, attribution you would defend to a CFO. A small budget with disproportionate leverage, because it makes the other ninety percent legible as investment.

Two governance rules travel with the split. Put owners on the seams rather than the stages. Value leaks at handoffs: where the goal becomes method, where the plan becomes work, where the work becomes evidence. The six stages have five such handoffs, so name five signatories. And keep a floor above the shoreline: if less than a third of AI spend lands at stages 01, 02, 05, and 06, you are buying execution throughput and leaving the tideline exactly where it was.

Four questions for your AI program:

  1. For what share of our work does a machine-checkable definition of done exist, and who is writing the next one?
  2. Is our method in a document an agent can read, and who owns keeping it true?
  3. Which agent grades the work, and along which axis is it independent of the agent that did it?
  4. If we hit the number, could we prove it was us?

Then draw the two lines for one live workflow, with the operators who run it in the room. The distance between where your agents stand and where your tideline sits is next quarter's roadmap, and raising the tideline is work you can start on Monday without waiting for anyone's model release.

References

Article byRahul Parundekar

Rahul Parundekar

San Francisco-based consultant specializing in cutting-edge Generative AI (GenAI). I partner with organizations to pinpoint high-impact opportunities, streamline AI operations, and accelerate the launch of innovative products—efficiently, cost-effectively, and with controlled risk. Founder of Elevate.do and A.I. Hero, Inc.