Blog
What fourteen prototypes taught us about building AI-native software

What fourteen prototypes taught us about building AI-native software

We ran a design sprint and built fourteen interactive prototypes, each a real workflow redesigned around what agents can now do. Laid side by side, they show one change in what an interface is for: when the software can do the work, the screen stops being where a person performs the job and becomes where they supervise it. This piece states that position, walks the split it draws through all fourteen prototypes, and names the two things the design depends on.

Updated
July 21, 2026
Reading Time
8 min

The position, up front

We ran a design sprint and built fourteen interactive prototypes, each one a real workflow rebuilt around what agents can now do: a triage console that acts overnight, a planning tool where the spec is the source of truth, a compliance watch that never sleeps, a drafting surface, an access layer, and nine more. To be plain about what they are: design-sprint outcomes you can open and click through, not software running for customers. They span very different domains, and they were drawn as separate explorations. But designing them back to back surfaced one thing they all share, and it is sharper than “these feel AI-native.”

For thirty years, business software was a place you went to do the work. When the software can do the work itself, the interface’s job changes: it becomes the place you supervise it. That is the position this piece argues, and the fourteen prototypes are the evidence. It rests on three arguments laid out in their own essays: that how much work you can hand to an agent depends on which stage of the work you’re looking at, that a written method is what makes an agent conform, and that AI-native software needs a product team rather than a handoff. First we show the split all fourteen prototypes draw, then what a supervision surface has to do, then the two things the design depends on.

The same split runs through all fourteen prototypes

The reason this arrived now is mundane: agents can finally carry the work forward between a person’s decisions, which nothing before could. What remains for the person is setting intent, taking the escalations, and checking what got done. The industry’s default response has been to treat AI as a feature: a chat panel bolted onto the screens you already had, the bolted-on copilot the problem page calls out, which makes the old process a little faster and changes nothing about how the work runs. The prototypes treat the agent as staff. The clearest way to see the difference is to lay the fourteen side by side and split each one down the same seam: what the agent took over, and what the person is left to supervise.

PrototypeWhat the agent takes overWhat you supervise
HiroOvernight ops triage; the safe, reversible actionsThe calls it escalates; where the may-act line sits
DriftWriting the code and reconciling it to the specThe spec, the ordering, and the drift it flags
SagaRunning the production conversationsThe ranked exceptions; the tests they're judged against
FidéChecking every policy, continuouslyWhat a rule means; who owns a breach
Data RoomsReading the pile into facts; drafting the replyThe reading; whether the draft goes out
AuthorShaping the draftThe beats, the voice, and the scope of each edit
FlowRunning the steps you hand itThe process; each step's manual/assisted/automated dial
Autonomous HumansThe paperwork between decisionsThe approvals where the bar is defended
NorthReconciling the inputs; rolling work up to the planThe conflicts; every decision on the plan
CozyRemembering, filing, and remindingNoticing what matters; choosing to reach out
LemurTutoring, strictly from the lessonThe learner checks each answer against the page
AuthComputing access live; acting under an agencyTeam membership and the agencies you grant
PrimerLoading and following the methodsAuthoring the methods; watching which ones fire
Design SystemBuilding the UIThe taste, written down as rules the builder follows

The same split holds across fourteen unrelated domains: the agent does the work, and the person supervises at the handoff. That split is the finding. Everything people notice on the screens (the inboxes, the dials, the provenance marks) falls out of drawing it.

What a supervision surface has to do

Once the person supervises rather than performs, a handful of things the old interfaces never needed become mandatory, and they’re the same across the gallery. The screen has to initiate. A supervisor can’t patrol every queue, so the work has to come to them: Hiro’s agent reaches you at 3am, Drift’s grilling agent opens the decision, North surfaces the $40M/$38M conflict, Fidé pushes the failing rows up. A dashboard you watch and a chat box that waits both leave the finding to the person; here the decision arrives addressed, with its evidence and a recommendation attached.

Control moves to the boundary. A supervisor stops deciding individual steps and instead decides how far the worker’s authority runs. That’s Hiro’s may-act/must-ask line, Flow’s per-step manual/assisted/automated dial, Author’s edit scopes, and Auth’s agencies: autonomy granted as an adjustable boundary rather than an on/off mode. The surface has to prove who did what, because two kinds of actor now share it: the rose color reserved for the agent, the evidence chips, the provenance on every fact, the audience printed on every Cozy pane. And the work has to be checkable without redoing it: Saga pins every claim to a named test, Data Rooms links every figure to its document, Drift shows a live number for how far the code has drifted from the spec.

One more thing separates a surface you’d trust from one you wouldn’t, and it’s a posture rather than a feature: these designs earn trust by what they refuse. The tutor in Lemur answers only from the lesson. Cozy drafts the message and won’t send it. Hiro logs every autonomous action, because an action taken silently is one a supervisor can never check. Drift’s agent must not average two owners’ disagreement into something confident-sounding. When the software can plausibly do anything, the boundaries a supervisor can read on the surface are what make the rest believable.

Two things the inversion depends on

A supervisor is only as good as what the worker knows how to do, which is why the hard constraint here isn’t model capability. It is capture: the binding limit on autonomy is how much of your method is written down. We measured it in our skills benchmark: hand the agent the written method and conformance jumps to roughly 90%; withhold it and the agent improvises. The prototypes keep building for this. Flow lifts a process out of people’s heads, Primer keeps the methods as a library, and Saga turns recorded runs into the next skill. Model releases will keep coming; none of them writes down how your team works.

The second dependency is why this has to be bespoke. A worker is only useful inside one organization’s actual thresholds, exceptions, and vocabulary: the $500 auto-approve line, the contractor-hold-SSO rule, the policy code that maps to a disclosure. Generic software can’t carry those. And because agents made the build cheap, the expensive part is no longer writing the code; it’s knowing which software to build, fitting it to one team, and keeping it changing after it ships. That is the case the problem and solution pages make in full.

From prototypes to production

Prototypes are the cheap way to find the shape of a thing before you commit to building it. The next step is pointing them at real teams: taking the capture idea from Flow and Primer into actual workflows, hardening the pieces every build shares, and finding out which of these designs survive contact with a customer’s Tuesday. The designs are all open to walk through. Every entry in the gallery opens as an interactive prototype with a case study beside it, and each case study carries a design-notes section on what’s genuinely different about its screens and why. If you’re midway through building something of your own, the useful question is the one that runs through the whole gallery: treat the agent as staff, and ask of each screen whether the person is doing the work or supervising it.

Wherever the answer is still “doing,” ask whether the software could now do that part and hand the person the judgment instead. Every design in the gallery started from exactly that question, asked about one screen of one workflow.

The company shape follows from the software shape. Software fitted to one team, and kept changing as the team changes, needs a product team that stays; that is why AI Hero sells a subscription to a team rather than a tool or an engagement. If something in the gallery looks like a workflow you run, the invitation from the bespoke SaaS post stands: send us a paragraph about the work and where it hurts, and we’ll come back with a one-page design within a week.

Article byRahul Parundekar

Rahul Parundekar

San Francisco-based consultant specializing in cutting-edge Generative AI (GenAI). I partner with organizations to pinpoint high-impact opportunities, streamline AI operations, and accelerate the launch of innovative products—efficiently, cost-effectively, and with controlled risk. Founder of Elevate.do and A.I. Hero, Inc.