Notes from the studio.
The paradigm, prototypes, and use cases behind bespoke SaaS. Most posts are open; a few ask you to sign in.
Opus 5 looks great on paper. In practice, a lot of people are disappointed.
Anthropic shipped Claude Opus 5 on July 24, 2026, and it took the top of the Artificial Analysis intelligence index at 61, five points above the Opus 4.8 it replaced. Two weeks later a fair number of developers have gone back to 4.8. The complaints converge on one behavior: the model widens the scope of whatever you asked for, then verifies the extra work it invented. Anthropic's own prompting guide concedes both, and tells you to delete the verification instructions you wrote for the last model. Here is what the threads actually say, sourced and linked, plus the guardrails that cut down the wandering.
What fourteen prototypes taught us about building AI-native software
We ran a design sprint and built fourteen interactive prototypes, each a real workflow redesigned around what agents can now do. Laid side by side, they show one change in what an interface is for: when the software can do the work, the screen stops being where a person performs the job and becomes where they supervise it. This piece states that position, walks the split it draws through all fourteen prototypes, and names the two things the design depends on: a written method (which lifted conformance to roughly 90% when we measured it) and thresholds so specific to one company that only bespoke software can carry them.
Bespoke SaaS: the build stopped being the hard part
We built fourteen prototypes in short design sprints, and that speed is the evidence under this essay: agents now do the volume work of building, so the build has stopped being the hard part of custom software. What stays hard is that fitted software has to keep changing, because usage, priorities, scope, and your own people keep moving, and a handover can't absorb any of that. This piece compares the four ways to buy software (traditional SaaS, consulting, forward-deployed engineering, bespoke SaaS) on fit, who runs day two, how changes happen, and what you pay for, and makes the case for a product team that stays, priced like SaaS.
Hiro: an ops triage console where the agent can reach its human
We built Hiro in a design sprint: a prototype ops triage console where the agent clears the overnight pile itself, acts inside the bounds you set, and reaches out when a decision needs its owner. The demo run takes 214 overnight items down to 5 held for a person, with every autonomous action logged (demo data throughout). The interesting part isn't the triage. It's the direction of the conversation: instead of you checking in on the agent, the agent has a way to reach you.
Drift: the spec is the desired state, the code is the actual state
Drift is a design-sprint prototype of planning software for teams whose code is written by agents: the spec is the desired state, main is the actual state, and the difference between them is the work, measured. In the demo the Code layer reads 92% in sync, 2 gaps, 2 building. The product team tends a three-layer board and an ordered graph of spec PRs with the reasoning attached; coding agents pull main toward the target branch by branch, and scheduled checks flow drift on main back in as new work. We call the shape the flow model. If you plan software for a living, this is a look at what your job might become.
Saga: a control room for a workforce of agents
We prototyped Saga in a design sprint: one console for the record of how a workforce of agents did its work, built for the three people who need that record. The auditor gets policy-linked evidence, and the EU AI Act makes the record law for high-risk systems from August 2, 2026, with logs kept at least six months. The engineer gets the trajectory, the only honest explanation of how the work got done. The owner, the reader most teams miss, gets hill-climbing: captured routes become skills, so this quarter's exploration becomes next quarter's reliable pathway. The demo's numbers are staged and labeled as demo data.
Fidé: a continuous watch for the policies no auditor certifies
We prototyped Fidé in one of our design sprints: a compliance platform for the policies that are yours alone, from pre-commit configs to the cafeteria kitchen standard. Certification platforms automate SOC 2 and ISO 27001 well, and most now accept custom controls, but a control the vendor has no automated test for falls back to manual evidence, so internal policy still runs on the honor system. Fidé's model is small: controls group into checklists, checklists apply to targets, and agents keep every check current, so the supervisor's screen is a list of what's wrong right now, not a report assembled for an audit.
Data Rooms: software that does the reading before you ask
We prototyped Data Rooms in one of our design sprints: a room that reads a pile of personal records into structured data before you ask it anything. The demo room reads six lab reports of mixed provenance (two labs' PDFs, a photo of a printout, an email attachment) into an LDL trend that climbs from 118 to 142 and stands on the overview before the cardiologist asks. The room admits what it hasn't read (81 of 84 documents structured, the three unread ones named), drafts the appeal for a denied claim without sending it, and flags the referral letter that isn't in the room. The same shape fits fundraising data rooms, buying a house, and customer onboarding.
Author: send a paragraph, get a finished draft
We designed Author in one of our design sprints: a drafting prototype where you supply the beats (a dictated paragraph qualifies), a style guide mined from your published writing carries the voice, and the agent does the shaping. The voice lives in a named, inspectable profile with a reach-for list and a ban list, and every edit path scopes what the agent may change: the whole draft, a block, a line, or a phrase. We revised this post with the same loop. This is a design case study, not a product announcement.
Flow: capturing the real process into software the team works in
We spent a design sprint on Flow, a prototype that captures a company's real process into software the team works in, not a document that goes stale the day it's written. You describe a step by talking, and an agent drafts the step, its owner, a control, and a memory note; nothing saves until a person approves. Every step wears a manual, assisted, or automated dial that turns both directions, and the exceptions (revoke on a failed background check, hold a contractor's SSO) sit in the step where everyone can see them. A non-technical walk through the design decisions.
Autonomous Humans: running an events program end to end
We prototyped the operations software for Autonomous Humans, an initiative that teaches non-technical people to use AI through courses, cohorts, and a rolling program of online events. Event platforms host a single event well; the program itself, a dozen sessions in flight at different stages, lives in a spreadsheet. The prototype treats the program as the unit you operate: every session on one screen, each walking the same planned-to-live route, with agents drafting the routine messages between an operator's decisions.
North: making strategy something you can operate
We prototyped North in a design sprint: one surface where every piece of work, human or agent, rolls up through function KPIs into the company's objectives. The OKR ritual assumes work arrives at human speed and gets checked against the plan a few times a year; agent work lands continuously, so both direction and verification need a live surface. This walkthrough shows a plan drafted by voice, a $40M-versus-$38M conflict surfaced with both sources attached instead of averaged away, and the discipline that keeps the map honest. Every number on the screens is sample data.
Cozy: a relationship manager that does the remembering and refuses to send
We spent a design sprint on Cozy, an interactive prototype of a personal relationship manager that takes over the remembering so the reaching out can stay human. You speak a note and the app files it to the right person, creating the profile as it goes; reminders arrive carrying reasons (a promise you made, a rhythm you set, an event worth marking) against a cadence you set per person; and every surface says who can see it. The app drafts, files, summarizes, and reminds, and it will not send a warm message on your behalf. This is a UX exploration from a sprint; nothing here is a shipped product.
Lemur: a course platform where the tutor only knows the lesson
We spent a design sprint on Lemur, an AI-native LMS prototype, after one company refused generic AI training outright and asked for every lesson to be rebuilt around its own processes. The design turns on one decision: a learner can't grade the answers they're getting, so the tutor on every lesson page answers only from the material in front of them, states that scope before the first question, and flips from explaining to checking one tap away. Current tutor products draw the line at the platform's whole library; Lemur draws it at the lesson, which makes every answer checkable against the page beside it. If you're planning to upskill a team on AI-native tools, this is the design thinking, laid out so you can borrow it.
When agents build the UI, taste becomes pass/fail rules
We wrote the design system before designing any of the fourteen prototypes in our gallery: a fourteen-step ink ramp, one indigo accent, four fixed status meanings, and a composer on every surface, each rule phrased so a screen passes it or fails it. Coding agents write most of the prototypes' UI, and a builder with no memory of last week's taste decisions drifts on any rule that leaves room for judgment. The system travels as a shadcn registry, so an agent scaffolding a new screen pulls the real components and tokens instead of improvising its own. Fourteen separate explorations open looking like one studio made them because of it.
Auth: don't store the access answer, compute it
We spent a design sprint on identity for the agentic era and built Auth, an interactive prototype of the access surface: a portal that shows everything you can open, entitlement recomputed from your teams at every page load so a stale grant has nowhere to live, and an agency for each agent, a defined and auditable set of actions on specific projects tied to who it works for. Identity-governance products reconcile stored grants with reviews after the fact; this design removes the stored answer instead. The screens are a prototype, and the open questions are named.
Primer: a library for your team's agent skills
We designed Primer in a sprint: a prototype of the library a team needs once its agent skills outgrow the git repo, one place to define, update, share, and watch them. The design turns on one mechanism: a skill runs because the agent matched the request's words against the skill's one-line description, which makes the description the trigger and makes usage measurable. The prototype counts fires, and its most actionable number is a zero: a skill whose wording never matches how people ask. This is the walkthrough, on sample data, with the thinking behind each screen.
Agent Skills transfer to goose, but the model decides the payoff: a SkillsBench replication
I ran SkillsBench on goose, the open-source agent the Agentic AI Foundation stewards, because the benchmark didn't include it: a matched, paired no-skill/with-skill comparison against Claude Code and OpenHands, on both Claude Opus 4.8 and GPT-5.5, over 40 of the 87 tasks per pairing. On Opus 4.8, goose's +17.7-point uplift matches the strongest harness in the paper, with zero regressions. On GPT-5.5 the identical configuration gets nothing from skills, even though it opens them almost as often. The payoff is mediated twice by the model: whether it reaches for the skill, and whether reaching for it changes anything. Single attempt, one machine, 40 of 87 tasks; the caveats are inside.
Why agents need skills: doing the job the way you want, unsupervised
We measured what it takes for an agent to follow a Standard Operating Procedure unsupervised: five agent configurations across three benchmarks with deterministic oracles, all on claude-haiku-4-5. Ship the SOP as a skill the agent can discover by itself and single-turn task success rises by +0.35 (95% CI [+0.26, +0.45], p<0.001), while multi-turn conformance on τ²-bench climbs from roughly 50 to 59% of prescribed actions to about 90%. The delivery channel turns out to be a wash, and externalized run-tracking costs 2 to 4× the tool calls for no gain. The case for skills is auto-discovery: with no user present, a user-invoked prompt never reaches context.
The tideline: where AI belongs in the enterprise
All work runs through the same six stages, from setting the goal to measuring the impact. Two lines cross them: the shoreline, where your agents stand at each stage today, which every model release pushes up, and the tideline, the mark showing how high agentic work can reach in your enterprise, which moves only when you write the method down, make verification independent, and state objectives in falsifiable form. Everything above the mark stays with a human at the helm, at Replit as much as anywhere. With the one-question diagnostic and a 60/30/10 budget that funds the tideline rather than the shoreline.
Agents Are Looking Inward. The Work Is Outward.
Two months of sabbatical, and one thing kept nagging me: almost everything we build for agents points inward—at the agent's own tools, context, and memory. Context engineering is the flagship of it. What's missing is the outward view: an agent that understands how the people and the organization around it actually do the work. Here's where my head is.
The Icon Saga: From LLM Generation to Lucide Lookup
We spent 40+ commits and a six-hour sprint making an LLM draw icons for grocery items, todos, and weather cards. Then we deleted the pipeline and replaced it with a 1,666-icon library and a keyword match.
Before bespoke SaaS: what building consumer agents taught us
The original January 2026 launch post for Hiro, our consumer agent app, rewritten in hindsight. Four voice-first agents, cards that changed with the conversation, and one persona named Elise: what the B2C investigation taught us, and what carried into the bespoke-SaaS model.
The Agent Factory: A Monorepo Built for AI Coding
Treat your development environment as a product for the coding agent — one monorepo plus disciplined CLAUDE.md recipes — and cross-service features collapse into single-session, single-commit changes. Three examples from one week of building AI Hero Studio.
The Rise of Agentic Commerce: What Every Brand Needs to Know About AI-Powered Shopping
A guide for e-commerce business owners on AI-powered shopping: OpenAI's Instant Checkout, Shopify's Commerce for Agents, AEO, and a practical roadmap for making your products visible when AI agents go shopping.
Natural Language to SQL: A Production Guide for Enterprise Data Access
LLMs already write syntactically correct SQL. Production text-to-SQL fails on meaning. A field guide to the semantic layer, pinned-query memory, and per-tenant context that make natural-language data access reliable.
Building GenAI Voice Agents: Voice Bar
Why push-to-talk beat always-on voice for a screen-first webapp: PTT mechanics, transcript design, card-based tool output, and a dated changelog of what weeks of daily use changed, including the drawer that became a persistent conversation panel.
Building GenAI Voice Agents: Implementation
The hardest part of a phone-based voice agent isn't the voice—it's retrieving the right answer fast enough to say it. The model spends 200-300ms of a budget that runs out around 500ms, which leaves roughly 100-200ms for tools. This guide covers how we solved that with an in-process index pre-loaded during the greeting, plus the full architecture for OpenAI Realtime SIP integration.
Building GenAI Voice Agents: Evaluation
A voice agent that passes every text-style metric can still fail badly on a live call, because voice failures are temporal and acoustic. The metrics worth tracking on both axes, how to run automated simulation with platforms like Hamming and Coval, and the open/axial coding loop that turns real user failures into test cases.
Building GenAI Voice Agents: Architecture Guide
A working guide for teams putting voice agents into production: the two architectures and how to choose between them, where to run the agentic loop, the vendor ecosystem, conversation design, safety, day-2 operations, and a cost model you can defend to your CFO. The short version: the decision is a latency-versus-governance tradeoff, and most real deployments run both architectures and route per conversation.
AI Hero Achieves SOC 2 Type II Compliance
AI Hero has completed its SOC 2 Type II audit: an unqualified opinion with no exceptions, covering security, availability, and confidentiality controls over a three-month window. What the audit covered, and how to request the report for your vendor assessment.