The MarginAnalysis

Where the Human Belongs in an AI Workflow: What the Evidence Says

I made a video about one person running a five-specialist AI crew. The idea holds up. The title does not, and the best measurement anyone has published says something uncomfortable about the person in the middle: sixteen experienced developers were 19 percent slower with AI tools and believed they had been 20 percent faster. If you cannot feel the difference, you cannot manage by feel. Here is where the human actually belongs, with the dated studies.

The idea in the video is one I still believe. AI is not the shortcut around being a creator, it is what lets one person work like a crew: research, script, edit, design and voice, five jobs running from one chair with a person conducting.

The question the video does not really answer is the one that decides whether any of it works. Where exactly is that person, and what are they doing that the crew cannot?

"Conducting" is a lovely word and it is not a job description. So I went looking for what has actually been measured, and the honest answer is more specific and considerably less flattering than the version I told on camera.

First, the title

The video is called "The AI Workflow That's Replacing Entire Creative Teams." That is a good title and I do not think it is true.

The most serious ongoing measurement of this question is the Stanford Digital Economy Lab's work on AI and entry-level employment, built on "high-frequency administrative payroll data from ADP covering millions of U.S. workers." The current revision is dated 12 August 2026 and covers payroll through June 2026, which makes it about as fresh as evidence on this gets.

Its first finding is the one nobody repeats:

"We find no evidence of widespread, economy-wide job displacement."

Its second finding is the one everybody repeats, and it is narrower than the headlines suggest. Employment of young workers aged 22 to 25 in AI-exposed occupations "now stands 19% below" where it would otherwise be expected to sit. And the mechanism matters enormously:

"operates primarily through reduced hiring of young workers rather than increased separations"

Read that carefully. Teams are not being cleared out. Entry doors are closing. Nobody is being marched off; the junior role that used to exist is simply not posted. That is a different phenomenon with different consequences, and it is worse in a way that is easy to miss, because a redundancy is visible and a job that was never advertised is not.

There is a direct implication for the workflow in this article, and I would rather say it than dodge it. The tasks the AI crew handles well are exactly the tasks that used to be somebody's first job. First-pass research, first drafts, rough cuts, basic packaging. If you run this model you are not replacing a creative team. You are removing the rung you personally climbed.

I am not going to resolve that here. I am going to stop pretending the title was neutral.

The measurement that decides where the human belongs

Now the study that changed how I think about my own workflow, and I would put it in front of anyone about to build one.

METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real issues, each averaging about two hours, published 10 July 2025. The developers worked on repositories they already knew, with and without AI assistance.

"When developers are allowed to use AI tools, they take 19% longer to complete issues"

That result is startling on its own. What makes it the most important finding in this whole area is what the same people believed.

Before the study, they predicted AI would speed them up by 24 percent. After it, having actually been slowed down, they still believed it had sped them up by 20 percent.

Sit with the size of that. These were experienced practitioners, working in their own codebases, on real tasks, and their felt experience of their own productivity was wrong by roughly forty points and pointing the wrong way. Not slightly optimistic. Inverted.

The honest caveats, because they are real. The authors state plainly that they "do not claim that our developers or repositories represent a majority or plurality of software development work." They flag possible sampling bias, since developers who were confident AI helped them may have declined to take part. And they note they cannot rule out "learning effects for AI tools like Cursor that only appear after several hundred hours of usage." Sixteen people is a small study, it is software rather than creative work, and it is from 2025.

Take all of that seriously and one conclusion still survives, because it does not depend on the direction of the result at all. Practitioners were unable to perceive their own throughput accurately. Whether AI helps or hurts you, your feeling about it is not evidence.

That is the first and most concrete answer to where the human belongs: wherever the measurement is, because the intuition is demonstrably unreliable, and nobody selling you a workflow is going to tell you that.

What the five seats actually do

The video moves through five jobs quickly and confidently. Having run all five for months, here is the audit, seat by seat, with the overstatements marked.

Five seats, honestly rated

  • Research: genuinely does the job. Going wide across sources and angles in seconds is real and it is the largest time saving in the whole workflow. What it cannot do is pick the angle worth telling, which is the one decision that determines whether the video is worth making.
  • Script: genuinely does the job, as a first draft. You never face a blank page again. The draft is also, reliably, one-third generic. Cutting the lines that sound like everyone else is not optional polish, it is the work.
  • Voice: genuinely does the job. This is the seat that has improved most and needs the least supervision.
  • Edit: overstated in my video. I said "the edit assembles itself." It does not. Automatic assembly produces a rough cut with correct structure and no timing instinct, and timing is most of what editing is.
  • Design: overstated in my video. "A thumbnail generates on demand" is true and misleading. Generating one takes seconds. Generating one that is legible at small size, matches the title and is not the fourth variation of a look everyone else has takes several rounds and a person deciding.

There is a sixth thing missing from that list, and it is missing from the video too. The video names five seats and omits the one that checks the work. Our own blueprint specifies a Supervisor whose entire job is catching what is wrong before anything ships, and the film I made about this workflow skips straight past it to the montage.

I do not think that was carelessness. I think the reviewing seat is genuinely hard to make look impressive on screen, which is exactly why it drops out of every video about AI workflows including mine, and exactly why people build workflows without one.

The four places the human is not optional

Here is the replacement for "conducting," in four specific jobs that do not delegate.

  1. The decision. Which angle, which idea, which of the three hooks. The crew can generate twenty options and rank them; it cannot know which one is worth your name. This is the highest-leverage thing you do and it takes the least time, which is why people undervalue it.
  2. The standard. Rejecting adequate work. A crew produces competent output by default and competent output loses. The single behaviour that separates people who get results from people who get volume is a willingness to throw away a finished draft.
  3. The verification. Every number, quotation and date checked against a primary source. Not because the model is careless, but because it is confident, and confident wrongness published at speed is the characteristic failure of this entire way of working.
  4. The accountability. Your name is on it. The crew has no stake in your reputation and cannot acquire one. When something false goes out, there is exactly one person who pays, and the workflow diagram never shows them.

Notice what those four have in common. Not one of them is a task. All four are judgements, and they are the parts of creative work that were always the actual job, hidden under the hours of production that used to bury them.

That is the real reframing, and it is more demanding than the one in my video. The AI does not free you to be a director. It removes your excuses. When production stops taking the week, the only thing left explaining a mediocre result is your taste, and you now have to look at that directly.

The compounding claim, corrected

In the video I say the loop compounds: the more you run it, the less each round takes, the AI starts pre-filling from your taste, and your corrections shrink.

The outcome is real. The mechanism I implied is not. Nothing in a standard project-based setup learns from your corrections on its own. It does not observe you rejecting a line and adjust. Each session starts from the same instructions and the same knowledge files it started from yesterday.

The compounding happens because you write the correction back into the instructions. When it opens on a rhetorical question and you dislike that, the improvement is permanent only if you add a line saying not to. Otherwise you correct the same thing every week forever and call it a workflow.

So the honest version: the loop compounds if it is maintained, and decays if it is not. That is a small correction and it changes what you do on a Tuesday, which is the only kind of correction worth publishing.

The cost nobody puts in the diagram

One more thing belongs here, from Anthropic's engineering post on their multi-agent research system, published 13 June 2025. It is the clearest public account of how this architecture behaves, and it contains two sentences that no workflow video mentions.

"agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats"

"token usage by itself explains 80% of the variance, with the number of tool calls and the model choice as the two other explanatory factors"

Performance is mostly bought. The crew architecture is not a clever trick that makes a cheap plan behave like an expensive one, it is a way of spending considerably more compute to get a better result. Worth knowing before you plan capacity around it.

The same post also names where the shape fails: domains "that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today." Content production splits well. Plenty of work does not.

The test to run on yourself

Given that practitioners in a controlled trial could not perceive their own throughput, the only defensible position is to measure rather than feel. This takes almost no effort.

A two-week self-check

  • Count outputs, not sessions. Finished, published things. Not drafts, not chats, not hours spent feeling productive.
  • Log the wall-clock time from brief to published, including every correction round. The correction rounds are where the slowdown hides.
  • Track how many drafts you threw away. Rising is a good sign. It means your standard is intact.
  • Note every factual error you caught before publishing. That number is the value of the verification gate, stated in the only currency that matters.
  • Compare against the fortnight before you started. Imperfect, honest, and infinitely better than a feeling.

If the numbers say it is working, you have evidence rather than enthusiasm. If they say it is not, you have found that out in two weeks rather than two years, which is the entire point.

Come and build it properly

If you want to feel the difference between prompting a chatbot and directing a specialist, start free. Our Build Your AI Team course walks you through building the first seat, the Writer, on a free account, and the prompt pack is yours to keep with no account needed. The Claude prompting course sharpens the skill underneath all of it, which is describing a job precisely enough that someone else can do it.

When you want the whole crew and, more importantly, the gates that keep you in the chair, that is at ideasrepay.com, and the walkthrough is The One-Person Studio.

It is the build rather than an article about the build. Seventeen steps across five phases, click by click. You get the role prompts to paste straight in, so the crew exists in an afternoon instead of a month of trial and error: a Researcher, a Scriptwriter, a Voice, an Editor, a Designer and a Supervisor whose whole job is catching what is wrong before it ships. You get the three-way niche fit test that sets your earnings ceiling before you make anything, the full production run from a blank idea to a published video through the four quality gates you never hand over, the packaging and quality control checklists, the disclosure rules in plain language, and the route from your first upload to monetization and diversified revenue. The Crew Build Kit and the Studio Operating Kit come as downloads, and the whole thing follows one operator from month zero to month fourteen so you can see the shape of a year rather than guessing at it.

One payment of $99 opens that walkthrough and every other one we have published, plus every one we publish after it, across all three verticals: online businesses, YouTube and content, and offline work in the real world. Downloads, templates and scripts are included in every blueprint, there are no renewals and no upsells, and you can email us while you build. That is launch pricing, and it moves to $199 after the first 500 members. There is a community building alongside you and I mentor through it personally.

If you want the related audits, What $100 a Month of AI Actually Replaces is the honest cost picture, and Build an AI Writer That Sounds Like You is the free build of the first seat, click by click.

The crew is not the interesting part any more. Everyone can have one by Friday. The interesting part is whether the person in the middle is making four real decisions or watching a machine work and calling it directing.