The MarginGuide

How to Train an AI on Books: What Works, and What You May Not Upload

Feeding five books from one field into an AI genuinely changes what comes back, and the reason is not the one I gave on camera. It also runs straight into two problems nobody mentions: the knowledge container is smaller than five books, and by uploading them you personally warrant you had the right to. Here is the version that is both legally clean and technically better, and it is not the version in my video.

I made a video about putting five of the best marketing books into one AI and getting back something that argued with me instead of agreeing with me. The effect is real and I still use it daily.

Writing this article I found three problems with how I explained it. One is a mechanism I described wrongly. One is a contradiction inside the video itself. And one is a legal exposure I waved at in a single sentence when it deserved a section, because the advice I gave, use books you actually own, does not do the work I implied it does.

The good news is that fixing all three produces a better build than the one I demonstrated.

What actually changed, and what I said changed

In the video I say the agent "develops taste," that it "picks sides," that it stops being an assistant.

Nothing inside the model changed. It is the same model on the same day with the same weights. What changed is the material sitting in the context it retrieves from when it answers, and a model given five specific frameworks will apply a framework, while the same model given nothing will produce the average of everything ever written on the subject. That average is what people call the safe answer, and it is safe precisely because it is an average.

This distinction is not pedantry, because it tells you where the leverage is. You are not upgrading an intelligence, you are curating a library, and the quality of what comes back tracks the quality and specificity of what you put in far more than it tracks which model you are running. That part of the video I stand behind completely: the person running an ordinary model against a real library is getting better answers than the person running the best available model against an empty one.

It also means the effect is fragile in a way I did not mention. Retrieval is not guaranteed. Nothing forces the model to reach for the framework on any particular answer, which is why the instruction layer matters as much as the knowledge layer. If you want it to argue, you have to tell it in the standing instructions that it is allowed to disagree with you and required to name the trade-off. Loading the books does not make it brave.

Where the idea is genuinely right

The strongest claim in the video is the one about overlap, and I have not seen it made elsewhere, so let me sharpen it rather than repeat it.

Where the five authors agree, you are looking at the closest thing the field has to settled practice. Not truth, but consensus among people who each independently thought hard about it and had no obligation to concur. That is a meaningfully different signal from one confident author asserting something.

Where they disagree, you have located the judgement call, and that is more valuable still, because the disagreements mark exactly the decisions that cannot be reduced to a rule. Any field's hard part lives in its unresolved arguments. An agent holding all five can tell you that four would do this, one would do the opposite, and here is the situation where the odd one out is right.

A single author cannot show you that. They wrote their book and made their case. This is genuinely something the format does that reading the books one at a time does not, and it is the honest core of the idea.

The contradiction in my own video

Two minutes apart, the video makes two claims that cannot both stand.

First: that what you build is "in one specific and slightly unsettling way, smarter than any of the authors in it," because none of them had to read the other four and yours did.

Then: "you are not getting the author. You do not get their judgment, and you do not get their taste."

Both cannot be the framing. The second one is correct and the first one is a flourish I liked too much. What you have built is better cross-referenced than any single author, which is a real and useful property, and it is not the same as being smarter than them. It has read more of the field and none of the practice. Every author in that stack learned their material by doing the work and getting it wrong for years; your agent has the output of that process with none of the process.

The version that survives is the smaller one I already said in the same video, and I would keep only this: you get their playbook applied to your exact situation at two in the morning for pennies. That is a genuinely valuable thing and it does not need inflating.

The part you can get badly wrong

Now the section the video gave one sentence and should have given five minutes. In it I say: use books you actually own, and never hand the result to anybody else. That is good instinct and it is not sufficient, in three separate ways.

You are the one making the warranty

Anthropic's consumer terms, effective 8 October 2025, are unusually direct about this. Two clauses matter.

"You are responsible for all Inputs you submit to our Services and all Actions."

"By submitting Inputs to our Services, you represent and warrant that you have all rights, licenses, and permissions that are necessary for us to process the Inputs under our Terms."

Read that second one slowly, because it is the whole point. The vendor is not deciding whether you may upload a book. You are, and you are making a formal representation that you may, every time you drag a file into a project. The Acceptable Use Policy separately prohibits using the service "to infringe, misappropriate, or violate intellectual property or other legal rights."

On outputs, the terms assign to you all of the company's right, title and interest in what comes back, and the qualifier they attach to that assignment is the interesting part: it applies to their interest "if any." An assignment of whatever rights the vendor holds is not a warranty that the output infringes nobody else's.

Owning a book and having a file are different things

This is the one I expect most readers to get wrong, because the phrase "books you actually own" sounds like it settles the matter.

For most people, a purchased ebook is a licence to read it in a particular application, delivered as an access-controlled file. Getting that text into an AI project generally requires either a copy sold without access controls, which many publishers do offer and which is worth seeking out deliberately, or getting around the control.

That second path is its own legal question, separate from whether you paid for the book. Section 1201 of the DMCA makes it unlawful to circumvent technological measures that control access to a copyrighted work, and the Copyright Office's own explanation of it does not name ownership of a lawful copy as a defence. Exemptions exist, but only those granted through a triennial rulemaking by the Librarian of Congress, which is a narrow and specific list rather than a general allowance.

So "I bought it" answers a different question than the one being asked. If you want a book in an AI project, the clean route is to buy an edition sold without access controls, or to use the notes method below, which sidesteps the whole issue.

The $1.5 billion reminder

There is a reason this is not theoretical, and the timing makes it hard to ignore.

A federal judge in the Northern District of California granted final approval on 20 July 2026 to a $1.5 billion settlement between Anthropic and authors, over pirated books used to train Claude. Reported terms work out at roughly $3,000 per book, with around 350 authors opting out and payments expected in August 2026. It is the largest copyright settlement of its kind.

The distinction that mattered in that case is precisely the one that matters to you. The dispute was not really about whether a machine may learn from a book. It was about where the copies came from. Provenance was the issue. If you are building a personal library for an agent, provenance is the thing to be able to account for.

What is genuinely free

The video says some of the best books are old enough to be free, and that is true and more specific than I made it sound.

On 1 January 2026, works published in 1930 entered the US public domain. The rule advances a year at a time, so the line to remember in 2026 is 1930 and earlier. Sound recordings run on a different clock, with 1925 recordings entering at the same moment, which is irrelevant here but worth knowing before you assume one rule covers everything.

For a marketing library this is not a consolation prize. Claude Hopkins published Scientific Advertising in 1923, and it is both firmly inside that window and one of the most procedural books ever written on the subject, which as you will see below is exactly the property that makes a book work as an agent.

Two cautions. Public domain status is jurisdictional, and many countries run on life of the author plus seventy years rather than a publication-year rule, so a work that is free in one place may not be in another. And a modern edition can carry its own new copyright in its introduction, annotations and typesetting even when the underlying text is free. Take the text, not the edition.

Here is where the two problems in this article collapse into one solution, and it is the thing I would actually tell you to do.

You cannot fit five books into the knowledge container anyway. Anthropic's documentation states that enhanced project knowledge is "only available to users with paid Claude plans" and, when available, expands capacity "by up to 10x." Five full books is an enormous amount of text against any of those limits, and on a free account you are working inside the smaller container.

Our own Writer build says the same thing from the other direction, and I said it on camera in that video: do not dump one giant file in, keep each one tight, roughly one to three pages, because it pulls a much stronger signal out of a few sharp pages than out of one big messy document. Two of my own videos give opposite advice about the same box, and the Writer video has it right.

So the build is this.

  1. Read the book, or read your own highlights from when you read it. This step does not delegate. You are the one who knows which parts apply to your work.
  2. Write one page per book, in your own words. The named moves, the conditions under which each applies, and the author's stated reasoning. Procedures, not prose.
  3. Write a sixth page, and make it the most important one. Where the five agree, and where they conflict, with the situations that decide which side wins. This page is the actual asset and it does not exist in any of the books.
  4. Load those six pages as the project knowledge. They will fit anywhere, on any plan.
  5. Put the argumentative behaviour in the instructions, not the knowledge. Tell it to name the weaker option and the trade-off you are avoiding, and to cite which framework it is applying.

That version fits the container, retrieves more reliably because the signal is dense rather than buried, contains your synthesis rather than anybody's text, and leaves you with nothing you would be uncomfortable explaining. It is better on every axis than uploading the books, which is a happy outcome and not one I expected when I started checking.

And the honest cost: it takes a few hours instead of a few minutes. That is the entire trade.

The filter, which is the best rule in the video

One rule from the video survives completely intact and I would put it above everything else here.

If you cannot name three moves from that book that a person could do tomorrow morning, it does not become an agent.

A book that hands you moves you can run becomes a genuinely useful specialist. A three-hundred-page memoir becomes four bullet points and a lot of atmosphere, because there was never a procedure in it to extract. This is why the field matters less than the shape of the writing: negotiation, copywriting, direct response and hiring are full of procedural books, while a great deal of business writing is narrative and produces nothing when compressed.

Run that filter before you spend the hours. It will disqualify about half the books you were excited about, which is the filter working.

The best version of this idea faces inward

The strongest application in the video is the one that gets thirty seconds, and it has no copyright dimension at all.

Every business is sitting on a pile of its own writing. The sales playbook. The onboarding documents. The proposals that actually won, and ideally the ones that lost. Support transcripts. The internal note explaining why you stopped doing something.

Feed that in and you do not get a generic assistant with a reading list. You get something that reasons the way your company specifically reasons, on material nobody else has, that you own outright, that no licence question touches, and that improves every quarter as you add to it.

It is also the version with a business behind it, which is worth saying plainly. A generic well-read agent is a nice personal tool. An agent built from a specific company's own winning documents is worth paying for, and building those for other people is a service business rather than a party trick.

Come and build the crew around it

If you want to build your first specialist properly, start free. Our Build Your AI Team course walks the Writer build end to end on a free account, and the prompt pack is yours to keep with no account needed. It teaches the instruction layer this whole article depends on, which is the half that makes an agent argue rather than agree. The Claude prompting course is also free and sharpens the skill underneath it.

When you want the whole crew rather than one well-read seat, that is at ideasrepay.com, and the walkthrough is The One-Person Studio.

It is the build rather than an article about the build. Seventeen steps across five phases, click by click. You get the role prompts to paste straight in, so the crew exists in an afternoon instead of a month of trial and error: a Researcher, a Scriptwriter, a Voice, an Editor, a Designer and a Supervisor whose whole job is catching what is wrong before it ships. You get the three-way niche fit test that sets your earnings ceiling before you make anything, the full production run from a blank idea to a published video through the four quality gates you never hand over, the packaging and quality control checklists, the disclosure rules in plain language, and the route from your first upload to monetization and diversified revenue. The Crew Build Kit and the Studio Operating Kit come as downloads, and the whole thing follows one operator from month zero to month fourteen so you can see the shape of a year rather than guessing at it.

One payment of $99 opens that walkthrough and every other one we have published, plus every one we publish after it, across all three verticals: online businesses, YouTube and content, and offline work in the real world. Downloads, templates and scripts are included in every blueprint, there are no renewals and no upsells, and you can email us while you build. That is launch pricing, and it moves to $199 after the first 500 members. There is a community building alongside you and I mentor through it personally.

If you want the related builds, Build an AI Writer That Sounds Like You is the free first seat click by click, and Where the Human Belongs in an AI Workflow is the evidence on what the person in the middle is actually for.

The reading is the moat. It just turns out the reading has to be yours.