be/brief
Request access
← The Debrief
The craft

How to brief an AI for first-draft research.

The deep-research agent hands back twelve pages. Clean prose, a dozen citations that all resolve when you click them. It looks like a week of someone’s work. And you still cannot forward it to the person who asked for it, because you have no real idea which parts are true.

That gap has a number now. Across fourteen frontier models, deep-research agents produce reports whose links work more than 94% of the time and read as on-topic more than 80% of the time, yet only 39 to 77 percent of their claims are actually supported by the source they cite (Onweller et al., 2026). The links are real. The pages are real. The sentence the report built on top of them is often not in there. And it gets worse the harder the agent works: factual support fell about 42% as it ran more searches, from a couple of lookups up to a hundred and fifty.

So here is the opinion worth arguing with. A good AI research brief is written to make the assistant easy to fact-check, not to make it autonomous. The point of the brief is to turn checking the sources into a ten-minute job instead of a two-hour one. You can hand off the reading and the drafting. You cannot hand off deciding what is true, and a brief that pretends otherwise has delegated the appearance of research and kept none of the safety.

Here is the shape of a brief that gets that right.

The outcome

Done is a first-draft research memo you could stand behind by nine tomorrow morning: a defensible synthesis where every load-bearing claim traces to a source you can open in one click, and where the assistant told you plainly which questions it could not answer. What you are describing is a memo you would forward the moment you finished reading it, its every claim already traceable to something you can open.

What you hand off

The gathering and the grind. Finding candidate sources, reading at a volume you never would, pulling the relevant passages, drafting a first synthesis with real structure. The machine is genuinely good at this, and none of it is where the danger lives.

The danger lives in the two jobs you keep: deciding which claims are load-bearing, and making the final call on what is true. So say where the work starts and where it stops. It starts with a sharp question and a scope boundary. It stops at a draft with its receipts attached, handed to you for a final audit.

How to brief it

This is the whole game, and it is written in plain language, the way you would ask a capable colleague, not in parameters.

Give it three things. The actual question, plus the decision it feeds, so the word “relevant” means something: “I need to know whether we build on top of X now or wait, so I care about maturity and breaking changes, not launch hype.” The source boundary: “Use peer-reviewed work, standards, and primary documentation. If the only thing you can find for a claim is a vendor blog, say so rather than dressing it up.” And the output contract, which is the single most important line in the whole brief:

Every claim links to its source and quotes the sentence that backs it. Anything you cannot verify gets flagged, not filled.

Cite or flag. That instruction is what converts a two-hour audit into a ten-minute one, because it forces the assistant to show its work exactly where the work is weakest.

What to check

Not everything. The brief bought you the right to check narrowly. Start with the claims your decision actually rests on: click the links, read the quoted line, confirm it says what the draft says it says. This is the 39-to-77 gap made personal, and it is where fabrication survives even in the strong models. In one peer-reviewed audit, GPT-4 still invented about 18% of its citations and got a substantive detail wrong in a quarter of the ones that were real (Walters and Wilder, 2023).

Then read the flags it raised, which are the most honest thing in the document. Finally, check that it did not grow more sure of itself the deeper it dug, because that is the exact direction the evidence says confidence and accuracy part ways.

What stays yours

The judgment. Which claims matter, and the verdict on each one: true, or not yet. An assistant can read a thousand pages and still not tell you which single fact is worth betting the decision on. That was never yours to give away. A brief that pretends otherwise does not save you the audit. It buries it, and it resurfaces at the worst possible moment, in front of the person you forwarded the memo to.

So you stay the director of the research rather than its operator. In practice that means the last thing you read is not the whole report. It is the four or five claims the decision turns on, each with its source open in the next tab, waiting for one word from you.

Brief is opening to a small group at a time. Direct a team instead of operating one more tool.

Request access