When an AI agent reads a redline as Markdown, it uses 44% fewer tokens than when it reads the same redline as a Word tracked-changes file, and 69% fewer than when it reads it as a PDF. It also finishes 1.8 times faster than Word and 2.7 times faster than PDF.
Those figures come from 100 real redlines, each read by a separate Claude Opus 5 agent in each of the three formats.
Headline results
| Per redline, average | Word | Markdown | |
|---|---|---|---|
| Time | 58.3 s | 39.8 s | 21.7 s |
| Total tokens | 608,406 | 337,907 | 189,975 |
| Model calls | 8.3 | 5.0 | 2.6 |
| Agent steps, median | 9 | 4 | 1 |
- Markdown used 43.8% fewer tokens than Word and 68.8% fewer than PDF.
- Markdown took 45.6% less time than Word and 62.8% less time than PDF.
- Markdown was faster than Word in 99 of 100 redlines, and faster than PDF in all 100.
- Markdown used fewer tokens than Word in 92 of 100 redlines, and fewer than PDF in all 100.
At Claude Opus 5's list price of $5 per million input tokens, before any prompt caching, that is about $0.74 less per Word redline an agent reviews and about $2.09 less per PDF redline.
Why the format matters to an agent
Lawyers read redlines in Word and PDF. Those formats were designed for people and for printers.
An agent asked to review a Word redline has to unzip the file, find the tracked-change elements in its XML, and write a script to pull them out. An agent asked to review a PDF redline has to extract text page by page and work out from color and strikethrough what was inserted and what was deleted.
Every one of those steps is a call to the model, and every call re-sends the conversation so far. The agent spends most of its tokens getting to the changes, not reasoning about them.
So Version Story, a deterministic document comparison engine, returns every redline in a third format alongside Word and PDF: Markdown, with the changes marked inline as insertions, deletions and moves. Lawyers and agents can review the same comparison in parallel, each in the format built for them.
Where the savings come from
The saving comes from steps, not from file size.
The typical Markdown agent read the file once and wrote its summary. The typical Word agent needed four steps. The typical PDF agent needed nine. Measured as content pulled into the agent's context, the Markdown redline was not smaller than the Word one. What it removed was the calls the agent had to make before it could see the changes at all.
That is also why the saving grows with the size of an agent's own instructions. The more context an agent carries into every call, the more each avoided call is worth.
Results by document size
| Markdown size | Redlines | PDF tokens | Word tokens | Markdown tokens |
|---|---|---|---|---|
| Under 25 KB | 34 | 496,710 | 316,743 | 130,388 |
| 25 to 75 KB | 41 | 642,037 | 335,893 | 165,555 |
| Over 75 KB | 25 | 705,156 | 369,994 | 311,061 |
| Markdown size | Redlines | PDF time | Word time | Markdown time |
|---|---|---|---|---|
| Under 25 KB | 34 | 45.6 s | 35.1 s | 16.9 s |
| 25 to 75 KB | 41 | 61.1 s | 40.4 s | 20.5 s |
| Over 75 KB | 25 | 71.0 s | 45.4 s | 30.2 s |
The advantage is largest on short and mid-length documents, where Markdown used less than half the tokens of Word. On the largest quarter of the sample it narrowed to 16% against Word and stayed above 55% against PDF.
Get the Markdown redline
- Over the REST API: request
format=md. See Formats & limits, and Redline format for how changes are marked. - In Claude and ChatGPT: the Version Story plugin gives the assistant this rendering to read when you ask it to explain what changed. See Compare Documents in Claude and Compare Documents in ChatGPT.
Before an agent can read a redline, something has to produce it. See What It Costs Claude to Compare Two Documents Without a Comparison Tool.
Method
- Documents. 100 comparisons drawn at random, with a fixed seed, from 1,173 successful comparisons created through the Version Story API in production. They are contracts of many kinds, including master services agreements, addenda, employment letters and purchase agreements. Markdown sizes ranged from 3 KB to 342 KB, with a median of 38.5 KB.
- Formats. Three files for each redline, all produced by the same Version Story comparison: the Markdown redline, the Word tracked-changes redline, and the PDF redline.
- Agents. 300 agents, one per redline per format, so no agent saw a second format of the same redline. Each ran Claude Opus 5 with ordinary file and shell tools and chose its own approach.
- Task. The same instruction for every agent: summarize the changes shown in this redline, note how extensive they are, keep the summary under 400 words, and use only this file. The prompt was identical across formats except for the file name.
- Measurement. Wall-clock time, total tokens across every model call, model calls and agent steps, all read from each agent's transcript. Total tokens is the sum over every call of input and output tokens, so an extra step adds tokens even when it reads little new content.
- Run date. September 14, 2026.
Limits of this benchmark
- Very large documents. In a separate run on our internal test suite, once the Markdown redline exceeded about 300 KB the Word agents used slightly fewer tokens than the Markdown agents, because a file that size no longer fits in a single read. Only two redlines in the production sample were that large.
- Comments. The Markdown redline does not carry Word comments today. An agent that needs the comment thread should read the Word file.
- Quality was not scored. This benchmark measures cost and speed. We spot-checked the summaries but did not grade them.
- Timings are relative. Agents ran 10 to 16 at a time, and the PDF batch ran separately from the other two. Token counts do not depend on load.
- Document mix. Most of the sample came from two customers, so contract types are not evenly spread. The documents and the customers are not published.
- One model and one agent runtime. Other models and agent frameworks will produce different absolute numbers.
Frequently asked questions
What is the most token-efficient way to give an LLM the changes between two versions of a contract?
Of the three redline formats tested, Markdown with the insertions, deletions and moves marked inline. In this benchmark an agent read a Markdown redline with 44% fewer tokens than a Word tracked-changes file and 69% fewer than a PDF redline.
How many tokens does an AI agent use to read a redline?
In this benchmark, an agent used an average of 190,000 tokens to summarize a redline delivered as Markdown, 338,000 as Word tracked changes, and 608,000 as PDF.
Is Markdown faster for an AI agent than a Word document?
Yes. Agents summarized Markdown redlines in 21.7 seconds on average, against 39.8 seconds for Word and 58.3 seconds for PDF. Markdown was faster than Word in 99 of 100 redlines.
Why is a PDF redline expensive for an agent to read?
A PDF has no structure that says what changed. The agent has to extract text page by page and infer insertions and deletions from formatting. In this benchmark that took a median of nine steps, against one for Markdown.
Does the Markdown redline replace the Word redline?
No. Lawyers review in Word and PDF. The Markdown redline is for the agent working alongside them, and all three come from the same comparison.
