Skip to content
Benchmark

What It Costs Claude to Compare Two Documents Without a Comparison Tool

We asked Claude to compare 25 real contract revisions and produce a tracked-changes redline for each, with no comparison tool available. It averaged 5.7 minutes and 858,000 tokens per comparison. The Version Story API produced a redline for each of the same 25 revisions in a median of 2.8 seconds, with no model tokens.

When Claude is asked to compare two versions of a contract and has no comparison tool, it writes one. Across 25 real contract revisions, that took an average of 5.7 minutes, 858,000 tokens and 50 model calls per comparison, at a cost of $1.37 each on Claude Opus 5.

Version Story, a deterministic document comparison engine, produced a redline for the same 25 pairs through its API in a median of 2.8 seconds, using no model tokens.

Headline results

Per comparisonClaude, no comparison toolVersion Story API
Time, average5.7 minutes5.3 seconds
Time, median4.9 minutes2.8 seconds
Time, slowest14.3 minutes17.5 seconds
Model tokens, average858,0890
Model calls, average500
Model cost, average$1.37$0

Across all 25 comparisons:

  • The agent spent 143 minutes of model time. The API spent 132 seconds. That is 65 times faster in total, and between 17 and 335 times faster on individual pairs.
  • The agent consumed 21.5 million tokens, costing $34.31 with prompt caching. Without caching, the same tokens would have cost about $121, or $4.85 per comparison.
  • The cheapest agent run still took 2.4 minutes and 255,000 tokens. The most expensive took 14.3 minutes and 3.35 million tokens.

Results by document length

Pages are counted at 350 words each.

Document lengthPairsAgent timeAgent tokensAgent costAPI time
Up to 2 pages33.8 min401,856$0.861.2 s
3 to 7 pages44.6 min582,792$1.051.3 s
8 to 17 pages45.6 min838,613$1.352.0 s
18 to 37 pages45.5 min916,934$1.373.4 s
38 to 74 pages45.6 min789,768$1.325.1 s
75 to 149 pages38.9 min1,639,069$2.2311.5 s
150 pages and over36.5 min939,009$1.5615.6 s

Each figure is the average for that band.

What the agent actually does

In every run the agent followed the same path. It looked for a library or an office program that could compare documents, found none, and then wrote its own comparison program: several hundred lines that unzip the Word files, align paragraphs and words between the two versions, and write the differences back as tracked changes. It then wrote further scripts to check its own output, and revised the program until those checks passed.

That explains three things in the data.

Length barely predicts cost. A 306-page credit agreement cost $1.25, about the same as a 4-page document. The agent does not read the documents end to end. Its cost comes from writing and debugging a program, and a short document needs the same program as a long one.

The floor is high. No run finished in under 2.4 minutes or under 250,000 tokens, including on documents of one or two pages.

The cost is unpredictable. Two credit agreements of almost identical length, 36 and 37 pages, cost $2.30 and $0.95. One run on a 103-page contract took 90 model calls and 3.35 million tokens. How long the agent spends depends on how many problems its own program runs into, which nobody can know in advance.

A deterministic comparison engine has none of these properties. Its time grows with document length and is otherwise the same every time.

Why the tokens add up

An agent works in steps. Each step is a call to the model, and each call re-sends the conversation so far: the instructions, every command the agent has run, and every result it has seen. Fifty steps means the growing transcript is sent fifty times.

Prompt caching reduces the price of the repeated portion to a tenth, and all figures here include it. It does not reduce the time, and it does not reduce the token count that rate limits are measured against.

What to use instead

Give the agent a comparison tool. It hands over the two versions, gets the redline back, and writes no comparison code.

Version Story also returns each redline as Markdown, the cheapest format for an agent to read. See Markdown, Word or PDF: What a Redline Costs an AI Agent to Read.

Method

  • Documents. 25 pairs of consecutive versions of real contracts, drawn from a corpus of 100 production comparisons from ten law firms and legal teams in the United States and the United Kingdom. The pairs were chosen to spread evenly across length, from 1 to 306 pages, and to mix lightly and heavily edited revisions. They include M&A agreements, credit and loan documents, leases, fund documents, license agreements and corporate documents. 15 of the 25 contain tables, 10 contain footnotes, and 7 arrive with tracked changes or comments already in the earlier version.
  • Agent. Claude Opus 5, run through Claude Code with shell access and file read, write and edit tools. One fresh agent per pair. No limit on steps or time. Prompt caching on.
  • Task. Each agent received the two versions and this instruction: produce a redline of the later version against the earlier one as a Word document with native tracked changes that a lawyer can accept or reject in Microsoft Word, using only these two files, working efficiently.
  • Environment. Each agent ran in a sealed container holding the two documents and a standard Python installation, with no document libraries, no office software and no network access except to the model. This is the setting the benchmark is about: an agent with no comparison tool. Every run attempted between two and four times to install a library or find an office program before writing its own code. Each attempt was blocked and logged.
  • API. The same 25 pairs were submitted to the Version Story comparison API in production, one at a time. API time is measured from submission until the Markdown redline was available to download. The Word and PDF redlines were ready at an average of 6.5 seconds.
  • Measurement. Agent time is model time as reported by the agent runtime. Tokens are the sum over every model call of input, cached input and output tokens. Cost is at Claude Opus 5 list prices: $5 per million input tokens, $25 per million output tokens, with cached input at a tenth of the input price.
  • Run date. October 1, 2026.

Per-comparison results

DocumentPagesAgent minAgent tokensCallsAgent costAPI seconds
Other14.4485,36440$0.970.8
Corporate24.7465,48244$1.021.8
License22.4254,72132$0.581.1
Corporate32.4353,73038$0.611.2
Credit34.6606,04147$1.071.0
M&A45.2557,82541$1.131.3
Lease76.2813,57344$1.401.6
License84.7727,58342$1.192.1
Credit85.7743,08653$1.282.0
Fund97.21,348,51966$1.861.4
M&A154.7535,26347$1.062.3
Fund183.7499,27741$0.872.8
Other257.3939,07358$1.644.4
M&A292.8334,79635$0.682.1
Credit368.21,894,59180$2.304.4
Credit373.9510,49541$0.953.5
Lease406.2514,86338$1.214.3
M&A594.7653,64143$1.134.7
Fund737.71,480,07166$1.997.9
Lease783.5314,85436$0.725.6
Other10314.33,350,73790$3.9311.5
M&A1449.01,251,61767$2.0417.4
Other1599.01,357,26161$2.1617.5
M&A1644.9752,19245$1.2717.5
Credit3065.6707,57546$1.2511.8

Document is the type of contract, inferred from its file name: M&A agreements, credit and loan documents, leases, fund documents, license agreements, corporate documents, and other contracts. Calls are model calls. The documents themselves, and the firms they came from, are not published.

Limits of this benchmark

  • This measures cost and speed, not accuracy. We did not grade the agents' redlines here. Whether an agent-built redline is complete and correct is a separate question, and a separate benchmark.
  • One model and one agent runtime. Other models and other agent frameworks will produce different absolute numbers. An agent with a larger system prompt will use more tokens per step than this one did.
  • One run per pair. Agent runs vary, so a rerun of any single pair could land higher or lower. The averages across 25 pairs are the stable figures.
  • A sealed environment. An agent that is allowed to install a document library would do less work from scratch. It would still be writing and debugging comparison code on every request.
  • API time includes upload. It was measured from a single client submitting one comparison at a time, and includes about half a second of upload per pair.
  • 25 pairs. The sample was chosen to cover a range of lengths and document types, not to represent any one firm's workload.

Frequently asked questions

Can Claude compare two Word documents?
Yes, but without a comparison tool it does so by writing its own comparison program. In this benchmark that took Claude an average of 5.7 minutes and 858,000 tokens per comparison across 25 real contract revisions. The benchmark measured time and cost, not the accuracy of the result.

How long does it take an AI agent to compare two contracts?
In this benchmark, Claude with no comparison tool took an average of 5.7 minutes to compare two versions and produce a tracked-changes redline, with a range of 2.4 to 14.3 minutes across 25 real contract revisions.

How many tokens does an AI agent use to compare two documents?
An average of 858,000 tokens and 50 model calls per comparison, ranging from 255,000 to 3.35 million tokens.

How much does it cost for an AI agent to compare two documents?
An average of $1.37 per comparison at Claude Opus 5 list prices with prompt caching, and about $4.85 without caching.

Is a longer contract more expensive for an agent to compare?
Only slightly. The agent's cost comes from writing and debugging a comparison program, so a 306-page agreement cost about the same as a 4-page one. The cost varied more between documents of the same length than between short and long documents.

What tool should I give an AI agent for document comparison?
A deterministic comparison engine the agent can call, so that it does not write comparison code on every request. Version Story is one: it is available as a REST API, as an MCP server, and as a plugin for Claude and ChatGPT. On the same 25 pairs it produced a redline in a median of 2.8 seconds and used no model tokens.

How fast is the Version Story API by comparison?
A median of 2.8 seconds per comparison and an average of 5.3 seconds on the same 25 pairs, with the longest taking 17.5 seconds. It uses no model tokens.

Two lawyers look out over the city from a corner office

See it in action

Enter the epoch of AI with redlining that scales