Research · Benchmark
ConcreteTakeoff-Bench v0.2: text-only models measure L-shaped slabs as rectangles.
Synthetic sheets
Every sheet in this benchmark is synthetic: a generated vector PDF with known quantities. They are clean and CAD-like, with a text layer, every quantity written or dimensioned, and no clutter, revisions or scanning. They are not real drawings, and nothing here says how any model does on real foundation plans.
In short
- The task: read a foundation plan and report slab area, footing length, footing counts and concrete volumes, or send the sheet to a human when its information contradicts itself.
- 40 synthetic tasks, 34 to quantify and 6 planted controls, three runs each.
- Reading the PDFs as text, Claude Haiku 5.5 passed 38% of the quantity sheets (mean of three runs, sd 8 points). Given the sheet images too, it passed every one, in every run.
- Its largest single text-only error is geometry: 28 of its 63 failed trials read the slab as a rectangle.
Text only, and with the sheet images
Claude Haiku 5.5: quantity sheets passed
Mean of three runs, by tier. 34 synthetic sheets.
| Tier | Text only | With the sheet images |
|---|---|---|
| All tiers | 38.2% | 100% |
| Original | 62.7% | 100% |
| Harder | 16.7% | 100% |
| Hardest | 9.5% | 100% |
Source: our benchmark write-up, docs/CONCRETETAKEOFF-BENCH.md, v0.2 (run on 8 October 2026).
The gap is +61.8 points, 95% interval +47.1 to +75.5 (tasks resampled), and harder sheets widen it: text-only fell from 62.7% on the original tier to 16.7% on the harder tier and 9.5% on the hardest. With the images, Haiku made no errors and no false alarms, in any run.
We haven't yet built a synthetic sheet that image input gets wrong, so this tier can't say where image input stops helping.
Example: ctb-syn-0018
An L-shaped slab from the original tier. The outline is 90 × 60 ft with a 30 × 20 ft notch, and every edge is dimensioned. The callouts give a 5" slab, and the continuous footing as 20" wide by 12" deep.
- The slab: 4,800 SF
- What text-only Haiku measured, in all three runs: the 90 × 60 ft rectangle, 5,400 SF
- The notch it counted: 600 SF, 12.5% too much slab
| Answer | Slab SF | Footing LF | Slab CY | Total CY | Result |
|---|---|---|---|---|---|
| Expected | 4,800 | 300 | 74.07 | 97.46 | |
| Haiku 5.5, text only (all three runs) | 5,400 | 300 | 83.33 | 106.72 | Fail |
| Haiku 5.5, with images (all three runs) | 4,800 | 300 | 74.07 | 97.46 | Pass |
| DeepSeek V4.1 Flash, text only (all three runs) | 4,800 | 300 | 74.07 | 97.46 | Pass |
The text-only answer's footing length is right, because an L's perimeter equals its bounding rectangle's, and so are its footing counts and footing volumes. Only the area misses the notch, and the slab and total volumes with it.
DeepSeek V4.1 Flash, an open-weights model also reading the text, got this sheet right in all three runs. It isn't immune: it read 7 other slabs as rectangles.
How text-only fails
Every failed trial, all runs pooled, under the first cause that applies:
| Cause | Haiku 5.5 | DeepSeek V4.1 Flash |
|---|---|---|
| Slab read as its bounding rectangle | 24 | 6 |
| Slab opening not deducted | 4 | 1 |
| A consistent sheet sent to a human (false alarm) | 26 | 9 |
| Other slab, length or volume errors | 9 | 0 |
| Failed trials | 63 | 16 |
- Rectangles. On the non-rectangular sheets, text-only Haiku got the slab area right in 10 of 63 attempts. With the images, in 63 of 63.
- False alarms come from the same root. Text-only Haiku took a dimension line, drawn 0.6" outside the slab, for the slab edge; added up dimension strings from different chains; and trusted lines it measured over the dimension strings that disagreed with them. Each time, lines and labels were matched to the wrong geometry.
- So is it rectangles? Partly. For Haiku it's the largest single error, 28 of 63 failures. For DeepSeek, false alarms outnumber it, 9 to 7 of 16. With the images, Haiku made neither.
- The planted controls didn't separate the systems. Every system sent every planted control to a human with a reason, in every run it ran.
All results
| System | All | Original | Harder | Hardest | Cost per task |
|---|---|---|---|---|---|
| Claude Haiku 5.5, text only | 38.2% | 62.7% | 16.7% | 9.5% | $0.0080 |
| Claude Haiku 5.5, with images | 100% | 100% | 100% | 100% | $0.0100 |
| DeepSeek V4.1 Flash, text only | 80.9%* | 90.2% | 93.3% | 50.0%* | $0.0610 |
* Two complete runs: DeepSeek's third run of the hardest tier was stopped to stay under the spending cap. Tiers: original, 17 sheets to quantify (rectangles and L-shapes); harder, 10 (L, T and U slabs, openings, thickened edges, a two-sheet set); hardest, 7 (three-sheet sets, plans not to scale, an opening dimensioned only in an enlarged plan).
Method, briefly
- Harness: Harbor's Terminus-2 agent. Text only, the model works in a shell with
pdftotext,pdfplumberand Python, for at most 40 turns. With images, the same agent's first prompt also carries each sheet as an overview and six overlapping tiles. - Scoring: slab area and footing length within 2%, every concrete volume within 3%, footing counts exact, and the route (quantify, or send to a human) right. A pass is every check right.
- Leakage check, before the headline: no sheet states an area, a volume, a total or any answer value. Every quantity has to be computed.
What it doesn't show
- Synthetic sheets only. Real plan sets — scanned, revised and multi-sheet — are next.
- 40 tasks, few per kind. Each harder or hardest kind appears once or twice, so a rate for one kind rests on 3 to 6 attempts.
- Three runs. The spread is a rough estimate, and the interval on the gap resamples tasks, not runs.
- One harness family. The images go in once, in the first prompt, as tiles we chose. A harness that lets the model zoom, or gives it a takeoff tool, could score differently.
- The traps target measuring. A plan not to scale catches a solver that measures the drawing. One that reads the dimension strings isn't caught, which is what Haiku with images did.
- Two models, chosen for price and availability, not as a survey.
The data
The write-up, docs/CONCRETETAKEOFF-BENCH.md, and the per-trial data with every answer, review reason, failed check and cost, packages/takeoff-bench/baselines/synthetic-v0.2.json, are in our repository with the sheet generator, the grader and the commands to reproduce the run. The task export's manifest SHA-256 is bd6ad45aa5477b45427de820afc114cd79389a874f84742edf182386465d5cc6.
Talk to us about a benchmark collaborationRequest construction takeoff environments