Research · Benchmark

ConcreteTakeoff-Bench v0.2: text-only models measure L-shaped slabs as rectangles.

Synthetic sheets

Every sheet in this benchmark is synthetic: a generated vector PDF with known quantities. They are clean and CAD-like, with a text layer, every quantity written or dimensioned, and no clutter, revisions or scanning. They are not real drawings, and nothing here says how any model does on real foundation plans.

In short

  • The task: read a foundation plan and report slab area, footing length, footing counts and concrete volumes, or send the sheet to a human when its information contradicts itself.
  • 40 synthetic tasks, 34 to quantify and 6 planted controls, three runs each.
  • Reading the PDFs as text, Claude Haiku 5.5 passed 38% of the quantity sheets (mean of three runs, sd 8 points). Given the sheet images too, it passed every one, in every run.
  • Its largest single text-only error is geometry: 28 of its 63 failed trials read the slab as a rectangle.

Text only, and with the sheet images

Claude Haiku 5.5: quantity sheets passed

Mean of three runs, by tier. 34 synthetic sheets.

Claude Haiku 5.5's pass rate on the quantity sheets, mean of three runs, reading the PDFs as text and with the sheet images
TierText onlyWith the sheet images
All tiers 38.2% 100%
Original 62.7% 100%
Harder 16.7% 100%
Hardest 9.5% 100%

Source: our benchmark write-up, docs/CONCRETETAKEOFF-BENCH.md, v0.2 (run on 8 October 2026).

The gap is +61.8 points, 95% interval +47.1 to +75.5 (tasks resampled), and harder sheets widen it: text-only fell from 62.7% on the original tier to 16.7% on the harder tier and 9.5% on the hardest. With the images, Haiku made no errors and no false alarms, in any run.

We haven't yet built a synthetic sheet that image input gets wrong, so this tier can't say where image input stops helping.

Example: ctb-syn-0018

An L-shaped slab from the original tier. The outline is 90 × 60 ft with a 30 × 20 ft notch, and every edge is dimensioned. The callouts give a 5" slab, and the continuous footing as 20" wide by 12" deep.

ctb-syn-0018: an L-shaped slab, 90 by 60 feet less a 30 by 20 foot notch at the top right. The dashed 90 by 60 rectangle is what text-only Haiku measured; the hatched notch is the 600 square feet it counted that isn't there. 60'-0" 30'-0" 20'-0" 40'-0" 90'-0" 60'-0" Slab 4,800 SF +600 SF
  • The slab: 4,800 SF
  • What text-only Haiku measured, in all three runs: the 90 × 60 ft rectangle, 5,400 SF
  • The notch it counted: 600 SF, 12.5% too much slab
ctb-syn-0018: expected and submitted quantities
AnswerSlab SFFooting LFSlab CYTotal CYResult
Expected4,80030074.0797.46
Haiku 5.5, text only (all three runs)5,40030083.33106.72Fail
Haiku 5.5, with images (all three runs)4,80030074.0797.46Pass
DeepSeek V4.1 Flash, text only (all three runs)4,80030074.0797.46Pass

The text-only answer's footing length is right, because an L's perimeter equals its bounding rectangle's, and so are its footing counts and footing volumes. Only the area misses the notch, and the slab and total volumes with it.

DeepSeek V4.1 Flash, an open-weights model also reading the text, got this sheet right in all three runs. It isn't immune: it read 7 other slabs as rectangles.

How text-only fails

Every failed trial, all runs pooled, under the first cause that applies:

Why text-only trials failed
CauseHaiku 5.5DeepSeek V4.1 Flash
Slab read as its bounding rectangle246
Slab opening not deducted41
A consistent sheet sent to a human (false alarm)269
Other slab, length or volume errors90
Failed trials6316
  • Rectangles. On the non-rectangular sheets, text-only Haiku got the slab area right in 10 of 63 attempts. With the images, in 63 of 63.
  • False alarms come from the same root. Text-only Haiku took a dimension line, drawn 0.6" outside the slab, for the slab edge; added up dimension strings from different chains; and trusted lines it measured over the dimension strings that disagreed with them. Each time, lines and labels were matched to the wrong geometry.
  • So is it rectangles? Partly. For Haiku it's the largest single error, 28 of 63 failures. For DeepSeek, false alarms outnumber it, 9 to 7 of 16. With the images, Haiku made neither.
  • The planted controls didn't separate the systems. Every system sent every planted control to a human with a reason, in every run it ran.

All results

Quantity sheets passed, mean of three runs
SystemAllOriginalHarderHardestCost per task
Claude Haiku 5.5, text only38.2%62.7%16.7%9.5%$0.0080
Claude Haiku 5.5, with images100%100%100%100%$0.0100
DeepSeek V4.1 Flash, text only80.9%*90.2%93.3%50.0%*$0.0610

* Two complete runs: DeepSeek's third run of the hardest tier was stopped to stay under the spending cap. Tiers: original, 17 sheets to quantify (rectangles and L-shapes); harder, 10 (L, T and U slabs, openings, thickened edges, a two-sheet set); hardest, 7 (three-sheet sets, plans not to scale, an opening dimensioned only in an enlarged plan).

Method, briefly

  • Harness: Harbor's Terminus-2 agent. Text only, the model works in a shell with pdftotext, pdfplumber and Python, for at most 40 turns. With images, the same agent's first prompt also carries each sheet as an overview and six overlapping tiles.
  • Scoring: slab area and footing length within 2%, every concrete volume within 3%, footing counts exact, and the route (quantify, or send to a human) right. A pass is every check right.
  • Leakage check, before the headline: no sheet states an area, a volume, a total or any answer value. Every quantity has to be computed.

What it doesn't show

  • Synthetic sheets only. Real plan sets — scanned, revised and multi-sheet — are next.
  • 40 tasks, few per kind. Each harder or hardest kind appears once or twice, so a rate for one kind rests on 3 to 6 attempts.
  • Three runs. The spread is a rough estimate, and the interval on the gap resamples tasks, not runs.
  • One harness family. The images go in once, in the first prompt, as tiles we chose. A harness that lets the model zoom, or gives it a takeoff tool, could score differently.
  • The traps target measuring. A plan not to scale catches a solver that measures the drawing. One that reads the dimension strings isn't caught, which is what Haiku with images did.
  • Two models, chosen for price and availability, not as a survey.

The data

The write-up, docs/CONCRETETAKEOFF-BENCH.md, and the per-trial data with every answer, review reason, failed check and cost, packages/takeoff-bench/baselines/synthetic-v0.2.json, are in our repository with the sheet generator, the grader and the commands to reproduce the run. The task export's manifest SHA-256 is bd6ad45aa5477b45427de820afc114cd79389a874f84742edf182386465d5cc6.

Talk to us about a benchmark collaborationRequest construction takeoff environments