Research
What our runs show.
Results from our own task batches and benchmarks, with the method and the limits. Every number comes with its source.
Posts
-
We planted three broken tasks. Here's what caught them.
Grader probes caught the loose grader. Blind solves caught the wrong reference answer and the ambiguous prompt.
-
ConcreteTakeoff-Bench v0.2: text-only models measure L-shaped slabs as rectangles.
On synthetic sheets, Claude Haiku 5.5 passed 38% reading the text, and every sheet with the images.
-
Sample report: our first batch, task by task.
Pass rates, exploit attempts and the planted controls, from our own demo shop.
Earlier hackathon projects
Two projects that came before the task sets. They are still online.
-
Applied Will Guard
An AI control layer that sits between agents and everything they touch, and decides before any action runs.
Read about Guard -
Agent Mandates
A way for an AI agent to pay on your behalf within limits you set, without trusting the agent.
Read about Agent Mandates