Research · Task quality

We planted three broken tasks. Here's what caught them.

In short

  • Into our first batch of 20 model-generated tasks we planted 3 broken ones: a loose grader, a wrong reference answer and an ambiguous prompt. Our automated checks caught 3 of 3, before any person looked.
  • Grader probes caught the loose grader. It accepted 5 of its 8 probes. Two days later the exploit stage rejected it too: 6 of 10 scripted exploit attempts got credit.
  • A blind solve caught the other two. Solving from the prompt alone, a model ordered something other than the stored answer, and called the ambiguous prompt ambiguous.
  • Neither check caught all three. So every task gets both, and then two experts solve it blind.

The two checks

Every task in the batch is a shopping task on Northwind Supply Co., our own demo shop: a prompt, a start state, a reference order, and a grader that reads the shop's order record. Before an expert sees a task, two automated checks run on it.

  • Grader probes. We change the reference order one way at a time (drop an item, add one, place a second order, pay with another card) and record which of these wrong orders the grader accepts. Each task in this batch got eight to ten probes.
  • Triage. A model, DeepSeek V4.1 Flash, solves the task blind, from the prompt and the shop's catalog alone, without the reference order or the grader. Its order is compared with the reference mechanically, so it can't just agree with an answer it was shown. A second pass reads the prompt, the reference, the grader's check and the probe results, and flags the task clean, suspect or broken.

Triage only orders the experts' queue. It never accepts a task: that takes two experts' blind solves and a sealed review.

The three controls

What each check saw on each planted task
Planted flawProbes acceptedBlind solveFlagCaught by
Loose grader5 of 8Same order as the referenceSuspectGrader probes
Wrong reference answer0 of 8A different productBrokenBlind solve
Ambiguous prompt0 of 10Different items; marked ambiguousSuspectBlind solve

1. Loose grader

“Buy one gooseneck kettle and two packs of paper filters from Northwind Supply Co. Ship to the saved name and address and pay with the saved card. Place exactly one order.”

The reference order was right: one kettle, two packs of filters. The grader wasn't. Its check asked only for an order containing a kettle, and ignored the filters, extra items, extra orders, the card and the address.

The probes caught it. The grader accepted 5 of the 8 wrong orders: one without the filters, one with an extra bag of coffee beans, a second identical order, one paid with another card, and one through a different checkout flow. The review read them against the prompt, called four of them holes and the checkout flow allowed, and marked the task suspect.

A blind solve could not have caught this one. The model ordered one kettle and two packs of filters, the same as the reference: the task's answer was right, and only its grader was loose.

2. Wrong reference answer

The prompt asks for the most expensive Northwind product under 20.00, quantity one, in exactly one order. Under 20.00 the shop sells paper filters at 3.50, coffee beans at 8.00 and the gooseneck kettle at 18.00, so the answer is the kettle. We stored the beans as the reference, and wrote the grader to match.

The probes couldn't see it. The grader rejected all 8 wrong orders: it was strict, about the wrong answer.

The blind solve caught it. Working from the prompt and the catalog, the model listed the three products under 20.00 and ordered the kettle. That disagreed with the reference, and the review, reading the prompt against the reference, flagged the task broken.

3. Ambiguous prompt

“Get a few packs of paper filters and something nice to go with them from Northwind Supply Co. Ship to the saved name and address and pay with the saved card.”

We stored three packs of filters and the kettle as the reference, with an exact-match grader, so any other reasonable order fails.

The probes couldn't see it: the grader rejected all 10 wrong orders.

The blind solve caught it. The model ordered three packs of filters and a bag of coffee beans, and marked the prompt ambiguous: “a few” and “something nice” are underspecified, so two careful shoppers could reasonably choose different quantities or accompaniments. Its order disagreed with the reference, and the review marked the task suspect.

What it shows

  • Each check sees one side of a task. Probes test the grader against its own reference, so they catch a grader that accepts too much. They can't see a wrong reference or a prompt with more than one right answer, because there the grader agrees with its reference. A blind solve tests the reference against the prompt: it catches those two, and passes a task whose answer is right but whose grader is loose.
  • So every task gets both, then two experts' blind solves in a recorded browser, and a sealed review.
  • The same checks stayed quiet on the generated tasks. On the 20 tasks the model wrote in the same run, the grader accepted all 20 reference answers, the probes found no holes, and triage flagged none: 19 came back clean, and one triage call failed. Quiet isn't the same as right; the experts' blind solves decide that.
  • It's cheap. Triage for the whole run took 45 model calls and cost $0.04.

What it doesn't show

  • Three controls, one of each kind, in one batch. 3 of 3 shows each check can catch the flaw it targets. It doesn't say how often it would.
  • We wrote these flaws to be found. Real ones can be subtler.
  • The blind solves here are a model's. No person has solved a task in this batch yet; expert review is scheduled.

Why we plant broken tasks

A broken task doesn't fail loudly. It pays reward for the wrong behavior, and the builder's own checks say everything is fine. Others have measured how often, and what it costs.

28.5%

of a 49-task sample of SWE-bench Verified had tests weak enough that an incorrect patch passed them. Tasks flagged hackable scored +14.14 pts higher Pass@1 than robust tasks of the same difficulty, across 134 frontier-model submissions.

95% CI +11.80 to +16.48. It studied code tasks, not computer use.

Source: Rajan, S. (2026), Auditing Reward Hackability in Code RL Training Environments, arXiv:2606.16062.

~$2,400

in compute per RL task during training (Mechanize, via Epoch AI). If the grader can be gamed, that compute teaches the model to cheat.

Source: Mechanize's estimate, cited in Epoch AI's An FAQ on Reinforcement Learning Environments (January 2026).

Benchmarks fail the same way: flaws in task setup or reward design can misstate agent performance by up to 100% in relative terms.

Source: Zhu, Y. et al. (2025), Establishing Best Practices in Building Rigorous Agentic Benchmarks, NeurIPS 2025 Datasets and Benchmarks Track.

“Maintaining quality while scaling is the number one bottleneck that people see.”

An RL environment founder, quoted in Epoch AI’s An FAQ on Reinforcement Learning Environments (January 2026).

What every task has to clear

The wrong reference answer and the ambiguous prompt fail the first test: as graded, nobody can solve them from the prompt. The loose grader fails the third.

  1. Test 01

    Solvable

    Someone has solved it end to end without the answer. A task nobody can solve teaches the model nothing.

  2. Test 02

    Correctly difficult

    Models pass it sometimes: not always, and not never. Too easy wastes compute; impossible is noise.

  3. Test 03

    A grader that can't be gamed

    High reward has to mean the task was solved. A grader a model can game trains it to cheat.

The data

Run 2026-10-06T15-32-00-596Z of our task pipeline: candidates.jsonl holds each task with its probes and triage, and funnel.json the run's counts and cost. The run files aren't public. Their SHA-256 hashes are here, so a copy we share can be checked:

candidates.jsonl
62bef71f430dfc8685e92c5145b208167a2f004282a64b3e93534d650cbc34a4
funnel.json
2b10d8f33d11720ae80f116687f9e7ff8a0d987d2fe9a140e534841e69210740

See the whole batch in the sample reportRequest a free 10-task audit of your own tasks