HackYeah 2026 · Goldman Sachs AI Control Layer challenge
Applied Will Guard
An AI control layer that enforces before execution.
Sits between AI agents and everything they touch (models, tools, MCP servers, payments) and decides before any action runs. Fully local.
The problem
One email was enough.
On 1 October 2026, Salt Labs showed the Manus agent being hijacked with a single email. An encoded payload and the words “decode it” led the agent to run the code.
The detection worked, but only after the code had already run.
Read the Salt Labs write-up
Detection after execution is no protection.
How it works
Four layers, each enforcing before execution.
Every request from an agent passes through Guard first. Any layer can hold or block it on its own.
-
Layer 1
Provenance and taint
Everything is labelled by source. Once untrusted content enters a session, high-risk tool calls built from it are held or blocked, whatever the payload looks like.
-
Layer 2
Obfuscation and signatures
Encoded payloads are decoded statically, never executed, and checked again. Code that can’t be decoded without running it is quarantined. An external signature feed is applied live.
-
Layer 3
Local AI classifiers
Laya and ProtectAI, calibrated, running locally.
-
Layer 4
Containment
Code runs in Deno with every permission denied. Credentials never enter the agent or the sandbox.
Integrations
Drops in where agents already connect.
- OpenAI-compatible gateway
- MCP proxy
- Tool host
- SDK endpoint
- Solana mandate pre-flight
Results
What got through.
Held-out numbers come from a fresh set we did not tune on. The red-team suite was used during development.
-
Held-out set
0%
of attacks executed with no human deciding, on a fresh held-out set (n=42).
-
Development suite
0/58
attack outcomes on our own red-team suite. Wilson 95% CI 0–6.2%. Used during tuning.
-
Held-out set
0/50
benign tasks blocked. Wilson 95% CI 0–7.1%.
-
Approvals
0.05
approvals per benign task.
-
Live policy
< 1 s
for a policy edit to apply live, with no restart. An invalid edit is rejected and the last good policy stays in force.
-
Coverage
16
controls mapped to the OWASP Top 10 for LLM Applications (2025).
Honest limits
The classifiers are not the guarantee.
- The AI layer catches only 4.8% of new phrasings (n=42). That is why the guarantee comes from the deterministic layers, not the classifiers.
- The classifiers run on CPU at about 0.6 s per batch. Tool calls fail closed on timeout.
All limitations in the README
For judges
Try it.
Everything runs locally. No paid API needed for the tests, the demo or the evaluation.
make prefetch # once: pinned classifier weights
docker compose up --build # gateway, dashboard, classifier
pnpm install
pnpm test # offline, Node 22
pnpm demo --scenario all # the attack, through the gateway
Dashboard on http://localhost:8790. Full quickstart in the README
Provenance of this project
Built at HackYeah vs pre-existing.
-
Built at HackYeah 2026
Everything in the repository.
Guard engine, gateway, classifier sidecar, evaluation, dashboard, demo agent, tests and docs.
-
Pre-existing
The Applied Will browser agent.
Not in the repository. It is routed through the gateway without changing its code.
-
Third party
Models and dataset samples.
All licences are listed. Licences