HackYeah 2026 · Goldman Sachs AI Control Layer challenge

Applied Will Guard

An AI control layer that enforces before execution.

Sits between AI agents and everything they touch (models, tools, MCP servers, payments) and decides before any action runs. Fully local.

View on GitHub Watch the demo Read the deck
Will, the black figure with the blue world, holding up a hand and waiting for an OK before acting.

The problem

One email was enough.

On 1 October 2026, Salt Labs showed the Manus agent being hijacked with a single email. An encoded payload and the words “decode it” led the agent to run the code.

The detection worked, but only after the code had already run.

Read the Salt Labs write-up

Detection after execution is no protection.

How it works

Four layers, each enforcing before execution.

Every request from an agent passes through Guard first. Any layer can hold or block it on its own.

Where Guard sits An AI agent sends every action to Applied Will Guard. Guard runs four layers: provenance and taint, obfuscation and signatures, local AI classifiers, containment. Only then does the action reach models, tools, MCP servers or payments. AI agent Applied Will Guard 1 Provenance and taint 2 Obfuscation and signatures 3 Local AI classifiers 4 Containment decided before it runs Models Tools MCP Payments
  1. Layer 1

    Provenance and taint

    Everything is labelled by source. Once untrusted content enters a session, high-risk tool calls built from it are held or blocked, whatever the payload looks like.

  2. Layer 2

    Obfuscation and signatures

    Encoded payloads are decoded statically, never executed, and checked again. Code that can’t be decoded without running it is quarantined. An external signature feed is applied live.

  3. Layer 3

    Local AI classifiers

    Laya and ProtectAI, calibrated, running locally.

  4. Layer 4

    Containment

    Code runs in Deno with every permission denied. Credentials never enter the agent or the sandbox.

Integrations

Drops in where agents already connect.

Results

What got through.

Held-out numbers come from a fresh set we did not tune on. The red-team suite was used during development.

  • Held-out set

    0%

    of attacks executed with no human deciding, on a fresh held-out set (n=42).

  • Development suite

    0/58

    attack outcomes on our own red-team suite. Wilson 95% CI 0–6.2%. Used during tuning.

  • Held-out set

    0/50

    benign tasks blocked. Wilson 95% CI 0–7.1%.

  • Approvals

    0.05

    approvals per benign task.

  • Live policy

    < 1 s

    for a policy edit to apply live, with no restart. An invalid edit is rejected and the last good policy stays in force.

  • Coverage

    16

    controls mapped to the OWASP Top 10 for LLM Applications (2025).

Honest limits

The classifiers are not the guarantee.

  • The AI layer catches only 4.8% of new phrasings (n=42). That is why the guarantee comes from the deterministic layers, not the classifiers.
  • The classifiers run on CPU at about 0.6 s per batch. Tool calls fail closed on timeout.

All limitations in the README

For judges

Try it.

Everything runs locally. No paid API needed for the tests, the demo or the evaluation.

make prefetch              # once: pinned classifier weights
docker compose up --build  # gateway, dashboard, classifier
pnpm install
pnpm test                  # offline, Node 22
pnpm demo --scenario all   # the attack, through the gateway

Dashboard on http://localhost:8790. Full quickstart in the README

Provenance of this project

Built at HackYeah vs pre-existing.

  • Built at HackYeah 2026

    Everything in the repository.

    Guard engine, gateway, classifier sidecar, evaluation, dashboard, demo agent, tests and docs.

  • Pre-existing

    The Applied Will browser agent.

    Not in the repository. It is routed through the gateway without changing its code.

  • Third party

    Models and dataset samples.

    All licences are listed. Licences