Your agents learn to write the programs that decide when work is done
verify → review → docs hooks that fail closed your own make verbs lessons that become tests about 0.1 s of overhead Claude Code, Codex & more
A prompt can ask a coding agent to run the tests, mock nothing and clean up its debug prints, but it cannot make the agent do it. defuss-vae moves the checks out of the prompt and into code: hooks deny git commit and send the agent back to work when it stops early, until the current code passes verify → review → docs. With wrap, the agent codifies each lesson into the verifier, as a test or a rule that every later gate runs; only a lesson no program can check stays a memory line.
Method, plugin and paper by Aron Homberg (kyr0) (opens in a new tab) – Open Source (MIT (opens in a new tab))
A gate the agent can't talk past
Hooks deny git commit until the exact code fingerprint is verified, reviewed and documented. A prompt can be ignored; a denied commit cannot. Any edit changes the fingerprint and restarts the gate.
Fails closed
A crashing gate denies the commit and blocks the stop once, instead of waving changes through. Hooks run on plain python3 3.9 or newer, stdlib only: no dependencies, no daemon, no uv in the hook path.
Your checks, your toolchain
The gate runs your Makefile verbs (lint, test, coverage, e2e) and the rules in .agents/VERIFY.py. Defaults for Go, Rust, JVM, .NET, JS/TS and Python fill the gaps, and an existing toolchain stays.
Lessons become verifier code
wrap codifies each lesson into a test or a VERIFY.py rule, so every later gate checks for it, whatever the agent remembers. Only a lesson no program can check stays one memory line, within 4 KiB.
Docs are gated too
Every changed page passes a static prose check and a review, window by window, against a 58-rule catalog: evidence, logic, terms, relevance, structure, language and typography.
You stay in charge
The agent starts six skills when their step comes. A risky task without your plan gets one for your review first; only you start wrap, which commits, and nothing is pushed or released without you.
Plan agentic
Agent swarm implements
Verifier binds agent to quality
Tiny review for humans
Agentic self-learning, semantic commits, changelog updates
Human in control for every release
🔁 Two loops: one fixes the work, one codifies the lessons
Programs check the work and tell the agent what remains; the agent fixes it and runs the gate again. The second loop codifies what the work taught into the verifier itself, so every later gate checks it; memory holds only what no program can check. Figure 1 of the paper (opens in a new tab) describes the general mechanism; this is how defuss-vae 0.8.0 implements it.
- your make verbs
- goal
- load
- remaining
- next call
- green
- outcomes
- leads
- codify
- every gate runs it
- else one line
- next session
Scroll the diagram sideways to see the whole loop.
What the agent reads
The gate's output is plain text the harness hands to the model's next call: which checks passed, which remain, and the command that fixes each one. Checking, selecting the next window of a page and formatting the output make no model call.
The output lines are captured verbatim from defuss-vae 0.8.0 on a small fixture, as printed in the paper's Figure 2: a temporary probe marker on line 2 fails the built-in hygiene rule, and the same check passes once the agent removes it.
python3 vae.py gate --repo .
VERIFIED[hygiene.probes]=false BC glob='*' files=1; hits=['component.ts:2']
PROVEN: ∅
REMAINS: hygiene.probes
AGENT_CMD: FIX hygiene.probes: component.ts:2
the agent removes the marker and calls the gate again
python3 vae.py gate --repo .
VERIFIED[hygiene.probes]=true BC glob='*' files=1
PROVEN: hygiene.probes
REMAINS: ∅
AGENT_CMD: ∅
Lessons become code, not just memory
A lesson the agent only remembers depends on the next session reading it and acting on it. A lesson codified into the verifier is checked by every later gate, in every session, whatever the agent remembers. So wrap promotes each lesson to the strongest form that fits: a test or a .agents/VERIFY.py rule first, a memory line only when no program can check it, and it deletes the written lesson once a check enforces it.
The paper records this in its study: one human correction (the rendered API pages lacked each state's configuration) became an end-to-end audit of all 59 component pages, so later work checks that requirement without the original exchange (section 5.1). Over the study the verifier grew from 33 to 87 checks, while MEMORY.md held eight entries in 2,311 bytes from October 6 to the end (sections 2.4 and 4.1).
- Test or
VERIFY.pyruleA program: every later gate runs it, and its failure tells the agent what to fix. MEMORY.mdlineOne line with its reason, loaded at every session start: 240 characters, 4 KiB in all.EPISODES.mdentryA lead the gate logs;wrappromotes it to a rung above or deletes it with evidence.
every rule has this shape; this one ships in the template
(pattern shortened)
RULES = [{
"id": "tests.no-mocks",
"kind": "not_regex",
"glob": "*",
"pattern": r"unittest\.mock|MagicMock\(|…",
"claim": "tests exercise real subsystems, not mock frameworks",
}]
🧭 What runs by itself, what the agent starts, what stays yours
Seven skills, one gate the hooks enforce, and two steps no program takes for you. Skill names show the skills-CLI form; the Claude Code plugin namespaces them (/defuss-vae:plan) and Codex writes $plan.
Hook · plugin
Runs by itself. SessionStart loads rules and memory; the Stop hook runs the same gate the agent loops in-turn and blocks once per turn; PreToolUse denies git commit until that gate is green for the current code.
Skill
A slash command with its own procedure. You can start any of them; the agent starts six when their condition holds.
Manual
Your call. You read the commits and decide what is pushed, tagged and released.
| Stage | Kind | Who drives it | What it does |
|---|---|---|---|
| /plan | Skill | you; the agent only for a risky task you gave no plan for | Traces the real code path, probes unknowns instead of guessing and writes the smallest plan, in which every acceptance criterion is an executable check, to plans/. A plan the agent started waits for your review. |
| /implement | Skill | you or the agent, fully agentic | Understands first, fixes the root cause and its sibling callers, writes the minimum code, adds tests against real subsystems and loops the gate in-turn. |
| /verify | Skill | the agent at a goal or milestone; you any time | Reviews against requirements, callers and tests, and fixes the defects it confirms: the whole change when it is large, otherwise its paths and tests. |
| gate | Hook | hooks, automatic | verify runs the project's own make verbs; review checks requirements, every changed path and its callers; docs records why this design beats the plausible alternative. Doc pages get the prose check and catalog review instead of the test suites. |
| /doc | Skill | you any time; the agent after implementing | Writes and checks documentation pages: claims grounded in code and tests, Mermaid where the content is schematic, then the static prose check and a walk through each page against every catalog rule. |
| /doc-edit | Skill | you any time; the agent after doc |
Edits the named pages as instructed with the same grounding, then walks each one against every catalog rule, the whole page unless you name parts. |
| /wrap | Skill | only you, since it commits | Splits the work into coherent Conventional Commits, updates CHANGELOG.md and codifies what the work taught into the verifier as tests or VERIFY.py rules; only a lesson no program can check becomes a concise memory line with its reason. Then it audits agent memory against the current code. |
| /status | Skill | you or the agent, any time | Shows what runs (sub-agents, the service, free disk, RAM and GPU), reconciles the sub-agent registry with the process table and names the next step per agent. |
| human review | Manual | you | You read the commits. Nothing has been pushed yet. |
| release | Manual | you, via CI/CD | Push, tag, publish to package managers. On GitHub, init adds a workflow that runs the same make setup and make verify, so CI runs the same gate. |
Observed, not promised*
From the technical report Verified Agentic Engineering in Practice (opens in a new tab) (Homberg, October 2026): the defuss-shadcn component system, developed under the VAE method from 2026-09-05 to 2026-10-07, measured from git, verifier output and session transcripts.
* One project, one maintainer, one model family, no control group; the task windows include contemporaneous work and are not controlled benchmarks (report, section 6). A green gate establishes its encoded checks, not unspecified behavior, and a review attestation records the agent's claim, not its quality. The gate's own share of a cold run is about 105 ms (make bench, defuss-vae 0.6.0, Apple M4, medians of 7 runs; see the README).
Four repository-wide changes, one session
2026-10-06, from each request to its reported endpoint (report, Table 2).
| Task | Minutes | Tool calls | Human corrections |
|---|---|---|---|
| Typed API docs, 59 components | 241.9 | 437 | 1 |
| Test coverage 31.6% → 75.6% | 20.6 | 48 | 0 |
| Layout: 228 folders, 19 sections | 43.4 | 98 | 0 |
| Type errors 1,147 → 0 | 27.3 | 130 | 0 |
Templates to counter agentic derails
A symbolic dialect to save tokens
Semantic rules to reduce hallucinations
Expert coding guidelines
Expert architecture guidelines
Increased efficiency and scalability

Want verified agentic engineering in your team?
Aron is a freelance AI researcher with over 25 years of professional software engineering experience. His interest in machine learning and AI began long before ChatGPT. He is an O’Reilly author (2011), conference speaker, and longtime mentor to software engineering teams.
...if you're looking for a remote training/mentoring session for you or your engineering team.
🧑💻 Up and running in a minute
Requirements: python3 3.9 or newer, git and make. Hooks and CLI are stdlib-only.
Claude Code recommended
The full experience: the commit gate, the Stop-hook gate and memory injected at session start.
- Run the two commands inside Claude Code, or the shell form from your terminal.
- Start a new session afterwards.
- Then describe a task, e.g.
/defuss-vae:planfollowed by your request in plain words.
/plugin marketplace add kyr0/defuss-vae
/plugin install defuss-vae@defuss-vae
claude plugin marketplace add kyr0/defuss-vae
claude plugin install defuss-vae@defuss-vae
Codex
The same plugin folder ships an Agent Plugins 1.0 manifest (plugin/plugin.json) and plugin/.codex-plugin/plugin.json.
- Install defuss-vae from
/plugins, then trust its hooks in/hooks; the README marks this route as untested. - Untested: hook parity with Claude Code, including the Stop-hook block; see COMPATIBILITY.md.
- Start a skill with
$planfollowed by your request.
/plugins
install defuss-vae
/hooks
review and trust the defuss-vae hooks
Cursor, Gemini CLI, Copilot, Windsurf and more
The open skills CLI detects your agents and installs the seven skills into any Agent Skills host.
- Skills only: no hooks and no CLI, so nothing blocks
git commit. - Clone the repository once and run the gate yourself, as in the second window.
- Pick agents with
--agent; add--globalfor all projects.
npx skills add kyr0/defuss-vae --skill '*'
git clone https://github.com/kyr0/defuss-vae ~/defuss-vae
python3 ~/defuss-vae/plugin/scripts/vae.py gate --repo .
Let your agents learn to write programs that review your code and decide when work is done
Seven skills, a gate the hooks enforce, and lessons codified into checks that every later gate runs. The agent does the work; you review the risky plans and the commits.
MIT licensed · python3 ≥ 3.9 · git · make