v0.8.0Verified Agentic Engineering: no commit until the gate is green

Your agents learn to write the programs that decide when work is done

verify → review → docs hooks that fail closed your own make verbs lessons that become tests about 0.1 s of overhead Claude Code, Codex & more

A prompt can ask a coding agent to run the tests, mock nothing and clean up its debug prints, but it cannot make the agent do it. defuss-vae moves the checks out of the prompt and into code: hooks deny git commit and send the agent back to work when it stops early, until the current code passes verify → review → docs. With wrap, the agent codifies each lesson into the verifier, as a test or a rule that every later gate runs; only a lesson no program can check stays a memory line.

Method, plugin and paper by Aron Homberg (kyr0) (opens in a new tab) – Open Source (MIT (opens in a new tab))

A gate the agent can't talk past

Hooks deny git commit until the exact code fingerprint is verified, reviewed and documented. A prompt can be ignored; a denied commit cannot. Any edit changes the fingerprint and restarts the gate.

Fails closed

A crashing gate denies the commit and blocks the stop once, instead of waving changes through. Hooks run on plain python3 3.9 or newer, stdlib only: no dependencies, no daemon, no uv in the hook path.

Your checks, your toolchain

The gate runs your Makefile verbs (lint, test, coverage, e2e) and the rules in .agents/VERIFY.py. Defaults for Go, Rust, JVM, .NET, JS/TS and Python fill the gaps, and an existing toolchain stays.

Lessons become verifier code

wrap codifies each lesson into a test or a VERIFY.py rule, so every later gate checks for it, whatever the agent remembers. Only a lesson no program can check stays one memory line, within 4 KiB.

Docs are gated too

Every changed page passes a static prose check and a review, window by window, against a 58-rule catalog: evidence, logic, terms, relevance, structure, language and typography.

You stay in charge

The agent starts six skills when their step comes. A risky task without your plan gets one for your review first; only you start wrap, which commits, and nothing is pushed or released without you.

Plan agentic

Agent swarm implements

Verifier binds agent to quality

Tiny review for humans

Agentic self-learning, semantic commits, changelog updates

Human in control for every release

How it works

🔁 Two loops: one fixes the work, one codifies the lessons

Programs check the work and tell the agent what remains; the agent fixes it and runs the gate again. The second loop codifies what the work taught into the verifier itself, so every later gate checks it; memory holds only what no program can check. Figure 1 of the paper (opens in a new tab) describes the general mechanism; this is how defuss-vae 0.8.0 implements it.

You goal · plan · approval
SessionStart hook Task context rules · stack defaults · memory
Model in harness Coding agent plan · implement · fix
Artifacts Code · tests · pages one content fingerprint
Stop hook · vae.py gate Gate verify → review → docs
  • your make verbs
Program output FIX · REMAINS the next instruction
git commit denied until green
EPISODES.md FAIL · DONE · FINDING
You start wrap Reconcile codify · audit · commit
Codified lessons Verifier code tests · VERIFY.py rules
Only if no check fits MEMORY.md one line with its reason
  1. goal
  2. load
  3. remaining
  4. next call
  5. green
  6. outcomes
  7. leads
  8. codify
  9. every gate runs it
  10. else one line
  11. next session

Scroll the diagram sideways to see the whole loop.

What the agent reads

The gate's output is plain text the harness hands to the model's next call: which checks passed, which remain, and the command that fixes each one. Checking, selecting the next window of a page and formatting the output make no model call.

The output lines are captured verbatim from defuss-vae 0.8.0 on a small fixture, as printed in the paper's Figure 2: a temporary probe marker on line 2 fails the built-in hygiene rule, and the same check passes once the agent removes it.

python3 vae.py gate --repo .
VERIFIED[hygiene.probes]=false BC glob='*' files=1; hits=['component.ts:2']
PROVEN: ∅
REMAINS: hygiene.probes
AGENT_CMD: FIX hygiene.probes: component.ts:2
the agent removes the marker and calls the gate again
python3 vae.py gate --repo .
VERIFIED[hygiene.probes]=true BC glob='*' files=1
PROVEN: hygiene.probes
REMAINS: ∅
AGENT_CMD: ∅

Lessons become code, not just memory

A lesson the agent only remembers depends on the next session reading it and acting on it. A lesson codified into the verifier is checked by every later gate, in every session, whatever the agent remembers. So wrap promotes each lesson to the strongest form that fits: a test or a .agents/VERIFY.py rule first, a memory line only when no program can check it, and it deletes the written lesson once a check enforces it.

The paper records this in its study: one human correction (the rendered API pages lacked each state's configuration) became an end-to-end audit of all 59 component pages, so later work checks that requirement without the original exchange (section 5.1). Over the study the verifier grew from 33 to 87 checks, while MEMORY.md held eight entries in 2,311 bytes from October 6 to the end (sections 2.4 and 4.1).

  1. Test or VERIFY.py ruleA program: every later gate runs it, and its failure tells the agent what to fix.
  2. MEMORY.md lineOne line with its reason, loaded at every session start: 240 characters, 4 KiB in all.
  3. EPISODES.md entryA lead the gate logs; wrap promotes it to a rung above or deletes it with evidence.
every rule has this shape; this one ships in the template
(pattern shortened)
RULES = [{
    "id": "tests.no-mocks",
    "kind": "not_regex",
    "glob": "*",
    "pattern": r"unittest\.mock|MagicMock\(|…",
    "claim": "tests exercise real subsystems, not mock frameworks",
}]
Workflow

🧭 What runs by itself, what the agent starts, what stays yours

Seven skills, one gate the hooks enforce, and two steps no program takes for you. Skill names show the skills-CLI form; the Claude Code plugin namespaces them (/defuss-vae:plan) and Codex writes $plan.

Hook · plugin

Runs by itself. SessionStart loads rules and memory; the Stop hook runs the same gate the agent loops in-turn and blocks once per turn; PreToolUse denies git commit until that gate is green for the current code.

Skill

A slash command with its own procedure. You can start any of them; the agent starts six when their condition holds.

Manual

Your call. You read the commits and decide what is pushed, tagged and released.

The stages of a task, from plan to release.
Stage Kind Who drives it What it does
/plan Skill you; the agent only for a risky task you gave no plan for Traces the real code path, probes unknowns instead of guessing and writes the smallest plan, in which every acceptance criterion is an executable check, to plans/. A plan the agent started waits for your review.
/implement Skill you or the agent, fully agentic Understands first, fixes the root cause and its sibling callers, writes the minimum code, adds tests against real subsystems and loops the gate in-turn.
/verify Skill the agent at a goal or milestone; you any time Reviews against requirements, callers and tests, and fixes the defects it confirms: the whole change when it is large, otherwise its paths and tests.
gate Hook hooks, automatic verify runs the project's own make verbs; review checks requirements, every changed path and its callers; docs records why this design beats the plausible alternative. Doc pages get the prose check and catalog review instead of the test suites.
/doc Skill you any time; the agent after implementing Writes and checks documentation pages: claims grounded in code and tests, Mermaid where the content is schematic, then the static prose check and a walk through each page against every catalog rule.
/doc-edit Skill you any time; the agent after doc Edits the named pages as instructed with the same grounding, then walks each one against every catalog rule, the whole page unless you name parts.
/wrap Skill only you, since it commits Splits the work into coherent Conventional Commits, updates CHANGELOG.md and codifies what the work taught into the verifier as tests or VERIFY.py rules; only a lesson no program can check becomes a concise memory line with its reason. Then it audits agent memory against the current code.
/status Skill you or the agent, any time Shows what runs (sub-agents, the service, free disk, RAM and GPU), reconciles the sub-agent registry with the process table and names the next step per agent.
human review Manual you You read the commits. Nothing has been pushed yet.
release Manual you, via CI/CD Push, tag, publish to package managers. On GitHub, init adds a workflow that runs the same make setup and make verify, so CI runs the same gate.

Observed, not promised*

From the technical report Verified Agentic Engineering in Practice (opens in a new tab) (Homberg, October 2026): the defuss-shadcn component system, developed under the VAE method from 2026-09-05 to 2026-10-07, measured from git, verifier output and session transcripts.

55 → 231components in 33 days, across 220 commits and 17 releases, while the verifier grew from 33 to 87 checks
59 / 59interactive components got typed API documentation in one session, with one human correction
31.6% → 75.6%line coverage in 20.6 minutes; two contract tests exercised every interactive component
85 / 85minified artifacts byte-identical after the agent removed 1,147 type errors

* One project, one maintainer, one model family, no control group; the task windows include contemporaneous work and are not controlled benchmarks (report, section 6). A green gate establishes its encoded checks, not unspecified behavior, and a review attestation records the agent's claim, not its quality. The gate's own share of a cold run is about 105 ms (make bench, defuss-vae 0.6.0, Apple M4, medians of 7 runs; see the README).

Four repository-wide changes, one session

2026-10-06, from each request to its reported endpoint (report, Table 2).

Task Minutes Tool calls Human corrections
Typed API docs, 59 components241.94371
Test coverage 31.6% → 75.6%20.6480
Layout: 228 folders, 19 sections43.4980
Type errors 1,147 → 027.31300

Templates to counter agentic derails

A symbolic dialect to save tokens

Semantic rules to reduce hallucinations

Expert coding guidelines

Expert architecture guidelines

Increased efficiency and scalability

Portrait of Aron Homberg

Want verified agentic engineering in your team?

Aron is a freelance AI researcher with over 25 years of professional software engineering experience. His interest in machine learning and AI began long before ChatGPT. He is an O’Reilly author (2011), conference speaker, and longtime mentor to software engineering teams.

Talk to Aron on LinkedIn (opens in a new tab)

...if you're looking for a remote training/mentoring session for you or your engineering team.

Install

🧑‍💻 Up and running in a minute

Requirements: python3 3.9 or newer, git and make. Hooks and CLI are stdlib-only.

Claude Code recommended

The full experience: the commit gate, the Stop-hook gate and memory injected at session start.

  • Run the two commands inside Claude Code, or the shell form from your terminal.
  • Start a new session afterwards.
  • Then describe a task, e.g. /defuss-vae:plan followed by your request in plain words.
/plugin marketplace add kyr0/defuss-vae
/plugin install defuss-vae@defuss-vae
claude plugin marketplace add kyr0/defuss-vae
claude plugin install defuss-vae@defuss-vae

Let your agents learn to write programs that review your code and decide when work is done

Seven skills, a gate the hooks enforce, and lessons codified into checks that every later gate runs. The agent does the work; you review the risky plans and the commits.

MIT licensed · python3 ≥ 3.9 · git · make