Who does what, and how a single automatic development attempt (a “run”) unfolds — in two diagrams.
01The cast and their roles
AutoDev LG runs on four layers. The local LLMs at the bottom build the apps. Claude Code grows the system that makes them (the framework), and a human sets the direction.
👤 Human the owner
Says what app to build (e.g. “Make me an Invaders game”)
Judges only the finished app (the final output), and gives a quick instruction if a feature is missing. Honestly, never even looked at the app spec
Decides AutoDev LG's goals, policy and spec changes (when Claude Code asks; e.g. “3 PASSes in a row on INVADERS”, “no app-specific logic in the framework”)
🛠 Claude Code develops the framework (ChatGPT did too, earlier)
① Writes the spec for AutoDev LG from the human's request
② Starts and monitors AutoDev LG, then analyzes the logs to find why it failed
③ Fixes AutoDev LG itself: write the tests first → bump the version → release gate (1,200+ tests) → deploy AutoDev LG → back to ② until app generation succeeds reproducibly
④ Writes handoff notes in case the Claude Code session gets cut off
spec · start · fix ↓↓↑↑ run results · logs
⚙️ AutoDev LG an automatic development workflow built on LangGraph
It takes a spec and hands out work to the agents below, in order. When an agent gets stuck, it automatically swaps in a different model. At the end of a run, it also writes a user guide to go with the finished app.
Leaderrequirements, design, test plan Qwen3.8 Flash-Nextif stuck → Qwen3.8 27B → Gemma 4 → gpt-oss-120b
Reviewerchecks that the design and tests really mean what the spec says Gemma 4 26B
Test writerwrites the test code that decides pass or fail Qwen3.8 27B
Workerwrites and fixes the actual code through OpenHands Qwen3.8 Flash-Nextif stuck → Gemma 4
prompts ↓↓↑↑ answers (JSON · code)
🖥 Local LLMs running on my home PC with llama.cpp (llama-server)
One model at a time is loaded across five GPUs and swapped whenever the role changes. No cloud AI is used.
RTX 5070 Ti 16GBRTX 5060 Ti 16GBRTX 5060 Ti 16GBRTX 5060 Ti 16GBRTX 5060 Ti 16GB= 80GB VRAM
02The flow of one run
It starts from a spec, and if every test passes, it's a PASS. If something fails along the way, it diagnoses the cause, fixes it according to its type, and verifies again.
Match code against designImplementation Contract Gate
Run the testsQuality: Public + Acceptance tests
all passed
🎉 PASS
failures
Diagnose the causeDiagnose (Public / Acceptance)
A bug in the testsfixture / test bug
Safely repair or regenerate the teststest repair / regenerate
A gap in the designcontract gap
Review the designcontract reconcile / review
A bug in the codeimplementation bug
Fix the codecode repair (coding)
Verify again → continue if it made progressre-validate → validated progress?
Repeats the same failure → switch to another model. Still no progress → the run ends (NOT PASS)
at the end of the run
📦 Output: the app files + 📘 user guideproject files + USAGE_GUIDE.md
The user guide is written by a program, not an LLM: it states only what can be verified from the spec, the design and the generated code (the launch command comes from the real code, so “I did what it says and it didn't start” is unlikely).
* The real workflow has more than 20 steps, including test-plan audits, harness recovery and scope checks. The diagram shows only the main path.
03After a run
If it's NOT PASS, Claude Code analyzes the logs. Then it fixes the framework in a general way — never with a fix tailored to one specific app — bumps the version, and starts the next run. Thanks to this loop, the version number has passed 140.
The daily results are updated every morning in the dev log.