I wanted to “build by vibe” using nothing but my own PC.
It all started with a single phrase I came across on YouTube — and turned into a system that builds apps automatically using only local LLMs.
01It started with a YouTube recommendation
One day, while idly browsing YouTube, the words “Vibe Coding” jumped out at me.
Instead of writing code line by line, you tell an AI “I want an app like this,” run whatever comes out, and ask again. You build by vibe rather than by the details of the program — a whole new style of development.
Just describe it in words, and an app takes shape. Honestly, it blew my mind a little.
02But could I do it “entirely on my own PC”?
The videos used huge cloud AI services. They're convenient, no doubt about it. But then a thought struck me.
Could this be done with just the local LLMs on my own machine?
No AI on the other side of the internet — just open models running on the GPUs in my home PC, building an app from a spec. No usage fees to worry about, so I could try again and again all night long. If it worked, it would obviously be fun.
What is AutoDev LG?
Give it a spec, and a team of local LLMs automatically repeats design → write tests → implement → verify → fix until a working app is finished. “LG” comes from LangGraph, which forms the skeleton of the workflow. The coding itself is done through OpenHands.
03The GPU-multiplying disease
The first attempt was at the end of July. I paid just $2 to DeepSeek, famous for being cheap, and had its AI build things. But every time a feature was added, the old ones vanished — regression hell — and I got bored in one day.
About a month later, I swapped my GPU for a plausible reason: “smoother gaming.” That was the beginning of the end.
08-27RTX 5070 → 5070 Ti. ¥169,800 on a flea-market app. Same day, “I want to try bigger LLMs too,” so a 5060 Ti 16GB as well. (I've lost it.)
09-03Two more 5060 Ti cards. (I've lost it.)
09-04One more 5060 Ti. Five in total. No PC to plug them into yet. (Completely out of my mind.)
09-06While grumbling “more GPUs, but still DeepSeek...,” news of a free month of ChatGPT Plus. Signed up instantly. Talking it over, the idea of today's AutoDev LG slowly took shape
09-08Switched to a LIAN LI O11D EVO XL case (huge). Crammed five cards into my trusty gaming motherboard with riser cables: 80GB of VRAM, complete
A server motherboard was out of reach. So: riser cables, and cram them in.
From that day on began the days of having ChatGPT build AutoDev LG and testing it with local LLMs. ...Or so I thought, until ChatGPT quickly hit its usage limit. The same day I signed up for Claude Pro too, and the long journey of run → fail → fix began.
04A team of small models
LLMs that run locally aren't as smart as the cutting-edge cloud models. Hand everything to one of them and it gets lost halfway, or misreads the tests.
So AutoDev LG splits the roles across several models: a leader that directs the whole thing, a worker that does the hands-on work, and a reviewer that checks the results. When one gets stuck, it passes the baton to another model — just like a small dev team.
It all runs on five GPUs: one GeForce RTX 5070 Ti and four 5060 Ti. Open models such as Qwen, Gemma and gpt-oss take turns working on them.
And outside that team, there's another division of labor.
Who does what: the human, Claude Code and AutoDev LG
👤 Human: says what app to build, like “Make me an Invaders game.” Judges only the finished app, and just gives a quick instruction if a feature is missing. Honestly, never even looked at the app spec (lol). Changes to AutoDev LG's own goals, policy and spec are decided by the human when Claude Code asks.
🛠 Claude Code: writes the spec for AutoDev LG from that request, runs AutoDev LG, and when it fails, fixes AutoDev LG itself.
⚙️ AutoDev LG: takes the spec and builds the app with its team of local LLMs.
The first assignment was a small taskboard (a task-management app) built with Tkinter — just add, delete and change status. Even so, at first it never made it to the end.
Run it (one automatic development attempt), analyze where it stumbled, fix the framework, bump the version. Repeat that, and the version number passed 100 in no time.
2026-07-30The prequel. Paid DeepSeek $2 to try it, but regression hell — every new feature erased the old ones — made me give up in a day
2026-08-27 – 09-04GPUs multiply: 1, 2, 4, 5
2026-09-08The 80GB VRAM setup is complete. Development starts with ChatGPT and Claude as partners
2026-09-09First prototype. Still a single Python script
2026-09-13Rebuilt as the LangGraph-based “LG” line
2026-09-14First PASS on taskboard — the moment an app was finished automatically. ...Followed by a long string of failures
2026-09-21A PASS after a week. From here, the battle to turn “by luck” into “every time”
2026-09-22Bought a Mac mini, telling myself “I'll make iPhone apps someday.” Not used once so far (lol)
2026-09-263 PASSes in a row on taskboard
2026-09-27Jumped on Hermes Agent after a YouTube video — useless with local LLMs. A day wasted, then back to AutoDev LG (meanwhile the local LLMs quietly made 9 attempts)
2026-09-28The day the whole dev environment (WSL) vanished. Restored from backups within a day
NowTaking on the next assignment, an Invaders game. The goal: 3 PASSes in a row
06How far we've come
140+
versions shipped
100+
AutoDev LG runs
1,200+
tests for the framework itself
5
GPUs at work
The goal is to build more kinds of apps from even shorter descriptions. We're not there yet, but every step of the way is published daily in the dev log.
The good days and the spectacular faceplants alike. Come take a look.