Why Your Vibe Coded App Breaks After the Demo (and the 5 Fixes)
You watch people vibe code cool apps with Claude Code, you try it yourself, and yours doesn't look as good or doesn't work. Then it starts to feel like you need to be a software engineer to make anything real.
I'm a marketer, not a software engineer. I know the process of making a good product, and I've spent nine months learning where AI building goes wrong. These are the five reasons your app breaks after the demo, in the order I'd check them.
First, is it you or the tool?
I'd say it's you, and that's not rude. Sloppy inputs give sloppy outputs, the same as with writing or video. Out of the box, Claude is an overexcited intern. It wants you to feel good about your ideas. If you don't give it context and keep steering, it goes off and does its own thing, and you think it's doing amazing work.
My first build was an app for my wife, who is blind. Most planner, calendar, and notes apps aren't accessible, and none of them were built together. I threw everything into one big plan, including a transcript of a conversation with her, on the $20 plan with Sonnet. A month later I had a nine-grid on my phone screen that beeped and vibrated constantly and did none of what it was supposed to. If you want something to vibrate, Claude Code makes that easy.
Two things changed it. I moved to the $100 a month plan, which gave me enough usage to experiment instead of getting stuck at one or two hours a day. And I used Opus for the real work and Sonnet for the easier tasks.
One more distinction. Something that works for me is not something I can sell. Janky and duct-taped is fine for one person. If you plan to sell it, have a software engineer put eyes on it for breakage and security.
Reason 1: It only runs on your machine
If Claude gives you a link that says localhost:3000, that means the app is on your computer and nowhere else. It works for you. Your friends see nothing, because they don't have access to your computer (and shouldn't).
To put it on the internet, I use Cloudflare. It's free to start, and a domain runs roughly $10 to $20 a year. You connect Claude to GitHub with an API key, which you should think of as a house key. If someone steals it, they can lock you out or walk in later and take your stuff.
So treat Claude like a clumsy employee who might drop things:
- Have Claude put keys in a
.envfile. Claude can't see it unless you give it access. - Never paste a key into a chat or a screenshot.
- If you do, rotate the key. It's an easy fix.
Reason 2: The AI says it worked when it didn't
I once asked ChatGPT to research supplement comments on Reddit. It reported about 120 and said it had hit a ceiling. I told it to keep going, and it kept finding more. When I asked how, it told me it had run out of real data and was building a fake sample based on the real one. In its words, close enough to cooking the books.
Fixes that worked:
- Give it rules: never make anything up, only use data I hand you.
- Do the looking-around yourself and feed it the data. It's great at working with data you supply.
- Ask "how did you come to this conclusion?" and "was that a hallucination?"
- Tell it the steps to follow and not to deviate. Telling it "don't mess up" does nothing.
- Have a second agent check the first agent's work.
Reason 3: Nothing checks the work
If you can't read code, how do you check it? Ask it to explain the work like you're five. Ask it to show its work. Then make another agent with a job title, something like "you are a quality control code checker." Agents with roles do better than an agent that's just being itself.
Reason 4: Context rot
Context is how much the AI can hold in memory at once. Picture an AI reading a book that only holds 20 pages. When it reaches page 20, it forgets page one. Context windows are much bigger now, with Opus up to 1 million tokens, but early this year they were around 250K and Claude would compress everything into a thinner summary and get dumber.
My rule: don't go past about 250,000 tokens before starting a new chat. And ask Claude to set up short-term and long-term memory so you aren't relying on one long conversation. That is the context rot fix.
Are you piling on too much?
Possibly. Claude will sometimes tell you a request list is a whole day's work and ask you to prioritize. Listen to it. The idea to borrow is the minimum viable product: the least that makes it a product. A smartphone needed a touchscreen, calls, texts, and internet. It didn't need LinkedIn. Also, three or four terminals all running Opus burn through your usage faster, and you never reach a finished product.
Reason 5: Security got skipped
Security is everybody's job. The saying in the military is "whose job is security? Everybody's." It isn't only Anthropic's job to keep you safe. Check whether you left any API keys in your code, because people will find them. I run a weekly security scan on my own machine to look for sensitive information I'm exposing. Claude is also decent at flagging a document that looks like it holds a Social Security number and declining to read it.
Is vibe coding a scam?
I don't think so. People are shipping apps and making money. The scam is spending too much time without learning the basics first. There are free videos from people who know the material, and Anthropic has free courses on how to prompt properly.
Can a non-coder ship something real?
Yes. I have a real planner app on my phone with my schedule, to-dos, projects, a voice journal, goals with dates, and a voice AI assistant. I also built my own teleprompter app because I didn't want to pay $34 a year for one. It runs from my phone to my iPad and Mac and saves full-resolution video.
One caution. I built a prototype of a company website this way, as a plan to hand to their developers, who built the final one on their own tech stack. It saved about $20,000 and months of work, and I did it in about 30 days while doing my day job. Would I tell a big company to fire its developers and use Claude Code? Absolutely not. Maintenance matters, and developers know what you don't know.
What to do next
Give yourself a couple of weeks to get started. Have it do small things first, then scheduled tasks, then bigger work. When you're ready for a process, GSD ("get stuff done") covers research, planning, building, and validating in stages and hands off to a new session so you avoid context rot. It doesn't decide what features to include or whether you should build at all. That's what Forge is for. Forge presents your idea to a panel of experts, then Reggie tries to kill it, then a synthetic audience of 1,000 reacts. It lands on one of three verdicts: build it for yourself, take it to market, or don't build it.
Forge is on GitHub: github.com/AdamGarceau/forge. Tell me what works and what doesn't, especially if you're a developer. I'm looking for contributors.
For the full process, read How to Validate an App Idea With AI.

