How to Validate an App Idea With AI (Before You Build It)
Most AI will tell you your idea is the biggest breakthrough yet. It says that every time, and most of the time it's wrong. I wanted the opposite: a system that tells me no before I burn a month building the wrong thing.
I built one. I ran it on my own idea, and it told me not to ship it. This post is the written version of how that works, so you can steal the process even if you never press play.
Why building with AI keeps breaking
Here is how most people make software with AI. You have an idea, you spitball it to the AI, and you hope what comes back is close. You get excited for a few minutes. Then you try to use it and there are problems. You fix one and it breaks something else. Fix, break, fix, break.
Most people quit there. Plenty never start. I see it in forums all the time: "somebody should vibe code that."
There are four reasons it goes sideways out of the box:
- It has no standard. It doesn't know what good looks like because you never told it.
- It forgets. The conversation scrolls away from where you actually are, and the more it has to remember, the more it compresses your instructions into a watered-down version. That is where the fix-break-fix loop comes from.
- It's a yes man. It tells you what you want to hear, then deepens your confirmation bias.
- It has no process. It brute-forces the job like a cheerful intern on day one, learning on your dime.
If you've felt any of those, the tool isn't broken. It's missing context and a process.
Where this started: a land navigation app
The idea came from my Army Reserve unit, where we were learning land navigation with a paper map and compass. A paper map is reliable and I know how to use it, but I wanted a calculator that did it better.
So I made an app called Dead Reckon. You take a picture of a military map, plot three known points, and it triangulates your position using your phone's GPS. I started it on my phone in the morning during class, and it was finished by the end of lunch with very little prompting.
What I did differently was this:
- I told the AI to find a group of experts and ask them about the problem.
- I gave it the actual field manual, so it had the right context.
- I had it ask a synthetic audience that matched the real one: ROTC cadets, soldiers, anyone who gets lost in the woods with a paper map.
- Then I had it run user testing on the result.
I never asked for the feature that shows you where you are on the map. I never asked it to point me in the right direction when I tap two spots. It built both because it asked the experts, then the audience. Dead Reckon was about 90% done by the time that process finished.
What Forge is
After Dead Reckon I took that literal chat, gave it to Claude, and said "turn this into a skill and call it Forge." Forge is not one genius prompt. It's a set of skills I built over nine months of using Claude Code every day: skills that run the survey audience, skills that work out who the audience is and who matters most in it, and skills that do market research and tell you whether to add a feature or just change the words you use to sell it.
The most important piece is a character named Reggie. His only job is to tell you no. He has seen what a failed idea looks like, and he won't sugarcoat it. If there's a feature you love that is actually pointless, he'll save you from it. And if Reggie can't kill the idea himself, there's a swarm of five or six skeptics who each attack a different angle and score it for one of three verdicts:
- Build it for myself.
- Take it to market.
- Don't build it at all.
Knowing what not to build is worth more than knowing what to build, because the wrong road costs you weeks.
Forge pairs with a second skill called GSD ("get stuff done"). GSD manages the AI's context, breaks the work into stages, and has a second AI check the first one at every stage. GSD fixes the forgetting. Forge decides whether you should be building at all. GSD never stops to say "this is a bad idea."
Why a synthetic audience is worth trusting
Forge also borrows from the Google Ventures design sprint. Their problem was that finding out whether anyone wanted a product took 18 months. What they found is that five people from the target audience, a sticks-and-glue version that looks finished, and about five days will tell you whether something mostly works. Five is enough to know if it is 80% there, and 80% is enough to launch.
You probably don't have a decision maker, a designer, and an engineer sitting next to you. Forge runs that same sprint with those roles through five synthetic users.
Can that really show you what a real person would? I was skeptical too. Stanford ran a study where people took a survey, then took it again two weeks later and matched their own answers about 85% of the time. AI models role-playing those same people were about 85% accurate too.
I tested it on real data. I took call data from a company I worked with and turned it into a synthetic audience of 1,000 respondents. I had it weight them to find the ideal customer profile: 40% of the audience was one type of person, 30% another, 20% a third, along with the problems each was trying to solve. I built a website from that, and it cut their churn in half and increased calls for new services.
Does it just flatter you? On one of my products it took the score from a 6.2 out of 10 to an 8, then stopped and said that was the ceiling. It did not give me the 9 I asked for. It told me the exact steps to get there, and those steps were on me. That is the whole point of the tool.
The cheap way to run it
If you have the hardware, the synthetic surveys are free and run locally. I run a Gemma 4 12B model on a MacBook with 24 GB of RAM. Claude orchestrates and the local model does the volume. If I want speed, I let Claude do all of it. If I want it cheap, I run it overnight. I once ran a 10-hour survey and had results by morning.
The day Forge killed Forge
I ran Forge through Forge. It was the scariest run of my life, because it's my baby. It told me not to ship it: the only person who had ever used it was me, and nobody was asking for something like it.
I'm putting it on GitHub anyway, because I think people who need it don't know it yet. Take a project you already built, or one that frustrated you because you couldn't get Claude to do it, and run it through Forge. See if the output beats what you had.
One honest note on Dead Reckon: it doesn't make any money. Forge's verdict on it was "don't sell it, this is just for you," and that was right.
The moat is you
AI can't build anything original for you. Forge can't, GSD can't, nothing can. The part that can't be copied is your imagination, your experiences, and your data. Someone could rebuild Dead Reckon in a weekend. They'd still be slower than me, because they don't have the process. Someone could screenshot my planner app and tell Claude to copy it. They can't copy my brain.
And if the thing stopping you is "what if someone else vibe codes my idea," the risk is low. Most people are too lazy to even open Claude Code and start asking questions, and your idea probably isn't on their radar.
What to do next
Run your own idea through the same sequence this week:
- Write down what good looks like before you build anything.
- Ask the AI to play a panel of experts and give them the real source material.
- Have it play your audience, then attack your idea from five or six angles.
- Ask for a verdict, and ask what it would take to reach 9 out of 10.
Forge is on GitHub: github.com/AdamGarceau/forge. Try it and tell me your verdict in the comments, including the "it didn't work for me" ones. I want the objections.

