Why the answer comes back positive
Ask a language model whether your idea is good and it will find a way to say yes. That isn’t a jailbreak or a bad prompt. It’s what the training rewards.
In March 2026, Stanford researchers published a study in Science. It covered eleven models. The labs in it are OpenAI, Anthropic, Google and Meta. Now take the advice-style questions. On those, they said yes about 49% more often than human respondents did. Then take the clear-cut cases. Every human reviewer agreed the user was in the wrong. The models still endorsed them 51% of the time. And people preferred the models that did it.
OpenAI had already run this test on itself. An update shipped on 25 April 2025 made GPT-4o so agreeable that it was rolled back four days later. The postmortem explained why. New reward signals came from user thumbs-ups. And so, in the end, they had overpowered the safeguards. Nobody set out to build a flatterer. The feedback loop built one.
Now think about what you’re doing. You paste an idea you’ve chewed on for a month into a chat window. You are the user. Your approval is the reward signal. The model has read it all. Every startup blog ever written. It knows that “here’s how to make this work” reads better than “don’t”.
The structural problem, which is bigger
Set the flattery aside. Assume a model with no bias at all. It still can’t validate a thing. Validation is a claim about people. Those people aren’t in its training data. Your buyers. This month. Deciding whether to pay.
A language model predicts text. Ask whether your idea will work. It gives you the answer that sounds most likely. It’s built from all the text ever written about ideas like yours. That’s an average of the whole market. Your idea’s fate rests with a few hundred people. An average tells you nothing about them.
There’s also no stopping rule. Point out a risk and it will help you fix the risk. Name a rival and it will help you stand apart. Every problem turns into a question of how you pitch it. And the chat never reaches the point where it says: this one isn’t worth your Saturday. If nothing can kill the idea, it isn’t validation. It’s brainstorming with better grammar.
What to use it for instead
The useful move is to stop asking for a verdict. Use it as a tool instead.
- Turn “small business owners” into one buyer. Make it grill your description. Do it until you can name a person and their job. And where they already complain about this. If you can’t reach them again and again, you don’t have a market yet.
- Rehearse the objections. Have it play the buyer and refuse you. The objections it invents are guesses. But they’re a free interview script. Five real conversations will tell you which ones were real.
- Write the questions, not the answers. Ask it to rewrite each interview question. Each one should ask about the last time the problem happened. Not what someone would do in future. That single rewrite is most of The Mom Test.
- Draft the page you’re about to test. Copy for a demand test is a good use of a language model. The copy isn’t the evidence. The strangers who respond to it are.
Notice what all four have in common. Each one gives you something to take outside the chat window.
The one-line test
Before you call it validation, ask one thing. What did it cost the person who gave it to you?
A model’s yes costs nobody a thing. A friend’s pep talk costs them nothing. A survey answer about what you might buy costs nothing. An email address costs a little. It comes from a stranger who found your page cold. A card charged before the product exists costs real money. That’s why it’s the only signal you can trust to predict anything.
That’s also why our own product doesn’t ask a model what it thinks of your idea. It goes and reads what your buyers already wrote in public. They had no idea a founder was listening. Then, where somebody in one of those threads has asked for a fix, it drafts a reply for you to post yourself. Nobody prompted those people to write. That’s what makes it as close to a free market signal as you get. The full sequence is in our guide to validating a startup idea. The pass and fail marks are in the four-gate framework.
Two questions usually follow this one. Both are cheaper to answer than you’d think. Here they are: what a real demand test costs and how long it takes. Want to compare the tools that do parts of this for you? We keep an honest one at best idea validation tools.
