ProofMachine is an AI idea validation tool that answers one question first: has anyone already said, in public, that they have this problem? It searches the forums and review threads your buyers post in, quotes them word for word with a link back to each thread, and returns a verdict that is allowed to be no.
The useful part of a validation report is the part that disagrees with you.
Most tools in this category can’t produce it. You paste in an idea, a language model reasons about it, and a score comes back somewhere north of seven. That number is a summary of how the idea reads, not of whether anyone wants it. It is the same answer you’d get from a friend who doesn’t want to upset you, delivered faster, and it is the same thing that happens when you ask ChatGPT directly. We compared six of these tools and five of the six work this way.
This page describes the other approach, and where it runs out.
What a run actually does
You describe an idea in a sentence or two. The run then goes looking for people who have already complained about the problem it solves.
That search is the whole product. It reads public threads: subreddits, review sites, niche forums, the places where someone annoyed enough to type three paragraphs at midnight has already typed them. It pulls the passages where a real person describes the pain, and it keeps the URL of the thread each one came from.
A run takes about four minutes. It happens on our side, so you can close the tab.
What comes back
Four checks, each answered against the evidence found rather than in the abstract:
- The problem is real. People describe it unprompted, in their own words, and here are the threads.
- You can find these people. They cluster somewhere reachable, and here is where.
- They’d pay to fix it. Someone said so, or bought something adjacent. This is the check that fails most often.
- The field is wide open. Or it isn’t, and here are the six tools already doing it.
Underneath sits the verdict, which is a sentence with a “because” in it. Above the checks sit the numbers a run turned up: how many live threads, how many people actively asking for a fix right now, how many of the businesses already in this space clear the side-project line.
That last figure comes from a set of about 145,000 anonymised business records rather than from the model’s guess. The records carry no company name, website or contact details of any kind — they exist to produce a range, and a range is what you get back. Category base rates are unglamorous and they resolve arguments. If four in ten businesses in a category report meaningful revenue, that is a different world from one in forty, and no amount of reasoning about your idea moves the number.
Claims that arrived without a source are penned separately from claims that carry one. You can always see which is which.
Where the evidence comes from
Public writing by people who did not know they were being read for market research, which is exactly what makes it worth reading. Nobody performs for a survey they don’t know they’re taking.
Every quote in a report carries a link. Click it and you land on the thread. If a quote looks too convenient, check it, and if the thread doesn’t say what the report says it says, that is a bug we want to hear about.
There is a fair objection here, so let’s state it. Forum complaints skew toward the annoyed and the online. People who quietly tolerate a problem don’t post about it, and neither do the people who already pay someone to solve it. Evidence of complaint is evidence that a problem exists, not proof that a market does. The validation framework covers the gates that come after this one, and the honest ones involve asking someone for money.
What it won’t tell you
It won’t tell you whether you should build it. That reduces to your appetite and your runway, and no tool can see either.
It won’t find evidence in a market that doesn’t write things down. Some buyers, particularly in trades and enterprise procurement, transact almost entirely offline. A run against that kind of idea comes back thin, and thin is the honest answer rather than a padded one.
It won’t stop you from building the wrong thing carefully. No market need is the most commonly cited cause of startup failure, and it is a cause that survives good engineering.
It won’t replace talking to people. It shortens the list of who to talk to and gives you their words to open with, which is a real saving of weeks. It is not the conversation.
And it won’t reward you for a good pitch. Rewriting your idea more persuasively changes nothing about what strangers posted last year.
Who should skip it
If you have already shipped and have users, you have better data than this. Read your own support tickets.
If your idea depends on a market shift that hasn’t happened yet, there is nothing to find. Absence of complaint about a problem nobody has yet is not a signal, and a run will say so.
If you want a document to show an investor, buy a report generator instead. They are cheaper, faster, and prettier, and several are free. A verdict that says no is worth something to you and nothing to a pitch deck.
Running one
Sign in, describe the idea, and come back in about four minutes. New accounts get one full run free, which is enough to see the whole output shape and decide whether the honest version of the answer is what you wanted. After that, runs cost credits; the pricing page has the current numbers.
Bring the fuzziest version of the idea. Precision at this stage is usually invented, and the search works better against a plain description of the problem than against a positioning statement.