A chatbot can say it would buy your product tomorrow. That is not a customer. It is a sentence.
That is why Bootstrap OS just changed. The playbook is for people starting a company alone with AI. We keep it current. When new evidence shows up about using AI to find customers, we read it. If it holds, we change the rules. If it is a vendor ad, we leave it on the floor.
Version 2.8.3 is on github.com/ivelin/bootstrap (PR #5).
You will hear three kinds of yes. Someone said they would buy. A chatbot guessed they would buy. A real person spent time or money. The first two are notes. The third is evidence. When they disagree, believe the person who showed up. A spoken “I would buy” can sit on the board. It cannot graduate the idea.
Cheap AI research is still useful. You can ask the same question two ways and see which story falls apart. You can change one thing — price, or time, or the thing they already use — and see if the yes survives. What you cannot do is open with “what would you pay?” and treat the number as demand. If you change price, copy, and persona at once, you learned fog.
We did not invent that from a slogan. In 2024, Bisbee and colleagues asked ChatGPT-3.5 to fill out surveys like a person. The answers looked too sure. Ask the same thing later and they moved. We kept the warning: chatbot surveys are jumpy. We did not keep a fantasy that you can replace your customers.
A later paper from Brand, Israeli, and Ngwe at Harvard (revised April 2026) taught a model on people shopping for laptops. It did not transfer to tablets. Asking the model how many dollars it would pay was useless. They still used simple A-or-B choices that included prices. So we did not write a law that says never mention a dollar. We wrote: do not treat a chat price as demand.
There is a paper that goes the other way. Maier ranks “I’d buy it” talk and treats that ranking as useful. A ranking of talk is still talk. It is not time. It is not money.
The ads are louder than the papers. “95% of papers agree.” “We’re 90% accurate.” “The AI said they’d buy, so we have PMF.” “85% accurate twins.” Those last ones usually mean something much narrower than “this is your market.” Newer models are not a free upgrade either. A 2023 score can rot. A pattern like “the answers are all the same” still matters until someone runs that test on today’s model.
We review the new facts as they land and we filter the hype. Aaru is one public record we read — a hit, a miss, an EY recreation, a gap on heirs. They are trying to model what people do, not what they say. The public tests go both ways. Semafor reported that they called the 2024 New York-16 primary closely — Latimer 58.7, Bowman 41.3, within a few hundred votes. They missed the presidential race that November. They had Harris winning the Electoral College and lost Michigan, Nevada, Pennsylvania, and Wisconsin. They missed Ohio’s Senate race too. EY later published a wealth-survey recreation: 3,600 people, 53 questions, a “median correlation of 90%.” The extra decimals you see on Aaru’s site are theirs, not EY’s. On heirs keeping the family advisor, the survey said 82%, the sim said 43%, and a separate industry study sat around 20–27%. EY notes the sim was closer to observed retention than the survey was, and still too high. Ned Koh (cofounder) told Fortune that their sales pitch is “Do not trust us. Do not trust our model.” They want you to test them on a decision you already know.
That is why we keep the playbook honest. If you have no real customers, write none yet. Do not invent a warm start. Do not borrow a list from a last job. Those people are not yours. A made-up persona is a sketch. It does not fill the empty box. Do not seed a persona from a demographic one-liner. Seed from traces of what they already do, pay for, or said in public. “Pretend you are 34” is the weak case.
This week, get one real hour or one real dollar. Ask someone who has the problem for twenty minutes. Or ask them to pay something, or wait in a way that costs them. If they will not, write that down. That no is more useful than a chat that said yes.
The old books still apply. The Mom Test, Continuous Discovery Habits, Testing Business Ideas. Compliments were never demand. Cheap chat does not retire that. I already wrote the research loop in How Bootstrap OS Actually Runs Research. This is the update on the sentence everyone wants to skip.
The model can be a harsh critic of your copy. It cannot sign a check.
