Here is a scene most solo founders know too well.
It is Monday. You open three chat threads from last week. Each one ends with a confident next step. None of them agree with each other. Your notes app has a page titled "strategy v7." Somewhere an agent offered to "keep shipping" while you slept. And when you ask the simple question — where is the company, really? — you get energy instead of an answer.
In How to See Exactly Where Your Startup Stands, we put a name on the fix: a portable Company Operating System. Two clocks. Founder gates. An honest snapshot. Thin rules so AI cannot quietly promote your favorite story into strategy.
That piece is the map. This one is about what happens when you actually try to drive with it.
Because a beautiful operating system document has the same failure mode as a beautiful business plan: it sits there looking mature while the week runs on vibes. So in a real product repo we made the OS executable — two workflows that implement the same discipline under load:
company-operating-loop— the outer loop: where are we, what should run next, who has to approveuser-research— the deep research pass: rank several customer groups honestly, then wait for a human decision
They live in the open Totbox repo under .grok/workflows/. You will see Totbox details in the report snapshots below — home-service job PM, HVAC, cleaning — only as a live instance illustration. Steal the workflow discipline. Do not steal someone else's beachhead and call it your destiny.
The map
Company OS blueprint
Two clocks, founder gates, a weekly snapshot, evidence over pep talks — in How to See Exactly Where Your Startup Stands.
Monday cockpit
company-operating-loop
Where are we? What is honest next? Research handoff when early. Journey only with your OK.
The hard filter
user-research
Five peer groups, scorecards, dialogues that can refuse, a skeptic, then your decision.
Monday morning: the outer loop
Most founders do not need another framework lecture. They need a calm first question that does not depend on mood: Where are we? What is honest to do next? Do I need to decide something before the machine keeps going?
That is what company-operating-loop is for. Not a second brain. More like a cockpit check. You run it so you do not re-explain the whole company to a new chat every time — and so you do not confuse "we generated text" with "we advanced the business."
- 1
Look
Re-read the living company. Say where the journey and weekly loop actually are.
- 2
Hand off
Still early on customers? Stop. Research is due — not optional busywork later.
- 3
Do one honest thing
Status, restart week, continue, or open research — never a silent strategy leap.
- 4
Ask you
Moving the slow journey still needs your explicit yes. AI recommends; you decide.
What "status" should feel like
On every run, the outer loop re-reads the living company — usually something like company/state/company-state.json plus short instance notes — and answers in plain language:
- Where you are on the slow prove-it journey (phases 1–9 from the Company OS blueprint)
- Where you are in this week's learning loop (stages 1–7)
- Whether the gate is open, waiting on you, or blocked
- What it recommends next — continue, research, hold, or ask you to approve a phase move
The important part is emotional as much as technical: the system is not allowed to "helpfully" climb the journey ladder because the slides would look better on phase 6. Journey only moves when you say so.
A week in ordinary language
Here is how it shows up in real life — not as a CLI manual, as a rhythm:
| You are thinking… | You ask the loop to… | What should happen |
|---|---|---|
| "I just want an honest picture." | status | A read-only snapshot: phase, stage, gate, whether research is still due. Nothing silently rewrites strategy. |
| "Keep going" — but you are still early on customers | continue | Often blocked on purpose. You get a handoff: run research until you agree_ready (or consciously skip — not by accident). |
| "I need the research pack now." | user-research | Stops the theater of progress and points you at the sibling workflow that ranks customer groups. |
| "I think this phase is done." | advance-journey | Pauses until you explicitly approve. One phase at a time — not a strategy hopscotch. |
| "Restart the weekly cycle." | start | The fast loop goes back to stage 1. The slow journey does not fake a promotion just because you restarted work. |
Research handoff · when continue is blocked
Triggers
- · Live loop stage 1 or 2
- · Journey phase 1–3
- · No
READY_FOR_REAL_WORLD.mdyet
Sibling workflow
user-research
Produces ROUND reports + FOUNDER_FEEDBACK. Founder decides iterate / agree_ready / kill.
The handoff rule is almost parental, and that is the point. If the live loop is still in early research/validation stages — or the journey is still in the first three phases — "continue" without a ready marker is how founders skip the hard part and call the skip momentum.
You can read the outer loop as open source here: .grok/workflows/company-operating-loop.rhai.
When the outer loop stops being enough
Status tells you the weather. It does not tell you which customers are real.
Early on, the most dangerous AI habit is not laziness — it is favorite protection. You have a story about who will buy. The model learns the story. Then it writes research that politely confirms the story, and you feel scientific because there were bullets and scores.
The user-research workflow is built to make that harder. Not by being mean for sport — by forcing several peer customer groups onto the table, scoring reward and risk, running synthetic conversations that can say no, and then putting a skeptic in the room whose job is to puncture overclaims.
user-research · seven phases per round
- 1
Context
Load thesis, OS rules, existing ICPs, prior feedback
- 2
Propose
Exactly five peer ICP candidates (not one favorite)
- 3
Scorecards
Parallel reward/risk scores per ICP
- 4
Synthetic
Parallel synthetic dialogues and verdicts
- 5
Challenge
Adversarial skeptic on ranking and promotion
- 6
Report
Write ROUND_*_report.md + FOUNDER_FEEDBACK.md
- 7
Founder gate
iterate · agree_ready · kill — founder only
What a round actually feels like
Think of one round as a disciplined week of thinking — compressed:
- Context — Remember the thesis, the OS rules, what you already claimed, and any notes you left yourself last time. No inventing files that are not there.
- Propose five peers — Not one beloved beachhead and four straw men. Five real candidates. Your current favorite stays on the slate as a peer, not a crowned king.
- Scorecards — Pain, will they pay or act, can you reach them, channel fit, risk, time-to-signal, kill criteria, and the brutal field:
promote_lean. - Synthetic dialogues — Composite people who can refuse. Status quo. Trust barriers. Reasons to say no. A verdict that can land weak even when the pain sounds high.
- Challenge — A separate skeptic attacks ranking soundness, missing challengers, and whether anyone is allowed to claim
ready_for_real_world. - Report — A round document under
research/icps/ROUND_*_report.md, plus a feedback template waiting for you. - Your gate — You fill FOUNDER_FEEDBACK.md with
iterate,agree_ready, orkill. Until you write a decision, the process is not finished. That is a feature.
Founder gate · FOUNDER_FEEDBACK.md
After every report the AI waits. You write one decision word — or the week stays open on purpose.
Keep filtering. Drop/add ICPs, dispute scores, re-run next round.
Write READY_FOR_REAL_WORLD.md. Outer loop may continue past research stages.
Stop protecting this thesis or slate. Record why.
promote_lean stays hold. ready_for_real_world stays false on synthetic alone.
The rules that keep you honest
These are not bureaucratic decorations. They are the difference between research and cosplay:
- Every pack is labeled SYNTHETIC ONLY. It is not product-market fit. It is not a letter of intent. It is not twenty real households.
- AI
ready_for_real_worldstays false unless the whole system and the founder open that door on purpose. promote_leanstays hold until synthetic evidence, real evidence, and manageable risk line up. The model does not get to promote your primary focus because the narrative was pretty.- The file
READY_FOR_REAL_WORLD.mdappears only when you chooseagree_ready. - Public-safe composites only. No real names, emails, or private addresses turned into fake proof.
If you want to see the workflow itself: .grok/workflows/user-research.rhai.
What it looked like when we actually ran it
Theory is cheap. So here are condensed snapshots from a real public instance — Totbox — after the research workflow ran. The numbers and customer-group names come from ROUND_1_report.md and ROUND_1-r2_report.md.
Read them as a worked example of discipline under load. Do not read them as a suggestion that your startup should become an HVAC company.
Round 1 — ranking without a victory lap
The first pack did something founders rarely do voluntarily: it demoted hard. One group earned a strong synthetic fit. Four others did not. Every seat stayed on hold. The AI refused to call the company ready for real-world proof.
No primary-focus promotion — every seat stayed promote_lean=hold. Only dual-income earned strong_fit, and even that meant "test first," not "we won."
| # | ICP | Verdict | promote |
|---|---|---|---|
| 1 | Busy dual-income household (HVAC + cleaning chore PM) | strong_fit | hold |
| 2 | Remote / multi-site service coordinator | weak_fit | hold |
| 3 | Seasonal tree / arborist decision households | weak_fit | hold |
| 4 | Local HVAC/multi-trade operator — Agentic Ready (PAYER) | weak_fit | hold |
| 5 | SMS/phone-native grounds services households | weak_fit | hold |
Scorecard composites (AI judgment · not multi-N proof)
| ICP | pain | pay_or_act | channel | risk | TTS | verdict |
|---|---|---|---|---|---|---|
| Dual-income | 0.82 | 0.58 | 0.88 | 0.58 | 0.74 | strong_fit |
| Remote multi-site | 0.85 | 0.68 | 0.82 | 0.70 | 0.55 | weak_fit |
| Seasonal tree | 0.68 | 0.58 | 0.68 | 0.72 | 0.55 | weak_fit |
| Operator payer | 0.72 | 0.58 | 0.60 | 0.74 | 0.38 | weak_fit |
| SMS grounds | 0.58 | 0.42 | 0.38 | 0.72 | 0.58 | weak_fit |
The decision trace for that pack (traces/decisions/2026-user-research-round-1.md) is almost boring on purpose: publish the synthetic filter; treat dual-income as a test priority, not a promoted destiny; keep promote_lean=hold; wait for the founder. That boredom is honesty. Useful ranking. No champagne.
Round 1-r2 — when the founder says "try again, harder"
Here is the human part people skip in demos: after Round 1, nothing auto-closes. In this instance the founder path was iterate — keep the strong dual-income seat, throw out the weak seats, force new challengers with different buyer seats and cadences. Then the whole filter ran again.
After Round 1 the founder said iterate: keep dual-income, replace every weak seat with a real challenger. The skeptic then reordered the board — single decision-maker first — and refused ranking_sound if dual stayed #1 only because it won last time.
| # | Recommended rank | Verdict | promote |
|---|---|---|---|
| 1 | Single decision-maker household (recurring home-service PM) | strong_fit | hold |
| 2 | Busy dual-income household (HVAC + cleaning chore PM) | strong_fit | hold |
| 3 | New-homeowner / move-in concurrent job burst | weak_fit | hold |
| 4 | Small landlord / tenant-split home-service coordinator | weak_fit | hold |
| 5 | Reactive emergency HVAC repair households | weak_fit | hold |
Scorecard composites (r2)
| ICP | pain | pay_or_act | channel | reach | risk | TTS | verdict |
|---|---|---|---|---|---|---|---|
| Single-DM | 0.76 | 0.62 | 0.87 | 0.72 | 0.54 | 0.78 | strong_fit |
| Dual-income | 0.82 | 0.58 | 0.88 | 0.68 | 0.58 | 0.74 | strong_fit |
| Move-in burst | 0.86 | 0.56 | 0.76 | 0.56 | 0.72 | 0.62 | weak_fit |
| Small landlord | 0.81 | 0.63 | 0.75 | 0.48 | 0.73 | 0.50 | weak_fit |
| Emergency HVAC | 0.88 | 0.60 | 0.50 | 0.50 | 0.78 | 0.52 | weak_fit |
What the skeptic refused to claim
- · "This group is the business now" (no PMF for any ICP)
- · Measured fewer touchpoints than the messy status quo
- · Proven willingness to pay (every pay_or_act score stayed mid/soft)
- ·
READY_FOR_REAL_WORLDstamped by the AI alone
What got better: two strong_fits instead of one, and a cleaner A/B between single decision-maker households and dual-income coordination chaos. The skeptic even refused ranking_sound when Round 1's favorite tried to stay #1 by retention narrative alone.
What did not get better in the fake way founders crave: ready_for_real_world stayed false. Every promote_lean stayed hold. Nobody got to claim product-market fit because the second report was longer.
And at the time of writing, FOUNDER_FEEDBACK.md still has an open decision line after r2. That is not a bug. An empty gate is not a secret agree_ready. The system waits for a person.
How this maps back to the two clocks
If you read the blueprint article, you already know the slow journey and the fast weekly loop. Here is the same idea without the machinery cosplay:
| What the OS is trying to protect | What the workflows actually do |
|---|---|
| Early journey phases still about finding truth | Outer loop insists on research handoff; user-research produces ranked evidence |
| Weekly stages still in research / validation | Same block — do not "continue" into build theater yet |
| Journey only moves with founder judgment | advance-journey requires explicit approval |
| Research only opens with founder judgment | iterate / agree_ready / kill in FOUNDER_FEEDBACK |
| Evidence over favorite stories | Skeptic section, hold-by-default promote_lean, decision traces under traces/ |
| "Where are we?" in under two minutes | Outer loop status re-reads live state before it recommends anything |
What we are not covering yet
Synthetic research is a sharp filter. It is not the whole movie.
After you agree_ready — or once the slate is stable enough to test in parallel — two more chapters matter. We will give them their own article soon:
- A simulated sandboxed end-to-end prototype feasibility test — a tiny slice of product, clear pass/fail numbers, no burning real households just to feel productive.
- A parallel real user interest test campaign — light-touch real signals on the one or two groups the filter prioritizes: conversations, waitlists, shadow jobs, pay-or-act. Not a multi-segment GTM parade because the slides wanted five logos.
This piece stops at the outer loop and the research filter on purpose. If you cannot yet survive an honest synthetic ranking without promoting your favorite, real-world theater will only make the self-deception more expensive.
If you want to look under the hood
| What you want | Where to go |
|---|---|
| The Company OS map | How to See Exactly Where Your Startup Stands · docs/company-os/ |
| Outer loop | .grok/workflows/company-operating-loop.rhai |
| User-research workflow | .grok/workflows/user-research.rhai |
| Round reports (illustration only) | research/icps/ |
Closing
AI will happily write you a research novel that ends in applause. Markets will not read it.
The outer loop exists so you can ask "where are we?" without cosplay. The user-research workflow exists so early ranking can hurt a little — before real people and real weeks pay the bill for a protected favorite.
Together they are the executable half of the system in How to See Exactly Where Your Startup Stands.
Get status. Accept the handoff. Rank five peers. Let the skeptic talk. Then write one honest word: iterate, agree_ready, or kill.
You still keep the keys.
FAQ
How does this relate to the Company OS article?
That piece is the map: two clocks, gates, snapshot, evidence rules. This piece is what running the map feels like — outer loop + user-research — with real Totbox report snapshots as illustration only.
When does the outer loop force research?
When you are still in early loop stages (1–2) or early journey phases (1–3), and you have not created a founder READY_FOR_REAL_WORLD.md yet. "Continue" stays blocked until then unless you intentionally skip — and skipping should feel like a decision, not a glitch.
What are the user-research phases, simply?
Load context, propose five peer groups, score them, talk to synthetic users who can refuse, attack the ranking with a skeptic, write the report, then founder gate: iterate / agree_ready / kill.
Do these Totbox tables prove product-market fit?
No. They are labeled SYNTHETIC ONLY, with promote_lean=hold and ready_for_real_world=false. Steal the discipline. Do not steal the market.
What comes after this research filter?
Two follow-ups for a separate article: simulated sandboxed end-to-end prototype feasibility tests, and a parallel real user interest test campaign.
Related insights
- How to See Exactly Where Your Startup Stands — Company OS blueprint this article implements
- Synthetic Users & Outcome Pricing — broader synthetic research + pricing playbook
- Agent Skills & MCP distribution — how founders distribute agent capabilities
- Non-tech founder AI agent team — building with AI without a traditional eng org
- Totbox on GitHub — workflows + research pack (instance illustration)
