Back to Insights

How Your Company OS Actually Runs Research

From "where are we?" to honest synthetic research — without letting AI declare victory for you

By Ivelin Ivanov18 min readRSS
Outer company operating loop orbiting synthetic user-research phases with a founder-held gate

🎯 Key Takeaways

  • Know where you stand before you build more. Last week we shared the Company OS. This week is how to actually run the research part of it.
  • Shipping is not the same as learning. If you still do not know who the customer is, "keep going" is often just motion. Good systems pause you for research first.
  • Look at five possible customers, not one favorite. Score them. Argue with the results. Then you pick: try again, talk to real people, or walk away.
  • A good report can still say no. In our live example, the top ideas stayed on hold. That was not failure. That was honesty.
  • Real tests come next — not in this article. After research: a small safe prototype test, and a real interest check with people. We will cover those soon.

Here is a scene most solo founders know too well.

It is Monday. You open three chat threads from last week. Each one ends with a confident next step. None of them agree with each other. Your notes app has a page titled "strategy v7." Somewhere an agent offered to "keep shipping" while you slept. And when you ask the simple question — where is the company, really? — you get energy instead of an answer.

In How to See Exactly Where Your Startup Stands, we put a name on the fix: a portable Company Operating System. Two clocks. Founder gates. An honest snapshot. Thin rules so AI cannot quietly promote your favorite story into strategy.

That piece is the map. This one is about what happens when you actually try to drive with it.

Because a beautiful operating system document has the same failure mode as a beautiful business plan: it sits there looking mature while the week runs on vibes. So in a real product repo we made the OS executable — two workflows that implement the same discipline under load:

  • company-operating-loop — the outer loop: where are we, what should run next, who has to approve
  • user-research — the deep research pass: rank several customer groups honestly, then wait for a human decision

They live in the open Totbox repo under .grok/workflows/. You will see Totbox details in the report snapshots below — home-service job PM, HVAC, cleaning — only as a live instance illustration. Steal the workflow discipline. Do not steal someone else's beachhead and call it your destiny.

The map

Company OS blueprint

Two clocks, founder gates, a weekly snapshot, evidence over pep talks — in How to See Exactly Where Your Startup Stands.

Monday cockpit

company-operating-loop

Where are we? What is honest next? Research handoff when early. Journey only with your OK.

The hard filter

user-research

Five peer groups, scorecards, dialogues that can refuse, a skeptic, then your decision.

First you need a map of where you stand. Then you need something that actually runs on Monday morning — without inventing a new process from scratch every week.

Monday morning: the outer loop

Most founders do not need another framework lecture. They need a calm first question that does not depend on mood: Where are we? What is honest to do next? Do I need to decide something before the machine keeps going?

That is what company-operating-loop is for. Not a second brain. More like a cockpit check. You run it so you do not re-explain the whole company to a new chat every time — and so you do not confuse "we generated text" with "we advanced the business."

Outer loop · company-operating-loop
  1. 1

    Look

    Re-read the living company. Say where the journey and weekly loop actually are.

  2. 2

    Hand off

    Still early on customers? Stop. Research is due — not optional busywork later.

  3. 3

    Do one honest thing

    Status, restart week, continue, or open research — never a silent strategy leap.

  4. 4

    Ask you

    Moving the slow journey still needs your explicit yes. AI recommends; you decide.

statusstartcontinueuser-researchadvance-journey
Status first. Research when you are still in early stages. Continue only when it is honest to continue. Journey advances only with your signature.

What "status" should feel like

On every run, the outer loop re-reads the living company — usually something like company/state/company-state.json plus short instance notes — and answers in plain language:

  • Where you are on the slow prove-it journey (phases 1–9 from the Company OS blueprint)
  • Where you are in this week's learning loop (stages 1–7)
  • Whether the gate is open, waiting on you, or blocked
  • What it recommends next — continue, research, hold, or ask you to approve a phase move

The important part is emotional as much as technical: the system is not allowed to "helpfully" climb the journey ladder because the slides would look better on phase 6. Journey only moves when you say so.

A week in ordinary language

Here is how it shows up in real life — not as a CLI manual, as a rhythm:

You are thinking…You ask the loop to…What should happen
"I just want an honest picture."statusA read-only snapshot: phase, stage, gate, whether research is still due. Nothing silently rewrites strategy.
"Keep going" — but you are still early on customerscontinueOften blocked on purpose. You get a handoff: run research until you agree_ready (or consciously skip — not by accident).
"I need the research pack now."user-researchStops the theater of progress and points you at the sibling workflow that ranks customer groups.
"I think this phase is done."advance-journeyPauses until you explicitly approve. One phase at a time — not a strategy hopscotch.
"Restart the weekly cycle."startThe fast loop goes back to stage 1. The slow journey does not fake a promotion just because you restarted work.

Research handoff · when continue is blocked

Triggers

  • · Live loop stage 1 or 2
  • · Journey phase 1–3
  • · No READY_FOR_REAL_WORLD.md yet

Sibling workflow

user-research

Produces ROUND reports + FOUNDER_FEEDBACK. Founder decides iterate / agree_ready / kill.

The outer loop does not quietly "research for you" inside another process. It stops, hands you the work, and waits — so the hard part cannot disappear into a status update.

The handoff rule is almost parental, and that is the point. If the live loop is still in early research/validation stages — or the journey is still in the first three phases — "continue" without a ready marker is how founders skip the hard part and call the skip momentum.

You can read the outer loop as open source here: .grok/workflows/company-operating-loop.rhai.

When the outer loop stops being enough

Status tells you the weather. It does not tell you which customers are real.

Early on, the most dangerous AI habit is not laziness — it is favorite protection. You have a story about who will buy. The model learns the story. Then it writes research that politely confirms the story, and you feel scientific because there were bullets and scores.

The user-research workflow is built to make that harder. Not by being mean for sport — by forcing several peer customer groups onto the table, scoring reward and risk, running synthetic conversations that can say no, and then putting a skeptic in the room whose job is to puncture overclaims.

user-research · seven phases per round

  1. 1

    Context

    Load thesis, OS rules, existing ICPs, prior feedback

  2. 2

    Propose

    Exactly five peer ICP candidates (not one favorite)

  3. 3

    Scorecards

    Parallel reward/risk scores per ICP

  4. 4

    Synthetic

    Parallel synthetic dialogues and verdicts

  5. 5

    Challenge

    Adversarial skeptic on ranking and promotion

  6. 6

    Report

    Write ROUND_*_report.md + FOUNDER_FEEDBACK.md

  7. 7

    Founder gate

    iterate · agree_ready · kill — founder only

Five groups get scored and stress-tested side by side. Then a human — you — decides what happens next. That last step is not optional.

What a round actually feels like

Think of one round as a disciplined week of thinking — compressed:

  1. Context — Remember the thesis, the OS rules, what you already claimed, and any notes you left yourself last time. No inventing files that are not there.
  2. Propose five peers — Not one beloved beachhead and four straw men. Five real candidates. Your current favorite stays on the slate as a peer, not a crowned king.
  3. Scorecards — Pain, will they pay or act, can you reach them, channel fit, risk, time-to-signal, kill criteria, and the brutal field: promote_lean.
  4. Synthetic dialogues — Composite people who can refuse. Status quo. Trust barriers. Reasons to say no. A verdict that can land weak even when the pain sounds high.
  5. Challenge — A separate skeptic attacks ranking soundness, missing challengers, and whether anyone is allowed to claim ready_for_real_world.
  6. Report — A round document under research/icps/ROUND_*_report.md, plus a feedback template waiting for you.
  7. Your gate — You fill FOUNDER_FEEDBACK.md with iterate, agree_ready, or kill. Until you write a decision, the process is not finished. That is a feature.

Founder gate · FOUNDER_FEEDBACK.md

After every report the AI waits. You write one decision word — or the week stays open on purpose.

iterate

Keep filtering. Drop/add ICPs, dispute scores, re-run next round.

agree_ready

Write READY_FOR_REAL_WORLD.md. Outer loop may continue past research stages.

kill

Stop protecting this thesis or slate. Record why.

hold (AI)

promote_lean stays hold. ready_for_real_world stays false on synthetic alone.

The AI can write a beautiful report. It cannot declare victory. Promote and "ready for real world" stay closed until you say otherwise.

The rules that keep you honest

These are not bureaucratic decorations. They are the difference between research and cosplay:

  • Every pack is labeled SYNTHETIC ONLY. It is not product-market fit. It is not a letter of intent. It is not twenty real households.
  • AI ready_for_real_world stays false unless the whole system and the founder open that door on purpose.
  • promote_lean stays hold until synthetic evidence, real evidence, and manageable risk line up. The model does not get to promote your primary focus because the narrative was pretty.
  • The file READY_FOR_REAL_WORLD.md appears only when you choose agree_ready.
  • Public-safe composites only. No real names, emails, or private addresses turned into fake proof.

If you want to see the workflow itself: .grok/workflows/user-research.rhai.

What it looked like when we actually ran it

Theory is cheap. So here are condensed snapshots from a real public instance — Totbox — after the research workflow ran. The numbers and customer-group names come from ROUND_1_report.md and ROUND_1-r2_report.md.

Read them as a worked example of discipline under load. Do not read them as a suggestion that your startup should become an HVAC company.

Round 1 — ranking without a victory lap

The first pack did something founders rarely do voluntarily: it demoted hard. One group earned a strong synthetic fit. Four others did not. Every seat stayed on hold. The AI refused to call the company ready for real-world proof.

Live instance · Totbox · ROUND 1 · SYNTHETIC ONLY
ready_for_real_world=falseranking_sound=trueROUND_1_report.md

No primary-focus promotion — every seat stayed promote_lean=hold. Only dual-income earned strong_fit, and even that meant "test first," not "we won."

#ICPVerdictpromote
1Busy dual-income household (HVAC + cleaning chore PM)strong_fithold
2Remote / multi-site service coordinatorweak_fithold
3Seasonal tree / arborist decision householdsweak_fithold
4Local HVAC/multi-trade operator — Agentic Ready (PAYER)weak_fithold
5SMS/phone-native grounds services householdsweak_fithold

Scorecard composites (AI judgment · not multi-N proof)

ICPpainpay_or_actchannelriskTTSverdict
Dual-income0.820.580.880.580.74strong_fit
Remote multi-site0.850.680.820.700.55weak_fit
Seasonal tree0.680.580.680.720.55weak_fit
Operator payer0.720.580.600.740.38weak_fit
SMS grounds0.580.420.380.720.58weak_fit
From Totbox research/icps/ROUND_1_report.md. Even the winner of this round carried a soft pay_or_act (~0.58). The report refused to dress that up as promote.

The decision trace for that pack (traces/decisions/2026-user-research-round-1.md) is almost boring on purpose: publish the synthetic filter; treat dual-income as a test priority, not a promoted destiny; keep promote_lean=hold; wait for the founder. That boredom is honesty. Useful ranking. No champagne.

Round 1-r2 — when the founder says "try again, harder"

Here is the human part people skip in demos: after Round 1, nothing auto-closes. In this instance the founder path was iterate — keep the strong dual-income seat, throw out the weak seats, force new challengers with different buyer seats and cadences. Then the whole filter ran again.

Live instance · Totbox · ROUND 1-r2 · SYNTHETIC ONLY
ready_for_real_world=falseranking_sound=falseROUND_1-r2_report.md

After Round 1 the founder said iterate: keep dual-income, replace every weak seat with a real challenger. The skeptic then reordered the board — single decision-maker first — and refused ranking_sound if dual stayed #1 only because it won last time.

#Recommended rankVerdictpromote
1Single decision-maker household (recurring home-service PM)strong_fithold
2Busy dual-income household (HVAC + cleaning chore PM)strong_fithold
3New-homeowner / move-in concurrent job burstweak_fithold
4Small landlord / tenant-split home-service coordinatorweak_fithold
5Reactive emergency HVAC repair householdsweak_fithold

Scorecard composites (r2)

ICPpainpay_or_actchannelreachriskTTSverdict
Single-DM0.760.620.870.720.540.78strong_fit
Dual-income0.820.580.880.680.580.74strong_fit
Move-in burst0.860.560.760.560.720.62weak_fit
Small landlord0.810.630.750.480.730.50weak_fit
Emergency HVAC0.880.600.500.500.780.52weak_fit

What the skeptic refused to claim

  • · "This group is the business now" (no PMF for any ICP)
  • · Measured fewer touchpoints than the messy status quo
  • · Proven willingness to pay (every pay_or_act score stayed mid/soft)
  • · READY_FOR_REAL_WORLD stamped by the AI alone
From Totbox research/icps/ROUND_1-r2_report.md. Two strong fits this time — and still every seat on hold. The recommended first real test shifted toward single decision-maker, with dual-income as the A/B, not the automatic crown.

What got better: two strong_fits instead of one, and a cleaner A/B between single decision-maker households and dual-income coordination chaos. The skeptic even refused ranking_sound when Round 1's favorite tried to stay #1 by retention narrative alone.

What did not get better in the fake way founders crave: ready_for_real_world stayed false. Every promote_lean stayed hold. Nobody got to claim product-market fit because the second report was longer.

And at the time of writing, FOUNDER_FEEDBACK.md still has an open decision line after r2. That is not a bug. An empty gate is not a secret agree_ready. The system waits for a person.

How this maps back to the two clocks

If you read the blueprint article, you already know the slow journey and the fast weekly loop. Here is the same idea without the machinery cosplay:

What the OS is trying to protectWhat the workflows actually do
Early journey phases still about finding truthOuter loop insists on research handoff; user-research produces ranked evidence
Weekly stages still in research / validationSame block — do not "continue" into build theater yet
Journey only moves with founder judgmentadvance-journey requires explicit approval
Research only opens with founder judgmentiterate / agree_ready / kill in FOUNDER_FEEDBACK
Evidence over favorite storiesSkeptic section, hold-by-default promote_lean, decision traces under traces/
"Where are we?" in under two minutesOuter loop status re-reads live state before it recommends anything

What we are not covering yet

Synthetic research is a sharp filter. It is not the whole movie.

After you agree_ready — or once the slate is stable enough to test in parallel — two more chapters matter. We will give them their own article soon:

  1. A simulated sandboxed end-to-end prototype feasibility test — a tiny slice of product, clear pass/fail numbers, no burning real households just to feel productive.
  2. A parallel real user interest test campaign — light-touch real signals on the one or two groups the filter prioritizes: conversations, waitlists, shadow jobs, pay-or-act. Not a multi-segment GTM parade because the slides wanted five logos.

This piece stops at the outer loop and the research filter on purpose. If you cannot yet survive an honest synthetic ranking without promoting your favorite, real-world theater will only make the self-deception more expensive.

If you want to look under the hood

What you wantWhere to go
The Company OS mapHow to See Exactly Where Your Startup Stands · docs/company-os/
Outer loop.grok/workflows/company-operating-loop.rhai
User-research workflow.grok/workflows/user-research.rhai
Round reports (illustration only)research/icps/

Closing

AI will happily write you a research novel that ends in applause. Markets will not read it.

The outer loop exists so you can ask "where are we?" without cosplay. The user-research workflow exists so early ranking can hurt a little — before real people and real weeks pay the bill for a protected favorite.

Together they are the executable half of the system in How to See Exactly Where Your Startup Stands.

Get status. Accept the handoff. Rank five peers. Let the skeptic talk. Then write one honest word: iterate, agree_ready, or kill.

You still keep the keys.

FAQ

How does this relate to the Company OS article?

That piece is the map: two clocks, gates, snapshot, evidence rules. This piece is what running the map feels like — outer loop + user-research — with real Totbox report snapshots as illustration only.

When does the outer loop force research?

When you are still in early loop stages (1–2) or early journey phases (1–3), and you have not created a founder READY_FOR_REAL_WORLD.md yet. "Continue" stays blocked until then unless you intentionally skip — and skipping should feel like a decision, not a glitch.

What are the user-research phases, simply?

Load context, propose five peer groups, score them, talk to synthetic users who can refuse, attack the ranking with a skeptic, write the report, then founder gate: iterate / agree_ready / kill.

Do these Totbox tables prove product-market fit?

No. They are labeled SYNTHETIC ONLY, with promote_lean=hold and ready_for_real_world=false. Steal the discipline. Do not steal the market.

What comes after this research filter?

Two follow-ups for a separate article: simulated sandboxed end-to-end prototype feasibility tests, and a parallel real user interest test campaign.

Related insights