If you're a solo founder right now, you already know the feeling.
Your chat threads are full of clever half-answers. Your notes app has three competing "strategies." The product is either overbuilt or underbuilt, and you can't tell which. Someone on the internet is shipping faster than you — or at least posting like they are. And when you ask an AI assistant how the startup is doing, you get a pep talk instead of a dashboard.
This article is about fixing that.
Not with another framework cult. Not with a forty-page operating plan nobody opens twice. With a simple Company Operating System: a way to see where you are, what evidence you actually have, and what you should decide next — clear enough that a teenager can follow it, honest enough that weak ideas can't hide behind activity.
I use this system with founders I mentor (Founder Institute, SCORE, and similar). The portable blueprint lives in the open as a mentorship pack inside a real product repo — github.com/ivelin/totboxapp — under docs/company-os/. You can point your AI at that pack and say: take the operating system, not the product thesis, and apply it to my startup.
Portable template vs one live instance
Two different things sit next to each other in that open Totbox repo. Mixing them up is how mentees accidentally "adopt HVAC home services" as their business.
Portable template
Steal this: principles, two clocks, gates, scores, thin AI rules
In the Totbox open repo: docs/company-os/ — especially operating-system.md, live-runtime.md, ai-instructions.md
Live instance (Totbox)
Illustration only: their thesis, beachhead, filled state, gap analysis
applied-here.md, instance/, company/state/ — not your default market or stack
| Layer | What it is | What you steal |
|---|---|---|
| Portable template | How any solo founder should run the company: principles, 9 journey phases, 7-stage weekly loop, gates, scores, thin AI rules | Process and control |
| Live instance | One company's filled-in state: their thesis, ICPs, scores, open questions, code, market | Only as a worked example |
Golden rule: Copy process and control. Do not copy another founder's market, ICP list, feature roadmap, or "current hypothesis" as yours.
Later we fill the same blank board with Totbox's public instance numbers (from company/state/company-state.json) so you can see the method under load — still labeled as illustration only.
True north
AI is the greatest leverage tool entrepreneurs have ever had.
Today's AI is the worst it will ever be.
That is not a slogan. It is a planning assumption.
Treat AI as a slightly faster intern and you get slightly faster mediocrity. Treat it as a co-founder who never sleeps — and who still needs your judgment on truth, risk, and taste — and you get something closer to unfair speed.
You supply the insight. AI supplies the speed.
Skillful tool use is table stakes now. Real advantage comes from noticing opportunities most people (and current models) still undervalue or dismiss. Plenty of good companies win on a simple truth others thought was too messy, too small, or not worth the effort.
Solo is a smart starting position because you can move at AI speed without waiting for permission. It is not a religion. When you find force multipliers — people who use AI even better in an area of shared passion — weigh the risks of bringing them on against the risks of staying alone, and decide with eyes open.
Blueprint vs live system
Founders constantly confuse these:
| Layer | What it is | Failure if you confuse them |
|---|---|---|
| Blueprint | How a company OS should work: phases, gates, rules | Planning theater — beautiful docs, no proof |
| Live runtime | Actual state, scores, traces, open questions, results | Automation theater — agents running with no control |
| Product | What customers touch | Building features while the business hypothesis is still fog |
Writing a plan is not proving a business. Running agents is not running a company. You need both a clear blueprint and a live system you can query in plain language.
The two clocks
Most founders run one vague clock called "progress." That is how you stay busy and lost at the same time. A useful company OS runs two clocks:
Where is the company on the prove-it path? You gate every move.
- 1Thesis + ICPs
- 2Success defs
- 3Synthetic research
- 4Real research
- 5Tiny architecture
- 6Build (EDD)
- 7Real users
- 8Learn / improve
- 9Scale (after proof)
Advance · Iterate · Hold · Kill — founder only
What did we learn this week? Many cycles can live inside one journey phase.
- 1Synthetic research
- 2Validation / concept
- 3Product building
- 4Testing
- 5Evaluation
- 6Real feedback
- 7Memory update
Stage 7 writes memory, then wraps back to stage 1
Bootstrap journey (slow, you gate it)
| # | Simple name | Exit signal |
|---|---|---|
| 1 | Form thesis and list possible customer groups | Written thesis + ≥3 ICP candidates |
| 2 | Define what success looks like for each group | Clear “done means…” per group |
| 3 | Synthetic research and first validation | Ranked groups with evidence notes |
| 4 | Real-world research + monetization stress | Conversations/tests; weak groups demoted |
| 5 | Design the simplest system that can test the winner | Tiny slice + pass/fail rules + human gates |
| 6 | Build a tiny slice and test it hard | Slice runs end-to-end; gate scores |
| 7 | Try it with real or realistic users | Observed behavior, not only compliments |
| 8 | Learn from what happens and improve | Decision traces + score movement |
| 9 | Grow only after it clearly works | Proof of value or payment; then expand |
Monetization stress (will they pay / which path?) lives mainly in phases 4–5 and in ongoing reward/risk scorecards — not as a separate tenth phase.
Slow clock = where is the company on the prove-it path? Fast clock = what did we learn this week? Many fast loops inside one slow phase is healthy. Advancing the slow clock because you are tired of the phase — or because the slides look better on phase 9 — is not.
The snapshot every founder should fill in
Steal this. Print it. Paste it into your notes. Make your AI maintain it.
Your company
Hypothesis: [one sentence — subject to evidence]
Journey · slow
N / 9
[phase name]
Live loop · fast
M / 7
[stage name]
Top risks / open questions
- • …
- • …
System recommendation
Advance / Iterate / Hold / Kill — founder decides
Copy-paste version for notes or AI memory:
┌─────────────────────────────────────────────────────────────────┐
│ YOUR COMPANY · operating system instance │
│ Hypothesis: [one sentence — subject to evidence] │
├─────────────────────────────────────────────────────────────────┤
│ JOURNEY (slow clock) N / 9 [phase name] │
│ LIVE LOOP (fast clock) M / 7 [stage name] │
│ GATE OPEN | WAITING | BLOCKED │
│ LAST ACTION [what just happened] │
└─────────────────────────────────────────────────────────────────┘One-line read: Where are we on the slow clock, where are we on the fast clock, is the gate open, and what proof is still missing?
Journey and loop rails (blank pattern)
JOURNEY
1 ██░░░░░░░░ Thesis + ICPs
2 ░░░░░░░░░░ Success definitions
3 ░░░░░░░░░░ Synthetic research
4 ░░░░░░░░░░ Real research
5 ░░░░░░░░░░ Tiny architecture
6 ░░░░░░░░░░ Build (EDD)
7 ░░░░░░░░░░ Real users
8 ░░░░░░░░░░ Learn / improve
9 ░░░░░░░░░░ Scale blocked until proof
LIVE LOOP
1 Synthetic research ██░░░░░░
2 Validation / concept ░░░░░░░░
3 Product building ░░░░░░░░
4 Testing ░░░░░░░░
5 Evaluation ░░░░░░░░
6 Real feedback ░░░░░░░░
7 Memory update ░░░░░░░░
└── wraps → stage 1 (journey stays put until you decide)Rule: Moving the journey is a visible decision. Many loop cycles inside one journey phase is discipline. Never skip stage 7 — without memory update you are generating activity, not running a company OS.
Founder gates · journey only moves with your OK
The AI can recommend. It must never advance the journey alone.
Exit criteria met. Move the slow clock one phase.
Stay in this phase. Fix the weak evidence or slice.
Pause the journey. Finish memory, clear a blocker.
Hypothesis is weak. Stop protecting it.
Core beliefs
- You supply the insight. AI supplies the speed. AI without personal agency is not an edge.
- Do not fall in love with your first idea. Form a thesis, test it hard, learn, change or kill it.
- Stay small until it works. Tiny proof first. Platform later.
- You stay in control. AI does heavy work. You decide truth, build, go, and stop.
- Everything important must be visible and explainable. "Where are we?" must always have an honest answer.
- Evidence beats narrative. Time spent is not proof. Preference is not proof. Synthetic research is a filter. Real-world action is the gate.
- Build evaluation-first when you build. Spec success criteria and a harness before (or with) the build — not after a big unmeasured project.
Contrarian insight
Before serious time, answer:
- What do I believe that most smart people and current AI undervalue or disagree with?
- Is the disagreement about can, when, or whether anyone cares?
- If I am right, why is this not already fully taken?
- Upside if right vs downside if wrong?
- Can I test this with a tiny, fast experiment?
You do not need rocket science. Many good companies win on a simple truth others thought was too messy, too small, or not worth it.
Early phases exist to stop you from building the wrong thing
- Form a clear thesis (problem, whom, why now, why you).
- List several possible customer groups — not one favorite.
- Run synthetic research (AI simulations of realistic people).
- Rank by pain, willingness to act or pay, and reachability.
- Do real conversations and small tests.
- Only then lock a primary focus and a tiny first slice.
- Keep testing. Change focus if evidence is weak.
Any "current focus" is only a hypothesis that has survived so far.
Synthetic research (fast filter)
Ask, across several groups:
- How painful is this?
- What do they do instead?
- Would they pay or change behavior?
- What would make them trust a new solution?
- What would make them say no?
Real-world research
Watch what people do, not only what they say. Small tests beat big decks: price conversations, manual concierge delivery, simple landing pages, waitlists with friction.
Ranking rule
Promote a group to primary focus only if both synthetic and real evidence support it and the reward/risk scorecard looks manageable for a solo founder at your stage. Improving scores on a weak hypothesis is less valuable than finding a stronger hypothesis.
Reward / risk
Reward: pain frequency, willingness to pay/act, reachable market for a solo founder, channel fit.
Risk: hard to reach, hand-holding, messy counterparties, legal/ops complexity, time-to-signal.
A one-row scorecard beats "this group seems good."
ICP:
Reward notes (pain, pay, reach, channel fit):
Risk notes (reach cost, hand-holding, messiness, legal, time-to-signal):
Synthetic evidence (date, summary):
Real evidence (date, summary):
Rank (1 = best) / Promote? (yes/no/hold):
Kill criteria for this ICP:The founder control plane
You must always be able to see:
- Journey phase and live loop stage
- Whether the next gate is open, ready for review, blocked, or waiting for you
- Top open questions and risks
- Scores for active hypotheses and slices (not vanity metrics)
- Moments that need your judgment
- Recent important actions and why (plain language)
Questions worth asking every week
- Where are we right now?
- What is blocking the next step?
- What evidence do we actually have?
- Why is this customer group strong?
- What should I decide today?
- Challenge the current ranking of customer groups.
- Show me the weakest assumptions we are still carrying.
Decisions only you should make
- Primary customer group
- Success thresholds
- Weak evidence = wrong idea vs needs more work
- When to add people
- Enough proof to grow — or enough weakness to kill
- Which monetization path to test next
- Autonomy levels (what AI may draft vs what needs your OK)
Decision traces
Important actions leave a simple record: what was done, why, what was observed, what happens next. Traces do three jobs — honesty, learning, and feedback into personas, thresholds, rankings, and agent behavior.
Date:
Journey phase / loop stage:
Decision:
Options considered:
Evidence used (synthetic | real | mixed):
Choice:
Expected outcome:
Actual outcome (fill later):
Next review:Company-as-code
In an AI-native company the repository is not only product code. It is how the company thinks. If strategy lives only in chat history, you do not have a company — you have vibes with commits.
| Category | Holds |
|---|---|
research/ | ICPs, personas, validation notes |
product/ | What you ship |
evals/ | Tests, scores, harnesses, stress scenarios |
traces/ | Decision records |
docs/ | OS, thesis, open questions |
AGENTS.md | Thin always-on AI rules |
Exact folder names can match your stack. The principle matters more than the labels: one living source of truth the founder and agents share.
The continuous live loop
Even after you have a product, the learning loop keeps running:
PERSISTENT STATE
personas · product knowledge · decision traces
hypotheses · real feedback · scores · loop cursor
│
▼
1 Synthetic research → 2 Validation → 3 Build
→ 4 Test → 5 Evaluate → 6 Real feedback → 7 Memory update
│
└── back to 1| Stage | Goal |
|---|---|
| 1 Synthetic research | Explore several ICPs cheaply |
| 2 Validation | Stress concepts and the tiny slice |
| 3 Build | Smallest thing that can fail pass/fail rules honestly |
| 4 Testing | Prove the slice works under control |
| 5 Evaluation | Score quality vs thresholds; recommend Advance/Iterate/Hold/Kill |
| 6 Real feedback | Bring reality into state |
| 7 Memory update | Write back so the next cycle is smarter |
Evaluation-driven development
When you build, prefer this over "code first, measure later":
Spec + success criteria
→ Harness that can fail the slice repeatedly
→ Implement the smallest increment
→ Evaluation gate (scores vs thresholds)
→ Integrate + decision tracesThe next increment uses the same discipline. Do not invent a looser process because the feature "feels smaller."
Tiny slice checklist
Before building, write:
- End-to-end goal
- Pass/fail numbers (not vibes)
- How you re-run the slice (harness)
- Expected artifacts
- Human gates
- Whether synthetic users can run the path
Scoreboard
Use bars, not adjectives:
EVIDENCE
Problem evidence ████████░░
Willingness to pay ██░░░░░░░░
OPERATIONS
Completion ████████░░
Extraction ████░░░░░░
Escalation discipline████████░░
Trace completeness ████████░░
TRUST
Approval friction ████████░░
Trust feel ██████░░░░
ECONOMICS
Reward vs risk ██████░░░░
Early revenue █░░░░░░░░░Rule: Raising operational scores on a weak hypothesis is less valuable than proving or killing the hypothesis.
AI instructions (portable)
Paste a short block into Cursor / Claude / Grok / root AGENTS.md. Keep it thin. Point it at your operating system copy. Hard rules should include:
- Never advance a journey phase without your explicit approval.
- Never treat an idea as proven until synthetic and real evidence support it.
- Always report journey phase, loop stage, evidence, and next gate.
- Surface high-leverage human decisions.
- Prefer small honest tests over big protected builds.
- Record decision traces; close memory update after meaningful work.
- Keep reward/risk visible when ranking ICPs or monetization paths.
- Do not import another company's product thesis or market as yours unless you explicitly adopt it.
- Plain language. No strategy guessing. No protecting weak ideas.
The full pasteable block lives in the template file docs/company-os/ai-instructions.md in the Totbox open repo.
Near-term checklist
- Confirm the focus and contrarian edge still feel right to you.
- Write thesis + several customer groups.
- Synthetic research across several groups.
- Reward/risk scorecards for top candidates.
- Rank by evidence, not preference.
- Real conversations and small tests.
- Lock primary focus + tiny slice + pass/fail numbers.
- Install thin AI instructions.
- Minimal eval harness so the slice can re-run.
- Keep asking: "What evidence do we actually have?" and "What would make us kill this?"
What the system explicitly avoids
- Decks as proof
- Platforms before one complete user loop
- Protecting a favorite ICP when scores are weak
- Optimizing ops scores on a weak hypothesis
- AI silently changing strategy or journey phase
- Copying another startup's market because their OS docs lived in a public repo
- Expanding channels before the thin slice works
- Skipping memory update (stage 7)
Living example: Totbox instance
This section is not the portable template. It shows how one real public product repo — github.com/ivelin/totboxapp — fills the same boxes so you can see the discipline under load.
Totbox is building a household job project manager for home-service chores (beachhead examples in that repo include HVAC preventive maintenance and house cleaning), primarily via host LLMs + MCP — not a vendor directory. That beachhead is their hypothesis under test. It is not advice that your startup should do home services.
Mentees should steal the discipline, not the beachhead. Start from the portable pack in docs/company-os/, not the product README.
Hypothesis (instance product thesis)
Home-service job PM for busy households — scheduling/coordination, not a vendor directory. Subject to evidence.
Steal the discipline below. Do not adopt HVAC, cleaning, or MCP as your default business.
Journey · slow
6 / 9
Build (EDD)
Live loop · fast
4 / 7
Testing
Journey rails (where the instance sits)
Honest scores
Completion ~0.86 on engineering smoke/tests — not business PMF.
Trace completeness improving. Business Phase 1 exit (real household jobs with touchpoint drop) still open.
Open questions
- • What numeric Phase 1 business thresholds lock before more build scope?
- • When does the next real household job produce a redacted stage-6 feedback note?
| Template idea | How Totbox uses it (public-safe summary) |
|---|---|
| Thesis as hypothesis | Product thesis docs labeled subject to evidence |
| Journey + loop state | Durable machine state in company/state/company-state.json |
| Founder gates | CLI refuses journey advance without explicit --approve |
| Company-as-code | research/, product/, evals/, traces/, company/, docs/company-os/ |
| Eval discipline | Engineering smoke/tests strong; business Phase 1 exit still open |
| Gap honesty | applied-here.md and instance/ score notes admit missing formal scorecards and real-job feedback ritual |
If you point an AI at that public repo, say something like:
Read docs/company-os/operating-system.md and live-runtime.md first. Apply the portable Company OS to MY startup. Use applied-here.md and docs/company-os/instance/ only as examples of discipline and gap analysis. Do not adopt Totbox's product thesis, beachhead, or stack unless I explicitly ask. Start by restating my thesis, candidate ICPs, journey phase, loop stage, and the smallest next experiment.
Where to find the template
All of the following live in github.com/ivelin/totboxapp:
| Want | Look in Totbox open repo |
|---|---|
| Full blueprint | docs/company-os/operating-system.md |
| Weekly loop + state shape | docs/company-os/live-runtime.md |
| Thin AI paste block | docs/company-os/ai-instructions.md |
| Separation rules | docs/company-os/README.md |
| One filled instance (example only) | docs/company-os/applied-here.md, docs/company-os/instance/, company/state/ |
Closing
AI tools change fast. Principles last longer.
The goal is simple: help a solo founder use AI to move faster than ever — while staying honest, in control, and unwilling to protect weak ideas.
If you cannot fill the blank snapshot board this week, that is not a failure. That is your first honest status report.
Fill the board. Challenge the board. Then decide: Advance, Iterate, Hold, or Kill.
You keep the keys.
FAQ
What is a Company Operating System for solo founders?
A simple way to see where your startup is, what evidence you have, and what to decide next — clear enough for a non-technical founder, strict enough that weak ideas cannot hide behind activity. It combines a slow bootstrap journey, a fast weekly learning loop, honest scores, decision traces, and thin AI instructions.
What are the two clocks?
Slow clock: bootstrap journey phases 1–9 (prove the business; founder gates every advance). Fast clock: live loop stages 1–7 (research → validate → build → test → evaluate → real feedback → memory update). Many loops can run inside one journey phase.
Who advances a journey phase?
Only you. Labels: Advance, Iterate, Hold, Kill. The AI can recommend; it must not move the journey alone.
Template vs Totbox instance?
Template = portable method in docs/company-os/. Instance = Totbox's filled product thesis and scores. Steal discipline, not the beachhead.
Where is the open template?
github.com/ivelin/totboxapp — start with operating-system.md, live-runtime.md, and ai-instructions.md.
Should I copy Totbox's product?
No, unless you explicitly choose to. Point your AI at the OS docs and apply them to your thesis, ICPs, and next experiment.
Related insights
- Synthetic Users & Outcome Pricing — synthetic research and pricing discipline for AI-native founders
- Founder Institute agentic environments — mentorship context for FI builders
- Agent Skills & MCP distribution — how founders distribute agent capabilities
- Non-tech founder AI agent team — building with AI without a traditional eng org
- Totbox Company OS on GitHub — portable template + one live instance
