Why we built it
AI has made code and documents cheap. It hasn't made it cheap to know whether something is worth building, or to keep a design tied to the evidence behind it. The failure modes are familiar: jumping to a solution before the problem is understood, "ideas" that are all variations on the same thing, designs that quietly rest on a premise someone changed last week, and coding agents that "pass" by editing the tests.
Forge OS is the toolchain we use to fight those failure modes. It's a local-first desktop app that runs an idea through research, discovery, ideation, strategy, design, a test-gated build and launch, with a human sign-off at every gate. It's an internal tool today. This post is about how it works.
The shape of it
There are two Electron apps sharing one design. Forge runs the pipelines. Telic is a research substrate that maps how a problem domain evolved, era by era, through pressures, pains, outcomes, concepts and realisations. Forge launches Telic as a headless sidecar and talks to it over HTTP on the loopback interface.
Skills: prompts with schemas and validators
Every AI step is a skill, defined in one place (defineSkill.ts) as an id, an output schema, prompt builders and a deterministic validator. Running one means: build the prompt, call generateObject with the Zod schema, then run the validator, which returns a list of issues graded error, warning or info. Validators carry the opinions. For example:
- A discovery verdict of
buildis an error if any of the four worthiness checks failed. Every verdict (build,narrow,reject,research_more,manual_test_first) must carry its own payload, such as a narrowed scope or a manual test to run. - An evidence ledger's contradictions must cite claims that exist. A ledger with no tensions gets an info flag, because a record without tensions usually hasn't been examined.
- Any idea that can't cite its evidence is quarantined.
Revisions come in two flavours. Repair feeds validator issues back to the model. Refine applies a human's notes, keeps unflagged content byte-identical, and never renumbers ids, so references downstream stay valid.
The core never imports a model provider. The model is injected as an executor built on the Vercel AI SDK, with Gemini, Anthropic, OpenAI, Groq and Mistral available per stage. Keys are encrypted with the operating system keychain through Electron's safeStorage, and no IPC call ever returns a key.
Artifacts and the change graph
Every output is an artifact with a kind (about 130 of them), a status (generated, human_confirmed or stale), a payload, provenance (who or what created it, from what) and its validation result. Entities inside carry stable ids, and validators check every cross-reference.
When an upstream artifact changes, everything downstream of it is marked stale. The graph is deliberately conservative: over-marking costs a regeneration, while under-marking would let a design ship on a premise that no longer holds. Two details keep it practical. The cascade only fires when the payload actually changed, so re-confirming something doesn't wipe the work below it. And auto-run regenerates only the missing or stale stages.
Projects are stored as a single JSON document, written crash-safe: write to a temp file, fsync, keep the previous version as a backup, then atomically rename. Schema migrations upgrade old projects on load.
Ideation that resists sameness
Ask a model for ten ideas and you often get one idea ten times. The ID8 ideation pipeline is 12 stages with 9 separate agent roles: an opportunity synthesiser, an incumbent teardown analyst, an operator-driven ideator, a concept composer, a segment simulator, a red-team critic, an evaluator, an experiment designer and a PRD drafter. The ideator is forced across a library of 60 operators in 10 families. Generation, simulation, critique and evaluation are kept as separate agents. Near-duplicates (token-set similarity of 0.7 or more) are flagged. Ideas are killed before scoring if they amount to parity with an incumbent without a structural advantage. Two gates in the middle can only be passed by a human.
Governance that scales with the stakes
Nine gates cover scope, customer and evidence, strategy, UX, domain invariants, architecture, security and privacy, release, and amendments to the rules themselves. Gates switch on along a release ladder from R-2 (a sketch) to R5+ (general availability), and an operating mode (lightweight, standard, strict or discovery-reset) adjusts the scope further. A prototype gets the scope and worthiness gates; a general release gets the full ladder. The security gate is a floor that no mode can waive.
Gates have five outcomes: approve, approve with conditions, revise, send back or reject. AI can give an advisory verdict, labelled as such, but only a person can approve. Downstream pipelines are locked in both the UI and the main process until their gate clears.
The coding loop: red, then green, without touching the tests
Phase 5 turns the design into a test strategy, test cases, tickets and context packs for coding agents. Build execution then drives Claude Code or Codex as child processes, one ticket at a time:
- Red first. The first run may only write tests, and Forge runs them scoped to the files the agent wrote. A test that passes before any implementation is flagged as suspicious.
- Never weaken the tests. After the implement run, Forge stages everything and lists the changes with
git diff --cached --name-only --no-renames. If any path matches a test-file pattern (covering JavaScript, TypeScript, Go, Python, Ruby and Rust conventions), the run is rejected astests_modified. The--no-renamesflag closes the loophole of renaming a test file so it no longer looks like one. - Judge the ticket, not the world. Green means the ticket's own red tests now pass, decided by a small pure function that is unit-tested. One flaky test elsewhere can't block every ticket.
- Merge carefully. A merge needs a passing implement run, a clean tree and no detached HEAD. It uses
merge --no-ff, aborts on conflict, re-runs the full suite, and only deletes the branch if the base is still green.
Measurement that has to mean something
Forge's measurement rail has one rule we like a lot: a metric that doesn't declare which decision it triggers is rejected by the schema. Phase 6 ends in a next-cycle decision that reopens a phase by marking its entry artifact stale, which closes the loop mechanically rather than in a retro document.
Stack
| Layer | What we use |
|---|---|
| Apps | Electron, electron-vite, React 19, TypeScript |
| Schemas | Zod 4 for every artifact and structured output |
| Models | Vercel AI SDK generateObject with Gemini, Anthropic, OpenAI, Groq and Mistral, chosen per stage |
| Storage | JSON per project (fsync, backup, atomic rename, migrations); Telic adds per-version history and a JSONL log of every model call |
| Agents | Claude Code or Codex CLI in git worktrees, with timeouts and process-group kill |
| CI | Format, lint, typecheck, offline tests, build, dependency audit and a secret scan of the full history |
Honest limits
Forge OS is our own tool and is still maturing. The coding agent runs in an isolated worktree, but not yet in a sandbox: it has the same filesystem and network access as the user. Proper sandboxing is the blocker before we'd hand Forge to anyone else. Evidence today is mostly synthesised by models from public knowledge; ingesting real interviews, reviews and tickets is next. So is a judge that calibrates itself by comparing past predictions with what actually happened.
Forge is meant to sit behind Bridge, which would become its front door. See how the products fit together.