Forge OS: an evidence chain from idea to merge, with agents that can't weaken the tests

The internal toolchain behind the studio: schema-checked AI skills, a change graph that marks downstream work stale, governance that scales with release stage, and a red-then-green coding loop that rejects any run that touches a test file.

Why we built it

AI has made code and documents cheap. It hasn't made it cheap to know whether something is worth building, or to keep a design tied to the evidence behind it. The failure modes are familiar: jumping to a solution before the problem is understood, "ideas" that are all variations on the same thing, designs that quietly rest on a premise someone changed last week, and coding agents that "pass" by editing the tests.

Forge OS is the toolchain we use to fight those failure modes. It's a local-first desktop app that runs an idea through research, discovery, ideation, strategy, design, a test-gated build and launch, with a human sign-off at every gate. It's an internal tool today. This post is about how it works.

The shape of it

Forge OS pipeline and governanceAn idea enters intake and is normalised. Telic, a research sidecar, maps how the problem evolved. Discovery ends in a verdict. Ideation runs with a red-team critic and two human gates. Strategy, product definition and system design follow, each behind a gate. A build plan feeds test-gated coding agents, then launch and learnings, which can reopen an earlier phase. Cross-cutting: a governance resolver decides which gates are live, and a change graph marks downstream work stale.Intake7 intents · normalised · confirmedTelic research maphow the problem evolved · sidecarPhase 1 · Discoveryevidence ledger → verdict⛔ gateID8 · Ideation60 operators · red-team critic⛔ 2 human gatesPhases 2–4strategy · definition · system design⛔ gatesPhase 5 · Build plantests · tickets · context packsBuild executiontest-gated coding agents⛔ security floorPhase 6 · Launch & learnnext-cycle decision⛔ release gatereopens a phase by marking it stale
Seven pipelines, 104 steps. Which gates are live depends on the release stage and operating mode. Tap to zoom.

There are two Electron apps sharing one design. Forge runs the pipelines. Telic is a research substrate that maps how a problem domain evolved, era by era, through pressures, pains, outcomes, concepts and realisations. Forge launches Telic as a headless sidecar and talks to it over HTTP on the loopback interface.

Skills: prompts with schemas and validators

Every AI step is a skill, defined in one place (defineSkill.ts) as an id, an output schema, prompt builders and a deterministic validator. Running one means: build the prompt, call generateObject with the Zod schema, then run the validator, which returns a list of issues graded error, warning or info. Validators carry the opinions. For example:

  • A discovery verdict of build is an error if any of the four worthiness checks failed. Every verdict (build, narrow, reject, research_more, manual_test_first) must carry its own payload, such as a narrowed scope or a manual test to run.
  • An evidence ledger's contradictions must cite claims that exist. A ledger with no tensions gets an info flag, because a record without tensions usually hasn't been examined.
  • Any idea that can't cite its evidence is quarantined.

Revisions come in two flavours. Repair feeds validator issues back to the model. Refine applies a human's notes, keeps unflagged content byte-identical, and never renumbers ids, so references downstream stay valid.

The core never imports a model provider. The model is injected as an executor built on the Vercel AI SDK, with Gemini, Anthropic, OpenAI, Groq and Mistral available per stage. Keys are encrypted with the operating system keychain through Electron's safeStorage, and no IPC call ever returns a key.

Artifacts and the change graph

Every output is an artifact with a kind (about 130 of them), a status (generated, human_confirmed or stale), a payload, provenance (who or what created it, from what) and its validation result. Entities inside carry stable ids, and validators check every cross-reference.

When an upstream artifact changes, everything downstream of it is marked stale. The graph is deliberately conservative: over-marking costs a regeneration, while under-marking would let a design ship on a premise that no longer holds. Two details keep it practical. The cascade only fires when the payload actually changed, so re-confirming something doesn't wipe the work below it. And auto-run regenerates only the missing or stale stages.

Projects are stored as a single JSON document, written crash-safe: write to a temp file, fsync, keep the previous version as a backup, then atomically rename. Schema migrations upgrade old projects on load.

Ideation that resists sameness

Ask a model for ten ideas and you often get one idea ten times. The ID8 ideation pipeline is 12 stages with 9 separate agent roles: an opportunity synthesiser, an incumbent teardown analyst, an operator-driven ideator, a concept composer, a segment simulator, a red-team critic, an evaluator, an experiment designer and a PRD drafter. The ideator is forced across a library of 60 operators in 10 families. Generation, simulation, critique and evaluation are kept as separate agents. Near-duplicates (token-set similarity of 0.7 or more) are flagged. Ideas are killed before scoring if they amount to parity with an incumbent without a structural advantage. Two gates in the middle can only be passed by a human.

Governance that scales with the stakes

Nine gates cover scope, customer and evidence, strategy, UX, domain invariants, architecture, security and privacy, release, and amendments to the rules themselves. Gates switch on along a release ladder from R-2 (a sketch) to R5+ (general availability), and an operating mode (lightweight, standard, strict or discovery-reset) adjusts the scope further. A prototype gets the scope and worthiness gates; a general release gets the full ladder. The security gate is a floor that no mode can waive.

Gates have five outcomes: approve, approve with conditions, revise, send back or reject. AI can give an advisory verdict, labelled as such, but only a person can approve. Downstream pipelines are locked in both the UI and the main process until their gate clears.

The coding loop: red, then green, without touching the tests

Phase 5 turns the design into a test strategy, test cases, tickets and context packs for coding agents. Build execution then drives Claude Code or Codex as child processes, one ticket at a time:

Test-gated coding agent loopFor each ticket Forge creates a git worktree on its own branch. The agent first writes only tests, which must fail. Then a fresh run implements. Forge diffs the changes without rename detection, and rejects the run if any test file changed. The ticket's own tests must pass. Merging re-runs the full suite and keeps the branch if the base breaks.git worktree addbranch forge/<project>/<ticket>Agent: write tests onlymust FAIL before implementationAgent: implementfresh worktree, same branchgit diff --no-renamesany test file touched → rejectedTicket's own testsmust now pass (red → green)merge --no-ffthen re-run the full suiteBase still green?delete branch · else keep itpassingalready?suspicious
Agents run under a 15-minute timeout and a per-repo lock. Every run is recorded in the build dispatch log. Tap to zoom.
  • Red first. The first run may only write tests, and Forge runs them scoped to the files the agent wrote. A test that passes before any implementation is flagged as suspicious.
  • Never weaken the tests. After the implement run, Forge stages everything and lists the changes with git diff --cached --name-only --no-renames. If any path matches a test-file pattern (covering JavaScript, TypeScript, Go, Python, Ruby and Rust conventions), the run is rejected as tests_modified. The --no-renames flag closes the loophole of renaming a test file so it no longer looks like one.
  • Judge the ticket, not the world. Green means the ticket's own red tests now pass, decided by a small pure function that is unit-tested. One flaky test elsewhere can't block every ticket.
  • Merge carefully. A merge needs a passing implement run, a clean tree and no detached HEAD. It uses merge --no-ff, aborts on conflict, re-runs the full suite, and only deletes the branch if the base is still green.

Measurement that has to mean something

Forge's measurement rail has one rule we like a lot: a metric that doesn't declare which decision it triggers is rejected by the schema. Phase 6 ends in a next-cycle decision that reopens a phase by marking its entry artifact stale, which closes the loop mechanically rather than in a retro document.

Stack

LayerWhat we use
AppsElectron, electron-vite, React 19, TypeScript
SchemasZod 4 for every artifact and structured output
ModelsVercel AI SDK generateObject with Gemini, Anthropic, OpenAI, Groq and Mistral, chosen per stage
StorageJSON per project (fsync, backup, atomic rename, migrations); Telic adds per-version history and a JSONL log of every model call
AgentsClaude Code or Codex CLI in git worktrees, with timeouts and process-group kill
CIFormat, lint, typecheck, offline tests, build, dependency audit and a secret scan of the full history

Honest limits

Forge OS is our own tool and is still maturing. The coding agent runs in an isolated worktree, but not yet in a sandbox: it has the same filesystem and network access as the user. Proper sandboxing is the blocker before we'd hand Forge to anyone else. Evidence today is mostly synthesised by models from public knowledge; ingesting real interviews, reviews and tickets is next. So is a judge that calibrates itself by comparing past predictions with what actually happened.

Forge is meant to sit behind Bridge, which would become its front door. See how the products fit together.

See Forge OS All posts