What we set out to fix
Note apps make you organise, either before you write or after. Flashcard apps make you copy what you wrote into a second system, and then the two drift apart: you fix a note and the card built from it keeps teaching the old answer. AI tools add a third problem. They tend to rewrite things all at once, which is hard to check and harder to undo.
Chamber puts notes, cards and spaced repetition in one local-first app, and treats every AI change as a proposal you can inspect and reverse. This post walks through how it is built.
The shape of the system
Chamber is a static single-page app that installs as a PWA and boots offline. There is no Chamber server in the critical path. The browser's IndexedDB is the source of truth, and everything else (AI providers, sync, the agent companion) is optional and plugged in from the edges.
Stack
| Layer | What we use |
|---|---|
| App | TypeScript, React 18, Vite 5, TanStack Router, Zustand, Zod, Tailwind and Radix UI |
| Editor | TipTap (ProseMirror), KaTeX for math, lowlight for code, stored as one row per block with stable block ids |
| Storage | Dexie 4 over IndexedDB (schema v44), with an append-only operations table |
| On-device ML | tesseract.js for OCR (self-hosted, no CDN), Whisper through transformers.js on ONNX Runtime Web, pdf.js |
| AI | Bring your own key: Anthropic, OpenAI-compatible providers (OpenAI, Groq, Mistral, DeepSeek, xAI, OpenRouter, Cohere, Ollama), Gemini, and Voyage for embeddings |
| Sync (optional) | Supabase Auth, Postgres with row-level security and Storage, called with plain fetch; one Deno edge function for reminders |
| Agents | A Node MCP companion built on the official MCP SDK |
| Tests | Vitest on fake-indexeddb, Playwright in three CI shards |
One write path, recorded
Every write goes through repository functions in lib/db/repositories, which update the canonical tables (pages, blocks, cards, reviewStates, reviewLogs and friends) and append to the operations log in the same transaction. That one rule pays for a lot:
- The change feed built on the log drives the semantic indexer and background jobs, without polling the whole database.
- The sync loop pushes operations and dirty rows, rather than diffing snapshots.
- A test (
syncCoverage) fails the build if any write path forgets to reach the sync queue.
Multi-tab editing is handled with a per-block three-way merge, and a small unload journal recovers a save that was interrupted by a closing tab.
FSRS-6, implemented and checked
The scheduler in lib/scheduler is a pure, deterministic implementation of FSRS-6 with its 21 weights. We wrote it ourselves rather than vendoring it so it could be tested at the level we wanted:
- Parity tests. Golden vectors from an independent reference implementation are replayed through the real
applyReview. Those tests caught a real bug: an early "default" weight vector turned out to be overfitted to one person's history. - A per-user optimiser. Once you have 1,000 reviews, a log-loss coordinate-descent optimiser fits weights to your memory, with versioned parameters so every past interval can be traced to the exact weights that produced it.
- Answers are withdrawn, never deleted.
reviewLogsis append-only and syncs, so undoing an answer marks it withdrawn and restores the state the log recorded. A test sweep checks that every reader respects withdrawals. - An exam planner that says whether a study plan is realistic, plus load balancing, easy days, leech handling and custom study.
There's a full write-up of the parity tests and the optimiser in our FSRS deep dive.
CI runs the unit suite a second time with the clock shifted two years forward, which flushes out tests that silently depend on today's date. That matters a lot in a scheduler.
Cards that follow their notes
Because notes are stored per block with stable ids, every card remembers the block it came from. When you accept an edit to that block, Chamber works out what it means for the cards:
The design choice we care about most is the split of responsibility. A model is good at describing what changed between two versions of a sentence. It should not decide what happens to your review schedule. So the model's classification is only an input, and the consequence comes from a small, fixed table in code that's easy to read. That makes the behaviour predictable, testable and explainable.
AI as proposals, not edits
All chat calls go through one function (complete()) and all embedding calls through another (embed()). Requests go straight from the browser to the provider you configure, with your key encrypted at rest under a non-extractable WebCrypto key. Nothing generated is saved until you accept it:
- Edits arrive as a word-level diff. Each proposal snapshots the version of its target, and if the target has changed since, the proposal turns stale instead of overwriting your work.
- Accepted changes become action batches with before and after snapshots. Undo is all-or-nothing and dry-runs first to detect conflicts.
- Generated cards, summaries, quizzes and tutor explanations are drafts until approved. Free-recall grading only suggests a rating.
Search is hybrid: lexical matching plus vector similarity computed in a Web Worker, fused with reciprocal rank fusion. The same index finds duplicate cards before they reach your reviews.
On-device by default
Image OCR runs with a self-hosted tesseract build, so no third-party CDN is contacted. Spoken answers in free-recall mode are transcribed with Whisper running locally. Silence is detected before the model is downloaded, and Whisper's well-known hallucinated filler ("Thank you.") is filtered out. Cloud transcription exists, but only if you turn it on.
Imports are local too. The Anki importer reads .apkg files with a small hand-written read-only SQLite reader and the browser's own decompression, and maps Anki's review state onto FSRS so your history comes with you.
Sync that encrypts before it uploads
Sync is optional. When enabled, a loop runs every 30 seconds: push operations, push dirty rows, pull, then move file chunks. With end-to-end encryption on, a passphrase is stretched with PBKDF2-SHA256 (310,000 iterations, per-vault salt) into a non-extractable AES-GCM-256 key, with a fresh IV per row. A locked vault pushes nothing at all. We are upfront about the limits: some metadata stays visible to the server, and a lost passphrase means lost cloud data.
Review reminders are designed the same way. The server stores only a wake-up time and sends an empty push. The service worker then counts what's due from local data and writes the notification itself, so no card content passes through the server.
Honest status: sync, encryption and reminders are implemented and unit-tested, but haven't yet been verified end to end against a live deployment. That is the next milestone before we call sync generally available.
Letting agents in, safely
The MCP companion lets an agent such as Claude Desktop search notes, fetch a page, capture text, create a card or read the due summary. It can't read the vault directly. It relays tool calls over a WebSocket bound to 127.0.0.1, checks a token with a constant-time compare, exposes a closed list of tools, and every write lands as an AI-attributed batch you can undo.
What's still early
The Workspace Composer, a single input that files what you type into the right page, ships in a conservative mode. It places captures when retrieval evidence is strong and sends anything uncertain to an inbox. Fully automatic placement stays switched off until it proves itself in shadow mode. Extracting knowledge from long sources works, but its quality hasn't yet been measured against labelled data, so we treat it as a preview. Native apps are deferred while the PWA does the job.
Chamber shares design ideas with the rest of the studio, such as keeping AI configuration separate from keys, which comes from Forge OS. See how the products fit together.