Engineering

How this thing is actually built.

Every claim on this site is a consequence of a decision in the code. This page shows the decisions.

Six diagrams of the real mechanism — the pipeline that replaces a bank login, a categorizer that learns without a server, how sync stays honest across two backends, and why some of the safest parts of the system are the parts that don't exist.

4native surfaces: iOS, macOS, Android, web
0bank credentials asked for, stored or brokered
0analytics or ad SDKs linked into the app
$0infrastructure cost per Apple-only customer
01 / INGEST

Importing your books without a bank login

Almost every accounting app opens with “connect your bank.” We never built that door — so the work had to happen somewhere else.

Scroll the diagram sideways to read it →

Pipeline: a statement file is fingerprinted, parsed on device, and matched against existing records. A crossed-out lane shows the bank-login route that was never built. THE DOOR WE NEVER BUILT your bank login→ Plaid / aggregator→ their serversnever implemented statement file CSV · OFX · QFX · QIF · IIF dragged in by you — no credentials header fingerprint sniffer detectFormat(header:) → 5 dialects chase bofa wells capone generic quote-aware CSV scanner hand-rolled state machine — commas inside quoted fields survive OFX/SGML tag reader <TRNAMT> <DTPOSTED> <NAME> unterminated tags handled BankTransaction { date · description · amount · isCredit } reconciliation matchermatched|Δamount| ≤ $0.01 ∧ |Δdate| ≤ 2dpossiblesame amount, window widens to 7dunmatched→ ReceiptClassifier.classify()returns a Schedule-C category+ a confidence score, for review
The ingest pipeline. A statement file is fingerprinted, parsed and matched entirely on your device. The crossed-out lane at the top is the route we declined to build.

“Connect your bank” means handing your banking credentials to an aggregator, which then holds a token against your account indefinitely and turns your transaction history into an asset on someone else’s balance sheet. Plaid paid $58 million settling a case about how much it collected that way.

What happens instead

You drop in a statement file. Everything after that runs locally.

A header fingerprinter reads row one and picks a dialect. Chase, Bank of America, Wells Fargo and Capital One each export a different column order, so detectFormat(header:) sniffs the header and routes to the matching row parser. Anything unrecognised falls through to a generic reader rather than failing outright.

CSV goes through a hand-rolled, quote-aware scanner — not a split(","). Merchant names contain commas: "SMITH, J & SONS LLC" is one field, not two, so the parser tracks quote state character by character. OFX and QFX are SGML rather than XML, and real exports routinely ship unterminated tags, so the tag reader falls back to a newline boundary when there’s no closing tag.

Both paths converge on one value type: BankTransaction { date, description, amount, isCredit }.

Then the matcher runs

Each transaction is compared against what you’ve already recorded. Amount within a cent and date within two days is treated as matched — no action needed. The same amount inside a seven-day window is a possible match, flagged for you to confirm rather than merged silently. Anything left over is unmatched, and an on-device classifier proposes a Schedule C category with a confidence score.

What this costs you

This is more manual than an automatic bank feed. You export a file instead of a connection doing it quietly in the background. In exchange there is no stored credential, no third party holding a token against your account, and no copy of your transaction history in an aggregator’s warehouse. Nothing in this pipeline makes a network call.

  • Services/BankImportService.swift
  • Services/ReceiptClassifier.swift
  • web/src/screens/StatementImport.tsx
02 / INTELLIGENCE

A categorizer that learns, and never phones home

“Use AI to categorize expenses” usually means shipping every transaction you have to someone else’s GPU. We wanted the opposite: something that improves with use while transmitting nothing.

Scroll the diagram sideways to read it →

A five-stage cascade for categorising an expense, from learned vendor memory through Naive Bayes to an on-device language model, with user corrections feeding back into training. input: "SQ *BLUE BOTTLE $6.40" learned vendor memory exact match on vendors you already corrected1conf 1.00 seeded vendor table 111 known US SMB brands → Schedule C category2instant token Naive Bayes P(category | tokens), trained only on YOUR edits3posterior keyword heuristics the deterministic floor — always answers4fallback Apple Foundation Models @Generable guided decode · iOS 26 / macOS 265on-device LLMyou correct it → train() the model is a file on your machine Swift: UserDefaults · TypeScript: localStorage — same algorithm, two languages, byte-compatible No training server. No shared corpus. Your corrections improve nobody's model but yours.
The cascade. Five stages, cheapest first, stopping at the first confident answer. Your corrections feed straight back into stages one and three.

1 — Learned vendor memory. You corrected “Blue Bottle” to Meals once. That mapping is permanent and returns at confidence 1.0. Most repeat transactions never get past this stage.

2 — A seeded vendor table. 111 well-known US small-business vendors pre-mapped to Schedule C categories, so a fresh install is useful on day one instead of after a month of training. Deliberately excluded: Amazon, Walmart, Target. A big-box charge is genuinely ambiguous, and a confident wrong answer is worse than no answer.

3 — Token-level Naive Bayes. For vendors nobody has seen, P(category | tokens) over a model trained exclusively on your own corrections. This is what handles the long tail of local businesses no seed table will ever contain.

4 — Keyword heuristics. The deterministic floor. Always returns something.

5 — Apple Foundation Models. On iOS 26 and macOS 26, @Generable with guided decode constrains the model to emit exactly one valid category — no parsing, no invented labels outside the enum. It runs on the Neural Engine, offline.

Where the model lives

UserDefaults on Swift, localStorage on the web. The same algorithm implemented twice, byte-compatible across both. There is no training server and no shared corpus: your corrections make your copy smarter and nobody else’s.

Federated learning is a fine paper. Not sending the data anywhere is a stronger guarantee.

  • Services/SmartCategorizer.swift
  • Services/IntelligenceService.swift
  • web/src/lib/smartCategorize.ts
03 / SYNC

One app, two backends, one source of truth

An economics problem disguised as an architecture problem — and the reason “free” can stay free.

Scroll the diagram sideways to read it →

One local store routed to either CloudKit or Postgres, with a merge protocol covering delta push, filtered pull, last-write-wins merge, and deletion tombstones in both directions. SwiftData — local store the source of truth, always SyncCoordinator routes on BusinessProfile.syncBackend bound write-once at onboarding CloudKit private zone runs in the customer's own iCloud no account, no password, no login the 2nd device is already signed in infra cost to us: $0 Postgres + RLS web, Android, mixed-device teams anonymous session upgrades to an account row-level isolation per org_id infra cost to us: we pay per rowApple-only fleetseveryone else the merge protocolpushdelta only — newer than an updated_at watermarkpull.gt("updated_at", since) — RLS returns only your rowsmergelast-write-wins on updated_at, applied idempotentlydelete ↑SyncTombstone rows, retried until the server confirmsdelete ↓AFTER DELETE trigger → record_sync_deletion()that trigger fn: SECURITY DEFINER, search_path pinned, revoked from /rpc
Routing and merge. One local store, two possible destinations, and a delete path that gets the same rigour as writes.

The app is free. So every customer whose data sits in our database is a customer who costs money — rows, storage, egress, indefinitely. Scale that and “free” quietly becomes “free until it isn’t,” which is the exact bait-and-switch this app exists to avoid.

The way out: for Apple-only customers, sync runs through their own iCloud. A CloudKit private database. Their storage, their Apple ID, our infrastructure cost of zero. No account, no password, no login screen — the second device is already signed in, because it’s the same Apple ID.

Customers on web, Android, or mixed-device teams can’t use CloudKit; a browser has no path to it. They ride Postgres with row-level security instead, and that’s the tier we actually pay for.

SwiftData remains the local source of truth either way. A SyncCoordinator routes to the chosen backend, bound write-once at onboarding.

The guardrail matters more than the saving

Onboarding tells genuinely cross-platform users to pick Postgres up front, so nobody gets stranded on a backend that can’t reach their Android tablet and needs a painful migration later. A cost-saving default that traps a user isn’t a saving.

The merge protocol

Push sends deltas only — everything newer than an updated_at watermark. Pull asks for .gt("updated_at", since) and lets row-level security decide what “your rows” means. Merge is last-write-wins on updated_at, applied idempotently so a re-fetch is always safe.

Deletes are where sync engines quietly lose data, so they carry equal weight. Deleting locally writes a tombstone that retries until the server confirms it. Deleting remotely fires an AFTER DELETE trigger into a sync_deletions table that other devices drain — SECURITY DEFINER, search_path pinned, and revoked from the REST layer so it’s reachable only by the trigger.

One rule we’ll defend to anyone: an empty pull must never be read as “everything was deleted elsewhere.” That single assumption is how sync engines eat someone’s books.

For the user-facing side of this — how you choose a backend and what sync-off means — see the security page.

  • Services/SupabaseClientShim.swift
  • Services/CloudSyncService.swift
  • supabase/migrations/0024_record_sync_deletion_fn.sql
04 / HANDOFF

Desktop → phone camera → desktop, with no backend

You’re on your laptop. The receipt is on your desk. The camera is in your pocket. Most apps solve this with an account on the phone, a websocket and a service to broker it.

Scroll the diagram sideways to read it →

A six-step cross-device scan: the desktop mints a single-use signed upload URL into a QR code, the phone uploads with no account, and the desktop polls until the file lands. DESKTOP — signed in PHONE — no account, no session createCaptureSlot(orgId) Storage returns { path, token } a single-use upload grant1 render a QR code /capture#path=…&token=… # = fragment, never sent anywhere2 phone opens the route no auth, no cookie, nothing kept on that page3 uploadToSignedUrl(…) capture="environment" → straight to Storage4 poll every 2500 ms ticks > 120 → the slot expires and the UI says so5 insert the document row it appears in the cabinet, live no refresh needed6scannedobject lands 0 edge functions · 0 logins · 0 lines of server code
The handoff. A single-use upload grant travels to your phone inside a QR code, and the desktop watches for the result. No second login, no server code.

The desktop asks Storage for a single-use signed upload URL and receives a path and a token. It renders them into a QR code pointing at a public /capture route — with the secret in the URL fragment.

The fragment is the whole trick. Everything after the # is never transmitted in an HTTP request, so the upload grant reaches the phone’s JavaScript without ever appearing in a server log, a proxy, or a Referer header.

The phone opens that page with no account, no cookie and no session. It’s an <input capture="environment"> that fires the rear camera and posts straight to the pre-authorized path. The page itself stores nothing.

The desktop polls every 2500 ms for the object to appear. After 120 ticks — five minutes — the slot expires and the interface says so plainly instead of spinning forever. When the file lands, a document row is inserted and it shows up in your filing cabinet live.

No edge function. No websocket. No login on the second device. No server code at all, because the grant is the authorization.

Where it stops

On-device OCR — reading the amount and vendor off the receipt automatically — stays in the native app. A browser can’t run Vision privately. So the web moves the photo across devices and then points you at the app, rather than pretending it can do the rest.

  • web/src/components/ScanCard.tsx
  • web/src/screens/Capture.tsx
  • Utilities/OCRService.swift
05 / ISOLATION

The safest table is the one you never created

Multi-tenant isolation usually means every query gets a WHERE clause and everyone hopes nobody forgets one. That’s layer one. Layer two is the one we actually rely on.

Scroll the diagram sideways to read it →

Tenant isolation: a token hook stamps org_id into every JWT and row-level security scopes on it, while the financial tables are absent from the schema entirely. auth.users identity only memberships user → org → role access-token hook fires on issue JWT claims org_id · employee_id · is_ownerROW-LEVEL SECURITYevery policy scopes on org_id from the token — not from anything the client sendsworkers narrow further: clock_entries.employee_id · appointments.assigned_employee_idTABLES THAT DELIBERATELY DO NOT EXIST HEREthe worker-sync schema allowlists 6 record types. these were never created:invoicesincomeexpensesclientscontractorsproposalsjobsmileagebusiness_profilebank_statementsA worker cannot read your books through this backend. Not by policy — by absence. contractor Tax IDs: only the last four digits ever leave the device
Two layers. Claims minted server-side drive every policy — and the records that matter most were never given a table to leak from.

Layer one — claims the client can’t forge

An access-token hook fires when a JWT is issued, reads the memberships table, and stamps org_id, employee_id and is_owner into the token’s claims. Every row-level policy scopes on those claims — on what the server minted, never on anything the client sends. A client that lies about its org_id gets an empty result set, because the filter reads the token, not the request.

Layer two — absence

When worker accounts were added, we first wrote down which record types a worker can ever touch. Six. Then we made the schema match that list — not with policies, with absence.

There is no invoices table in that backend. No income. No expenses. No clients, contractors, proposals, jobs, mileage or business profile. They were never created.

A misconfigured policy on a table that doesn’t exist leaks nothing. A future migration can’t accidentally widen access to a column that was never written. The migration file opens with the invariant stated in a comment, so the next person to touch it sees the rule before they see the schema.

The same reasoning on Tax IDs

Full TINs and signed W-9s never leave the device. Only the last four digits sync — enough to render a 1099-NEC, useless to anyone who breaches the database. There’s no encryption story to get right, because the plaintext was never uploaded.

Policies are a control. Absence is a guarantee. Where we could choose, we chose absence.

The Apple sync tier reaches the same guarantee through a hard allowlist enforced at three points instead — that one is written up on the security page.

  • supabase/migrations/0001_core.sql
  • Services/CloudSyncPolicy.swift
  • Models/Contractor.swift
06 / BACKUPS

An encrypted backup we couldn’t open if we wanted to

Every backup feature has a quiet question behind it: who else can open this? Here the answer had to be nobody — enforced, not promised.

Scroll the diagram sideways to read it →

The backup file layout — magic bytes, salt, and an AES-GCM sealed box — with the key derived from a passphrase through 600,000 PBKDF2 iterations. THE FILEmagic 'BZLY2'salt · 16 random bytesnonce ‖ ciphertext ‖ GCM tagTHE KEY your passphrase never stored, never sent PBKDF2-HMAC-SHA256 600,000 iterations + salt AES-GCM-256 seal / openTHE VERSION TAG EARNS ITS KEEPmagic == BZLY2 → PBKDF2 · every backup written todaymagic == BZLY1 → legacy HKDF · restore path only, never written againONE CONTRACT, TWO RUNTIMESSwift CryptoKit · CommonCrypto CCKeyDerivationPBKDFTypeScript crypto.subtle.deriveKey · crypto.subtle.encryptA backup made on an iPhone restores in a browser. Byte for byte. we hold no key, no escrow, no reset link Lose the passphrase and the backup is unrecoverable — including by us. That is the feature. A recovery path we control is a recovery path someone else can take.
The envelope. Magic bytes, a per-backup salt, and an authenticated sealed box — with a version tag that lets the crypto improve without orphaning old files.

The file has three parts: five magic bytes, a 16-byte random salt, then an AES-GCM sealed box carrying nonce, ciphertext and authentication tag.

The key comes from your passphrase through PBKDF2-HMAC-SHA256 at 600,000 iterations, salted per backup — so two backups of identical data produce completely different bytes. The iteration count is deliberately expensive: it’s the difference between a leaked backup being brute-forced in an afternoon and it not being worth attempting.

GCM matters as much as AES here. It’s authenticated encryption, so a corrupted or tampered backup fails to open rather than silently restoring garbage into your books. A wrong passphrase is rejected by the tag, not by hoping the JSON parses.

The version tag has already earned its keep

An earlier build derived keys with HKDF — fine for key material, wrong for a human passphrase, because there’s no work factor. Rather than orphan those files, restore branches on the magic bytes: BZLY2 goes to PBKDF2, BZLY1 to legacy HKDF. Old backups still open; new ones are never written the old way. You get to correct a mistake without abandoning the people who trusted the earlier version.

One contract, two runtimes

CryptoKit and CommonCrypto in Swift, crypto.subtle in TypeScript. A backup made on an iPhone restores in a browser, byte for byte.

The part that isn’t a technical decision

There is no key escrow and no reset link. Lose the passphrase and that backup is gone — for you and for us. We could add recovery, but we’d have to hold something to do it, and any recovery path we control is a recovery path that can be subpoenaed, breached or handed over.

Encryption of individual fields inside the app — full SSNs and Tax IDs, with a device-bound Keychain key — is a separate scheme, covered on the security page.

  • Services/BackupService.swift
  • Services/PortableArchive.swift
  • web/src/lib/portableExport.ts
— / RULES

The rules underneath all six

Read the diagrams together and the same three decisions keep showing up.

01

Prefer absence to policy

A control can be misconfigured. A component that was never built cannot be. No bank connection, no financial tables in the worker backend, no key escrow — each one is a class of failure that simply has nowhere to happen.

02

Push work to the edge

Parsing, categorizing, OCR and encryption all run on your device. That’s a privacy property and an economic one at the same time: work that happens on your hardware is work nobody has to fund by monetizing you.

03

Say where it stops

Statement import is more manual than a bank feed. The web can’t do private OCR. A lost passphrase is unrecoverable. Every one of those is stated in the product, not just here — a limit you hide is a limit your user discovers at the worst moment.

— / STACK

The stack, per platform

Four native surfaces sharing one data model, one feature-flag system and one sync contract.

SurfaceBuilt withLocal storeSync
iPhone & iPadSwift, SwiftUI, WidgetKit, App Intents, Vision, FoundationModelsSwiftData on deviceCloudKit private zone, or Postgres
MacSwift, SwiftUI — a native app, not a wrapped web viewSwiftData on deviceCloudKit private zone, or Postgres
AndroidKotlin, Jetpack ComposeRoom, encrypted at restPostgres with row-level security
WebReact, TypeScript, Vite — no server rendering, no framework backendBrowser sessionPostgres with row-level security
BackendPostgres, row-level security, a small set of edge functions—Token claims minted by an access-token hook

Feature toggles live on the business profile and sync across all four, so turning off Payroll on your Mac turns it off on your phone. Sensitive material — full Tax IDs, W-9 documents, receipt images — is excluded from sync by design rather than by setting.

Read it, then check it.

Everything above is visible in the app itself — the “what leaves your device” panel shows the same boundaries these diagrams describe.