App reference

ADR-0035: Stored-data evolution rules

Metadata

  • Status: Accepted
  • Date: 2026-08-20
  • Deciders: Eagraí Clainne Team
  • Related: ADR-0004 (the one data migration to date), ADR-0021 (dual-backend storage), ADR-0023 (import/restore)

Context

Schema DDL evolves itself: ApplySchema runs the embedded idempotent scripts on every boot, so a new table or index needs no migration step. The rows inside those tables do not evolve. Entities are protojson documents, and the layer reads them with a strict protojson.Unmarshal (rows.go, table.go) — no DiscardUnknown, no version column, no backfill hook anywhere in boot.

Two operations therefore have no home:

  1. Field rename. protojson matches by field name. Rename a proto field and every stored row still carries the old key — which the new binary silently ignores. The data is not corrupted; it is invisible, which on a family server is the same as lost.
  2. One-time data fix. "For all existing rows, move X to Y" has no place to run and no schema-version table to record that it ran. ADR-0004 needed exactly this once and dodged it: a write-path shim (ReferenceUser auto-inverts CHILD edges) made old and new rows both work, and a hand-run SQL file (internal/database/migrations/0004-…) cleaned history at leisure. That worked because the fix was optional.

Downgrade behaviour is already acceptable and stays so: an old binary ignores tables it does not know, and nothing destructive runs on boot. The strict unmarshal does mean an old binary can fail to read a row that gained a new field — upgrades are safe, rollbacks are not — and that is a deploy-ordering concern, not a migration mechanism.

Decision

The implicit rules become explicit:

  • Proto fields are added, never renamed and never re-numbered. A wrong name is lived with or superseded: add the new field, keep the old one (reserved in prose, populated for compatibility, or shimmed at the seam as ADR-0004 did). Enum values likewise: add values, never rename or renumber existing ones (ADR-0034 is the model).
  • A mandatory data backfill requires building a mechanism first. Any change whose correctness depends on rewriting existing rows must not ship until a tracked migration mechanism (a schema-version table and a boot-time runner, or equivalent) exists — designed in its own ADR. Until then, every data fix must be optional the way ADR-0004's was: the code handles both old and new row shapes, and any cleanup script is idempotent and hand-run.

No mechanism is built now. The decision is the constraint, recorded so that the absence of a migration runner reads as a rule and not a gap.

Consequences

  • Renames stop being a refactor. Reviewers reject a proto field or enum rename on sight; the commit that wants one must instead add-and-shim.
  • Feature work that needs a true backfill grows a prerequisite: design the migration mechanism first. That friction is deliberate.
  • Old field names accumulate in the schema. Acceptable: protojson rows are invisible to users, and reserved names cost nothing at runtime.
  • Rollback across a row-shape change remains unsafe (strict unmarshal). Deploys that add fields to hot records stay upgrade-only, as already learned with the ADR-0029 preference fields.

Alternatives considered

  • Build the migration runner now. A version table plus ordered boot scripts is well-trodden, but nothing needs it yet, and dual-backend DDL (ADR-0021) doubles the surface to get right. Deferred until a change actually requires a mandatory backfill.
  • DiscardUnknown on read to soften rollbacks. Rejected: it would make a downgraded binary silently rewrite rows without the newer fields on the next save, turning a loud read error into quiet data loss. Strict is the safer failure.
  • Do nothing. The rules already held in practice, but only as folk knowledge. A rename is an easy, innocent-looking change; the point of this ADR is that future contributors — human or agent — meet the rule before the mistake.