ADR-0035: Stored-data evolution rules
Metadata
- Status: Accepted
- Date: 2026-08-20
- Deciders: Eagraí Clainne Team
- Related: ADR-0004 (the one data migration to date), ADR-0021 (dual-backend storage), ADR-0023 (import/restore)
Context
Schema DDL evolves itself: ApplySchema runs the embedded idempotent
scripts on every boot, so a new table or index needs no migration step.
The rows inside those tables do not evolve. Entities are protojson
documents, and the layer reads them with a strict protojson.Unmarshal
(rows.go, table.go) — no DiscardUnknown, no version column, no
backfill hook anywhere in boot.
Two operations therefore have no home:
- Field rename. protojson matches by field name. Rename a proto field and every stored row still carries the old key — which the new binary silently ignores. The data is not corrupted; it is invisible, which on a family server is the same as lost.
- One-time data fix. "For all existing rows, move X to Y" has no
place to run and no schema-version table to record that it ran.
ADR-0004 needed exactly this once and dodged it: a write-path shim
(
ReferenceUserauto-inverts CHILD edges) made old and new rows both work, and a hand-run SQL file (internal/database/migrations/0004-…) cleaned history at leisure. That worked because the fix was optional.
Downgrade behaviour is already acceptable and stays so: an old binary ignores tables it does not know, and nothing destructive runs on boot. The strict unmarshal does mean an old binary can fail to read a row that gained a new field — upgrades are safe, rollbacks are not — and that is a deploy-ordering concern, not a migration mechanism.
Decision
The implicit rules become explicit:
- Proto fields are added, never renamed and never re-numbered. A wrong name is lived with or superseded: add the new field, keep the old one (reserved in prose, populated for compatibility, or shimmed at the seam as ADR-0004 did). Enum values likewise: add values, never rename or renumber existing ones (ADR-0034 is the model).
- A mandatory data backfill requires building a mechanism first. Any change whose correctness depends on rewriting existing rows must not ship until a tracked migration mechanism (a schema-version table and a boot-time runner, or equivalent) exists — designed in its own ADR. Until then, every data fix must be optional the way ADR-0004's was: the code handles both old and new row shapes, and any cleanup script is idempotent and hand-run.
No mechanism is built now. The decision is the constraint, recorded so that the absence of a migration runner reads as a rule and not a gap.
Consequences
- Renames stop being a refactor. Reviewers reject a proto field or enum rename on sight; the commit that wants one must instead add-and-shim.
- Feature work that needs a true backfill grows a prerequisite: design the migration mechanism first. That friction is deliberate.
- Old field names accumulate in the schema. Acceptable: protojson rows are invisible to users, and reserved names cost nothing at runtime.
- Rollback across a row-shape change remains unsafe (strict unmarshal). Deploys that add fields to hot records stay upgrade-only, as already learned with the ADR-0029 preference fields.
Alternatives considered
- Build the migration runner now. A version table plus ordered boot scripts is well-trodden, but nothing needs it yet, and dual-backend DDL (ADR-0021) doubles the surface to get right. Deferred until a change actually requires a mandatory backfill.
DiscardUnknownon read to soften rollbacks. Rejected: it would make a downgraded binary silently rewrite rows without the newer fields on the next save, turning a loud read error into quiet data loss. Strict is the safer failure.- Do nothing. The rules already held in practice, but only as folk knowledge. A rename is an easy, innocent-looking change; the point of this ADR is that future contributors — human or agent — meet the rule before the mistake.