Cerbero

Offline safety checks for Ecto migrations, judged against a committed snapshot of your production database's catalog metadata — for PostgreSQL and CockroachDB.

The claim, worded precisely: cerbero detects a specific catalog-derivable class of unsafe migrations, judged at export-time scale. It does not certify migrations as safe; it judges the statement, not the moment.

How it works

  1. mix cerbero.snapshot exports catalog metadata (schema shapes, row/byte estimates, traffic counters, index/constraint validity) to priv/repo/cerbero_snapshot.json — a canonical, checksummed, human-diffable artifact you commit (or refresh on a CI cron — recommended past one deploy/week; a manually-refreshed artifact rots).
  2. mix cerbero.check parses pending migrations (static AST — your code never runs), folds their effects into the catalog model, and judges each operation: lock mode × cost class × your production scale and traffic.
  3. CI reads the exit code: 0 clean, 1 unsafe at --fail-on threshold, 2 cerbero misconfigured. No database is reachable from CI.

Example finding:

[error] unsafe_index_creation: SHARE lock blocks writes on public.events
(~412M rows, stats 2026-07-01) for a full-table scan; use concurrently: true
with @disable_ddl_transaction and @disable_migration_lock
(priv/repo/migrations/20260801000000_add_events_payload_index.exs:5)

Privacy boundary

The snapshot contains identifiers, type names, enumerated keywords, booleans, numbers, and timestamps — never expression text, never literals, never row data. pg_stats (histograms, most-common values) is permanently excluded. Every SQL statement the exporter can run fits on one reviewable screen (Cerbero.Snapshot.Exporter.Queries); mix cerbero.snapshot --emit-sql prints it as a script a DBA can run with their own credentials, and --from-file ingests the result.

What is not claimed: table/column names are exported (they already appear in your migration files); row counts and byte sizes are business metrics and will be visible to everyone with repo access, forever — opt into precision: :order_of_magnitude (in .cerbero.exs, or mix cerbero.snapshot --precision order_of_magnitude) to bucket every count and byte to its power-of-ten floor: a 4.1M-row subscriptions table exports as 1_000_000, and findings read "~1M rows". The default severity tiers (100k/1M rows) are powers of ten, so verdicts survive bucketing; the byte tier defaults to 1 GiB, so teams using this mode should prefer a power-of-ten bytes_error. The snapshot records its own precision and the check summary line says so. A hot-standby snapshot has degraded traffic stats (recorded and warned); a scrubbed subset copy produces confidently wrong scale and is incompatible with scale judgment.

The checksum detects corruption and hand-edits. It is not tamper-proofing: anyone who can commit can regenerate it. For tamper-proofing, sign snapshots (mix cerbero.snapshot --sign-key, keygen via --gen-signing-key) and pin the public keys in snapshot_verify_keys — then a regenerated checksum no longer verifies.

Why not excellent_migrations?

We like excellent_migrations. It is AST-only by identity: it judges the migration text with no knowledge of the database, so it must assume the worst everywhere — and teams drown in warnings on tables with 40 rows, annotate everything with skips, and stop reading. Cerbero's differentiation is catalog knowledge:

excellent_migrationscerbero
Non-concurrent indexalways warnsseverity from actual rows/bytes/traffic; silent on small cold tables; correct per-partition recipe on partitioned parents
SET NOT NULLalways warnssilent if already NOT NULL; recognizes the validated-CHECK scan-skip (PG ≥ 12), including its own recommended two-step in raw SQL
Type changecannot see the current typevarchar(50)→varchar(255): info note only (lock_timeout caveat), never blocks CI; int→bigint on 412M rows: error, with the index rebuilds named
FK without indexcannot knowrule impossible without the catalog
ADD FOREIGN KEYflags the statementnames the referenced table whose writes block while the referencing table is scanned
CockroachDBn/aengine-conditional verdicts + the CRDB limitation table (verified against a live node: catches true in-transaction rejections, flags the rest — e.g. type changes — as limitation/cost caveats, not blanket rejections)
Invalid index in prod, stale stats, history divergencen/asnapshot_health
Adoption on a 400-migration repore-annotate historypending-only: zero findings at adoption

Three of cerbero's ten rules are impossible without catalog knowledge; five are severity upgrades where the catalog changes whether CI fails; one is kept for self-consistency, plus snapshot_health, which judges the snapshot itself rather than a migration. If you want AST-only checks with zero setup, excellent_migrations remains the right tool — cerbero's no-snapshot mode (--no-snapshot) gives you a comparable structural baseline plus a trial path.

Requirements

Elixir ≥ 1.18. PostgreSQL ≥ 13 or CockroachDB ≥ v23.1. Postgres and CockroachDB only — no adapter behaviour in v1.

Configuration

.cerbero.exs (all optional; mix cerbero.gen.config writes one with the defaults spelled out and commented):

[
rows_warning: 100_000,
rows_error: 1_000_000,
bytes_error: 1_073_741_824,
hot_ops_per_sec: 1.0,
deploy_cadence: 1, # days between deploys; pending migrations this close to the snapshot aren't flagged as aged
headroom_days: 14, # past this, thresholds shrink by headroom_multiplier
headroom_multiplier: 0.5,
stale_warn_days: 30,
stale_degrade_days: 90, # past this, all scale is treated as unbounded
fail_on: :error,
strict_concurrent_index: false,
lock_timeout_attested: false,
skip_checks: [],
extra_checks: [], # third-party modules implementing Cerbero.Check, run after the built-ins
severity_overrides: %{}, # e.g. %{snapshot_health: :error}
start_after: nil,
precision: :exact, # :order_of_magnitude buckets exported counts/bytes to powers of ten
schemas: ["public"],
snapshot_path: "priv/repo/cerbero_snapshot.json",
migrations_paths: ["priv/repo/migrations"],
repos: [], # umbrella multi-repo: [[name: "app_a", migrations_paths: [...], snapshot_path: "..."]]; check runs all, or one via --repo NAME
snapshot_verify_keys: [] # base64 Ed25519 public keys; when set, snapshots must be signed (mix cerbero.snapshot --sign-key, keygen via --gen-signing-key)
]

Per-migration escape hatch (reason required, findings still shown at info):

@cerbero_skip [{:unsafe_index_creation, "maintenance window 2026-07-20, comms sent"}]

Known limitations

A snapshot is point-in-time; findings carry their stats dates; no wall-clock duration estimates, ever. Pending vs. applied-after-snapshot is offline- indistinguishable — scheduled re-export is the real mitigation. Lock queues, long transactions, and concurrent DDL at deploy time are invisible offline: cerbero judges the statement, not the moment. Dynamically-built operations surface as unknown_operation, never silence. down bodies are judged only on request (mix cerbero.check --down); rollback judgments use the post-up catalog state in version order, not a one-at-a-time unwind.