Documentation

Known limitations — ZigBase

Current caveats in ZigBase — auth/email, framework hooks, schema migrations, fields, the scheduler, platform/UI gaps, and deferred work.

ZigBase v0.13.0 is an early release. The gaps below are known and tracked for future releases.

Auth & email

  • Per-device session list/revoke requires opt-in .auth.session.store = .table; the default is still .epoch. The default App(.{ .auth = .{ .session = .{ .store = .epoch } } }) revokes per principal via the token epoch — revokeAllSessions() (“log out everywhere”), refresh(), rotate() — stateless, with zero extra DB work and unchanged token format. Opting into App(.{ .auth = .{ .session = .{ .store = .table } } }) adds a server-side _sessions store enabling per-device list/revoke, at the cost of one extra read per authenticated request (the documented trade-off; flipping the default awaits real-world perf data). The surface now spans three layers, all gated the same way: the ctx.auth() verbs (listActiveSessions() / revoke(sessionId)), the REST endpoints (GET/DELETE /api/collections/:col/auth/sessions[/:sid]), and the TypeScript SDK (listSessions() / revokeSession(id)). In .epoch mode, ctx.auth() returns error.SessionStoreNotEnabled and the REST/SDK surface returns a non-oracle 404.
  • A mail delivery path must be configured for production. Verification and password-reset email is delivered once a mailer is configured. Zero-code options selectable by env: SMTP (ZIGBASE_SMTP_HOST + friends; none / starttls / implicit TLS — certificates verified by default, ZIGBASE_SMTP_INSECURE disables verification for self-signed relays only), or a local sendmail/msmtp-style command (ZIGBASE_SENDMAIL_COMMAND, e.g. msmtp -t or sendmail -t -i; whitespace-split into argv, and it takes precedence over SMTP when set). Embedded consumers can also swap in the built-in SES / Postmark backends or a custom Mailer plugin via App(.{ .mailer = T }). With none configured, those tokens are logged to the server (the dev/CI default) rather than emailed — a real deployment must configure one of the above.
  • Rate limiting keys on proxy headers only when --trust-proxy is set. Login / verification / password-reset are rate limited. X-Forwarded-For / X-Real-IP are ignored by default (they are spoofable on direct exposure); set --trust-proxy / ZIGBASE_TRUST_PROXY=true only behind a trusted reverse proxy to honor them. Without a trusted proxy the limiter keys on the submitted identity/email (not header-spoofable). The limiter also fails open under memory pressure (it never becomes a self-DoS), so it is a throttle, not a hard guarantee.
  • Durable queue rates use shared database windows. Processes sharing a database share each queue’s .rate allowance, charged atomically with claims. Use consistent queue definitions and synchronized clocks across instances; the ceiling is per integer-second window, not a rolling-second rate.
  • Durable mail delivery is at-least-once: a crash can produce one duplicate send. Both the built-in "mail" job kind and bulk sendBulk delivery persist to a durable queue and mark their row sent/done only after the backend accepts the message; a crash between backend-accept and that row update replays the job on restart, producing one duplicate send to the recipient. Handlers are otherwise idempotent (bulk redelivery of an already-sent/suppressed/canceled recipient row is a no-op) — this window is the one unavoidable exception.
  • No CSS inliner or cid: inline images. ZigBase does not inline <style> rules into style="…" attributes or support MIME cid:-referenced inline images — author inline styles directly (or run a build-time inliner over your template sources) and host images at absolute HTTPS URLs. (Regular download attachments — e.g. a .ics invite — are supported via MailMessage.attachments.) See “HTML that renders everywhere” in framework.md.

Framework / hooks

  • Comptime index declarations are authoritative, including their empty default. Metadata indexes omitted from code are removed on startup. Equivalent ordinary migration-created indexes can be adopted; conflicting or unsupported external definitions (including partial indexes) refuse startup with named diagnostics and require an explicit migration. Failed unique builds also refuse startup after rolling back index changes.
  • Comptime raw-SQL checking (checkSql/checkedSql) validates tables strictly but qualified columns only best-effort. zigbase.checkedSql(Backend.collections, sql) fails the build on an unknown table (after FROM/JOIN/INTO/UPDATE — the UPDATE in an upsert’s ON CONFLICT ... DO UPDATE SET is recognized as a conflict clause, not a table reference) and on a mistyped qualified column (alias.col / table.col) whose qualifier resolves to a known collection. By design — to guarantee zero false compile errors on valid SQL — it does not check unqualified columns, function names, computed AS aliases, alias.*, or the contents of string literals, comments, and FROM (subquery) bodies. It catches the common typo classes, not every mistake; it is a strict-tables / best-effort-columns net, not a full SQL type-checker. Internal/migration-owned tables are opted in via .extra_tables.
  • before-hook side-writes are atomic with the triggering write; copy ctx.records() results into ev.arena before storing them in ev.record. On the HTTP create/update/delete path a before-hook runs inside the triggering write’s transaction, so its ctx.records() side-writes commit atomically with the primary write and roll back together if the hook errors or the access rule denies (fail closed). The one caveat: values returned by ctx.records() (e.g. a .get result) live on the hook’s per-invocation arena, which is freed when the hook returns — not on ev.arena, which owns ev.record and outlives the hook. Copy such a value into ev.arena before storing it into ev.record, or it will dangle.

Schema / migrations

  • Consumer migration coordination is cooperative. PostgreSQL apply/rollback batches require a direct or session-pooled connection, not transaction/statement pooling. SQLite batches use a permanent <canonical-main-database-path>.migrations.lock sidecar and fail fast with MigrationBusy on contention. Use a writable, trusted local directory with working file locks; never replace/unlink the sidecar or database during use, or access the database through hard-link aliases. Individual commit boundaries and non-transactional callbacks are preserved. Old binaries, external SQL writers, automatic collection provisioning and distributed filesystems still need a single leader. Coordination does not make external callback effects atomic or incompatible rolling schema upgrades safe.
  • Comptime auto-migration is additive-only. Startup provisioning of a comptime .collections schema creates missing collections and adds new fields, preserving data. Non-additive changes (rename, drop, or type-change a field) are detected, logged, and skipped — they require an explicit .migrations entry. (Access rules are the exception: a changed rule is re-applied on every startup, as a metadata-only write.)
  • Schema rule preflight is structural, not a proof of authorization behavior. schema apply (including dry runs) and REST collection mutations reject invalid syntax and unresolved field/relation references against the prospective schema. Request-dependent values and application policy still require runtime tests. Trusted low-level collection writers can stage incomplete graphs; run zigbase schema check-rules --data-dir PATH against their final state before cutover.
  • A live server picks up another process’s schema change within about five seconds, not instantly. schema apply, migrate, and import run as separate CLI processes; the server notices via a background observer that polls the _schema_state generation marker (~5s cadence) and then drops its in-process collection cache. Requests in that window still see the pre-change view. A cutover that must be seen immediately should still restart the server. See Migrating an existing backend to ZigBase.
  • Collection rename requires an explicit offline migration. m.renameCollection(from, to, .{ .offline = true }) preserves identities, database/auth/search references, and immutable file namespaces; document diffs do not infer renames. Stop all processes and update compiled name references; old URLs/topics are not aliased. Namespace reservations survive deletion, so reserved physical prefixes cannot be claimed by a new collection. See offline collection rename.
  • import --manifest’s deferred relation values (cross-collection cycles, self-relations) require the source row to carry its own id, and a required field on one of those relations is refused up front — see Migrating an existing backend to ZigBase.

Scheduler

  • Retained memory-job byte budgets cover copies, not total memory. Optional .admission.max_job_bytes (with -Dcoordinated-admission=true) limits queue-owned payload and app.submit name lengths through retries. It can run without HTTP admission (max_requests omitted). It excludes caller serialization, inline borrowed payloads, task/allocator overhead, handler allocations, stacks and durable queues. Empty payloads consume no byte budget; use the independent work-count ceiling and existing ring/worker bounds to constrain task counts.
  • Cron/interval scheduling is per-process by default. Opt-in comptime .distributed jobs share persisted scheduling and expiring ownership across a database, with crash recovery and stale-completion fencing. Leases do not heartbeat or cancel application code: overruns can overlap replacement handlers, and external effects still require idempotency. Reactive jobs remain per-process. See distributed scheduling.
  • Cron is UTC, minute-granularity, and ANDs day-of-month with day-of-week (unlike Vixie cron, which ORs them when one is *). Fields accept case-insensitive 3-letter month (JAN..DEC) and day-of-week (SUN..SAT) names as well as numbers, plus <lo>-<hi>/<step> ranges; day-of-week 7 is a Sunday alias (0 and 7 both mean Sunday). A comptime .cron string is validated at compile time — a malformed or out-of-range expression is a build error, not a job that silently never fires.
  • Interval jobs fire measured from completion, so a long-running interval job drifts from a fixed wall-clock cadence. Runs never overlap (single-flight).
  • app.submit and memory-queue jobs run on a small bounded in-process worker pool (drained and joined at shutdown, so tasks submitted before shutdown complete). The ring is bounded: overflow is rejected with error.QueueFull rather than blocking the caller, and a retrying memory job holds one pool worker for its whole backoff — put sustained high-volume or long-retry work on a durable queue. Optional coordinated admission shares a process-local work-count limit with HTTP, not a byte/RSS budget; retries retain a permit. There is no fairness guarantee or reserved HTTP capacity, so outstanding jobs can reject HTTP until capacity returns. Serial durable poll batches also reserve one work permit; retained durable rows use a separate optional database capacity. Durable claim payload allocations, transport buffering and payload serialization before enqueue are outside the memory-byte budget.
  • Background-job queue delivery semantics (.queues/ctx.enqueue). Durable queues are at-least-once: a job whose claim exceeds its queue’s visibility_timeout_s, or that crashes after a side effect but before completion, may run more than once — make durable handlers idempotent (set visibility_timeout_s above the entire serial batch’s worst-case processing time, and use an idempotency key for external effects). Durable workers poll (~0.5s cadence), so jobs drain with low but non-zero latency, and a worker’s concurrency is a per-cycle batch processed serially (parallelism = more workers). There is no visibility heartbeat. Memory queues are at-most-once across restart (in-RAM only; lost on crash/shutdown) and run on the bounded worker pool (a full ring rejects new jobs with error.QueueFull); use a durable queue when a job must survive a restart or for sustained throughput.
  • TTL record physical deletion is eventually consistent (~5 minutes), but expired rows are hidden from reads immediately. Expired rows in a .ttl_field collection are automatically excluded from every read (list, get, relation expand — via the HTTP API and ctx.records()) by a read-time predicate that is ANDed with your filter, access rule, and keyset cursor, so no manual expires_at > @now filter is needed. Physical deletion is still handled by an internal GC job that runs once at startup and then on the .ttl_gc_interval cadence (default every 5 minutes, tunable via the comptime .ttl_gc_interval App config key), so an expired row may persist in the table until the next sweep — but it is never returned by a read in the meantime. A row whose ttl field is null never expires; a row with an unparseable ttl value is fail-safe (stays visible, never reaped).

Testing & determinism

  • The ZIGBASE_FAKE_NOW test clock freezes the framework’s clock AND consumer SQL 'now'. Setting ZIGBASE_FAKE_NOW to an ISO-8601 UTC instant (dev builds only) freezes every “now” the framework controls — token iat/exp, the scheduler’s next-fire math, the auth rate-limiter wall clock, and the auth-challenge / keyset-cursor TTL and expiry checks. As of the consumer-SQL determinism work, it also freezes a consumer’s own raw datetime('now') / unixepoch('now') / strftime(…, 'now') (and date/time/julianday, including their zero-argument implicit-'now' forms): those date/time builtins are shadowed on every reader and writer connection so they resolve to the frozen instant, while every other input (explicit datetimes, '+1 day' modifiers, the strftime format string) passes through to genuine SQLite. It also freezes the SQL keywords CURRENT_TIMESTAMP / CURRENT_TIME / CURRENT_DATE and column DEFAULT CURRENT_TIMESTAMP timestamps: those read SQLite’s clock through the VFS rather than the SQL-function layer, so (dev builds only) connections open against a wrapping VFS that is a byte-for-byte copy of the default VFS with only its current-time hooks overridden to return the frozen instant — all file I/O still delegates to the genuine OS VFS unchanged. There are no remaining unfrozen 'now' paths.
  • The test clock is impossible to enable on a production build. It is compiled in only when the dev_mode build option is true (on in Debug, off in any release build; the release script ships it off) — the same gate that covers all dev-only seams (frozen clock, seeded entropy, test-capture). A production binary never reads ZIGBASE_FAKE_NOW — the override folds to a comptime no-op — so time can never be frozen in production.

Platform & UI

  • No Windows build — Linux and macOS only (the embedded HTTP server depends on facil.io/zap). The official Docker image (ghcr.io/valthon/zigbase, see docs/docker.md) is the supported path on Windows hosts.
  • Admin UI: The Logs & realtime view is capability-gated — it appears only when the app enables .analytics (so it’s hidden on the stock zigbase serve binary).

Static file serving

  • Static files are served without authentication — collection access rules do not apply to the static root; use file storage for access-controlled delivery.
  • No directory listings; directories resolve to index.html or 404.
  • Path safety is lexical (.., backslashes, and NUL bytes are rejected) and symlink-aware: the request path is percent-decoded (single-pass, in-house) before those lexical checks — so files whose names need encoding (my file.pdf → /my%20file.pdf) are servable, while encoded traversal (%2e%2e, %2f, %00, %5c) is decoded and then rejected fail-closed, and double-encoding is never recursively decoded (%252e → the literal %2e, not .). A served file is additionally canonicalized and refused if its real path escapes the configured static root, so a symlink inside the root pointing outside it is not followed out (F10).
  • No on-the-fly compression; pre-compress at the CDN or reverse proxy if needed.
  • In dir mode, conditional requests (If-None-Match/If-Range) are delegated to facil.io’s sendFile, which uses its own exact-match ETag semantics (an unquoted base64 size^mtime tag), not RFC 7232 list/weak comparison. A ranged request with a matching If-Range correctly resumes (206) — zigbase neutralizes an inverted branch in the vendored facil.io that otherwise deleted the Range on a match and forced a full 200 (RFC 9110 §13.1.5). The one residual: a ranged request with a stale/mismatched If-Range still returns 206 rather than the RFC-mandated full 200, because facil.io’s dir-mode ETag is seeded from process-local (ASLR’d) addresses and so cannot be recomputed and compared in zigbase’s layer. Owned file serving is fully RFC-correct here: record-file downloads (/api/files/…) and embedded static assets mint and strong-compare their own ETag, so a mismatched If-Range there ignores the Range and returns 200.

Postgres backend

  • verify-full hostname checks match DNS names only. Dialing an IP literal under the default sslmode=verify-full generally fails hostname verification even when the certificate carries an iPAddress SAN — connect by DNS name, or use sslmode=verify-ca on an otherwise-trusted path. Client certificates (mTLS), CRL/OCSP, and SCRAM channel binding (SCRAM-SHA-256-PLUS) are not supported.
  • SCRAM passwords that require NFKC normalization are rejected, not normalized. The driver implements RFC 4013 SASLprep except the final NFKC normalization step: a password whose SASLprep output would need real NFKC normalization (e.g. U+2168 ROMAN NUMERAL NINE → IX, or U+00AA FEMININE ORDINAL INDICATOR → a) fails at connect with error.PasswordNeedsNormalization rather than being normalized. Real PostgreSQL’s pg_saslprep would normalize and accept such a password; ZigBase deliberately hard-errors instead — a loud, actionable failure that names the fix is chosen over silently deriving a wrong SCRAM hash and returning an opaque auth failure. Supply the password pre-normalized to NFKC, or use an ASCII password. (Everything else is correctly prepped or intentionally matches PostgreSQL’s own use-verbatim behavior.)
  • migrate-db: the superuser fast path is faster; the non-superuser path is fully supported. A superuser target suspends FK enforcement wholesale; a non-superuser target provisions cycle-edge FKs as deferrable and defers them to COMMIT — correct for cyclic and self-referential graphs, verified against live Postgres in CI.

Resumable uploads (-Dresumable-uploads)

  • RAM by default, always fully buffered. Default sessions are lost on restart and unknown to other instances. An additional -Ddurable-resumable-uploads=true plus .files.resumable.durable = true persists bounded SQLite/local sessions across process restarts, with one owner and no uncertain hook replay. It does not provide power-loss durability, cross-instance resume, or streaming. Retained-payload limits do not bound RSS or DB/WAL size. Staged bytes enter database backups without file-field encryption; logical deletion is not secure erasure. See persistence guarantees and tradeoffs.
  • Terminal acknowledgements consume slots until expiry. Completed/failed sessions immediately free payload bytes but cannot be aborted; DELETE returns 409 and never undoes a mutation. Their slots remain until the original fixed TTL expires, reclaimed lazily on API calls. With defaults, two quick commit attempts exhaust a principal’s two slots for up to 900 seconds. Tune session counts and TTL for expected throughput and acknowledgement retention.
  • Resource bounds are not fairness or abuse protection. Per-principal session count times maximum upload size is an implicit byte bound, not an independent byte quota. Multiple principals can exhaust aggregate slots/bytes (four principals with two 4 MiB reservations each exhaust the default eight slots and 32 MiB). Restrict upload authorization and request rates. Receiving sessions can be aborted to release capacity; otherwise capacity returns on expiry, and an in-flight commit stays pinned even beyond TTL. See the protocol and sizing tradeoffs.

S3 storage (-Ds3)

  • Upload bytes and database references are not one distributed transaction. HTTP record uploads PUT bytes before acquiring the writer, then validate and commit their references transactionally. Failed requests best-effort remove uploaded bytes; a crash between PUT and commit can still leave an orphan. Upload POST returns 409 on collection-definition changes during transfer; PATCH also returns 409 on record changes. Reload and retry. Replacement/deletion cleanup is synchronous by default; opt into durable HTTP cleanup to enqueue it transactionally. The worker holds a database write lock during each storage deletion, so slow storage can still delay concurrent writes.
  • Rejected creates can incur storage requests. File constraints and locked create rules are checked before transfer, but ordinary field validation, before-hooks and data-dependent create rules run afterward in the transaction. Even an anonymous request allowed through preflight can incur billable PUT and cleanup DELETE operations before returning 400 or 403. Bound upload sizes and apply request rate limits; storage plugins must not assume a record exists during PUT.
  • Deletes are best-effort by default. Opt-in durable HTTP cleanup retries replaced-object and deleted-record-prefix deletion but does not cover raw SQL, Data, cascade/TTL, collection deletion, or failed uploads. Prefix cleanup requires physical row absence; reused IDs retain their objects, including expired rows. Exhausted retries require operator attention; removed/recreated collections are conservatively skipped. Opt-in files inventory reports current sizes and reference candidates, not proof that objects are safe to remove (in-flight uploads can be unreferenced). Local scans stop at 100,000 entries and exclude symlinks/deeper-than-record-layout paths; S3 reports current objects in the configured prefix, not versions or multipart state. See storage inventory.
  • Orphan deletion is offline local/SQLite maintenance, not a background job. Opt-in files reconcile defaults to read-only dry-run. Explicit apply requires an exclusive storage lease and SQLite writer transaction; cooperating apps (including serve --ignore-lock) hold one shared boot-lifetime descriptor even in default builds. Stop older/external writers yourself and use a root dedicated to one database; age alone cannot close the PUT-before-commit race. Never replace the permanent lock file. S3/PostgreSQL/custom storage, unknown metadata, encrypted file fields and dropped-collection prefixes are outside deletion scope. Batches are bounded and cursors are not approvals/snapshots. Partial filesystem deletion cannot be rolled back, including when a later output failure prevents the report from being emitted.
  • Proxy serving is the default; presigned-redirect is opt-in. By default downloads flow through the server’s spool cache (§D.6), so every download consumes server bandwidth/CPU even though the bytes ultimately come from S3. An opt-in presigned-redirect mode is now available via the comptime App(.{ .files = .{ .s3_presign_redirect = true } }) option (S3 backend only; default off = proxy): an authorized download is answered with a 302 to a time-limited presigned GET URL (s3_presign_ttl_s, default 900s) instead of proxying the bytes. Tradeoff: per-request authorization still runs before the redirect is issued, but the issued URL is a bearer capability — valid for s3_presign_ttl_s and not bound to the authorized requester, so anyone holding it can fetch the object until it expires. Keep the TTL short.
  • Multipart does not make uploads memory-streaming. S3 transfers large objects in bounded sequential parts, but Storage.put and HTTP ingestion still require the entire input in memory. Incomplete uploads can survive crashes, lost initiation responses, or failed aborts; configure bucket lifecycle expiration. A lost completion response can leave a completed orphan object. Opt-in client-resumable uploads only resume inbound transfers within one running ZigBase process; they are not durable S3 upload sessions.
  • Spool cache disk usage + eviction approximates last-access LRU, not strict LRU. The local spool cache (ZIGBASE_S3_CACHE_DIR, default <data-dir>/storage_cache) needs its own disk budget on top of the database; size it via ZIGBASE_S3_CACHE_MAX_BYTES. Eviction still sorts by file mtime, but mtime is now bumped on a cache hit as well as on the miss-fill that creates the entry, so eviction approximates last-access LRU — a frequently-read file survives over a rarely-read one. It is still not a strict LRU: there is no separate access-time bookkeeping or journaling, and the mtime touch is best-effort and coarse-grained.

Other deferred work

  • Idempotency is opt-in custom-operation support, not global middleware. zigbase.Idempotency coordinates database-only effects and saved result bytes on SQLite/PostgreSQL; callbacks must be trusted, authorization read-only, and current access checked on every replay. PostgreSQL serializes each namespace with fail-fast transaction locks. Retention begins at the trusted attempt timestamp, not commit, and expired keys may execute again. Idle receipts are not physically deleted by a timer; database copies refuse non-empty receipt ledgers. No automatic REST hook/rule wrapper, external exactly-once guarantee, or database/WAL/RSS hard limit is provided. See the contract.
  • Image transforms beyond named local ImageMagick thumbnails, including formats other than PNG/JPEG/WebP, remote source backends and persistent derivative storage. Thumbnail-enabled deployments must install and maintain a trusted ImageMagick executable; process limits are not a hard RSS sandbox. Cross-instance or streaming resumable uploads; durable/cross-instance realtime replay and per-event-guard load-tuning. Opt-in SQLite record invalidation backfill is process-local, bounded, and requires full reloads for gaps and authorization/query-dependency changes (see docs/api.md).

These are tracked for upcoming releases. Contributions welcome.