Documentation
Known limitations — ZigBase
Current caveats in ZigBase — auth/email, framework hooks, schema migrations, fields, the scheduler, platform/UI gaps, and deferred work.
ZigBase v0.13.0 is an early release. The gaps below are known and tracked for future releases.
Auth & email
- Per-device session list/revoke requires opt-in
.auth.session.store = .table; the default is still.epoch. The defaultApp(.{ .auth = .{ .session = .{ .store = .epoch } } })revokes per principal via the token epoch —revokeAllSessions()(“log out everywhere”),refresh(),rotate()— stateless, with zero extra DB work and unchanged token format. Opting intoApp(.{ .auth = .{ .session = .{ .store = .table } } })adds a server-side_sessionsstore enabling per-device list/revoke, at the cost of one extra read per authenticated request (the documented trade-off; flipping the default awaits real-world perf data). The surface now spans three layers, all gated the same way: thectx.auth()verbs (listActiveSessions()/revoke(sessionId)), the REST endpoints (GET/DELETE /api/collections/:col/auth/sessions[/:sid]), and the TypeScript SDK (listSessions()/revokeSession(id)). In.epochmode,ctx.auth()returnserror.SessionStoreNotEnabledand the REST/SDK surface returns a non-oracle404. - A mail delivery path must be configured for production. Verification and password-reset email is delivered once a mailer is configured. Zero-code options selectable by env: SMTP (
ZIGBASE_SMTP_HOST+ friends;none/starttls/implicitTLS — certificates verified by default,ZIGBASE_SMTP_INSECUREdisables verification for self-signed relays only), or a local sendmail/msmtp-style command (ZIGBASE_SENDMAIL_COMMAND, e.g.msmtp -torsendmail -t -i; whitespace-split into argv, and it takes precedence over SMTP when set). Embedded consumers can also swap in the built-in SES / Postmark backends or a customMailerplugin viaApp(.{ .mailer = T }). With none configured, those tokens are logged to the server (the dev/CI default) rather than emailed — a real deployment must configure one of the above. - Rate limiting keys on proxy headers only when
--trust-proxyis set. Login / verification / password-reset are rate limited.X-Forwarded-For/X-Real-IPare ignored by default (they are spoofable on direct exposure); set--trust-proxy/ZIGBASE_TRUST_PROXY=trueonly behind a trusted reverse proxy to honor them. Without a trusted proxy the limiter keys on the submitted identity/email (not header-spoofable). The limiter also fails open under memory pressure (it never becomes a self-DoS), so it is a throttle, not a hard guarantee. - Durable queue rates use shared database windows. Processes sharing a database share each queue’s
.rateallowance, charged atomically with claims. Use consistent queue definitions and synchronized clocks across instances; the ceiling is per integer-second window, not a rolling-second rate. - Durable mail delivery is at-least-once: a crash can produce one duplicate send. Both the built-in
"mail"job kind and bulksendBulkdelivery persist to a durable queue and mark their rowsent/doneonly after the backend accepts the message; a crash between backend-accept and that row update replays the job on restart, producing one duplicate send to the recipient. Handlers are otherwise idempotent (bulk redelivery of an already-sent/suppressed/canceledrecipient row is a no-op) — this window is the one unavoidable exception. - No CSS inliner or
cid:inline images. ZigBase does not inline<style>rules intostyle="…"attributes or support MIMEcid:-referenced inline images — author inline styles directly (or run a build-time inliner over your template sources) and host images at absolute HTTPS URLs. (Regular download attachments — e.g. a.icsinvite — are supported viaMailMessage.attachments.) See “HTML that renders everywhere” in framework.md.
Framework / hooks
- Comptime index declarations are authoritative, including their empty default. Metadata indexes omitted from code are removed on startup. Equivalent ordinary migration-created indexes can be adopted; conflicting or unsupported external definitions (including partial indexes) refuse startup with named diagnostics and require an explicit migration. Failed unique builds also refuse startup after rolling back index changes.
- Comptime raw-SQL checking (
checkSql/checkedSql) validates tables strictly but qualified columns only best-effort.zigbase.checkedSql(Backend.collections, sql)fails the build on an unknown table (afterFROM/JOIN/INTO/UPDATE— theUPDATEin an upsert’sON CONFLICT ... DO UPDATE SETis recognized as a conflict clause, not a table reference) and on a mistyped qualified column (alias.col/table.col) whose qualifier resolves to a known collection. By design — to guarantee zero false compile errors on valid SQL — it does not check unqualified columns, function names, computedASaliases,alias.*, or the contents of string literals, comments, andFROM (subquery)bodies. It catches the common typo classes, not every mistake; it is a strict-tables / best-effort-columns net, not a full SQL type-checker. Internal/migration-owned tables are opted in via.extra_tables. before-hook side-writes are atomic with the triggering write; copyctx.records()results intoev.arenabefore storing them inev.record. On the HTTP create/update/delete path abefore-hook runs inside the triggering write’s transaction, so itsctx.records()side-writes commit atomically with the primary write and roll back together if the hook errors or the access rule denies (fail closed). The one caveat: values returned byctx.records()(e.g. a.getresult) live on the hook’s per-invocation arena, which is freed when the hook returns — not onev.arena, which ownsev.recordand outlives the hook. Copy such a value intoev.arenabefore storing it intoev.record, or it will dangle.
Schema / migrations
- Consumer migration coordination is cooperative. PostgreSQL apply/rollback batches require a direct or session-pooled connection, not transaction/statement pooling. SQLite batches use a permanent
<canonical-main-database-path>.migrations.locksidecar and fail fast withMigrationBusyon contention. Use a writable, trusted local directory with working file locks; never replace/unlink the sidecar or database during use, or access the database through hard-link aliases. Individual commit boundaries and non-transactional callbacks are preserved. Old binaries, external SQL writers, automatic collection provisioning and distributed filesystems still need a single leader. Coordination does not make external callback effects atomic or incompatible rolling schema upgrades safe. - Comptime auto-migration is additive-only. Startup provisioning of a comptime
.collectionsschema creates missing collections and adds new fields, preserving data. Non-additive changes (rename, drop, or type-change a field) are detected, logged, and skipped — they require an explicit.migrationsentry. (Access rules are the exception: a changed rule is re-applied on every startup, as a metadata-only write.) - Schema rule preflight is structural, not a proof of authorization behavior.
schema apply(including dry runs) and REST collection mutations reject invalid syntax and unresolved field/relation references against the prospective schema. Request-dependent values and application policy still require runtime tests. Trusted low-level collection writers can stage incomplete graphs; runzigbase schema check-rules --data-dir PATHagainst their final state before cutover. - A live server picks up another process’s schema change within about five seconds, not instantly.
schema apply,migrate, andimportrun as separate CLI processes; the server notices via a background observer that polls the_schema_stategeneration marker (~5s cadence) and then drops its in-process collection cache. Requests in that window still see the pre-change view. A cutover that must be seen immediately should still restart the server. See Migrating an existing backend to ZigBase. - Collection rename requires an explicit offline migration.
m.renameCollection(from, to, .{ .offline = true })preserves identities, database/auth/search references, and immutable file namespaces; document diffs do not infer renames. Stop all processes and update compiled name references; old URLs/topics are not aliased. Namespace reservations survive deletion, so reserved physical prefixes cannot be claimed by a new collection. See offline collection rename. import --manifest’s deferred relation values (cross-collection cycles, self-relations) require the source row to carry its ownid, and arequiredfield on one of those relations is refused up front — see Migrating an existing backend to ZigBase.
Scheduler
- Retained memory-job byte budgets cover copies, not total memory. Optional
.admission.max_job_bytes(with-Dcoordinated-admission=true) limits queue-owned payload andapp.submitname lengths through retries. It can run without HTTP admission (max_requestsomitted). It excludes caller serialization, inline borrowed payloads, task/allocator overhead, handler allocations, stacks and durable queues. Empty payloads consume no byte budget; use the independent work-count ceiling and existing ring/worker bounds to constrain task counts. - Cron/interval scheduling is per-process by default. Opt-in comptime
.distributedjobs share persisted scheduling and expiring ownership across a database, with crash recovery and stale-completion fencing. Leases do not heartbeat or cancel application code: overruns can overlap replacement handlers, and external effects still require idempotency. Reactive jobs remain per-process. See distributed scheduling. - Cron is UTC, minute-granularity, and ANDs day-of-month with day-of-week (unlike Vixie cron, which ORs them when one is
*). Fields accept case-insensitive 3-letter month (JAN..DEC) and day-of-week (SUN..SAT) names as well as numbers, plus<lo>-<hi>/<step>ranges; day-of-week7is a Sunday alias (0and7both mean Sunday). A comptime.cronstring is validated at compile time — a malformed or out-of-range expression is a build error, not a job that silently never fires. - Interval jobs fire measured from completion, so a long-running interval job drifts from a fixed wall-clock cadence. Runs never overlap (single-flight).
app.submitand memory-queue jobs run on a small bounded in-process worker pool (drained and joined at shutdown, so tasks submitted before shutdown complete). The ring is bounded: overflow is rejected witherror.QueueFullrather than blocking the caller, and a retrying memory job holds one pool worker for its whole backoff — put sustained high-volume or long-retry work on a durable queue. Optional coordinated admission shares a process-local work-count limit with HTTP, not a byte/RSS budget; retries retain a permit. There is no fairness guarantee or reserved HTTP capacity, so outstanding jobs can reject HTTP until capacity returns. Serial durable poll batches also reserve one work permit; retained durable rows use a separate optional database capacity. Durable claim payload allocations, transport buffering and payload serialization before enqueue are outside the memory-byte budget.- Background-job queue delivery semantics (
.queues/ctx.enqueue). Durable queues are at-least-once: a job whose claim exceeds its queue’svisibility_timeout_s, or that crashes after a side effect but before completion, may run more than once — make durable handlers idempotent (setvisibility_timeout_sabove the entire serial batch’s worst-case processing time, and use an idempotency key for external effects). Durable workers poll (~0.5s cadence), so jobs drain with low but non-zero latency, and a worker’sconcurrencyis a per-cycle batch processed serially (parallelism = more workers). There is no visibility heartbeat. Memory queues are at-most-once across restart (in-RAM only; lost on crash/shutdown) and run on the bounded worker pool (a full ring rejects new jobs witherror.QueueFull); use a durable queue when a job must survive a restart or for sustained throughput. - TTL record physical deletion is eventually consistent (~5 minutes), but expired rows are hidden from reads immediately. Expired rows in a
.ttl_fieldcollection are automatically excluded from every read (list, get, relation expand — via the HTTP API andctx.records()) by a read-time predicate that is ANDed with your filter, access rule, and keyset cursor, so no manualexpires_at > @nowfilter is needed. Physical deletion is still handled by an internal GC job that runs once at startup and then on the.ttl_gc_intervalcadence (default every 5 minutes, tunable via the comptime.ttl_gc_intervalApp config key), so an expired row may persist in the table until the next sweep — but it is never returned by a read in the meantime. A row whose ttl field isnullnever expires; a row with an unparseable ttl value is fail-safe (stays visible, never reaped).
Testing & determinism
- The
ZIGBASE_FAKE_NOWtest clock freezes the framework’s clock AND consumer SQL'now'. SettingZIGBASE_FAKE_NOWto an ISO-8601 UTC instant (dev builds only) freezes every “now” the framework controls — tokeniat/exp, the scheduler’s next-fire math, the auth rate-limiter wall clock, and the auth-challenge / keyset-cursor TTL and expiry checks. As of the consumer-SQL determinism work, it also freezes a consumer’s own rawdatetime('now')/unixepoch('now')/strftime(…, 'now')(anddate/time/julianday, including their zero-argument implicit-'now'forms): those date/time builtins are shadowed on every reader and writer connection so they resolve to the frozen instant, while every other input (explicit datetimes,'+1 day'modifiers, thestrftimeformat string) passes through to genuine SQLite. It also freezes the SQL keywordsCURRENT_TIMESTAMP/CURRENT_TIME/CURRENT_DATEand columnDEFAULT CURRENT_TIMESTAMPtimestamps: those read SQLite’s clock through the VFS rather than the SQL-function layer, so (dev builds only) connections open against a wrapping VFS that is a byte-for-byte copy of the default VFS with only its current-time hooks overridden to return the frozen instant — all file I/O still delegates to the genuine OS VFS unchanged. There are no remaining unfrozen'now'paths. - The test clock is impossible to enable on a production build. It is compiled in only when the
dev_modebuild option is true (on inDebug, off in any release build; the release script ships it off) — the same gate that covers all dev-only seams (frozen clock, seeded entropy, test-capture). A production binary never readsZIGBASE_FAKE_NOW— the override folds to a comptime no-op — so time can never be frozen in production.
Platform & UI
- No Windows build — Linux and macOS only (the embedded HTTP server depends on facil.io/zap). The official Docker image (
ghcr.io/valthon/zigbase, see docs/docker.md) is the supported path on Windows hosts. - Admin UI: The Logs & realtime view is capability-gated — it appears only when the app enables
.analytics(so it’s hidden on the stockzigbase servebinary).
Static file serving
- Static files are served without authentication — collection access rules do not apply to the static root; use file storage for access-controlled delivery.
- No directory listings; directories resolve to
index.htmlor 404. - Path safety is lexical (
.., backslashes, and NUL bytes are rejected) and symlink-aware: the request path is percent-decoded (single-pass, in-house) before those lexical checks — so files whose names need encoding (my file.pdf→/my%20file.pdf) are servable, while encoded traversal (%2e%2e,%2f,%00,%5c) is decoded and then rejected fail-closed, and double-encoding is never recursively decoded (%252e→ the literal%2e, not.). A served file is additionally canonicalized and refused if its real path escapes the configured static root, so a symlink inside the root pointing outside it is not followed out (F10). - No on-the-fly compression; pre-compress at the CDN or reverse proxy if needed.
- In dir mode, conditional requests (
If-None-Match/If-Range) are delegated to facil.io’ssendFile, which uses its own exact-match ETag semantics (an unquoted base64 size^mtime tag), not RFC 7232 list/weak comparison. A ranged request with a matchingIf-Rangecorrectly resumes (206) — zigbase neutralizes an inverted branch in the vendored facil.io that otherwise deleted theRangeon a match and forced a full200(RFC 9110 §13.1.5). The one residual: a ranged request with a stale/mismatchedIf-Rangestill returns206rather than the RFC-mandated full200, because facil.io’s dir-mode ETag is seeded from process-local (ASLR’d) addresses and so cannot be recomputed and compared in zigbase’s layer. Owned file serving is fully RFC-correct here: record-file downloads (/api/files/…) and embedded static assets mint and strong-compare their ownETag, so a mismatchedIf-Rangethere ignores theRangeand returns200.
Postgres backend
verify-fullhostname checks match DNS names only. Dialing an IP literal under the defaultsslmode=verify-fullgenerally fails hostname verification even when the certificate carries an iPAddress SAN — connect by DNS name, or usesslmode=verify-caon an otherwise-trusted path. Client certificates (mTLS), CRL/OCSP, and SCRAM channel binding (SCRAM-SHA-256-PLUS) are not supported.- SCRAM passwords that require NFKC normalization are rejected, not normalized. The driver implements RFC 4013 SASLprep except the final NFKC normalization step: a password whose SASLprep output would need real NFKC normalization (e.g. U+2168 ROMAN NUMERAL NINE →
IX, or U+00AA FEMININE ORDINAL INDICATOR →a) fails at connect witherror.PasswordNeedsNormalizationrather than being normalized. Real PostgreSQL’spg_saslprepwould normalize and accept such a password; ZigBase deliberately hard-errors instead — a loud, actionable failure that names the fix is chosen over silently deriving a wrong SCRAM hash and returning an opaque auth failure. Supply the password pre-normalized to NFKC, or use an ASCII password. (Everything else is correctly prepped or intentionally matches PostgreSQL’s own use-verbatim behavior.) migrate-db: the superuser fast path is faster; the non-superuser path is fully supported. A superuser target suspends FK enforcement wholesale; a non-superuser target provisions cycle-edge FKs as deferrable and defers them to COMMIT — correct for cyclic and self-referential graphs, verified against live Postgres in CI.
Resumable uploads (-Dresumable-uploads)
- RAM by default, always fully buffered. Default sessions are lost on restart and unknown to other instances. An additional
-Ddurable-resumable-uploads=trueplus.files.resumable.durable = truepersists bounded SQLite/local sessions across process restarts, with one owner and no uncertain hook replay. It does not provide power-loss durability, cross-instance resume, or streaming. Retained-payload limits do not bound RSS or DB/WAL size. Staged bytes enter database backups without file-field encryption; logical deletion is not secure erasure. See persistence guarantees and tradeoffs. - Terminal acknowledgements consume slots until expiry. Completed/failed sessions immediately free payload bytes but cannot be aborted; DELETE returns
409and never undoes a mutation. Their slots remain until the original fixed TTL expires, reclaimed lazily on API calls. With defaults, two quick commit attempts exhaust a principal’s two slots for up to 900 seconds. Tune session counts and TTL for expected throughput and acknowledgement retention. - Resource bounds are not fairness or abuse protection. Per-principal session count times maximum upload size is an implicit byte bound, not an independent byte quota. Multiple principals can exhaust aggregate slots/bytes (four principals with two 4 MiB reservations each exhaust the default eight slots and 32 MiB). Restrict upload authorization and request rates. Receiving sessions can be aborted to release capacity; otherwise capacity returns on expiry, and an in-flight commit stays pinned even beyond TTL. See the protocol and sizing tradeoffs.
S3 storage (-Ds3)
- Upload bytes and database references are not one distributed transaction. HTTP record uploads PUT bytes before acquiring the writer, then validate and commit their references transactionally. Failed requests best-effort remove uploaded bytes; a crash between PUT and commit can still leave an orphan. Upload POST returns
409on collection-definition changes during transfer; PATCH also returns409on record changes. Reload and retry. Replacement/deletion cleanup is synchronous by default; opt into durable HTTP cleanup to enqueue it transactionally. The worker holds a database write lock during each storage deletion, so slow storage can still delay concurrent writes. - Rejected creates can incur storage requests. File constraints and locked create rules are checked before transfer, but ordinary field validation, before-hooks and data-dependent create rules run afterward in the transaction. Even an anonymous request allowed through preflight can incur billable PUT and cleanup DELETE operations before returning
400or403. Bound upload sizes and apply request rate limits; storage plugins must not assume a record exists during PUT. - Deletes are best-effort by default. Opt-in durable HTTP cleanup retries replaced-object and deleted-record-prefix deletion but does not cover raw SQL, Data, cascade/TTL, collection deletion, or failed uploads. Prefix cleanup requires physical row absence; reused IDs retain their objects, including expired rows. Exhausted retries require operator attention; removed/recreated collections are conservatively skipped. Opt-in
files inventoryreports current sizes and reference candidates, not proof that objects are safe to remove (in-flight uploads can be unreferenced). Local scans stop at 100,000 entries and exclude symlinks/deeper-than-record-layout paths; S3 reports current objects in the configured prefix, not versions or multipart state. See storage inventory. - Orphan deletion is offline local/SQLite maintenance, not a background job. Opt-in
files reconciledefaults to read-only dry-run. Explicit apply requires an exclusive storage lease and SQLite writer transaction; cooperating apps (includingserve --ignore-lock) hold one shared boot-lifetime descriptor even in default builds. Stop older/external writers yourself and use a root dedicated to one database; age alone cannot close the PUT-before-commit race. Never replace the permanent lock file. S3/PostgreSQL/custom storage, unknown metadata, encrypted file fields and dropped-collection prefixes are outside deletion scope. Batches are bounded and cursors are not approvals/snapshots. Partial filesystem deletion cannot be rolled back, including when a later output failure prevents the report from being emitted. - Proxy serving is the default; presigned-redirect is opt-in. By default downloads flow through the server’s spool cache (§D.6), so every download consumes server bandwidth/CPU even though the bytes ultimately come from S3. An opt-in presigned-redirect mode is now available via the comptime
App(.{ .files = .{ .s3_presign_redirect = true } })option (S3 backend only; default off = proxy): an authorized download is answered with a 302 to a time-limited presigned GET URL (s3_presign_ttl_s, default 900s) instead of proxying the bytes. Tradeoff: per-request authorization still runs before the redirect is issued, but the issued URL is a bearer capability — valid fors3_presign_ttl_sand not bound to the authorized requester, so anyone holding it can fetch the object until it expires. Keep the TTL short. - Multipart does not make uploads memory-streaming. S3 transfers large objects in bounded sequential parts, but
Storage.putand HTTP ingestion still require the entire input in memory. Incomplete uploads can survive crashes, lost initiation responses, or failed aborts; configure bucket lifecycle expiration. A lost completion response can leave a completed orphan object. Opt-in client-resumable uploads only resume inbound transfers within one running ZigBase process; they are not durable S3 upload sessions. - Spool cache disk usage + eviction approximates last-access LRU, not strict LRU. The local spool cache (
ZIGBASE_S3_CACHE_DIR, default<data-dir>/storage_cache) needs its own disk budget on top of the database; size it viaZIGBASE_S3_CACHE_MAX_BYTES. Eviction still sorts by file mtime, but mtime is now bumped on a cache hit as well as on the miss-fill that creates the entry, so eviction approximates last-access LRU — a frequently-read file survives over a rarely-read one. It is still not a strict LRU: there is no separate access-time bookkeeping or journaling, and the mtime touch is best-effort and coarse-grained.
Other deferred work
- Idempotency is opt-in custom-operation support, not global middleware.
zigbase.Idempotencycoordinates database-only effects and saved result bytes on SQLite/PostgreSQL; callbacks must be trusted, authorization read-only, and current access checked on every replay. PostgreSQL serializes each namespace with fail-fast transaction locks. Retention begins at the trusted attempt timestamp, not commit, and expired keys may execute again. Idle receipts are not physically deleted by a timer; database copies refuse non-empty receipt ledgers. No automatic REST hook/rule wrapper, external exactly-once guarantee, or database/WAL/RSS hard limit is provided. See the contract. - Image transforms beyond named local ImageMagick thumbnails, including formats other than PNG/JPEG/WebP, remote source backends and persistent derivative storage. Thumbnail-enabled deployments must install and maintain a trusted ImageMagick executable; process limits are not a hard RSS sandbox. Cross-instance or streaming resumable uploads; durable/cross-instance realtime replay and per-event-guard load-tuning. Opt-in SQLite record invalidation backfill is process-local, bounded, and requires full reloads for gaps and authorization/query-dependency changes (see
docs/api.md).
These are tracked for upcoming releases. Contributions welcome.