Documentation
Framework — ZigBase
Embed ZigBase as a Zig library — comptime record hooks, custom routes, scheduled jobs, a comptime schema with additive auto-migration, pluggable storage/mailer backends, and footprint levers.
ZigBase is an embeddable Zig framework for building custom applications with integrated backend services and explicit control over resources. You zig fetch --save it, @import("zigbase"), and configure zigbase.App(.{...}) with comptime hooks, custom routes, scheduled jobs, and lifecycle/auth/file event handlers. Your app is the ZigBase server, plus your extensions.
Write these integrations directly or build them with a coding agent: the same typed API, compiler checks, and local tests support both. Start with Why ZigBase for the design rationale, or jump to footprint levers for resource profiles, effective settings, and measurement-driven tuning.
For runnable, end-to-end usage of these APIs (hooks, a custom route with a path param, and a DB-touching cron job), see the tutorial and the recipes. This page is the framework reference.
1. Overview
The shipped binary is, in its entirety:
const std = @import("std");
const zigbase = @import("zigbase");
pub fn main(init: std.process.Init) !void {
return zigbase.App(.{}).runCli(init); // no extensions = the stock server
}
App(cfg) is a comptime application builder. Everything you register is assembled and validated when your program compiles:
- An unknown config key (e.g.
.hookinstead of.hooks) is a compile error. - A typo’d hook phase (e.g.
.beforeCreat) is a compile error. - A route or job spec missing a required field, or with a wrong-typed handler, is a compile error.
- An unknown key on a
.collectionscollection or field spec (e.g..requied,.encrypte,.ttl_filed, or a misspelled rule under.rules) is a compile error — the message lists the recognized keys for that spec. - An unknown/typo’d key on any other list-shaped config surface is a compile error too, rather than being silently ignored: route specs (
.rate_limit/.rate_limit_key/etc.), background job specs,.auth.methodsand each built-in method’s options (.magic_link/.otp/.password/.webauthn),.auth.oauth2and its provider literals (e.g..tokenURL), collection.indexesentries (.unique/.collation/.where), the.poolstuning group, and.migrationsentries. A duplicate or empty.migrationsid is also a compile error.
So a misconfigured extension never reaches runtime — it fails the build loudly.
runCli(init) parses argv and dispatches the usual CLI (serve, migrate, superuser creation, help), wiring your assembled extensions into the running server. (App(...).run(init, cfg) starts the HTTP server directly with an explicit zigbase.Config, skipping CLI parsing.)
2. Add the dependency
The fastest path is to scaffold — it writes a build.zig already wired with the helpers below, a build.zig.zon, a src/main.zig with a comptime schema and in-process tests, and an AGENTS.md:
npx zigbase init --framework --dir myapp && cd myapp
zig fetch --save git+https://github.com/valthon/zigbase
To wire it by hand instead:
zig fetch --save git+https://github.com/valthon/zigbase
zig fetch --save writes both the dependency URL and its content hash into build.zig.zon. Never hand-write either, and never copy the relative .path = "../.." form out of this repository’s examples/ — that only resolves inside this repo.
In your build.zig:
const zigbase = @import("zigbase"); // the dependency's own build.zig
const dep = b.dependency("zigbase", .{ .target = target, .optimize = optimize });
zigbase.addTo(dep, exe_mod);
zigbase requires libc — it carries the bundled SQLite amalgamation and zap transitively. addTo adds the "zigbase" import and sets link_libc on your module in one call, so your build never depends on remembering either. Wiring it by hand works too, but then .link_libc = true on your module is yours to remember:
const exe_mod = b.createModule(.{
.root_source_file = b.path("src/main.zig"),
.target = target,
.optimize = optimize,
.link_libc = true, // required; addTo sets this for you
});
exe_mod.addImport("zigbase", dep.module("zigbase"));
zigbase is pinned to a single Zig minor series (currently 0.16.0, see build.zig.zon’s minimum_zig_version and mise.toml). Building against an unsupported Zig version fails at compile time with a clear required-vs-actual message (e.g. zigbase requires Zig 0.16.x … but you are building with 0.17.0) instead of an opaque deep-compilation error, so a toolchain mismatch is obvious. Install the pinned toolchain with mise install, or invoke mise exec zig@0.16.0 -- zig build.
Minimal src/main.zig:
const std = @import("std");
const zigbase = @import("zigbase");
/// Structured logging: this routes every std.log call through the same JSON-capable
/// encoder as request logging. Access lines and --log-format/--log-level work either
/// way; omitting this line just leaves std.log output in Zig's default format, mixed
/// with JSON access lines under --log-format=json. `std_options` is resolved from the
/// root source file, so every consumer binary declares this itself.
pub const std_options = zigbase.std_options;
pub fn main(init: std.process.Init) !void {
return zigbase.App(.{}).runCli(init); // no extensions = the stock server
}
std_options is a Zig language feature resolved from your program’s root source file — it cannot be inherited from the zigbase module, so every consumer binary (including all three examples/*) declares that one line itself. Omitting it does not disable --log-format/--log-level or request logging — those are applied directly by the server regardless. It means any std.log call (from your own code or ZigBase internals) keeps rendering in Zig’s default ANSI, timestamp-free format, so under --log-format=json you get a mixed stream: JSON access lines interleaved with un-encoded std.log lines.
3. The App(.{...}) config keys
App(.{...}) accepts exactly these optional keys. Any other key is a compile error.
| Key | Purpose | Unset ⇒ in your binary? |
|---|---|---|
hooks | Per-collection record lifecycle hooks (before/after create/update/delete). | always — dispatch plumbing is core; empty when you register nothing. |
onError | Consumer error handler, runs before the built-in backstop. | always — the backstop itself always ships. |
routes | Custom HTTP routes. | always — routing plumbing is core; a route’s code is only in your binary because you wrote it. |
onAuth | Notify-only: fires after a session is issued (login / oauth2). | always — auth lifecycle plumbing is core. |
auth.two_factor | Selected TOTP/WebAuthn factors, recovery codes, and application requirement hook. See Two-factor authentication. | excluded until explicitly configured; enrollment and requirement data remain runtime-configurable. |
beforeAuthSuccess | Writable, transactional, abortable hook that runs before the session is issued (claim records on first login; veto a login). | always — auth lifecycle plumbing is core. |
auth | Auth config group: .hooks (lifecycle hooks — before/after register/logout/refresh/password-change), .methods (built-in auth-method set + custom AuthMethod types), .captcha (.{ .provider, .secret }), and .session (.store = .epoch|.table, .gc_cron, .rotation_grace_s). See Auth methods, CAPTCHA, and Revoking sessions. | mixed — the hook dispatch is core (empty when unset); a deselected .methods built-in, .captcha, and .session = .{ .store = .table }’s extra store/GC machinery are each excluded until configured. |
onFileServe | Fires before serving a file download (may deny). | always — file-serving plumbing is core. |
onFileUpload | Fires after a file upload. | always — file-upload plumbing is core. |
onBootstrap | Lifecycle: after bootstrap. | always — lifecycle dispatch is core. |
onBeforeServe | Lifecycle: just before the server starts serving. | always — lifecycle dispatch is core. |
onBeforeTerminate | Lifecycle: just before shutdown. | always — lifecycle dispatch is core. |
cron | Scheduled job table. | always — the scheduler is core; empty table when unset. |
jobs | The queue job-kind registry (kind → handler). See Background jobs & queues. | always — the registry is core; a job kind’s code is only in your binary because you wrote it. |
queues | Named background-job queues (backend memory/durable, priority, retry). A .default queue is always synthesized. See Background jobs & queues. | excluded — the durable-backend poller + GC jobs compile in only when a queue declares .backend = .durable; the synthesized in-memory .default queue is always present. |
workers | Named queue workers (queue subset drained in strict priority + concurrency). Omit → one implicit worker over all queues. | always — the queue-draining loop is core (it runs even for the implicit single worker). |
collections | Comptime schema: collections provisioned at startup (additive auto-migration). | always — the provisioner is core; empty when unset. |
migrations | Explicit migrations (the escape hatch for non-additive schema changes); a bare tuple or a typed slice. | always — the migration runner is core; empty when unset. |
static_files | Comptime static-file mode: absent (default flag), .disabled, .{ .dir = "..." }, .{ .embedded = ... }. | always — static serving is core; this key only selects its mode. |
storage | Storage plugin TYPE (defaults to local-disk storage). | always — you always get a storage plugin, default or custom. |
mailer | Mailer plugin TYPE (defaults to log/SMTP mailer). | always — you always get a mailer plugin, default or custom. Together with .mail, also enables the built-in "mail" job kind (see .mail below). |
reporter | Error-reporter plugin TYPE — the terminal backstop every framework-swallowed error routes through (defaults to SentryReporter when ZIGBASE_SENTRY_DSN is set, else LogReporter). | always — you always get a reporter plugin, default or custom. |
reporter_dedup | Error-report TTL dedup window: .{ .window_s = N } (seconds) suppresses a repeat of the same (message, phase) within N, or .off to report every swallowed error. Default (omitted): on, 60s. | always — dedup is on by default; .off compiles the dedup map out entirely (a single null-pointer branch, no allocation). |
pools | Footprint levers: reader pool, scheduler workers (.jobs), lazy memory-job/submit workers (.memory_jobs), thread stack size, SQLite page cache. | always — these are levers on core connection/thread machinery, not an optional subsystem. |
resource_profile | Optional .minimal, .balanced, or .throughput defaults for existing pool levers; explicit .pools fields win. | data-only — selects constants, never enables a subsystem. |
admission | Optional positive u32 .max_requests rejects excess synchronous HTTP work with 503. Optional positive u32 .max_work shares capacity with outstanding memory jobs/app.submit and requires .max_requests. Optional positive usize .max_job_bytes independently bounds retained payload/name copies, without requiring HTTP admission. At least .max_requests or .max_job_bytes is required; both job budgets require -Dcoordinated-admission=true. | HTTP checks compile out without .max_requests; all state/diagnostics excluded when .admission is omitted; job accounting excluded without its build flag; one null app pointer remains. |
rest_idempotency | Explicit collection allowlist and receipt limits for authenticated REST retries. Requires -Drest-idempotency=true. | excluded when disabled; keyed requests are refused when unavailable. |
query_workbench | Optional .{ .max_entries = 64, .slow_ms = 100 } budgets for the query workbench. Requires -Dquery-workbench=true. | build-flag gated — no statement fields, TLS, timestamps, state, or endpoints when off. |
pagination | Enable/disable offset & cursor list paging and pick the cursor token format. | always — core list-response plumbing. |
flags | Declared boolean feature flags. See Feature flags + experiments. | data-only — lowers to an empty slice + a zero-variant Flag enum when unset. |
experiments | Declared A/B/n experiments (variants + weights, optional .sticky). See Feature flags + experiments. | data-only — lowers to an empty slice + a zero-variant Experiment enum when unset. |
features | Public feature-state projection: .{ .public_route = "/api/state" } (custom path) or .{ .public_route = .disabled }. Default "/api/state" (auto-mounted). | always — the projection route is mounted by default; this key only customizes or disables it. |
onFeatureExposure | Notify-only hook fired when a declared flag/experiment is resolved for a subject (exposure logging). | excluded — zero-cost when absent: the resolver gates on the dispatch field and never even constructs the event. |
experiment_assignment_ttl | TTL in days for sticky _experiment_assignments rows (default 90). Only valid when a .sticky experiment is declared — setting it otherwise is a @compileError. | excluded — the sticky-assignment GC sweep installs no job/timer/writer touch unless a .sticky experiment is declared. |
ttl_gc_interval | Cadence (schedule.Interval) for the framework-internal _ttl_gc sweep that reaps expired .ttl_field rows (default .{ .minutes = 5 }). Only valid when a collection declares a .ttl_field — setting it otherwise, or to a degenerate .{ .minutes = 0 }, is a @compileError. Expired rows are hidden from reads immediately regardless of cadence. | excluded — the TTL GC sweep installs no job/timer/writer touch unless a collection declares a .ttl_field. |
enable_typegen | Enable the typegen CLI subcommand (default false). Set true only for client-generation builds. | excluded — off by default so production builds carry no codegen weight. |
realtime | .{ .max_connections = 10000, .canSubscribe = fn }: positive u32 shared WS/SSE connection cap and optional custom-topic guard. See ctx.realtime(). | always — realtime is core; omitted cap preserves 10,000, omitted guard keeps custom topics public. |
tenancy | Multi-tenant account scoping: .{ .enabled = true, .auth_collection = "...", .resolver = ..., .roles = .{...} }. | excluded — the account-activation endpoints and membership-scoping code exist only when .tenancy.enabled = true. |
abilities | Declarative relationship-based row abilities per collection: .{ .<col> = .{ .view = <rule>, … } }. Requires .tenancy.enabled = true. | data-only — the composition code is core policy plumbing that always runs; unset just means every collection’s ability predicate is null (a no-op). |
mail | Email-subsystem policy knobs (.require_verified_sender, .webhook_secret, …) threaded into app.mail. Together with .mailer, enables the built-in "mail" job kind backing ctx.mail().enqueue. | excluded — the senders API and inbound bounce/complaint webhook exist only when .mail is set. |
sms | SMS-subsystem knobs threaded into app.sms: .{ .default_region = .us, .queue = "default" }. Setting .sms (or .sms_provider) enables the built-in "sms" job kind backing ctx.sms().enqueue. Default .{} = US default region. See ctx.sms(). | excluded — the "sms" job kind is compiled in only when .sms/.sms_provider is set. |
sms_provider | SMS provider plugin TYPE (defaults to Twilio when the ZIGBASE_TWILIO_* env vars are set, else a logging no-op). Supply your own type implementing the SmsSender contract (create/interface/deinit) to swap providers. Also enables the built-in "sms" job kind. | excluded — like .sms, gates the "sms" job kind. |
analytics | Product analytics: .{ .rollups = .{ … } } declares rollup specs, each producing a scheduled _rollup:<name> aggregation job; also gates the analytics read API. ctx.track always captures to _events regardless. | excluded — the rollup jobs and analytics read API exist only when .analytics is set. |
static_routes | Tier-2 comptime static rewrites: a list of .{ .match = "/…", .serve = "/…" } entries mapping request paths to fixed served paths/files. | data-only — lowers to an empty list when unset. |
enable_spa_marker | Tier-1 .spa marker enablement for static serving (issue #183). Default: on when .static_routes is absent/empty (byte-identical to the shipped binary), off when routes are declared; an explicit bool always wins. | always — a bool feeding core static-serving logic, not an optional subsystem. |
collections_frozen | Assert that collections do not change after boot + migrations (default false, issue #234). When true: the collection-metadata cache runs on all backends — including Postgres, where it is otherwise skipped because a concurrent instance could ALTER collections unseen — and the runtime collection create/update/delete endpoints return 403 (schema then evolves only via .migrations + a redeploy). | always — a standalone bool feeding core collection-metadata caching and the collection-DDL endpoints, not an optional subsystem. |
static_cache_control | Comptime default Cache-Control for static responses (embedded + dir). Runtime --static-cache-control / ZIGBASE_STATIC_CACHE_CONTROL override it; unset → facil.io’s stock max-age=3600. | data-only — a ?[]const u8 default feeding core static serving, not an optional subsystem. |
admin | .admin = .disabled removes the embedded admin SPA (route dispatch and the ~58 KiB of @embedFile-d assets) from your binary — useful for headless/embedded consumers. Default: served at /_/. | excluded, opt-out — the admin SPA ships by default; set .admin = .disabled to exclude it (the only key in this table where “unset” means included). |
webhooks | .webhooks = true registers the built-in "webhook" job kind, compiling in ctx.webhook()’s managed outbound delivery (webhook.zig, ~689 LOC). Default: off, so a consumer that never sends managed webhooks pays nothing for it. See ctx.webhook(). | excluded — off by default; webhook.zig is not compiled in unless .webhooks = true. |
files | File-serving and cleanup knobs: .{ .s3_presign_redirect = false, .s3_presign_ttl_s = 900 }. Optional .cleanup_queue = "cleanup" selects a declared durable queue for transactional HTTP cleanup. Optional .resumable = .{ ... } tunes bounded upload budgets and requires -Dresumable-uploads=true; .resumable.durable = true additionally enables SQLite restart persistence with -Ddurable-resumable-uploads=true. Optional .thumbnails = .{ .imagemagick = .{ ... }, .profiles = .{ ... } } selects named image derivatives and budgets, requiring -Dimage-thumbnails=true. With .s3_presign_redirect = true and the S3 backend active, authorized downloads redirect to a presigned URL (s3_presign_ttl_s, 1..=604800 seconds); other backends keep proxy serving. Default .{} = proxy-only, synchronous cleanup. | cleanup handler excluded unless configured; resumable handlers/store/budgets and thumbnail subprocess/routes/admission compile out without their flags; redirect data is always compiled and engages with S3 at runtime. |
push | Web Push config (.{ .subject = "mailto:ops@example.com" }) — the VAPID sub contact. Registers the built-in "push" job kind backing ctx.push().enqueue, compiling in push/*.zig’s encrypted delivery. The VAPID keypair itself comes from ZIGBASE_VAPID_PUBLIC_KEY/_PRIVATE_KEY at runtime; without them ctx.push() is a no-op. See ctx.push(). | excluded — off by default; push/send.zig’s job handler is not compiled in unless .push is set. |
app_context | A type naming a consumer-owned app-scoped context struct. Set the handle once in onBootstrap (try ctx.setAppData(T, &value)) and read it anywhere via ctx.appData(T) (a *T). Declaring it makes setting it a boot contract: the server refuses to start if onBootstrap never calls ctx.setAppData. See App-scoped context. | data-only — two ?*anyopaque/?[]const u8 fields on App; apps that don’t declare it pay nothing. |
The files group also accepts .thumbnails = .{ .imagemagick = .{ ... }, .profiles = .{ ... } } with -Dimage-thumbnails=true. Thumbnail configuration defines the named PNG/JPEG/WebP profiles, bounded admission and ImageMagick process limits. The subprocess backend, derivative routes and admission state compile out with the flag off; no image codec is linked. ZIGBASE_IMAGEMAGICK_EXECUTABLE may override the compiled absolute executable path at deployment; it must match the compiled ImageMagick 6 (.convert) or 7 (.magick) command style.
Route gating: the built-in analytics, senders/mail-webhook, and tenancy (account activation) endpoints exist only when their config key is set (.analytics, .mail, .tenancy.enabled) — an unset key leaves no trace of those routes (or the handlers they’d call) in your binary; they aren’t merely 404 at runtime.
Inspect those compiled registrations and your custom route declarations with <your-app-binary> routes --json, without opening a database or starting the server. The command follows this application’s gates and feature-route mapping; declarative auth metadata is redacted and is not a runtime authorization decision. See Offline compiled routes. It is development tooling, omitted by -Ddev-tools=false.
Choosing a config plane
The assignment rule: structure/behavior = comptime App(.{…}) key; deploy-varying values & secrets = env var (+ CLI flag if path-like); alternatives with real binary cost (extra C sources, wire protocols, dev-only codepaths) = -D build flag. Never gate a whole subsystem on a runtime value alone.
The laziness contract: for every key marked excluded above, “unset” means the subsystem is not in your binary — not compiled, not routed, not registered. This holds because of the gating invariant: no unconditional fn-pointer registration for optional capabilities (fn-pointer tables defeat Zig’s lazy analysis; builtin_routes and the built-in job registry are comptime-assembled from your config, enforced by scripts/check-gating.sh in CI).
3b. The typegen gate (.enable_typegen)
Setting .enable_typegen = true in the App(.{ … }) literal gates in the typegen subcommand, which generates a typed TypeScript client from the server’s live schema (runtime introspection). It is false by default so that production binaries carry no codegen code or dependencies.
This is a different axis from the -Ddev-tools build flag above: .enable_typegen is an App(.{...})-level comptime key that a consumer’s own binary opts into (independent per build target — e.g. off in your main server, on in a dedicated codegen build); -Ddev-tools is a build.zig flag that governs whether development commands compile into the CLI at all: init, agents-md, typegen, capabilities, routes, migrate preview, tune and diagnostics, regardless of .enable_typegen. A binary needs both -Ddev-tools=true (the default) and .enable_typegen = true for typegen to actually run; either one off makes it unavailable, each with its own actionable error.
// client-generation build target — NOT your production binary
zigbase.App(.{
.collections = .{ /* … */ },
.enable_typegen = true,
}).runCli(init);
Invoke the subcommand against a provisioned data directory (no server required) or against a running instance:
# Offline — reads an already-provisioned data directory:
./myserver-gen typegen --data-dir ./zb_data --out src/zbase.gen.ts
# Live — against a running instance (superuser credentials required):
./myserver-gen typegen --url https://api.example.com --admin-email admin@x.io --admin-password '…' --out src/zbase.gen.ts
The generated output is the typed db / realtime / files surface. Because custom routes are not introspectable at runtime, rpc.* is not emitted — use the comptime generator (zig build gen-client) if you need typed RPC. See the TypeScript SDK docs for flags and a CI staleness-gate recipe.
Recommendation: keep
enable_typegen = falsein your main application binary and true only in a dedicated client-generation build step or a separatebuild.zigtarget.
4. Record hooks (.hooks)
Record hooks fire around collection record writes. The shape is a struct keyed by collection name, plus an optional any wildcard group:
.hooks = .{
.any = .{ // fires for EVERY collection (before the collection-specific group)
.beforeCreate = auditCreate,
},
.posts = .{
.beforeCreate = slugify,
.afterUpdate = reindex,
.beforeDelete = guardDelete,
},
},
The six valid phase fields are beforeCreate, afterCreate, beforeUpdate, afterUpdate, beforeDelete, afterDelete. Within a triggered write, the any group runs first, then the collection-specific group; only the field matching the current phase runs.
Every hook has the signature:
fn (ctx: *zigbase.Ctx, ev: *zigbase.RecordEvent) anyerror!void
The first parameter is the per-request capability object (ctx.records() for DB access, ctx.http() for outbound HTTP, etc. — see §5b). Add _ = ctx; when a hook doesn’t use it.
zigbase.RecordEvent fields:
rctx— the request context;ev.rctx.authis the authenticated record (if any). Namedrctx(notctx) —ctxalways means the*zigbase.Ctxparameter in a hook signature.arena: std.mem.Allocator— the request-scoped allocator that ownsrecord’s JSON storage.collection: []const u8— the collection name.record: *std.json.Value— mutable inbefore_*; the persisted record inafter_*.phase: RecordPhase.
Need the app itself (e.g. app.allocator, app.io)? Use the hook’s ctx.app — RecordEvent has no app field of its own.
Semantics
before*hooks may MUTATEev.recordand may return an error to REJECT the write (the request fails with400). They run AFTER access rules pass.after*hooks are post-commit. An error returned from an after-hook is swallowed and routed to the error backstop (it does not undo the committed write).
Always allocate record data with ev.arena.a
ev.arena is a typed RequestArena, not a bare std.mem.Allocator — the type exists so a long-lived general-purpose allocator cannot be handed to an arena-scoped API by accident, and so the dependency is visible in every signature. The allocator inside it is the field .a.
Any allocation that becomes part of ev.record MUST use ev.arena.a (the request-scoped allocator that owns the record’s JSON map). Mixing allocators on the arena-backed JSON map is undefined behavior.
From the worked example’s slugify (before_create on posts):
fn slugify(ctx: *zigbase.Ctx, ev: *zigbase.RecordEvent) anyerror!void {
_ = ctx;
if (ev.record.* != .object) return;
if (ev.record.object.get("slug") != null) return;
const title = if (ev.record.object.get("title")) |t| switch (t) {
.string => |s| s,
else => return,
} else return;
const buf = try ev.arena.a.alloc(u8, title.len); // <-- ev.arena.a
var len: usize = 0;
// ... build the slug into buf ...
try ev.record.object.put(ev.arena.a, "slug", .{ .string = buf[0..len] }); // <-- ev.arena.a
}
DB access from a hook (ctx.records())
A hook reaches the database through ctx.records() (the Records handle described in §5b). In a before* hook it is bound to the triggering write’s in-transaction connection, so side-writes commit atomically with the triggering write:
get(collection, id, .{}) !?std.json.Value— returnsnullfor both an unknown collection and a missing record (the 3rd arg isGetOptions, e.g..{ .expand = "author" }).create(collection, value) !std.json.Valueupdate(collection, id, value) !?std.json.Valuedelete(collection, id) !boollist(collection, opts) !ListResult
create/update/delete/list return error.UnknownCollection when the collection name does not resolve.
Result lifetime: a
std.json.Valuereturned byctx.records()is not part of theev.arena-ownedev.recordmap. It is fine to read for the duration of the hook, but do not store one intoev.recordwithout first copying it withev.arena.a(mixing allocators on the arena-backed JSON map is UB).
Atomicity: on the HTTP create/update/delete path a
before*hook runs inside the triggering write’s transaction. Side-writes a hook issues viactx.records()commit atomically with the triggering write, and a before-hook that returns an error — or a denied access rule — rolls the whole transaction back, so a rejected write persists nothing (fail closed).
5. Custom HTTP routes (.routes)
.routes = .{
.{ .method = .GET, .path = "/api/blog/ping", .handler = ping, .auth = .public },
},
Each spec needs .method, .path, and .handler (a missing field or wrong-typed handler is a compile error). .auth is optional and defaults to .superuser (the safe default) when omitted. The three auth levels are:
Route paths are canonical absolute paths. A :name segment captures exactly one path segment; capture names use identifier syntax ([A-Za-z_][A-Za-z0-9_]*) and cannot repeat within one route. Relative paths, trailing or repeated slashes, dot segments, percent escapes, literal OpenAPI braces, and .UNKNOWN methods are compile errors, so runtime routing and OpenAPI export use one grammar. Consumer routes use declaration-order precedence, including when two routes with the same method overlap. Put a more-specific literal route such as /api/items/new before a capture route such as /api/items/:id. A .path_secret carried in the path must name one of those captures, including when a helper returns the explicitly named RouteAuthGuard type. The generated OpenAPI operationId defaults to a name derived from the path and must be globally unique, including against built-in operation IDs. The derived name sanitizes punctuation and leading digits into a TypeScript-safe identifier ([A-Za-z_][A-Za-z0-9_]*); set an explicit .name = "..." when the result would collide or the path has no usable name components. HEAD handlers may return their representation body normally: the server preserves its Content-Length metadata and removes the bytes at the transport boundary.
.public— anyone (anonymous identity still provided)..authed— any authenticated principal (a token from any auth collection)..superuser— superusers only.
Collection-scoped .authed (#243). With more than one auth collection, a bare .authed accepts a principal from any of them. To require the principal belong to a specific auth collection, pass a struct instead of the bare enum:
// Only principals whose token was minted for the `customers` collection may reach this route.
.{ .method = .GET, .path = "/api/portal/me", .handler = portal.me,
.auth = .{ .authed = "customers" } },
// `operators` principals OR a superuser (opt-in via .allow_superuser).
.{ .method = .GET, .path = "/api/ops/x", .handler = ops.x,
.auth = .{ .authed = "operators", .allow_superuser = true } },
- Fail-closed, no oracle. A valid token from any other collection — and a superuser, unless
.allow_superuser = true— is rejected with the same401 "Not authenticated."as no token at all. The response never distinguishes “wrong collection” from “no token”, so a caller can’t probe which collection a route belongs to. An empty principal id is likewise rejected. .allow_superuser(optional, defaultfalse) — additionally accept a superuser token.- Comptime-validated. The named collection must be declared in
.collectionsand be of.type = .auth; a typo, an undeclared name, or a non-auth collection is a@compileErrorat build time (fail-fast), as is an unknown sibling key or a non-string.authed/non-bool.allow_superuser. - A plain
.authed(bare enum) is unchanged — it still accepts any authenticated principal.
The handler signature is:
fn (ctx: *zigbase.Ctx) anyerror!zigbase.http.Response
fn ping(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
_ = ctx;
return .{ .status = 200, .body = "{\"pong\":true}" };
}
The *zigbase.Ctx is the per-request capability object: reach the raw HTTP request via ctx.request.? (a *http.RequestCtx), the authenticated identity via ctx.user(), DB access via ctx.records(), and outbound HTTP via ctx.http(). The framework enforces .auth before calling the handler. Built-in routes always win over custom routes that would match the same method + path.
Route guards: path-secret + per-route rate limit
Routes run an ordered guard chain before the handler: (1) the .auth level above, (2) an optional path-secret guard, (3) an optional per-route rate limit. The first guard to deny short-circuits — the handler never runs.
Path-secret (#139). Instead of an AuthLevel, .auth may be a guard struct that gates the route on a shared secret the caller presents (think unguessable webhook/deploy URLs):
.{ .method = .POST, .path = "/api/hooks/deploy/:token", .handler = onDeploy,
.auth = .{ .path_secret = .{
.param = "token", // path param / query key / header name (per .in)
.source = .{ .kv = "deploy_secret" },// .kv | .settings | .config = "..."
.in = .path, // .path (default) | .query | .header
.on_mismatch = .not_found, // .not_found (default, bare 404) | .forbidden (403)
} } },
- Source —
.kv/.settingsresolve the secret from the named_kventry at request time (rotate it live withctx.kv().set(...)or the settings API);.configis a static compile-time string baked into the binary. - Constant-time + no oracle (security). The submitted value is compared to the stored secret in constant time (no
==timing leak), and a mismatch returns a bare 404 by default — indistinguishable from a route that doesn’t exist, so an attacker can’t probe for the endpoint’s existence. Set.on_mismatch = .forbiddenfor an explicit 403 when the path itself is already public knowledge. An empty/absent stored secret fails closed (never matches). - Rotation. Write a new secret to the source and every link carrying the old one immediately 404s. There is no grace window — old URLs break the moment the secret changes.
- A guard-gated route is
.publicat the AuthLevel layer (the secret is the gate). For ad-hoc checks,ctx.verifyPathSecret(param, stored)does the same constant-time path-param comparison inside a handler (returnctx.notFound()onfalseto mirror the bare-404 behavior).
Per-route rate limit (#142). Add a .rate_limit (and optional .rate_limit_key):
.{ .method = .POST, .path = "/api/contact", .handler = contact, .auth = .public,
.rate_limit = .{ .custom = .{ .max = 5, .window_s = 60 } } }, // 5 / 60s per client
.rate_limitis.default/.off(no per-route bucket — routes are unthrottled by default) or.{ .custom = .{ .max = N, .window_s = S } }. A denied request returns429with aRetry-Afterheader.- Bucket key (security). By default the bucket keys on the client IP, which honors
ZIGBASE_TRUST_PROXY: when proxies are untrusted a spoofedX-Forwarded-Forresolves to an empty IP, so a forged header cannot evade or poison another client’s bucket. Supply.rate_limit_key = fn(*Ctx) ?[]const u8to key on something you resolve instead (API key, tenant id, user id); returningnullfalls back to the IP key. - Both guards compose on one route — the path-secret check is ordered first, so a wrong secret 404s without consuming the rate-limit budget.
Ordering caveat. Because
path_secretruns beforerate_limit, a.rate_limiton a path-secret route throttles authorized traffic (requests that pass the secret check) — it does not throttle secret-guessing, since a wrong secret 404s before reaching the limiter. This ordering is intentional (a flood of bad guesses can’t exhaust a legitimate caller’s budget). Rely on a high-entropy secret for brute-force resistance — the constant-time compare already makes guessing infeasible — not on the per-route rate limit.
Reading the request (ctx.request.?)
In a route handler ctx.request is non-null (it is null only in job/hook contexts). Dereference it to reach the raw request data — a *http.RequestCtx with:
ctx.arena— the request-scoped arena, a typedRequestArenawhose allocator is the field.a. Allocate any dynamic responsebodywithctx.arena.a(thehttp.Response.bodyslice must outlive the handler return but is freed with the request; a string literal needs no allocation). The wrapper type is deliberate: an arena-scoped API cannot be handed a general-purpose allocator by accident, and.ais the explicit, greppable bridge to plain-Allocatorhelpers.ctx.request.?.param("id")—?[]const u8, a path param captured from a:id-style pattern segment.ctx.request.?.query/ctx.request.?.body— raw query string / request body.ctx.request.?.bearerToken()/ctx.request.?.cookie(name)— auth header / cookie helpers.
fn confirm(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
const req = ctx.request.?;
const id = req.param("id") orelse
return ctx.jsonError(404, "not_found", "Not found.");
const body = try std.fmt.allocPrint(ctx.arena.a, "{{\"id\":\"{s}\"}}", .{id});
return .{ .status = 200, .body = body }; // body lives in the request arena
}
See recipes.md → a custom business route with a path param for a full worked route.
Response builders + deferred cookies/headers (ctx)
The raw http.Response literal is always available — but for the common shapes the ctx carries builders (all allocating on ctx.arena.a) and a deferred-mutation accumulator. They are conveniences layered over the same http.Response; mixing them with a hand-built literal is fine.
Response builders:
ctx.json(status, value)— serialize any JSON-encodable value (incl. astd.json.Value) into anapplication/jsonresponse.ctx.jsonError(status, code, message)— the canonical{status,code,message,data}envelope with a caller-supplied machinecode(unvalidated against ZigBase’s frozen ledger — it’s your vocabulary). You can also name ZigBase’s own registry via the public re-exportszigbase.error_codes/zigbase.ErrorCode(e.g.zigbase.error_codes.s(.not_found)) when a custom route wants to emit one of the frozen codes instead of inventing its own.ctx.html(status, body)— atext/html; charset=utf-8response.ctx.redirect(status, location)— a redirect with aLocationheader.ctx.notFound()— the canonical404 Not found.envelope.
Reading the request:
ctx.query()— the URL query string, lazily parsed (and cached) into decoded key/value pairs:+→ space,%XXpercent-decoded.q.get("k")returns?[]const u8. In a job/hook context (no request) it is empty rather than an error.ctx.randomToken(n)/ctx.randomHex(n)— arena-owned random tokens (base36 / hex).
Deferred response mutation — ctx.setCookie(cookie) and ctx.addHeader(header) queue a cookie/header that the framework merges onto whatever response the handler returns. Crucially this happens on both the success and the error path, so a Set-Cookie you queue still reaches the client even if the handler then returns error.NotFound / ctx.fail(...):
fn track(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
// Read-or-mint an opaque, anonymous-friendly visitor id (one Set-Cookie, idempotent
// within the request). An existing well-formed cookie is returned verbatim.
const visitor = try ctx.subjectCookie("zb_subject", .{ .max_age_s = 60 * 60 * 24 * 365 });
const q = try ctx.query();
if (q.get("ref")) |ref| {
// ... record the referral for `visitor` ...
_ = ref;
}
return ctx.json(200, .{ .visitor = visitor });
}
ctx.subjectCookie(name, opts) reads name from the incoming request; if present and well-formed it returns that value (no Set-Cookie), otherwise it mints a fresh opaque id and queues a single Set-Cookie. It is not a signed/authenticated identity — just a stable handle for anonymous attribution. SubjectCookieOpts defaults to a secure, SameSite=Lax, http-only, root-path cookie (max_age_s = 0 is a session cookie); set max_age_s/domain for a persistent or cross-subdomain id. An explicit ?subject= query param still wins where the framework consumes it.
DB access from a route (ctx.records())
An untyped route handler reaches the database through ctx.records() — the same Records handle described in §5b. It lazily checks out a pooled reader for reads and acquires the pool writer per write call, releasing both when the framework tears the ctx down, so there is no manual acquireReader / acquireWriter + Data wiring.
For custom HTTP routes, authentication and tenant resolution release their lookup reader before route guards and the handler run. Identity, session and membership values remain request-owned; handler reads acquire a pooled reader via ctx.connForRead(), typically the one authentication just returned. This does not reserve that connection for the handler, create a shared transaction, or cap concurrent database connections.
// reads + writes both go through the same handle
const created = try ctx.records().create("posts", value);
const rec = try ctx.records().get("posts", id, .{});
A typed route reaches the identical handle via req.ctx.records() (see Typed routes).
For raw SQL on a migration-owned table (not a collection), acquire the pooled writer directly and hand it back:
const w = ctx.app.pool.acquireWriter();
defer ctx.app.pool.releaseWriter();
try w.exec("INSERT INTO plugin_audit_log(note) VALUES('...');");
Name-mapped raw reads — data.queryAs
A complex read (join, aggregate, window scan) that outgrows the collection query API drops to raw SQL, where the row layer is positional — st.columnText(28) — which silently corrupts the moment a column is inserted into the SELECT. data.queryAs decodes each result row into a struct T by matching every field to the result column of the same name (respecting AS aliases), not by position:
const OrderRow = struct {
id: i64,
total: f64,
customer_email: []const u8, // ← maps to the `AS customer_email` column
note: ?[]const u8, // ← optional over a nullable column
};
// Ergonomic wrapper — pooled reader + invocation arena supplied for you:
const rows = try ctx.records().queryAs(OrderRow,
\\SELECT o.*, c.email AS customer_email
\\FROM "orders" o JOIN "customers" c ON c.id = o.customer_id
\\WHERE o.total > ?1
, .{min_total}); // []OrderRow — rows + string fields live on the invocation arena
// Or the free function with an explicit connection (raw writes / migration-owned tables):
const conn = try ctx.connForRead(); // *db.Db reader
// const conn = ctx.app.pool.acquireWriter(); // …or the writer for a raw write
const also = try zigbase.data.queryAs(OrderRow, conn, ctx.arena.a, sql, .{min_total});
Contract:
- Args bind positionally — tuple element
ibinds to placeholder?{i+1}(rewritten to$non Postgres by the same chokepoint the rest of the stack uses). Supported arg types:[]const u8(including string literals), any integer, any float,bool,?T(binds the payload or SQLNULL), andnull. - Fields decode by Zig type —
[]const u8(duped onto the allocator), any integer, any float,bool, and?Tover a nullable column (SQLNULL→null, else the value). - Optionals ↔ nullable/absent — an optional field maps to a nullable column; a
NULLvalue decodes tonull. An optional field with no matching result column at all also decodes tonull. - Missing non-optional column is an error — a non-optional
Tfield with no matching result column returnserror.ColumnNotFound(the field name is logged, since a Zig error value can’t carry text). This fires off the column description, so it surfaces even for an empty result. - A SQL
NULLin a non-optional field is coerced to the type’s zero value (""/0/0.0/false), matching the raw column accessors. Make the field?Tif you need to distinguishNULL. - Extra result columns are ignored — the common
SELECT o.*, …case just works; only the columns named byTare read. - Use an arena on the free-function path — if a decode fails partway through a run given a non-arena (GPA) allocator, the per-row string dups already made are not individually freed; pass an arena so a mid-decode error reclaims them wholesale. The
ctx.records().queryAswrapper already uses the invocation arena, so this only concerns the raw free function.
Get a conn for the free function from ctx.connForRead() (a pooled reader) or ctx.app.pool.acquireWriter() (for a raw write); inside a ctx.tx / before* hook the ctx.records().queryAs wrapper binds the active transaction connection automatically.
Comptime-checked raw SQL — checkSql / checkedSql
Raw SQL (queryAs, conn.prepare(...)) names tables and columns as bare strings, so a typo — FROM postts, p.stauts — is invisible until the query runs. zigbase.checkedSql / zigbase.checkSql close that gap: they validate the identifiers in a SQL string against your comptime .collections schema and fail the build on an unknown table or a mistyped qualified column, with a message naming the offending identifier and the valid options.
To reach the lowered schema, name the App type (rather than building it inline in main) so you can pass Backend.collections:
const Backend = zigbase.App(.{ .collections = .{ .posts = .{ ... }, .orders = .{ ... } } });
pub fn main(init: std.process.Init) !void { return Backend.runCli(init); }
fn report(ctx: *zigbase.Ctx) !void {
const Row = struct { id: []const u8, title: []const u8 };
// checkedSql validates AND returns the string, so it wraps the argument in place:
const rows = try ctx.records().queryAs(Row,
zigbase.checkedSql(Backend.collections,
"SELECT p.id, p.title FROM posts p WHERE p.status = ?1"),
.{ "published" });
// ...
}
Three forms are exported:
checkSql(cols, sql)— comptime check only (returnsvoid); use it as acomptimestatement.checkedSql(cols, sql) [:0]const u8— checks then returnssql, so you can wrap the exact string handed toqueryAs/prepare.checkSqlOpts(cols, sql, opts)— the same check withSqlCheckOptions:extra_tables: []const []const u8— tables that legitimately appear in raw SQL but are not comptime collections: engine internals (_kv,_events,_sessions,_queue_jobs) and migration-owned tables (a table your own.migrationscreated). List them here so they count as known instead of being flagged.check_columns: bool = true— disable qualified-column checking (tables are always checked).
// A migration-owned table opted into the known set:
comptime zigbase.checkSqlOpts(Backend.collections,
"INSERT INTO plugin_audit_log(note) VALUES('x')",
.{ .extra_tables = &.{"plugin_audit_log"} });
What is validated (and what is not) — by design. The overriding rule is zero false positives: a checker that rejects valid SQL gets deleted, so it validates only a narrow, high-confidence subset and skips anything it cannot confidently classify.
- Tables are checked strictly — the name after
FROM/JOIN/INTO/UPDATE. Known = your collection names ∪<collection>_fts(for a.searchablecollection’s shadow table) ∪ CTE names (WITH x AS (...)) ∪extra_tables. A subqueryFROM (SELECT ...)is never treated as a table, and upserts are understood: theUPDATEinON CONFLICT ... DO UPDATE SETis a conflict clause with no table operand, so it checks cleanly with no opt-out needed. - Qualified columns are best-effort —
alias.column/table.columnis checked only when the qualifier resolves to a known collection (via theFROM/JOINalias map). The valid-column set isid, created, updated+ your fields, plus the auth system columns (email,username,passwordHash,tokenKey,verified,token_epoch) for anauthcollection.viewcollections are not column-checked. - Never checked (to avoid false errors): unqualified columns, function names, computed
ASaliases,alias.*, expression tokens, and the contents of string literals, comments, andFROM (subquery)bodies. So this catches the common typo classes, not every possible mistake — it is a strict-tables / best-effort-columns net, not a full SQL type-checker.
Typed SELECT builder — Query.select
Where checkedSql validates SQL you hand-wrote, zigbase.Query.select constructs it from a config struct — validating every table and column against the schema as it builds, and emitting a validated [:0]const u8 plus positional binds. It does not execute or decode; the output feeds the same queryAs/prepare path:
const Backend = zigbase.App(.{ .collections = .{ .orders = .{ ... } } });
const Q = zigbase.Query.select(Backend.collections, "orders", .{
.columns = &.{ "id", "total" },
.where = &.{ .{ .col = "total", .op = .gt }, .{ .col = "status", .op = .eq } },
.order = &.{ .{ .col = "created", .dir = .desc } },
.limit = 50,
});
// Q.sql == `SELECT "id", "total" FROM "orders" WHERE "total" > ?1 AND "status" = ?2 ORDER BY "created" DESC LIMIT 50`
// Q.bind_count == 2
const rows = try ctx.records().queryAs(OrderRow, Q.sql, .{ min_total, "open" });
select(cols, table, spec) returns a type exposing pub const sql: [:0]const u8 (the validated, ?N-parameterized string) and pub const bind_count: usize (== spec.where.len). SelectSpec:
columns: []const []const u8— projected columns; empty ⇒SELECT *. Use an explicit list when you decode into a struct so the projection is stable.where: []const Cond—{ col, op }conditions, ANDed. Each binds one placeholder?1..?Nin slice order (so binds are positional by construction — the tuple you pass toqueryAslines up left-to-right).Opiseq | ne | lt | le | gt | ge | like; all single-bind (likeemits"col" LIKE ?N).order: []const OrderBy—{ col, dir }terms (asc/desc), emitted with an explicit direction.limit,offset—offsetis only valid withlimit(SQLite has no bareOFFSET); setting it alone is a build error.distinct→SELECT DISTINCT.
Compile-time guarantees. An unknown table, or any column (in columns/where/order) that isn’t a valid column of that collection, is a @compileError naming the identifier and listing the valid options — the same valid-column set as checkSql (built-ins id/created/updated + auth system columns + your fields). Identifiers are double-quoted in the emitted SQL.
Scope (v1). Single-table SELECT reads only — no joins, no writes (INSERT/UPDATE/DELETE), no GROUP BY/aggregates, and no in (a variadic-bind predicate, deferred as future work). For anything past this, drop to hand-written SQL guarded by checkedSql.
Auth collections:
ctx.records().create(collection, fields)on an auth collection runs the same credential transforms as the HTTP layer (generates the per-recordtokenKey, forcesverified=false, and hashespasswordif one is supplied). Apasswordis optional, so a passwordless flow can provision a credential-less account thatzigbase.auth.issueSession/mintLinkTokencan immediately operate on. Non-auth collections take the plain insert path. (The lower-level enginerecords.createdoes not provision — reach for it directly only for raw import/migration.)
Typed routes and the generated rpc surface
For routes that carry a structured input or output, ZigBase supports typed routes declared with a Req(Input) / Output handler signature instead of the raw untyped fn(*zigbase.Ctx) anyerror!zigbase.http.Response form. A typed route handler looks like:
// Input and Output are Zig types in the bounded Zig→TS subset: bool, int/float,
// []const u8 (string), enums (→ string-literal union), std.json.Value (→ unknown),
// optionals ?T, slices []T, and nested structs — applied recursively, so ?Struct
// and []Struct are allowed. Anything else is a comptime error. (Caveat: a GET/HEAD/DELETE
// query Input must be a flat struct of scalars/enums/strings/optionals-of-those.)
fn handler(req: *zigbase.Req(InputType)) zigbase.RouteError!OutputType {
const id = req.param("id"); // ?[]const u8 — a :param from the path
const input = req.input; // InputType — parsed request body (POST/PUT/PATCH/OPTIONS) or query (GET/HEAD/DELETE)
// ...
return OutputType{ ... };
// or: return req.fail(404, "not found"); // → RouteError propagated as HTTP 404
}
A typed handler reaches the per-request capability object through req.ctx: req.ctx.records() for DB access, req.ctx.http() for outbound HTTP, req.ctx.arena.a for allocations (req.ctx.arena is a typed RequestArena; .a is the allocator inside it), and req.ctx.app for the runtime. req.ctx is the only route to those capabilities — Req itself carries just input, params, auth_id, failure, and ctx.
Register typed routes in .routes identically to untyped ones (.method, .path, .handler, optional .auth defaulting to .superuser). The generator (zig build gen-client) reads the comptime App declaration and emits a zb.rpc.* method for each typed route — named by camel-joining the path segments (:param segments omitted):
| Path | Method | Generated name |
|---|---|---|
/api/bookings/:id/confirm | POST | bookingsConfirm |
/api/bookings/:id/cancel | POST | bookingsCancel |
/api/listings/:id/availability | GET | listingsAvailability |
/api/golfsim/health | GET | golfsimHealth |
The generated TypeScript method signature mirrors the route shape:
paramsobject (e.g.{ id: string }) when the path has:paramsegments.inputargument when the ZigInputtype is non-void (POST/PUT/PATCH/OPTIONS routes serialize as the request body; GET/HEAD/DELETE routes pass as query parameters).- Output is the TypeScript equivalent of the Zig return type;
std.json.Valuemaps tounknown. .authdefaults to.superuserwhen omitted — typed routes are locked to superusers unless explicitly set.
Untyped routes (the raw fn(*zigbase.Ctx) anyerror!zigbase.http.Response form) own their full response — status, cookies, redirect, content-type — and carry no typed Input/Output, so the generator deliberately skips them: they do not appear in zb.rpc.*. Call them with your own fetch/HTTP client. This is what lets an untyped handler set a session cookie, return a 307 redirect, or serve a non-JSON body (e.g. text/calendar) that a typed zb.rpc.* method could not express.
See the TypeScript SDK docs and examples/golfsim/ for the full worked example.
5b. Handler capabilities (ctx)
Every handler type — route, hook, and job — receives a *zigbase.Ctx as its first parameter. It provides curated DB access, an outbound HTTP client, outbound mail (ctx.mail()), and a standard error model, without manually acquiring connections from the pool.
ctx.records() — records access
ctx.records() returns a Records handle. For reads it lazily checks out and caches a pooled reader (released by the framework when the ctx is torn down); write methods (create/update/delete) acquire the pool writer and release it per-call. This replaces hand-built pool.acquireWriter() + Data wiring — e.g. a cron job that expires stale booking holds (the form used by examples/golfsim):
fn expireHolds(ctx: *zigbase.Ctx, ev: *zigbase.events.JobEvent) anyerror!void {
const stale = try ctx.records().list("bookings", .{
.filter = "status = \"pending\" && starts_at < @now",
.perPage = 200,
});
// stale.items: []std.json.Value | stale.totalItems: ?i64 (null in cursor mode)
for (stale.items) |item| {
const id = item.object.get("id").?.string;
var patch: std.json.ObjectMap = .empty;
defer patch.deinit(ev.app.allocator);
try patch.put(ev.app.allocator, "status", .{ .string = "cancelled" });
_ = try ctx.records().update("bookings", id, .{ .object = patch });
}
}
Available Records methods:
| Method | Notes |
|---|---|
list(col, opts) !ListResult | opts: filter, filter_args, sort, page, perPage, limit, cursor, expand |
get(col, id, opts) !?Value | opts: expand; returns null for a missing record |
create(col, value) !Value | acquires + releases the pool writer |
update(col, id, value) !?Value | acquires + releases the pool writer |
delete(col, id) !bool | acquires + releases the pool writer |
createAs(T, col, input) !T | typed create; input is an anon literal of writable fields |
getAs(T, col, id) !?T | typed read; parses the record into T, null if missing |
updateAs(T, col, id, patch) !?T | typed patch; patch is an anon literal of the changed fields |
Numeric fields accept numbers on write. Reads of an .int/.fixed-mode .number field return a JSON string (to preserve precision beyond f64), and on write those modes now accept a JSON number as well as the string form — symmetric with reads. price_cents = 500 (an integer) and price = 5.0 (a float on a fixed(scale=2) field, scaled to 500 exactly as the string "5.0" would be) both bind correctly. A fractional float on an .int field is still rejected. This removes the previous surprise where passing a number as a number failed validation.
Typed record I/O (createAs / getAs / updateAs). These layer over the std.json.Value methods (which remain for dynamic callers) so a handler can speak plain structs instead of hand-assembling ObjectMaps and unwrapping union tags:
const Order = struct {
id: []const u8,
ref: []const u8,
price_cents: i64, // an `.int` number field — written and read back as an i64
confirmed: bool,
note: ?[]const u8, // optional ↔ a nullable/absent column
};
// create from an anon literal of the writable fields; returns the full created record as Order
const created = try ctx.records().createAs(Order, "orders", .{
.ref = "A-1", .price_cents = 500, .confirmed = true,
});
// read it back (null if the collection or record is missing)
const order: ?Order = try ctx.records().getAs(Order, "orders", created.id);
// patch just the changed fields
const updated = try ctx.records().updateAs(Order, "orders", created.id, .{ .confirmed = false });
- Struct ↔ schema mapping is by name: an
input/patch/Tfield maps to the collection column of the same name. - Optionals in
Tmap to nullable/absent columns (nullround-trips asnull). .int-mode fields map to Zig integers (e.g.price_cents: i64above);.fixed-mode fields must be[]const u8(orf64) inT, since they round-trip as decimal strings (e.g."5.00") — parsing that string into an integer field errors.Tmay be a subset of the record’s columns — id/autodate columnsTdoesn’t name are skipped on the read back (ignore_unknown_fields).- Comptime field-subset guard: every field of a
createAs/updateAsliteral is verified to exist onTat compile time, so a typo’d key is a build error, not a silent runtime no-op. (The mapping to the collection is checked at runtime — the collection is named at runtime — so aTfield the collection doesn’t declare surfaces as the usual schema-validation error.) getAstakes no opts; expand/projection over typed reads is a future addition.
Binding runtime values (filter_args) — to splice a runtime value into a filter, do not string-concatenate it into the filter text (a value containing a quote, ||, (, etc. would be re-parsed as grammar). Instead put a ? placeholder in filter and pass the value in .filter_args. Each ? binds its value as a literal SQL parameter — exactly like a SQL bind parameter — and is never re-parsed as filter grammar, so it is the injection-safe way to bind untrusted input:
// find posts whose title is EXACTLY this (attacker-controlled) string
const page = try ctx.records().list("posts", .{
.filter = "author = ? && title = ?",
.filter_args = &.{ .{ .string = user_id }, .{ .string = untrusted_title } },
});
FilterArg is a union(enum) of .string, .int, .float, .bool, and .null (.null binds SQL NULL, so col = ? with a null arg matches nothing, per SQL three-valued logic — identical to writing the inline literal col = null, which also binds SQL NULL regardless of the field type). A null has no text representation, so a null operand in a LIKE/~ term (inline col ~ null or a .null arg) is rejected with a loud error.BadValue rather than matching everything. A placeholder binds a value exactly as if you had written it inline as a literal on that field — the same field-type coercion applies, so price = ? with .{ .float = 5.0 } scales to the same stored value as the inline literal price = 5.00 (no need to pre-convert to storage form). An arg whose kind is incompatible with a typed field (e.g. a .string for a numeric column, or a number for a bool column) is a loud error.BadValue, just as the equivalent inline literal would be. Placeholders bind 0-based, left-to-right, and the number of ? in filter must equal filter_args.len — a mismatch is a loud error.BadFilter, never a silent misbind. (The REST ?filter= query string supplies no filter_args, so a ? there also fails closed with the same count-mismatch error — it can never smuggle an unbound placeholder.)
Expanding relations — pass .expand in opts to inline related records under an "expand" key:
// get a post and expand its author relation
const post = (try ctx.records().get("posts", id, .{ .expand = "author" })).?;
const author_name = post.object.get("expand").?.object.get("author").?.object.get("name").?.string;
// list with expand
const page = try ctx.records().list("posts", .{
.filter = "status = \"published\"",
.sort = "-created",
.perPage = 50,
.expand = "author",
});
Response projection (fields=) — the read endpoints (list and single-record view) also accept a fields= query param that narrows which keys are serialized: a comma-separated list of dot-paths (fields=id,title,expand.author.name), with * for all keys at a level and a leading - to exclude (*,-secret). It descends into expanded relations (objects and arrays) and is strict — only listed paths appear, id is not auto-included. Projection is a pure output filter applied after expand and access rules, so it can only narrow a response, never reveal a field the record wouldn’t otherwise return. Full semantics in api.md → fields.
ctx.mail() — send application mail
ctx.mail() sends outbound application email from any route, hook, or job. The framework owns the security-critical parts so a consumer never re-rolls them: recipient (and reply_to) address validation and CRLF / control-char header-injection rejection in to / subject / reply_to happen before any byte reaches a backend or a queue row.
A MailMessage is { to, subject, text?, html?, reply_to?, attachments? } — supply text, html, or both (at least one is required). When both are present the message is built as multipart/alternative (plain-text part first, HTML last) so capable clients render the HTML and the rest fall back to text.
File attachments (attachments). Attach files — the canonical case is a .ics calendar invite — by setting attachments to a slice of { filename, content_type, data } (raw bytes; the framework base64-encodes them). When present, the message body is wrapped in a multipart/mixed around the existing body, with one base64 part per attachment; when absent (the default &.{}) the message is byte-for-byte what it was before. filename/content_type are CRLF/control-char checked like every other header field. The total assembled size (body + attachments at their base64-expanded size) is bounded by .mail.max_message_bytes (default 10 MiB) — an over-cap send/enqueue fails loudly at the call site with error.MailTooLarge, never a silent truncation. Attachments ride through every backend: SMTP and the sendmail/ Command backend get the raw MIME directly, Amazon SES switches to Raw MIME content, and Postmark uses its native Attachments array. This is for download attachments only — there is no cid: inline-image support (see below).
try ctx.mail().send(.{
.to = "user@example.com",
.subject = "Your tee time is booked",
.text = "Calendar invite attached.",
.attachments = &.{.{
.filename = "appointment.ics",
.content_type = "text/calendar; charset=utf-8; method=REQUEST",
.data = ics_bytes, // raw .ics you rendered
}},
});
// Synchronous: build + deliver through the configured mailer right now.
try ctx.mail().send(.{
.to = "user@example.com",
.subject = "Welcome",
.text = "Thanks for signing up!",
.html = "<h1>Thanks for signing up!</h1>",
.reply_to = "support@example.com",
});
// Background: hand it to the queue (the built-in "mail" job kind). Routed to the
// queue's backend — durable (survives restart) or memory (in-process). Pick the queue
// by name; defaults to the always-present "default" queue.
try ctx.mail().enqueue(.{ .to = "user@example.com", .subject = "Digest", .html = "<p>…</p>" }, .{ .queue = "emails" });
enqueue/deliverLaterrequire mail to be configured — a.mailerplugin TYPE, or.mail = .{}(defaults) to enable background delivery with the env-configured mailer. Without either, the"mail"kind is not compiled in andenqueuefails at call time witherror.UnknownJobKind(logged with a hint).sendis unaffected either way — it delivers directly through the mailer, not the queue.
send delivers through the same Mailer.send vtable seam every backend (Log / SMTP / Command / a custom .mailer plugin) routes through — including the dev-only testcapture.mail outbox, so consumer mail is assertable in tests exactly like the framework’s own auth mail (see Test-mode capture). When no mailer is wired (CLI/tests), send logs a fallback line. enqueue validates the message up front, so a malformed or injection-bearing message fails at the call site rather than later inside a worker; it requires a wired queue (see §7b Background jobs & queues). Errors: error.InvalidAddress, error.HeaderInjection, error.EmptyBody, error.MailTooLarge.
Email subsystem (#154): templates, providers, verified senders, suppression
The transactional-mail core layers on top of ctx.mail(). Everything below is additive and off by default — an app that only calls the existing mailer is unaffected.
Templates (zigbase.mail_template). A small, safe renderer for transactional mail — no arbitrary code, no loops/conditionals, just variable interpolation, named partials, and a shared layout. Interpolation is HTML-escaped by default ({{ name }}); raw output is an explicit opt-in ({{{ name }}}). {{> partial }} includes a named partial; renderInLayout wraps a body in a layout (which references the child via {{{ body }}}). renderHtml escapes; renderText does not (text/plain has no markup). Render both parts and pass them as .html / .text.
const tpl = zigbase.mail_template;
const html = try tpl.renderHtml(ctx.arena.a, "<p>Hi {{ name }}</p>", &.{ .{ .key = "name", .value = user_name } }, &.{});
HTTP-API providers. Two first-class backends implement the same Mailer vtable as SMTP, so call sites are provider-agnostic — select one via App(.{ .mailer = MyMailerPlugin }):
zigbase.SesMailer.init(region, access_key, secret_key, from)— Amazon SES v2SendEmail, AWS SigV4-signed.zigbase.PostmarkMailer.init(server_token, from)— Postmark/emailAPI.
A message may set a per-message from (additive, null default) to override the backend’s global sender — the seam verified per-account senders ride on. For tests, zigbase.CaptureMailer records messages in memory so you can assert subject/recipient/both body parts with no network.
Verified per-account sender identities. Prove an account controls a From address before it may send as it. Enforcement engages only when you set .mail.require_verified_sender = true AND the send is account-scoped (ctx.mail() attributes mail to the request’s active account automatically); a system/superuser send (no account) bypasses, so existing simple-SMTP apps never start rejecting. Routes (tenant-scoped, fail closed):
POST /api/senders{ "email": "from@acct.com" }— request verification (emails a single-use token).POST /api/senders/:id/verify{ "token": "…" }— confirm.GET /api/senders— list the active account’s identities:{ "items": [ … ] }.
A send whose From is not a verified identity for the account is rejected (error.SenderNotVerified). Addresses are compared case-insensitively (normalized/lowercased on store and lookup), so From@Acct.com and from@acct.com are one identity. Verification-email (re)sends are rate-limited per (account, email) (a repeat within ~60s returns 429) so an authenticated member cannot amplify mail at an arbitrary recipient. The verification token is matched in constant time.
Bounce/complaint suppression. POST /api/mail/webhooks/:provider (ses | postmark) ingests delivery events. It verifies a shared-secret HMAC-SHA256 signature with a constant-time compare — the signed string is "<X-Webhook-Timestamp>.<provider>.<X-Account-Id>.<body>", so the body AND the target account are authenticated (a replayer can’t redirect a captured event to another tenant). A stale X-Webhook-Timestamp (outside ±5m) and a wrong/missing signature are rejected (401), and with no .mail.webhook_secret set the route is disabled (404) — ingestion is opt-in. A hard bounce or complaint upserts a _suppressions row. When .mail.check_suppression = true, a send to a suppressed recipient (case-insensitive) is blocked (error.RecipientSuppressed).
Attribution honesty (per-tenant vs global). A genuine provider webhook (SES SNS / Postmark) cannot compute our HMAC and cannot add
X-Account-Id, so it can only reach this endpoint via an operator-run relay that decides the owning account, injectsX-Account-Id, and signs the string above. A suppression with an empty account is GLOBAL (blocks the address for every tenant). Do not assume per-tenant isolation from a raw provider webhook — only the signed relay provides it.
Enforcement boundary. Verified-sender + suppression checks are a policy of the
ctx.mail()layer, not of theMailer.sendvtable seam. Code that bypassesctx.mail()and calls a backend directly skips them (header/CRLF rejection still applies). Always send application mail throughctx.mail().
const App = zigbase.App(.{
.mailer = MyProviderPlugin, // SES / Postmark / SMTP
.mail = .{
.require_verified_sender = true, // tenant sends must use a verified From
.check_suppression = true, // block hard-bounced / complained recipients
.webhook_secret = "…", // enable the bounce/complaint webhook (signed relay)
},
});
The data model: migration 0016_email seeds two system collections — _sender_identities(account, email, verified_at, verification_token, status) (UNIQUE (account,email)) and _suppressions(account, email, reason, source) (UNIQUE (account,email)).
Bulk list sends (sendBulk). ctx.mail().sendBulk(BulkSend) fans one templated message out to many recipients over the durable queue, rendering per-recipient personalization at delivery time. BulkSend is { subject, text?, html?, from?, reply_to?, recipients, list = "", queue = "default", at?, account? } — recipients is []const BulkRecipient{ to, vars = &.{} }. Subject/text/html are template sources, rendered per recipient against vars with the same engine as transactional mail: {{ name }} is HTML-escaped, {{{ name }}} is the explicit raw opt-in. A bulk queue must be durable (error.BulkRequiresDurable) — memory jobs can’t survive a restart, carry a report, or be rate-throttled.
const batch_id = try ctx.mail().sendBulk(.{
.subject = "Hi {{ name }} — May updates",
.html = "<p>Hi {{ name }},</p><p>…</p>",
.list = "newsletter",
.queue = "ses_mail",
.recipients = &.{
.{ .to = "a@example.com", .vars = &.{.{ .key = "name", .value = "Ann" }} },
.{ .to = "b@example.com", .vars = &.{.{ .key = "name", .value = "Bo" }} },
},
});
const report = try ctx.mail().batchStatus(batch_id); // {total, pending, sent, suppressed, invalid, failed, canceled}
Under the hood, sendBulk writes one _mail_batches row (the templates, once) and one _mail_batch_recipients row per distinct recipient, then enqueues one durable "mail_batch_item" job per distinct recipient with the tiny payload {"batch":"…","to":"…"} — N small jobs rather than one driver job, so per-recipient retry/backoff, priority ordering, visibility-timeout crash recovery, and queue rate throttling all come from the existing job engine for free. ctx.mail().batchStatus(id) reads the durable send-report back as a BatchReport; a superuser can also browse _mail_batches / _mail_batch_recipients directly over the records API (they’re Locked system collections — superusers only, same as every other _-prefixed table). ctx.mail().cancelBatch(id) flips every still-pending recipient row to canceled and returns the count; it’s idempotent, and a stray in-flight "mail_batch_item" job for an already-canceled batch drains as a no-op rather than an error.
Suppression on list mail is always enforced, independent of the .mail.check_suppression knob (which only gates transactional send/enqueue) — sending past a suppression is a compliance issue, not a tuning choice. A suppressed recipient is a reported outcome (status = "suppressed" on its row), not a job failure. Account attribution mirrors send(): b.account defaults to the request’s active account scope (explicit "" is a system send), and when .require_verified_sender is on and the batch is account-scoped, b.from is asserted against _sender_identities once at submit (delivery re-checks anyway). Like every durable job, delivery is at-least-once: a crash between the backend accepting the message and the recipient row being marked sent replays the job, producing one duplicate send — identical to the "mail" job kind, and the reason durable handlers must tolerate replays.
Scheduled sends + the drip recipe (deliverAt / cancel). EnqueueOpts gained at: ?i64 and delay_s: ?u32 (unix-seconds absolute / relative-from-now) — set at most one (error.ConflictingSchedule), and either requires a durable queue (error.ScheduleRequiresDurable — a memory job can’t survive to a future time). ctx.mail().deliverAt(msg, opts) schedules a message and returns the durable job id; ctx.mail().cancel(id) calls a still-pending send off, returning false when it already ran, is running, or was already canceled. sendBulk‘s own .at schedules every recipient job’s earliest delivery the same way.
Persisting the id deliverAt returns is the whole drip-sequence primitive — there’s no separate campaign machinery:
// On trigger (e.g. signup): schedule each step and remember the ids.
const step1 = try ctx.mail().deliverAt(welcomeMsg(user), .{ .delay_s = 0 });
const step2 = try ctx.mail().deliverAt(tipsMsg(user), .{ .delay_s = 3 * 86_400 });
const step3 = try ctx.mail().deliverAt(nudgeMsg(user), .{ .delay_s = 7 * 86_400 });
// Persist step1/step2/step3 on your own record, or ctx.kv().
// On conversion: call off whatever hasn't fired yet.
_ = try ctx.mail().cancel(step2);
_ = try ctx.mail().cancel(step3);
For sequences whose steps depend on runtime state (branch on a later event, vary count per segment), reach for a cron job plus a query instead — deliverAt/cancel are deliberately a recipe, not a scheduling DSL.
One-click unsubscribe (RFC 8058). Set .mail.unsubscribe_base_url (comptime key or ZIGBASE_UNSUBSCRIBE_BASE_URL, env wins) to your app’s public origin, e.g. "https://app.example.com". Empty (the default) is off: no List-Unsubscribe headers are emitted and the endpoint 404s — the same default-off pattern as webhook_secret. Once configured, every sendBulk recipient automatically gets a signed, stateless unsubscribe token — no new secret to manage: the token is HMAC-SHA256’d with a labeled derivation of the existing JWT secret ("zigbase.mail.unsub.v1"), so it can never be confused with an auth token. The token never expires by design: an unsubscribe link must still work from a five-year-old inbox, and the worst-case “attack” is unsubscribing an address whose mail the attacker already possesses.
The endpoint is POST /api/mail/unsubscribe?t=<token> (mail clients / providers call this on the user’s behalf) — success is 204, including on a repeat call (idempotent upsert; no “was this address known” oracle), and any invalid token (bad signature, malformed, tampered) is one generic 400 (no oracle distinguishing failure reasons). GET on the same URL renders a minimal HTML confirmation page whose button POSTs back — it never mutates, so link-prefetchers and mail-scanner bots can’t unsubscribe someone by fetching the link. Both verbs are per-IP rate limited. A successful one-click unsubscribe records a reason = "unsubscribe" suppression that blocks list mail only for that recipient (transactional mail is unaffected) — it’s account-global, with the originating list name recorded in source = "one_click:<list>" for audit. Per-list preference centers (“resub to just the newsletter”) are an app-level UX layer you can build on top of that audit trail; the framework guarantees the compliance floor — one unconditional opt-out per address.
For hand-rolled list mail that doesn’t go through sendBulk, ctx.mail().unsubscribeUrl(account, list, recipient) mints the same signed URL to set on MailMessage.list_unsubscribe yourself (error.UnsubscribeNotConfigured when the base URL is unset). Transactional mail (ctx.mail().send / .enqueue) never carries these headers — only bulk list mail auto-wires them.
const App = zigbase.App(.{
.mail = .{
.unsubscribe_base_url = "https://app.example.com", // "" (default) = feature off
},
});
HTML that renders everywhere. ZigBase ships no CSS inliner or cid: inline-attachment support — author your HTML with inline style="…" attributes directly, or run a build-time inliner (MJML, juice, premailer, …) over your template sources and paste its output into your mail_template strings; {{ }} / {{{ }}} interpolation passes through untouched either way. Host images at absolute HTTPS URLs — your app’s file storage or static assets work well — rather than embedding them. Keep a rendered HTML body under roughly 100 KB; ctx.mail() warns (never errors) above that, since Gmail clips messages around ~102 KB and hides the rest behind “[Message clipped]”. CSS inlining is left out because a half-baked inliner without a real CSS selector engine is a wrong-styling footgun — worse than shipping no inlining at all. Regular download attachments (e.g. a .ics invite) are supported via MailMessage.attachments (see above); only cid: inline images are left out — they need per-backend raw-MIME plumbing keyed by a Content-ID the HTML references, for a problem hosted images at absolute HTTPS URLs solve more simply.
ctx.sms() — transactional SMS (#224)
ctx.sms() is the SMS analog of ctx.mail(): a framework-owned outbound text-message channel with a pluggable provider. Twilio is the first provider; when it is not configured the channel is a network-free logging no-op, so dev/CI/tests need no credentials.
// Synchronous send (delivers now, on the request path):
try ctx.sms().send(.{ .to = "+15551234567", .body = "Reminder: tomorrow 10:00" });
// Durable background send (retry/backoff on the queue engine):
try ctx.sms().enqueue(.{ .to = "555-123-4567", .body = "Your code is 123456" }, .{});
E.164 normalization is framework-owned and runs BEFORE any byte reaches a provider or a durable queue row — a consumer never re-rolls it (mirroring how ctx.mail() owns recipient-address validation). Common separators (spaces, dashes, parens, dots) are stripped; a leading + is preserved; a national number with no + gets the .sms.default_region country code prefixed (US → +1, so 555-123-4567 → +15551234567). The result must be + followed by 1–15 digits with a non-zero first digit (the E.164 shape); anything else (letters, an embedded +, a leading +0, too many digits) is rejected with error.InvalidPhoneNumber at the call site — so a bad number never reaches the provider or the queue.
Provider selection. The default provider plugin reads ZIGBASE_TWILIO_ACCOUNT_SID, ZIGBASE_TWILIO_AUTH_TOKEN, and ZIGBASE_TWILIO_FROM (secrets are ENV only, like SMTP): with all three set it delivers via Twilio’s Messages.json HTTP API (HTTP Basic auth, application/x-www-form-urlencoded body — each value percent-encoded); otherwise it logs the message and performs no network I/O. Supply your own provider with App(.{ .sms_provider = MyProvider }) — any type implementing the SmsSender contract (create/interface/deinit, like the mailer plugin).
Queued delivery. enqueue rides the same background-queue engine as ctx.mail().enqueue: it registers a built-in "sms" job kind (compiled in only when .sms or .sms_provider is configured) and honors the queue’s per-queue rate throttling, retry, and backoff. Pick a queue with .{ .queue = "texts" }.
Testing. CaptureSms is an in-memory SmsSender test double: wire it via App(.{ .sms_provider = ... }) (or call its sender() directly) and assert on the recorded messages (to/body/from) with no Twilio account and no network — the SMS analog of CaptureMailer.
ctx.http() — outbound HTTP client
Returns an HttpClient bound to the Ctx’s arena and the app’s io:
const client = ctx.http();
const res = try client.get("https://api.example.com/data");
// res: HttpResponse{ status: u16, headers: []const Header, body: []const u8 }
const res2 = try client.post("https://webhook.example.com/notify", .{
.body = "{\"event\":\"booking_confirmed\"}",
.headers = &.{ .{ .name = "Content-Type", .value = "application/json" } },
});
ctx.webhook() — managed outbound webhooks (#144)
ctx.http() is a one-shot client — fire-and-handle-the-result yourself. ctx.webhook() is its managed, retrying counterpart: it serializes payload to JSON, enqueues a background "webhook" job (a built-in queue kind, like "mail"), and a worker POSTs it with automatic retries and back-off.
Requires
.webhooks = truein yourApp(.{...})config — without itwebhook.zigis not compiled into your binary andctx.webhookfails at enqueue time witherror.UnknownJobKind(logged with a hint pointing at the missing key).
try ctx.webhook("https://hooks.example.com/booking", .{
.event = "booking_confirmed",
.id = booking_id,
}, .{
.queue = "outbound", // null → the always-present "default" queue
.retries = 5, // max delivery attempts (1 = no retry)
.backoff = .exponential, // .fixed | .exponential (queue back-off math)
.timeout_s = 10, // per-attempt request timeout
.sign = .{ .secret = "whsec_…" }, // optional HMAC-SHA256 body signature
// .idempotency = true, // default: stable Idempotency-Key across retries
});
Response classification. A 2xx is delivered. A network/transport error, any 5xx, or a 429 (honoring an integer Retry-After in preference to the configured back-off) is retryable up to retries. Any other 4xx (and 1xx/3xx) is terminal — the receiver rejected it, so retrying is pointless. When delivery is terminally rejected or attempts are exhausted, the framework fires your .onError hook with phase .webhook; the job itself then succeeds so the queue does not double-retry.
Signing (opts.sign). When set, each attempt adds X-Signature: hex(HMAC-SHA256(secret, "<timestamp>.<body>")) and X-Webhook-Timestamp: <unix>. The timestamp is bound into the signed string (and is fresh per attempt) so a captured request cannot be replayed indefinitely; a receiver recomputes the digest over "<timestamp>.<raw body>" and compares. Header names are overridable on the Hmac struct.
Idempotency (opts.idempotency, default on). A single Idempotency-Key is minted once at enqueue time and frozen onto the (durable) job row, so every retry — and any at-least-once replay after a crash — reuses the same key, letting the receiver dedupe.
Worker-stall caveat: retries back off by sleeping in the worker thread for the full retry duration. Under the default single-worker topology a slow or failing endpoint therefore stalls draining of every queue (including
"mail"). For production, give webhooks a dedicated queue + worker so their backoff never blocks other jobs —.queues = .{ .webhooks = .{ .backend = .durable } }plus a worker bound to it (.workers = .{ .hooks = .{ .queues = .{"webhooks"} } }), then pass.queue = "webhooks".
Security note: a durable, signed webhook persists the signing secret inside the
_queue_jobs.payloadcolumn (your own DB). Prefer amemoryqueue, a shortdone_ttl_s, or DB-at-rest encryption if that is a concern. TLS verification is always on. Requires a wired queue (see §7b Background jobs & queues).
ctx.push() — Web Push notifications (#223)
ctx.push() sends a browser Web Push notification to a subscription your frontend obtained from pushManager.subscribe(...). It composes an RFC 8291 aes128gcm-encrypted payload and an RFC 8292 VAPID Authorization header, then POSTs to the subscription’s endpoint.
Requires
.push = .{ .subject = "mailto:ops@example.com" }in yourApp(.{...})config (the VAPIDsubcontact) plus a VAPID keypair in the environment. Generate one withzigbase vapid-keygenand setZIGBASE_VAPID_PUBLIC_KEY/ZIGBASE_VAPID_PRIVATE_KEY. Without the keys,ctx.push()is a network-free logging no-op — dev and CI need no keys, and the public key is also the browser’sapplicationServerKey.
const sub = zigbase.PushSubscription{
.endpoint = row.get("endpoint"), // from pushManager.subscribe()
.p256dh = row.get("p256dh"), // base64url UA public key
.auth = row.get("auth"), // base64url UA auth secret
};
const result = try ctx.push().send(sub, .{
.title = "Booking confirmed",
.body = "Tee time 09:20 is locked in.",
.data = "{\"url\":\"/bookings/42\"}", // opaque JSON passed to the service worker
.ttl_s = 120, // push-service TTL header
});
switch (result) {
.delivered => {},
.gone => try ctx.records().delete("push_subscriptions", row.id), // prune a dead endpoint
.failed => {}, // transient; try again later
}
Tri-state result — prune on .gone. send returns .delivered (2xx), .gone (404/410 — the subscription is dead: delete your stored row and never retry), or .failed (5xx / 429 / timeout / transport error — transient). A failing push never throws into the request that triggered it; only an un-decodable subscription key surfaces as a Zig error (a corrupt stored row — terminal).
Background delivery (enqueue). ctx.push().enqueue(sub, msg, .{ .queue = "push" }) hands the delivery to the built-in "push" job kind. On .failed the queue retries with back-off; on .gone it simply stops retrying (a queued job cannot reach back into your DB to prune the row — the synchronous send path is how you learn .gone to prune); on .delivered it completes.
vapid-keygen. zigbase vapid-keygen prints a fresh base64url keypair. The public key doubles as the browser applicationServerKey; keep the private key secret.
ctx.verifyCaptcha() — CAPTCHA verification (#140)
Verify a browser-submitted CAPTCHA token against one of four supported providers: recaptcha_v2, recaptcha_v3, hcaptcha, turnstile.
Configure the provider + secret in App(cfg):
pub const app = zigbase.App(.{
.auth = .{
.captcha = .{
.provider = .recaptcha_v3, // .recaptcha_v2 | .recaptcha_v3 | .hcaptcha | .turnstile
.secret = "6LeXXXXXXXXX", // server-side site-verify secret
},
},
// ...
});
Verify in a route handler:
// Untyped route handler: a single `*zigbase.Ctx` argument returning an `http.Response`.
fn submitHandler(ctx: *zigbase.Ctx) anyerror!http.Response {
// Read the token however your frontend submits it — here from the query string;
// for a JSON/form POST, read `ctx.request.?.body` / `ctx.request.?.form_fields`.
const token = (try ctx.query()).get("captcha") orelse "";
const r = try ctx.verifyCaptcha(.recaptcha_v3, token);
if (!r.ok) return ctx.jsonError(403, "captcha_required", "Captcha required.");
// reCAPTCHA v3: score 0.0 (bot) → 1.0 (human); block suspicious traffic.
if (r.score) |score| if (score < 0.5) return ctx.jsonError(403, "suspicious_request", "Suspicious request.");
// ... proceed with the submission ...
}
CaptchaResult fields:
| Field | Type | Present | Notes |
|---|---|---|---|
ok | bool | always | true = token accepted |
score | ?f32 | reCAPTCHA v3 only | 0.0 = bot, 1.0 = human |
action | ?[]const u8 | reCAPTCHA v3 only | Action name from the frontend call |
hostname | ?[]const u8 | all providers | Domain that issued the token |
errors | []const []const u8 | when ok=false | Provider error codes |
Provider URL mapping:
| Provider | Verify URL |
|---|---|
recaptcha_v2 | https://www.google.com/recaptcha/api/siteverify |
recaptcha_v3 | https://www.google.com/recaptcha/api/siteverify |
hcaptcha | https://hcaptcha.com/siteverify |
turnstile | https://challenges.cloudflare.com/turnstile/v0/siteverify |
Dev-bypass: when app.captcha_secret is empty (the default when .auth.captcha is not configured), ctx.verifyCaptcha returns .{.ok = true} immediately — no network call, no live key needed. This lets local development and unit tests work without a real provider. A configured-provider-with-empty-secret logs a loud startup warning.
Never deploy with an empty secret — every
verifyCaptchacall returnsok=truewithout contacting the provider.
Errors propagate as a Zig error so the handler can tell a verdict (ok=false) apart from a non-verdict (an error) and decide fail-open vs fail-closed: error.TransportFailed (network), error.CaptchaProviderError (non-2xx reply), error.CaptchaParseError (malformed/non-object JSON):
const r = ctx.verifyCaptcha(.turnstile, token) catch |e| {
std.log.warn("captcha provider unreachable: {s}", .{@errorName(e)});
// fail-open: proceed; or return ctx.jsonError(503, "captcha_unavailable", "Captcha unavailable.") to fail-closed
return process(ctx);
};
ctx.realtime() — broadcast on custom channels
The realtime layer auto-publishes record-change events on <collection> / <collection>/<id> topics. ctx.realtime() lets a handler — a route or a background job — publish its own events on custom (non-record) channels over the same WebSocket, using the same subscribe/unsubscribe protocol clients already use:
// Signal-only (no payload). Subscribers receive {"type":"signal","topic":"availability"}
// and should re-fetch over an authenticated GET. The recommended default for private /
// per-subject state, since the channel carries nothing sensitive.
ctx.realtime().signal("availability");
// Payload-carrying. Subscribers receive
// {"type":"message","topic":"orders","data":{"type":"order.shipped","id":"REC1"}} verbatim.
try ctx.realtime().broadcast("orders", .{ .type = "order.shipped", .id = id });
A client subscribes to a custom topic exactly like a collection topic:
const ws = new WebSocket(`ws:///api/realtime`);
ws.onopen = () => ws.send(JSON.stringify({ action: "subscribe", topic: "availability" }));
ws.onmessage = (e) => { const m = JSON.parse(e.data); if (m.topic === "availability") refreshSlots(); };
Both publish entry points are a no-op when the realtime reactor isn’t running (tests/CLI), so they are safe to call unconditionally and from a background job (ctx.app.submit / a queue handler) where there is no HTTP request.
A custom topic is any topic name that is not a collection. Subscribing to a collection name always goes through that collection’s normal record-channel authorization (per-record viewRule); the custom-topic path can never be used to reach a collection’s records.
Connection capacity (.realtime.max_connections)
Compile .realtime = .{ .max_connections = 256 } into a small deployment, or raise the positive u32 cap for a larger machine. Omission preserves the existing 10,000 default; zero is a compile error, not an unlimited mode. Like .pools and .admission, this is a comptime capacity lever: changing it requires a rebuild. Resource profiles do not silently change this cap. resources --json reports the compiled value as realtime_connection_cap; superuser GET /api/realtime/stats reports the effective max_connections and current connections.
WebSocket and SSE share one process-wide atomic counter. Reservations include upgrades in progress and are released on failure or disconnect. New upgrades at capacity receive 503 before connection allocation; existing sessions are not evicted. This is not a per-user quota, cluster-wide cap, or total memory budget: transport buffers and each connection’s subscriptions still consume resources. Keep the outbound high-water mark enabled and tune against representative workloads.
Who may subscribe (.realtime = .{ .canSubscribe = fn })
By default, custom topics are public signal channels — anyone, including an anonymous socket, may subscribe (exactly the framework’s own __features signal). To gate a private channel, supply a predicate:
const App = zigbase.App(.{
.realtime = .{
// Return true to allow the subscription, false to deny it. `ctx` carries the
// socket's resolved identity (ctx.user() / ctx.rctx); `topic` is the requested
// custom channel name.
.canSubscribe = struct {
fn f(ctx: *zigbase.Ctx, topic: []const u8) bool {
if (std.mem.startsWith(u8, topic, "admin:")) {
const u = ctx.user() orelse return false;
return u.is_superuser;
}
return true; // other custom topics stay public
}
}.f,
},
});
Security guidance. Because a custom topic’s frame is delivered verbatim to every subscriber (no per-record viewRule), keep private/per-subject state signal-only: signal(topic) carries no data, so a subscriber learns only that something changed and must re-fetch the actual state over an authenticated GET. Use the payload-carrying broadcast(topic, payload) only for data that is safe for every subscriber of that topic, and use .canSubscribe to restrict who may join a private channel. The guard is reevaluated before custom-topic delivery as well as subscription; keep it read-only and inexpensive. Retained sessions are reverified before delivery, so changed two-factor requirements or revocation require reauthentication.
For contributor-only fanout measurements and their scope, see the realtime benchmark guide.
Multi-instance realtime (Postgres)
Realtime delivery is in-process: a record write publishes to facil.io’s pub/sub inside the writing instance, which is complete parity for the single-process SQLite story. When you run several app instances against one shared PostgreSQL (the reason to use the Postgres backend), ZigBase additionally fans record-change events across instances over Postgres LISTEN/NOTIFY — a write on instance A reaches subscribers connected to instances B and C. This is automatic when the active backend is Postgres (no configuration); each instance runs a dedicated listener connection (which auto-reconnects with backoff, logging loudly, if the connection drops on a PG restart/failover), and on a notification it runs its own per-subscriber viewRule/ability/tenant authorization before delivering. The notification payload is minimal and carries no row data — only {origin, collection, action, id}; create/update re-fetch the live row on the receiving instance, and a delete carries a random token that keys the deleted row’s snapshot in a small server-side table (_rt_delete_snapshots), which the receiver reads back over its own DB connection. This is deliberate: putting the deleted row in the NOTIFY payload would broadcast its column data — including the decrypted plaintext of .encrypted fields — to any DB role that can LISTEN, so the snapshot is stored at-rest (ciphertext) in the side table and never transits the wire. Delete authorization likewise evaluates the deleted row’s viewRule against the at-rest (ciphertext) snapshot on every path — local and cross-instance — matching the live create/update path (which compares the ciphertext column); a rule that references an .encrypted field therefore authorizes a delete identically everywhere. On SQLite (single-process) nothing changes — there is no cross-process step, no side table, and the in-process path is byte-identical (for the common case of no encrypted fields the delete takes no extra read). Custom-channel ctx.realtime().broadcast/signal events also fan out cross-instance on Postgres (best-effort, at-most-once, unordered): a signal NOTIFYs only the topic name (payload-less), while a broadcast stores its enveloped frame in the _rt_broadcasts side table keyed by a random token and NOTIFYs only the token — receivers read the frame back over their own connection and re-deliver it through the same per-subscriber authorization path (a forged/expired token finds no row and is dropped). The __features flag/experiment signal is cross-instance for the same reason. On SQLite (single-process) these stay in-process and byte-identical.
Test-mode capture — assert sent mail + mock outbound HTTP (zigbase.testcapture)
For deterministic e2e/integration tests, the framework can capture what it sent — an in-memory mail outbox and a record of every outbound ctx.http() call — and inject canned HTTP responses instead of hitting the network. It mirrors the determinism seam (ZIGBASE_FAKE_NOW) and shares the same comptime gate: it is compiled in only on a dev_mode build (on in Debug, off in any release build), so a production binary is byte-for-byte unaffected — zigbase.testcapture.enabled is comptime false there, the seams fold away, and there is no runtime branch or perf cost. Every API below is a no-op / returns empty when the gate is off.
Mail outbox. Mailer.send (the single seam every backend — Log/SMTP/Command/your own plugin — routes through) records each email when capture is on:
const tc = zigbase.testcapture;
tc.mail.enable(true); // capture + SUPPRESS real delivery (deterministic e2e mode)
defer tc.mail.reset(); // clear + free
// ... trigger a flow that sends mail (signup verification, password reset, your route) ...
try expectEqual(@as(usize, 1), tc.mail.count());
const e = tc.mail.get(0).?; // { from, to, subject, body }
try expectEqualStrings("user@example.com", e.to);
try expect(tc.mail.find("Verify") != null);
mail.enable(suppress): suppress = true records and skips real delivery; false records and still delivers (assert against a live MailHog/log run). from is the backend’s configured sender (empty for LogMailer, which has none). Beyond the indexed accessors, mail.entries() returns the whole captured slice ([]const MailEntry), mail.disable() stops capturing while keeping already-captured entries readable, and mail.reset() clears
- frees.
HTTP capture / mock. HttpClient.request (what ctx.http() returns) consults the capture before touching the network — recording the request and, if a mock matches the URL substring, returning the canned response with no network at all:
tc.http.enable(true); // capture; block_unmocked=true → un-mocked URLs error (no network)
defer tc.http.reset();
tc.http.mock("api.stripe.com", .{
.status = 200,
.headers = &.{.{ .name = "Content-Type", .value = "application/json" }},
.body = "{\"id\":\"ch_123\",\"paid\":true}",
});
// ... trigger a route/hook that calls ctx.http().post("https://api.stripe.com/...") ...
try expectEqual(@as(usize, 1), tc.http.requestCount());
const rq = tc.http.requestAt(0).?; // { method, url, headers, body }
try expectEqual(zigbase.HttpMethod.POST, rq.method);
http.enable(block_unmocked): with block_unmocked = true (recommended) a request to an un-mocked URL fails with error.TransportFailed rather than silently hitting the network; false lets un-mocked URLs pass through to the real network. Mocked responses are matched newest-registration-first on URL substring. Beyond the indexed accessors, http.requests() returns the whole captured slice ([]const HttpRequest), http.disable() stops capturing while keeping recorded requests + mocks readable, and http.reset() clears + frees.
Error helpers (ctx.fail, ctx.invalid, ctx.errorResponse)
For untyped route handlers, use the Ctx helpers to produce standard error responses rather than building raw HTTP error responses by hand:
fn confirmBooking(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
const req = ctx.request.?;
const id = req.param("id") orelse
return ctx.errorResponse(ctx.fail(400, "Missing booking id."));
const booking = (try ctx.records().get("bookings", id, .{})) orelse
return ctx.errorResponse(error.NotFound);
if (!isOwner(ctx.user(), booking))
return ctx.errorResponse(ctx.fail(403, "Not the owner."));
// ... update booking ...
}
ctx.fail(status, message) stashes the error and returns error.Handled. ctx.errorResponse(err) maps any error to a {code, message, data} JSON response:
anyerror | HTTP status |
|---|---|
error.Handled | renders the stashed ctx.fail / ctx.invalid error |
error.NotFound | 404 |
error.Forbidden | 403 |
error.Unauthorized | 401 |
error.BadRequest | 400 |
error.Conflict | 409 |
| anything else | 500 (no detail leaked) |
ctx.invalid(fields) stashes a 400 validation error with per-field detail:
return ctx.errorResponse(ctx.invalid(&.{
.{ .field = "email", .code = "invalid", .message = "Must be a valid email." },
}));
Idempotent custom mutations (opt-in SQLite/PostgreSQL)
zigbase.Idempotency(.{ .namespace = "booking-cancel-v1", .max_entries = 1024, .retention_seconds = 86400, .max_payload_bytes = 65536, .max_result_bytes = 4096, .cleanup_batch = 64 }) is a lazy generic helper, not global HTTP middleware. Instantiating and calling it opts in; unused applications have no receipt schema, state, timer, callback registration, or request overhead. There is no build flag or App configuration field for this helper. All its resource limits are comptime.
Call Receipts.execute(allocator, writer, input, callbacks). Acquire the pool writer once and defer its release; do not call from an existing ctx.tx or hook transaction. Caller-owned transactions are refused before doing any work. PostgreSQL requires the existing -Dpostgres=true build. IdempotencyInput contains:
principal = .{ .collection = authenticated_collection, .record = authenticated_id }: both come from verified server identity, never from submitted JSON.operation: a stable, versioned server operation name;key: the client’s retry key.payload: exact bytes covering all mutation inputs, including resource IDs in the URL. Different serialization means a different payload; JSON is not canonicalized.now: trusted server Unix seconds captured for this attempt. Expiry is exactlynow + retention_seconds, not commit time. Writer wait and mutation duration consume retention; choose an adequate interval, use short callbacks, and keep participating process clocks synchronized. Expiry permits the operation to run again.
Every principal component, operation and key must be 1..128 bytes. The namespace is 1..128 bytes, entries 1..1,000,000, retention 1..31,536,000 seconds, payload and result budgets 1..1,048,576 bytes, cleanup batch 1..1024. Runtime inputs are checked before hashing or SQL. Namespace separates capacity and cleanup; collection, record, operation and key are length-framed then SHA-256 hashed, so ambiguous concatenations cannot cross identities. Only hashes and result bytes are stored, not raw keys or payloads. Results may contain secrets: protect the database/backups.
IdempotencyCallbacks holds an opaque context and two required callbacks:
fn authorize(conn: *zigbase.Db, context: *anyopaque) anyerror!void;
fn mutate(conn: *zigbase.Db, context: *anyopaque, output: []u8) anyerror!usize;
authorize must be read-only and must recheck current access using the supplied writer, including on replay. Authentication and token validation still happen on every HTTP request before entering the helper. mutate uses only that connection, writes its response bytes into the supplied bounded buffer, and returns bytes used. An optional authorize_replay(conn, context, stored_result) read-only callback can additionally authorize the stored result before a replay commits; its error rolls back and releases the owned replay buffer. It does not run on a new mutation. No callback may end the transaction or reacquire the writer. These are trusted application callbacks, not a SQL sandbox; arbitrary callback memory/CPU is not bounded. The return owns its allocation: read result.body() / result.replayed, then result.deinit(allocator). Copy the body into your response’s lifetime before deinit.
An SQLite BEGIN IMMEDIATE covers current authorization, lookup, DB mutation and receipt insertion. PostgreSQL uses a helper-owned READ COMMITTED transaction and a fail-fast, transaction-scoped advisory lock per namespace. Namespace-lock contention returns IdempotencyBusy before authorization or mutation; retry the same key later. All operations within one namespace serialize, including distinct keys, to enforce strict capacity. Different namespaces can proceed independently after first-use schema creation. Cold creation uses a separate fail-fast advisory lock, so a concurrent first attempt in another namespace may also return IdempotencyBusy after read-only authorization but before mutation. Locks release on commit, rollback or connection loss; no session lock leaks into the pool. Hash collisions can cause conservative contention, never shared receipts. Independent writers cannot both execute an uncommitted key; lock contention may return a DB error and is safe to retry. If commit succeeds but the response is lost, the same authorized attempt returns the saved bytes without calling mutate. Changed payload under a live key returns PayloadConflict. Mutation, allocation, oversized-output and commit errors roll back effects and receipt together. A rollback failure is logged while the original operation error is returned. The helper does not discard, replace or reset a connection, and does not evict a pooled writer. Its transaction state is uncertain: the caller must stop reuse and recover the connection or pool (for example, stop the serving process and reopen it after diagnosing the storage failure), not merely release that writer back to normal service.
Live receipts are never evicted to admit new keys. Capacity exhaustion returns CapacityExceeded before mutation. Each attempt deletes at most cleanup_batch expired receipts plus its own expired target, using an expiry index. Capacity refusals commit housekeeping only, so lowering capacity cannot permanently wedge cleanup. Other failures roll housekeeping back. No timer deletes idle data: expiry ends replay protection but physical receipts can remain until later attempts. All processes for a namespace should use the same limits. Lowering capacity refuses new keys until rows fit; lowering retention never changes stored expiries. Lowering result budget below an existing result refuses its replay, rather than rerunning it. The logical bound is per namespace (including retained result bytes); number of namespaces, database pages/WAL, disk reclamation and total process RSS are not bounded. PostgreSQL stores namespace and result bytes as BYTEA (including NUL and non-UTF8 bytes) and expiry as BIGINT. Its text-protocol binding uses bounded hexadecimal scratch (twice the result length); replay queries limit result bytes before the driver buffers a row, even after lowering the configured result budget. The ledger must resolve to a permanent ordinary table in current_schema(); temporary/unlogged objects, views and fallback schemas fail closed with InvalidReceiptLedger, before mutation. Database copying applies the same check to both connections; configure their search paths deliberately. Quoted schema names are supported; cold creation revalidates the new ledger’s ownership before cleanup or mutation, rolling back temporary-schema creation. SQLite requires an ordinary ledger table in main; temporary objects, views, virtual tables and any attached-database ledger with the same name (even when main also owns one) are refused before mutation or copying. Name checks are case-insensitive, matching SQLite identifier resolution.
This is database-only atomicity, not exactly-once email, HTTP, files or realtime. Never perform these effects in a callback; use a transactional outbox if needed. The helper does not automatically invoke REST hooks/rules or serialize responses. The golfsim example shows an authenticated custom route with current guest authorization and a bound Ctx. Built-in REST retries are separately opt-in; see REST record idempotency. External side-effect orchestration remains outside this helper.
Database copying: migrate-db/dumpload.run refuse any source or target with non-empty _idempotency_receipts, even with --force or same-backend copies, before writing target schema or data (IdempotencyReceiptsPresent). Lazy ledgers cannot otherwise be copied safely by the target-driven transfer. Drain all participating applications, wait out stored retention windows, verify receipt expiry against trusted server time, and explicitly delete only verified-expired receipts before retrying the copy. The copier never expires or deletes them for you. An empty ledger does not prevent migration.
REST record idempotency
Compile with -Drest-idempotency=true, then opt specific base collections into persistent retry receipts:
const App = zigbase.App(.{
.rest_idempotency = .{
.collections = .{ "posts", "tasks" },
.limits = .{
.namespace = "my-app-rest-v1",
.max_entries = 1024,
.retention_seconds = 86400,
.max_payload_bytes = 65536,
.max_result_bytes = 4096,
.cleanup_batch = 64,
},
},
});
An authenticated POST, PATCH, or DELETE to the built-in records endpoint can carry Idempotency-Key: <opaque-key>. Keep the same key and exact body bytes when retrying the same logical operation. No automatic retry is performed. The receipt and database mutation commit in one transaction on SQLite or PostgreSQL; a retry after a lost response or process restart returns the original status/body without repeating the write. Concurrent PostgreSQL attempts can return retryable 503 while another process holds the namespace lock. Unkeyed requests keep their ordinary behavior. Keyed requests are refused with 400 if support is not compiled or configured, rather than silently executing without retry protection.
Keys are scoped by the authenticated collection’s stable ID, principal record ID, operation, target collection ID and target record ID. The fingerprint also binds the exact request body, active account and full current collection metadata. A key is 1–128 bytes. A changed body, schema, account or existing record representation returns 409; changes that make the collection ineligible (such as adding a hidden field) are refused with 400 before receipt lookup. Changing the namespace or waiting past retention permits new execution. All replicas sharing a namespace must agree on limits. Live receipts are never capacity-evicted: a full namespace returns 503 before mutation. Expired cleanup is bounded by cleanup_batch; retention starts at attempt time, including lock wait. Receipt storage is bounded per namespace by max_entries and max_result_bytes plus ledger metadata; request/schema scratch and database page overhead are separate. max_payload_bytes also covers the schema/target fingerprint input, so large schemas may require a larger budget. Limits follow the custom Idempotency helper’s ranges.
Authentication and current collection rules are checked inside every transaction, including replay. Replay view rules use an ordinary GET context (no mutation body), while action rules retain the mutation method and body. The transaction holds the schema-generation lock through commit, so concurrent schema/rule writers cannot invalidate that authorization mid-mutation. Create/update replay additionally requires current action and view permission and an unchanged surviving record; deleted or no-longer-authorized targets return 404. Deletes only support a current rule that permits the operation without per-row predicates (for example @public, or a superuser). Predicate-constrained deletes are refused because the deleted row cannot establish current authorization. This intentionally makes replay stricter than ordinary unkeyed CRUD.
Supported keyed POST/PATCH requests use JSON objects; keyed DELETE requires an empty body. All keyed requests are without query parameters, on allowlisted base collections without TTL, file, hidden or encrypted fields. Apps with any record hooks (including hooks on unrelated collections) refuse keyed operations: the adapter does not skip hooks or pretend arbitrary effects are transactionally replayable. Auth collections, multipart uploads and resumable commits retain their own workflows. Use custom Idempotency callbacks for trusted application operations with richer transactional authorization.
The ordinary record validation, transactional writes, tenant stamping and composed access rules are reused. Successful first attempts emit realtime notifications once after commit; replay emits none. A crash between commit and notification can lose the notification. With -Ddurable-realtime=true, the first mutation also captures a journal entry in the record/receipt transaction; journal failure rolls back both, and retries append no duplicate. Clients can recover that invalidation through durable backfill. This is database retry protection, not exactly-once external delivery. If transaction recovery cannot restore an idle, healthy writer, the pool refuses further SQL on that writer and keyed requests return 503. Operator recovery requires replacing the pool or restarting the process; the framework does not silently reopen the connection or discard an in-memory database.
Receipt response bodies are stored as plaintext in the database; use database backup and access policies appropriate to those records. Removing or hiding fields invalidates existing fingerprints rather than replaying the old response.
For browser applications, configure your reverse proxy’s CORS policy to allow Idempotency-Key where cross-origin requests are permitted. ZigBase does not add a new CORS policy for this feature. The TypeScript SDK accepts an explicit idempotencyKey option on record create/update/delete.
ctx.tx() — multi-write transactions
To write several records atomically — all commit or all roll back — define a file-scoped callback and pass it to ctx.tx:
fn transferCredits(t: *zigbase.Tx) anyerror!void {
var debit: std.json.ObjectMap = .empty;
try debit.put(t.arena(), "balance", .{ .integer = new_balance });
_ = try t.records().update("accounts", from_id, .{ .object = debit });
var credit: std.json.ObjectMap = .empty;
try credit.put(t.arena(), "balance", .{ .integer = credited });
_ = try t.records().update("accounts", to_id, .{ .object = credit });
}
// in a route / job / hook handler:
try ctx.tx(void, transferCredits);
The callback receives a *Tx — a thin wrapper whose t.records() returns the same Records API as ctx.records(). All writes inside the callback reuse the single in-transaction writer connection (no second acquire, no deadlock). If the callback returns any error the transaction rolls back automatically; on success it commits.
The first type parameter is the callback’s return type (use void when you do not need to surface a value from the transaction):
fn countAndTag(t: *zigbase.Tx) anyerror!u64 {
_ = try t.records().create("tags", value);
return 1;
}
const n = try ctx.tx(u64, countAndTag);
Attempting to call ctx.tx from a callback that is already running inside a transaction returns error.NestedTransaction immediately (without beginning a new transaction).
Do not perform long network calls (
ctx.http()) inside actx.tx/ctx.txWithcallback. The writer connection is held for the entire duration of the callback. Long I/O stalls all other writes on the server — complete any external HTTP calls before or after the transaction block.
In a
before*hook,ctx.records()is bound to the hook’s in-transaction connection and never acquires from the pool, so side-writes commit atomically with the triggering write. In route and job handlersctx.records()lazily checks out a pooled connection that the framework releases on ctx teardown.
ctx.txWith() — threading request data into a transaction (#237)
ctx.tx’s callback is a bare *const fn(*Tx) anyerror!T — Zig has no closures, so a callback that needs data computed in the handler (an id, a validated input struct, a priced total) has nowhere to get it from. Reaching for a threadlocal to smuggle it in works but is a footgun: a stale value from a previous request, a name collision between two routes, and an easy-to-forget set/defer-clear ritual around every call site.
ctx.txWith is the same transaction scope with one addition: a caller-supplied payload threaded straight into the callback as its second argument — no globals:
const ConvertParams = struct { hold_id: []const u8, booking: std.json.Value };
fn convertHoldTxn(t: *zigbase.Tx, p: *const ConvertParams) anyerror!std.json.Value {
const created = try t.records().create("bookings", p.booking);
if (!try t.records().delete("holds", p.hold_id)) return error.HoldVanished;
return created;
}
// in the route handler:
const params = ConvertParams{ .hold_id = id, .booking = booking_value };
return try ctx.txWith(std.json.Value, ¶ms, convertHoldTxn);
payload is typically a pointer to a stack struct built right before the call — it only needs to stay valid for the (synchronous) duration of txWith. Everything else matches ctx.tx: an IMMEDIATE transaction on the writer, commit on success, automatic rollback on any callback error, and error.NestedTransaction if the Ctx is already bound. Reach for ctx.tx when the callback needs no external data; reach for ctx.txWith the moment it does.
ctx.markSchemaChanged() — announce a hand-written metadata change
ZigBase caches parsed collection definitions, and a background observer drops that cache whenever collection metadata changes — including changes made by a different process (a zigbase migrate, zigbase import, or another instance sharing the database). It notices via a schema-generation marker that every engine write to _collections (create / update / delete / rule updates) bumps inside its own transaction.
You never need to think about this unless you write collection metadata yourself. ctx.records().queryAs() runs unrestricted SQL on the transaction’s writer, so a hook or custom route can do this:
_ = try t.records().queryAs(struct {}, "UPDATE \"_collections\" SET \"listRule\" = ?1 WHERE \"name\" = ?2;", .{ "@request.auth.id != \"\"", "posts" });
try t.markSchemaChanged(); // ← without this, the OLD rule keeps being enforced
That write is invisible to the engine, so without the call the stale definition — including a stale access rule — keeps being enforced until the process restarts. Call it in the same transaction as the write; the marker then commits or rolls back with it.
Detection is deliberately manual. PRAGMA schema_version would catch a raw DROP/ALTER but not a rule-only UPDATE — partial coverage for substantial complexity, which is worse than one explicit call on the rare path that needs it.
The observer polls every 5 seconds, so a cross-process metadata change is picked up within that window. It is not a configuration knob: there is no config key and no environment variable for it.
ctx.kv() — built-in key/value settings store
Not every piece of mutable server state deserves its own collection. For small, server-managed values — a maintenance toggle, a cached external token, a counter, a JSON settings blob — ctx.kv() is a built-in key→value store backed by an internal _kv system table. No schema, no access rules, no ceremony:
try ctx.kv().set("welcome_banner", "Closed for maintenance");
const banner = try ctx.kv().get("welcome_banner"); // ?[]const u8 (null if unset)
_ = try ctx.kv().delete("welcome_banner"); // bool: did a row exist?
Values are opaque TEXT. To store structured data, stringify it yourself (e.g. std.json.Stringify.valueAlloc(ctx.arena.a, value, .{})) and parse it back on read. set is an upsert that preserves the original created timestamp and bumps updated.
Reads use the read connection; writes use the bound connection inside a before-hook or ctx.tx (so they commit atomically with the surrounding transaction) and otherwise acquire/release the pool writer for you — the same lifetime rules as ctx.records().
The same primitive is on the curated Data facade as data.kvGet / data.kvSet / data.kvDelete (and data.kvList) for internal/test consumers.
Security: KV/settings are superuser-managed and never public by default. The
_kvtable is not a collection, so it is invisible to the record API, query engine, and access-rule system. To expose a value publicly, write a custom route that returns exactly what you intend (see feature flags below).
App-scoped context (.app_context / ctx.appData)
ctx.kv() is for persisted, mutable settings. When what you need instead is in-memory, process-lifetime state — parsed config, a preloaded content bundle, a shared client handle, a lookup table computed once at boot — reach for the app-scoped context. It replaces the “module-level var globals + a bootstrap setter ritual” pattern with one explicit, typed handle that every handler, hook, job, and cron reads the same way.
Declare the context type at comptime, set it once in onBootstrap, and read it anywhere as a *T:
const AppData = struct { cfg: Config, content: SiteContent };
// 1. Declare the type on the App config:
pub const App = zigbase.App(.{
.app_context = AppData,
.onBootstrap = boot,
// ... your other keys
});
// 2. Build the value and install the handle ONCE, in onBootstrap.
// `app_data` must outlive the app. A clean, safe pattern is to allocate it on
// the heap with `ctx.app.allocator` (process lifetime) — no module-global state,
// no dangling pointer from a stack local.
fn boot(ctx: *zigbase.Ctx, ev: *zigbase.events.LifecycleEvent) anyerror!void {
_ = ev;
const gpa = ctx.app.allocator;
const app_data = try gpa.create(AppData);
app_data.* = .{ .cfg = try loadConfig(gpa), .content = try loadContent(gpa) };
try ctx.setAppData(AppData, app_data);
}
// 3. Read it from any handler / hook / job / cron.
fn handler(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
const d = ctx.appData(AppData); // *AppData
return ctx.json(.{ .site = d.content.title });
}
Caution — allocate for process lifetime, not from
ctx.arena.a. Thectx.arenainside a lifecycle hook is a per-invocation arena the framework frees the instant the hook returns (before any request is served). Theapp_datapointer above survives because it’s heap-allocated viactx.app.allocator, not a stack local or actx.arena.aallocation. Allocate everything the handle transitively owns (config strings,content.title, …) fromctx.app.allocatortoo (the App allocator, which lives for the process) — or makeAppDatafully by-value/static. Anything allocated fromctx.arena.abecomes a dangling read on the first request. The pointer and everything it reaches must outlive the app.
Semantics
- Single-set.
setAppDatainstalls the handle exactly once. A second call returnserror.AppContextAlreadySetand overwrites nothing — the first value stays live. - Startup contract (fail-fast). Declaring
.app_context = Tmakes setting it mandatory: ifonBootstrapreturns without callingctx.setAppData, the server refuses to start with an actionable error (it never boots into a state where the firstctx.appDataread would fault). Apps that don’t declare.app_contextpay nothing — the twoAppfields default to null and nothing enforces anything. - Lifetime is yours. The framework stores only a type-erased pointer; the value must outlive the running app — heap-allocate it via
ctx.app.allocator(as in the example above), not a stack local (destroyed the momentonBootstrapreturns) or ad hoc module-global state. The handle is read-shared across request threads — guard interior mutability yourself if you mutate it after boot. onBootstrapreturnsanyerror!void, sotry ctx.setAppData(...)composes naturally (as do the siblingonBeforeServe/onBeforeTerminatelifecycle hooks — an error fromonBootstrap/onBeforeServefails the boot; anonBeforeTerminateerror is logged during shutdown).
Type-guard tradeoff. App is a concrete (non-cfg-parameterized) struct and Ctx lives in a separate file, so ctx.appData(T) cannot perform a cross-file comptime check of T against the declared .app_context type. Instead it enforces the type at runtime: setAppData records @typeName(T), and appData asserts the stored name matches the T you ask for. Calling appData with the wrong T (or before it is set) trips a loud std.debug.assert in safe builds. This is a deliberate choice — it keeps the shared Ctx/App types uncontorted; the comptime .app_context = T declaration must be a type (a non-type value is a @compileError), and the boot contract catches the declared-but-unset case before any request runs.
Feature flags + experiments (declared)
Breaking in 0.8.0. Flags are now declared-only: you list them in the
App(cfg)literal and access them through a typed accessor whose.nameis checked at compile time. The old runtime-stringctx.flag("arbitrary")(KV-or-false) API is removed — see the migration note at the end of this section.
Declare flags and experiments alongside the rest of your config:
pub const App = zigbase.App(.{
.flags = .{
.checkout_enabled = true, // bare bool = default
.new_dashboard = .{ .default = false, .description = "Dark mode shell" },
},
.experiments = .{
.checkout_layout = .{ .variants = .{ "control", "compact" }, .weights = .{ 50, 50 } },
},
});
A flag is either a bare bool (its default) or .{ .default = <bool>, .description = "…" }. An experiment is .{ .variants = .{…}, .weights = .{…} } (parallel tuples; weights need not sum to 100) with optional .sticky and .description. Malformed declarations (unknown sub-key, length mismatch, empty/duplicate variants, all-zero weights) are compile errors.
Typed accessors live on the App type — a typo’d .name is a compile error, not a silent miss:
if (App.flag(ctx, .checkout_enabled)) { // bool; unset → declared default
// ... new path
}
try App.setFlag(ctx, .checkout_enabled, false); // operator kill switch: writes the override
const variant = try App.experiment(ctx, .checkout_layout, user_id); // []const u8 ("control"/"compact")
App.flag returns the flag:<name> override from the KV store if one is set, otherwise the declared default (so a default-ON kill switch stays on until an override sets "false"); it swallows read errors back to the default so a check never fails the request. App.experiment buckets subject deterministically over the weights (FNV1a-64(name ++ 0x00 ++ subject)), so the same (name, subject) always lands on the same variant; an empty subject maps to the first variant. A weight override in _kv under exp:<name>:weights (JSON, e.g. [90,10]) changes the split without a redeploy.
Sticky assignments (.sticky = true). By default an experiment is a pure function of (name, subject, weights) — change the weights and a subject can re-bucket to a different variant. Declare .sticky = true to persist a subject’s first assignment so it survives later weight changes (#129):
.experiments = .{
.checkout_layout = .{ .variants = .{ "control", "compact" }, .weights = .{ 50, 50 }, .sticky = true },
},
The first App.experiment(ctx, .checkout_layout, subject) for a non-empty subject buckets it and writes the result to the internal _experiment_assignments table (keyed by (experiment, subject)); every later resolve returns the stored variant, even after you edit exp:<name>:weights. New subjects still follow the current weights. An empty subject is never persisted (it always maps to the first variant). The sticky read/write needs the writer, so sticky resolution costs one extra write only on the first sighting of a subject; pure-hash experiments touch no table at all.
Stale assignments are reaped by a framework-internal _experiment_gc job that is installed only when at least one declared experiment is .sticky (zero overhead otherwise) and runs hourly, deleting rows older than .experiment_assignment_ttl (in days, default 90) in bounded batches on the writer:
zigbase.App(.{
.experiments = .{ .checkout_layout = .{ .variants = .{ "control", "compact" }, .weights = .{ 50, 50 }, .sticky = true } },
.experiment_assignment_ttl = 30, // reap sticky assignments older than 30 days
});
Setting .experiment_assignment_ttl without any .sticky experiment is a @compileError (it would be a silent no-op).
Dynamic names + batch resolution on ctx (the runtime escape hatches):
const on = ctx.flagByName("checkout_enabled"); // ?bool — null when the name is undeclared
const all = try ctx.flags().resolveAll(user_id); // every declared flag + experiment
// all.flags: []{ name, value: bool } all.experiments: []{ name, variant: []const u8 }
Zero-DB steady-state resolution. Flag/experiment resolution (App.flag, ctx.flagByName, ctx.flags().resolveAll, the exp:*:weights override read, and the /api/state projection) is served from an in-process override cache of the flag:* / exp:*:weights _kv entries — so a hot request costs no _kv reads. A same-instance override write (App.setFlag, the admin settings verbs) invalidates the cache instantly, so a kill-switch flip takes effect on the very next request. On Postgres, another instance’s override write self-heals within a 5 s staleness bound. (The sticky _experiment_assignments reads are separate and remain reader-first + batched.)
Public read plane (GET /api/state). A built-in, unauthenticated projection returns every declared flag + experiment resolved for a caller-supplied subject:
// GET /api/state?subject=user-42 (no auth)
{ "flags": { "checkout_enabled": true }, "experiments": { "checkout_layout": "compact" } }
It serves resolved values only (never the _kv keys, defaults, weights, or the superuser settings verbs). A .sticky experiment returns its persisted assignment here too — resolved reader-first, so a repeat call for a known subject is read-only and only a subject’s first resolve briefly takes the writer (an unauthenticated caller can’t storm the writer lock). It auto-mounts at /api/state; the .features knob remaps or disables it:
App(.{
.features = .{ .public_route = "/state" }, // remap (default "/api/state")
// .features = .{ .public_route = .disabled }, // turn off entirely (404)
});
The typed TypeScript SDK surfaces this as zb.flags.resolveAll(subject) (named-boolean flags + variant string-unions) — see TypeScript SDK → Typed feature state. For a bespoke shape you can still write your own custom route over ctx.flagByName / ctx.flags().resolveAll.
Migrating from 0.7: declare each flag you used in .flags, then replace try ctx.flag("x") with App.flag(ctx, .x) (compile-checked) or ctx.flagByName("x") (dynamic), and try ctx.setFlag("x", v) with try App.setFlag(ctx, .x, v).
Exposure events (.onFeatureExposure)
Register .onFeatureExposure to feed an analytics/exposure pipeline: it fires every time a declared flag or experiment is resolved (the App.flag/ctx.flagByName and App.experiment read paths).
fn onExposure(ev: *zigbase.ExposureEvent) void {
switch (ev.kind) { // .flag | .experiment
.flag => log.info("flag {s} = {}", .{ ev.name, ev.value }),
.experiment => log.info("exp {s} [{s}] -> {s}", .{ ev.name, ev.subject, ev.variant }),
}
}
// .onFeatureExposure = onExposure,
ExposureEvent is { app, kind: enum { flag, experiment }, name, subject, value (flag), variant (experiment) }. For a .flag exposure value is the resolved boolean and subject is empty (flags are global); for an .experiment variant is the resolved variant and subject is the bucketing subject. The handler is notify-only — it cannot abort or write. It is zero-cost when unregistered: the resolver short-circuits before constructing the event, so an app without .onFeatureExposure pays nothing on the read path.
Realtime signal (__features)
Any override change — ctx.setFlag/App.setFlag, or an admin PUT/DELETE of a flag:<name> / exp:<name>:weights setting — broadcasts a single frame {"type":"signal","topic":"__features"} on the fixed public realtime channel __features (over the existing WebSocket). It is signal-only: no per-subject state or experiment assignment is ever pushed. Clients subscribe anonymously to __features and re-GET /api/state (or call ctx.flags().resolveAll) on receipt to pull fresh resolved values:
const ws = new WebSocket(`ws:///api/realtime`);
ws.onopen = () => ws.send(JSON.stringify({ action: "subscribe", topic: "__features" }));
ws.onmessage = (e) => {
const m = JSON.parse(e.data);
if (m.type === "signal" && m.topic === "__features") refetchState();
};
The signal is transport-agnostic — WebSocket and SSE subscribers receive the identical frame.
Changed in 0.10.0: this channel previously emitted a bespoke {"type":"features.changed"} frame; it now uses the same standard signal frame as every custom topic.
Superuser settings HTTP API
For administrative management there is a built-in superuser-only HTTP surface over the KV store (every endpoint requires a valid superuser token):
| Method & path | Effect |
|---|---|
GET /api/settings | List all settings ({"items":[{key,value,created,updated}, …]}). |
GET /api/settings/:key | Fetch one ({key,value}); 404 if absent. |
PUT /api/settings/:key | Upsert; body {"value":"…"}. |
DELETE /api/settings/:key | Remove; 204, or 404 if absent. |
The embedded admin UI exposes these endpoints as a “Settings” section where superusers can view, create, edit, delete entries, and toggle boolean flags with a checkbox — no API client required.
This is the management plane; the public read plane is the built-in GET /api/state endpoint (above), or a custom route if you need a bespoke shape.
Admin UI
The embedded admin UI (/_/) covers Collections (schema editor), Records (browse/query/edit), Users, Files (browse/upload, below), Email (senders, suppressions, batches, below), Logs (analytics events + rollups + realtime health, below), Features (flags & experiments, below), and Settings (raw KV, above). It ships as browser-native ES modules — no build step — and every asset is served with a CRC32 ETag.
Users (/_/#/users) manages both superusers and auth-collection users: list and search (filter-based) either, create/edit/delete records, an admin password reset (superuser PATCH), and a read-only OAuth-providers panel per auth collection. It composes the existing collections/records/auth REST API — no new endpoints ship for it. Per-device sessions are not yet exposed here.
Files (/_/#/files) covers per-collection file-field browsing: a read-only storage-backend strip (local disk vs S3, from the new superuser GET /api/files/config — non-secret backend info only, never the S3 credentials), a collection picker limited to collections with at least one file field, a records browser with image thumbnails and download links, and a per-record file drawer to upload, replace, or remove a file’s contents. It composes the existing records + file-serve API — see api.md → Files.
Email (/_/#/email) covers the email subsystem for superusers, in three tabs plus a read-only policy strip:
- Senders — list, invite (
POST /api/senders), and delete verified sender identities. - Suppressions — list, add, remove, and filter by reason (including one-click-unsubscribe entries).
- Batches — read-only bulk-send progress; there is no send/cancel action since no such endpoint exists yet.
The policy strip reads the new superuser-only GET /api/mail/config — booleans only (require_verified_sender, check_suppression, whether a webhook secret and an unsubscribe base URL are configured) — never secrets. See api.md → Email.
Logs (/_/#/logs) is a Logs & realtime view: a filterable, cursor-paginated browser for the raw analytics events feed (name/actor/since filters, expandable payloads), a viewer for an app-declared rollup’s aggregated series (by name, with from/to bounds), and a read-only realtime health strip (live connection count plus the connection/subscription caps and configured outbound high-water-mark, from the new superuser GET /api/realtime/stats). The tab is capability-gated: it only appears when the app configures .analytics (the stock zigbase serve binary doesn’t enable it, so the tab is hidden there). It composes the existing analytics API — see api.md → Analytics and api.md → Realtime.
The embedded admin UI (/_/) has a dedicated Feature Flags & Experiments screen (/_/#/features) that exposes the comptime-declared registry to superusers without any API calls:
- Declared flags table — every declared flag with its name, default, description, current effective value, and source (override vs default). A checkbox sets or clears the
flag:<name>override; a Use default button appears when an override is active and removes it on click. - Declared experiments panel — one card per experiment showing variants and weight inputs. Edit the weights and click Save weights to write the
exp:<name>:weightsoverride; Reset to declared deletes it. When no override is active the inputs show the declared weights.
The screen reads the declared registry via GET /api/features (superuser-only) and all mutations go through the existing PUT/DELETE /api/settings/:key verbs, so the raw Settings screen (also in the sidebar) shows the same override rows if you need lower-level inspection.
6. Auth / file / lifecycle events
One handler each, registered by the matching config key:
| Key | Signature | When |
|---|---|---|
onAuth | fn (ev: *zigbase.events.AuthEvent) void | Notify-only, after a session is issued (login / oauth2). |
beforeAuthSuccess | fn (ctx: *zigbase.Ctx, ev: *zigbase.events.AuthSuccessEvent) anyerror!void | Writable + abortable, before the session is issued. See Auth lifecycle. |
auth.hooks | struct of fn (ctx: *zigbase.Ctx, ev: *zigbase.events.AuthLifecycleEvent) anyerror!void | before/after register/logout/refresh/password-change (under the .auth config group). See Auth lifecycle hooks. |
onFileServe | fn (ev: *zigbase.events.FileEvent) anyerror!void | Before serving a download; return an error to deny (framework → 404). |
onFileUpload | fn (ev: *zigbase.events.FileEvent) void | After a successful upload. |
onBootstrap | fn (ctx: *zigbase.Ctx, ev: *zigbase.events.LifecycleEvent) anyerror!void | After bootstrap. An error fails the boot. Typical home of try ctx.setAppData(...) (see App-scoped context). |
onBeforeServe | fn (ctx: *zigbase.Ctx, ev: *zigbase.events.LifecycleEvent) anyerror!void | Just before serving starts. An error fails the boot. |
onBeforeTerminate | fn (ctx: *zigbase.Ctx, ev: *zigbase.events.LifecycleEvent) anyerror!void | Just before shutdown. An error is logged (it fires in a shutdown defer). |
onFeatureExposure | fn (ev: *zigbase.ExposureEvent) void | Notify-only, each time a declared flag/experiment is resolved. Zero-cost when unset. See Exposure events. |
AuthEvent carries app, ctx, collection, record: ?std.json.Value, and method (.password | .oauth2 | .magic_link | .otp | .webauthn | .custom | .refresh). FileEvent carries app, ctx, collection, record_id, and filename. LifecycleEvent carries app.
Auth lifecycle (beforeAuthSuccess)
onAuth is notify-only and fires after a session exists — perfect for logging or audit, useless for mutating state as part of the login. beforeAuthSuccess is the writable, abortable counterpart: it runs after the credentials/token are verified (and, for magic-link, the link token is consumed) but before the session JWT is issued, with a *Ctx bound to the login’s in-transaction writer.
fn claimGuestPosts(ctx: *zigbase.Ctx, ev: *zigbase.events.AuthSuccessEvent) anyerror!void {
// ev.record is the just-authenticated record; ev.record_id / ev.method are also set.
const email = ev.record.object.get("email").?.string;
// ctx.records() reuses the login transaction — this write commits WITH the session.
var patch: std.json.ObjectMap = .empty;
try patch.put(ctx.arena.a, "author", .{ .string = ev.record_id });
for (try guestPostIds(ctx, email)) |id| _ = try ctx.records().update("posts", id, .{ .object = patch });
}
// App(.{ .beforeAuthSuccess = claimGuestPosts, ... })
Guarantees:
- Transactional & atomic. The consume path runs in
BEGIN IMMEDIATE … COMMIT. The hook’sctx.records()writes commit together with the login. - Abortable, fail-closed. Return any error to block the login: the transaction rolls back (the hook’s side-writes are discarded, and a magic-link token is un-consumed so the link still works) and no session is issued. Use
ctx.fail(status, msg)for a chosen status,error.Forbidden/error.Unauthorizedfor 403/401; any other error → 500. - Bound connection. Do not call
ctx.txinside the hook (you are already in a transaction →error.NestedTransaction); usectx.records()directly.ctx.user()reflects the just-authenticated principal.
Where it fires: the unified POST /api/collections/:col/auth/:method/complete endpoint (password / otp / webauthn / oauth2 / custom), the magic-link GET …/auth/magic-link/consume link, and — since 0.10.0 — the legacy POST …/auth-with-password (tag .password) and POST …/auth-refresh (tag .refresh, in the same transaction as beforeRefresh, lifecycle phase first). This includes _superusers: the admin SPA logs in through _superusers/auth-with-password, so a hook that errors unconditionally will lock superusers out of the admin UI. That is deliberate fail-closed behavior; recovery is fixing the hook and rebuilding (hooks are comptime-compiled — operator == developer). onAuth still fires once, after issuance; on refresh it now reports .refresh (previously mislabeled .password). The surrounding lifecycle phases (register / logout / refresh / password-change) have their own before/after hooks — see below.
Auth lifecycle hooks (register / logout / refresh / password-change)
The .auth.hooks config group adds before/after hooks for the surrounding auth lifecycle, all following the beforeAuthSuccess discipline. Each handler is fn (ctx: *zigbase.Ctx, ev: *zigbase.events.AuthLifecycleEvent) anyerror!void; the event carries .collection, .record_id, .phase, and a writable .record where applicable.
zigbase.App(.{
.auth = .{
.hooks = .{
.beforeRegister = gateSignup, // validate / gate; abort blocks the account
.afterRegister = seedProfile, // post-create side effects
.beforeLogout = onBeforeLogout,
.afterLogout = onAfterLogout,
.beforeRefresh = onBeforeRefresh,
.afterRefresh = onAfterRefresh,
.beforePasswordChange = onBeforePwChange,
.afterPasswordChange = onAfterPwChange,
},
},
});
A typo’d hook name (e.g. .beforeRegsiter) or a wrong-typed handler is a compile error, never a silently-dead hook. beforeAuthSuccess and onAuth are separate keys and unchanged.
| Phase | Fires on | before: writable | before: transactional | before: abortable (fail closed) |
|---|---|---|---|---|
register | record create for an auth collection (POST /api/collections/:col/records) | ✅ (mutate the new account) | ✅ (in the create txn) | abort rolls back → no account created |
logout | POST /api/collections/:col/auth-logout | ✅ (bound writer) | — (no write txn) | abort → mapped response, cookies not cleared |
refresh | POST /api/collections/:col/auth-refresh | ✅ | ✅ | abort rolls back → no new session |
password-change | POST /api/collections/:col/confirm-password-reset and PATCH …/records/:id with a password (non-superusers must also send a verifying oldPassword) | ✅ | ✅ | abort rolls back → password unchanged (reset token un-consumed, on the reset-confirm path) |
Semantics, consistent across phases:
- Before-hooks run with a
*Ctxbound to the action’s connection (in-transaction for register/refresh/password-change).ctx.records()writes participate in the action’s transaction. Returning any error aborts the action and fails closed, mapped via theCtxerror model (ctx.fail(status, msg),error.Forbidden→403, else 500). Where a write transaction exists, the abort rolls it back. Do not callctx.tx(you are already in a transaction). - After-hooks run post-commit (post-action for logout) and are notify-only — an error is routed to the framework error backstop; it never fails the request.
registerfires only for auth collections;before_registerhas norecord_idyet (the account isn’t created), andev.recordis the writable to-be-created data (edits are persisted).ev.recordis read-only forbefore_refresh/before_password_change. It is a snapshot for reading; mutatingev.record.*is unsupported — the change is not isolated (the same value backsonAuth, the after-hook, and the HTTP response body) and is not persisted. For mutations use abefore*record hook or the writable register phase. (Onlybefore_registerexposes a writable record.)logoutkeeps a no-writer fast path: it only resolves the caller and acquires the writer when an.auth.hookshook is actually registered (an empty.auth = .{}or.auth = .{ .hooks = .{} }installs nothing).- On the
PATCH /records/:idpassword-change path,beforePasswordChangeruns inside the update transaction, after the record’s ownbeforeUpdatehooks — so a record hook can still veto the write first, and the auth hook sees the already-validated patch.
// Seed a profile row atomically with the new account.
fn seedProfile(ctx: *zigbase.Ctx, ev: *zigbase.events.AuthLifecycleEvent) anyerror!void {
if (ev.phase != .after_register) return;
var p: std.json.ObjectMap = .empty;
try p.put(ctx.arena.a, "user", .{ .string = ev.record_id });
_ = try ctx.records().create("profiles", .{ .object = p });
}
Both formerly-deferred items are now wired (0.10.0): self-service password change rides PATCH /records (see the lifecycle table above — the password-change hooks fire there too, and the endpoint enforces oldPassword for non-superusers), and beforeAuthSuccess fires on the legacy /auth-with-password / /auth-refresh endpoints (tag .password / .refresh). The ctx.auth() session verbs also gained a REST surface — see §6 Session verbs.
Two-factor authentication
Two-factor support is opt-in for embedded applications. Omission or .disabled excludes its routes and factor implementations. Select only the factors your application needs; primary authentication methods are configured independently. The shipped CLI explicitly selects both built-in second factors.
const App = zigbase.App(.{
.auth = .{
.two_factor = .{
.factors = .{ .totp, .webauthn },
.recovery_codes = true, // default when enabled
.policy = requireSecondFactor, // optional, selected at comptime
},
},
.collections = .{
.users = .{
.type = .auth,
.auth = .{ .two_factor = .optional },
},
},
});
fn requireSecondFactor(ctx: *zigbase.TwoFactorPolicyContext) !bool {
// Application-owned roles/groups remain runtime data.
const id = ctx.record.object.get("id").?.string;
const requirements = try ctx.list("security_requirements", .{
.filter = "principal = ? && required = true",
.filter_args = &.{.{ .string = id }},
.perPage = 1,
});
return requirements.items.len > 0;
}
The example assumes an application-owned security_requirements collection; the golfsim example provides its complete schema. The policy context offers findById and list read helpers. Its record and query results have request-arena lifetimes. Keep policy callbacks read-only and cheap: they run during login, completion, and session verification. Errors fail closed.
Collection .two_factor is disabled, optional, or required. Requirements combine: collection-required, application-required, or enrolled voluntarily. A hook returning false cannot waive another requirement. Applications authorize group leaders’ policy writes with their normal access rules; ZigBase does not impose a group model. A group change affects existing primary-only sessions on their next authorization check. Runtime data cannot enable a factor omitted from the binary, and a configured requirement with an unavailable subsystem fails closed.
No ordinary session is issued while another factor is pending. beforeAuthSuccess and onAuth run at final session issuance, after factor verification. Custom primary routes call zigbase.auth.beginAuthentication(request, writer, collection, record_id) before issueSession, returning the non-null pending response. The existing issuer independently enforces the requirement.
TOTP secrets use a dedicated authenticated-encryption key domain derived from the application JWT secret. Preserve that secret when moving an installation. Recovery codes are random, stored as digests, and consumed atomically. WebAuthn uses the configured RP/origin and separates second-factor credentials from primary passkeys. See the wire contract for enrollment, management, recovery, expiry, and client handling.
Auth methods overview
ZigBase ships with a pluggable auth-method system that lets every auth collection enable built-in or custom login methods via config, with no route code. The system is built around a two-phase contract:
- Initiate — challenge delivery (email a link/code, return WebAuthn options, or a no-op for API-key flows).
- Complete — proof verification. On success the method returns a
Resolution.record(record_id)and the framework mints the session via the one seam (issueSession+emitAuth), firingonAuth. The method never mints sessions itself.
There are three tiers:
| Tier | How | When |
|---|---|---|
| Built-in | magic_link, otp, password, webauthn in .auth.methods | Most apps — no Zig code needed |
| Custom plugin | Register a TYPE in the app-level .auth.methods; reference by slug in the collection’s .auth.methods | Non-standard verification (corp SSO, API keys, etc.) |
| Escape hatch | Custom routes + ctx.issueSession / zigbase.auth API (see below) | Exotic flows where the two-phase contract doesn’t fit |
Backward compatibility: an auth collection with no .auth.methods config defaults to password — existing deployments are unaffected.
What lives under
.auth(comptime) vs env. The comptime.authconfig group holds the structure/behavior knobs:.hooks,.methods,.captcha, and.session(.store/.gc_cron/.rotation_grace_s). The deploy-varying runtime auth knobs stay env-configured and are deliberately not under.auth: token TTLs (ZIGBASE_AUTH_TOKEN_TTL), the server-side OAuthstatestore (ZIGBASE_OAUTH_STATE_*), cookie security (ZIGBASE_COOKIE_SECURE/--insecure-cookies), and rate-limiting (ZIGBASE_RATE_LIMIT_*). See the env table in the README.
Enabling auth methods per collection (.auth.methods)
Each auth collection in .collections can declare a .auth.methods struct enabling one or more built-in methods:
| Method | Key | Options | Notes |
|---|---|---|---|
| Password | password | (none beyond rate_limit) | Built-in default when no .methods specified |
| Magic link | magic_link | ttl_s: i64 (default 900), auto_create: bool (default false) | when auto_create=true, provisions an account for unknown identities; initiate emails link; complete verifies+consumes |
| OTP | otp | length: u8 (default 6), ttl_s: i64 (default 300), auto_create: bool (default false) | when auto_create=true, provisions an account for unknown identities; initiate emails code; complete verifies code |
| WebAuthn | webauthn | rp_id: []const u8, rp_name: []const u8, origin: []const u8, credentials_collection: []const u8, require_uv: bool (default false) | all four string fields required; require_uv rejects assertions without user-verification (biometrics/PIN) |
| OAuth2 | (see below) | gated by .auth.oauth2.enabled + .auth.oauth2.providers | 5th built-in; uses the contract but is NOT listed in .auth.methods |
Per-collection auth options (set directly on .auth, independent of which methods are enabled):
require_verified: bool(defaultfalse) — whentrue, any login attempt is rejected with403if the auth record’sverifiedfield isfalse. This gate applies to all auth methods, including WebAuthn passkeys and OAuth2 accounts created from providers that did not confirm the email address (those are createdverified=false). Enable it only after ensuring existing users have verified accounts, or after setting up a verification flow.
auto_create(onmagic_linkandotp, defaultfalse) — whentrue,initiateautomatically provisions a passwordless auth record (verified = false) for email addresses not yet in the collection, then sends the link/code as usual. Enables “sign up or sign in” with a single step. Note: auto-created accounts haveverified = false; if the collection also setsrequire_verified = true, those accounts cannot log in until verified. Consider whether to pair these settings. Works best with single-fieldidentityFields(the default
Each method accepts a rate_limit field:
.default— uses the global env-var rate-limiter (ZIGBASE_RATE_LIMIT_MAX/ZIGBASE_RATE_LIMIT_WINDOW)..off— disables rate-limiting for this method (logged at startup like a@publicrule)..{ .custom = .{ .max = 5, .window_s = 60 } }— per-method override: the configuredmax/window_sare honored against a dedicated bucket scoped by collection + method, so different methods (and the same method on different collections) never share a budget. The bucket subject (client IP, else the submitted identity) matches the global limiter. A custom limit applies even when the global limiter is disabled (ZIGBASE_RATE_LIMIT_MAX=0).
Example — two collections, each with different methods:
.collections = .{
.members = .{ .type = .auth, .auth = .{
.methods = .{ .magic_link = .{ .ttl_s = 900 } },
} },
.staff = .{ .type = .auth, .auth = .{
.methods = .{ .password = .{}, .webauthn = .{
.rp_id = "app.example.com",
.rp_name = "My App",
.origin = "https://app.example.com",
.credentials_collection = "webauthnCredentials",
} },
} },
},
This yields the following endpoints (no route code):
POST /api/collections/members/auth/magic_link/initiate
POST /api/collections/members/auth/magic_link/complete
POST /api/collections/staff/auth-with-password (legacy alias)
POST /api/collections/staff/auth/webauthn/initiate
POST /api/collections/staff/auth/webauthn/complete
POST /api/collections/staff/auth/webauthn/register/begin (authed)
POST /api/collections/staff/auth/webauthn/register/finish (authed)
The dispatch enforces enablement: a disabled or unknown method slug returns 404.
Typed TS client. zig build gen-client emits a typed client.auth.<collection>.<method> surface for every enabled non-password method (camelCased, e.g. client.auth.members.magicLink). For the three built-in methods the generated initiate/complete carry precise input/result types (the server contracts are fixed):
| Method | initiate(input) | initiate result | complete(input) | complete result |
|---|---|---|---|---|
magic_link | { identity } | void (204) | { token } | { token } |
otp | { identity } | void (204) | { identity, code } | { token } |
webauthn | { identity? } | { challenge, rpId, ceremonyId, timeout } | { ceremonyId, credentialId, authenticatorData, clientDataJSON, signature } | { token } |
Every built-in complete resolves to { token } | PendingAuthentication (AuthMethodResult). Only the authenticated result sets session cookies (zb_auth/zb_csrf). Custom methods can be typed too: a bare-string slug (.custom = .{"corp-sso"}) stays on the untyped Record<string, unknown> / unknown stubs, while the struct form declares comptime I/O types: .{ .slug = "corp-sso", .Initiate = .{ .Input = …, .Output = … }, .Complete = .{ .Input = …, .Output = … } }. Generated custom completion types also include PendingAuthentication, since they share the same policy gate. A void Input omits the input argument. See typed auth methods.
Custom AuthMethod plugin (app-level .auth.methods)
For verification logic that the built-ins don’t cover, implement an AuthMethod plugin type and register it at the app level, under the .auth config group:
// App-level registration of custom method TYPES:
.auth = .{ .methods = .{ WebAuthnMethod, CorpSsoMethod } },
// Then reference by slug in a collection's .methods:
.staff = .{ .type = .auth, .auth = .{
.methods = .{ .password = .{}, .custom = &.{"corp-sso"} },
} },
Selecting the built-in set. The bare-tuple form above always keeps all five built-ins (.password, .magic_link, .otp, .webauthn, .oauth2) — non-breaking, unchanged from before. To drop built-ins you don’t use — WebAuthn’s CBOR/COSE stack is the big one at ~3.2k LOC — switch to the named form and list exactly the built-ins you want:
.auth = .{ .methods = .{
.builtins = .{ .password, .otp }, // exactly these two; the other three are never
// analyzed, so their code (and routes) is
// absent from the binary
.custom = .{ CorpSsoMethod },
} },
Omitting .builtins keeps all five (equivalent to the bare-tuple form); .builtins = .{} drops every built-in and leaves only .custom. A deselected built-in’s method-specific route group (webauthn/magic_link/oauth2) is dropped too — the flat core auth routes (auth-with-password, auth-refresh, auth-logout, verification, password reset) are always present regardless of the built-in set, since they don’t go through the method registry.
The contract — every plugin type must implement three functions:
// Required on every AuthMethod plugin type:
pub fn create(gpa: std.mem.Allocator, io: std.Io, cfg: zigbase.Config) !Self;
pub fn method(self: *Self) zigbase.AuthMethod; // returns the vtable view
pub fn deinit(self: *Self) void;
zigbase.AuthMethod vtable:
pub const AuthMethod = struct {
slug: []const u8, // URL slug → /auth/<slug>/initiate | /complete
ctx: *anyopaque,
vtable: *const VTable,
pub const VTable = struct {
initiate: *const fn (ctx: *anyopaque, ac: *AuthCtx) anyerror!InitiateResult,
complete: *const fn (ctx: *anyopaque, ac: *AuthCtx) anyerror!Resolution,
};
};
AuthCtx — what the framework passes to each phase:
pub const AuthCtx = struct {
app: *Runtime,
ctx: *http.RequestCtx, // body / query / cookies / remote_ip
collection: schema.Collection, // the :col auth collection (validated .type == .auth)
// No ambient conn — each method acquires its own via writer()/reader().
// Connection RAII handles:
pub fn writer(ac: *AuthCtx) WriterData; // acquires pool writer (mutex); call deinit()
pub fn reader(ac: *AuthCtx) !ReaderData; // checks out a pooled reader; call deinit()
// Blessed helpers (each takes the conn you acquired):
pub fn rateLimit(ac: *AuthCtx, scope: []const u8, ident: []const u8) !?http.Response;
pub fn mintLinkToken(ac: *AuthCtx, conn: *db.Db, record_id: []const u8, ttl_s: i64) ![]const u8;
pub fn verifyLinkToken(ac: *AuthCtx, conn: *db.Db, token: []const u8) !?Claims;
pub fn consumeLinkToken(ac: *AuthCtx, conn: *db.Db, claims: Claims) !void;
pub fn deliverMail(ac: *AuthCtx, to: []const u8, subject: []const u8, body: []const u8) !void;
pub fn findByIdentity(ac: *AuthCtx, conn: *db.Db, identity: []const u8) !?[]const u8;
};
pub const InitiateResult = struct { status: u16 = 200, body: ?[]const u8 = null };
pub const Resolution = union(enum) {
record: []const u8, // record id → framework issues session + fires onAuth
fail: struct { status: u16, message: []const u8 },
};
A plugin type missing create/method/deinit is a compile error. Built-in methods (password, magic_link, otp, webauthn) implement the exact same contract — no privileged private path.
WebAuthn — login + passkey registration
Login flow (via the two-phase contract):
initiate: POST body{ "identity": "user@example.com" }(optional for discoverable credentials). Returns{ "challenge": "...", "rpId": "...", "ceremonyId": "..." }.- Browser calls
navigator.credentials.get()with the options. complete: POST body{ "ceremonyId": "...", "credentialId": "...", "authenticatorData": "...", "clientDataJSON": "...", "signature": "..." }. On success, the framework mints the session and firesonAuth(.webauthn).
Passkey registration (requires an existing session):
POST /api/collections/:col/auth/webauthn/register/begin→ returns{ "challenge": "...", "rpId": "...", "rpName": "...", "ceremonyId": "..." }.- Browser calls
navigator.credentials.create()with those options. POST /api/collections/:col/auth/webauthn/register/finishwith{ "ceremonyId": "...", "id": "...", "rawId": "...", "response": { "clientDataJSON": "...", "attestationObject": "..." } }. Stores the credential in_webauthnCredentialsbound to the authenticated user.
Notes:
- Supports ES256 (P-256, COSE -7) and Ed25519 (COSE -8) public keys; the COSE key curve is validated against the algorithm (ES256→P-256, EdDSA→Ed25519).
- Attestation format:
fmt:"none"(v1; other formats are future work). signCountis tracked; a count regression (possible credential clone) fails authentication closed.- Each credential is bound to the collection it was registered on; presenting it on a different collection returns
401. require_uv: truerejects assertions where the authenticator did not perform user verification (biometrics or PIN). Defaultfalse.- Requires
webauthn.credentials_collectionto name a collection that stores credentials.
ChallengeStore
ChallengeStore is a TTL’d, GC’d server-side store backed by _authChallenges. It is used internally by magic_link (the link token), otp (the emailed code), and webauthn (ceremony challenges). Custom plugins may access it via AuthCtx.challengeStore(), which returns a handle with:
store(id, data, ttl_s)— store a challenge underidwith the given TTL.consume(id) ![]const u8— retrieve and delete in one atomic step (single-use).
OAuth2 providers (.auth.oauth2)
Configure OAuth2 sign-in for an auth collection by adding .auth.oauth2 to its .auth block:
.users = .{ .type = .auth, .fields = .{}, .auth = .{
.oauth2 = .{
.enabled = true,
.providers = .{
// Preset provider — endpoints come from ZigBase's built-in table.
.{ .name = "google", .redirectUrls = .{"https://app.example/oauth/callback"} },
// Generic provider — all three endpoint URLs are required (all https://).
.{ .name = "myprovider",
.authURL = "https://myprovider.example/oauth/authorize",
.tokenURL = "https://myprovider.example/oauth/token",
.userinfoURL = "https://myprovider.example/api/user",
.redirectUrls = .{"https://app.example/oauth/callback"} },
},
},
} },
Built-in presets (endpoint URLs supplied automatically): google, github, microsoft, discord. A generic provider requires authURL, tokenURL, and userinfoURL (all must be https://).
Per-provider fields (all optional except .name):
| Field | Default | Notes |
|---|---|---|
name | — | Required. Valid identifier; becomes the slug and is uppercased into the env var name. |
enabled | true | Set false to temporarily disable without removing the declaration. |
redirectUrls | &.{} | Tuple of allowed OAuth2 redirect URLs. |
clientId | "" | Prefer env (see below); literal is accepted but bakes the value into the binary. |
clientSecret | "" | Do not set in the literal. Use env — the literal embeds the secret in the binary. |
authURL | null | Generic provider only. |
tokenURL | null | Generic provider only. |
userinfoURL | null | Generic provider only. |
discoveryURL | null | Generic provider only; alternative to authURL/tokenURL/userinfoURL. See below. |
scopes | null | Override default scopes for a preset, or supply scopes for a generic provider (default openid email profile when using discoveryURL). |
Generic OIDC discovery (discoveryURL). Any OIDC-compliant IdP (Auth0, Okta, Keycloak, Entra custom tenants, Zitadel, …) is one line — point at its discovery document instead of hand-copying three endpoint URLs:
.providers = .{
.{ .name = "okta",
.discoveryURL = "https://acme.okta.com/.well-known/openid-configuration",
.redirectUrls = .{"https://app.acme.com/oauth/callback"} },
},
discoveryURL is mutually exclusive with explicit authURL/tokenURL/userinfoURL (compile error), https-only, and rejected for the built-in preset names. Resolution happens once, at startup: the framework fetches the document over the same TLS transport every OAuth call uses, requires all three of authorization_endpoint/token_endpoint/ userinfo_endpoint (https-only) plus an issuer that prefixes the discovery URL, and refuses to start on any failure (a half-configured IdP must never silently disable login). The resolved endpoints are persisted exactly like literal generic endpoints, so the usual provisioning caveat applies (re-resolution needs a migration or an admin-API PATCH — there is no per-request discovery dependency). Scopes default to openid email profile (override with .scopes); the claim mapping is the fixed OIDC standard set; PKCE is already unconditional. id_token/JWKS validation remains out of scope — identity is verified via the userinfo endpoint over TLS, as for every provider. (Provider fields keep their documented camelCase style — discoveryURL matches authURL/tokenURL.)
Runtime secrets via environment variables
The clientId and clientSecret are NOT read from the comptime literal at runtime (a comptime literal cannot hold a secret safely). Instead, set environment variables at provisioning time:
ZIGBASE_OAUTH_<UPPER(NAME)>_CLIENT_ID=<your-client-id>
ZIGBASE_OAUTH_<UPPER(NAME)>_CLIENT_SECRET=<your-client-secret>
For example, for name = "google":
ZIGBASE_OAUTH_GOOGLE_CLIENT_ID=123456789-abc.apps.googleusercontent.com
ZIGBASE_OAUTH_GOOGLE_CLIENT_SECRET=GOCSPX-...
The framework uppercases the provider name character-by-character to form the env var key. The clientSecret is encrypted (AES-256-GCM, HKDF key derived from the JWT secret) before being persisted in the database — the plaintext never reaches disk.
Provisioning caveat (applied on first creation only)
The env-secret injection and collection creation happen at startup, but only on the first boot (when the collection does not yet exist in the database). On subsequent boots the live database options are preserved — the env vars are ignored for an already-provisioned collection. This means:
- Rotating a secret: update the value via the admin API (PATCH
/api/collections/:col) or the admin UI. The admin path re-encrypts under the new plaintext. - Changing a redirect URL or adding a provider: update the comptime literal AND use an explicit
.migrationsentry to apply the change (provisioning is additive-only; it does not re-apply options on existing collections).
CI/e2e caveat: this feature cannot be exercised in automated CI without real provider credentials. All verification is at the unit level (comptime lowering + encrypt-on-inject with a stub env getter and a fake app secret).
OAuth2 — contract method
OAuth2 (Google, GitHub, etc.) is the fifth built-in AuthMethod (slug oauth2). It is gated by the existing .auth.oauth2.enabled + .auth.oauth2.providers config, not by .auth.methods. When OAuth2 is enabled for a collection the framework auto-mounts three endpoints:
| Method | Path | Description |
|---|---|---|
| GET | /api/collections/:col/auth/oauth2/providers | Discovery — list enabled providers (name, authURL, clientId, scopes) as {"items":[…]}. Secrets never returned. |
| POST | /api/collections/:col/auth/oauth2/initiate | Return provider metadata so the client can drive the authorization redirect. |
| POST | /api/collections/:col/auth/oauth2/complete | Exchange the authorization code for a session. |
oauth2 initiate — body { "provider": "<name>" }:
// response (200)
{
"authURL": "https://accounts.google.com/o/oauth2/v2/auth?...",
"clientId": "my-client-id.apps.googleusercontent.com",
"scopes": ["openid", "email", "profile"],
"state": "<server-issued-state>" // always present (ZIGBASE_OAUTH_STATE_SERVER defaults to true)
}
oauth2 complete — body:
{
"provider": "google",
"code": "<authorization-code>",
"codeVerifier": "<pkce-verifier>",
"redirectUrl": "https://app.example.com/callback",
"state": "<server-issued-state>" // required by default (ZIGBASE_OAUTH_STATE_SERVER defaults to true)
}
// response (200) — sets zb_auth and zb_csrf cookies
{ "token": "<jwt>" }
On success the framework fires onAuth(.oauth2) through the shared session seam — identical to every other method. All paths share one implementation and enforce the same security: single-use TTL’d CSRF state consumed before the code exchange, PKCE required, redirect allow-list, and https-only provider URLs.
Connection model: each auth method manages its own DB connection for the duration of its call.
completeacquires a writer, consumes the CSRFstate(single-use, before the provider exchange), then releases the writer before the outbound provider HTTP round-trip. A new writer acquisition completes the record lookup and session mint. This meanscompletedoes not hold the writer across blocking I/O — high OAuth2 concurrency does not stall writes. Passwordcompleteuses a reader (argon2 verification is read-only); the writer is acquired only at session-mint time.
Tier 3: escape hatch for exotic flows
For flows where the built-in two-phase AuthMethod contract doesn’t fit, ZigBase exposes a zigbase.auth helper surface so you can build custom login flows while staying in the same session-issuance seam that built-in logins use.
The seam guarantee: every login — including custom flows via ctx.issueSession — fires your onAuth handler. There is no way to mint a session that bypasses it. Built-in flows (password, OAuth2) and custom flows both go through the same issueSession+emitAuth path, so the onAuth hook is the single, reliable chokepoint for cross-cutting session logic (audit logging, account-state checks, etc.).
zigbase.auth API reference
// src/auth_helpers.zig (imported as zigbase.auth.*)
pub const Issued = struct { token: []const u8, cookies: [2]http.Cookie };
pub const LinkToken = struct { token: []const u8 };
// Issue a full session (sets cookies, fires onAuth).
// Prefer ctx.issueSession (below) from a route — it acquires the writer for you.
pub fn issueSession(
ctx: *http.RequestCtx, conn: *db.Db,
collection: []const u8, record_id: []const u8,
) !Issued
// Single-use magic-link token helpers.
// opts.payload (default "") binds a small opaque, signed, tamper-proof string into the
// token — returned by verifyLinkToken as claims.pl. Use it to carry e.g. a post-login
// redirect target in the one token instead of an unsigned &next= URL param. Signed, not
// encrypted: readable-but-tamper-proof; keep it small (base64url'd into the JWT payload).
pub const MintOptions = struct { payload: []const u8 = "" };
pub fn mintLinkToken(
ctx: *http.RequestCtx, conn: *db.Db,
collection: []const u8, record_id: []const u8, ttl_s: i64, opts: MintOptions,
) !LinkToken
// Returns null when the token is expired, wrong collection, or has a bad signature.
// claims.id is the record id stored in the token; claims.pl is the bound payload ("" if none).
pub fn verifyLinkToken(
ctx: *http.RequestCtx, conn: *db.Db,
collection: []const u8, token: []const u8,
) !?jwt.Claims
// Marks the token consumed. Returns error.AlreadyConsumed on replay.
pub fn consumeLinkToken(conn: *db.Db, claims: jwt.Claims) !void
// Clear the session cookies (logout), mirroring issueSession. Returns arena-owned
// zb_auth/zb_csrf cookies built from the framework's own cookie policy, so they match
// the built-in logout exactly. Equivalent to ctx.auth().clearSession() (below).
pub fn clearSession(ctx: *Ctx) ![]const http.Cookie
// Session revocation (#99). Equivalent to the ctx.auth() verbs below.
pub fn revokeAllSessions(ctx: *Ctx) !void // "log out everywhere" (bump epoch)
pub fn refresh(ctx: *Ctx) !Issued // re-mint, same epoch (sliding refresh)
pub fn rotate(ctx: *Ctx) !Issued // bump epoch + re-mint (kill other sessions)
// Send auth email via the configured mailer (SMTP or log in dev).
pub fn deliverAuthMail(
app: *App, alloc: std.mem.Allocator,
to: []const u8, subject: []const u8, body: []const u8,
) !void
// Rate-limit helper — returns a ready-to-return 429 Response when the caller
// is over the limit, null otherwise. scope identifies the endpoint; ident is the
// key (e.g. email or raw request body).
pub fn rateLimit(
ctx: *http.RequestCtx, scope: []const u8, ident: []const u8,
) !?http.Response
ctx.issueSession — the ergonomic route helper
From inside a route handler, use ctx.issueSession instead of calling zigbase.auth.issueSession directly. It acquires and releases the DB writer for you (it is only valid in a route — ctx.request must be non-null):
// src/ctx.zig
pub fn issueSession(
ctx: *Ctx, collection: []const u8, record_id: []const u8,
) !Issued // acquires+releases the writer, fires onAuth(.custom)
Warning:
ctx.issueSessionacquires the pool writer for you internally. Do NOT call it while you are already holding the writer — insidectx.tx, from abefore*hook, or with a bound connection — it would deadlock on the single non-reentrant writer. In those cases callzigbase.auth.issueSession(ctx.request.?, conn, collection, record_id)directly with the connection you already hold.
Typical usage in a confirm handler that does NOT separately hold the writer:
fn myConfirm(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
// ... verify your token, resolve the record id ...
const issued = try ctx.issueSession("members", record_id);
return .{ .status = 200, .body = "{\"ok\":true}", .cookies = &issued.cookies };
}
When your handler already holds the writer (e.g. for verifyLinkToken and consumeLinkToken), reuse it:
fn myConfirm(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
const w = ctx.app.pool.acquireWriter();
defer ctx.app.pool.releaseWriter();
// ... verifyLinkToken, consumeLinkToken using w ...
// Use zigbase.auth.issueSession, NOT ctx.issueSession, to avoid double-acquiring the writer.
const issued = try zigbase.auth.issueSession(ctx.request.?, w, "members", record_id);
return .{ .status = 200, .body = "{\"ok\":true}", .cookies = &issued.cookies };
}
issued.cookies is a [2]http.Cookie containing zb_auth (httpOnly) and zb_csrf (readable). Pass &issued.cookies as the cookies field on the http.Response and the framework writes both Set-Cookie headers.
For a complete, copy-pasteable example (two-route magic-link flow with rate limiting, enumeration safety, and single-use token replay protection) see recipes.md → magic-link login.
ctx.auth() — session management
The ctx.auth() namespace is the session-management surface.
clearSession mirrors issueSession for logout: it returns the cleared session cookies built from the framework’s own cookie policy (the same one the built-in authLogout uses), arena-owned so they slot straight into Response.cookies. A logout handler is one line:
fn logout(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
return .{ .status = 204, .body = "", .cookies = try ctx.auth().clearSession() };
}
zigbase.auth.clearSession(ctx) is the equivalent free-function form.
Revoking sessions (#99)
Sessions are stateless JWTs, but they are still revocable via a per-auth-record token epoch. Each .auth token embeds the record’s token_epoch; verification rejects a token whose epoch no longer matches the record’s current value (fail closed). This is the default model (App(.{ .auth = .{ .session = .{ .store = .epoch } } })) and costs no extra query on either the verify hot path or login — the epoch is folded into the single tokenKey SELECT each already performs (one extra column, not an extra round-trip).
| verb | effect |
|---|---|
ctx.auth().revokeAllSessions() | bump the epoch → every outstanding token for the principal stops verifying (“log out everywhere”). Pair with clearSession to also drop this browser’s cookie. |
ctx.auth().refresh() | re-mint a token (new exp, same epoch) — a sliding refresh that leaves other sessions valid. Returns Issued (JWT + cookies). |
ctx.auth().rotate() | bump the epoch then re-mint — keep this session, kill every other. Returns Issued. |
fn changePassword(ctx: *zigbase.Ctx) anyerror!zigbase.http.Response {
// ... verify + update the password ...
const issued = try ctx.auth().rotate(); // invalidate all old sessions, keep this one
return .{ .status = 200, .body = body, .cookies = try ctx.arena.a.dupe(zigbase.http.Cookie, &issued.cookies) };
}
Free-function forms: zigbase.auth.revokeAllSessions(ctx) / refresh(ctx) / rotate(ctx).
Back-compat is total: tokens minted before the epoch existed (no claim) and freshly created records (NULL column) both read as epoch 0, so all existing valid tokens keep working.
Per-device sessions (Variant B, .auth.session.store = .table). Opting into the table model maintains a server-side _sessions row per session, enabling a full per-device UI:
| verb (table mode only) | effect |
|---|---|
ctx.auth().listActiveSessions() | the current principal’s active sessions (id, created, last_seen, user_agent, ip, is_current), newest first. |
ctx.auth().revoke(sessionId) | “log out THIS device” — delete one session row. Authorized: only the owning user (or a superuser) may revoke a given session; a non-owner gets error.NotFound (indistinguishable from an absent id, so revoke can’t probe other users’ session ids). |
Since 0.10.0 the table-mode verbs also have a canonical REST surface (the SDK’s listSessions/revokeSession/revokeAllSessions ride it):
| Route | Mode | Behavior |
|---|---|---|
GET /api/collections/:col/auth/sessions | .table only | 200 {"items":[{id,created,last_seen,user_agent,ip,is_current},…]} newest-first |
DELETE /api/collections/:col/auth/sessions/:sid | .table only | 204; a non-owned or absent sid is an indistinguishable 404 |
DELETE /api/collections/:col/auth/sessions | both modes | “log out everywhere”: epoch bump (+ row wipe in table mode) → 204 with cleared cookies |
In .epoch mode the two per-device routes return 404 (feature not enabled). :col must match the caller’s authenticated collection (else 401). The sessions segment under /auth/ is reserved — a custom auth-method slug named sessions is a compile error.
In table mode, login/refresh/oauth issuance records a session row and embeds its id in the token (sid claim); verify checks the row exists and is unexpired (one extra indexed read per authenticated request — the accepted cost of this mode). Logout and revoke delete the row; revokeAllSessions clears all of the principal’s rows (and bumps the epoch). refresh/ rotate rotate the current device’s row (no accumulation), carrying the original created (session start) forward and stamping last_seen = now. By design last_seen reflects the last token refresh, not every request — verify never writes the session table, so an authenticated request stays one read (no per-request write amplification on the single writer).
Rotation grace (.rotation_grace_s, HTTP auth-refresh only). When the POST /api/collections/:col/auth-refresh endpoint rotates a session, the predecessor row is not deleted outright — its expires is clamped to now + rotation_grace_s so requests already in flight with the old cookie don’t 403 through the rotation window (an SPA whose guard refreshes on navigation races its own data fetches). The graced rows are reaped by the GC sweep below. The default is 30 seconds; set .auth.session = .{ .store = .table, .rotation_grace_s = 0 } to restore an immediate delete (validated at comptime: negative is a @compileError, and the key has no effect — also a @compileError — outside .table mode). The in-process ctx.auth().refresh() / ctx.auth().rotate() helpers do not use the grace path; they delete the old row immediately.
Expired-session GC. Enabling .table auto-installs a framework-internal recurring job (_session_gc) that deletes expired _sessions rows (those whose expires has passed; NULL = never expires) in bounded batches on the writer — no opt-in needed, and nothing is installed in .epoch mode (no job, no timer). The default cadence is hourly ("0 * * * *"); override it with .auth.session.gc_cron (UTC, minute-granularity cron syntax), e.g. App(.{ .auth = .{ .session = .{ .store = .table, .gc_cron = "*/30 * * * *" } } }) for every 30 minutes. (Expired rows are inert before collection — verify already rejects them — so GC is housekeeping, not a correctness gate.)
Switching an existing app to
.tableis not retroactive. Tokens already minted under.epochcarry nosid, so they skip the per-device session check and stay valid until they expire —revoke(sessionId)can’t kill them (no row exists). UserevokeAllSessions()(an epoch bump) once after enabling.tableto invalidate every pre-existing session.
.epoch is the default and is unchanged: it issues ZERO session-table queries, and enabling .table does not alter the .epoch-mode token shape — the sid claim is simply omitted when absent. In .epoch mode listActiveSessions/revoke return error.SessionStoreNotEnabled (there is no per-session inventory); revokeAllSessions/refresh/rotate work in both modes. See the session-management design spec.
7. Scheduled jobs (.cron + .jobs)
.pools = .{ .jobs = 2 }, // scheduler worker-pool size; defaults to 2 when unset
.cron = .{
.{ .name = "heartbeat", .schedule = zigbase.schedule.Schedule{ .interval = .hourly }, .handler = heartbeat },
},
.pools = .{ .jobs = N }is the ONE lever for the scheduler worker-pool size..jobsis purely the queue job-kind handler registry (kind = handlerFnbindings) — see §7b Background jobs & queues. The pre-0.10.jobs = .{ .pool_size = N }spelling was removed; it is now a compile error naming.pools = .{ .jobs = N }.
Each cron spec needs .name, .schedule, and .handler (missing/wrong-typed → compile error). The three schedule modes (zigbase.schedule.Schedule):
.{ .cron = "0 3 * * *" } // 5-field cron, UTC, JAN/MON names ok
.{ .interval = .hourly } // also .daily / .weekly
.{ .interval = .{ .minutes = 15 } } // every N minutes
.reactive // handler decides its own next fire
Handler signatures depend on the mode. Cron and interval jobs use a ctx-first handler (zigbase.JobEvent is re-exported at the top level for convenience; it is the same type as zigbase.events.JobEvent):
fn (ctx: *zigbase.Ctx, ev: *zigbase.events.JobEvent) anyerror!void
fn heartbeat(ctx: *zigbase.Ctx, ev: *zigbase.events.JobEvent) anyerror!void {
_ = ctx;
std.log.info("blog heartbeat job '{s}' ran", .{ev.name});
}
A .reactive job’s handler instead returns its next schedule:
fn (ctx: *zigbase.Ctx, ev: *zigbase.events.JobEvent) anyerror!zigbase.schedule.Reactive
It returns either .{ .after = <Interval> } (re-run after that interval, e.g. .{ .after = .{ .minutes = 5 } } or .{ .after = .daily }) or .stop (retire the job). JobEvent carries app and name. Touch the database from a job through ctx.records() (see DB access from a route); for raw SQL on a migration-owned table acquire the pooled writer via ctx.app.pool.acquireWriter() / ctx.app.pool.releaseWriter().
Caveats
Shared cron/interval jobs across replicas
Opt a job into database-backed coordination at compile time:
.cron = .{
.{
.name = "billing-sweep-v1", // durable identity shared by every replica
.schedule = zigbase.schedule.Schedule{ .interval = .{ .minutes = 5 } },
.distributed = .{ .lease_seconds = 120 },
.handler = billingSweep,
},
},
Only opted-in wrappers reference the coordination runtime. Jobs without .distributed retain their per-process behavior, and reactive jobs cannot opt in. The typed configuration is zigbase.DistributedJobConfig.
All replicas sharing a database and job name observe one persisted cadence. A short transaction claims due work with an owner token, generation and expiry; handlers run outside that transaction. PostgreSQL uses row locks and its server clock after lock acquisition, while SQLite uses its immediate writer transaction. SQLite coordination applies to processes sharing the same local database file, not independent database copies or unsupported network filesystems.
A new schedule first fires at its next scheduled time, never immediately on boot. Restarts retain that time. Missed cron occurrences coalesce into one pending run; after success the next cron time is calculated from completion. Intervals likewise advance from completion. Errors persist capped exponential retry backoff. The JobEvent.scheduled_at value remains stable across retries and crash recovery; combine it with JobEvent.name for an application idempotency key. generation identifies the current claim attempt; both fields are null for ordinary jobs.
Expired leases can be recovered by another replica. An expired owner can still complete if no replacement has claimed that occurrence. Completion from a replaced owner cannot advance the schedule, reset failures, or clear the newer lease. This fences dispatcher state, not arbitrary application writes or external effects: a crashed/paused handler may execute more than once, and an overrun can overlap its replacement. Make handlers idempotent and set lease_seconds above worst-case handler duration. There is no automatic lease heartbeat or cancellation of application code.
Distributed job names must be unique, nonempty, at most 200 bytes, and not start with _. Idle replicas probe readiness using readers every 15 seconds, acquiring the writer only when registration or claiming may be needed. Due jobs and expired leases can therefore wait up to roughly 15 seconds plus worker availability. Stopped schedules retire locally rather than polling forever. Claimed work still rechecks readiness under the database row lock.
Replicas must deploy identical schedule definitions for a name; a mismatch fails closed with DistributedScheduleMismatch. To change a definition, drain old replicas and explicitly remove that job’s _scheduler_jobs row in an application migration before restarting, or use a new versioned name after retiring the old job. Removing a declaration does not delete its saved state. Lease duration may differ between replicas and can be tuned during a rolling deployment; each claim stores its own expiry. All declared minute intervals (including per-process jobs) must be positive; reactive handlers can still request an immediate next tick explicitly.
The scheduler is intentionally simple (see ../KNOWN_LIMITATIONS.md):
- Per-process by default — use
.distributedfor coordinated cron/interval jobs. - UTC — all cron/interval evaluation is in UTC.
- Cron — UTC, minute-granularity; supports the full standard grammar:
*,a,a,b,c,a-b,*/n,a-b/n(a step over a range, e.g.0-23/2for every other hour), and case-insensitive 3-letter month (JAN..DEC) / day-of-week (SUN..SAT) names (steps stay numeric). Day-of-week7is a Sunday alias —0and7both mean Sunday. A.cronstring is validated at compile time: a malformed expression (wrong field count, a full day name likeMONDAY, a trailing or doubled space) or an out-of-range value (minute > 59, hour > 23, month > 12, day-of-month0, day-of-week > 7) is a build error, not a job that silently never fires. The same validation covers.auth.session.gc_cron. - Day-of-month and day-of-week are ANDed (not Vixie cron’s OR semantics).
- Minute granularity — the smallest schedule resolution is one minute.
- Interval drift — interval schedules measure the next fire from the previous run’s completion, not a fixed wall clock, so long-running jobs drift.
- Single-flight — a job never overlaps itself.
The framework may also register internal jobs of its own alongside your .cron table. Declaring a TTL collection (.ttl_field, see §8) registers an interval job named _ttl_gc that reaps expired rows on the .ttl_gc_interval cadence (default every 5 minutes); it runs on the same worker pool and starts the scheduler even when you have no .cron configured.
Ad-hoc background work: app.submit
To offload one-off background work from a route or hook onto the worker pool:
try ctx.app.submit("reindex", reindexTask);
// reindexTask: fn (ctx: *zigbase.Ctx, ev: *zigbase.events.JobEvent) anyerror!void
submit copies the name before returning, so the caller retains ownership of the input slice. The task borrows ev.name from the queue for that invocation; it must not free or retain the slice after returning.
submit returns error.SchedulerUnavailable when the server is not running (unit tests / CLI); while serving it is always available — no scheduler configuration required (0.10.0).
Submitted tasks run on the bounded background worker pool (shared with memory-queue jobs). At shutdown the pool is drained and joined, so a task submitted before shutdown completes. When the pool’s ring is full,
submitreturnserror.QueueFullrather than blocking — sustained high-volume background work belongs on a durable queue.
7b. Background jobs & queues (.queues / .workers / .jobs)
.cron/app.submit (§7) are for scheduled and fire-and-forget work. For enqueued background jobs with retries, priorities, and optional durability, ZigBase ships a generic multi-queue / worker / job engine. The same engine backs the built-in mail and webhook job kinds.
Three config keys lower at comptime:
App(.{
// Named queues. `.default` (memory/normal) is ALWAYS synthesized if you don't declare it.
.queues = .{
.emails = .{
.backend = .durable, // .memory (default) | .durable
.priority = .high, // .high | .normal (default) | .low
.retry = .{ .max_attempts = 5, .backoff = .exponential, .base_ms = 1000, .max_ms = 300_000, .jitter = true },
// Durable-only reliability knobs (defaults shown). Set the visibility timeout
// ABOVE this queue's longest job runtime so the reclaim sweep never
// double-dispatches a still-running job:
.visibility_timeout_s = 300, // reclaim a claim older than this (>= 1)
.done_ttl_s = 604_800, // GC done/failed rows older than this (7d)
.rate = .{ .per_second = 14 }, // durable only: per-queue send-rate ceiling (e.g. SES default 14/s)
},
.reports = .{ .backend = .durable, .priority = .low },
},
// Named workers, each draining a subset of queues in STRICT priority order.
// OMIT `.workers` entirely → ONE implicit worker drains ALL queues, strict priority.
// `concurrency` = per-poll-cycle BATCH SIZE processed SERIALLY (NOT parallel handler
// threads); run handlers in parallel by declaring MULTIPLE workers.
.workers = .{
.mailer = .{ .queues = .{"emails"}, .concurrency = 4 },
.general = .{ .queues = .{ "reports", "default" } },
},
// Job-kind registry: kind name → handler `fn(*Ctx, payload: []const u8) anyerror!void`.
// (Scheduler pool size is set separately via `.pools = .{ .jobs = N }` — see §7.)
.jobs = .{
.resize_image = resizeImage,
.reindex = reindex,
},
})
A job handler receives the JSON payload as bytes and deserializes it:
fn resizeImage(ctx: *zigbase.Ctx, payload: []const u8) anyerror!void {
const parsed = try std.json.parseFromSlice(struct { id: []const u8 }, ctx.arena.a, payload, .{});
// …do the work; touch the DB via ctx.records() / ctx.app.pool.acquireWriter()…
}
Enqueueing
From any route, hook, or job:
// Compile-checked (a typo'd queue or kind is a COMPILE ERROR), mirrors App.flag:
try App.enqueue(ctx, .emails, .resize_image, .{ .id = "abc123" });
// Runtime-validated escape hatch (errors UnknownQueue / UnknownJobKind):
try ctx.enqueue(.emails, .resize_image, .{ .id = "abc123" });
payload is JSON-serialized (a []const u8 is treated as raw JSON and passed through unchanged), and the job is routed to the queue’s backend.
The framework can register two built-in job kinds on the same engine, each gated by its own config key so an unconfigured one costs nothing: "mail" (needs .mail/.mailer) backs ctx.mail().enqueue (it deserializes a MailMessage payload and delivers it); "webhook" (needs .webhooks = true) backs ctx.webhook(). Each is reached via its helper (or ctx.enqueueByName(queue, kind, payload)), not the compile-checked Job enum, which reflects only your declared .jobs. The kind names mail/webhook are reserved either way, so a consumer .jobs entry with either name is a compile error even when the built-in itself is gated off.
Durable backlog capacity
Bound retained jobs separately from in-process execution:
.queues = .{
.emails = .{
.backend = .durable,
.capacity = .{ .max_jobs = 10000, .max_payload_bytes = 32 * 1024 * 1024 },
},
},
Omit capacity for legacy unbounded admission. It requires durable storage and a max_jobs in 1..1000000; optional max_payload_bytes must be positive and fit i64. These limits apply to each database/queue name across producers using the same configuration. All producers must enable the same limits before relying on this contract: older binaries, unconfigured producers, or direct SQL can bypass admission. Apply migrations before starting the upgraded fleet.
Every retained _queue_jobs row counts: pending, delayed, claimed, retrying, done, failed/dead-letter, and canceled. Payload bytes are the stored UTF-8 byte length, including embedded NUL bytes on SQLite. Claim/retry/reclaim/cancel/completion does not free capacity; the existing terminal-history GC or deliberate operator deletion must remove rows. Choose done_ttl_s and capacity together. No automatic eviction of unfinished jobs occurs. Shrinking a limit below existing occupancy fails closed until enough retained rows are removed. This is not a disk-size or process-memory limit; indexes, row metadata, WAL, serialization, and claimed payload allocations are outside it.
ctx.enqueue, built-in mail/webhook enqueueing, scheduled mail, and transactional file cleanup all use this admission point. error.QueueFull means no row was inserted. SQLite serializes the check and insert in a writer transaction. PostgreSQL uses a per-queue transaction advisory lock and fresh READ COMMITTED snapshot; contention returns error.QueueAdmissionBusy without waiting. Standalone enqueues select READ COMMITTED explicitly. Enqueues inside ctx.tx/hooks retain atomicity with the caller transaction through a savepoint; a PostgreSQL caller using stronger snapshot isolation receives error.QueueCapacityIsolation. SQLite busy/connection errors still propagate. Use bounded retry/backoff outside the transaction, or map capacity refusal to an application-specific overload response. Never acknowledge an enqueue that failed, and do not busy-wait while holding HTTP or database resources. A successful bound enqueue remains subject to caller commit.
Inspect from a suitably authorized operator route or trusted handler:
const snapshot = try ctx.queueCapacity(.emails);
// Dynamic equivalent: try ctx.queueCapacityByName("emails")
return ctx.json(200, snapshot);
QueueCapacitySnapshot reports limits, retained_jobs, retained_payload_bytes, exact, and full; it reserves no capacity. full means no row or byte headroom; an empty payload may still fit at an exact byte ceiling if row capacity remains. Snapshots and admission first inspect at most max_jobs + 1 indexed rows. When that count exceeds the limit, exact is false, the count is a lower bound, and bytes are null without scanning payload lengths. Otherwise bytes are summed over at most max_jobs rows in the same statement snapshot. No payload content is returned. These are database-derived observations, so restart and worker recovery cannot lose accounting; a snapshot can become stale immediately under concurrent work. Disabled capacities return error.QueueCapacityDisabled rather than scanning an unbounded queue. This API is a trusted framework capability, not a built-in public HTTP endpoint; the fixture demonstrates an operator-authenticated route.
Backends, priority, and reliability
- Backend
memory(default): the job runs in-process on the bounded background worker pool with backoff retry. At-most-once across restart — an enqueued memory job lives only in RAM, so a crash/shutdown before it completes drops it. Zero schema; great for best-effort work. - Backend
durable: the job is persisted to the_queue_jobstable and drained by a per-worker poller. At-least-once — a crash after a side effect but before the row is marked done replays the job, so durable consumers must tolerate replays (idempotency keys are the antidote). A reclaim sweep resets jobs stranded by a crashed worker (claimed longer than that queue’svisibility_timeout_s), and a GC sweep reaps done/failed rows older than itsdone_ttl_s— both are per-queue (setvisibility_timeout_sabove the queue’s longest job runtime, or a long job gets reclaimed mid-flight and re-dispatched). The poller and GC jobs are installed only when a durable queue is declared — pure-memory and no-queue apps install nothing (zero overhead). - Priority (
high/normal/low) is a per-queue property. A worker bound to several queues drains them in strict priority order: all readyhigh-priority jobs first, thennormal, thenlow. There is no weighted-fair scheduler — priority plus the worker/concurrencytopology is the throughput and anti-starvation lever. - Throughput vs. parallelism: a worker’s
concurrencyis the per-poll-cycle batch size processed serially within that worker — it is NOT a count of parallel handler threads. To run handlers concurrently, declare multiple workers (each is its own scheduler job on its own pool thread). - Retry/backoff: on a retryable failure the durable job’s
attemptsis bumped and its next run is pushed out by the queue’s backoff (fixedorexponential, with optional jitter, capped atmax_ms); exhaustingmax_attemptsmarks itfailedand fires your.onErrorhandler (phase.job). Memory jobs retry in-process the same way. - Rate throttling (
.rate = .{ .per_second = N }, durable only): a per-second window per rated queue, capacity =per_second(max burst = one second), enforced at claim time — a job without remaining capacity this tick simply waits for the next ~500ms scheduler tick rather than sleeping in the worker. Integer-second windows are stored in the database: all instances sharing that database and queue name share the ceiling. Claims and their rate charge commit atomically; empty claims spend no capacity, and restarts do not reset the current window. Deploy the same rate configuration on each instance (mixed configurations apply each claimant’s own ceiling to the window’s already-used capacity)..rateon amemoryqueue is a compile error (memory jobs dispatch inline and can’t be throttled).
Caveats
- Durable workers poll (roughly every scheduler tick, ~0.5s), so durable jobs drain with low but non-zero latency. PostgreSQL workers use
FOR UPDATE SKIP LOCKED, so multiple instances can drain the same durable queues without claiming the same pending row. Completion is fenced by a claim generation: a handler that outlives its visibility timeout cannot overwrite a newer attempt’s state. Delivery stays at-least-once; fencing does not undo external side effects, so handlers still need idempotency and an appropriate visibility timeout. A worker claims its whole batch before running handlers serially: size the visibility timeout for the entire batch’s worst-case processing time, not just one handler, or reduce.concurrency(the batch size). No lease heartbeat extends that timeout. The scheduler itself is per-process (see §7 caveats); ordinary scheduled handlers are not coordinated. - Upgrade coordinated workers together: drain and stop older worker binaries before applying migration
0025_queue_ratesfrom one process, then start the upgraded instances. Older binaries neither enforce shared rate windows nor fence completion, so mixed-version workers do not provide these guarantees. Older binaries still need a single migration leader; current PostgreSQL consumer and system migrations, and SQLite consumer batches, are serialized automatically (see §8). - Memory jobs run on a bounded worker pool that is drained and joined at shutdown (like
app.submit); jobs still queued at a hard crash are lost (at-most-once), and a full ring rejects witherror.QueueFull.
8. Define your schema in code (.collections + .migrations)
Declare collections at comptime instead of provisioning them through the REST API (see recipes.md).
Built-in system migrations coordinate concurrent startups automatically. PostgreSQL takes a database-scoped transaction advisory lock in a READ COMMITTED transaction before ledger bootstrap and each migration’s applied-state check; SQLite takes its immediate writer lock. Each migration retains its own commit boundary. Failure rolls back that migration and releases its lock, so another startup can retry; earlier committed migrations remain applied. A failed lock acquisition aborts startup rather than proceeding unlocked (PostgreSQL deployment lock_timeout applies). Old binaries that lack the lock still require a single migration leader; continue draining old workers before incompatible queue upgrades. The system runner and public ledger bootstrap (ensureLedger, also used by status/rollback/doctor) share this lock and must run outside a caller-owned transaction. Nested use logs an explicit refusal and preserves the caller’s transaction, returning ExecFailed for compatibility with the existing database error surface.
Consumer migrations on PostgreSQL also coordinate automatically. Forward application and rollback acquire a separate database-scoped session advisory lock before reading the ledger, including rollback selection/preflight. The lock spans the entire batch, while each migration keeps its existing commit boundary and .transactional = false callbacks remain outside a transaction. This requires a direct PostgreSQL connection or a session-pooling proxy; transaction/statement-pooling proxies cannot preserve this lock. SQLite consumer batches coordinate through an exclusive, nonblocking file lock. For a file-backed main database, apply and rollback acquire <canonical-main-database-path>.migrations.lock before reading the consumer ledger, including rollback selection/preflight, and retain it across individual commits and non-transactional callbacks. Contention returns MigrationBusy immediately; retry the command after the other batch finishes. There is no background waiter, lease table, or per-request cost. Private memory/temporary databases need no cross-process lock (shared-cache SQLite is not compiled in).
For framework apply/rollback commands and startup with declared consumer migrations, the SQLite lease is acquired after opening the pool but before prerequisite system migration or ledger setup, and retained through the consumer batch’s cleanup. Pool opening/WAL initialization and later automatic collection provisioning are outside this lease. PostgreSQL retains its existing system-setup-before-consumer-lock order.
The sidecar is a permanent, empty file; never remove or replace it while migrations may run. Closing its descriptor releases the lock, including when the process dies. Its parent directory must be writable, trusted, and on a local filesystem with working file locks. A symlink at the sidecar filename is refused, not followed. Open/locking failures propagate distinctly from MigrationBusy, without running the consumer callbacks. Canonicalization covers relative paths and symlinks, not hard-link aliases: all participants must use the same canonical database path. Do not rename/replace the database or lock file during use. This only coordinates cooperating current-version consumer runners, not old binaries, raw SQL writers, automatic collection provisioning, shared-cache memory databases, or distributed filesystems. These other workflows still need an external single-leader procedure.
Consumer runners refuse caller-owned transactions with MigrationCallerTransaction. A callback that leaves a transaction open stops the batch at that migration with MigrationUnfinishedTransaction — checked immediately after the callback returns, so the offender is never recorded as applied and no later migration joins the transaction it left behind — and the transaction is rolled back before the lock is released. A callback’s original error is preserved when transaction cleanup succeeds. Advisory unlock is attempted even if rollback fails, and real cleanup failures take precedence; the error a cleanup failure displaces is logged. The unlock’s own result is checked, so a connection that does not hold the lock it acquired — the signature of a transaction-pooling proxy — fails with MigrationLockSessionLost instead of reporting coordination that silently did not happen. Locks are released on normal completion and callback errors; connection loss releases PostgreSQL session locks, and process exit releases SQLite file locks. SQLite retains its lock until transaction cleanup finishes. Lock/cleanup failures propagate and abort framework startup; low-level callers must discard the connection after a cleanup failure. Set PostgreSQL lock_timeout to bound lock waits. This does not make external side effects atomic, coordinate old binaries, or make incompatible mixed-version schemas safe.
ZigBase provisions the declarations at startup:
zigbase.App(.{
.collections = .{
.users = .{ .type = .auth, .fields = .{
.{ .name = "display_name", .type = .text },
} },
.posts = .{ .fields = .{
.{ .name = "title", .type = .text, .required = true },
.{ .name = "author", .type = .relation, .target = "users" }, // by NAME
.{ .name = "status", .type = .select, .values = .{ "draft", "published" } },
}, .rules = .{ .list = "status = \"published\"" } },
},
}).runCli(init);
The .collections shape
.collections is a struct literal whose field name is the collection name. Each value is a struct with:
.type—.base(default) /.auth/.view(aschema.CollectionType)..fields— a tuple of field literals (see below)..rules— optional.{ .list, .view, .create, .update, .delete }(any subset; each is a filter-expression string). Safe-by-default: an omitted rule,null, or""is Locked (superuser only). Use the explicit sentinel"@public"to open an operation to everyone (ZigBase logs a startup warning for every@publicrule). Any other string is a filter expression checked per record (full grammar in api.md → Filter grammar — comparisons, relation traversal, theinset-membership operator, and the@request.auth.*/@request.account.*macros).
Each field literal needs .name and .type, plus optional .required, .unique, .hidden, and type-specific options:
.{ .name = "title", .type = .text, .required = true, .min = 1, .max = 200 }
.{ .name = "price", .type = .number, .mode = .fixed, .scale = 2 } // .float (default) / .int / .fixed
.{ .name = "owner", .type = .relation, .target = "users", .maxSelect = 1, .cascadeDelete = false }
.{ .name = "status", .type = .select, .values = .{ "draft", "published" }, .maxSelect = 1 }
.{ .name = "avatar", .type = .file, .maxSelect = 1, .mimeTypes = .{ "image/png" } }
.{ .name = "meta", .type = .json }
The full field-type catalog (text / email / url / editor / date / autodate / bool / number / json / select / relation / file) and their options is in fields.md. A relation field’s .target is the target collection name — provisioning resolves it to the target’s id (no need to capture ids as you would over the REST API). Mistakes are caught at compile time: an unknown field type, a select without .values, a fixed number without a valid .scale = 1..8, or a relation without .target is a compile error. Field names reserved by the engine (id, created, updated, email, username, passwordHash, tokenKey, verified, token_epoch) are also rejected at compile time with a clear error message. A malformed text field .pattern (invalid regex syntax) or a malformed date field .min/.max bound in a comptime .collections literal is likewise a @compileError at build time — consistent with the rest of the comptime-validated surface.
Indexes
A collection may declare .indexes — a tuple of index literals provisioned as CREATE INDEX statements when the collection is created, and reconciled on later startups even when no fields change. Added, removed, or changed declarations update only their indexes and metadata; index-only changes do not copy the table or remove unrelated indexes created by explicit migrations. A failed unique-index build rolls back index changes and metadata together and refuses server startup, with collection/index diagnostics. The declaration is authoritative, including the default empty .indexes: indexes added through REST/admin metadata but omitted from code are removed at the next startup. Every DROP/CREATE is logged. Existing ordinary migration-created indexes with matching table, columns, uniqueness, direction and collation are adopted into metadata ownership rather than recreated. Conflicting names or unsupported definitions (such as partial indexes) are preserved and refuse startup; resolve them with an explicit migration. Once adopted, an index follows the declaration, including later removal. Unchanged declarations do not rebuild indexes or bump schema generation:
.indexes = .{
// unique, case-sensitive (default collation)
.{ .name = "idx_users_handle", .fields = .{"handle"}, .unique = true },
// case-insensitive: SQLite emits ("email" COLLATE NOCASE); Postgres a lower("email") functional index
.{ .name = "idx_users_email", .fields = .{"email"}, .unique = true, .collation = .nocase },
// partial / conditional-unique: emits ... WHERE deleted_at IS NULL
.{ .name = "idx_active_slug", .fields = .{"slug"}, .unique = true, .where = "deleted_at IS NULL" },
},
.collation (.binary default / .nocase) is applied to every indexed column; .where is an optional partial-index predicate emitted verbatim as WHERE <where>. The index .name and .fields are validated as identifiers; the .where predicate is raw SQL authored in the schema. Index .fields reference fields by their declared name (not an internal id); .name and each field are validated as identifiers at compile time.
Row expiry (TTL) — .ttl_field
A collection may name an existing date/autodate field as the row’s expiry timestamp via .ttl_field. A framework-internal garbage-collector then reaps rows whose timestamp is in the past — automatically, with no cron of your own:
.sessions = .{
.fields = .{
.{ .name = "token", .type = .text },
.{ .name = "expires_at", .type = .date }, // ISO-8601 UTC, e.g. "2026-07-01T00:00:00Z"
},
.ttl_field = "expires_at", // rows with expires_at <= now are deleted
},
Semantics:
.ttl_fieldmust name a field declared on the same collection and of type.dateor.autodate— anything else is a compile error (those types hold an ISO-8601 instant). Both the read-exclusion predicate and the GC normalize both sides via SQLitestrftime(...)before comparing, so non-canonical.datevalues (timezone offsets, space separator, date-only) are handled correctly — do not assume lexical comparison.- A row is reaped when its ttl value is non-null and at/before “now”. A row whose ttl field is
nullnever expires (so an optional, never-set expiry is a permanent row). A row with an unparseable ttl value is treated the same asnull— fail-safe: it remains visible and is never reaped by the GC. - Read-time exclusion: expired rows are automatically hidden from every read (list and get, via the HTTP API,
ctx.records(), and relation expand). The predicate is ANDed with your filter, access rule, and keyset cursor, so it composes transparently. You do not need to add a manualexpires_at > @nowfilter. - The GC sweep still runs once at startup and then on the
.ttl_gc_intervalcadence (aschedule.Interval, default.{ .minutes = 5 }; framework-internal job_ttl_gc, see §7), deleting expired rows so the table does not grow unboundedly. Tune it with the top-level.ttl_gc_intervalApp config key (e.g..hourly); the read-time exclusion above hides expired rows immediately regardless of sweep cadence. Declaring a TTL collection starts the scheduler even if you have no.cronof your own. - With an
.autodatettl field, prefer one whose value you set explicitly to the intended expiry (autodate defaults to write-on-create “now”, which would expire the row immediately).
Multi-tenancy — account-scoped collections (.tenancy + .tenant_field)
ZigBase has built-in account-scoped multi-tenancy (#156): a collection can be tenant-owned, and every read/write of it is automatically narrowed to the request’s active account — you do not write account = @request.account.id on every rule, and a client can never read or write across tenants.
Turn it on in App(.{ ... }) and mark each owned collection’s owning-account column with .tenant_field:
zigbase.App(.{
.tenancy = .{
.enabled = true,
.resolver = .header, // read the active account from X-Account-Id / the zb_account cookie
.auth_collection = "users", // the auth collection whose records are members
.roles = .{ "viewer", "editor", "admin", "owner" }, // optional; this is the default ladder
},
.collections = .{
.projects = .{
.fields = .{
.{ .name = "title", .type = .text },
.{ .name = "account", .type = .relation, .target = "_accounts" }, // owning account
},
.tenant_field = "account", // <- makes `projects` tenant-owned
.rules = .{ .list = "@public", .view = "@public" },
},
},
});
.tenant_field accepts any TEXT-storage field — a plain .text column works too — but a .relation is recommended: it’s the form .abilities’ .via requires (below), so the same account column can back both tenancy scoping and relationship-based abilities.
Data model. Migration 0014_tenancy creates three system collections (visible in _collections, system = 1):
_accounts(id, created, updated, name, slug UNIQUE, owner_user, status)— one row per tenant._memberships(id, …, account, user_collection, user, role, status)— the principal↔account edge.(account, user_collection, user)is unique;user_collectionis part of the key because several auth collections may exist. Indexed on(user_collection, user, status)(“my accounts”) and(account, status)._invitations(id, …, account, email, role, token UNIQUE, invited_by, expires, accepted_at)— pending invites. (Invite/accept/remove lifecycle is a later PR; PR2 ships the tables only.)
Resolution (fail closed). On each request the active account is resolved from X-Account-Id (or the signed zb_account cookie) and verified against an active _memberships row for the authenticated principal — one indexed SELECT, cached on the request. No/invalid/inactive membership ⇒ no account context, and tenant-owned data is invisible. The resolution also fills the rule macros @request.account.id, @request.account.role, and the list-valued @request.account.ids (the in operator’s membership-bounded source).
Activation for browsers. POST /api/accounts/:id/activate verifies membership and sets a signed, HttpOnly zb_account cookie, so a browser SPA selects an account once instead of sending X-Account-Id on every call. API clients can skip this and just send the header. 403 without an active membership; 404 when tenancy is disabled.
Scoping & the write path. A tenant-owned collection is auto-scoped at all CRUD chokepoints and realtime delivery via a bound "<col>"."<tenant_field>" = ? predicate composed into the guard stack — WHERE (filter) AND (rule) AND (tenant_field = ?) AND (ttl). A locked rule (null/"") still denies first (the fail-closed floor is unchanged). On create the owning account is stamped onto the row (clients can’t spoof it); on update a cross-tenant move is rejected by the in-transaction guard, and a cross-tenant target is rejected before the before_update hook runs (so hooks never fire against another account’s row). Creating in a tenant-owned collection with no active account is denied.
Realtime is scoped too. A WebSocket connection resolves its active account at the handshake (the signed zb_account cookie or X-Account-Id header) and verifies a membership at auth-frame time; delivery of a tenant-owned collection’s create/update/delete frames (including the delete snapshot) is then filtered to the subscriber’s account. With tenancy enabled, a connection with no resolved account receives nothing from a tenant-owned collection — even one with a @public viewRule — so realtime never leaks across accounts (fail closed).
Roles form a total order (tenancy/roles.zig), default viewer < editor < admin < owner, configurable via .tenancy.roles. (PR3 consumes the ranking for ability checks.)
Superuser & cross-tenant tooling. Superusers bypass tenancy entirely (consistent with the access-rule engine). For admin tooling that must legitimately span accounts (an ops dashboard, a maintenance job), zigbase.crossTenant(rctx) returns a context with the override enabled — the explicit, never-silent way to widen scope.
Back-compat is byte-identical. With .tenancy absent/enabled = false, or for any collection without a .tenant_field, the composed SQL and authorization decisions are identical to the pre-tenancy engine (pinned by tests in policy.zig). Enabling tenancy never changes a non-tenant collection.
Relationship-based row abilities (.abilities)
Tenancy scopes a collection to one active account. Abilities (#155) authorize a CRUD action by the principal’s relationship to the row — “you may edit a project if you are an editor (or higher) of the account it belongs to” — without writing that membership join by hand in every rule. Abilities are declared at the top level of App(.{ … }), keyed by collection name, with a per-action relationship rule:
const App = zigbase.App(.{
.tenancy = .{ .enabled = true, .auth_collection = "users",
.roles = .{ "viewer", "editor", "admin", "owner" } },
.collections = .{
.projects = .{
.fields = .{
.{ .name = "title", .type = .text },
.{ .name = "account", .type = .relation, .target = "_accounts" }, // owning account
},
.rules = .{ .list = "@public", .view = "@public" },
},
},
// A row of `projects` is authorized when the principal holds a membership (role ≥ floor) of the
// account named by the `account` relation field.
.abilities = .{
.projects = .{
.view = .{ .relationship = .{ .via = "account" } }, // any active member
.update = .{ .relationship = .{ .via = "account", .min_role = .editor } },
.delete = .{ .relationship = .{ .via = "account", .min_role = .admin } },
.create = .{ .relationship = .{ .via = "account", .min_role = .editor } },
},
},
});
Lowering. Each rule compiles to a bound IN predicate over the principal’s qualifying membership account-ids — "projects"."account" IN (?,?,…) — exactly the shape the in operator and @request.account.ids macro already emit. min_role filters the membership set through the configured role ladder (.tenancy.roles); .via must name a relation field (the field whose column holds the owning account id). list reuses the view ability.
Composition. The ability predicate is AND-ed into the same guard stack as the access rule and the tenant scope, on every chokepoint — view/create/update/delete, expand, realtime delivery, AND the bulk list endpoint: WHERE (filter) AND (rule) AND (ability) AND (tenant_field = ?) AND (ttl). An ability forces a per-row check even when the access rule alone would allow, so a .rules.list = "@public" collection with a view ability returns the ability-narrowed set (HTTP 200) — not every row, and not a 400.
Fail closed. No qualifying membership ⇒ the constant-false predicate 0 (the row is denied; never SQLite’s invalid IN ()). A locked rule (null/"") still denies first. Account-ids are always bound parameters, never interpolated; superusers bypass abilities entirely.
Comptime validation. An ability naming an unknown collection, a .via that is not a relation field, or a .min_role not in .tenancy.roles is a loud @compileError. .abilities also requires .tenancy.enabled = true (abilities authorize by account membership, which only resolves under tenancy) — configuring abilities with tenancy disabled is a @compileError rather than a silent deny-all at runtime.
Custom routes. try ctx.can(.update, "projects", id) authorizes a specific record through the same policy (rule + ability + tenant scope) the REST chokepoints use — use it instead of re-implementing the check in a handler.
Introspection. GET /api/collections/:col/records/:id/abilities returns a JSON object of booleans like {"view": true, "update": false, "delete": false} — the actions the current principal may perform on that record. The endpoint itself requires view access (404 otherwise), so it never leaks a record’s existence (and "view" is therefore always true on a 200 response).
Back-compat is byte-identical. A collection with no .abilities entry composes a null predicate, so its decisions and compiled SQL are identical to the pre-abilities engine (pinned in policy.zig).
Product analytics (.analytics + ctx.track)
Built-in event capture and declarative rollups. Two halves:
1. Capture — ctx.track(name, payload). From any hook, route, or job, append one immutable event to the built-in _events collection:
fn afterSignup(ctx: *zigbase.Ctx, ev: *zigbase.RecordEvent) anyerror!void {
try ctx.track("user.signup", .{ .plan = "pro" });
}
The actor / actor_collection (the authenticated principal), the account (the request’s active tenant scope, "" when tenancy is off), and the occurred_at timestamp are all resolved server-side — a client cannot forge any of them. payload is any JSON-serializable value, stored as opaque JSON text (a []const u8 is taken as raw JSON). It is a single cheap INSERT; inside a hook / ctx.tx it reuses the in-transaction connection.
For multiple events, ctx.trackBatch(&.{ .{ .name = "event.one", .payload_json = "{}" }, .{ .name = "event.two" } }) persists up to 1024 events atomically with one writer acquisition and prepared statement. It stamps the same server-side context as track. A failed batch rolls back its own inserts; a successful batch inside a hook/ctx.tx is still rolled back if that outer transaction fails. No background buffer or shutdown flush is introduced. See analytics.md.
2. Rollups — declarative, scheduled aggregation. Declare named rollups; each registers one job on the existing scheduler that aggregates _events into a _rollup_<name> summary table:
const App = zigbase.App(.{
.analytics = .{
.rollups = .{
.signups_daily = .{
.event = "user.signup",
.every = .{ .interval = .hourly }, // a cron/interval Schedule
.group_by = .{ .account = true, .time_bucket = .day },
.metric = .count,
},
},
},
});
.group_by keys: .account / .actor (bools) and .time_bucket (.none / .day / .hour). .metric is .count (default). Aggregation is incremental and idempotent: a persisted watermark (in _kv) tracks the monotonic _events.rowid already aggregated, and each run aggregates the disjoint window watermark < rowid <= max_rowid. Because the watermark is the rowid (not the timestamp) and the job holds the exclusive writer for the pass, a run neither double-counts nor drops — even an event inserted in the same wall-clock second as a prior run still has a strictly-greater rowid and is counted next pass. Summary-table / column identifiers are gated through schema.isValidIdentifier. Misconfiguration fails loudly at compile time — an unknown .group_by/.metric, a missing or empty .event, or a rollup name that is not a valid identifier is a @compileError.
Read API (tenant-scoped, fail-closed). Both endpoints are authenticated and never leak across accounts:
GET /api/analytics/events?name=&actor=&since=&limit=&cursor=— the raw activity feed; paginates with the house cursor (nextCursor/hasNextin the response).GET /api/analytics/rollups/:name?from=&to=— a rollup’s summary rows (for charts).
A superuser sees everything; a member sees only their active account’s data (resolved from a verified _memberships row, the same path the records chokepoints use); with tenancy disabled the feed is scoped to the caller’s own events and a (global) rollup is superuser-only. A member can never read another account’s events or rollups.
Visibility is account-level, not role-level. Any active member of an account — whatever their role — can read the entire account’s event feed (including other members’ events and payloads) and all of its rollup buckets. The trust boundary is the tenant, not the role; there is deliberately no intra-account role gating. Treat event payloads as readable by every member of the account.
Back-compat. With no .analytics config the _events table is still seeded (harmless) but no rollup job is scheduled, and ctx.track works standalone.
Field encryption at rest (.encrypted)
Mark a text, editor, or json field .encrypted = true to store it encrypted at rest. The records layer encrypts on write and decrypts on read, so your handlers, the records API, and the HTTP responses always see plaintext — only the SQLite file holds ciphertext:
.fields = .{
.{ .name = "ssn", .type = .text, .encrypted = true },
.{ .name = "notes", .type = .json, .encrypted = true },
},
Envelope. AES-256-GCM, versioned: each value is stored as v<N>: + base64url(nonce ‖ ciphertext ‖ tag) with a fresh random nonce per write, where N is the key generation (default 1 — see Key rotation below).
Key (required). The key comes only from the ZIGBASE_FIELD_KEY environment variable (HKDF-derived; the raw value may be any length). Unlike the JWT secret it is never auto-generated, persisted, or logged — losing or rotating it determines whether the data is recoverable, so you must manage it. If any collection declares an .encrypted field and ZIGBASE_FIELD_KEY is unset, the server refuses to start (fail-closed — it never silently stores plaintext). This holds for both comptime .collections and collections created at runtime (via the admin/collections API): startup scans the live database schema after provisioning, so a restart without the key is refused even for a runtime-added encrypted field.
Constraints (enforced). Encrypted values are per-row-nonce ciphertext, so they cannot be indexed, marked .unique, or used in a ?filter/?sort:
- Indexing an encrypted field, marking it
.unique, or setting.encryptedon a non-text/editor/jsonfield is a compile error. - A request that filters or sorts by an encrypted field gets a 400.
- Access rules that compare an encrypted field will compare against ciphertext and effectively never match — don’t reference encrypted fields in rules.
Strict reads / enabling on existing data. Reads are strict: a stored value that is not a valid v<N>: envelope (e.g. legacy plaintext) or that fails authentication (wrong key, tamper, or an unconfigured key generation) fails closed — there is no plaintext passthrough. Therefore, turning .encrypted on for a column that already holds plaintext rows requires running zigbase rewrap first (see below); “encrypted means encrypted”.
Key rotation
The v<N>: envelope prefix is the key generation: a v<N>: value is decrypted with generation N’s key. Writes always use the primary generation and stamp its version; reads dispatch on the prefix. This lets you write under a new key while still reading old data, then migrate the old data forward.
Configuration (env only — keys are never auto-generated, persisted, or logged):
| Env var | Meaning |
|---|---|
ZIGBASE_FIELD_KEY | The primary (current/write) key. Required if any field is encrypted. |
ZIGBASE_FIELD_KEY_GENERATION | Integer 1..64, default 1 — the generation of the primary key (= the v<N>: version written). |
ZIGBASE_FIELD_KEY_V<M> | A read-only key for an older generation M, used to decrypt existing v<M>: data. |
The default (just ZIGBASE_FIELD_KEY, generation 1) is identical to the single-key build — it writes and reads v1:. Each generation derives an independent key (HKDF, domain-separated by generation), so generations never share key material. Setting ZIGBASE_FIELD_KEY_V<M> for the primary generation M is a fatal config error (ambiguous — the primary key already comes from ZIGBASE_FIELD_KEY).
To rotate from generation 1 (key oldkey) to generation 2 (key newkey):
Restart with
ZIGBASE_FIELD_KEY=newkey,ZIGBASE_FIELD_KEY_GENERATION=2,ZIGBASE_FIELD_KEY_V1=oldkey. New writes arev2:; oldv1:rows still read.Run the rewrap command to re-encrypt every
v1:cell asv2::ZIGBASE_FIELD_KEY=newkey ZIGBASE_FIELD_KEY_GENERATION=2 ZIGBASE_FIELD_KEY_V1=oldkey \ zigbase rewrap --data-dir ./zb_dataOnce rewrap completes, drop
ZIGBASE_FIELD_KEY_V1— nov1:data remains.
zigbase rewrap walks every .encrypted field of every collection, decrypts each cell with whichever generation matches its envelope version (or, for legacy plaintext, takes it as-is), and re-encrypts it under the primary key. It is the supported path both to finish a rotation and to migrate existing plaintext into ciphertext when first enabling .encrypted. It is idempotent (cells already at the primary version are skipped), transactional per collection, and fail-closed: a cell it cannot decrypt (missing generation key, wrong key, tamper) aborts the run with the offending row reported and that collection’s transaction rolled back — no data is lost. --dry-run reports counts without writing. Run it with the primary key plus every older generation present in your data configured.
Memory note: rewrap buffers a collection’s rewritten cells in memory before writing them back, so peak memory is O(rows) in the collection being processed. This is fine for a one-off maintenance command on typical tables; chunked rewrapping for very large encrypted tables is a possible future option.
Note: the envelope hides a value’s contents but not its length — ciphertext length is proportional to plaintext length. A single long-lived key suits typical volumes; for very high write volumes, periodic key rotation is recommended.
Startup provisioning + additive auto-migration
On every startup, ZigBase diffs each declared collection against the live database and applies the minimal safe change set (running it twice is a clean no-op):
- A collection that doesn’t exist yet is created.
- A field present in the spec but missing from the live collection is added, rebuilding the table while preserving existing data (the new column is null for old rows).
- A changed access rule (
.list_rule/.view_rule/.create_rule/.update_rule/.delete_rule) is re-applied to the live collection. This is a metadata-only write — no table rebuild, no data copy — so tightening a rule in code and redeploying really does tighten it in the running server. A rule left unset (null) in the literal means “unspecified: leave whatever is live alone”; to lock a rule explicitly, spell it""(blank and unset are the same rule at evaluation time — superusers only — so switching between them is not treated as a change). - A non-additive change — a field rename, drop, or type/storage-class change — is detected, logged, and SKIPPED (never applied, so no data loss). Relation targets must reference a known collection (a comptime collection or a pre-existing live one such as
_superusers); an unknown target is a startup error.
For the changes auto-migration won’t do, use the .migrations escape hatch.
.indexeschanges on an existing collection are not yet re-applied. Adding or removing an entry in a live collection’s.indexesonly takes effect if that same startup also adds a field (which rebuilds the table). Use an explicit.migrationsentry with rawCREATE INDEX/DROP INDEXto change indexes on a collection that already exists.
Explicit migrations (.migrations)
.migrations is a list of zigbase.Migration records, each run once (recorded in _migrations under a prov: prefix) before provisioning. It accepts either a bare tuple (lowered at comptime, like every other list-shaped config key — .routes, .cron, .static_routes) or a typed slice &[_]zigbase.Migration{ ... }, which coerces directly:
.migrations = .{
.{ .id = "0001_rename_title", .up = renameTitle },
},
// or, equivalently, the typed-slice form:
// .migrations = &[_]zigbase.Migration{
// .{ .id = "0001_rename_title", .up = renameTitle },
// },
// .up signature: fn (m: *zigbase.Migrator) anyerror!void
// `m` carries the active SQL dialect, the writer connection (`m.db`), a scratch
// arena (`m.arena`), and `m.io`. It exposes exec/execLowered/prepare/rawFor.
fn renameTitle(m: *zigbase.Migrator) anyerror!void {
// RENAME COLUMN is identical on SQLite and Postgres, so plain raw SQL is fine here.
try m.exec("ALTER TABLE \"posts\" RENAME COLUMN \"headline\" TO \"title\";");
}
Each migration has an .id (used for the once-only record) and a forward step. Use migrations for renames, drops, type changes, and data backfills. The full shape:
pub const Migration = struct {
id: []const u8,
/// Auto-reversible forward change (Rails-style). Written once with the DSL; the same body
/// inverts for rollback. Set EXACTLY ONE of `change` or `up`.
change: ?*const fn (m: *zigbase.Migrator) anyerror!void = null,
/// Explicit forward step — use when the change isn't auto-reversible (raw SQL, data transforms).
up: ?*const fn (m: *zigbase.Migrator) anyerror!void = null,
/// Explicit reverse step for rollback (`migrate rollback`). Pairs with `up`, or overrides a
/// `change`'s auto-derived inverse.
down: ?*const fn (m: *zigbase.Migrator) anyerror!void = null,
/// Per-migration transactional opt-out (default `true`) for DDL that can't run in a transaction.
transactional: bool = true,
};
Exactly one of change/up is required (comptime-enforced); down needs a forward step.
The schema DSL & auto-reversible change
A change migration describes the forward schema edit once, using a dialect-aware DSL, and ZigBase derives the inverse for rollback (Rails’ change). The DSL emits backend-correct SQL on SQLite and Postgres, so you don’t hand-write ALTER TABLE per dialect:
.migrations = .{
.{ .id = "0002_add_comments", .change = addComments },
},
fn addComments(m: *zigbase.Migrator) anyerror!void {
try m.createTable("comments", &.{
.{ .name = "id", .type = .integer, .pk = true, .null = false },
.{ .name = "post_id", .type = .integer, .null = false },
.{ .name = "body", .type = .text },
.{ .name = "created", .type = .timestamp },
});
try m.addColumn("posts", .{ .name = "comment_count", .type = .integer, .null = false, .default = "0" });
try m.addIndex("comments", &.{"post_id"}, .{});
}
Applied forward, that creates the table, adds the column, and builds the index. A rollback (migrate rollback) re-runs the same function with m.direction == .reverse, where each op emits its inverse.
Sharp edge — ops do NOT reorder within a
change. Reverse re-runs the body top-to-bottom, inverting each op in place (the DSL emits SQL eagerly — there is no recorded op list to replay backwards). So a singlechangethat doescreateTable("t", …)then a dependentaddColumn("t", …)reverses asDROP TABLE tthenDROP COLUMN t.…— and the second step fails because the table is already gone. Put acreateTableand a dependentaddColumn(or any op that depends on an earlier op in the same body) in separate migrations: rollback processes migrations newest-first, so the dependent one unwinds before its dependency. Independent ops in onechange(e.g. touching different tables) are fine.
The v1 op set (each m.<op>(...) anyerror!void):
| Op | Forward | Auto-inverse |
|---|---|---|
createTable(name, cols) | CREATE TABLE | DROP TABLE |
dropTable(name, .{ .was = cols }) | DROP TABLE | re-CREATE from .was |
addColumn(table, col) | ADD COLUMN | DROP COLUMN |
dropColumn(table, name, .{ .was = col }) | DROP COLUMN | re-ADD from .was |
addIndex(table, cols, .{ .name?, .unique? }) | CREATE [UNIQUE] INDEX | DROP INDEX (same derived name) |
dropIndex(name, .{ .was = .{ table, cols, unique } }) | DROP INDEX | re-CREATE from .was |
renameTable(from, to) | RENAME TO | rename back (self-inverse) |
renameCollection(from, to, .{ .offline = true }) | Atomic collection metadata/table/dependency rename | Rename back; revoked sessions are not restored |
renameColumn(table, from, to) | RENAME COLUMN | rename back (self-inverse) |
addForeignKey(table, col, ref_table, .{ .ref_column?, .on_delete_cascade?, .name? }) | ADD CONSTRAINT … FOREIGN KEY (Postgres only) | DROP CONSTRAINT (same derived name) |
Col is .{ .name, .type, .null = true, .pk = false, .default = null } and ColType is one of .text / .integer / .real / .boolean / .blob / .timestamp (mapped to the backend-native keyword). An index/FK name defaults deterministically (idx_<table>_<cols…>, fk_<table>_<col>) so the inverse can name the object it created.
Auto-reversibility rule. An op is auto-reversible when its inverse is unambiguous. The drop ops (dropTable/dropColumn/dropIndex) need a .was snapshot to rebuild from — without .was they are irreversible and a reverse pass fails loudly (error.MigrationNotReversible). Raw SQL (m.raw/m.rawFor/m.exec) and data transforms (m.records(), below) have no general inverse and are likewise irreversible in reverse mode. addForeignKey is Postgres-only (SQLite has no ALTER TABLE ADD CONSTRAINT; it returns error.WrongBackend — declare the FK inline in createTable, or use m.raw). Whenever a step is irreversible, write it as an explicit up and supply your own down.
Offline collection rename
Use try m.renameCollection("posts", "articles", .{ .offline = true }); in an explicit migration, not renameTable: the latter changes only a physical table. Both arguments must be collection names, not collection IDs. Stop every serving process and worker before migrating. The flag acknowledges that coordination; it cannot stop remote processes for you. Deploy changed .collections declarations together with the migration, and update all name literals in policies, auth configuration, hooks, routes, SDKs, and integrations. SQLite (with legacy_alter_table disabled) rewrites dependent trigger/view SQL references during its table rename; PostgreSQL maintains catalog-bound dependencies. The helper does not rewrite arbitrary SQL name literals, serialized payloads, or rule strings; review and migrate those explicitly in the same maintenance window. There are no old-name URL or realtime-topic aliases.
Use a migration connection without pre-existing temporary relations: they can shadow data or engine metadata tables, so the operation rejects them before taking the schema lock. PostgreSQL collections and their registry must resolve in the current persistent schema. Ambiguous legacy relation names that also identify another collection’s stable ID are rejected instead of silently retargeted. Destination-name relation metadata or dangling SQLite foreign keys likewise return Conflict and require explicit repair before retrying.
The operation takes the schema lock and commits as one transaction (or joins the caller’s transaction through a savepoint). It preserves collection, field, and record IDs. If rollback or savepoint cleanup itself fails, the offline migration terminates the process rather than returning an unsafe writer to its pool; recover the database before restarting migrations. Normal failures roll back and return their original error. The successful operation rewrites incoming/self relation metadata to stable IDs, retains user index names, and rebuilds generated search indexes. Subsequent additive provisioning under the new declaration retains existing field IDs. Destination names must be valid identifiers of at most 55 bytes (as must the source, so reversal remains possible); existing names and generated search-object collisions are rejected rather than overwritten.
Generated search objects must match the current searchable schema exactly before the rename removes them. Stale objects after search was disabled, or custom objects using generated names, cause a conflict and remain untouched; reconcile those objects explicitly before retrying the rename. SQLite also rejects arbitrary trigger/view definitions mentioning the old or destination generated FTS name (conservatively including comments or literals), because rebuilding that table cannot rewrite those references safely. Verified engine auth indexes are recreated under the new name on SQLite and renamed in place on PostgreSQL; application-owned indexes retain their names.
OAuth links, WebAuthn/TOTP credentials, rate budgets, memberships, and actor attribution follow the renamed identity. Signing keys rotate, sessions and pending auth challenges are revoked, and stateful cursors are deleted. Users must sign in again; a reverse rename does not restore those capabilities. Stateless/signed cursor fingerprints now include collection identity, name, and an engine-owned monotonic rename epoch. Renaming back cannot revive a previous cursor, while unrelated collection renames leave it valid. Upgrading to this version invalidates previously issued cursors, and clients must restart pagination.
Unexpired _idempotency_receipts return PendingIdempotencyReceipts before any mutation. Their scopes hash application-selected principal/operation names, so the engine cannot prove a receipt unrelated to a rename. Stop new operations and wait out the existing retention windows; receipts are not discarded or silently rebound. Rename uses the same ledger-ownership validation as idempotent execution and database copying; ambiguous or non-durable ownership returns InvalidReceiptLedger. The existing offline precondition still rejects any temporary relation with Conflict before this check.
File-bearing collections keep an immutable storage namespace. Upgrading seeds the engine-owned reservation ledger from existing collection IDs and names; no local, S3, or custom-backend objects are copied or renamed. File reads/writes, thumbnails, presigning, cleanup, inventory, and reconciliation resolve that physical prefix independently of the current public collection name. Generic collection input, schema import, and update cannot override another collection’s namespace.
Durable uploads targeting the collection or authenticated through it, and unfinished file-cleanup jobs, are relinked in the same database transaction. Both the stored name and stable collection ID must agree. Malformed or inconsistent metadata fails closed with InvalidPendingMetadata; orphan destination-name dependencies return PendingStorageDependency. Payload bytes and upload offsets are untouched. In-memory uploads are lost when their process stops, as usual. Metadata scans use 64-row keyset batches with row-scoped parsing and cleared statement bindings, including on PostgreSQL; payloads larger than 64 KiB fail closed.
Namespace reservations survive collection deletion. Creating a collection whose name is already reserved as a physical prefix fails with StorageNamespaceConflict (HTTP 409), even if the former collection was renamed or deleted. There is no automatic reclamation: leftover objects must never become a new collection’s files. New reservations also reject ASCII case variants of existing prefixes, on every backend, to protect case-insensitive local filesystems. Existing physical prefixes and inventory lookups remain exact; this does not rename or normalize legacy objects. Take a backup and test both migration directions on a copy before production.
Data transforms (m.records())
m.records() is a records-aware facade for data migrations (backfills, reshapes) that routes through the SAME engine primitives the HTTP API uses — so encrypted fields decrypt on read and re-encrypt on write, and typed fields coerce exactly as a normal write would (no re-implemented crypto/coercion):
.migrations = .{
.{ .id = "0003_slugify", .up = slugify, .down = unslugify },
},
fn slugify(m: *zigbase.Migrator) anyerror!void {
const r = try m.records();
for (try r.all("posts")) |rec| {
const id = rec.object.get("id").?.string;
const title = rec.object.get("title").?.string;
var patch: std.json.ObjectMap = .empty;
try patch.put(m.arena, "slug", .{ .string = try slugFrom(m.arena, title) });
_ = try r.update("posts", id, .{ .object = patch });
}
}
r.all(collection) reads every row (unscoped) through the engine; r.get(collection, id) and r.update(collection, id, patch) (partial merge) do single-row I/O. Data transforms are irreversible — m.records() returns error.MigrationNotReversible in reverse mode, so pair a data-transform change with an explicit down (or write it as up/down, as above).
Transactions & the .transactional opt-out
Each migration’s forward step + its _migrations bookkeeping run inside one transaction — a failure rolls the whole thing back and the migration stays un-applied (re-runnable), never half-done. Some statements can’t run inside a transaction (e.g. SQLite VACUUM, Postgres CREATE INDEX CONCURRENTLY); set .transactional = false on that migration and it runs without a wrapping transaction (it owns its own atomicity; it is recorded only after it succeeds).
The migrate CLI (apply + status + preview + rollback + dump)
zigbase migrate applies pending SYSTEM migrations, then the app’s comptime .migrations, and exits — so a deploy step can migrate ahead of starting the server. It does not provision the .collections tables; those are created when the server starts (serve), not by migrate:
# Apply pending SYSTEM migrations, then the app's comptime `.migrations`, then exit.
zigbase migrate --data-dir ./zb_data
Each applied migration is recorded once in the _migrations ledger (system migrations by name, consumer .migrations under a prov:<id> name), so migrate is idempotent: already-applied migrations are skipped.
zigbase migrate status reads that ledger and reports your comptime .migrations, in declared order, as applied (at <ts>) or pending — without applying anything:
$ zigbase migrate status --data-dir ./zb_data
Consumer migrations (2 declared):
0001_widgets applied (at 2026-07-06 12:00:00)
0002_slugify pending
1 applied, 1 pending, 0 orphaned
An orphaned row — an applied prov: migration whose id is no longer compiled into the binary (you deleted it from .migrations) — is listed separately under “Orphaned (in ledger, not in binary)” and counted in the summary, so a divergence between the ledger and the source is visible at a glance.
--json emits the same information as one object on stdout instead (see Machine-readable CLI output):
$ zigbase migrate status --json --data-dir ./zb_data
{"migrations":[{"id":"0001_widgets","applied":true,"applied_at":"2026-07-06 12:00:00"},{"id":"0002_slugify","applied":false,"applied_at":null}],"orphaned":[],"summary":{"declared":2,"applied":1,"pending":1,"orphaned":0},"ok":false}
Both forms share the same exit code: migrate status exits 1 when anything is pending or orphaned, 0 when the database is fully up to date (ok in the JSON body carries the same signal) — so it can gate a deploy step: zigbase migrate status || zigbase migrate.
Offline migration declaration preview
zigbase migrate preview [--json] emits one versioned JSON object on stdout, without loading deployment configuration, opening a database, or invoking migration callbacks. It requires -Ddev-tools=true (the default). CLI logging still reads logging preferences. Both forms emit JSON; success exits 0, invalid arguments or output failures exit nonzero. --data-dir, --out and rollback counts are rejected.
The report has protocol_version: 1, scope: "compiled-consumer-migrations", and items in declaration order. Each item contains id, transactional, forward_callback (up or change), reverse_callback (down, change, or none), and rollback_declaration:
explicit_down: a reverse callback was declared, not proven correct.change_requires_runtime_verification: transactionalchangecan be attempted in reverse; its operations may still be irreversible or backend-incompatible.missing_reverse: anupwithout a reverse callback.nontransactional_change_rejected: rollback requires an explicitdownfor this declaration.
unknown is ["pending_state", "sql", "effects", "runtime_reversibility"] and includes_system_migrations is false. No SQL is predicted or callback body analyzed; neither transactional: true nor an explicit down proves safe rollback or reversible external effects. This is a declaration inventory, not an execution plan or deploy gate. Use migrate status --json for ledger state (that command can initialize database state). Consumers should reject unsupported protocol versions and ignore unknown fields. Output is streamed from borrowed declarations, with constant-sized writer buffering and no per-migration allocations. There is no runtime catalog or database-dependent row growth.
Consumer migration rollback
zigbase migrate rollback [N] reverses the N most-recently-applied consumer migrations, newest first (N is a positional integer, default 1); system migrations are never touched:
# Reverse the single most-recent consumer migration.
zigbase migrate rollback --data-dir ./zb_data
# Reverse the three most-recent, newest first.
zigbase migrate rollback 3 --data-dir ./zb_data
The reverse of a migration is down orelse change (the mirror of the forward change orelse up): an explicit down runs as-is; otherwise the change re-runs with m.direction == .reverse, each op emitting its inverse. A migration with only an up (and no down/change) has no reverse and is refused. Each migration’s reverse body and its _migrations ledger delete commit in one transaction (honoring .transactional), so a rollback is atomic per migration and re-applying afterwards works (the ledger row is gone). Reversing across migrations is newest-first, so dependent ops split across separate migrations unwind in the right order.
Rollback fails loudly and changes nothing it cannot undo:
- A lone-
upmigration (nodown, nochange), or a non-transactional migration that would reverse via achange(no transaction to undo a partial reverse), is refused up front before anything is touched — supply an explicitdown. - A
changewhose inverse hits an irreversible op (m.raw/m.records(), a.was-lessdropTable/dropColumn/dropIndex, oraddForeignKeyon SQLite) is only knowable by running: its per-migration transaction rolls the partial reverse back, then the run fails, naming the offending migration. - An orphaned ledger row (an applied
prov:migration no longer compiled into the binary) cannot be reversed — it is named and the run fails, leaving the ledger row intact.
N larger than the number of applied consumer migrations rolls back all of them (min(N, applied)).
zigbase migrate dump [--out <file>] introspects the live database and writes a canonical, dialect-native structure.sql. Output goes to stdout by default; --out <path> writes it to a file (parent directories are created):
# Print the live schema to stdout.
zigbase migrate dump --data-dir ./zb_data
# Or write it to a file for review / test-DB setup.
zigbase migrate dump --out db/structure.sql --data-dir ./zb_data
It dumps the true live structure of both the system tables and any collection/migration-created tables — on SQLite it emits the exact stored DDL from sqlite_master; on Postgres it reconstructs the DDL from the system catalogs (information_schema, pg_get_constraintdef, pg_indexes) — no external pg_dump binary is ever invoked. Tables come first, then constraints (as ALTER TABLE … ADD CONSTRAINT, so foreign-key ordering never breaks), then indexes; finally the applied-migration ledger is emitted as a single INSERT INTO "_migrations" … so a restored dump lands at the same migration state. The output is deterministic — deterministic ordering and no dump-time timestamps or other volatile data — so it diffs cleanly across runs and in review, and it re-runs to recreate the schema (handy for a fast test database).
The dump is a snapshot for inspection, review-diffing, and test-DB setup — it is NOT a schema source (that is your comptime .collections) and it is never loaded at boot. Treat it as generated output, not authoritative configuration.
Replay scope. The dump reconstructs tables, constraints, and indexes for ZigBase-shaped schemas (text ids, IDENTITY keys); it does not emit
CREATE SEQUENCE/CREATE TYPE/CREATE EXTENSION(Postgres) or views/triggers/FTS virtual tables (SQLite). ZigBase’s own schema replays cleanly, but a hand-rolled schema usingserial/enum/extension columns or those SQLite objects may not re-run into a bare database.
Preflight checks (zigbase doctor)
zigbase doctor [--production] [--json] [--data-dir PATH] runs nine checks over a deployment — JWT-secret persistence, every @public access rule (enumerated by name), cookie security, bind address, reverse-proxy coherence, mailer configuration, pending/orphaned migrations, data-dir writability, and legacy password hashes still pending re-hash.
See serve.md → zigbase doctor for the full worked example (prose and NDJSON output, captured from a real run), the per-check severity table, exit codes, both deploy-gate idioms, and doctor’s database-mutation contract (it only ever creates the _migrations ledger table if absent — the same write migrate status already makes).
Offline bulk import (zigbase import)
zigbase import bulk-loads records from an NDJSON file (one JSON object per line) into a collection offline — the HTTP server is not running — yet routes every row through the record engine, so a hand-written INSERT never bypasses validation or the encryption envelope again:
# Import into an existing collection (created via .collections or the admin UI).
zigbase import --collection posts --data-dir ./zb_data seed.ndjson
# Idempotent re-import: match each row on a (recommended-unique) field and UPDATE it if present.
zigbase import --collection users --upsert-key email --data-dir ./zb_data users.ndjson
# Read NDJSON from stdin with a file path of `-`.
cat dump.ndjson | zigbase import --collection posts --data-dir ./zb_data -
# Rehearse a 500k-row migration without writing anything, then run it for real, tolerating
# a handful of bad rows and keeping a machine-readable record of what happened.
zigbase import --collection posts --dry-run --data-dir ./zb_data seed.ndjson
zigbase import --collection posts --continue-on-error --error-log errs.ndjson \
--progress 10000 --json --data-dir ./zb_data seed.ndjson
| Flag | Meaning |
|---|---|
--collection NAME | Target collection (required). Must already exist. |
--upsert-key FIELD | Match each row on FIELD; UPDATE if present, else create. FIELD must be a scalar, non-encrypted, existing field; a unique constraint is recommended (a non-unique key logs a warning). |
--batch-size N | Rows per transaction (default 500; must be ≥ 1). |
--data-dir PATH | Data directory (also ZIGBASE_DATA_DIR, default ./zb_data). |
--dry-run | Validate and execute every row through the full engine, then roll back instead of committing. Nothing is written; the summary reports what a fresh import would do. Because nothing commits, an --upsert-key lookup never sees rows “created” earlier in the same dry run. |
--continue-on-error | Skip a failing row (wrapped in its own SAVEPOINT) instead of aborting the whole import. The good rows in that batch still commit. |
--error-log FILE | NDJSON sink for per-row failures, one object per line: {"line":N,"code":"…","detail":"…"}. Truncated up front so a re-run never appends to a stale log. |
--progress N | Print a progress line to stderr every N rows read (0 = off, the default). |
--json | Print the run summary as one JSON object on stdout (snake_case keys: zigbase_import, collection, dry_run, created, updated, failed, total, error_log). |
<file.ndjson> | Positional NDJSON path; - reads stdin. |
Exit codes. 0 — every row imported. 1 — the import failed outright (bad input, refused operation, DB error); batches committed before the failure persist. 3 — the import completed but skipped one or more rows under --continue-on-error: a lossy import is never reported as a success, so a driving script or agent must treat 3 as “needs a look,” not “done.”
Through-engine guarantees. Import performs the SAME full boot as serve (minus the socket): system migrations run, your comptime .collections are provisioned (so the target collection is guaranteed to exist), and the .encrypted field cipher is stamped onto the writer connection. Each row therefore gets: field validation and required checks, autodate defaults (created/updated), the .encrypted at-rest envelope (the same ZIGBASE_FIELD_KEY is required — plaintext is never written, and a missing key fails closed), and, for an auth collection, the credential transforms (password hashing, tokenKey generation, verified=false — a client-supplied verified is never trusted).
Streaming + batching + fail-fast (the default). Lines are read one at a time (the whole file is never slurped; a single record line must fit a 1 MiB buffer). Rows commit in batches of --batch-size. A bad row — malformed JSON, a validation failure, or a duplicate id — fails fast: the in-flight (uncommitted) batch is rolled back and the error names the offending 1-based line. Batches committed before the failure persist, so a large import that trips near the end is a resumable checkpoint (fix the file and re-run, ideally with --upsert-key). Blank lines are skipped. On completion it logs created, updated, total, and failed counts (and, with --json, the same numbers as one JSON object on stdout).
Migration-scale imports (--continue-on-error). One bad row out of half a million is not worth aborting a 40-minute import over. With --continue-on-error, each row runs inside its own SAVEPOINT: a failure rolls back only that row (not the batch it’s part of) and is recorded — to --error-log as NDJSON, and always in the final failed count — while every good row still commits. The exit code becomes 3 (not 0) whenever failed > 0, so a lossy import can never be mistaken for a clean one. --progress N prints a heartbeat every N rows so a long run isn’t a black box, and --dry-run rehearses the whole thing (full validation, defaults, encryption, auth transforms — everything except the final commit) so you can find every bad row before touching real data.
Id preservation (import-only). By default each row’s own id is preserved — essential for migrations, because relations reference records by id and an exported dataset must keep those ids intact. Seed the referenced rows (owners) first, then the rows that reference them. This id-honoring behavior is reachable only from import: the HTTP / custom-route / hook create path always generates a fresh id and ignores any client-supplied id, so a client can never choose a record’s id. (An imported id is sanity-checked — non-empty, ≤ 255 printable non-space bytes — and binds as a parameter, never interpolated.)
Library entrypoint. The same engine is exposed as zigbase.Import for a small consumer migration/seed binary that wants to load data programmatically:
const zigbase = @import("zigbase");
// ... obtain a booted App + writer (e.g. inside a custom CLI) ...
var reader = std.Io.Reader.fixed(ndjson_bytes); // or a file/stdin reader
const report = try zigbase.Import.run(app, writer, io, &reader, .{
.collection = "posts",
.upsert_key = "slug", // optional
.batch_size = 500,
.preserve_ids = true,
.continue_on_error = true, // optional: skip bad rows instead of aborting (see below)
});
std.log.info("imported {d} ({d} new, {d} updated, {d} skipped)", .{
report.total, report.created, report.updated, report.failed,
});
See examples/golfsim/seed/ for a worked seeding example (hosts + the simulators that reference them by preserved id).
Cross-backend migrations (SQLite and Postgres)
The same migration runs on whichever backend the server was started with (SQLite by default, or PostgreSQL when built with -Dpostgres and pointed at a postgres:// URL). *zigbase.Migrator makes that explicit — it is pass-the-dialect, not an SQL transpiler:
m.execLowered(sql)— run a curated statement written in the SQLite flavor; the dialect lowers the portable migration-level SQLite-isms (INTEGER→BIGINT,datetime('now')→a text now,INSERT OR IGNORE→ON CONFLICT DO NOTHING) on Postgres while leaving SQLite byte-identical. Covers most additive DDL + seeds.m.exec(sql)— run raw, backend-specific SQL verbatim. You own dialect correctness: SQLite-only SQL run on Postgres fails loud at startup (the database rejects it and the migration aborts — migrations are fail-fast).m.dialect.kind(.sqlite/.postgres),m.rawFor(.postgres, sql)/m.rawFor(.sqlite, sql)(run only on the matching backend), andm.requireBackend(.sqlite)(assert + fail loudly on the wrong backend) let a migration branch when a statement genuinely differs per backend.
ZigBase deliberately does not transpile SQL between dialects — it is fragile and silently mis-handles the edges that matter (collation, strftime, GLOB). A SQLite-only consumer that never builds with -Dpostgres keeps working unchanged. The framework’s own comptime-schema provisioning is fully cross-backend, so most apps need no raw migrations.
Postgres collation. Provisioned TEXT columns are pinned to COLLATE "C" so text ordering / keyset pagination matches SQLite’s BINARY byte order across backends. A comptime index marked .collation = .nocase is case-INSENSITIVE on both backends: SQLite uses COLLATE NOCASE, while Postgres (which has no built-in NOCASE collation) provisions a lower("col") functional index — a built-in, no citext/extension dependency. So a .nocase UNIQUE index rejects case-variant duplicates (Bob@x.com vs bob@x.com) on Postgres exactly as on SQLite (#159). Lookups and comparisons are case-insensitive on BOTH backends, so a .nocase column behaves identically everywhere: identity/email lookups (findByIdentity/findByEmail) and filter/rule equality (=/!=/in) against a .nocase column emit lower("col") = lower($1) on Postgres (the lower() functional index) and "col" COLLATE NOCASE = ?1 COLLATE NOCASE on SQLite (the COLLATE NOCASE index) — so a user registered as Bob@x.com can log in as bob@x.com on either backend (the uniqueness ⇔ lookup consistency holds). The built-in auth identity uniqueness (the partial unique index auto-created for each identityFields entry) is a plain CASE-SENSITIVE index on both backends; case-insensitive identity is opt-in by declaring a .nocase index on the field (the pattern the example apps use). One nuance: Postgres lower() is locale-aware (folds non-ASCII, e.g. É→é), whereas SQLite NOCASE folds ASCII A–Z only.
Migrating an existing SQLite instance to Postgres (migrate-db)
Once you have built a -Dpostgres binary, the migrate-db subcommand copies an existing SQLite-backed instance into a fresh PostgreSQL database — schema and data:
Stop the source and apply this version’s system migrations first. The copy requires the immutable storage namespace ledger and preserves its live and retired reservations; an older source missing the ledger is rejected before modifying the target.
# Build with the PostgreSQL backend compiled in.
zig build -Dpostgres=true
# Point it at your existing data.db and an empty target database.
zigbase migrate-db \
--from ./zb_data/data.db \
--to "postgres://user:pass@db.example.com:5432/zigbase?sslmode=require"
What it does:
- Provisions the equivalent schema on the target via the same code paths the server uses — it runs the system migrations, then creates each collection’s record table (with indexes and relation foreign keys) from the source’s
_collectionsmetadata. This includes collections you created at runtime via the admin UI, not just comptime ones. - Bulk-loads every row atomically. The entire data load runs in one transaction, so a mid-migration failure rolls the target back to a clean state (schema present, zero migrated rows) rather than leaving it half-populated. It preserves record ids, timestamps, and
_collectionsmetadata (collection ids survive verbatim). The target’s freshly-applied_migrationsledger is not overwritten, and server-generated columns (the Postgres_migrationsidentity PK and_events._seq) are regenerated. - Carries encrypted-field envelopes verbatim. Encrypted cells are backend-neutral
vN:<base64>TEXT envelopes;migrate-dbcopies the ciphertext byte-for-byte and never decrypts or re-encrypts, so you do not passZIGBASE_FIELD_KEYto the tool. The same key that read the SQLite data reads it on Postgres afterward. - Resets the analytics rollup watermarks. A rollup’s incremental watermark is an absolute
_eventssequence value; since_events._seqis regenerated on the target, the watermarks are cleared so each rollup recomputes from the migrated events on its next scheduled pass (a verbatim watermark would otherwise silently stop counting).
Safety:
- A non-empty target is refused (a target that already has a ZigBase schema) unless you pass
--force. With--forceevery target table is truncated and reloaded. - It reports per-table row counts measured on the target and fails loudly if any table’s target count does not match the source — rolling back the whole load.
- The command is present in every binary, but the PostgreSQL side requires
-Dpostgres; a stock binary fails with a clear error rather than silently doing nothing.
Not migrated: SQLite-only physical artifacts (the FTS5 full-text shadow tables) and the _rollup_<name> summary tables (rebuilt from the migrated _events after the watermark reset). The collection’s .searchable metadata is preserved, so the Postgres full-text index is provisioned the first time you zigbase serve against the migrated database.
9. Pluggable storage & mailer backends (.storage / .mailer)
.storage and .mailer each select a comptime plugin type. The defaults reproduce the built-in wiring:
.storagedefaults tozigbase’sDefaultStoragePlugin— local-disk storage rooted at<data_dir>/storage..mailerdefaults toDefaultMailerPlugin— config-driven with a fixed precedence: aCommandMailer(pipes the message to a local MTA’s stdin) whenZIGBASE_SENDMAIL_COMMANDis set, else anSmtpMailer(STARTTLS / implicit TLS / plaintext) whenZIGBASE_SMTP_HOSTis set, else aLogMailer(logs the email). Switching is config-driven; no code change is needed to upgrade from logging to a local sendmail/msmtp relay or real SMTP.
A plugin is a type with this uniform contract (built from the runtime zigbase.Config):
pub fn create(gpa: std.mem.Allocator, io: std.Io, cfg: zigbase.Config) !Self;
pub fn interface(self: *Self) zigbase.Storage; // or zigbase.Mailer — the type-erased vtable view
pub fn deinit(self: *Self) void; // release owned resources
create builds the backend from config; interface returns the type-erased vtable handle stored on the app; deinit tears it down (the instance outlives the server). Supply your own to back storage or mail with a different system. A custom mailer hands back a zigbase.Mailer view built from a static VTable whose send receives a zigbase.Email:
const AuditMailer = struct {
sent: usize = 0,
pub fn create(gpa: std.mem.Allocator, io: std.Io, cfg: zigbase.Config) !AuditMailer {
_ = gpa; _ = io; _ = cfg;
return .{};
}
pub fn interface(self: *AuditMailer) zigbase.Mailer {
return .{ .ptr = self, .vtable = &vtable };
}
pub fn deinit(self: *AuditMailer) void { _ = self; }
const vtable = zigbase.Mailer.VTable{ .send = send };
fn send(ptr: *anyopaque, io: std.Io, alloc: std.mem.Allocator, email: zigbase.Email) anyerror!void {
_ = io; _ = alloc;
const self: *AuditMailer = @ptrCast(@alignCast(ptr));
self.sent += 1;
std.log.info("to={s} subject={s}", .{ email.to, email.subject });
}
};
zigbase.App(.{ .mailer = AuditMailer }).runCli(init);
A custom storage plugin follows the same shape, returning a zigbase.Storage view from interface(). Its col argument is an immutable physical namespace, not necessarily the current collection name. Forward it unchanged; do not derive it from request URLs. Hooks and public routes still use the logical collection name. Inventory keys must use the same physical namespace. The zigbase.Storage vtable has four required methods — put / fetch / delete / deleteRecord — plus optional presignGetUrl and inventory (both default to null, so existing four-method backends stay valid) — so a custom backend wraps or replaces them. fetch(ctx, io, alloc, col, record_id, filename) returns a local filesystem path whose contents ARE the file, materializing it locally first if necessary (a remote backend spools to a local cache); null means the backend has no such object. (0.10.0: localPath(ctx, alloc, …) → fetch(ctx, io, alloc, …) — rename + one new parameter; return a local path, materializing the file locally if necessary; null = object missing.)
Record uploads call put before acquiring the database writer, using server-generated record IDs and freshly generated file names. A plugin must not assume the destination record already exists. The subsequent transaction runs hooks, validation, and access rules before committing references to those bytes; failed requests best-effort delete their uploads, including a PUT whose reply failed. Upload POST returns 409 when the collection definition changed during transfer. An upload PATCH rechecks its record snapshot under the write transaction (with a PostgreSQL row lock) and returns 409 if another write changed the record during upload, or if the collection definition changed. PATCH checks update access to the existing row before transfer and to the updated row before commit. Reload the record and retry a conflict. A crash between PUT and commit can still leave unreferenced bytes; this is not a distributed transaction.
Read-only inventory (opt-in CLI)
Build with -Dfile-inventory=true to include zigbase files inventory [--limit 1..1000] [--cursor KEY] [--data-dir PATH]. Add -Ds3=true for S3 storage and -Dpostgres=true when the database uses PostgreSQL. This is a binary-cost decision, not an HTTP policy: there is no inventory endpoint, and the tool requires the operator’s existing database/filesystem/S3 credentials. There is deliberately no duplicate App(.{}) runtime switch. A consumer dependency can select it with b.dependency("zigbase", .{ .target = target, .optimize = optimize, .@"file-inventory" = true }). The default build retains neither built-in inventory callbacks nor the CLI implementation. Existing custom storage vtables remain valid; without the optional capability the command reports InventoryUnsupported.
The command prints one JSON page with items (key, bytes, reference), nextCursor, hasNext, and usage (scope: "page", object/byte counts and candidate/unknown counts). Pass nextCursor unchanged to the next invocation, using the same backend, database, and key prefix. usage is not a global total; sum pages only when a best-effort live observation is sufficient. Object keys are relative to the configured local root or S3 key prefix; S3 credentials, signed URLs, file contents, and other record fields are never included in the report.
reference is referenced, candidate_unreferenced, or unknown (unexpected layout or metadata lookup failure). References include hidden file fields and expired TTL rows that still physically exist; normal record visibility does not determine blob ownership. Candidates include uploads between PUT and COMMIT, failed cleanup, and concurrent record/schema changes. No candidate is proven safe to delete. The command has no deletion mode, runs no migrations or provisioning, and opens SQLite read-only or PostgreSQL with read-only transactions. Builtin storage initialization performs no writes; custom plugin initialization and its inventory callback must honor the same read-only contract.
Inventory and reconciliation require this binary’s current collection metadata schema. After upgrading, run migrations explicitly before these commands. An incompatible or unreadable metadata schema fails the command with CollectionMetadataUnavailable and no JSON page, rather than silently reporting every file as unknown. Neither command applies migrations automatically.
Keys and cursors must be UTF-8; otherwise the command fails without a partial JSON page and directs operators to byte-safe backend-native tooling. This keeps their JSON types stable rather than emitting integer arrays for invalid bytes. Local inventory scans regular files in the three-level record-file layout and shallower paths. The operator-configured storage root may be a symlink (for an external volume); descendant symlinks are never followed. Unknown filesystem entry types are resolved with no-follow metadata reads. Each invocation rescans up to 100,000 directory entries, retaining at most limit + 1 keys in memory; exceeding this bound fails with InventoryScanLimit instead of returning a misleading partial total. Deeper directories and nonregular files are outside this scope. S3 inventory requests one ListObjectsV2 page with max-keys=limit, needs bucket listing permission, and observes only the configured prefix. It lists current objects, not old object versions, delete markers, or unfinished multipart uploads. Neither backend offers snapshot isolation across pages; concurrent writes can change the observation. Inventory does not change record download authorization.
Offline orphan reconciliation (opt-in CLI)
The same -Dfile-inventory=true build includes a separate maintenance command for the built-in local storage + SQLite combination only:
# Observe a bounded page without changing the database or storage.
zigbase files reconcile --data-dir ./zb_data --min-age-seconds 86400 --limit 100
# Stop all application instances and other writers before explicitly deleting.
zigbase files reconcile --data-dir ./zb_data --min-age-seconds 86400 --limit 100 --apply
Both modes refuse PostgreSQL, S3 configuration and custom storage plugins; they do not fall back to a local database or bucket copy. They require an existing database and storage root and run no migrations/provisioning. Dry-run opens SQLite read-only and creates no maintenance file. The default grace period is 86,400 seconds; --min-age-seconds accepts 1–31,536,000. --limit defaults to 100 and accepts 1–1,000. Each invocation uses the inventory scan bound of 100,000 directory entries and retains at most one bounded page.
Age is not proof that an upload has finished. Every booted application using built-in local storage holds a shared lease on the permanent storage/.zigbase-maintenance.lock file, including default builds without this CLI and serve --ignore-lock. This costs one open descriptor per booted app, with no per-request lock operation. Apply takes an exclusive, nonblocking lease using a writable descriptor; ordinary shared leases open existing lock files read-only. Exclusive apply retains write permission because some flock implementations require it. Neither mode replaces an existing lock inode. Apply holds its lease on that same actual storage root, then a SQLite writer transaction for the whole batch. An active app or database writer makes it refuse rather than wait. This covers an HTTP upload paused between its storage PUT and database commit; an age threshold alone would not. The lease also blocks new app boots until maintenance releases it. Root-directory symlink aliases share the lease; descendant symlinks are not followed. Never remove or replace the lock file: doing so breaks coordination between holders of different inodes.
The root must belong exclusively to this database. Stop older ZigBase binaries, direct Storage users, raw filesystem/SQL writers and other processes that do not participate in the lease protocol before apply, and keep them stopped until it finishes. The lease cannot detect or restrain those writers. Use a local filesystem with working advisory locks; this is not an online garbage collector, shared-bucket reconciler, or distributed maintenance protocol.
Only ordinary collection/record/filename files with known collection metadata can become candidates. Current physical references include hidden fields and expired TTL rows that have not been deleted. Missing/malformed collection metadata, encrypted file fields and unfamiliar layouts stay unknown; dropped collection prefixes are not automatically removed. Failed uploads and files left by non-HTTP mutations can be reclaimed when their collection remains known and no physical row refers to them. Immediately before each single-file unlink, apply checks the inode/size/timestamps, grace period and current references again under the writer transaction. It never recursively deletes a prefix.
Each invocation emits one versioned JSON object with mode, backend, items, minAgeSeconds, nextCursor, hasNext, and failures. Items include key, bytes, modifiedAt (Unix seconds), outcome and an optional failure error name. Outcomes are candidate, referenced, recent, unknown, changed, deleted, missing, or failed. Pass nextCursor unchanged with the same database/root/settings to inspect the next page. A cursor is only a pagination position, not a saved deletion approval or snapshot: apply recomputes its own observations, and changes between runs can change the results.
Exit 0 means the page completed without I/O failures, not that every file was deleted (unknown and protected files are retained). Exit 1 reports refusal or failure. All page keys and the output cursor must be UTF-8; invalid bytes reject the entire page before any deletion. Use byte-safe filesystem tooling to inspect such names. Per-file I/O failures appear in the JSON and processing continues within the bounded page. Filesystem deletions are irreversible and cannot be rolled back by SQLite. A later failure—including report output failure—can leave some files deleted; retain backups, inspect the report when available, and recompute a dry-run before retrying.
Durable HTTP file cleanup (opt-in)
Move HTTP record replacement/deletion storage requests to the existing durable queue engine:
const MyApp = zigbase.App(.{
.queues = .{ .cleanup = .{ .backend = .durable } },
.files = .{ .cleanup_queue = "cleanup" },
});
The named queue must be durable and drained by a worker (the implicit worker qualifies); invalid configuration fails compilation. Absent this setting, no cleanup handler or additional worker is registered and existing synchronous, best-effort cleanup remains unchanged. The reserved job kind is file_cleanup.
Removed file references and their cleanup jobs commit in the same database transaction. Enqueue failure aborts the mutation. Non-upload requests run existing rule/tenancy prechecks before taking their cleanup row lock. If the locked record differs from that snapshot, the request returns 409; reload and retry. Uploads retain their preflight and locked snapshot checks. These extra snapshots are skipped for collections without file fields. PATCH jobs delete individual objects and recheck physical references, including hidden fields and TTL-expired rows. DELETE jobs remove the whole record prefix only if the physical row is absent; this also removes unknown keys left by earlier schemas and local record directories. A reused record ID suppresses prefix deletion even when its row has expired. DELETE still queues cleanup when the current schema has no files. Local and S3 backends tolerate repeated deletion after a lost acknowledgement; custom backends must also make deleting an absent object or record prefix succeed. Storage failures follow the selected queue’s retry policy and eventually appear as failed jobs; monitor those jobs and retry them after repairing storage access. S3 record sweeps attempt every listed key and log each remote deletion failure before returning an error for retry. Local S3 spool-cache removal failures are logged but do not fail a job whose remote objects have been removed.
For safety, each deletion holds the SQLite writer or a PostgreSQL table write lock across the reference check and storage call. This removes remote I/O from the initiating request, but slow storage can still delay concurrent writes. PostgreSQL locks the data table before collection metadata, matching schema DDL, and revalidates the collection identity afterward. A transaction-local 250ms lock timeout makes contention retryable without changing the connection’s usual limit. Storage plugins must not acquire database connections or invoke record hooks from delete or deleteRecord, must propagate storage failures, and should bound their I/O timeouts. Use a low-concurrency cleanup worker and an appropriate visibility timeout. Stored object keys must remain immutable; custom out-of-band uploads must not overwrite keys pending deletion or write into a deleted record’s prefix without first creating its database row.
Collection identity is checked: deleted or recreated collections are conservatively skipped, leaving their objects for separate operator review. Coordinated renameCollection migrations relink pending jobs by stable ID while preserving their physical namespace; uncoordinated name changes are not supported. This first version covers HTTP record PATCH/DELETE only, not raw SQL, Data mutations, cascade/TTL deletes, collection deletion, failed upload cleanup, or orphan reconciliation. A crash before upload references commit can still orphan bytes. Inventory remains read-only and is not authority to delete objects.
Presigned-redirect serving (S3)
By default every download is proxied: the server calls fetch (spooling to a local cache for S3) and streams the bytes itself. With App(.{ .files = .{ .s3_presign_redirect = true } }) and the S3 backend active (-Ds3 + ZIGBASE_S3_*), an authorized download is instead answered with a 302 redirect to a time-limited presigned GET URL (s3_presign_ttl_s, default 900s), offloading the byte transfer to the object store. The optional presignGetUrl vtable method backs this; local disk and non--Ds3 builds return null and keep the unchanged proxy path. Authorization still runs per-request before the redirect is issued, but the issued URL is a bearer capability valid until it expires and not bound to the authorized requester — keep the TTL short. See KNOWN_LIMITATIONS.
const MyStorage = struct {
backend: zigbase.LocalStorage, // or your own (S3, etc.)
pub fn create(gpa: std.mem.Allocator, io: std.Io, cfg: zigbase.Config) !MyStorage {
_ = io;
const root = try std.fmt.allocPrint(gpa, "{s}/storage", .{cfg.data_dir});
return .{ .backend = zigbase.LocalStorage.init(root) };
}
pub fn interface(self: *MyStorage) zigbase.Storage {
return self.backend.storage(); // or build a Storage{ .ctx = self, .vtable = &vt }
}
pub fn deinit(self: *MyStorage) void { _ = self; }
};
zigbase.App(.{ .storage = MyStorage }).runCli(init);
The default zigbase.DefaultStoragePlugin wraps zigbase.LocalStorage; the zigbase.Mailer vtable — a single send(io, alloc, zigbase.Email) — backs mail (the default zigbase.DefaultMailerPlugin selects zigbase.LogMailer or zigbase.SmtpMailer from config). A plugin type that omits any of create/interface/deinit is a compile error with a contract-specific message. See examples/plugins/ for a full, compiling custom mailer (AuditMailer) and custom storage (AuditStorage, which wraps zigbase.LocalStorage and logs each of the four vtable calls before delegating).
S3-compatible storage (-Ds3)
DefaultStoragePlugin also backs an opt-in S3-compatible object storage backend (AWS S3, MinIO, Cloudflare R2, and similar) — selected by configuration alone, no code change:
- Build with
-Ds3=true. Off by default; a stock binary has no S3 code compiled in at all. - On an
-Ds3binary, setZIGBASE_S3_BUCKET(and credentials) to switchDefaultStoragePluginfromLocalStoragetoS3Storageat startup — the same “switch via configuration alone” contract asZIGBASE_DB_URL(postgres). A stock (-Ds3=false) binary handedZIGBASE_S3_BUCKETdoes not silently keep writing to local disk unnoticed — it logs a loud warning and falls back to local storage.
| Env var | Default | Purpose |
|---|---|---|
ZIGBASE_S3_BUCKET | "" (off) | non-empty enables S3 storage (on an -Ds3 binary) |
ZIGBASE_S3_REGION | "us-east-1" | AWS region (used for SigV4 signing and the default endpoint) |
ZIGBASE_S3_ENDPOINT | "" | "" → https://s3.<region>.amazonaws.com; set for MinIO/R2/other S3-compatible endpoints |
ZIGBASE_S3_ACCESS_KEY_ID | "" | SigV4 access key id (required) |
ZIGBASE_S3_SECRET_ACCESS_KEY | "" | SigV4 secret access key (required) |
ZIGBASE_S3_FORCE_PATH_STYLE | auto | true/1 forces path-style addressing; unset auto-selects path-style when ZIGBASE_S3_ENDPOINT is set, virtual-hosted otherwise |
ZIGBASE_S3_KEY_PREFIX | "" | prefix prepended to every object key (<prefix><col>/<rid>/<name>) — namespace multiple apps in one bucket |
ZIGBASE_S3_CACHE_DIR | "" | "" → <data_dir>/storage_cache; the local spool cache directory (see below) |
ZIGBASE_S3_CACHE_MAX_BYTES | 1073741824 (1 GiB) | spool cache size cap; eviction reclaims down to a 3/4 low-water mark |
ZIGBASE_S3_MULTIPART_THRESHOLD_BYTES | 67108864 (64 MiB) | upload objects at or above this size using multipart; range 5 MiB–5 GiB |
ZIGBASE_S3_MULTIPART_PART_BYTES | 8388608 (8 MiB) | preferred multipart part size, 5 MiB–5 GiB; increased automatically if needed to stay within 10,000 parts |
Large uploads use sequential multipart transfers. Parts borrow slices of the already resident input; HTTP response scratch is reclaimed between attempts and the completion manifest is bounded to 10,000 ETags. The HTTP request and Storage.put still hold the complete input in memory: this is neither streaming request ingestion nor resumable client uploads. Keep ZIGBASE_MAX_UPLOAD_SIZE within the deployment’s memory budget, even though the backend no longer imposes the single-PUT 5 GiB ceiling. AWS allows 5 MiB–5 GiB parts (the last may be smaller) and at most 10,000 parts, which makes its stated maximum object size 48.8 TiB — exactly 10,000 × 5 GiB, so the part limits and the object ceiling are one constraint, not two (AWS prose elsewhere restates the same ceiling in decimal units as “50 TB”). validateConfig therefore refuses a ZIGBASE_MAX_UPLOAD_SIZE above that ceiling at startup. Compatible providers may impose lower limits that ZigBase cannot know at boot; those surface as a provider error on upload. See AWS multipart limits.
Part uploads and transient completion failures retry once using the same upload ID and part numbers. Completion checks the XML response: HTTP 200 can still carry an error, as documented by AWS CompleteMultipartUpload. Initiation is not blindly retried because a lost response can leave an unknown upload ID. On subsequent failure the backend attempts to abort the known upload (also one retry). Grant multipart upload/abort permissions and configure the bucket’s incomplete-multipart lifecycle expiration: crashes, allocation/transport failures, or failed aborts can leave unfinished uploads. An ambiguous completion may have created an object even if its response was lost; orphan reconciliation remains necessary. Retry bounds are attempt counts, not wall-clock deadlines (the shared HTTP client’s socket-timeout limitation still applies).
Downloads are never proxied straight from S3: fetch spools an object to a local cache file on first read (an atomic download-then-rename, safe under concurrent misses) and every subsequent read is served from that local file — so GET /api/files/:col/:rec/:name gets Range requests, conditional requests, ETag, and per-collection cacheability exactly as with local storage (§9’s fetch contract). Startup runs a fail-fast HeadObject probe (200 or 404 both prove DNS/TLS/SigV4/bucket/permissions end-to-end; anything else refuses to start) so a bad S3 config is caught at boot, not on first upload. Record cleanup follows every ListObjectsV2 continuation page, decodes XML key text, and refuses malformed or out-of-prefix listings. See Known limitations for best-effort cleanup, crash orphans and presigned-URL trade-offs.
zigbase.S3Storage is exported alongside zigbase.LocalStorage (an empty placeholder type on a stock, non--Ds3 build — its create/interface methods only exist when compiled in), so a custom storage plugin built on an -Ds3 binary can wrap or delegate to it the same way the example above wraps LocalStorage.
10. Footprint levers (.pools)
.pools tunes ZigBase’s memory/connection footprint at comptime. All fields are optional; without a resource profile each defaults to the historical value:
| Field | Default | Meaning |
|---|---|---|
.readers | 16 | requested retained idle-reader cap, currently clamped to 16 by both backends; not a bound on active overflow connections. |
.jobs | 2 | scheduler worker-pool size. |
.memory_jobs | 4 | lazy worker count shared by memory queues and app.submit; positive 1..64, separate from scheduled jobs. |
.stack_size | 1 MiB | per-thread stack for scheduler/job/submit threads (vs std.Thread’s 16 MiB default). Clamped up to a safe floor — the lever can only raise the stack, e.g. for unusually deep job handlers. |
.cache_kib | 1024 | SQLite per-connection soft page-cache target (KiB), including overflow readers; 0 preserves SQLite’s default, larger values clamp to i32 max. |
zigbase.App(.{
.pools = .{ .readers = 4, .jobs = 2, .memory_jobs = 1, .stack_size = 2 << 20, .cache_kib = 256 },
}).runCli(init);
.pools = .{ .jobs = N } is the ONE lever for the scheduler worker-pool size. The legacy .jobs = .{ .pool_size = N } spelling was removed; it is now a compile error naming the replacement.
The memory-job pool allocates its worker-handle slice and starts threads only on the first enqueue or app.submit; unused pools allocate neither. .memory_jobs sets its requested worker count, while .stack_size applies to these threads as well as the scheduler. The queue ring stays at 256 tasks and rejects overflow with QueueFull rather than waiting. Handler allocations and queued payload sizes are not bounded by worker or stack settings. Lower counts save requested thread stacks but can increase queueing; retry backoff occupies a worker. Shutdown drains work and joins the actual started threads. If startup spawns only some requested workers, that subset continues serving the pool until shutdown; if none start, the push fails and a subsequent push may retry.
Resource profiles and inspection
HTTP admission and backpressure
Opt into .admission = .{ .max_requests = 3 } to cap concurrent synchronous HTTP callbacks per application process. This is independent of resource profiles and does not change the transport’s four HTTP threads. Choose 1–3 to leave a thread available to reject excess work; a limit of 4 or higher cannot shed work with this transport configuration. Excess requests receive 503, code overloaded, and Retry-After: 1 (HEAD has no body). There is no additional waiting queue. Clients should use bounded retries with backoff/jitter; do not blindly retry non-idempotent operations after ambiguous transport failures.
Admission happens before ZigBase’s request arena, multipart parsing, and auth. A permit is released after synchronous response handling and teardown, including errors and file/HEAD responses. Bytes may reach a client before that teardown finishes, so even an immediate sequential request can briefly receive 503 at a limit of one. Rejections are access-logged without a request allocation.
The exact built-in GET /api/health liveness probe is exempt and ignores request bodies, preventing intentional load shedding from triggering an orchestrator restart. Other methods (including HEAD), nearby paths, and diagnostics are not exempt. This is not a dedicated health worker: transport/thread saturation can still delay probes. Counter synchronization parks contending threads instead of busy-spinning, including on single-vCPU deployments.
GET /api/admission/stats requires a superuser and returns limit, active, high_water, and rejected (saturating u64). Snapshots are coherent and process-local, reset on restart; active includes the diagnostics request itself. With byte-only admission, limit is null and HTTP counters stay zero instead. Additive shared-budget fields are work_limit (nullable), jobs, work_high_water, and jobs_rejected (saturating u64); these job counters remain zero when shared admission is disabled. Diagnostics obey admission too and can receive 503. Application hooks may inspect ctx.app.admission, an optional borrowed state pointer, and call snapshot() without HTTP or database access. Do not mutate/release framework-owned permits. When omitted, the diagnostics route is absent and no state, atomics or per-request checks are retained (only the app’s null pointer). The plugins example enables it.
This is not a total RSS cap or protection against buffered request bodies: facil.io receives/buffers transport data before invoking ZigBase. Reverse-proxy connection/body/time limits remain necessary. WebSocket upgrades, long-lived WS/SSE connections, asynchronous file transmission and background jobs are outside this limit (an SSE setup callback, if routed through HTTP, only holds a permit during setup). It neither cancels slow accepted handlers nor limits durable queue depth. Existing memory-job ring (256 queued tasks, configurable workers, QueueFull on overflow) and scheduler bounds remain independent and unchanged.
Coordinated HTTP and job admission
Build with -Dcoordinated-admission=true and configure both ceilings:
.admission = .{ .max_requests = 3, .max_work = 16 },
For an embedded consumer, forward .@"coordinated-admission" = true in the b.dependency("zigbase", .{ ... }) options. The flag compiles the integration; the optional positive u32 max_work activates its per-application ceiling. max_requests remains independently required. Omitting the flag excludes shared accounting code; specifying max_work without it is a compile error.
One work permit covers each admitted synchronous HTTP callback, each queued or running memory job, each app.submit task, and each serial durable poll batch. A memory job retains its permit through retry backoff and final cleanup. HTTP still obeys max_requests, while HTTP plus outstanding memory work and active durable batches must also fit max_work. Full shared capacity returns HTTP 503 overloaded with Retry-After: 1, or error.QueueFull from a memory-job enqueue/submit. Durable polls defer claiming until a later tick. No waiting queue is added and work is never evicted.
An HTTP handler that enqueues a memory job needs a second permit while it still owns its HTTP permit. Size the shared ceiling for that overlap and handle enqueue failure; do not spin or wait for capacity from inside a handler. Jobs may use all shared permits: there is no reserved HTTP capacity or fairness guarantee. The exact GET /api/health exemption still applies, but diagnostics can be rejected.
work_limit reports the configured ceiling; jobs counts outstanding admitted memory work plus active serial durable batches; work_high_water is the peak of active + jobs. jobs_rejected counts memory-job refusals and deferred durable poll batches, not independent ring-full refusals or lost durable jobs. Reservation precedes the retained payload/name copy and is returned on allocation or enqueue failure. The inline memory-job fallback also holds a permit through execution. Ctx payload serialization occurs before this reservation.
A durable poller reserves one permit before any claim, and holds it through its serial batch and cleanup (not one permit per prefetched row). Saturation leaves jobs pending without spending attempts or rate tokens; that poll also skips its reclaim pass, while the independent scheduled GC/reclaim sweep remains active. Polls with no available work briefly acquire/release a permit. Durable claimed payload allocations are not charged to max_job_bytes.
This is a process-local work-count limit, not a byte or RSS cap. Other scheduler work itself, transport buffers, long-lived realtime sessions and arbitrary plugin work are not counted. Existing ring/worker bounds still apply. The offline resources envelope exposes coordinated_admission_max_work, or null when not configured; it does not measure current occupancy.
Retained memory-job byte budget
Under the same -Dcoordinated-admission=true build gate, opt into a separate positive usize ceiling with .admission.max_job_bytes:
.admission = .{ .max_job_bytes = 64 * 1024 },
This charges exactly the lengths of the queue-owned payload copies for memory jobs and name copies for app.submit, across queued and running tasks. Work-count and byte reservations are atomic: either both fit or neither is acquired. Reservations precede copying; allocation, startup, ring-full and shutdown errors return them. Retries retain bytes until final task cleanup. The inline no-pool fallback borrows its payload and therefore charges zero bytes, while retaining any configured work-count permit. Empty payloads/names also charge zero bytes.
Byte exhaustion returns error.QueueFull, including a single copy larger than the entire budget. This byte-only configuration neither counts nor caps HTTP: HTTP admission calls compile out, while the superuser diagnostics route remains. Its limit is null, and active, high_water, and rejected remain zero. Embedder migration: admission.Config.max_requests (exposed through App.admission_config) and admission.Snapshot.limit (returned by Runtime.admission.snapshot()) changed from u32 to ?u32. Use optional capture (if (snapshot.limit) |limit| { ... }) and handle null as no HTTP cap. Existing HTTP-capped integrations may use .? only when their configuration guarantees a limit. Do not substitute zero for null: zero is not a valid configured cap. Omit .max_requests to keep byte-only admission; omitting .admission disables all its budgets. Add .max_requests = 3 to cap HTTP separately, and optionally .max_work = 16 to share work-count capacity; max_work still requires max_requests because it coordinates HTTP and job counts. Without max_work, the bounded ring/workers remain the task-count protection. When both job ceilings are full, the work-count check runs first and records the rejection only in jobs_rejected.
Diagnostics add job_bytes_limit (nullable), job_bytes, job_bytes_high_water and saturating job_bytes_rejected. The last counter records only byte-budget refusals, not independent ring-full errors. Without the byte ceiling its counters remain zero; without max_work, jobs and work_high_water remain zero even when bytes are tracked. resources.envelope.coordinated_admission_max_job_bytes reports the configured ceiling, not current occupancy.
This is not an RSS or total job-memory bound: it excludes task headers, allocator overhead, handler arenas/allocations, stacks, caller-owned data, durable queues and transport buffers. Ctx serializes payloads before admission; that serialization is not covered. No new per-task accounting field or allocation is added. The build flag off excludes queue accounting calls; an omitted byte ceiling leaves byte counters untouched and adds no separate lock to work-count admission. HTTP-only configuration performs no job-accounting lock. Byte-only zero-byte reservations (including inline borrowed payloads) also take no admission lock.
Bounded query workbench (opt-in)
Build with -Dquery-workbench=true to measure synchronous SQLite and PostgreSQL prepared statements inside matched built-in and consumer HTTP route handlers and declared durable/scheduled job handlers. The default build compiles out measurement calls, statement fields, thread-local attribution, counter storage and inspector routes. This is a diagnostic build cost, not an always-on logging feature. Enabled instrumentation has timestamp, fingerprinting and synchronization overhead; it is not zero-cost when enabled.
pub const App = zigbase.App(.{
.query_workbench = .{ .max_entries = 64, .slow_ms = 100 },
});
These are also the standalone defaults when the build flag is on. Configuration is comptime: 1–256 entries and a positive slow_ms; unknown keys fail compilation. Each entry aggregates attribution kind + backend + method + route template (or declared job name) + structural query shape. Only the first configured number of distinct entries is retained until restart; new keys are dropped (counted), never allowed to grow a map. Templates longer than 192 bytes and unsupported or greater-than-16-KiB statements are omitted. Only the first 32 distinct shapes per request are tracked for repetition. This can undercount repeats after saturation, but never expands storage. Counters saturate instead of wrapping. No telemetry is persisted or exported automatically.
No raw SQL, parameter values, SQL literals, identifier text or request-path values are retained. A bounded lexer hashes only a fixed SQL-keyword allowlist and punctuation: identifiers become one marker, literal/bind contents another, comments disappear. Named binds ($, :, @, #, including SQLite Tcl suffix syntax) are omitted entirely for SQLite; anonymous/numbered ? binds are supported. PostgreSQL supports numbered $N binds and :: casts. Dollar-quoted strings, escape/Unicode string prefixes, backslashes in quoted tokens and nested block comments are conservatively omitted rather than risking literal-content capture. Thus different tables/columns and parameter values can share one opaque shape ID. repeatedShapes means repeated structural shapes within one matched handler, not proof of an N+1 query bug. Route templates themselves are operator-supplied code metadata and are visible to the inspector; do not put secrets in route definitions. Reports are operator-only, not tenant-scoped.
GET /api/query-workbench/stats requires a current superuser bearer token. Cookies alone do not authorize it. It returns bounded {items} plus limits, backend scope, threshold and dropped-execution count. Each item has attribution (http, scheduled_job, or durable_job), backend, measurement, method, routeTemplate, opaque hexadecimal shape, executions, stepNanoseconds, maxStepNanoseconds, slowExecutions, repeatedShapes, and failedExecutions. For job SQL items, method is JOB, routeTemplate is null, and jobName holds the declared job name. Timing is the sum of SQLite step() call durations per execution, including SQLite busy wait, but excluding preparation, binding, pool wait, row decoding, application processing and response transmission. Completion, error, reset or early finalization closes an execution; work retained past the originating scope is omitted, not attributed to the next request. Inspector requests exclude their own auth and plan queries. State snapshots are coherent and bounded.
For PostgreSQL, those same stepNanoseconds fields measure client elapsed time for the first step’s entire extended-protocol exchange, including parameter marshalling, network/server waits and result materialization. Buffered subsequent steps add no timestamps or execution counts. This is not server CPU or pool-wait time. Zero-row results, errors and partial consumption still count one exchange; reset permits another execution. The report’s top-level measurement is backend-specific-see-items; item labels distinguish prepared-statement-step-time from client-extended-protocol-exchange-time. Top-level backend/activeBackend identify the application pool; individual items can also describe another backend opened by a custom handler. Backend identity also separates repeat accounting. Telemetry storage is bounded; PostgreSQL’s pre-existing whole-result buffering is not bounded by these telemetry limits, and instrumentation copies no rows.
The additive routes array measures completed synchronous matched-dispatch scopes, including routes that execute no SQL. Each method/template aggregate has completedScopes, totalNanoseconds, maxNanoseconds, and slowScopes (elapsed at least slowMilliseconds). routeMeasurement is matched-handler-scope. maxScopeEntries equals the configured max_entries: HTTP routes and jobs share this scope table, independently of the query-shape table. maxRouteEntries remains an upper bound on HTTP routes, not a separate reservation. droppedRouteScopes and droppedJobScopes count omitted HTTP and job completions separately. Each table retains its first keys until restart; a full table still updates known keys. Route labels over 192 bytes are omitted. Storage adds at most max_entries fixed-size scope aggregates, including six status counters, one error counter, and four three-counter pool-wait buckets per entry; capture allocates nothing per completed scope. HTTP and job SQL shapes also share the existing query-entry cap.
The awake-clock interval starts after route matching and ends as synchronous dispatch unwinds, before taking the aggregation lock. It includes in-scope auth, guards, pool waits, handler processing, and deferred cleanup. Consumer dispatch also includes its own error conversion; built-in dispatch ends before the outer server error backstop. It excludes parsing/routing and admission before dispatch, response transmission, streaming connection lifetime, and detached background work. Failed or denied matched handlers count. responseStatusClasses classifies returned responses as informational, success, redirection, clientError, serverError, or other. handlerErrors counts errors escaping the measured handler separately: it does not assume the status an outer error handler will eventually send. A consumer handler error mapped to a response inside dispatch increments both the error counter and the actual returned status class. Inspector scopes read no timing clock and do not enter either table.
This is elapsed time, not CPU time or end-to-end request latency. Two additional clock reads and a bounded locked update occur per measured scope. Nested scopes are inclusive; do not subtract aggregated query/lifecycle totals to infer exclusive application time, and do not add route totals to obtain process busy time. Concurrent requests overlap. Active/incomplete scopes are absent until they finish.
Each completed scope also reports four fixed poolWaits buckets: SQLite/PostgreSQL reader/writer mutex acquisition counts, total nanoseconds, and maximum nanoseconds. poolWaitMeasurement is pool-mutex-acquisition. The timer brackets acquiring the pool mutex (including a reader free-list spin lock), not the whole connection acquisition operation. Connection creation, PostgreSQL handshake, database locks, SQL execution, and time holding a connection are excluded. Uncontended lock calls also count; a nonzero duration alone does not prove contention. No pool/store pointer is retained on a returned connection, and out-of-scope acquisitions are unmeasured.
The jobs array reports declared durable and scheduled handler attempts, including SQL-free handlers, under jobMeasurement: "handler-attempt-scope". Each aggregate identifies attribution (scheduled_job or durable_job) and jobName. An attempt’s handlerErrors does not mean the logical job exhausted its retries. Handler duration excludes queue residence, scheduler delay, claim/acknowledgment bookkeeping, and retry backoff. Compiled handler names are metadata visible to operators; payloads, tenant IDs, and dynamic submitted names are never retained. Memory-queue handlers and app.submit are not measured in this chapter. Inline memory execution masks the calling HTTP scope so its SQL cannot be mistaken for the request’s SQL.
Completed statements also report a separate lifecycle family:
finalizedStatements: successful prepares finalized in their originating scope, whether stepped or not. Reset/reuse still counts as one statement.statementLifetimeNanosecondsandmaxStatementLifetimeNanoseconds: sum/max elapsed time from entry to prepare through return from finalize (PostgreSQL prepare copies SQL locally; server parsing happens in the first step).measuredCallNanoseconds: elapsed time in the measured prepare, step, reset, and finalize calls for those finalized statements, including backend waits. PostgreSQL measures only the execution-bearing step; buffered row handling remains in held time.heldNanoseconds: lifetime minus those measured calls, clamped at zero. It includes application pauses, scheduling, binding, column access/decoding and instrumentation overhead. It is not CPU time, a pure application-time measurement, or proof that a connection is being misused.
The response identifies lifetimeMeasurement: "prepare-through-finalize" and measuredCalls: ["prepare", "step", "reset", "finalize"]. Existing step-time slow/repeat/failure counters keep their meanings; they are not lifecycle counts. A statement can have several executions, or none, and execution statistics can appear before its lifecycle is finalized. Do not subtract aggregate step timing from lifecycle timing: the populations can differ. Failed prepares, raw exec, pool acquisition and full request latency are not captured by statement metrics. A lifecycle whose measured calls or finalization leave the original scope is omitted rather than reattributed; no request/store pointers are retained in statements. Long-lived unfinalized statements do not appear in lifecycle aggregates.
Lifecycle keys share the existing entry capacity. droppedStatements counts finalizations omitted for unsupported SQL or full/oversized keys, independently of droppedExecutions; cross-scope omissions cannot safely update the departed scope and are not counted. Five extra saturating u64 counters per retained entry, one store-wide dropped counter, and fixed per-statement timestamp/identity state are the additional retained cost when enabled. Prepare/reset/finalize add clock reads, and each in-scope finalization performs one mutex-guarded store update with a bounded entry scan (at most max_entries). Step timing reuses existing reads, without extra per-row clocks. Successful in-scope prepares fingerprint the compiled SQL once; execution/reset cycles reuse that key. Statements prepared outside the active scope fall back to fingerprinting on execution, preserving execution attribution. For PostgreSQL, that fallback applies only to originally unscoped statements. A statement prepared inside a collecting scope loses execution attribution permanently when used in another scope (or outside a scope); reset or returning to its original scope does not revive it. These omissions do not increment the next scope’s dropped or execution counters. Default-off builds retain none of this storage or instrumentation.
POST /api/query-workbench/explain accepts only a structured SELECT shape:
{"collection":"posts","equalityField":"id","orderField":"created","descending":true}
Only collection is required. Field names must exist in the current collection schema (or be system id/created/updated fields). The generated shape selects id, optionally compares one field to a NULL parameter, optionally orders one field, and uses LIMIT 100. It runs EXPLAIN QUERY PLAN, never the SELECT or EXPLAIN ANALYZE; no caller SQL, functions, expressions, values or joins are accepted. Request body is capped at 4 KiB, output at 32 rows and 512 valid UTF-8 bytes per detail, with explicit truncated. Schema/index names are deliberately visible to authenticated operators. This is NOT a captured-query plan or a simulation of access rules, tenant predicates, cursor/expansion queries, or value-dependent optimizer choices. Use it to inspect simple index/scan choices, not assert a real request’s exact plan or cost. The same bearer-only superuser boundary applies before any planning. Inspect only trusted local/staging data when exposing schema names would be sensitive.
PostgreSQL plan inspection still returns 501; only SQLite supports plans. The raw exec() paths, memory jobs, app.submit, async work, WebSocket/SSE delivery, unmatched/static routes and remapped feature-state dispatch are not covered. Nested synchronous dispatch restores outer attribution; cross-thread/cross-scope statement retention does not transfer attribution. This first slice is not a general SQL profiler, an automatic index adviser or a distributed tracing system. /api/meta advertises capabilities.queryWorkbench and the optional stats URL; compiled route discovery includes both inspector endpoints. The runnable fixtures/query-workbench exercises repeated bound queries and a deliberately held/reset statement: /held demonstrates why lifecycle and step durations answer different questions without retaining SQL or parameter text.
Profile defaults
Select .resource_profile = .minimal, .balanced, or .throughput as a comptime starting point. Every explicitly supplied .pools field wins over the profile; omitted fields inherit it. Omit the profile to keep historical defaults.
| Profile | Reader cap | Scheduler workers | Memory-job workers | SQLite cache KiB per connection |
|---|---|---|---|---|
minimal | 2 | 1 | 1 | 256 |
balanced | 16 | 2 | 4 | 1024 |
throughput | 64 | 8 | 8 | 4096 |
All profiles keep the 1 MiB stack default and its existing safety floor. These are explicit starting points, not benchmark-derived optimal settings, automatic CPU detection, or process memory caps. More concurrency can hurt a workload; measure before adopting throughput. SQLite caches apply only to SQLite connections; they do not tune PostgreSQL’s server cache. Reader counts request idle-retention caps, not total connection limits or eager allocation. Scheduler settings apply when scheduled jobs exist; memory-job settings apply lazily even without scheduled jobs. Durable queue-worker batches remain controlled by .workers, and HTTP server concurrency is unchanged.
const Backend = zigbase.App(.{
.resource_profile = .minimal,
.pools = .{ .readers = 4 }, // jobs=1, memory_jobs=1, cache_kib=256 remain inherited
});
// Backend.resource_report is a typed zigbase.ResourceReport constant.
Run the resulting binary with resources (or resources --json) for a single JSON object, schema_version: 1, containing its compiled pool settings, profile (null when omitted), scheduler/admin state, and PostgreSQL/S3/file-inventory build gates. The command does not load or validate server configuration or open a database; common CLI logging initialization still reads log-format/level variables. It includes no secret values. This is a compiled resource report, not a complete live configuration dump: it does not identify the runtime-selected storage/database backend, report current allocations, or enumerate every optional subsystem. memory_job_workers describes the requested lazy pool even when scheduler_enabled is false; job_stack_bytes is the shared effective stack size after the 1 MiB floor. The worker count is null when an older saved report omitted it, not zero workers. Actual worker count can be lower after partial thread-spawn failure; startup warns with the actual/requested counts. Stack sizes are requests to the platform, not resident-memory or total-process ceilings.
Additive resource fields retain schema_version: 1. Compatibility is backward: a newer advisor can read older reports using defaults for fields they lacked (for example, the historical realtime connection cap is 10,000). This is not a forward-compatibility promise: an older strict advisor may reject newer fields. Use an advisor that recognizes every captured resource field, normally one from the same or a newer ZigBase release.
The additive envelope object explains independent configured quantities, not RSS predictions or a hard total-memory bound. Older saved tune measurement documents without this field remain accepted (their envelope is null).
| Envelope field | Interpretation |
|---|---|
http_admission_max_requests | Concurrent admitted synchronous callbacks; null means admission disabled. Transport buffering happens before admission. |
coordinated_admission_max_work | Shared count of admitted synchronous HTTP callbacks plus queued/running memory jobs and app.submit tasks, including retry backoff; null when disabled. Not a byte/RSS budget and excludes durable jobs. |
coordinated_admission_max_job_bytes | Ceiling for queue-owned memory-job payload and app.submit name copy lengths; null when disabled. Excludes inline borrowed payloads, pre-enqueue serialization, task/allocator overhead, and handler allocations. |
http_body_limit_source | runtime_ZIGBASE_MAX_UPLOAD_SIZE: the body limit is deployment-configured and deliberately not read by this offline command. |
retained_reader_cap | Actual retained idle-reader cap after the shared SQLite/PostgreSQL clamp. The existing top-level reader_pool_cap remains the requested value. |
sqlite_cache_target_bytes_per_connection | Soft SQLite page-cache target after its clamp; null when 0 preserves the engine default. Not applicable to a PostgreSQL deployment. |
sqlite_writer_and_retained_readers_cache_target_bytes | Sum of those soft targets for one writer and a full idle pool; not a bound on total connections, cache memory or RSS. Overflow readers have additional caches. |
scheduler_stack_bytes | (workers + tick thread) × effective stack bytes after the 1 MiB floor when jobs exist, 0 otherwise. Overflow is a compile error. Virtual stack space, not resident memory; excludes named queue workers and other threads. |
memory_job_stack_bytes | Up to memory_job_workers × job_stack_bytes of virtual stack space for lazy memory-job/submit workers, using the shared effective stack floor. Overflow is a compile error. Not current allocation or RSS: unused pools start no threads, and partial startup can produce fewer workers. Independent of scheduler enablement. |
resumable | Compiled session, single-upload, total staged-payload and chunk ceilings, or null when excluded from the binary. No promise about commit copies, metadata or storage usage. |
exclusions | Costs outside these quantities: overflow connections/non-cache database memory, transport/realtime, request/response/commit allocations, queue workers/plugins/runtime/allocator overhead, and resumable metadata/storage. |
For example, minimal yields two retained readers and a 256 KiB soft cache target per SQLite connection: 768 KiB for the writer plus a full idle pool. throughput requests 64 readers but currently retains only 16 on either backend; its SQLite retained-cache target is 68 MiB, not 65 × 4 MiB and not a process memory cap. These quantities must not be summed into a total: measure actual peak RSS under representative workloads with tune instead. Reporting adds no server hot-path instrumentation and never changes admission, pool, stack or staging behavior.
Profiles neither enable nor disable features. Keep using explicit comptime gates (for example .admin = .disabled) and build flags to exclude unwanted code; deployment environment variables retain their existing roles and do not override these comptime pool settings. Embedded consumers can use Backend.resource_report without compiling CLI reporting into their application.
Offline measurement advisor (zigbase tune)
zigbase tune --input measurements.json [--as-of UNIX_SECONDS] [--json] compares supplied measurements, not predicted performance. It emits one JSON object, never starts workloads or edits configuration, and opens only the named input file (maximum 1 MiB). Common CLI logging initialization still reads log format/level variables. -Ddev-tools=false removes the advisor. JSON is the default output; --as-of makes freshness checks reproducible, otherwise the real wall clock is used (not the application test clock).
Input schema_version: 1 requires workload, revision, environment, max_age_seconds, memory_budget_bytes, p95_budget_ms, and candidates. There must be 1–128 uniquely named candidates. Each candidate requires:
id,workload,revision,environment, andmeasured_at_unix;- positive finite
throughput_rps,p95_ms, and positivepeak_rss_bytes; failed_requests(any unsuccessful or semantically incorrect responses);resources: the object emitted by that measured binary’sresourcescommand.
Identifiers are nonempty, at most 256 bytes, without control characters. Unknown fields, invalid/nonfinite metrics or budgets, duplicate IDs, and unsupported document/resource versions are errors (nonzero exit). Context labels and the resource report are caller-supplied provenance, not independently verified build or hardware identities; capture them from your actual experiment. The advisor does not validate or synthesize deployable settings from the report. Unknown-field rejection also applies inside resources; the version number alone does not guarantee an older advisor understands a newer report.
The output repeats the context/budgets and includes items, each containing the candidate and its first exclusion reason, in priority order: failed_requests, context_mismatch, future_measurement, stale, memory_budget, latency_budget, or feasible. Freshness and budget boundaries are inclusive. Among feasible observations, recommendation is the ID with highest throughput, then lowest peak RSS, then lowest p95, then lexically smallest ID. No feasible candidate produces recommendation: null with exit 0—not a made-up suggestion.
Match workload data/response assertions, request count, concurrency, CPU limits, machine, toolchain, optimization, and all non-profile options across candidates. Use the environment label to identify that controlled setup. CPU is a controlled experiment input here, not an automatically detected/enforced budget. Repeat runs, inspect variability, and retain headroom: this first advisor does not aggregate trials or infer confidence intervals. Measured RSS is not a guaranteed production maximum. Allocator peak_live from zig build bench is not RSS; neither p95 nor median latency can substitute for measured wall-clock throughput.
Reproducible local measurement example
The Linux-only tools/tuning/measure.py helper starts each explicitly supplied binary against fresh temporary SQLite state on loopback, clears ZIGBASE_* environment variables, warms up ten requests, then measures a concurrent /api/health workload. Every response must be HTTP 200 with status: ok and backend: sqlite; any error aborts collection with no comparison document. It reports throughput from completed requests / load wall time, nearest-rank p95, and the server process’s /proc RSS high-water mark (including startup). It stops its children and removes temporary state. Run only trusted test binaries: custom startup hooks can have side effects beyond their data directory.
zig build -Doptimize=ReleaseFast
python3 tools/tuning/measure.py \
--candidate stock=./zig-out/bin/zigbase \
--revision YOUR_SOURCE_REVISION \
--environment linux-machineA-zig016-releasefast-unrestricted \
--requests 1000 --concurrency 4 \
--memory-budget-bytes 134217728 --p95-budget-ms 20 > measurements.json
./zig-out/bin/zigbase tune --input measurements.json
To compare profiles, build the same consumer revision separately with .resource_profile = .minimal and .throughput, keeping all other options identical, then supply both --candidate minimal=/path/to/minimal and --candidate throughput=/path/to/throughput. The helper collects each sequentially. The health workload is an executable smoke example, not evidence about database, job, or application throughput; adapt the collection harness and response checks to representative seeded application traffic before choosing production settings. Use an external OS/container CPU limit consistently when CPU capacity is part of the experiment. The helper and advisor do not apply that limit themselves.
10b. Pagination (.pagination)
.pagination chooses, at comptime, which list-pagination modes the records list endpoint exposes and which cursor token format it mints. All fields are optional; the stock binary behaves as .{ .offset = true, .cursor = true, .cursor_token = .stateless }.
zigbase.App(.{
.pagination = .{
.offset = true, // page/perPage offset paging (default true)
.cursor = true, // cursor (keyset) paging (default true)
.cursor_token = .stateless, // .stateless | .signed | .stateful (default .stateless)
},
}).runCli(init);
.offset = false— requests withpage/perPageget a 400; clients must walkcursor..cursor = false— requests withcursorget a 400; only offset paging is allowed.- Both
false— a@compileError(a list endpoint must have at least one mode).
The .cursor_token selector (security/statefulness tradeoffs):
| Value | Token | Tamper-evident | State | Use when |
|---|---|---|---|---|
.stateless (default) | base64url JSON payload, validated against the request’s sort/filter | No (rules + parameterized binding secure it) | None | Default; CDN-friendly; SDK-byte-compatible. |
.signed | stateless payload + HMAC-SHA256 keyed by the server JWT secret | Yes (400 on bad MAC) | None | You want tamper-evidence with no extra storage. |
.stateful | random opaque id; payload stored in _cursorStates with a TTL (GC’d) | N/A | A row per cursor | You want server-controlled validity/expiry (410 on expired). |
See the API reference for the request/response shape. The blog example uses .cursor_token = .signed; golfsim shows the explicit stateless default.
11. Errors + Sentry
.onError = handleError, // fn (ev: *zigbase.ErrorEvent) void
When the framework catches an error, your onError handler (if any) runs first, then a built-in backstop reports the error to Sentry when ZIGBASE_SENTRY_DSN is set, otherwise logs a structured line ([phase] err_name: message). ErrorEvent carries app, ctx (optional), err, phase (request / before_hook / after_hook / cron / job / file_serve / webhook / app), and message. The backstop never propagates.
This includes a built-in handler’s own failures, not just your routes’: an error escaping the built-in route table, the feature-state route, or custom-route dispatch itself (pool acquisition, authenticate) is logged and delivered to onError exactly like a consumer-route error, instead of surfacing as a silent, unexplained 500. Because no principal has been resolved yet at that layer (built-in handlers authenticate internally), the RequestContext on ev.ctx carries only method — every other field is its zero value. A static-file read failure is the one exception: it is logged at warn and still answers 404, since a browser-facing static miss is not a server incident.
Report your own errors: ctx.reportError. Route a swallowed-but-notable error from your own code — a hook, job, cron, or handler — through the same backstop the framework uses for its own caught errors:
ctx.reportError(err, "sync of order {s} failed", .{order_id});
It runs your onError handler first, then the configured reporter (Sentry or the log backstop), carrying the same TTL dedup and non-blocking queued delivery. The report is tagged with the app phase so a reporter can tell an app-reported error from a framework-caught one. It is best-effort and non-failing: it returns void, never blocks, and swallows even its own message-formatting allocation failure (falling back to the error name), so it can never fail or abort the caller.
Delivery is non-blocking and best-effort. The Sentry backstop does an HTTPS POST, so the framework never runs it inline on the thread that just swallowed the error (an HTTP worker, the scheduler, a queue worker). Instead it enqueues an internal "report" job on the in-process memory queue and performs the POST on a pool worker — it never touches the DB writer and never blocks the erroring path. A failed delivery is logged and dropped (never retried into a loop); the reporter is the terminal backstop, so its own failures never re-report.
TTL dedup (.reporter_dedup). By default the backstop suppresses a repeat of the same (message, phase) seen within a rolling window, so a hot error path reports once per window instead of flooding Sentry:
.reporter_dedup = .{ .window_s = 60 }, // the default when the key is omitted
.reporter_dedup = .off, // report EVERY swallowed error (no suppression)
Dedup is on by default with a 60-second window. The window is a comptime knob: with .off, no dedup map is ever allocated and the check compiles down to a single null-pointer branch. The suppression cache is in-process and bounded (best-effort — a different process instance dedups independently).
Dedup gates reporter delivery only — your onError handler still fires on every error (it is cheap and in-process; the reporter, which ships over the network, is the flood risk). And because delivery rides the shared memory-queue worker pool, a sustained error storm — or a slow/hung reporter endpoint — drops reports (the ring is bounded) rather than delaying the request path: error reporting is a best-effort backstop, never a guarantee.
12. The worked example
Three buildable examples form a ladder: examples/blog/ is the basic packaging proof (hooks + route + cron job + Zigapagos/Preact frontend served via --serve-static), examples/golfsim/ is a realistic app (hooks, routes, cron, a comptime-hardcoded .dir static mode, and a generated TypeScript typed client at clients/typescript/zbase.gen.ts — regenerate with zig build gen-client), and examples/plugins/ is the advanced framework-feature reference (custom mailer plugin + .collections schema + typed .migrations + .pools levers + fully embedded static assets via embedStaticDir).
The blog App(.{...}) block is the canonical basics reference (hooks + route + job):
pub fn main(init: std.process.Init) !void {
return zigbase.App(.{
.hooks = .{ .posts = .{ .beforeCreate = slugify } },
.routes = .{
.{ .method = .GET, .path = "/api/blog/ping", .handler = ping, .auth = .public },
},
.pools = .{ .jobs = 2 },
.cron = .{
.{ .name = "heartbeat", .schedule = zigbase.schedule.Schedule{ .interval = .hourly }, .handler = heartbeat },
},
}).runCli(init);
}
13. Serve a frontend: static files
Anything that misses /_/, the built-in API, and your custom routes falls through to the static-file server (GET/HEAD only; /api/* misses keep the JSON 404 envelope; static misses return a plain-text 404). / and directory paths resolve to index.html; embedded-mode responses carry a CRC32 content ETag (304 on If-None-Match); all responses include X-Content-Type-Options: nosniff. Both sources honor a single Range: bytes=a-b / bytes=a- / bytes=-n with a 206 (416 past EOF); a malformed or multi-range header is ignored (plain 200).
Pick a mode at comptime with .static_files:
| Mode | Config | --serve-static flag |
|---|---|---|
| runtime flag (default) | (field absent) | enabled |
| disabled | .static_files = .disabled | rejected |
| hardcoded dir | .static_files = .{ .dir = "frontend/dist" } | rejected |
| embedded | .static_files = .{ .embedded = &@import("static_assets").files } | rejected |
In embedded mode, assets are compiled into the binary. Generate the manifest from your build.zig using the helper exported by zigbase’s build.zig:
// build.zig
const zigbase_build = @import("zigbase");
const assets = zigbase_build.embedStaticDir(b, "frontend/dist");
exe_mod.addImport("static_assets", assets);
// main.zig
.static_files = .{ .embedded = &@import("static_assets").files }
The build fails with a clear error when frontend/dist is missing — build the frontend first (e.g. npm run build).
In dir mode (hardcoded .dir or --serve-static), ETag/Last-Modified and conditional-request handling (If-None-Match/If-Range) are managed by the underlying facil.io sendFile — using its own exact-match ETag semantics (an unquoted base64 size^mtime tag), not RFC 7232 list/weak comparison; zigbase adds X-Content-Type-Options: nosniff and the Range-normalization shim described above. The shim also neutralizes an inverted If-Range branch in the vendored facil.io so a matching If-Range correctly resumes (206) instead of forcing a full 200 (RFC 9110 §13.1.5); a mismatched If-Range in dir mode still resumes (206) because facil.io’s process-local ETag cannot be recomputed to detect the mismatch — owned record-file and embedded serving are fully RFC-correct on that corner. Cache-Control is covered next.
A hardcoded or --serve-static directory that is missing or unreadable at startup is a fatal startup error naming the path.
Cache-Control: the --static-cache-control knob
Every static response — embedded or dir — carries a Cache-Control header. The default is the stock max-age=3600 (byte-identical to pre-knob behavior); override it process-wide with, in precedence order:
--static-cache-control <value>(runtime flag; rejected as unknown when static serving is.disabled);ZIGBASE_STATIC_CACHE_CONTROL(env, same validation as the flag);- comptime
App(.{ .static_cache_control = "…" })(validated at compile time: non-empty, ≤ 256 bytes, no CR/LF — a violation is a@compileError).
The flag wins over the env var, which wins over the comptime default; leaving all three unset keeps the facil.io stock value. The knob is process-wide, not per-route: in dir mode it’s applied via a one-time facil.io FIOBJ swap at startup (so every sendFile-transported response inherits it), which also means a consumer route calling r.sendFile directly inherits the knob — that IS the knob, there is no separate per-call override. In embedded mode the same resolved value is threaded into every asset response directly.
The one exception is the SPA fallback shell (below), which always overrides the knob with Cache-Control: no-cache.
SPA fallback: the .spa marker
Client-routed apps (History-API SPAs) break on deep links: /app/orders/42 is a real URL in the browser but not a real file. Drop an empty file named .spa into any directory of your static tree to mark it as an SPA root: any GET/HEAD miss at or below that directory serves that directory’s index.html with status 200 (real files always win; every miss under the root gets the shell, including extension-bearing paths like a stale hashed asset). The marker is presence-only — its contents are ignored, and case-sensitively named .spa (matched case-insensitively against the request path, so it stays unservable on case-insensitive filesystems too — see below).
- Works in both static sources: a
--serve-static/.dirtree on disk, and anembedStaticDirembedded manifest (bundlers that emitdist/.spawork in both, but resolve differently — see “Dir mode: markers are live” below). - Markers nest: with
/.spaand/app/.spa, a miss at/app/orders/1servesapp/index.htmland a miss at/pricingserves the rootindex.html(longest/-bounded prefix wins —/applicationis not underapp/). - The
.spafile itself is never served, on any filesystem — the check on the request path’s final segment is ASCII case-insensitive, soGET /.SPAis refused even where dir-modestatis case-insensitive (macOS’s default APFS/HFS+ volumes). Other dotfiles (.well-known/…) are unaffected. The/apinamespace, the admin UI (/_/), and all custom routes always win over the fallback (checked on the normalized path too, so no raw-vs-normalized-path disagreement — e.g. a doubled leading slash — can route an api-looking miss to a fallback document); non-GET/HEAD methods never reach it. - Fallback responses ride the normal static pipeline for
nosniffand HEAD-mirrors-GET, but caching is special: the shell is always servedCache-Control: no-cachewith a revalidationETag(304 onIf-None-Match), overriding the--static-cache-controlknob — a stale cached shell after a redeploy would otherwise reference hashed assets that no longer exist. A direct hit on that sameindex.html(not via the fallback) is unaffected and keeps the knob’s normal value. In dir mode this one response is read and sent OWNED (not via facil.io’s delegatedsendFile), becausesendFilecannot emit a per-responseCache-Controldifferent from the process-wide knob.
Embedded mode: markers are resolved once, at startup. The manifest is comptime-static (baked into the binary), so the marker root set is derived once and reused for the process lifetime — there’s no live filesystem to go stale. A marked directory with no index.html logs a startup warning and is dropped (degrades to unmarked).
Dir mode: markers are resolved LIVE, per request. A --serve-static/.dir source has no cached marker set at all — every miss re-checks the filesystem (walking the missed path’s ancestor directories from deepest to the root, so the innermost marker still wins), which means adding, removing, or editing a .spa (or its index.html) takes effect on the very next request, with no restart and no cache to invalidate. Startup still validates the tree once, but only for the one mistake that must never reach production silently:
- A
.spa-marked directory with noindex.htmlis a fatal startup error naming the path — almost certainly a build/deploy mistake (nothing to serve as the shell), so it aborts boot rather than silently degrading. - An unreadable subdirectory encountered during that startup check (permission bits, a root-owned
0700dir,lost+found, …) is not fatal — it’s skipped with a startup warning naming the path, and behaves like it doesn’t exist both at startup and on every later live lookup, exactly like dir-mode serving already treats an ordinary unreadable file as a plain 404 rather than an error. - A directory that vanishes, becomes unreadable, or loses its
index.htmlafter boot never fails a request. If the deepest matching marker’sindex.htmlis gone, that miss resolves to a plain 404 (it does not fall through to an enclosing marker — a marked-but-indexless directory means “absent”, same rule the startup checks use); if the marker itself is gone, the walk simply continues outward to the next ancestor as if it had never been marked. - The per-miss ancestor walk is capped at 64 levels (far deeper than any real static tree); once exhausted it jumps straight to checking the static root as the final candidate, bounding the filesystem cost of a crafted, very-deep request path.
SPA fallback: comptime static_routes
Custom builds can declare explicit match → serve rewrites instead of (or on top of) the marker:
zigbase.App(.{
.static_files = .{ .embedded = &@import("static_assets").files },
.static_routes = &.{
.{ .match = "/app/orders/:id", .serve = "/app/orders/_shell.html" },
.{ .match = "/app/**", .serve = "/app/index.html" },
},
})
Patterns are matched segment-wise against the normalized path (trailing slashes never matter), first match wins in declaration order, and only on a real-file miss (real files always win). Routes are consulted before the marker.
| Segment | Matches |
|---|---|
| literal | exactly that segment |
:name | exactly one segment (capture discarded — serve is a fixed path) |
* | terminal only; one or more remaining segments (/admin/* does not match /admin) |
** | terminal only; zero or more remaining segments (/app/** does match /app) |
Malformed patterns (wildcards mid-pattern or mixed with text, //, bare :), entries without exactly .match + .serve, serve targets containing :/*/../empty segments (// or a trailing /), and .static_routes alongside .static_files = .disabled are all compile errors. In embedded mode every serve target is proven against the manifest at comptime; in dir mode targets are checked at startup (missing target = fatal, like a missing static dir), and declaring routes while starting with no static source at all is also fatal.
enable_spa_marker
.enable_spa_marker = true|false gates the marker tier — the startup derivation (embedded) / startup validation and live per-miss resolution (dir). Default: true when static_routes is absent/empty (the shipped-binary behavior), false when routes are declared (explicit config shouldn’t gain stray-dotfile behavior). An explicit value always wins; with both enabled, routes match first and the marker is the residual fallback. When false, a .spa file is just another never-served dotfile.
See examples/blog/ (runtime flag), examples/golfsim/ (hardcoded dir), and examples/plugins/ (embedded).
collections_frozen — cache collection metadata on every backend
.collections_frozen = true asserts that your collections do not change after boot + migrations. The framework parses each collection’s schema/indexes/options JSON on the record and realtime paths; a small in-process cache (colcache) keeps the parsed Collection so those reads skip the re-parse. By default that cache is installed only for SQLite (a single process, so the DDL endpoints’ in-process invalidation is complete). On Postgres it is skipped: a second instance could ALTER a collection without this process noticing, so reads stay direct rather than risk serving stale metadata.
Freezing removes that hazard by contract. When set, ZigBase:
- installs the collection-metadata cache on every backend, including Postgres — safe because nothing mutates collections after boot: provisioning +
.migrationsrun before serving, and - rejects the runtime collection create/update/delete endpoints with 403 (they never fire the cache’s invalidation, so the snapshot stays coherent for the process lifetime).
Schema then evolves the way a frozen app expects: edit your comptime .collections / add a .migrations entry and redeploy. Reads (GET collection, list) are unaffected. Leave it false (the default) if you rely on the admin UI / API to change collections at runtime.
14. Test / dev-mode determinism seams
Three env vars (dev builds only) make a test suite reproducible and encrypted-field apps debuggable. All three are compiled out entirely in production — a release binary never reads any of them.
Freezing time: ZIGBASE_FAKE_NOW
Set ZIGBASE_FAKE_NOW to an ISO-8601 UTC instant (e.g. 2029-03-07T16:00:00Z) to freeze the framework’s clock for the lifetime of the process. Every framework-controlled timestamp routes through the seam:
- Token
iat/exp(JWT auth tokens). - The scheduler’s next-fire math (cron / interval jobs).
- Auth rate-limiter wall clock.
- Auth-challenge and keyset-cursor TTL/expiry checks.
- A consumer’s own
datetime('now')/unixepoch('now')/strftime(…, 'now')/date('now')/time('now')/julianday('now')in raw SQL — those date/time functions are shadowed on every connection and resolve to the frozen instant when a freeze is active, including their zero-argument implicit-'now'forms (e.g.date()). Every other input passes through to genuine SQLite unfrozen: explicit datetimes,'+1 day'/'start of month'modifiers, and thestrftimeformat string. - The SQL keywords
CURRENT_TIMESTAMP/CURRENT_TIME/CURRENT_DATEand columnDEFAULT CURRENT_TIMESTAMPtimestamps. These read SQLite’s clock through the VFS, not the SQL-function layer, so (dev builds only) connections open against thezigbase_frozenwrapping VFS — a byte-for-byte copy of the default VFS with only its current-time hooks overridden to the frozen instant; all file I/O still delegates to the genuine OS VFS unchanged (src/clock_vfs.zig). There are no remaining unfrozen'now'paths.
On the Postgres backend (-Dpostgres, opt-in), the same freeze is achieved with one portable mechanism instead of the two SQLite shims (there is no in-process function registration or VFS on a remote server): when a freeze is active, every connection installs a session-level now() override — a zigbase_frozen.now() wrapper returning the frozen instant, placed on the connection’s search_path ahead of pg_catalog (src/backend/postgres/clock.zig). Because the framework and any consumer raw SQL both call now(), this freezes record autodate stamps, KV/metadata timestamps, and a consumer’s own now() alike. Same comptime dev_mode gate, so a production Postgres binary is unaffected.
ZIGBASE_FAKE_NOW="2029-03-07T16:00:00Z" ./zigbase serve ...
Seeding entropy: ZIGBASE_FAKE_SEED
Set ZIGBASE_FAKE_SEED to a decimal u64 (e.g. 12345) to plant a deterministic Xoshiro256++ PRNG as the entropy source for ID/token generation. Every record ID, field ID, and token key generated in the process comes from that seeded PRNG instead of the OS CSPRNG, so two runs with the same seed produce byte-for-byte identical IDs and tokens — useful for snapshot tests.
ZIGBASE_FAKE_SEED=12345 ./zigbase serve ...
The seam routes through src/id.generate, covering:
- Collection IDs and field IDs (provisioning, record creates).
tokenKeyper auth record and the OAuth2 CSRF state value.
Other randomness (AEAD nonces, OTP digits, WebAuthn challenges) is not routed through this seam — those are security-critical at runtime and seeding them is unsafe.
Fake field-crypto: ZIGBASE_FIELD_CRYPTO
Set ZIGBASE_FIELD_CRYPTO=fake to store .encrypted fields as a readable, self-labeling fake:<key>:<value> envelope instead of real AES-GCM ciphertext, so you can eyeball encrypted values right in the database while debugging (the label defaults to @test@, or ZIGBASE_FIELD_KEY if set). Fake and real envelopes are mutually unreadable, so a fake-encrypted database fails closed on a real binary. The in-process test harness (§15) exposes the same behavior via StartOptions.field_crypto / field_key (this is how a .encrypted-field app is booted under test). See Fields → Encryption at rest for the field-level details.
ZIGBASE_FIELD_CRYPTO=fake ./zigbase serve --insecure-cookies ...
Production gate
All three seams are compiled in ONLY when the dev_mode build option is true (the default in Debug builds). The release script forces -Ddev-mode=false for all shipped binaries, so a production binary has the override code comptime-eliminated:
ZIGBASE_FAKE_NOWis never read; the clock always returns wall time.ZIGBASE_FAKE_SEEDis never read; ID/token generation always uses the OS CSPRNG.ZIGBASE_FIELD_CRYPTOis never read;.encryptedfields always use real AES-GCM — fake mode is impossible on a release binary, and a fake-encrypted DB is unreadable.
You can also force the prod-safe behavior explicitly: zig build -Ddev-mode=false.
CI runs both passes: the default Debug pass (dev features on, prod-gate assertions skipped) and a -Ddev-mode=false prod-gate pass (dev features off, prod-gate assertions executed) to verify the compiled-out guarantee.
15. Testing your app (zigbase.testing)
zigbase.testing is an in-process test harness: it boots your App(.{...}) against a throwaway data directory and injects requests through the real pipeline — the same router, access rules, auth, hooks, and custom routes the socket server runs — with no socket, no port, and no background threads. Assertions run against genuine http responses, so a test exercises end-to-end behavior (rule evaluation, auth gating, record writes) that a pure-handler unit test cannot reach, without the cost of standing up a real server.
start / deinit
const std = @import("std");
const zigbase = @import("zigbase");
const MyApp = zigbase.App(.{
.routes = .{
.{ .method = .GET, .path = "/api/ping", .handler = ping, .auth = .public },
},
.collections = .{
.things = .{
.fields = .{ .{ .name = "name", .type = .text } },
.rules = .{ .list = "@public", .view = "@public", .create = "@public" },
},
},
});
test "ping" {
var t = try zigbase.testing.start(MyApp, .{}); // migrations run + onBootstrap fires
defer t.deinit(); // tears down the app + removes the tempdir
const r = try t.request(.GET, "/api/ping", .{});
try std.testing.expectEqual(@as(u16, 200), r.status);
}
start(comptime AppType, opts) boots AppType (the type returned by zigbase.App(.{...})) into a Harness. Migrations and comptime provisioning run, and the onBootstrap hook fires, before it returns. StartOptions (all defaulted, so .{} works):
| Field | Default | Meaning |
|---|---|---|
data_dir | null | Data dir to boot against. null mints a fresh tempdir that deinit removes; a supplied dir is used as-is and left in place. |
allocator | std.testing.allocator | Allocator backing the harness — the leak-checking default fails a test on any leak across boot → requests → deinit. |
io | std.testing.io | Async/IO handle threaded into the app. |
fake_now_unix | null | Freeze the dev clock (unix seconds) — token expiry, TTLs, and datetime('now') all read it (see §14). |
fake_seed | null | Seed the dev PRNG for reproducible IDs/tokens (see §14). |
field_key | null | Real AES-GCM field-encryption key for an app with .encrypted fields — closes #260 (see below). null and no field_crypto override defaults the app to fake-encrypt instead of failing closed. |
field_crypto | null (infer) | .real or .fake (field_policy.Mode). null infers .real when field_key is set, else .fake. .fake requires a dev-mode build (see below). |
deinit tears down the booted app, every request arena, the optional capture mailer, and the tempdir. The harness uses an on-disk tempdir, never :memory: — the multi-connection reader pool needs a shared file.
request and Response
const r = try t.request(.POST, "/api/collections/things/records", .{ .json = .{ .name = "widget" } });
try std.testing.expectEqual(@as(u16, 201), r.status);
const rec = try r.json(struct { id: []const u8, name: []const u8 }); // arena-owned parse
request(method, path, opts) builds an http.RequestCtx and runs it through the full routing/fallback chain. opts is an anonymous struct — every field is optional, so pass .{} for a bare request. An unrecognized key is a compile error (a typo like .jsn can never silently send an empty body):
| Field | Type | Effect |
|---|---|---|
json | anytype | Serialized to the body as JSON; sets Content-Type: application/json. |
body | []const u8 | Raw request body (ignored when json is present). A multipart/form-data body (set content_type) is pre-parsed into ctx.form_fields/ctx.files, exactly as on-socket — file-upload handlers work through the harness. |
auth | []const u8 | The Authorization header value — pass what an auth helper returns ("Bearer <tok>"). Wins over an Authorization entry in headers. |
headers | []const [2][]const u8 | Extra request headers as .{ name, value } pairs. They feed the generic ctx.headers list (the route-guard .header source) and, for the well-known names the real pipeline reads from dedicated RequestCtx fields — Cookie, If-None-Match, User-Agent, X-CSRF-Token, Content-Type, Authorization — the matching field. |
query | []const u8 | Raw query string (no leading ?); overrides a ?... embedded in the path. |
content_type | []const u8 | Overrides the request content-type. |
cookie | []const u8 | The Cookie request header (e.g. "zb_auth=<tok>") — the only way to send a request cookie, so cookie-based auth / token refresh / logout / CSRF double-submit flows are testable. |
The
.jsonergonomics. Because a struct field cannot itself beanytype,requesttakes the options asanytypeand introspects the literal. That makes.jsona true optional-anytype:.{ .json = .{ .name = "x" } }when present, simply absent otherwise — no sentinel, no separate method. Any serializable value works as the body.
The returned Response is owned by the harness request arena (freed at deinit), so it outlives the call:
.status: u16,.body: []const u8,.content_type: []const u8.header(name) ?[]const u8— case-insensitive;"content-type"resolves to the response type..cookie(name) ?[]const u8— aSet-Cookievalue by name.json(comptime T) !T— parse the body intoT(unknown fields ignored), into the request arena.
Auth helpers — mint vs. real login
Two ways to obtain an Authorization value, both returning "Bearer <tok>":
// (1) Direct JWT mint — deterministic, no HTTP. Reads the record's tokenKey and signs an
// `.auth` token. The default for the epoch session store.
const sess = try t.mintSession("users", user_id);
// (2) The REAL auth-with-password endpoint in-process — full fidelity (rate limiter, argon2,
// verification gate, beforeAuthSuccess/onAuth hooks).
const admin = try t.loginSuperuser("admin@example.com", "password123");
const user = try t.loginPassword("users", "user@example.com", "hunter2xx");
const r = try t.request(.GET, "/api/collections", .{ .auth = admin });
mintSession is the deterministic default; use loginPassword / loginSuperuser when a test must exercise the login endpoint itself (or under .session_store = .table, which needs a real _sessions row).
Seeding
const su_id = try t.createSuperuser("admin@example.com", "password123"); // returns the record id
const rec = try t.createRecord("things", .{ .name = "seeded" }); // via the Data facade
createSuperuser provisions a superuser exactly as zigbase superuser create does (argon2id hash
- random
tokenKey), sologinSuperuserworks against it.createRecordtakes any value serializable to a JSON object and creates through the sameDatafacade hooks and routes use — auth collections get a generatedtokenKeyand a hashedpassword.
Composing the mail + clock seams
Swap in a CaptureMailer to assert on outbound mail with no SMTP, and freeze the clock for deterministic tokens/TTLs:
test "verification email is sent" {
var t = try zigbase.testing.start(MyApp, .{ .fake_now_unix = 1_800_000_000 });
defer t.deinit();
const mail = try t.captureMail(); // installs an in-memory mailer; harness owns it
_ = try t.request(.POST, "/api/collections/users/request-verification",
.{ .json = .{ .email = "user@example.com" } });
try std.testing.expectEqual(@as(usize, 1), mail.messages.items.len);
try std.testing.expectEqualStrings("user@example.com", mail.messages.items[0].to);
}
captureMail is idempotent (repeated calls return the same instance) and the captured messages carry owned copies of every field (subject, recipient, both body parts, attachments). The fake_now_unix / fake_seed options drive the same dev-mode clock and seeded-entropy seams described in §14, so token exp, TTL math, and generated IDs are reproducible across runs.
Encrypted-field apps (#260)
An app whose comptime .collections declare an .encrypted field (see docs/fields.md) used to be impossible to boot through the harness at all: production boot fails closed with error.FieldKeyRequired when no ZIGBASE_FIELD_KEY is configured, and the harness had no way to set one. StartOptions now offers two ways to boot such an app:
// (1) Real AES-GCM — full fidelity, exercises the actual envelope format.
var t = try zigbase.testing.start(MyApp, .{ .field_key = "test-operator-key" });
// (2) Fake-encrypt (the default when field_key is omitted) — dev-mode only.
var t = try zigbase.testing.start(MyApp, .{});
With no field_key and no field_crypto override, the harness defaults to fake-encrypt: .encrypted values are stored as the plaintext-visible fake:<label>:<value> (label "@test@", or field_key if you set one without forcing .field_crypto = .real) instead of a real AES-GCM envelope, so assertions can compare against the plaintext directly. Fake-encrypt is compiled in only on a dev-mode build (zig build test, the default harness build) — see §14 and docs/fields.md for the production-safety gate. Pass .field_key to boot with real crypto instead (works in every build; not gated), e.g. to test key-rotation or envelope-format behavior end-to-end.
Wiring the test step (zigbase.addTest)
// build.zig
const zigbase = @import("zigbase");
const tests = zigbase.addTest(b, dep, .{ .root_module = app_mod });
b.step("test", "Run tests").dependOn(&b.addRunArtifact(tests).step);
Pass the same module you passed to addTo: rooting a second module at src/main.zig would place one file in two modules, which Zig rejects.
addTest wires a .simple-mode test runner that ZigBase ships (src/simple_runner.zig, resolved out of the dependency — you do not vendor anything). That matters for two reasons:
- It avoids a misleading Zig 0.16.0 diagnostic. The default server-mode runner (
--listen=-) can printfailed command:after a successful child leaves stderr output at exit. The reproduced trigger is facil.io’s destructor newline; the command exits zero and the summary reports passing tests. This does not establish a crash or a runner race. A.simplerunner reports through its exit code instead and does not print that stale label on success. - It rejects real failures. Each test runs under a fresh
std.testing.allocatorwith leak checking. Failed tests, logged errors, and abnormal process exits still fail the build.
For default-runner output, check both the command exit status and final build summary. Never dismiss a nonzero exit or signal as the newline diagnostic. See testing for the reproducible runner checks and #261 for the investigation. The standalone reproduction and candidate upstream patch are contributor diagnostics; keep using zigbase.addTest in consumer apps.
Compile-time build flags
zig build/zig build -Dname=value accepts these consumer-facing flags. Each folds its gated code to comptime-dead when off, so a build that doesn’t need a feature doesn’t pay for it:
| Flag | Default | Effect |
|---|---|---|
-Dfts5 | on | SQLite full-text search (FTS5). -Dfts5=false drops -DSQLITE_ENABLE_FTS5 from the SQLite build (~250-400 KB smaller) for lean binaries with no .searchable field; ?search= then 400s and the server refuses to start over a .searchable SQLite schema. Postgres full-text search is unaffected. → docs/search.md |
-Dvector | off | Opt-in nearest-neighbor ?vector= KNN search — sqlite-vec on SQLite, pgvector on Postgres. → docs/search.md |
-Dpostgres | off | Opt-in pure-Zig PostgreSQL wire-protocol backend, alongside the default SQLite one. → docs/postgres.md |
-Dfile-inventory | off | Read-only files inventory plus offline local/SQLite files reconcile (dry-run unless --apply). Bounded pages; no HTTP surface. Inventory supports S3 with -Ds3=true; reconciliation refuses S3/PostgreSQL/custom storage. The boot-lifetime local-storage maintenance lease remains in default builds. |
-Ddurable-realtime | off | Transactional built-in REST record replay on SQLite/PostgreSQL; implies realtime-backfill. Shared 4096-entry / 4 MiB journal, 64 KiB frame ceiling, 24-hour retention; restart/cross-instance checkpoints and current authorization. Raw SQL/Data/hook side-writes are outside capture scope. See API realtime section. |
-Dreplay-max-entries | 4096 | Durable journal global entry budget, 1..65536. All replicas must agree. |
-Dreplay-max-bytes | 4194304 | Durable journal encoded-frame byte budget, 1..1073741824. |
-Dreplay-max-frame-bytes | 65536 | Maximum individual durable frame bytes, 1..1048576, at most total byte budget. |
-Dreplay-retention-seconds | 86400 | Durable event lifetime stamped on capture, 1..31536000 seconds; physical expiry cleanup occurs on writes. |
-Drealtime-backfill | off | Single-process SQLite record invalidation backfill: 16 lazy collection slots, each 256 entries / 64 KiB (1 MiB encoded total), current authorization, explicit reset on gaps or slot replacement. No historical payloads or durable/cross-instance guarantee. See the API realtime section. |
-Dimage-thumbnails | off | Named PNG/JPEG/WebP derivatives on built-in local storage via a trusted external ImageMagick executable; configure .files.thumbnails. Subprocess support, routes and admission state are excluded when off; no image codec is linked. See thumbnails. |
-Dresumable-uploads | off | Principal-bound network-resume for file fields on existing records. Process-local by default, fully buffered, configurable session/byte/chunk/expiry budgets. SQLite restart persistence requires the additional -Ddurable-resumable-uploads flag and .files.resumable.durable = true; neither mode provides cross-instance durability. See resumable uploads. |
-Ddurable-resumable-uploads | off | Compile SQLite/local upload persistence; requires -Dresumable-uploads=true and .files.resumable.durable = true to activate. Single-owner process-restart recovery with atomic completion receipts; still fully buffered, not cross-instance or power-loss durability. See persistence limits. |
-Drest-idempotency | off | Compile opt-in persistent REST retry receipts. |
-Dquery-workbench | off | Bounded SQLite step and PostgreSQL client-exchange metrics, statement lifecycle timing, route attribution and repeated/slow shape counters. Structural EXPLAIN remains SQLite-only; no SQL/parameter capture. |
-Dcoordinated-admission | off | Compile memory/durable job admission: .admission.max_work shares HTTP/memory-job/durable-batch work capacity and requires .max_requests; independent .max_job_bytes bounds retained payload/name copy lengths without enabling HTTP admission. No RSS cap; durable claim payloads are not charged to .max_job_bytes. |
-Ddev-mode | on in Debug, off in release | The dev-only, never-in-prod seams: ZIGBASE_FAKE_NOW / ZIGBASE_FAKE_SEED (§14 above), test-capture, and fake field-crypto; the release script forces it off for shipped binaries. |
-Ddev-tools | on | The init/agents-md/typegen scaffolding/codegen verbs, capabilities/routes/migrate preview offline discovery, tune offline measurement advisor, and diagnostics structured doctor adapter (which can probe filesystem writability and initialize the migration ledger). Ordinary doctor and other migration actions remain available. Official release, Docker and npm artifacts include this tooling. Consumers can opt out for their deployment binary; stripped verbs exit nonzero with -Ddev-tools=true rebuild guidance. Distinct from .enable_typegen below — see §3b. |
-Dstrip | on except in Debug | Strip debug info from the binary (~7 MiB vs ~24 MiB unstripped in a release build). |
Version transparency & dependency auditing
A ZigBase binary bakes in a vendored SQLite C amalgamation, an optional sqlite-vec amalgamation (-Dvector), and the zap/facil.io server. The pinned versions of all of these — plus zigbase itself — are aggregated at build time (zero runtime cost) and surfaced four ways (#282):
zigbase --versionprints build provenance and a “Vendored/native components” block: SQLite (+ source id), sqlite-vec (+ a linked / not-linked note), zap (+ pinned commit), facil.io.zig build versionsruns the freshly-built binary with--version, so you can read exactly what a build would ship without launching a server.Startup log —
zigbase serveemits oneversions: …INFOline at boot (thesqlitevalue there is the live linked library version).GET /api/healthreturns aversionsobject next to the backend badge:{ "status": "ok", "backend": "sqlite", "versions": { "zigbase": "0.12.0", "commit": "…", "sqlite": "3.53.2", "sqliteVec": "v0.1.6", "zap": "0.10.6", "facil": "0.7.4" } }These are non-secret build provenance only — no connection string, host, or credential is exposed.
zig build audit compares the pinned versions against a curated in-repo advisory table (docs/security-advisories.md) and exits non-zero if any pin falls in a known-affected range. The transparency sources, the audit workflow, and the update process for a vendored C / dependency security fix are documented in docs/security-audit.md → “Dependency version transparency & supply-chain auditing”.
Exported names reference
The public surface (from src/root.zig):
zigbase.App— the comptime application builder.zigbase.Runtime— the runtime app context type (the*Appyou receive on events).zigbase.Config,zigbase.Server(a genericServer(comptime gates: Gates) type;App(cfg).runCli/serveinstantiate it for you from your config — see the.adminkey and route gating above).zigbase.http— HTTP types (http.Response,http.Method, …).zigbase.Ctx— the per-request capability object passed as the first parameter to every route / hook / job / lifecycle handler (ctx.records(),ctx.http(),ctx.user(),ctx.tx(),ctx.issueSession(),ctx.fail/ctx.invalid/ctx.errorResponse).zigbase.events— all event/handler types (events.AuthEvent,events.FileEvent,events.LifecycleEvent,events.JobEvent, …).zigbase.schedule—schedule.Schedule,schedule.Interval,schedule.Reactive.zigbase.RecordEvent,zigbase.ErrorEvent,zigbase.JobEvent— re-exported directly for convenience (the same types aszigbase.events.*); they are the second parameter of hook / job handlers.zigbase.Req,zigbase.RouteError— the typed-route handler input wrapper and error set.zigbase.Tx— the transaction scope passed to actx.tx(T, fn(*Tx) ...)callback (t.records(),t.arena()).zigbase.Migration— the.migrationsentry type (bare tuple or typed slice);zigbase.Db— the writer connection passed to a migration’s.up.zigbase.StaticFile— the embedded manifest entry type (path, bytes, etag); used by.static_files = .{ .embedded = ... }.zigbase.Storage/zigbase.Mailer/zigbase.Email— the storage & mailer plugin vtable types;zigbase.DefaultStoragePlugin/zigbase.DefaultMailerPlugin— the built-in defaults;zigbase.LocalStorage,zigbase.LogMailer,zigbase.SmtpMailer,zigbase.CommandMailer,zigbase.SmtpTls— the concrete backends.zigbase.MailMessage— the{ to, subject, text?, html?, reply_to? }message type taken byctx.mail().send/.enqueue.zigbase.QueueDef/zigbase.WorkerDef/zigbase.RetryPolicy/zigbase.Backend/zigbase.Priority/zigbase.Backoff/zigbase.QueueRegistry— the background-jobs config types named when declaring.queues/.workers. The compile-checked enqueue accessor (App.enqueue(ctx, .queue, .kind, payload)) and the generatedQueue/Jobenums live on theApp(cfg)type; the runtime escape hatch isctx.enqueue/ctx.enqueueByName.zigbase.AuthMethod— the auth plugin vtable type.zigbase.AuthCtx— the per-request auth context passed to plugin phases.zigbase.auth.Resolution/zigbase.auth.InitiateResult— the phase return types.zigbase.testing— the in-process test harness (testing.start(App, .{})→Harnesswithrequest/mintSession/loginSuperuser/createRecord/captureMail). See §15.
See also
- tutorial.md — build an app on ZigBase, end to end.
- recipes.md — copy-pasteable hook / route / job patterns (computed fields, owner rules, path-param routes, DB access in cron).
- fields.md — the field-type & options catalog.
- api.md — the HTTP REST + WebSocket reference.