Documentation

Full-text & vector search — ZigBase

Ranked ?search= queries on searchable fields, operators, how search composes with filters and rules, and opt-in -Dvector KNN.

ZigBase has first-class search on the list endpoint: ranked ?search= queries over fields you mark searchable, plus an opt-in -Dvector build for embedding KNN. Search is never a separate, unscoped query — it composes with the same filter + rules + tenant scoping as every other list request, so it can never widen visibility.

Make a field searchable

Mark one or more text/editor fields .searchable in the schema:

.posts = .{ .fields = .{
    .{ .name = "title", .type = .text,   .searchable = true },
    .{ .name = "body",  .type = .editor, .searchable = true },
} },

At startup ZigBase provisions the index automatically — no migration needed. On SQLite (the default backend) that’s an FTS5 external-content index per searchable collection ("<col>_fts", content='<col>') plus INSERT/UPDATE/DELETE triggers that keep it in lock-step with the base table — no doubled storage. On Postgres it’s a STORED tsvector generated column (to_tsvector('simple', …) over the searchable columns) plus a GIN index. searchable is mutually exclusive with encrypted — ciphertext is not searchable.

Startup reconciles only search objects whose catalog definitions prove they belong to ZigBase. A conflicting application table, column, index, trigger, or dependent object causes Conflict before search DDL; its data is preserved, even when the collection has no searchable fields. Proven engine indexes still support searchable-field changes, removal, and missing-index repair. Reconciliation is transactional.

Additive provisioning validates existing search objects before changing fields, then reconciles search after the new columns exist. The search reconciliation is transactional; this does not make the entire collection provision operation atomic. PostgreSQL resolves search tables only in the current schema, matching the registry schema when a registry exists. A conflicting temporary validation-probe relation is rejected without replacing it. Because runtime trigger functions can refer to generated columns without catalog dependencies, PostgreSQL refuses to remove or replace a search column while the table has user triggers. Unchanged search and missing-GIN-index repair remain supported; explicitly migrate affected triggers before changing the searchable field set. SQLite external-content search operands must still exist as text-storage columns in the physical base table; a missing or non-text operand requires explicit repair before its search objects can be adopted or removed.

Reserve <col>_fts and its suffixed object names for the engine. PostgreSQL ownership also requires the engine’s zbfts: column comment and matching generated expression. Legacy unmarked columns are no longer automatically deleted: inspect and back up the object, then explicitly migrate or rename conflicting application objects before retrying startup. Do not attach an ownership marker to an unverified column.

Build requirement (-Dfts5, default on)

SQLite FTS5 is compiled in by default — it’s a core feature, not an experiment. Lean custom builds that never declare a .searchable field can drop it with -Dfts5=false (~250-400 KB smaller binary); every FTS5 code path in src/search/fts.zig folds to comptime-dead. With -Dfts5=false, a ?search= request answers a clean 400 ("Full-text search is not enabled in this build."), and — more importantly — the server refuses to start if the comptime schema declares any .searchable field on the SQLite backend, with an actionable startup error, rather than silently skipping the index and surfacing a 500 on the first search. Postgres full-text search (the tsvector/GIN path above) is a server-native feature and is not gated by this flag.

Query it

Query with search (or its alias q):

GET /api/collections/posts/records?search=zig%20database
GET /api/collections/posts/records?q=alpha%20OR%20beta&filter=published=true

Results are ranked by relevance in offset mode — bm25 on SQLite, ts_rank(…) DESC on Postgres. The exact relevance order can differ between the two ranking functions (FTS5 bm25 length-normalizes; ts_rank does not), but the matched set is equivalent. On SQLite, terms support the basic FTS5 operators (AND, OR, NOT, and a trailing * for prefix search); on Postgres, plainto_tsquery parses the plain term and does not honor those operators. A search whose terms reduce to nothing (e.g. operator-only, ?search=AND) matches no rows rather than returning the whole collection. A search on a collection with no searchable field returns 400. The _fts collection-name suffix is reserved (it backs the per-collection shadow tables).

Compose with filters and rules

The search predicate is AND-ed into the same composed WHERE as your filter, the list rule, abilities, and tenant scope — search can never widen visibility. A search of a tenant-owned or ability-guarded collection returns only the rows the caller may already view. The whole term is passed as a bound parameter (never interpolated) and lowered to a guaranteed-valid query, so a malformed input is harmless — it can never become a SQL error or injection. This composed scoping is identical on both backends.

Vector search (opt-in)

Vector search is not compiled into the default binary. The single -Dvector=true flag enables KNN on both backends — on SQLite it vendors and links sqlite-vec (registered on every connection); on Postgres it emits the pgvector lowering. It enables KNN ordering over a field that stores a JSON embedding array:

GET /api/collections/docs/records?vector=embedding:cosine:[0.12,0.04,...]
GET /api/collections/docs/records?vector=embedding:l2:[0.12,0.04,...]&filter=lang="en"

The form is <field>[:cosine|:l2]:<json-embedding> (cosine is the default metric); rows are ordered nearest-first. The embedding is validated (a non-empty JSON array of finite numbers) and bound; a malformed or dimension-mismatched embedding returns a clean 400. Vector search runs in offset mode (cursor paging is rejected with 400). In the default build a vector query returns 400 ("Vector search is not enabled in this build."), and the binary is byte-for-byte unaffected. The same composed-WHERE scoping (filter + list rule + abilities + tenant) applies to vector queries exactly as it does to ?search= — identically on both backends.

Reference