User indexes¶
Use this guide to stand up your own Iceberg table: declare its fields, send the request that builds it, write to it over the Elasticsearch-compatible bulk API, and read it with SQL like any other table. Index templates do the same job from a glob pattern, so a shipper can write to a date-suffixed id without a request of its own.
If nothing matches a valid index id, ingest still accepts documents for it. The
compactor then holds filesystem segments sealed or releases catalog claims for
retry and increments
siglake_compactor_index_unresolved_total.
Create a user index: tokenizer, mapping mode, and how template matching and precedence work¶
You create a user index with one POST to the query server: it declares the
fields, gives each text field a tokenizer, and sets the mapping mode for
undeclared attributes.
curl -sX POST localhost:8089/api/v1/indexes \
-H 'Content-Type: application/json' \
-d '{
"index_id": "app-logs",
"doc_mapping": {
"mode": "dynamic",
"timestamp_field": "timestamp",
"field_mappings": [
{"name": "timestamp", "type": "datetime", "required": true},
{"name": "message", "type": "text", "tokenizer": "default"},
{"name": "level", "type": "text", "tokenizer": "raw"},
{"name": "service", "type": "text", "tokenizer": "raw"}
],
"tag_fields": ["level", "service"],
"default_search_fields": ["message"]
}
}'
A 201 returns the stored mapping, a duplicate id 409. There is no CLI
create command. Only text fields take a tokenizer, and only these three work:
| Tokenizer | What it does |
|---|---|
default |
Splits the value into words. |
raw |
The whole value is one term. |
stem |
Stemmed tokens, matched by word root. |
mode |
The undeclared attribute |
|---|---|
dynamic (default) |
Stays in attributes. |
lenient |
Is dropped silently. |
strict |
Also stays today; the row is counted, not rejected. |
A template creates the index instead, on the first drain of an id it matches;
matching reads the index id alone. Precedence is the highest priority, then
the smallest template_id, never the narrower pattern. all-logs
(*-logs-*, 100) and app-logs-daily (app-logs-*, 100) both match
app-logs-2026-09-06, so all-logs wins the tie; raise app-logs-daily above
100 to win.
Field types¶
Every entry in field_mappings has a name, a type, and an optional
required. These are the types:
| Type | Iceberg type |
|---|---|
text |
string |
long |
long |
double |
double |
bool |
boolean |
datetime |
timestamptz |
bytes |
binary |
json |
string |
The table schema reference is generated from the declarations and gives the Arrow type each of these compiles to.
datetime has microsecond precision and is required for the timestamp field.
bytes is an opaque payload. json is stored as a string for Arrow and
Iceberg compatibility. Only text takes a tokenizer: on any other type the
body fails to deserialize and nothing is stored. Omitting tokenizer on a
text field gives that field the same tokenization as default.
Pick default for messages and prose, raw for levels, ids and service names,
stem for prose you match by word root.
Choosing raw for low-cardinality fields like level or service is what
makes them useful as tag fields. A raw field contributes nothing to the
inverted index, because whole-string matching is already served by Parquet
blooms. See index_at_flush.
No unsigned integers, deliberately
The compiled schema must survive strict Iceberg conversion, and Iceberg only supports signed integer and floating-point primitives here.
Timestamp, tag, and search fields¶
timestamp_field is required. It must name a field that is declared,
datetime, and required. It becomes the partition and sort column. Everything
in Storage depends on it.
tag_fields are treated as low-cardinality dimensions. Like other typed
Parquet columns, they get statistics; strings use dictionary encoding, and tag
fields are eligible for group-count footers. Parquet-native per-column blooms
are a separate
opt-in, write-only feature,
not a default tag-field guarantee. Choose tag fields for values you often group
or filter by.
default_search_fields are the text fields a bare text query searches. They
must be text fields.
Strict mode preserves undeclared attributes today¶
strict is intended to reject documents carrying undeclared attributes. Today
it preserves them in attributes, exactly as dynamic does, and counts the
offending rows in siglake_compactor_strict_residual_rows_total.
Strict mode does not reject today
Don't rely on strict mapping as a data-quality gate. See Limitations.
An index id that matches no index and no template is a separate case. See When nothing matches.
Confirm the mapping the server stored¶
A GET returns the stored configuration, which is the mapping the compactor
applies from now on. Two parts of it decide what a later query can do: the
mapping mode, which fixes whether an undeclared attribute survives, and the set
of fields tokenized raw, which fixes whether a bare text query can match
inside them. Pull both out with jq:
curl -s localhost:8089/api/v1/indexes/app-logs \
| jq '.doc_mapping.mode, [.doc_mapping.field_mappings[] | select(.tokenizer == "raw").name]'
Naming rules¶
index_id must match [a-z0-9][a-z0-9_-]{0,127}. Names starting with _ are
system-reserved, as are names colliding with WAL layout directory names
(active, sealed, processing, committed, consumers).
Field names must match [a-zA-Z_][a-zA-Z0-9_]*, must be unique, and cannot
use reserved names (columns appended automatically).
Writing to an index¶
Via the Elasticsearch bulk API:
curl -s http://localhost:8088/api/v1/_elastic/_bulk \
-H 'Content-Type: application/x-ndjson' \
--data-binary $'{"index":{"_index":"app-logs"}}\n{"timestamp":"2026-07-28T12:00:00Z","message":"request completed","level":"info","service":"checkout","status":200}\n'
status is not in that mapping, so dynamic keeps it in the residual
attributes column instead of giving it a column of its own.
Or with the index in the path:
curl -s http://localhost:8088/api/v1/_elastic/app-logs/_bulk \
-H 'Content-Type: application/x-ndjson' \
--data-binary $'{"index":{}}\n{"timestamp":"…","message":"…"}\n'
Querying an index¶
It's just a table:
SELECT timestamp, service, message
FROM "app-logs"
WHERE level = 'error'
AND timestamp >= now() - INTERVAL '1 hour'
ORDER BY timestamp DESC
LIMIT 100;
SELECT service, count(*) AS n
FROM "app-logs"
GROUP BY service
ORDER BY n DESC;
Quote the name if it contains a hyphen.
Templates¶
Templates create indexes whose ids match a glob pattern. A shipper can then write to a date-suffixed index without a separate create request. The rule that decides between two matching templates is in Create a user index.
Template fields¶
PUT /api/v1/index-templates/{id} is an upsert; there is no POST, and
re-sending a template replaces the stored one wholesale.
| Field | Required | Meaning |
|---|---|---|
template_id |
yes | Must equal the path id, or the request is 400. Same syntax as an index id (see Naming rules). |
index_id_patterns |
yes | Glob patterns matched against the candidate index id. An empty list is 400. |
priority |
yes | Signed 32-bit integer. The highest matching value wins. |
doc_mapping |
yes | Exactly the mapping an index takes: field_mappings and timestamp_field required, mode, tag_fields, and default_search_fields optional. |
retention |
no | Copied onto each index the template creates. Omit for no retention policy. |
Any other field is rejected before the handler runs. The body fails to
deserialize and nothing is stored. In particular, index_at_flush is not a
template field: a template-created index always inherits the deployment default
(SIGLAKE_INDEX_AT_FLUSH).
Pattern syntax and the built-in templates¶
Patterns use Siglake's own glob, which is narrower than shell globbing: * is
the only metacharacter, ? and character classes are literal, and a pattern
with no * must equal the index id exactly. app-logs-* matches
app-logs-2026-09-06, *-logs matches any id with that suffix, and a**c
behaves like a*c. Only the index id takes part in the match; nothing in the
document does.
Two built-in templates always take part, stored or not:
template_id |
Pattern | priority |
Mapping |
|---|---|---|---|
siglake-logs |
siglake-logs-* |
0 | The events mapping |
siglake-traces |
siglake-traces-* |
0 | An OTLP span mapping |
Storing a template under one of those two ids replaces the built-in instead of
competing with it. priority is a signed 32-bit integer, so a negative value
loses to both built-ins. Where several patterns hit one id, the highest
priority wins and a tie goes to the lexicographically smallest template_id.
Resolution happens once, at index-creation time. Editing a template does not retroactively change the indexes it already created. Each carries its own copy of the mapping, so a change only reaches future indexes.
Stored templates are per tenant. A PUT writes one record per template id,
under _siglake/config/index_templates/<namespace>/<template_id>.json in the
warehouse, and resolution reads only the records of the index's own namespace.
The default namespace also reads two older warehouse-global layouts: the
single _siglake/config/index_templates.json document and records written
directly under _siglake/config/index_templates/. Templates stored by an
older version therefore reach the default tenant only.
A worked example¶
Pick a current timestamp and a matching index id first, so the rows the example writes are not immediately eligible for the template's retention horizon:
Store the template on the query server:
curl -sX PUT "http://localhost:8089/api/v1/index-templates/app-logs-daily" \
-H 'Content-Type: application/json' \
-d '{
"template_id": "app-logs-daily",
"index_id_patterns": ["app-logs-*"],
"priority": 100,
"doc_mapping": {
"mode": "dynamic",
"timestamp_field": "timestamp",
"field_mappings": [
{"name": "timestamp", "type": "datetime", "required": true},
{"name": "message", "type": "text", "tokenizer": "default"},
{"name": "level", "type": "text", "tokenizer": "raw"},
{"name": "service", "type": "text", "tokenizer": "raw"},
{"name": "status", "type": "long"}
],
"tag_fields": ["level", "service"],
"default_search_fields": ["message"]
},
"retention": {"period_secs": 2592000}
}'
A 200 echoes the stored template back. Now bulk-write to $INDEX, which does
not exist. There is no create call in between:
printf '{"index":{}}\n{"timestamp":"%s","message":"request completed","level":"info","service":"checkout","status":200}\n' "$TS" \
| curl -s "http://localhost:8088/api/v1/_elastic/$INDEX/_bulk" \
-H 'Content-Type: application/x-ndjson' --data-binary @-
The 201 in that response means the event reached the WAL. It does not mean
the index exists yet. The ingester never consults the index catalog or the
template list. The index
is created by the compactor when it drains that index's WAL lane. Wait for one
seal and one compactor cycle. This takes a few seconds with ingest-server
--with-compactor, or force it with siglake compactor --once against the same
--wal and warehouse. Then the index exists, with the template's mapping and
retention materialized onto it:
and it is queryable like any other table:
curl -s http://localhost:8089/api/v1/sql \
-H 'Content-Type: application/json' \
-d "{\"query\": \"SELECT timestamp, service, message FROM \\\"$INDEX\\\" ORDER BY timestamp DESC LIMIT 10\"}" \
| jq
This walkthrough stays in the default tenant. Ingest is single-tenant by
default, and the bulk request omits X-Scope-OrgID. The index lookup and SQL
request read the default namespace because the walkthrough does not configure
the query server's --oidc-tenant-claim, and the template is stored in that
same default namespace.
An X-Scope-OrgID naming another tenant is refused with 403 over HTTP or
PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. Naming default
has no effect. Set ingester.oidc.tenantClaim to route ingest by a verified
JSON Web Token (JWT) claim. Set ingester.trustScopeHeader to route by the
header on the client's word.
For a non-default tenant, follow the
custom tenant setup: configure the
same OIDC tenant claim on the ingester and query server, then send a JWT carrying
that claim to both. A present X-Scope-OrgID must agree with the claim. The
header never selects the tenant for a query request; query tenancy comes from
the verified JWT claim.
When nothing matches¶
An unmatched index ID stalls compaction but does not reject ingest. Ingest
still returns 200 with a per-item 201, and the events are durably written
to that lane.
The compactor is where it surfaces: it cannot resolve an index to commit into,
so it leaves the segments sealed and increments
siglake_compactor_index_unresolved_total by the number of pending segments
on that cycle and every later one. Catalog-claim mode releases the claim for
retry and increments the same counter.
Nothing is lost. Create the index, or a template that matches it, and the next
drain commits the accumulated backlog. Until then the rows are invisible to
every query and the id does not appear in GET /api/v1/indexes. That counter
is the only signal that a shipper is writing somewhere nothing will ever read,
so alert on it if you depend on templates.
Elasticsearch date suffixes with dots are not legal index ids
An index id must match [a-z0-9][a-z0-9_-]{0,127}, so the conventional
logstash-2026.09.06 is rejected. The bulk response is 200 with
errors: true and a per-item 400 illegal_argument_exception for each
such action. This is a validation failure rather than a template miss. No
pattern can rescue it. Configure the shipper's index pattern with -
separators.
Retention¶
Retention drops whole data files whose manifest maximum timestamp is older than the horizon. It is file- and day-granular. There is no row-level retention, so the effective horizon is the policy plus whatever span the last surviving file covers.
Only siglake retention-sweep applies the policy. The compactor's own
retention pass covers committed WAL segments, not index data, so a policy with
no scheduled sweep removes nothing. See
Retention and deletes.
index_at_flush¶
Whether this index's raw-text search acceleration is built inline with every flush or deferred to compaction. It controls when the work happens, not which indexes are enabled.
| Setting | Trade-off |
|---|---|
true (default) |
The enabled indexes are built on every flush, so fresh files are accelerated immediately. This costs about 30% of append time. |
false |
Better drain throughput. A file gets the enabled indexes when a compaction rewrite consolidates it. |
null |
Inherit the deployment default (SIGLAKE_INDEX_AT_FLUSH, inline unless 0). |
For high-volume streams, false usually gives better ingest throughput.
Queries over unindexed files remain correct because the funnel falls back to a
scan. Only latency changes.
Footer inverted indexes are enabled by default. The post-rewrite Puffin
backfill is not: it runs only under compactor.indexRebuild: true, so a
compaction rewrite that streams leaves its output to be searched by scan. That
switch is separate from this one: index_at_flush neither turns it on nor off,
and neither switch affects reads of an index a file already has. See Search
acceleration for the
three controls.
Which columns an inverted index covers is derived from this index's
doc_mapping: text fields tokenized default or stem. A raw-tokenized
text field contributes nothing. Whole-string matching is already served by
Parquet blooms, so a mapping whose only text fields are raw gets no inverted
index however the switches are set.
Update a managed index with ETag and If-Match¶
Use the ETag from GET /api/v1/indexes/{id} as the If-Match condition on
your PUT. This prevents your update from replacing a configuration changed by
another writer since the GET. The PUT body remains the full index
configuration.
Save the current configuration and its validator:
curl -sS -D index.headers -o index.json http://localhost:8089/api/v1/indexes/app-logs
etag="$(sed -n 's/^etag: //Ip' index.headers | tr -d '\r')"
Edit index.json, then send the full configuration with the saved validator:
curl -sS -D - -X PUT -H 'Content-Type: application/json' -H "If-Match: $etag" --data-binary @index.json http://localhost:8089/api/v1/indexes/app-logs
GET and a successful PUT return a strong ETag. Data-only commits keep that
validator, while a configuration change or delete and recreation changes it. A
false condition returns 412 Precondition Failed, the current configuration,
and its matching ETag. Another writer can replace that pair before you
receive it, so a retry can get another 412.
Omitting If-Match preserves the existing update behavior. It does not relax
the additive schema rules. See Index
management for the complete header
grammar and error response.
curl -s http://localhost:8089/api/v1/indexes
curl -sX DELETE http://localhost:8089/api/v1/indexes/app-logs
These operations have no CLI equivalent. siglake sql lists the tables it can
see when you type \d at its prompt, and stops there; --endpoint defaults to
http://localhost:8089.
Deleting an index records its table UUID and cleanup inventory before removing the catalog entry. The record is report-only, so the committed files remain and production cleanup reclaims nothing. Read the HTTP reference before treating a delete as a way to free storage or recreating the same index id. That reference also explains how the WAL isolates the old and new index incarnations.
Updates cannot change field types
Schema evolution is additive only: you can add nullable columns. You cannot drop, rename, reorder, or retype an existing field. If you need a different type, create a new index and migrate.
Limitations¶
- The WAL buffer serves a managed user index on the same terms as
events, and it is off until you set--query-wal-buffer-dir(Helm:query.walBuffer.enabled: true) and the query pods can read the shared WAL mount. Until then an index sees commit-cycle visibility, tens of seconds, rather than seconds-scale freshness.query_auditand the other system tables are not buffered. See Freshness. - Full-text pruning engages only on columns that have index blobs; others fall back to row evaluation.