Skip to content

Configure

engram uses env-first configuration with no viper: every setting is an environment variable (ENGRAM_*). A subset of settings also expose a --flag on engram serve — the server settings (listen address, MCP path), the OIDC/auth settings, and the web-UI lane settings (see the per-section Flag columns below); Qdrant, embedder, and logging variables are env-only. Where a flag exists, it takes precedence over the environment variable.

Environment variable Flag Default Description
ENGRAM_LISTEN_ADDR --listen-addr :8080 TCP address the HTTP server binds to
ENGRAM_MCP_PATH --mcp-path /mcp Path the MCP transport mounts at. / restores the legacy root catch-all (where the transport answered at every path). When the web UI is enabled and this is /mcp, the host root serves the console.
ENGRAM_MCP_RESOURCE_URL — (empty) The public URL clients reach the MCP endpoint on. When set, it becomes the resource field of the served protected-resource document (see Auth & Isolation) verbatim, and no request header can override it. When unset, the value is derived per request from the forwarded/Host headers and the MCP path.

Source: cmd/engram/serve.go (flag registration via internal/config)

Environment variable Flag Default Description
ENGRAM_QDRANT_ADDR — localhost:6334 Qdrant gRPC address (host:port)
ENGRAM_QDRANT_COLLECTION — mem_eval Qdrant collection name for memories (binary default; the Helm chart sets this to memory — see Deploy)
ENGRAM_EMBED_DIM — 1024 Vector dimension; must match the embedding model
ENGRAM_QDRANT_QUANTIZATION — int8 Vector quantization of the memory collection. int8 keeps an int8 scalar copy of every vector pinned in Qdrant’s RAM (quantile 0.99; about 3 KB per record), so a search does not stall re-reading evicted vectors from disk; final scores are still computed from the original vectors. off removes quantization from the collection. unmanaged leaves whatever quantization the collection has untouched — use it to tune quantization by hand. Only engram serve applies this setting; CLI commands never change an existing collection’s quantization, whatever this variable holds in their environment. Any other value fails startup.
ENGRAM_QDRANT_SCHEMA_TIMEOUT — 2m Time budget for startup schema provisioning: the quantization change above plus building any missing payload index. Separate from the fixed 15 s budget for connecting to Qdrant and creating the collection. The server opens its port only after both finish; exceeding either fails startup with an error naming the step. Must be a positive Go duration. The Helm chart derives its startup probe from this value — see Deploy.

Source: internal/server/tools.go (StoreAndEmbedderFromEnvNoEnsure).

Environment variable Flag Default Description
ENGRAM_OPENAI_BASE_URL — http://localhost:4000 OpenAI-compatible embeddings endpoint — point it at any backend that speaks the OpenAI /v1/embeddings API (e.g. Ollama, vLLM, TEI, LiteLLM, OpenAI)
ENGRAM_OPENAI_API_KEY — (empty) API key for the embeddings endpoint
ENGRAM_EMBED_MODEL — ollama/bge-m3 Model name forwarded to the endpoint
ENGRAM_OPENAI_CHAT_BASE_URL — (empty) Base URL for the chat/summarize lane only — the embedder always uses ENGRAM_OPENAI_BASE_URL regardless of this setting. Empty means the summarizer inherits ENGRAM_OPENAI_BASE_URL. Validated only when set: a malformed or non-HTTP(S) value fails startup; empty is always valid. See Auto-summary below for the URL-shape rule and the per-lane credential behavior.
ENGRAM_OPENAI_CHAT_API_KEY — (empty) API key for the chat/summarize lane only — the embedder always uses ENGRAM_OPENAI_API_KEY regardless of this setting. Empty means the summarizer inherits ENGRAM_OPENAI_API_KEY. Never validated at startup: unlike a base URL, an API key has no verifiable shape and empty is meaningful (inherit), so a wrong key fails at the provider rather than at boot. See Auto-summary below for the per-lane credential behavior.
ENGRAM_EMBED_TIMEOUT — 30s Per-request HTTP client timeout for the embeddings call. A non-positive value (including 0) no longer means “no timeout” — it resolves to the ENGRAM_EMBED_MAX_TIMEOUT ceiling below. See guides/upgrade.md if you rely on 0 today.
ENGRAM_EMBED_DRAIN_BYTES — 262144 Byte bound on draining the rest of the response body after a decode, so the underlying connection can be reused. Its only job is connection reuse: abandoning an over-bound body costs one new TCP handshake next time, nothing more. 0 skips the drain entirely — the body is closed immediately and the connection is not reused — for an operator who would rather burn a handshake than ever stall on a drain.
ENGRAM_EMBED_DRAIN_TIMEOUT — 2s Time bound on the same post-response drain, paired with the byte bound above — whichever limit is hit first ends the drain. 0 skips the drain entirely, same as ENGRAM_EMBED_DRAIN_BYTES=0.
ENGRAM_EMBED_MAX_TIMEOUT — 10m Ceiling a non-positive ENGRAM_EMBED_TIMEOUT resolves to. There is deliberately no value meaning “unbounded” — an explicit positive ENGRAM_EMBED_TIMEOUT is still honored uncapped, however large; only a non-positive one is clamped to this ceiling.

Source: internal/config (registry) + internal/server/tools.go (embedderFromConfig, embedTimeout, embedDrainBytes, embedDrainTimeout, embedMaxTimeout).

Environment variable Default Description
ENGRAM_MEMORY_MAX_SUMMARY_BYTES 512 Max byte length of a memory summary on store_memory/schedule_memory/supersede_memory/update_memory. A caller-supplied summary over this bound is rejected (field=summary hint=too_long) rather than silently truncated. 0 disables the bound.
ENGRAM_MEMORY_MAX_CONTENT_BYTES 65536 Max byte length of a memory content on store_memory/schedule_memory/supersede_memory/update_memory (when the content changes), on MCP, Connect and engram store. An over-cap write is rejected (field=content hint=too_long), never truncated.
ENGRAM_MEMORY_MAX_TAGS 128 Max number of tags on those same write paths. An over-cap write is rejected (field=tags hint=too_many).
ENGRAM_MEMORY_MAX_TAG_BYTES 128 Max byte length of one tag on those same write paths. An over-cap tag is rejected (field=tags hint=too_long).

This is separate from ENGRAM_SUMMARY_MAX_CHARS below: this bound is enforced at write time against a caller-authored summary; ENGRAM_SUMMARY_MAX_CHARS caps the length of a server-generated one.

Unlike ENGRAM_MEMORY_MAX_SUMMARY_BYTES, the three caps above are always enforced — 0 or a negative value fails startup, because the server sizes its bounded reads from them and a disabled cap would silently remove that provable ceiling. Existing records larger than a cap are never rewritten — they stay stored and readable (get_memory), and their owner can trim one with update_memory. store_discovery and store_rule keep their own, separate content bounds.

Source: internal/config (registry) + internal/server/tools.go (maxMemorySummaryBytes, memoryWriteCapsFromConfig, checkContentBytes, checkTags, validateStoreArgs/validateUpdateArgs).

When ENGRAM_SUMMARY_MODEL is set, the server digests memories that lack a summary using that chat model. By default it is served by the same OpenAI-compatible endpoint as the embedder (ENGRAM_OPENAI_BASE_URL + ENGRAM_OPENAI_API_KEY) — but the chat/summarize lane can be pointed at a different gateway by setting ENGRAM_OPENAI_CHAT_BASE_URL, independent of the embedder. This unblocks a common split deployment: a local embedder (no egress, no per-token cost) paired with a hosted chat model for summary quality — e.g. embeddings at http://localhost:4000 (a local TEI/Ollama/vLLM server) with summaries at https://api.openai.com/v1. Empty ENGRAM_SUMMARY_MODEL disables auto-summary entirely, and recall returns only client-authored summaries.

Each lane can carry its own API key. ENGRAM_OPENAI_API_KEY is always the embedder’s credential. The chat/summarize lane can carry its own via ENGRAM_OPENAI_CHAT_API_KEY; leaving it empty means the chat lane inherits the embedder’s key and sends it to whatever ENGRAM_OPENAI_CHAT_BASE_URL resolves to. That inherit-by-default behavior is safe for the local-embedder-plus-hosted-chat split above, because local embedding servers (Ollama, TEI, vLLM) simply ignore an Authorization header they don’t expect. It is worth knowing about before you point the chat lane at a hosted gateway, though: leaving ENGRAM_OPENAI_CHAT_API_KEY unset while setting ENGRAM_OPENAI_CHAT_BASE_URL sends your embedding API key to that gateway too. Set ENGRAM_OPENAI_CHAT_API_KEY to opt out and give the chat lane a credential of its own.

In a Helm deployment, set memory.summarize.chatApiKeySecret to render ENGRAM_OPENAI_CHAT_API_KEY into the pod spec as a secretKeyRef (unset omits it, matching the inherit-by-default behavior above).

URL shape matters. Supply the provider’s full OpenAI-compatible root, including its /v1 suffix (or /v1beta/openai for Gemini-compatible gateways), and engram appends the chat-completions path directly — e.g. ENGRAM_OPENAI_CHAT_BASE_URL=https://api.openai.com/v1 resolves to https://api.openai.com/v1/chat/completions. Supply a bare host with no /v1-shaped suffix and engram appends /v1/chat/completions itself — e.g. ENGRAM_OPENAI_CHAT_BASE_URL=http://litellm.internal:4000 resolves to http://litellm.internal:4000/v1/chat/completions. Getting this wrong (appending your own /v1/chat/completions onto a URL that already ends in /v1) is the most likely first failure when configuring this variable.

Environment variable Flag Default Description
ENGRAM_SUMMARY_MODEL — (empty) Chat model for auto-summary; empty disables auto-summary
ENGRAM_SUMMARY_MAX_CHARS — 280 Max generated-summary length (also the recall-truncation cap)
ENGRAM_SUMMARY_TIMEOUT — 30s Per-request HTTP client timeout for the summarize (chat-completions) call. A non-positive value (including 0) no longer means “no timeout” — it resolves to the ENGRAM_SUMMARY_MAX_TIMEOUT ceiling below. See guides/upgrade.md if you rely on 0 today.
ENGRAM_SUMMARY_DRAIN_BYTES — 262144 Byte bound on draining the rest of the response body after a decode, so the underlying connection can be reused. Its only job is connection reuse: abandoning an over-bound body costs one new TCP handshake next time, nothing more. 0 skips the drain entirely — the body is closed immediately and the connection is not reused — for an operator who would rather burn a handshake than ever stall on a drain.
ENGRAM_SUMMARY_DRAIN_TIMEOUT — 2s Time bound on the same post-response drain, paired with the byte bound above — whichever limit is hit first ends the drain. 0 skips the drain entirely, same as ENGRAM_SUMMARY_DRAIN_BYTES=0.
ENGRAM_SUMMARY_MAX_TIMEOUT — 10m Ceiling a non-positive ENGRAM_SUMMARY_TIMEOUT resolves to. There is deliberately no value meaning “unbounded” — an explicit positive ENGRAM_SUMMARY_TIMEOUT is still honored uncapped, however large; only a non-positive one is clamped to this ceiling.

(ENGRAM_OPENAI_CHAT_BASE_URL and ENGRAM_OPENAI_CHAT_API_KEY are documented in Embedder above, alongside ENGRAM_OPENAI_BASE_URL and ENGRAM_OPENAI_API_KEY — this section explains the per-lane credential behavior and the effect the base URL has on the chat/summarize lane specifically.)

In a Helm deployment, set memory.summarize.chatBaseURL to render this variable into the pod spec (unset omits it, matching the inherit-by-default behavior above).

Source: internal/config (registry) + internal/server/tools.go (summarizerFromConfig, summaryTimeout, summaryDrainBytes, summaryDrainTimeout, summaryMaxTimeout) + internal/openaiurl (the shape-aware endpoint join).

By default, store_memory/schedule_memory records without a client-authored summary stay unsummarized until the next engram summarize-missing sweep. Setting ENGRAM_SUMMARY_ON_WRITE=true enables a bounded in-process worker pool that fills the summary asynchronously right after each successful write — a memory typically gets a summary within seconds instead of waiting for the next sweep.

Two-step opt-in. Enabling this is a two-step, per-deployment decision, not a single flag flip:

  1. Set ENGRAM_SUMMARY_MODEL and run task eval:summary (ENGRAM_SUMMARY_EVAL=1 go test ./internal/summarize/ -run TestSummaryFidelity -v) to judge whether your configured model preserves caveats/negations well enough for your data. This is a manual, per-deployment gate — it is intentionally not run in CI, since it needs a live gateway + model and its verdict is judgment, not a pass/fail regression test.
  2. Once you’re satisfied with fidelity, set ENGRAM_SUMMARY_ON_WRITE=true to turn the worker pool on.

Both switches must be true at once for the pool to start (ENGRAM_SUMMARY_MODEL non-empty AND ENGRAM_SUMMARY_ON_WRITE parsing true) — setting ENGRAM_SUMMARY_MODEL alone only enables the summarize-missing sweep, not the async worker.

Environment variable Flag Default Description
ENGRAM_SUMMARY_ON_WRITE — false Enables the async-on-write summary worker pool (requires ENGRAM_SUMMARY_MODEL also set)
ENGRAM_SUMMARY_WORKERS — 2 Worker goroutine pool size draining the enqueue channel
ENGRAM_SUMMARY_QUEUE_SIZE — 256 Bounded enqueue channel capacity

Bounded and non-blocking. The enqueue channel is bounded (ENGRAM_SUMMARY_QUEUE_SIZE); a full queue drops the id and logs a warning instead of blocking the write — the next engram summarize-missing sweep reclaims any dropped or in-flight-at-shutdown records as a backstop. store_memory/schedule_memory always return success once the record is persisted, even when the summarizer is down, slow, or the queue is full: summarization is never on the synchronous write path.

Degradation. If the summary gateway is unreachable or erroring, the write still succeeds and the record simply has no summary yet (“no summary yet”) until a later fill (retry, or the next summarize-missing sweep) succeeds. Recall never fails because of a missing summary — it falls back to truncated content.

Source: internal/config (registry) + internal/server/tools.go (buildSummaryQueue, the D-01 AND-gate) + internal/server/summaryqueue.go (worker pool).

engram can optionally ask an external typed-decision provider (Jev, reached through OpenRouter’s Decisions API) yes/no, multiple-choice, and scored questions about a piece of state, and get back probabilities. It is off by default: set ENGRAM_DECISIONS_PROVIDER=jev to enable it. Answers are advisory — they are surfaced to you and never acted on automatically. Enabling it constructs and validates the client at startup. The provider is asked by engram spine-review consolidate, which attaches an advisory relation verdict to each candidate pair by default whenever ENGRAM_DECISIONS_PROVIDER is set (--no-verdicts skips it), and, when ENGRAM_SEARCH_RANKER=jev (see Search reranking (Jev) below), by every search_memory/search_discovery call, and, unless ENGRAM_SEARCH_UNDERSTANDING=off, by the console’s /search query understanding (see Query understanding (Jev) below).

Base URL. ENGRAM_DECISIONS_BASE_URL is required when the provider is enabled, and it deliberately does not inherit ENGRAM_OPENAI_BASE_URL — the embeddings/chat gateway does not serve Decisions. engram appends /alpha/decisions to whatever you set, so https://openrouter.ai/api resolves to https://openrouter.ai/api/alpha/decisions, and a LiteLLM pass-through at https://litellm.example.com/openrouter resolves to https://litellm.example.com/openrouter/alpha/decisions. The LiteLLM key needs the /openrouter/alpha/decisions pass-through route granted, or the gateway answers 403. The most likely first failure is a base URL ending in /v1 (or /api on the LiteLLM form) — that yields 404.

Key. An empty ENGRAM_DECISIONS_API_KEY inherits ENGRAM_OPENAI_API_KEY and sends it to the decisions host. This mirrors the chat-lane fallback above: if you don’t want the embedder or gateway key reaching the decisions host, set ENGRAM_DECISIONS_API_KEY explicitly. Startup logs api_key_source so you can see which key applied. In Helm this is memory.decisions.apiKeySecret.

Model. Pinned to typesafe/jev-1.13. Do not use the floating ~typesafe/jev-latest — it moves probability thresholds between releases.

What leaves your deployment. Each decision call sends the state and questions a feature builds (for curation and reranking features, that is memory record content) to OpenRouter, which routes Jev to TypeSafe (a service on the US West Coast). For spine-review consolidate, each request carries, per record, up to ENGRAM_DECISIONS_VERDICT_STATE_CHARS characters total of its summary followed by its content. For query understanding, it is the console query text — see Query understanding (Jev) below for exactly what and when. The provider’s policy: no training on inputs, standard retention, and zero data retention not confirmed. Enable this only if that is acceptable for the records in your store.

Failure behavior. Each call is bounded by ENGRAM_DECISIONS_TIMEOUT (one retry on 429/5xx inside that budget) and by response-size and drain bounds. Failures are reported as authentication, bad request, context too large (state plus questions over Jev’s 32k-token context), rate limited, unavailable, timeout, or response too large. A decision failure never fails the operation that asked for it.

Environment variable Flag Default Description
ENGRAM_DECISIONS_PROVIDER — (empty) Decision provider; empty disables typed decisions, jev enables it
ENGRAM_DECISIONS_BASE_URL — (empty) Decisions API base URL; required when the provider is set, never falls back to ENGRAM_OPENAI_BASE_URL
ENGRAM_DECISIONS_API_KEY — (empty) API key for the decisions host; empty inherits ENGRAM_OPENAI_API_KEY
ENGRAM_DECISIONS_MODEL — typesafe/jev-1.13 Decision model; pinned, do not use a floating alias
ENGRAM_DECISIONS_TIMEOUT — 10s Per-request HTTP client timeout for a decision call. A non-positive value (including 0) resolves to the ENGRAM_DECISIONS_MAX_TIMEOUT ceiling below
ENGRAM_DECISIONS_MAX_TIMEOUT — 10m Ceiling a non-positive ENGRAM_DECISIONS_TIMEOUT resolves to. There is deliberately no value meaning “unbounded”
ENGRAM_DECISIONS_DRAIN_BYTES — 262144 Byte bound on draining the rest of the response body after a decode, so the underlying connection can be reused. 0 skips the drain entirely
ENGRAM_DECISIONS_DRAIN_TIMEOUT — 2s Time bound on the same post-response drain, paired with the byte bound above. 0 skips the drain entirely
ENGRAM_DECISIONS_CONCURRENCY — 4 Caps how many decision calls one batch runs at once
ENGRAM_DECISIONS_VERDICT_THRESHOLD — 0.9 The probability below which a spine-review consolidate verdict is marked needs_review; a probability between 0 and 1. --verdict-threshold overrides it for one run
ENGRAM_DECISIONS_VERDICT_STATE_CHARS — 1500 How many characters of each record (its summary, then the head of its content) a spine-review consolidate verdict request sends; a positive integer

Source: internal/config (registry) + internal/decide/jev (the Jev client) + internal/server/decider.go (deciderFromConfig, the ENGRAM_OPENAI_API_KEY fallback).

search_memory and search_discovery can optionally be reordered by the typed-decision provider’s probability that each candidate record answers the query, instead of the default lexical-overlap ranking. It is off by default: set ENGRAM_SEARCH_RANKER=jev to enable it. Enabling it reorders results and adds a per-hit relevance value; it requires ENGRAM_DECISIONS_PROVIDER to already be set and reuses its base URL, key and model — there is no separate provider selector for search. Lexical stays the default ranker.

What leaves your deployment. On EVERY search while enabled: the query and, for up to 100 candidate records the caller can already read, each record’s summary followed by the head of its content, 600 characters each (shrunk further if needed to fit the provider’s context). This is the same provider and data policy as the Typed decisions section above — see its What leaves your deployment paragraph for the destination and retention policy.

Failure behavior. One request per search, bounded by ENGRAM_SEARCH_RERANK_TIMEOUT, with no retry (unlike the consolidate path’s single retry). On any failure — timeout, error, or malformed answer — the search still succeeds and falls back to the default lexical order, with no relevance values attached. Startup logs search reranking enabled when the ranker is jev.

Environment variable Flag Default Description
ENGRAM_SEARCH_RANKER — lexical Search-path reranker; empty or lexical keeps today’s ranking, jev reorders by relevance (requires ENGRAM_DECISIONS_PROVIDER)
ENGRAM_SEARCH_RERANK_TIMEOUT — 2s Per-search decision call timeout; dedicated to the search path (never shared with ENGRAM_DECISIONS_TIMEOUT). Must be strictly positive when the ranker is jev
ENGRAM_SEARCH_RERANK_AUDIT — false Opt-in audit capture: every reranked search logs its query text and candidate ids/ranks/scores (never content) for offline grading. See Telemetry below before enabling

Source: internal/config (registry) + internal/server/decider.go (searchDeciderFromConfig, searchRankHook, searchRerankAudit) + internal/decide/jev (WithNoRetry).

Two layers let an operator judge whether reranking is earning its cost.

Always on. Whenever a rank hook ran, the search’s own span (tool/search_memory, tool/search_discovery, or the Connect RPC span) carries engram.rerank.* attributes next to the decide child span’s cost and latency. None of them carries the query, a record id, or record content; when the ranker is off none of them is present.

Attribute Type Meaning
engram.rerank.outcome string applied (scores reordered the pool), fallback (hook failed or scores rejected; the pre-hook order shipped), skipped (empty pool, hook never called)
engram.rerank.fallback_class string Why it fell back: the hook’s own class word (timeout, unavailable, rate_limited, auth, context_too_large, malformed_response, state_budget, no_candidates, …) or the store’s rejected_scores / no_scores
engram.rerank.top1_changed bool The first hit the caller received differs from the pre-hook order’s first hit (applied only)
engram.rerank.moved int Positions within the caller’s k whose id differs from the pre-hook order (applied only)
engram.rerank.promoted int Ids within the caller’s k that the pre-hook order had beyond k — what the candidate over-fetch bought (applied only)
engram.rerank.relevance_max float Highest relevance returned; a value near zero across a search means nothing returned answered the query (applied only)

Opt-in audit capture. ENGRAM_SEARCH_RERANK_AUDIT=true additionally emits one search rerank audit info line per reranked search carrying surface, query (the text — this is the one place engram deliberately logs it), owner, k, pool, outcome, and candidates: a JSON array over the whole candidate pool of {id, before_rank, after_rank, cosine, relevance} (ranks 1-based; before_rank is the lexical order for search_memory and the vector order for search_discovery; relevance is absent on fallback). Record content, summaries and tags never appear — a grader fetches them by id via get_memory. It is meant to be turned on for a bounded window, graded offline, and turned off; startup logs a warning while it is on. It has no effect unless the ranker is jev.

The /search page can ask the typed-decision provider to turn a written query into advisory filter suggestions — a “Suggested” row of unapplied chips for category, time window, scope, and tag. Suggestions never change the results themselves until you click (or press Enter/Space on) a chip; accepting one applies the exact same filter a manual category/date-range/scope/tag control would. This runs over the Connect-only UnderstandQuery RPC — there is no MCP tool and no CLI verb for it.

On by default when a decisions provider is configured. Leaving ENGRAM_SEARCH_UNDERSTANDING unset makes query understanding follow ENGRAM_DECISIONS_PROVIDER: once a provider is set — for typed decisions or search reranking — console query text starts flowing to it for filter suggestions too. Set ENGRAM_SEARCH_UNDERSTANDING=off to keep the provider for other features without sending console query text. Setting ENGRAM_SEARCH_UNDERSTANDING=jev explicitly requires ENGRAM_DECISIONS_PROVIDER to already be configured (Config.Validate rejects the combination otherwise). Whenever understanding is effectively on — by default or explicitly — startup logs a Warn: search understanding enabled: console query text is sent to <host>, naming source (default or explicit), the endpoint host, the model, the timeout, and which env var supplied the API key.

What leaves your deployment. For a committed prose query the classifier reads as free text — two or more words, no cat:/scope:/tag:/operator syntax, and not an id or short_id lookup — engram sends: the query text, up to 2000 characters, and, only when no scope is already applied, the names of up to 254 of your own readable scopes as the options of one multiple-choice question (more scopes than that asks no scope question at all). It never records or sends record content or summaries. Tag names never leave your deployment either — a tag suggestion is matched locally against your own tag vocabulary, never asked of the provider. This is the same provider and the same data policy as Typed decisions above — see its What leaves your deployment paragraph for the destination and retention policy.

Failure behavior. Each query makes at most one decision call, bounded by ENGRAM_SEARCH_UNDERSTANDING_TIMEOUT, with no retry. A timeout, an error, or a malformed answer all produce zero suggestions for that call — the search itself never waits for query understanding and never fails because of it. A suggestion is only ever produced at or above a probability of 0.9.

Environment variable Flag Default Description
ENGRAM_SEARCH_UNDERSTANDING — (empty) Query-understanding switch; empty follows ENGRAM_DECISIONS_PROVIDER, off disables it explicitly, jev enables it explicitly (requires ENGRAM_DECISIONS_PROVIDER)
ENGRAM_SEARCH_UNDERSTANDING_TIMEOUT — 2s Per-query decision call timeout; dedicated to this path, one attempt, no retry
ENGRAM_SEARCH_UNDERSTANDING_AUDIT — false Opt-in audit capture: every understood query logs its query text and suggestion labels (never content). See Understanding telemetry below before enabling

Source: internal/config (registry) + internal/server/decider.go (understandingEnabled, understandDeciderFromConfig) + internal/understand (the suggestion core).

Whenever query understanding runs — decided, fell back, or skipped because no question needed asking — the UnderstandQuery RPC’s own span carries bounded engram.understand.* attributes; none of them carries the query, a scope, or a tag. They are absent entirely when understanding is off.

Attribute Type Meaning
engram.understand.outcome string decided (the call succeeded), fallback (the call or its answer failed), skipped (no question needed asking, or no decider configured)
engram.understand.fallback_class string Why it fell back (only present on fallback) — the same class-word vocabulary as engram.rerank.fallback_class
engram.understand.suggestion_count int Number of suggestions returned (locally-matched tag suggestions count too)
engram.understand.questions_asked int Number of questions the built request carried, regardless of outcome

Opt-in audit capture. ENGRAM_SEARCH_UNDERSTANDING_AUDIT=true additionally emits one query understanding audit info line per understood query, carrying query (the text — this is the one place this path deliberately logs it), outcome, questions_asked, fallback_class (only when set), and suggestions — a JSON array of every suggestion’s label in response order. Record content never appears. Startup logs a warning while it is on; setting it while query understanding is off logs a warning that nothing is actually audited.

Setting ENGRAM_OIDC_ISSUER enables bearer-token enforcement (JWKS signature + issuer + expiry validation). Without it, all requests are accepted and a loud warning is logged.

Environment variable Flag Default Description
ENGRAM_OIDC_ISSUER --oidc-issuer (empty) OIDC issuer URL; setting it enables bearer-token enforcement
ENGRAM_OIDC_AUDIENCE --oidc-audience (empty) Expected aud claim (optional; omit to skip audience check)
ENGRAM_OIDC_RESOURCE_METADATA --oidc-resource-metadata (empty) WWW-Authenticate resource metadata URL returned in 401 responses (optional). Left empty with ENGRAM_MCP_RESOURCE_URL set, this now defaults to the protected-resource document path engram itself serves; explicit config still wins.
ENGRAM_OWNER_CLAIM --owner-claim email OIDC claim whose value becomes the record owner (authz key); fail-closed if absent; requires email_verified when email

Source: cmd/engram/serve.go (init() flag registration and buildAuthChain, which composes the chain and wraps it in auth.EnforceExpiry so token expiry is enforced on the composed chain rather than only inside the MCP bearer wrapper).

A headless service principal (CI runner, batch job, another backend service) authenticates over a third lane composed alongside the human OIDC lane above — each mechanism activates independently, based on its own config being present. Env-only (secret-bearing, no --flag equivalents):

Environment variable Default Description
ENGRAM_SERVICE_AUTH_OIDC_ISSUER (empty — lane off) Client-credentials OIDC issuer URL (may reuse ENGRAM_OIDC_ISSUER’s IdP or be distinct)
ENGRAM_SERVICE_AUTH_OIDC_AUDIENCE (empty) Expected aud claim for the service lane, independent of ENGRAM_OIDC_AUDIENCE
ENGRAM_SERVICE_AUTH_OWNER_CLAIMS client_id,azp Ordered claim list the service lane resolves an owner from — never defaults to email
ENGRAM_SERVICE_AUTH_STATIC_TOKENS (empty — lane off) Comma-separated owner=token pairs, e.g. ci=tok-abc,batch=tok-def

A deployment with none of these set is unchanged: only the human OIDC lane (or no auth at all) is active. Static tokens have no revocation list — rotating means editing ENGRAM_SERVICE_AUTH_STATIC_TOKENS and restarting; see reference/auth.md (Service principals) for the fail-closed empty-owner guarantee, the no-revocation kill-switch, and the global cross-tenant shared-read decision (docs/adr/engram-svct-service-tenant-global-shared-read.md).

Source: internal/config (ServiceAuthConfig, service_auth.* registry rows) + cmd/engram/serve.go (buildAuthChain).

Environment variable Flag Default Description
ENGRAM_CONNECT_HEADLESS --connect-headless false Mounts the ConnectRPC lane on a deployment with the web UI disabled

ENGRAM_CONNECT_HEADLESS defaults off and is independent of every ENGRAM_UI_* and ENGRAM_SERVICE_AUTH_* variable — a deployment with no Connect surface today gains none on upgrade, including one that already has service-auth configured; configuring an auth lane never mounts Connect by itself.

Connect is mounted when the web UI is enabled or this flag is set. With either (or both), one lane serves both credential types: a well-formed Authorization: Bearer credential authenticates against the same verifier chain the MCP lane uses, and everything else falls through to the session-cookie lane.

Setting ENGRAM_CONNECT_HEADLESS with no auth lane configured (no ENGRAM_OIDC_ISSUER and no ENGRAM_SERVICE_AUTH_*) refuses to start — mounting would expose every write RPC unauthenticated into the anonymous empty-owner bucket. Configure at least one auth lane first: either ENGRAM_OIDC_ISSUER (the human OIDC lane) or ENGRAM_SERVICE_AUTH_* (client-credentials OIDC or static tokens, see Service principals above).

A bearer-authenticated Connect caller is exempt from the X-CSRF-Token double-submit check that cookie-authenticated browser callers must satisfy; the exemption is decided by which lane verified the request, never by which headers the caller sent.

Source: internal/config (the connect.headless registry row) + cmd/engram/serve.go (connectHeadlessGuard, connectResolverFor) + internal/server/connectbearer.go.

Environment variable Default Description
ENGRAM_LOG_LEVEL info Log level: debug, info, warn, error
ENGRAM_LOG_FORMAT json Log format: json or text
ENGRAM_LOG_STDOUT true Write logs to stdout; set false to suppress (requires OTLP endpoint)

Source: internal/telemetry/config.go (ConfigFromEnv).

Authorization decision diagnostics. At debug level, every authorization decision (both allow and deny) emits one "authz decision (bucket)" or "authz decision (record)" line carrying allow, action, the satisfied Cedar policy_ids, and policy_error_count; bucket-scoped decisions also carry bucket (own/shared). Volume is bounded: at most two lines per bulk recall call (one per bucket probed) and one per id-addressed operation (get/update/delete/set_visibility) — never one line per result row.

It does not carry a full Cedar expression trace, any policy error message text, or the caller’s owner/scope values — those are deliberately excluded so a decision line is always safe to leave in a log pipeline. Raise ENGRAM_LOG_LEVEL=debug to see these lines; they do not appear at info or above.

OTLP export is enabled when OTEL_EXPORTER_OTLP_ENDPOINT is set (standard OpenTelemetry env var, not ENGRAM_*). When it is empty, providers are no-ops.

The Helm chart exposes observability.otlpEndpoint, observability.otlpHeaders, and related values that the chart maps to the appropriate environment variables for the pod.

flag (--oidc-issuer) > environment variable (ENGRAM_OIDC_ISSUER) > built-in default

No viper, no config file. The Helm chart sets ENGRAM_* variables in the pod spec; use --set or a valuesObject override for cluster-specific values.