Report Generator (local LLM)¶
Docker is measured; Azure is not
Everything measured on this page was measured on Docker (one 2-vCPU VM). The Azure templates compile and are wired up, but the report generator has not yet been deployed to Azure: scale-to-zero, the prompt cache on Azure Files, the request time limits described under Azure and both network modes are unproven there. Deploy it on Azure only to try it out.
Experimental, and optional
The report generator is an extra container you choose to deploy. Without it, custom reports are still built by hand — see Custom Reports. The feature as a whole is off until an operator enables it under Admin → Experimental.
The table predates the change entity, grouping and the chat; both sets were re-run on 23 September 2026
The table below was measured against a system prompt without the change entity and
without counting per value. On 23 September 2026, on a 2-CPU host with real directory data
and the corrections described under What the mistakes look like:
- held-out set: 15/17 (13/15 with a non-empty answer), median 74 s, p90 109 s, one question over the 5-minute limit (a subset comparison that failed after two rounds);
- conversation set (
chat.json, 58 graded answers): 58/58 — the questions people typed when asking in plain language, in Dutch and English, 10 follow-ups in the same conversation, 10 out-of-scope requests declined, 8 questions about data we do not have asked back or declined, 2 about a person who does not exist asked back, 2 genuinely ambiguous ones asked back, 2 counts kept; median 75 s, p90 168 s, slowest 293 s, none over 5 minutes. The number is the sum of one full run (51/58) and a re-run of the seven rows whose fixes landed during it, on the same build.
The chat's first run that day scored 28/58 with a median of 77 s and two answers over the limit; everything between those two numbers is in the list below, none of it a change of model. The latency figures further down are from the tuning host.
Why this exists¶
Every customer asks the same kind of question about their own environment: guest accounts without a manager, groups nobody owns, accounts that still sit in a licence group. Each answer is a different query. Shipping a built-in report for every one of them does not scale, and teaching analysts a query builder covers only the people willing to learn it.
So the analyst describes the report and Identity Atlas builds it — but the model never writes the query. It fills in a report definition: which kind of record, which conditions, which columns. Identity Atlas validates that definition against its own catalog and compiles it to read-only SQL.
That split is the whole design, and it is what makes the feature defensible:
| Because the model only fills in a definition | |
|---|---|
| The model never sees your data | It gets the question, the list of fields it may use, and the type names that exist (e.g. Guest, ServicePrincipal). No rows, ever. |
| It cannot do damage | Unknown fields, unknown operators and unknown values are rejected. The query is parameterised, runs in a read-only transaction with a statement timeout, and touches only the tables the catalog names. |
| You can check it | The definition is shown back as editable criteria plus a plain-language sentence generated from what will actually run. |
| It is repeatable | A saved report is a definition, not a captured answer. It runs again after the next crawl and gives today's answer. |
| It works without the model | The same definitions can be built by hand, which is what every deployment without the extra container does. |
The model¶
| Model | Qwen3-4B-Instruct-2507, 4-bit quantised (Q4_K_M), ~2.5 GB |
| Licence | Apache 2.0 — commercial use permitted. The licence text ships inside the image at /models/MODEL-LICENSE.txt, with an attribution notice. llama.cpp (MIT) ships its licence next to it |
| Source | Pinned URL and SHA-256 in setup/docker/Dockerfile.report-generator; the build fails if the file does not match |
| Runtime | llama.cpp server, CPU only. No GPU, no external API, no telemetry |
| Chosen by | The release. There is no model picker: Admin → LLM shows which model this version ships and whether it is answering |
The model is part of the release so that a given Identity Atlas version always behaves the same way and we can state what it was measured at. Changing the model is a new release.
Models we evaluated¶
Measured against a set of real analyst questions on a real tenant (Fortigi: ~1,150 accounts, 158 groups, 10 business roles), plus a held-out set written before tuning and never used to improve prompts. A question counts as correct only when the generated report returns exactly the same rows as a hand-written reference definition.
Only the first row was measured on the setup that ships (llama.cpp, the final prompt, the bounded grammar); its raw results are kept with the evaluation tools. The other rows were measured earlier, on a different model runtime (Ollama) and an earlier version of the prompt, so read them as the reason each model was dropped, not as a like-for-like ranking.
| Model | Licence | Tuning set | Held-out set | Notes |
|---|---|---|---|---|
| Qwen3-4B-Instruct-2507 (shipped) | Apache 2.0 | 34/42 (81%) | 14/17 (82%) | Best of every model tested, commercial use allowed |
| Qwen2.5-Coder 3B | ⛔ Qwen Research (non-commercial) | 28/32 (88%) at the time | 9/14 (64%) | Strongest early candidate; cannot be shipped |
| Qwen2.5-Coder 1.5B | Apache 2.0 | 21/36 (58%) | 8/14 (57%) | Runs on 1 CPU / 2 GB, but clearly weaker |
| Qwen2.5-Coder 0.5B | Apache 2.0 | 7/32 (22%) | 5/14 (36%) | Unusable: copies the examples |
| Phi-4-mini (3.8B) | MIT | 11/32 (34%) | — | |
| Gemma 3 4B | Gemma terms | run abandoned | — | ~80 s per question: its attention design defeated prompt caching |
| Qwen2.5-Coder 7B | Apache 2.0 | spot checks only | — | Too slow on 2 CPUs for the benefit |
The sets grew during development (32 → 42 tuning, 14 → 17 held-out), so the percentages above are
comparable within a row, not exactly across rows. The shipped model's figures are the full sets,
measured on the release image as it ships — llama.cpp, the saved prompt cache, and the bounded output
grammar described below. Before the sign-in fields were added, the same setup scored 34/39 and 14/17; the
larger prompt answers the new sign-in questions but moved two unrelated answers (33/41 and 13/17), and each
prompt variant tried since traded one question for another. Looking up the names a question mentions
(see Privacy) won one held-out question back without losing a tuning question; the tuning
set's organisation question (42nd) was added with it and measured on its own. A run is repeatable — the
same prompt gives the same answers — so these differences are real, not noise; the sets are just too
small to tune further without fitting them. A few questions have an empty correct answer, which a wrong
definition can also produce; counting only questions with a non-empty answer, the shipped model scores
33/41 and 12/15. Re-run any time with tools/nl-reports/eval.mjs —
see Measuring it yourself.
The held-out run of 23 September 2026 (final build) missed two: a subset comparison ("business roles
whose members are all in group X"), which remains the weakest kind of question, and one over-specified
name filter (name and description must contain "License", one group differs). The conversation set
missed nothing on that build; the categories it grades — scope, unknown data, a person who does not
exist, an ambiguous question, follow-ups — are the ones a chat gets wrong in ways a report page never
shows, and each has its own row in tools/nl-reports/chat.json.
Larger models on more CPU (24 September 2026)¶
The question behind this run: is the 4B model the ceiling, or does a bigger model on a bigger
CPU budget answer more questions right? Measured on a fresh install in Azure (one
Standard_E8bds_v5 VM — 8 vCPU, 64 GB, Docker, the same compose files a customer gets), with a
copy of the same tenant, the same prompt and the same two sets: the held-out set (17 questions) and
the conversation set (72 graded turns: 64 questions in Dutch and English, 12 follow-ups, and the
scope / unknown-data / unknown-person / ambiguous rows). Every model runs through the same
pipeline, so the corrections it makes on the model's behalf count for all of them. The VM's cores are
about 1.3× faster than the reference host's (median 58 s against 74 s for the same run), so read the
times relative to each other.
| Model | RAM for the model server | Threads | Held-out | Conversation set | NL / EN | Follow-ups | Median | p90 | Slowest | Over 5 min |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3-4B-Instruct-2507 (shipped) | 4 GB | 2 | 14/17 | 67/72 | 35/36 · 32/36 | 11/12 | 68 s | 143 s | 352 s | 1 |
| Qwen3-8B (thinking off) | 10 GB | 4 | 13/17 | 63/72 | 32/36 · 31/36 | 10/12 | 45 s | 102 s | 250 s | 0 |
| Qwen3-14B (thinking off) | 22 GB | 8 | 14/17 | 64/72 | 31/36 · 33/36 | 11/12 | 53 s | 102 s | 160 s | 0 |
| Qwen3-30B-A3B-Instruct-2507 | 24 GB | 8 | 16/17 | 69/72 | 34/36 · 35/36 | 12/12 | 34 s | 80 s | 191 s | 0 |
| Gemma 3 12B | 20 GB | 8 | abandoned | — | — | — | ~355 s | — | — | every question |
| Qwen3-4B-Instruct-2507, more threads | 4 GB | 8 | 14/17 | 68/72 | 35/36 · 33/36 | 11/12 | 32 s | 72 s | 113 s | 0 |
What the table says:
- The 30B-A3B model is the one that helps. It is a mixture-of-experts model: 30B parameters on disk, 3B active per token, so it reads like a big model and writes at the speed of a small one — the fastest tier here and the most accurate, with every follow-up right and both languages level. Its three misses on the conversation set are two readings of "the groups of Anna's direct reports" (a relation inside a relation, which the definition language cannot express) and one "access packages" question answered as every resource — corrected by a pipeline rule since. What it needs: about 24 GB of memory for the model server (the 18.6 GB model file plus the context; measured 14 GB resident with the rest in the page cache) and 8 CPU threads; at 2 threads it would be about four times slower.
- 8B and 14B do not help. They are the older hybrid generation (April 2025) run with thinking off, because thinking tokens at CPU speed would blow the time limit; in that mode they are no more accurate than the 4B Instruct-2507 and make different mistakes (dropping the person from a change question, leaving out a principal type). More threads make them faster, not better.
- Gemma 3 does not work with this design at all: llama.cpp cannot reuse the saved prompt cache for its sliding-window attention, so every question re-reads the whole prompt — about six minutes each, the same finding as with Gemma 3 4B earlier.
- The 4B is not far behind. Its misses are model wobble: the same question passes on one run and fails on the next (13 to 15 of 17 across runs today; the conversation set gave 66, 67 and 68 of 72 on three runs), an invented condition, a name filter also applied to the description. That is what the correction rules in the pipeline exist for, and it is why the 30B's cleaner first attempts show up more in the follow-ups and in the repair count (five repair rounds in 72 turns against eight) than in the raw score.
The "License" question every model misses (name and description must contain the word — one group differs) and the subset comparison remain the two standing misses of the held-out set.
To run the same matrix: build the images with the other models' URL and checksum
(setup/docker/Dockerfile.report-generator takes them as build arguments), set
REPORT_GENERATOR_CPUS / REPORT_GENERATOR_MEMORY, and run tools/nl-reports/eval.mjs for
holdout.json and chat.json — see Measuring it yourself.
What the mistakes look like¶
About one question in seven comes back wrong, so the point is not perfection but visible mistakes. Typical failures, all of which show up in the plain-language reading:
- a condition dropped ("groups mostly like Sales, not in a business role" lost the second half);
- the wrong relation ("contains all members of business role X" read as "is part of business role X");
- a guess instead of a question: "everyone with admin rights" came back as one kind of admin. It asked on neither of the two deliberately ambiguous questions.
Comparisons are the weakest kind of question: 3 of 6 correct across both sets. The one this feature was demonstrated with ("groups with the same members as business role ACME - Algemeen - Partners, not part of it") is among the three that pass. For a comparison that matters, build it with + compare with… in the editor — once built, a comparison is exact; only the translation from words is uncertain.
Several failure modes are handled in code rather than left to the model. Two of them were once handed back to the model to fix and are now corrected on the spot — at about a token a second on the CPU box a correction round costs one to three minutes, and the model's correction was not reliably better (asked to fix "added AND removed" it dropped "removed" and the 90-day window with it). Every correction made this way is stated in the report's assumptions.
- A mistake with one sensible reading is corrected without asking — "action is Added AND action is Removed" becomes either; "owner count above zero" beside "has no owner" loses the count; "within the last 90 days" beside "more than 180 days ago" keeps the recent bound; and a count-per-value grouping the question never asked for is removed ("which groups was he added to" once came back as the number 5). A mistake with two readings ("empty and not empty") is still put to the model.
- The person asking is put where they belong — "van welke groepen ben ik eigenaar" whose definition names nobody gets the caller added (into an empty owners/members relation, or on the account itself); "which groups am I in" with the caller's id written inside the group condition is moved to the account; "groups I have that bram does not" with bram on both sides puts the caller on the side the question mentions first. Only when the definition names somebody else is the model asked once to add the caller.
- Follow-up bookkeeping is corrected, not trusted — the ids of "these groups" written on the members relation, on a user report's own id, or on a change's account are moved to where that kind of record is reached; "id is [a list]" is read as "one of"; a name written as an id is a name.
- A kind of thing is not a name — "in an access package" written as a comparison with a reference named "access package", or as a resource type on a group or its members, is the business-role relation; the column a question asks to see (members, owners, groups, access packages, manager) is added when the definition left it out.
- An invented condition goes, an asked-for one does not — a condition validation rejects is checked against the question: one nothing in the question asks for ("accountCount > 0" inside members) is dropped with a note; one the question did ask for ("MFA") goes to the correction round, and a definition without it is never run.
- A loop stops early — the reply grammar allows eight conditions per list (the largest measured answer needs four) and 450 tokens; a model repeating one condition ran 13 minutes to the old cap. Repeated conditions collapse to one.
- A correction may only fix what it was told — every correction round (invalid definition, "or", a name not used, the caller missing) compares the corrected definition with the original leaf by leaf; one that lost a condition it was not told about is refused and the original stands. The "or" correction once added "added or removed" and dropped the 90-day window, turning 2 rows into 149.
- A first name is one person — "bram" written as a name-contains filter is looked up before the report runs: one match is pinned to that person (the reading says who), several are offered as a choice, none asks for the exact name. On a small directory the substring happened to be right; on a large one it counts every Bram.
-
Out of scope is declined, not guessed — a request that is not about the data (general knowledge, small talk, writing) or that asks to change access is answered with one sentence and no report, in every front end, and filed as
declinedrather than as a failure. -
"or" read as "and" — if the question contains or but the definition has no any-group, the generator asks the model once to correct it, and keeps the original if the correction is no better.
- Invented values — a definition with an unknown field, operator or type value is rejected, the validator's own message is handed back to the model, and it gets one attempt to fix it. If the corrected definition is still invalid, the analyst gets an error — never a report quietly built from the conditions that did validate.
- Repeating itself — at temperature 0 a small model that starts repeating does not stop. Every list and every piece of free text in the reply has a hard length in the output grammar, so a loop ends where validation would have cut it anyway. Before that limit existed, one question listed the same ten columns until the token cap: 570 seconds and a reply that was no longer JSON. It now answers in 46.
For the chat on top of this — the flow of one question, the prompt verbatim, what is answered and what is refused, the corrections made on the model's behalf and the conversation-set results — see The Ask assistant.
Who can ask, and where¶
The same model answers questions in three places, and they are gated by two different permissions on purpose.
| Surface | Permission | What it is for |
|---|---|---|
| Ask tab | data.read.reports |
Type a question, read an answer. No definition editor, no saving. It knows who is asking, so "my groups" and "mijn medewerkers" mean you. |
| Custom reports builder | data.write.reports | Build, edit, save and delete report definitions everyone sees. |
Asking and building are separate rights and neither implies the other. A pilot manager who
should be able to find out who has access to what does not thereby get to delete the saved
reports an analyst depends on; an analyst holds both, so nothing they could do before
changes. Warming the model stays with data.write.reports — one model server, one slot,
shared by everyone, so spending its CPU is not a read action.
Every question asked in plain language is kept in the conversation store for the store's retention
period (90 days by default, NL_REPORTS_LOG_RETENTION_DAYS), per person: the question, what the
model was told, what it replied, the definition that ran and how the answer ended. That log is what
makes a wrong answer diagnosable after the fact, and what the evaluation reads. Result rows are
never stored — only the definition that produced them, so re-running a conversation re-runs the
query rather than replaying stale data. Nobody sees anyone else's conversations.
Asking is gated by data.read.reports and the custom reports feature flag, so a deployment
that has not switched the feature on has no plain-language endpoints at all.
Privacy¶
- Nothing leaves the deployment. The model runs in a container next to Identity Atlas. There is no cloud model, no API key to a provider, and no telemetry in this path.
- The model receives: the analyst's question, the catalog (entity, field and relation names with their descriptions), the type values that exist in this deployment (account types, resource types, system names), while refining the current report definition, for a name the question mentions — which fields that name occurs in — and, for an attribute the question names, the name of that attribute.
-
Attribute names, never attribute values. A deployment's own
extendedAttributeskeys (sfDepartmentID, an OU path) are not in the catalog: they differ per tenant, and there can be hundreds. When a question names one, the API matches it against the keys discovered in the data and tells the model the field name to use —"sfDepartmentID" is the field ext.sfDepartmentID. No value of that attribute is looked up, sent, or counted. An attribute the question does not name is not sent at all: it goes with the question, so the system prompt itself never changes between questions and the saved prompt cache keeps working (see below). -
That last one is the only thing looked up in the data for the model, and it is deliberately narrow. When a question names something ("guest accounts from Contoso"), the API checks, per name, whether it occurs as a whole word in a fixed set of text fields (name, email, company, department, job title, description) and in system names. The model is told the field names only —
"Contoso": user.companyName, user.email— never a row, a value from the row, or a count. It learns that a word the analyst typed exists in the data, which the analyst already implied by asking. Without it, the model guessed where an organisation name lives and filtered on an unrelated system. - The model never receives: rows, query results, counts, or any value it did not get from the analyst. Name lookups ("did you mean…") are done by the database, not by the model.
- Names do reach it in one way, and it is worth being exact about it. Whatever the analyst types is sent as typed, names included. And once an analyst confirms a "did you mean", the confirmed record's name is written into the definition — so if they then refine the report in words, that definition, with that name, goes back to the model. Nothing leaves the deployment either way; it is the model inside it that sees the name.
- On Docker the container has no published port, sits on an internal network shared only with the web container, and has no route to the internet. It does not need one: the model is inside the image.
- On Azure the Container App has public ingress, protected by a per-deployment API key (generated by the template, never entered by hand). In the default public network mode it is also narrowed to the web app's outbound addresses; in the private network mode it is protected by the key alone (see below). It holds no data and no credentials.
- Audit: every question is logged by the API twice — on arrival, with the user who asked, the model and the question; and when it is answered, with the outcome (report, clarification, "did you mean", error, or failed) and how long it took. The reply itself is not logged. Saved reports record who created and last changed them.
- The audit line contains the question as typed, which means it can contain a name an analyst typed ("groups like Jan de Vries"). That is deliberate — an audit trail of "someone asked something" is not worth keeping — but it is worth knowing when deciding how long to keep container logs. No query result is ever logged.
Sizing and consumption¶
Measured on a 2-vCPU VM (shared Proxmox host, Intel Core Ultra 5), with the shipped model:
| Value | |
|---|---|
| CPU | 2. Only 2 CPUs were measured; answer time scales roughly with CPUs, so fewer is slower |
| Memory | 3.2 GB in use and not reclaimable after all 56 questions; 3.65 GB peak including file cache. Limit 4 GiB on both Docker and Azure — the five heaviest questions were re-run at exactly 4 GiB with identical results and no out-of-memory kill. Lower is not measured |
| Disk | 2.76 GB image + 561 MB prompt cache |
| Restart → first answer | 76 s measured for "guest accounts without a manager, or whose manager is disabled": prompt cache restored in 0.1 s, 203 of 4,000 prompt tokens actually read, the rest is the answer being written. On Azure add the container start |
| Question once warm | median 49 s, p90 107 s, slowest 156 s over the tuning set (held-out: median 48 s, p90 78 s). The slow ones are the questions that needed a correction round |
| One-time preparation | 193–205 s measured, in the background, after an install or update |
| Idle | 10.6 MB once the model is unloaded — after 15 minutes unused by default. While it is loaded and idle: no CPU, ~3 GB |
| Unloaded → ready | 3.0 s with the model file in the disk cache, 15.6 s straight from disk. The builder shows "loading the model into memory" with a timer meanwhile, and the first question afterwards is as fast as ever (76 s measured, same prompt cache) |
Nearly all of an answer's time is the model writing the definition, at about 2.5 tokens a second on 2 CPUs; an average reply is ~105 tokens. More CPUs is what makes answers faster — more memory does not.
It only costs while it is used in the sense that matters for each platform:
- Docker: the model is loaded only while it is used. Opening the report builder loads it (3–16 s);
after
REPORT_GENERATOR_IDLE_SECONDSwithout a question (900 by default) it is unloaded and the container drops to ~10 MB. A host running the generator therefore needs the 4 GB only while someone is building reports. Set the idle time to0to keep the model loaded all the time. - Azure: the Container App scales to zero. Azure bills per second of activity, so an idle generator costs nothing beyond its share of the file share holding the prompt cache. At list prices, 2 vCPU + 4 GiB active costs in the order of tens of euro cents an hour, and Container Apps' monthly free grant covers light use — an estimate from list prices, not a measured bill; check current pricing. The trade-off is a slower first question after idle, because the container has to start (a 2.76 GB image pull included).
How the cold start was made survivable¶
The expensive part is not loading the model, it is the model reading its instructions (~4,000 tokens) — minutes on a small CPU. Three things fix that:
- The instructions are the same for every deployment of a release. Everything deployment-specific (your account types, resource types and system names) is sent with the question instead of being baked into them.
- The processed instructions are saved to disk (llama.cpp slot cache) and restored in ~0.1 s on every later start. That is measured on local disk; on Azure the 561 MB file lives on an Azure Files share, whose restore time has not been measured. A second file holds the instructions plus this deployment's value lists — the first thing in every user message — derived from the first file in seconds (restore it, read a few hundred tokens, save). Questions start from that one, so the server reads only what follows the lists: the caller, the name hints, the question. When the lists change, the next warm-up writes a fresh file; until then the server reuses what still matches, which is the instructions.
- Preparation runs in the background at API startup, and the builder says "preparing" instead of
blocking.
node tools/nl-reports/prepare-prompt-cache.mjsdoes it on demand and verifies a restore actually works.
Measured on the same hardware: 266 seconds for the first answer without a restored cache, 76 seconds with one.
One detail made the difference between this working and not: the prompt cache is re-checked before every question, never remembered. The model server restarts on its own — Azure scales it to zero between questions — and comes back empty. An earlier version remembered that it had warmed up, never restored after such a restart, and answered the next question in 266 seconds.
Deploying it¶
Docker¶
The model server is an opt-in profile in docker-compose.prod.yml:
# in .env
COMPOSE_PROFILES=report-generator
FEATURE_CUSTOM_REPORTS=true # or switch it on in Admin → Experimental
docker compose -f docker-compose.prod.yml up -d --pull always
Optional knobs (defaults shown): REPORT_GENERATOR_CPUS=2, REPORT_GENERATOR_MEMORY=4g,
REPORT_GENERATOR_IDLE_SECONDS=900 (seconds unused before the model is unloaded; 0 = always loaded).
How the model is loaded on demand: a small supervisor is the container's main process. It starts the model server (reachable only from inside the container) when a request needs it, passes every request through unchanged, and stops it when idle. It needs no extra privileges — it never touches Docker — and it checks the API key before it starts anything.
Leave COMPOSE_PROFILES unset and nothing extra is pulled or started; the builder then reports the
generator as unavailable and the definition editor still works.
Behind a reverse proxy, raise its read timeout for /api/nl-reports/interpret. A question is one
HTTP request that is answered when the model is done — a median of 49 s and up to several minutes — and
common defaults (nginx proxy_read_timeout 60 s) cut half of them off. The API itself waits up to
15 minutes (NL_REPORTS_LLM_TIMEOUT_MS).
Azure¶
The report generator is an opt-in Container App:
or set deployReportGenerator=true when deploying azure/main.bicep / the Deploy-to-Azure template.
The template then:
- creates the Container App with minReplicas 0 (scale to zero) and 2 vCPU / 4 GiB;
- generates an API key per deployment and gives it to both sides — llama.cpp refuses every request without it;
- mounts an Azure Files share for the prompt cache, so a scaled-to-zero app restarts fast;
- sets
FEATURE_CUSTOM_REPORTS=trueon the web app, because a deployment that paid for the container wants the feature.
Everything else (App Service, Postgres, worker) is unchanged. Leaving the switch off deploys exactly what it does today.
Known limits on Azure, not yet measured there:
- Request time limits. A question is one blocking HTTP request, and Azure closes those: App Service at about 230 seconds for the browser's request, Container Apps ingress at 240 seconds for the web app's call to the generator. Measured on Docker, a warm question takes 49 s at the median and 165 s at the slowest, so most fit; but the one-time prompt-cache preparation took 193–205 s, close to the 240 s limit, and Azure's vCPUs may be slower. If preparation is cut off, the cache is never saved and every question stays slow. Making questions asynchronous (submit, then poll) would remove this limit and is the fix if a deployment hits it.
- Private network mode is not supported for the generator yet. In that mode the web app routes all outbound traffic through the VNet without a NAT gateway, and whether it can then reach the generator's public address at all is untested. Deploy the generator only in the default public network mode until it can be made reachable from inside the VNet.
Updating an existing installation¶
Custom reports arrive with a normal update; the generator does not install itself.
| What the operator does | |
|---|---|
| Docker | Re-download docker-compose.prod.yml (it gained the service, the internal network and the new env vars), then set COMPOSE_PROFILES=report-generator and pull. Without the new compose file the app updates as usual and the generator is simply absent. |
| Docker, auto-update | The auto-update agent updates the services in IA_SERVICES (web worker by default). Add report-generator there so the model image follows the channel too. |
| Azure | Re-run the deployment with deployReportGenerator=true. Existing resources are updated in place. |
| Azure, auto-update | Set IA_REPORT_GENERATOR_APP=<prefix>-report-generator for the Azure update agent, next to IA_WORKER_APP, so the model image follows the channel. Without it the generator keeps the previous release's model — it still works, but it is no longer the combination that was measured. |
| Desktop (portable) | Not supported — the portable launcher runs no extra containers. The definition editor works. |
In all cases the feature stays off until someone enables it in Admin → Experimental, and the first warm-up after the update rebuilds the prompt cache in the background (a few minutes, once). An install that does nothing keeps working exactly as before: a new, empty table is added, and nothing tries to reach a model server while the feature is off.
Grant the permission. The Build custom reports permission (data.write.reports) is in the
built-in RoleMiner role, but only a deployment still on the default role mapping gets it from there. If
you have customised your role mapping, add the permission to the roles that should build reports under
Admin → Roles.
Security notes¶
- The API surface answers 404 while the feature is off and 403 without the
data.write.reportspermission (checked first, so nobody learns which installs have the feature). - Model output is treated as untrusted input: it is validated against the catalog before anything runs, and never interpolated into SQL.
- Reports execute in a
READ ONLYtransaction with a statement timeout and a row cap. - The model server accepts no input except from the web container (Docker: internal network; Azure: API key) and cannot reach the database.
- Outbound internet access differs per platform. On Docker it has none (an
internalnetwork). On Azure it has unrestricted outbound access, like any Container App without a VNet and egress rules. llama.cpp makes no outbound calls and the model is inside the image, so nothing is sent — but on Azure that rests on the software, not on the network. - One question at a time per analyst, and a conversation is capped at 10,000 characters. The model
server works on one question at a time, so without these one person (or a script) could hold it for
everyone. Anyone with
data.write.reportscan still keep it busy; grant that permission accordingly. - The generated SQL is shown to the analyst in the builder, so table and column names are visible to anyone who can build reports. They are the same names documented for the data model.
- Questions are logged; report definitions are stored with their author.
- The model server's monitoring endpoint is switched off (
--no-slots). It would otherwise let any caller read the prompt currently being processed. - The container runs as an unprivileged user (uid 1000), not root.
If you deploy it on Azure¶
Two details are worth knowing, because getting them wrong is silent:
- The API key is the control. llama.cpp reads it from
LLAMA_API_KEYand only that name — theLLAMA_ARG_prefix that every other option uses is ignored for this one, and a server started without a key answers everyone. The template sets the right name from a secret and a guard test keeps it that way; do not hand-edit it. - Ingress is narrowed on the second run.
deploy.ps1 -DeployReportGeneratoradds the web app's outbound addresses to the ingress allow-list, but those addresses only exist once the web app does, so the first deployment is protected by the key alone. Re-run the script (or passreportGeneratorAllowedCallerIps) to add the allow-list. Those are shared Azure addresses, so treat it as defence in depth rather than a boundary. - In the private network mode the allow-list is not applied. There the web app sends all outbound traffic through the VNet, so its calls do not come from the addresses on the list and the list would lock it out. The generator is protected by its API key alone. It still has public ingress in that mode — making it reachable only from inside the VNet is not done yet.
Measuring it yourself¶
# the expected answers are still right for your data
node tools/nl-reports/eval.mjs --check
# accuracy + latency, per model
node tools/nl-reports/eval.mjs --models qwen3:4b-instruct-2507-q4_K_M
node tools/nl-reports/eval.mjs --models qwen3:4b-instruct-2507-q4_K_M --file tools/nl-reports/holdout.json
# against a stack with sign-in on: a bearer for a signed-in analyst, or a command that mints one
node tools/nl-reports/eval.mjs --check --base https://<host> --token-cmd "az account get-access-token --resource api://<web app id> --query accessToken -o tsv"
Questions live in tools/nl-reports/questions.json (used while tuning), holdout.json (kept
untouched, so the number means something) and chat.json (the questions people typed into the
in plain language, in Dutch and English, with follow-ups asked in the same conversation; "my" means
whoever runs it, so it needs a signed-in stack). Each question carries a hand-written reference
definition; a model's answer counts only if it returns the same rows.