Azure deployment — Simple shape¶
Companion to the Bicep in
/azure. One-click install: see the README's Deploy to Azure button. Target: customers whose networking is owned by a central CCoE. The default public network mode needs no VNet, private endpoints or public IP that you provision; an opt-in private mode adds them for new deployments (see Network modes).
What you get¶
Internet
│
▼ https://<app>.azurewebsites.net
│ (Microsoft-managed TLS cert)
┌──────────────────────────────────────┐
│ App Service Plan (Linux) │
│ ┌──────────────────────────────┐ │
│ │ App Service │ │
│ │ Pulls ghcr.io/fortigi/ │ │
│ │ identity-atlas:latest │ │
│ │ Mounts /data/uploads │ │
│ │ KV refs for master key + │ │
│ │ DB password │ │
│ └──────────────────────────────┘ │
└──────────────────────────────────────┘
│
┌──────────────────────────┼─────────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌────────────────────┐ ┌──────────────────┐
│ Postgres Flex │ │ Azure Files share │ │ Key Vault │
│ Public endpoint, │ │ (Storage Account) │ │ Public endpoint, │
│ firewall: web │ │ /data/uploads │ │ access policies. │
│ app outbound IPs │ │ shared with worker │ │ master key + │
│ only │ └────────────────────┘ │ DB password. │
└──────────────────┘ └──────────────────┘
─── Container Apps Environment (Consumption, no VNet) ───
│
▼
┌──────────────────────────────────────┐
│ Container App: worker (always-on) │
│ Pulls ghcr.io/fortigi/ │
│ identity-atlas-worker:latest │
│ Mounts the same /data/uploads share │
│ Polls App Service /api for queued │
│ crawler jobs every 60s │
└──────────────────────────────────────┘
Inventory: 8 resource families (App Service Plan + App Service, Postgres, Storage, Key Vault, Log Analytics, ACA Environment, ACA App, managed identities). No VNet, no private endpoints, no NSG.
Sizing — pick at deploy time¶
The Deploy-to-Azure form shows a sizeProfile dropdown. Pick what matches your tenant.
| Profile | App Service | Postgres | Worker | ~€/mo | Use case |
|---|---|---|---|---|---|
| xs | B1 (1 vCPU / 1.75 GB) | B1ms Burstable (1 vCPU / 2 GB) | 0.25 vCPU / 0.5 GB | 45 | Demo / proof-of-concept / non-production |
| s ✅ default | B2 (2 vCPU / 3.5 GB) | B2s Burstable (2 vCPU / 4 GB) | 0.25 vCPU / 0.5 GB | 79 | Small production — single team of analysts, < 10k principals |
| m | S1 (+ staging slot) | B2s Burstable | 0.25 vCPU / 0.5 GB | 113 | Mid-sized, 10-25k principals, blue/green deploys |
| l | P1v3 (2 vCPU / 8 GB) | D2ds_v5 GeneralPurpose (2 vCPU / 8 GB) | 0.5 vCPU / 1 GB | 244 | Large tenant, 25-50k principals, sustained throughput |
| xl | P2v3 (4 vCPU / 16 GB) | D4ds_v5 GeneralPurpose (4 vCPU / 16 GB) | 0.5 vCPU / 1 GB | 469 | Enterprise — 50k+ principals, multiple concurrent analyst teams |
(All West Europe, ex VAT, single replica, no HA, Linux. Costs assume BYO Log Analytics; add ~€3-5/mo if a new workspace is created.)
Scaling later = re-click the button (or rerun deploy.ps1) with a different sizeProfile. Azure resizes the App Service Plan + Postgres in place — no data loss, ~2-3 min of Postgres downtime while the SKU swap happens.
One-way ratchet: l and xl use Postgres GeneralPurpose. You can't scale a Postgres Flex server from GeneralPurpose back to Burstable — only within the same tier. Pick l/xl only when you actually need the GP compute.
"Fast and snappy" design choices¶
- Always On on the App Service Plan (B1+) — no cold start on first request.
- Worker as ACA App, always-on —
Sync nowpicks up jobs within 60s (matches the docker-compose worker behaviour). - App Service health probe on
/api/health— drops sick replicas before users hit them. - Postgres autogrow on, HA off — no failover blip during a sync.
- Materialised matrix view refreshed at end of every crawl — matrix renders from indexed pre-computed rows.
xs has caveats under concurrent load (3+ analysts on the matrix); it's labelled "non-production." Default s is where snappy starts being free of caveats.
BYO Log Analytics¶
If your CCoE owns a central Log Analytics workspace and you want logs forwarded there instead of creating a new one, fill in one of these two parameter pairs:
Option 1 — workspace ID (preferred, 1 field):
existingLogAnalyticsWorkspaceId: /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.OperationalInsights/workspaces/<name>
existing resource lookup to derive the workspace customer ID + shared key. The deployer needs Log Analytics Reader on the workspace.
Option 2 — customer ID + key (fallback when you can't read the workspace, 2 fields):
existingLogAnalyticsCustomerId: <GUID> (shown in the LA "Overview" blade as "Workspace ID")
existingLogAnalyticsSharedKey: <primary or secondary key>
Leave all three empty → template creates a fresh workspace in your resource group (~€3-5/mo).
The deployment's outputs include logAnalyticsCreated: true|false so you can tell which path was taken.
How it deploys (timing)¶
- Storage, Log Analytics (or lookup), Managed Identities, Key Vault — all parallel, < 1 min.
- Bootstrap deployment script — runs as the deployScript managed identity. Generates the master key + a random Postgres admin password into KV when they do not exist yet (existing values are kept). ~30 s.
- Postgres Flexible Server — slowest single step, ~3-4 min.
- App Service Plan + App Service — App Service starts the container image (first pull from ghcr.io, ~1-2 min).
- Container Apps Environment + Worker App — ~2 min.
Total: ~5-7 minutes wall clock. Output: the public URL of your Identity Atlas.
First-run post-deploy steps¶
- Open the App URL. First paint takes ~20-30s while the App Service container warms up (subsequent loads are sub-second because Always On is enabled).
- Admin → Crawlers → "Load demo data" to explore, or "Add Crawler" to connect Microsoft Graph.
- (Optional) Admin → Authentication to switch on Entra ID sign-in. Until then the app is open (so an IP allow-list at the App Service level is the recommended interim).
Operational notes¶
Startup & migrations¶
The web container binds its port before running DB migrations — migrations
run in the background right after the server starts listening. This matters on
App Service: the platform kills a container that doesn't answer on its port
within the cold-start probe window (WEBSITES_CONTAINER_START_TIME_LIMIT, 230s
by default), so a migration slower than that used to leave the port closed and
crash-loop the container into an "Application Error" page.
- The app comes up immediately and
/api/healthreturns200(with aschemaReadyfield) while a migration is still running. - Crawler job-claim and ingest endpoints return
503 Retry-After: 30until the schema is ready, so no crawler runs against a half-migrated database. The worker just retries. - If a migration fails, the container no longer crash-loops — it stays up, logs the error, and retries with backoff, self-healing once the cause clears.
- The template also sets
WEBSITES_CONTAINER_START_TIME_LIMIT=1800(the 30-min platform max) as an extra backstop for very large one-time schema upgrades.
Track :latest, not :edge, in production. :latest is the stable
customer release; :edge is the latest merged commit on main and may carry
unvetted migrations that reach your production the moment they merge. Pin the
channel by redeploying with imageChannel=stable (the default) — pass
imageChannel=edge instead to opt into the latest main build.
Updating the container images¶
# Restart the App Service — re-pulls the :latest tag from ghcr.io
az webapp restart --name <namePrefix>-web --resource-group <rg>
# Restart the worker — same
az containerapp revision restart --name <namePrefix>-worker --resource-group <rg> --revision <latestRevisionName>
To pin to an exact version tag rather than a channel, edit _imageTag in main.bicep directly and redeploy — the CLI/portal imageChannel parameter only offers the stable / edge channels above.
Logs¶
# App Service stdout, live
az webapp log tail --name <namePrefix>-web --resource-group <rg>
# Worker container logs, live
az containerapp logs show --name <namePrefix>-worker --resource-group <rg> --follow
Or query Log Analytics (whether created here or BYO):
AppServiceConsoleLogs | where _ResourceId endswith "<namePrefix>-web"
ContainerAppConsoleLogs_CL | where ContainerAppName_s == "<namePrefix>-worker"
Network modes¶
The networkMode parameter picks the network shape.
| public (default) | private (new deployments only) | |
|---|---|---|
| Postgres | Public endpoint. Firewall allows only the web app's possible outbound IP addresses (one rule each). | Public network access disabled; private endpoint. |
| Key Vault | Public endpoint, default action Allow. Access still requires the managed identity's access policy. | Private endpoint. Default action Deny after the deployment finishes (trusted Azure services bypass). |
| Storage (Azure Files) | Public endpoint, default action Allow (App Service and a VNet-less Container Apps environment can only mount it that way). | Private endpoint. Default action Deny after the deployment finishes. |
| Web app | Public. | Public, with regional VNet integration and all outbound traffic routed through the VNet. |
| Worker | Container Apps environment without a VNet. | Container Apps environment in a /23 subnet. |
| Extra cost | — | Private endpoints and DNS zones, roughly €25/mo. |
Why private is for new deployments only: the VNet of an existing Container Apps environment cannot be changed, so switching an existing deployment fails at the worker environment step. The Key Vault and Storage lock-down steps run last and only after everything else succeeded, so such a failed attempt does not close access that is still in use, but the deployment does not complete. Create a new resource group for a private deployment.
During a private-mode redeploy, Key Vault and Storage are opened (default action Allow) while the bootstrap script runs and closed again at the end.
Postgres firewall (public mode). Earlier templates allowed "all Azure services" (0.0.0.0), which admits resources from every Azure tenant. Redeploying with the current template narrows the existing rule, keeping its name AllowAllAzureServicesAndResourcesWithinAzureIps, to one of the web app's outbound addresses and adds web-app-outbound-N rules for the rest. The outbound list changes only when the App Service Plan moves to another pricing tier; a redeploy refreshes the rules. If you need the old behaviour, set postgresAllowAllAzureServices=true (not recommended).
Web app access restrictions. webAccessDefaultAction (default Allow) and webAllowedIpCidrs control who can reach the web app. With Deny, only the listed CIDRs can — plus, in private mode, the worker subnet. In public mode the worker calls the web app's public URL from the Container Apps environment's outbound address, so add that address to webAllowedIpCidrs before you choose Deny.
TLS to Postgres. The web app sets PGSSLMODE=require. With node-postgres this already verifies the server certificate chain and hostname (it is treated as ssl: true, unlike libpq's require), so no separate verify-full setting is needed.
Postgres admin password¶
The admin password is random, generated once by the bootstrap script and stored in Key Vault as postgres-admin-password; the template passes it to the server with getSecret() and the web app references the exact secret version. Nothing about it is derived from resource names.
Deployments created before this change used a password derived from the subscription ID and resource group name, which anyone who learns both can compute. Redeploying keeps that password (nothing rotates silently). Rotate it once, deliberately:
az deployment group create -g <rg> --template-file azure/main.bicep --parameters rotatePostgresPassword=true
The bootstrap script writes a new random password to Key Vault, the Postgres server is updated, and the web app restarts onto the new secret version in the same deployment (expect a short interruption). Redeploy later without the parameter (it defaults to false); passing it again rotates again. In the portal, set Rotate Postgres Password to true on the Deploy to Azure form.
Entra ID authentication for Postgres (a managed-identity token instead of a password) is not enabled yet. It needs an Entra administrator on the server, a database role for the web app's managed identity, and token acquisition in the API; the schema is currently owned by the password-based admin role. Tracked as a follow-up.
Postgres access¶
In public mode Postgres accepts only the web app's outbound addresses. To run psql from a workstation:
1. Add a temporary firewall rule for your IP: az postgres flexible-server firewall-rule create --resource-group <rg> --name <pgname> --rule-name temp-yourip --start-ip-address <ip> --end-ip-address <ip>.
2. Fetch the admin password from Key Vault: az keyvault secret show --vault-name <kv> --name postgres-admin-password --query value -o tsv.
3. psql "postgres://identityatlas:<pw>@<pgFqdn>:5432/identity_atlas?sslmode=require".
4. Delete the firewall rule when done.
Tearing down¶
Key Vault has soft delete + purge protection (7-day retention). To redeploy with the same KV name within those 7 days, either wait or deploy into a different resource group — the name prefix is derived from the resource group ID, so it can no longer be chosen directly.
Decisions taken (and why)¶
| Decision | Choice | Why |
|---|---|---|
| Architecture | App Service + Postgres Flex + ACA worker | Matches the customer's CCoE pattern — no networking provisioning. |
| Postgres endpoint | Public + firewall limited to the web app's outbound IPs (default), or private endpoint (networkMode=private) |
"Allow Azure services" admitted every Azure tenant. The App Service publishes every outbound address it can use, so the rules are exact without a VNet. |
| Key Vault endpoint | Public, access policies (default); private endpoint + Deny (networkMode=private) |
The bootstrap script and App Service Key Vault references need network access; without a VNet that means the public endpoint. Access is gated by access policies either way. |
| App Service image source | Direct pull from public ghcr.io |
No ACR needed (~€5/mo saved + simpler deploy). Add ACR later if a tenant demands a private registry. |
| Worker model | ACA App (always-on), not ACA Job | "Sync now" should respond in <2 s. Job would have a 60-300 s schedule lag. |
| Master key + DB password | Random, generated by deployment script, stored in KV, exposed to the app via KV references | Zero secrets in the template, the deployment history, or ARM, and nothing derivable from resource names. App reads them via managed identity at startup. |
| Auth | AUTH_ENABLED=false by default |
Avoids requiring an App Registration before first login. Configurable post-deploy via Admin → Authentication. |
| HA | Off | Single-replica everywhere. Keeps cost predictable. |
| Application Insights | Skipped | Log Analytics covers stdout + system metrics. AI is for distributed tracing, not needed at this scale. |
| VNet integration | Opt-in networkMode=private for new deployments |
Keeps the default shape free of networking prerequisites while offering private endpoints when required. The template creates its own VNet; bringing a CCoE-provided subnet is not supported yet. |
Limitations¶
- No HA / zone-redundancy. Single AZ, single replica. Adding zone redundancy is two Bicep params away — future iteration.
- No custom domain on App Service. Free
*.azurewebsites.netcert covers v1. Custom domain + customer TLS cert is a 2-step manual setup post-deploy. - No GitHub Actions deploy workflow. Use Deploy-to-Azure button +
deploy.ps1for now. - No Entra App Registration auto-creation. Done manually if/when enabling Entra auth.
- No platform-level App Service authentication (Easy Auth). Entra sign-in is enforced inside the application, not by the App Service
authSettingsV2config. Microsoft Defender for Cloud therefore flags the web app with "App Service apps should have authentication enabled" on every deployment — expected, and cleared with a Defender exemption. See azure-deployment-walkthrough.md. - Public Key Vault / Storage endpoints in the default mode. Use
networkMode=privateon a new deployment for private endpoints. An existing public deployment cannot be switched in place. - Postgres uses password authentication. Entra ID authentication with the web app's managed identity is a follow-up.
File index¶
| File | Purpose |
|---|---|
azure/main.bicep |
Top-level orchestrator |
azure/main.json |
Compiled ARM (used by the Deploy to Azure button) |
azure/main.parameters.example.json |
Example parameter file |
azure/deploy.ps1 |
CLI deploy + post-deploy summary |
azure/modules/log-analytics.bicep |
New workspace OR BYO lookup OR pass-through |
azure/modules/storage.bicep |
Storage Account + Azure Files share |
azure/modules/identities.bicep |
2 user-assigned managed identities |
azure/modules/key-vault.bicep |
Key Vault + access policies + network ACL |
azure/modules/network.bicep |
Private mode: VNet, subnets, private DNS zones |
azure/modules/private-endpoints.bicep |
Private mode: private endpoints for Key Vault, Storage, Postgres |
azure/modules/postgres-firewall.bicep |
Public mode: Postgres firewall rules for the web app's outbound IPs |
azure/modules/bootstrap.bicep |
One-shot deployment script for secrets |
azure/modules/postgres.bicep |
Postgres Flex |
azure/modules/app-service.bicep |
App Service Plan + App Service for Containers |
azure/modules/aca-env.bicep |
Container Apps Environment (Consumption) |
azure/modules/aca-app-worker.bicep |
Worker Container App (always-on) |