Ingest API¶
The Ingest API is how every record gets into Identity Atlas. Crawlers — in any language — authenticate with an API key and POST batches of data to REST endpoints; the API handles validation, bulk merge, scoped delete detection, audit history, and sync logging. The worker container has no direct database access; everything flows through this layer.
Why a separate ingest layer¶
Before v5, Identity Atlas (then shipped as the FortigiGraph PowerShell module) had two tightly-coupled sync paths that both ran inside the module with direct SQL access: Start-FGSync for the Entra ID path and Start-FGCSVSync for CSV imports. That design had real problems:
| Problem (pre-v5) | Impact |
|---|---|
| Crawlers needed SQL credentials | Security risk; credentials spread across environments |
| Couldn't write crawlers in Python, Go, or anything non-PowerShell | Locked into PowerShell for all integrations |
| CSV files had to be on the same machine as the module | No remote ingestion; no wizard uploads; limited deployment flexibility |
Adding a new source system meant adding a new Sync-FG* function in the module |
High coupling; slow to extend |
| No way for third parties to push data | Only pull-based; no webhook path |
v5 replaced both paths with this Ingest API. Crawlers are now standalone processes (in the worker container or anywhere else with network access to the web container) that speak HTTP, not SQL. Adding a new source is a matter of writing a small crawler that targets the endpoints below — no module changes.
Reference Architecture¶
graph TB
subgraph Sources["Source Systems"]
S1[EntraID]
S2[AD OnPrem]
S3[Omada]
S4[CSV]
end
subgraph Crawlers["Crawler Scripts<br/><i>PowerShell, Python, Go, etc.</i>"]
C1[EntraID Crawler]
C2[AD Crawler]
C3[Omada Crawler]
C4[CSV Crawler]
end
S1 --> C1
S2 --> C2
S3 --> C3
S4 --> C4
C1 -->|HTTPS + Bearer Token| API
C2 -->|HTTPS + Bearer Token| API
C3 -->|HTTPS + Bearer Token| API
C4 -->|HTTPS + Bearer Token| API
subgraph API["Ingest API"]
direction TB
IA["POST /api/ingest/*<br/>Validation · Bulk MERGE · Scoped Delete"]
end
subgraph Analytics["Analytics Engines"]
direction TB
CR[Linking Dictionary] --> CE[Account Linking Engine]
RP[Risk Profile] --> CL[Classifiers]
CL --> HE[Heuristics Engine]
CE <-->|shared data| HE
CE --> CS[Link Confidence]
HE --> RS[Risk Scores]
end
API --> DB[(Database<br/>PostgreSQL + Audit History)]
Analytics <--> DB
DB --> ReadAPI["Read API<br/>GET /api/*"]
ReadAPI --> UI["Web UI<br/>React + Vite"]
The architecture has four layers:
-
Source Systems & Crawlers — Each source system has a dedicated crawler. Crawlers are lightweight HTTP clients that fetch data from their source and POST it to the Ingest API. They can be written in any language.
-
Ingest API — Receives data via REST endpoints. Handles validation, bulk merge, scoped delete detection, and audit history recording. Authenticates crawlers via self-contained API keys.
-
Analytics Engines — Account Linking and Risk Scoring run independently against the database. The Account Linking engine is deterministic: it uses an editable dictionary (weighted signals + regex account-type rules + a confidence threshold) to attach orphan accounts to identities with a
linkConfidence— no LLM. The Heuristics Engine uses risk profiles and classifiers to compute risk scores. -
Read API + Web UI — The existing Express read routes and React frontend remain unchanged.
Ingest API Design¶
Core Principle: Batch-Oriented Sync¶
Each ingest endpoint accepts a batch of records for a given entity type and system. The API then:
- Validates all records against the schema
- Normalizes data (type coercion, deterministic GUID generation for non-GUID IDs)
- Bulk MERGEs into the target table (INSERT new, UPDATE changed)
- Scoped delete detection — if
syncMode: "full", records in this system+scope that are NOT in the batch are deleted - Logs the sync operation to
GraphSyncLog - Returns a summary:
{ table, inserted, updated, deleted, records, durationMs }(plussystemIdson the entity types that create/resolve them)
Endpoint Pattern¶
All ingest endpoints follow the same pattern:
POST /api/ingest/{entity-type}
Authorization: Bearer <crawler-api-key>
Content-Type: application/json
{
"systemId": 3,
"syncMode": "full",
"scope": {
"resourceType": "Group"
},
"records": [
{ "id": "...", "displayName": "...", ... }
]
}
Response — 201 Created:
{
"table": "Resources",
"inserted": 142,
"updated": 38,
"deleted": 7,
"records": 180,
"durationMs": 2340
}
records echoes the number of records in the batch. systemIds is included only on the entity types that create or resolve systems (e.g. POST /api/ingest/systems) — it holds the resolved integer system IDs so a push-mode connector can capture the systemId for its follow-up calls. There is no syncId and no errors[] field on a single-batch response: a batch either succeeds as a whole (201) or fails validation up front (400, with the offending records under details).
Entity Endpoints¶
| Endpoint | Target Table | Key Column(s) | Scope Filters |
|---|---|---|---|
POST /api/ingest/systems |
Systems | id (INT, auto) |
— |
POST /api/ingest/principals |
Principals | id (GUID) |
principalType |
POST /api/ingest/resources |
Resources | id (GUID) |
resourceType |
POST /api/ingest/resource-assignments |
ResourceAssignments | (resourceId, principalId, assignmentType, governed) |
assignmentType |
POST /api/ingest/resource-relationships |
ResourceRelationships | (parentResourceId, childResourceId, relationshipType) |
relationshipType |
POST /api/ingest/identities |
Identities | id (GUID) |
— |
POST /api/ingest/identity-members |
IdentityMembers | (identityId, principalId) |
— |
POST /api/ingest/contexts |
Contexts | id (GUID) |
contextType |
POST /api/ingest/governance/catalogs |
GovernanceCatalogs | id (GUID) |
— |
POST /api/ingest/governance/policies |
AssignmentPolicies | id (GUID) |
— |
POST /api/ingest/governance/requests |
AssignmentRequests | id (GUID) |
— |
POST /api/ingest/governance/certifications |
CertificationDecisions | id (GUID) |
— |
Sync Modes¶
| Mode | Behavior | Use Case |
|---|---|---|
full |
MERGE all records + DELETE records in scope not in batch | Scheduled full sync |
delta |
MERGE only; no deletes | Real-time webhook, incremental changes |
POST /api/ingest/systems is never reconciled: systems are registered, not synced, so a full batch there is run as a delta. A reconcile delete that would have no system, scope or ownership bound at all is refused and reports deleted: 0.
Deterministic GUID Generation¶
For source systems that don't use GUIDs (e.g., Omada uses integer IDs):
{
"systemId": 3,
"idGeneration": "deterministic",
"idPrefix": "omada-resource",
"records": [
{ "externalId": "12345", "displayName": "Admin Role" }
]
}
When idGeneration: "deterministic", the API generates MD5(idPrefix + ":" + externalId) as UUID v3, matching the current CSV sync pattern.
Sync Sessions (Chunked Uploads)¶
For datasets larger than 50,000 records:
sequenceDiagram
participant C as Crawler
participant A as Ingest API
participant DB as Database
C->>A: POST /ingest/resources (syncSession: "start", records[0..10000])
A->>DB: Create temp table, MERGE batch 1
A-->>C: { syncId: "abc-123" }
C->>A: POST /ingest/resources (syncSession: "continue", syncId: "abc-123", records[10001..20000])
A->>DB: MERGE batch 2 into same temp table
A-->>C: { syncId: "abc-123" }
C->>A: POST /ingest/resources (syncSession: "end", syncId: "abc-123", records[20001..25000])
A->>DB: MERGE batch 3, run scoped delete, drop temp table
A-->>C: { syncId: "abc-123", inserted: 142, updated: 38, deleted: 7 }
A session holds one connection and one transaction for its whole life (at most 30
minutes), so it does not stretch to tens of millions of rows. Those scopes are sent
as independent delta batches followed by POST /ingest/reconcile, which removes
rows not touched since the run began — or, for a full re-read, as a staged load.
Staged full load¶
For a scope whose complete row set is sent every run — tens of millions of assignments, typically — the stage endpoints apply that set in one step instead of batch by batch:
| Call | Does |
|---|---|
POST /ingest/stages { entity, systemId, scope?, idGeneration?, idPrefix? } |
opens a stage → { stageId } |
POST /ingest/stages/{id}/rows { records } |
appends a batch; records get the same defaults, validation, id normalization and per-system boundary as /ingest/{entity} |
POST /ingest/stages/finalize { stageIds, deleteMissing?, maxDeleteShare? } |
applies a run's stages together |
POST /ingest/stages/{id}/finalize { deleteMissing?, maxDeleteShare? } |
applies one stage |
DELETE /ingest/stages/{id} |
abandons a stage — nothing was written |
Entities: resource-assignments, principals, resources,
resource-relationships. Nothing reaches the target table before finalize, which
takes one of two paths and reports it (path in the result, and the sync log).
Every result also carries rows (records received) and distinct (distinct keys
among them): after a finalize with deleteMissing, distinct is exactly what the
scope holds, which is what a caller should verify against. A finalize without
deleteMissing applies a window rather than a complete set, so the scope's total
says nothing about it; its result carries present instead, the number of those
distinct keys that are live in the table afterwards. The paths:
empty-table— the target table holds no rows at all (a first load on a new installation). Its non-constraint indexes are dropped, every stage is inserted bare, and the indexes are rebuilt once. Emptiness is checked under an exclusive lock taken in the same transaction; if the lock is not granted within 5 s, or the table is not empty, finalize merges instead.merge— each stage, in its own transaction: insert keys the table lacks, update only rows whose values changed, and withdeleteMissing: trueremove the system+scope's rows that are not in the stage (tombstoned on soft-delete tables). A tombstoned row that is back is revived.
A stage may carry only its key columns — a key sweep: finalize then inserts and
updates nothing and only removes what is missing. (systemId is stamped on every
staged row from the stage itself, so it does not count as content; a stage whose
records carried nothing but keys is still a sweep.)
maxDeleteShare caps that removal. With it, finalize measures what the
anti-join would remove against the scope's live rows — both inside the one
transaction — and above the share answers 409 and writes nothing at all
(the delete is rolled back, so the count is exact rather than an estimate). It
exists because a source read while it is being re-aggregated is
indistinguishable from a mass revocation, and a delete has no undo. Leave it out
for no ceiling, which is what every caller before the SQL connector's sweep does.
Staged load or timestamp reconcile? It is a size trade-off, not a verdict on
either. The timestamp reconcile needs every row the source still has to be touched
(its updatedAt rewritten) so the untouched ones can be found — cheap for a small
scope, and it needs no stage. At tens of millions of rows the touch is the cost:
updatedAt is indexed, so every touch rewrites the row and all of its index entries.
Measured on 4.1M unchanged assignments: 349 s with batches and reconcile, 99 s staged,
with nothing written (Scale Rehearsal).
How a crawler uses it: open one stage per (entity, system, scope) it reads in
full; stream the rows in batches of any size; finalize all of the run's stages in one
POST /ingest/stages/finalize, with deleteMissing: true. Finalizing together is
what lets a first load take the empty-table path for the whole run rather than
for its first system only. Stages live in memory: an API restart abandons them (their
tables are dropped at startup) and the crawler starts the run over. A stage expires
six hours after it was opened.
Crawler Authentication¶
Self-Contained API Keys¶
No external IdP dependency. The API manages its own crawler credentials.
erDiagram
Crawlers {
int id PK
string displayName
binary apiKeyHash
binary apiKeySalt
string apiKeyPrefix
string systemIds "JSON array"
string permissions "JSON array"
bool enabled
datetime expiresAt
int rateLimit
}
CrawlerAuditLog {
int id PK
int crawlerId FK
string action
string endpoint
int recordCount
int statusCode
string ipAddress
datetime timestamp
}
Crawlers ||--o{ CrawlerAuditLog : "tracked by"
Key format: fgc_<random-32-chars> — the fgc_ prefix makes keys recognisable in logs and auth middleware as crawler tokens (distinct from JWTs used for the read API). Only the hash is stored; the plaintext key is shown once at creation time.
Admin Endpoints (Entra ID Auth)¶
| Method | Endpoint | Purpose |
|---|---|---|
GET |
/api/admin/crawlers |
List all crawlers (without keys) |
POST |
/api/admin/crawlers |
Register new crawler, returns plaintext key once |
PATCH |
/api/admin/crawlers/:id |
Update name, description, enabled, systemIds, permissions |
DELETE |
/api/admin/crawlers/:id |
Disable (soft-delete) crawler |
GET |
/api/admin/crawlers/:id/audit |
View audit log |
POST |
/api/admin/crawlers/:id/reset |
Admin-initiated key reset |
Crawler Self-Service Endpoints (API Key Auth)¶
| Method | Endpoint | Purpose |
|---|---|---|
POST |
/api/crawlers/rotate |
Rotate own key (old key invalidated immediately) |
GET |
/api/crawlers/whoami |
Return crawler metadata |
Key Rotation Flow¶
# Example: Python crawler auto-rotation
new_key = requests.post("/api/crawlers/rotate",
headers={"Authorization": f"Bearer {current_key}"}).json()["apiKey"]
save_to_vault(new_key)
Auth Middleware Chain¶
graph LR
R[Request] --> D{Path?}
D -->|/api/ingest/*| CK[crawlerAuth<br/>API Key]
D -->|/api/crawlers/*| CK
D -->|/api/admin/crawlers/*| EA[Entra ID Auth]
D -->|/api/* other| EA
Ingest Engine¶
The server-side engine encapsulates all SQL complexity:
UI/backend/src/
├── ingest/
│ ├── engine.js — Core MERGE + delete detection
│ ├── validation.js — JSON Schema validation per entity type
│ ├── normalization.js — Type coercion, GUID generation
│ ├── schemas/ — JSON Schema per entity type
│ └── sessions.js — Sync session management
├── routes/
│ ├── ingest.js — Ingest endpoints
│ └── crawlers.js — Crawler management
├── middleware/
│ └── crawlerAuth.js — API key validation
Engine Operations¶
flowchart TD
A[Receive batch] --> B[Validate against JSON Schema]
B --> C[Normalize: type coercion, GUID generation]
C --> D[Create temp table]
D --> E[BulkLoad into temp table]
E --> F[MERGE into target table]
F --> G{syncMode?}
G -->|full| H[Scoped DELETE:<br/>systemId + scope + NOT IN temp]
G -->|delta| I[Skip delete]
H --> J[Write sync log]
I --> J
J --> K[Return summary]
Scoped Delete Detection¶
The engine preserves the same scoping patterns used by the current PowerShell sync:
- System-scoped:
WHERE systemId = @systemId - Attribute-scoped:
WHERE resourceType = @scope(if provided) - Current-state scoped: operates on the current table rows (no temporal filtering needed in v5)
- Batch-scoped:
AND NOT EXISTS (SELECT 1 FROM #temp WHERE ...)
Contexts and context members: owned, not system-scoped¶
Contexts and ContextMembers have no systemId column, and every crawler writes them.
A full sync of either is bounded by ownership, derived from the envelope systemId,
whatever scope the caller sends:
- Contexts: only
variant = 'synced'contexts whosescopeSystemIdis the sending system. Manual and generated contexts are never removed by a sync, and nor is another system's context. A crawler must therefore stampscopeSystemIdon the contexts it sends. A synced context without it is never reconciled, and the response carries awarningsentry that the shared crawler ingest prints in the job log. - Context members: only memberships of contexts the sender owns, and never one whose
addedByis'analyst'. A sync removes only memberships it added. An analyst's deliberate addition survives, even on a context the crawler owns, the same way the context algorithm runner replaces onlyalgorithmrows (Context redesign). If the context itself disappears from the source, its memberships go with it.
addedBy is therefore load-bearing, not decorative. Before this bound existed, a full sync
through an unrestricted key reconciled both tables whole. One SQL Server load removed another
source's 150 logical applications, and a routine refresh removed an analyst's tag
membership, while every count check passed. contextSyncScope.contract.test.js pins the
bound against real PostgreSQL.
Keys restricted to specific systems¶
A crawler key with systemIds set can only read and write data of those systems. On top of the envelope systemId check, every batch from such a key is refused (403) when:
- a record carries a
systemId(or a context'sscopeSystemId) outside the key's systems; - the batch would update an existing row owned by another system — ownership is the row's
systemId; for tables without one it is derived: a synced context'sscopeSystemId, a context member's context, an identity member's or activity row's principal, and an identity's linked principals; - it is a
fullsync ofidentities,identity-members,contexts,context-membersorprincipal-activitywithout ascope.
deletedIds and the full-sync reconcile only reach rows the key's systems own. Unrestricted keys (systemIds null, such as the built-in worker) are not subject to these checks. Columns the server or analysts own (deletedAt, riskScore, riskTier, linkConfidence, analystOverride, createdAt, updatedAt and similar) are never written from a crawler record; a field with such a name is kept in extendedAttributes like any other attribute.
A key can hold at most 3 multi-batch sessions open at once, and the API at most 7 in total (INGEST_MAX_SESSIONS_PER_CRAWLER, INGEST_MAX_SESSIONS_GLOBAL); a start beyond that returns 429. The built-in worker is only subject to the global limit. A record's extendedAttributes may hold at most 500 keys and 512 KB.
Validation Rules¶
| Field | Rule |
|---|---|
id (GUID) |
Valid UUID v4 format, or externalId + idGeneration: "deterministic" |
systemId |
Must exist in Systems table AND be in crawler's allowed systems |
displayName |
Required, max 255 chars |
principalType |
One of: User, ServicePrincipal, ManagedIdentity, WorkloadIdentity, AIAgent, ExternalUser, SharedMailbox |
resourceType |
Free-form text — no enum check. Values currently emitted by crawlers: Group, BusinessRole, Application, AppRole, DelegatedPermission, EntraDirectoryRole, GroupOwnership, ApplicationOwnership, ServicePrincipalOwnership |
assignmentType |
One of exactly: Direct, Indirect, Eligible — ingest rejects any other value. Everything that used to be its own type is modelled differently: ownership is a Direct membership on a GroupOwnership/*Ownership resource; governance is the governed=true flag on the assignment; the former source-detail types (OAuth2Grant, AppRole, AppRoleViaGroup, DirectoryRole, DirectoryRoleEligible) collapse to Direct/Indirect/Eligible with resourceType carrying the detail |
origin |
Optional, on assignments. One of exactly: Automatic (a rule or birthright policy assigned it), Requested (it was requested and approved), Discovered (the source found it on the target system) — ingest rejects any other value. Leave it out when the source does not say; that is not the same as Discovered. Not part of the row key or the reconcile scope, so a grant whose origin changes is the same row |
originDetail |
Optional, on assignments, max 100 chars. The source's own word for the origin (IdentityIQ: Rule, LCM, Aggregation, …), kept so the grouping into origin can be audited |
relationshipType |
One of: Contains, HasAppRole, DelegatesScope, HasApplicationPermission, HasAppOwnership, HasOwnership (GrantsAccessTo is accepted but reserved / not yet emitted) |
extendedAttributes |
Valid JSON object, max 64 KB |
OpenAPI / Swagger¶
The API serves an OpenAPI 3.0 spec and Swagger UI:
GET /api/docs— Swagger UI (interactive documentation)GET /api/docs/openapi.json— OpenAPI 3.0 spec file
From this spec, crawlers can auto-generate clients:
# Generate PowerShell client
npx @openapitools/openapi-generator-cli generate \
-i openapi.json -g powershell -o ./crawler-client-ps
# Generate Python client
npx @openapitools/openapi-generator-cli generate \
-i openapi.json -g python -o ./crawler-client-py
Observed performance¶
Measured against the committed load-test dataset (~2.17 M records, ~97 MB of CSV) on a VM with 6 cores / 16 GB RAM:
| Phase | Records | Duration | Throughput |
|---|---|---|---|
| Identities | 25,000 | 9 s | ~2,800 rows/s |
| Identity members | 76,000 | 40 s | ~1,900 rows/s |
| Certification decisions | 300,000 | 4 min 12 s (1 batch) | ~1,190 rows/s |
| Resource assignments | 1,500,000 | ~20 min (20 batches × 75 k) | ~1,250 rows/s sustained |
| Full run | ~2.17 M | ~30 min | ~1,200 rows/s overall |
See Scaling & Load Testing for the full analysis, including hardware utilisation (CPU 73 %, memory 87 %, disk 1 % — memory is the limiting factor) and reproduction instructions.
Future Extensions¶
- Webhook receiver — source systems push change events
- NDJSON streaming — for very large datasets
- Crawler SDK — npm/PyPI/PSGallery package with auth, chunking, retry
- Crawler templates — wizard in admin UI generates boilerplate
- Async ingestion — queue-based with job IDs
- Data quality scoring — completeness and consistency metrics per sync