Skip to content

Ingest API

The Ingest API is how every record gets into Identity Atlas. Crawlers — in any language — authenticate with an API key and POST batches of data to REST endpoints; the API handles validation, bulk merge, scoped delete detection, audit history, and sync logging. The worker container has no direct database access; everything flows through this layer.


Why a separate ingest layer

Before v5, Identity Atlas (then shipped as the FortigiGraph PowerShell module) had two tightly-coupled sync paths that both ran inside the module with direct SQL access: Start-FGSync for the Entra ID path and Start-FGCSVSync for CSV imports. That design had real problems:

Problem (pre-v5) Impact
Crawlers needed SQL credentials Security risk; credentials spread across environments
Couldn't write crawlers in Python, Go, or anything non-PowerShell Locked into PowerShell for all integrations
CSV files had to be on the same machine as the module No remote ingestion; no wizard uploads; limited deployment flexibility
Adding a new source system meant adding a new Sync-FG* function in the module High coupling; slow to extend
No way for third parties to push data Only pull-based; no webhook path

v5 replaced both paths with this Ingest API. Crawlers are now standalone processes (in the worker container or anywhere else with network access to the web container) that speak HTTP, not SQL. Adding a new source is a matter of writing a small crawler that targets the endpoints below — no module changes.


Reference Architecture

graph TB
    subgraph Sources["Source Systems"]
        S1[EntraID]
        S2[AD OnPrem]
        S3[Omada]
        S4[CSV]
    end

    subgraph Crawlers["Crawler Scripts<br/><i>PowerShell, Python, Go, etc.</i>"]
        C1[EntraID Crawler]
        C2[AD Crawler]
        C3[Omada Crawler]
        C4[CSV Crawler]
    end

    S1 --> C1
    S2 --> C2
    S3 --> C3
    S4 --> C4

    C1 -->|HTTPS + Bearer Token| API
    C2 -->|HTTPS + Bearer Token| API
    C3 -->|HTTPS + Bearer Token| API
    C4 -->|HTTPS + Bearer Token| API

    subgraph API["Ingest API"]
        direction TB
        IA["POST /api/ingest/*<br/>Validation · Bulk MERGE · Scoped Delete"]
    end

    subgraph Analytics["Analytics Engines"]
        direction TB
        CR[Linking Dictionary] --> CE[Account Linking Engine]
        RP[Risk Profile] --> CL[Classifiers]
        CL --> HE[Heuristics Engine]
        CE <-->|shared data| HE
        CE --> CS[Link Confidence]
        HE --> RS[Risk Scores]
    end

    API --> DB[(Database<br/>PostgreSQL + Audit History)]
    Analytics <--> DB

    DB --> ReadAPI["Read API<br/>GET /api/*"]
    ReadAPI --> UI["Web UI<br/>React + Vite"]

The architecture has four layers:

  1. Source Systems & Crawlers — Each source system has a dedicated crawler. Crawlers are lightweight HTTP clients that fetch data from their source and POST it to the Ingest API. They can be written in any language.

  2. Ingest API — Receives data via REST endpoints. Handles validation, bulk merge, scoped delete detection, and audit history recording. Authenticates crawlers via self-contained API keys.

  3. Analytics Engines — Account Linking and Risk Scoring run independently against the database. The Account Linking engine is deterministic: it uses an editable dictionary (weighted signals + regex account-type rules + a confidence threshold) to attach orphan accounts to identities with a linkConfidence — no LLM. The Heuristics Engine uses risk profiles and classifiers to compute risk scores.

  4. Read API + Web UI — The existing Express read routes and React frontend remain unchanged.


Ingest API Design

Core Principle: Batch-Oriented Sync

Each ingest endpoint accepts a batch of records for a given entity type and system. The API then:

  1. Validates all records against the schema
  2. Normalizes data (type coercion, deterministic GUID generation for non-GUID IDs)
  3. Bulk MERGEs into the target table (INSERT new, UPDATE changed)
  4. Scoped delete detection — if syncMode: "full", records in this system+scope that are NOT in the batch are deleted
  5. Logs the sync operation to GraphSyncLog
  6. Returns a summary: { table, inserted, updated, deleted, records, durationMs } (plus systemIds on the entity types that create/resolve them)

Endpoint Pattern

All ingest endpoints follow the same pattern:

POST /api/ingest/{entity-type}
Authorization: Bearer <crawler-api-key>
Content-Type: application/json

{
  "systemId": 3,
  "syncMode": "full",
  "scope": {
    "resourceType": "Group"
  },
  "records": [
    { "id": "...", "displayName": "...", ... }
  ]
}

Response — 201 Created:

{
  "table": "Resources",
  "inserted": 142,
  "updated": 38,
  "deleted": 7,
  "records": 180,
  "durationMs": 2340
}

records echoes the number of records in the batch. systemIds is included only on the entity types that create or resolve systems (e.g. POST /api/ingest/systems) — it holds the resolved integer system IDs so a push-mode connector can capture the systemId for its follow-up calls. There is no syncId and no errors[] field on a single-batch response: a batch either succeeds as a whole (201) or fails validation up front (400, with the offending records under details).

Entity Endpoints

Endpoint Target Table Key Column(s) Scope Filters
POST /api/ingest/systems Systems id (INT, auto) —
POST /api/ingest/principals Principals id (GUID) principalType
POST /api/ingest/resources Resources id (GUID) resourceType
POST /api/ingest/resource-assignments ResourceAssignments (resourceId, principalId, assignmentType, governed) assignmentType
POST /api/ingest/resource-relationships ResourceRelationships (parentResourceId, childResourceId, relationshipType) relationshipType
POST /api/ingest/identities Identities id (GUID) —
POST /api/ingest/identity-members IdentityMembers (identityId, principalId) —
POST /api/ingest/contexts Contexts id (GUID) contextType
POST /api/ingest/governance/catalogs GovernanceCatalogs id (GUID) —
POST /api/ingest/governance/policies AssignmentPolicies id (GUID) —
POST /api/ingest/governance/requests AssignmentRequests id (GUID) —
POST /api/ingest/governance/certifications CertificationDecisions id (GUID) —

Sync Modes

Mode Behavior Use Case
full MERGE all records + DELETE records in scope not in batch Scheduled full sync
delta MERGE only; no deletes Real-time webhook, incremental changes

POST /api/ingest/systems is never reconciled: systems are registered, not synced, so a full batch there is run as a delta. A reconcile delete that would have no system, scope or ownership bound at all is refused and reports deleted: 0.

Deterministic GUID Generation

For source systems that don't use GUIDs (e.g., Omada uses integer IDs):

{
  "systemId": 3,
  "idGeneration": "deterministic",
  "idPrefix": "omada-resource",
  "records": [
    { "externalId": "12345", "displayName": "Admin Role" }
  ]
}

When idGeneration: "deterministic", the API generates MD5(idPrefix + ":" + externalId) as UUID v3, matching the current CSV sync pattern.

Sync Sessions (Chunked Uploads)

For datasets larger than 50,000 records:

sequenceDiagram
    participant C as Crawler
    participant A as Ingest API
    participant DB as Database

    C->>A: POST /ingest/resources (syncSession: "start", records[0..10000])
    A->>DB: Create temp table, MERGE batch 1
    A-->>C: { syncId: "abc-123" }

    C->>A: POST /ingest/resources (syncSession: "continue", syncId: "abc-123", records[10001..20000])
    A->>DB: MERGE batch 2 into same temp table
    A-->>C: { syncId: "abc-123" }

    C->>A: POST /ingest/resources (syncSession: "end", syncId: "abc-123", records[20001..25000])
    A->>DB: MERGE batch 3, run scoped delete, drop temp table
    A-->>C: { syncId: "abc-123", inserted: 142, updated: 38, deleted: 7 }

A session holds one connection and one transaction for its whole life (at most 30 minutes), so it does not stretch to tens of millions of rows. Those scopes are sent as independent delta batches followed by POST /ingest/reconcile, which removes rows not touched since the run began — or, for a full re-read, as a staged load.

Staged full load

For a scope whose complete row set is sent every run — tens of millions of assignments, typically — the stage endpoints apply that set in one step instead of batch by batch:

Call Does
POST /ingest/stages { entity, systemId, scope?, idGeneration?, idPrefix? } opens a stage → { stageId }
POST /ingest/stages/{id}/rows { records } appends a batch; records get the same defaults, validation, id normalization and per-system boundary as /ingest/{entity}
POST /ingest/stages/finalize { stageIds, deleteMissing?, maxDeleteShare? } applies a run's stages together
POST /ingest/stages/{id}/finalize { deleteMissing?, maxDeleteShare? } applies one stage
DELETE /ingest/stages/{id} abandons a stage — nothing was written

Entities: resource-assignments, principals, resources, resource-relationships. Nothing reaches the target table before finalize, which takes one of two paths and reports it (path in the result, and the sync log). Every result also carries rows (records received) and distinct (distinct keys among them): after a finalize with deleteMissing, distinct is exactly what the scope holds, which is what a caller should verify against. A finalize without deleteMissing applies a window rather than a complete set, so the scope's total says nothing about it; its result carries present instead, the number of those distinct keys that are live in the table afterwards. The paths:

  • empty-table — the target table holds no rows at all (a first load on a new installation). Its non-constraint indexes are dropped, every stage is inserted bare, and the indexes are rebuilt once. Emptiness is checked under an exclusive lock taken in the same transaction; if the lock is not granted within 5 s, or the table is not empty, finalize merges instead.
  • merge — each stage, in its own transaction: insert keys the table lacks, update only rows whose values changed, and with deleteMissing: true remove the system+scope's rows that are not in the stage (tombstoned on soft-delete tables). A tombstoned row that is back is revived.

A stage may carry only its key columns — a key sweep: finalize then inserts and updates nothing and only removes what is missing. (systemId is stamped on every staged row from the stage itself, so it does not count as content; a stage whose records carried nothing but keys is still a sweep.)

maxDeleteShare caps that removal. With it, finalize measures what the anti-join would remove against the scope's live rows — both inside the one transaction — and above the share answers 409 and writes nothing at all (the delete is rolled back, so the count is exact rather than an estimate). It exists because a source read while it is being re-aggregated is indistinguishable from a mass revocation, and a delete has no undo. Leave it out for no ceiling, which is what every caller before the SQL connector's sweep does.

Staged load or timestamp reconcile? It is a size trade-off, not a verdict on either. The timestamp reconcile needs every row the source still has to be touched (its updatedAt rewritten) so the untouched ones can be found — cheap for a small scope, and it needs no stage. At tens of millions of rows the touch is the cost: updatedAt is indexed, so every touch rewrites the row and all of its index entries. Measured on 4.1M unchanged assignments: 349 s with batches and reconcile, 99 s staged, with nothing written (Scale Rehearsal).

How a crawler uses it: open one stage per (entity, system, scope) it reads in full; stream the rows in batches of any size; finalize all of the run's stages in one POST /ingest/stages/finalize, with deleteMissing: true. Finalizing together is what lets a first load take the empty-table path for the whole run rather than for its first system only. Stages live in memory: an API restart abandons them (their tables are dropped at startup) and the crawler starts the run over. A stage expires six hours after it was opened.


Crawler Authentication

Self-Contained API Keys

No external IdP dependency. The API manages its own crawler credentials.

erDiagram
    Crawlers {
        int id PK
        string displayName
        binary apiKeyHash
        binary apiKeySalt
        string apiKeyPrefix
        string systemIds "JSON array"
        string permissions "JSON array"
        bool enabled
        datetime expiresAt
        int rateLimit
    }
    CrawlerAuditLog {
        int id PK
        int crawlerId FK
        string action
        string endpoint
        int recordCount
        int statusCode
        string ipAddress
        datetime timestamp
    }
    Crawlers ||--o{ CrawlerAuditLog : "tracked by"

Key format: fgc_<random-32-chars> — the fgc_ prefix makes keys recognisable in logs and auth middleware as crawler tokens (distinct from JWTs used for the read API). Only the hash is stored; the plaintext key is shown once at creation time.

Admin Endpoints (Entra ID Auth)

Method Endpoint Purpose
GET /api/admin/crawlers List all crawlers (without keys)
POST /api/admin/crawlers Register new crawler, returns plaintext key once
PATCH /api/admin/crawlers/:id Update name, description, enabled, systemIds, permissions
DELETE /api/admin/crawlers/:id Disable (soft-delete) crawler
GET /api/admin/crawlers/:id/audit View audit log
POST /api/admin/crawlers/:id/reset Admin-initiated key reset

Crawler Self-Service Endpoints (API Key Auth)

Method Endpoint Purpose
POST /api/crawlers/rotate Rotate own key (old key invalidated immediately)
GET /api/crawlers/whoami Return crawler metadata

Key Rotation Flow

# Example: Python crawler auto-rotation
new_key = requests.post("/api/crawlers/rotate",
    headers={"Authorization": f"Bearer {current_key}"}).json()["apiKey"]
save_to_vault(new_key)

Auth Middleware Chain

graph LR
    R[Request] --> D{Path?}
    D -->|/api/ingest/*| CK[crawlerAuth<br/>API Key]
    D -->|/api/crawlers/*| CK
    D -->|/api/admin/crawlers/*| EA[Entra ID Auth]
    D -->|/api/* other| EA

Ingest Engine

The server-side engine encapsulates all SQL complexity:

UI/backend/src/
├── ingest/
│   ├── engine.js              — Core MERGE + delete detection
│   ├── validation.js          — JSON Schema validation per entity type
│   ├── normalization.js       — Type coercion, GUID generation
│   ├── schemas/               — JSON Schema per entity type
│   └── sessions.js            — Sync session management
├── routes/
│   ├── ingest.js              — Ingest endpoints
│   └── crawlers.js            — Crawler management
├── middleware/
│   └── crawlerAuth.js         — API key validation

Engine Operations

flowchart TD
    A[Receive batch] --> B[Validate against JSON Schema]
    B --> C[Normalize: type coercion, GUID generation]
    C --> D[Create temp table]
    D --> E[BulkLoad into temp table]
    E --> F[MERGE into target table]
    F --> G{syncMode?}
    G -->|full| H[Scoped DELETE:<br/>systemId + scope + NOT IN temp]
    G -->|delta| I[Skip delete]
    H --> J[Write sync log]
    I --> J
    J --> K[Return summary]

Scoped Delete Detection

The engine preserves the same scoping patterns used by the current PowerShell sync:

  • System-scoped: WHERE systemId = @systemId
  • Attribute-scoped: WHERE resourceType = @scope (if provided)
  • Current-state scoped: operates on the current table rows (no temporal filtering needed in v5)
  • Batch-scoped: AND NOT EXISTS (SELECT 1 FROM #temp WHERE ...)

Contexts and context members: owned, not system-scoped

Contexts and ContextMembers have no systemId column, and every crawler writes them. A full sync of either is bounded by ownership, derived from the envelope systemId, whatever scope the caller sends:

  • Contexts: only variant = 'synced' contexts whose scopeSystemId is the sending system. Manual and generated contexts are never removed by a sync, and nor is another system's context. A crawler must therefore stamp scopeSystemId on the contexts it sends. A synced context without it is never reconciled, and the response carries a warnings entry that the shared crawler ingest prints in the job log.
  • Context members: only memberships of contexts the sender owns, and never one whose addedBy is 'analyst'. A sync removes only memberships it added. An analyst's deliberate addition survives, even on a context the crawler owns, the same way the context algorithm runner replaces only algorithm rows (Context redesign). If the context itself disappears from the source, its memberships go with it.

addedBy is therefore load-bearing, not decorative. Before this bound existed, a full sync through an unrestricted key reconciled both tables whole. One SQL Server load removed another source's 150 logical applications, and a routine refresh removed an analyst's tag membership, while every count check passed. contextSyncScope.contract.test.js pins the bound against real PostgreSQL.

Keys restricted to specific systems

A crawler key with systemIds set can only read and write data of those systems. On top of the envelope systemId check, every batch from such a key is refused (403) when:

  • a record carries a systemId (or a context's scopeSystemId) outside the key's systems;
  • the batch would update an existing row owned by another system — ownership is the row's systemId; for tables without one it is derived: a synced context's scopeSystemId, a context member's context, an identity member's or activity row's principal, and an identity's linked principals;
  • it is a full sync of identities, identity-members, contexts, context-members or principal-activity without a scope.

deletedIds and the full-sync reconcile only reach rows the key's systems own. Unrestricted keys (systemIds null, such as the built-in worker) are not subject to these checks. Columns the server or analysts own (deletedAt, riskScore, riskTier, linkConfidence, analystOverride, createdAt, updatedAt and similar) are never written from a crawler record; a field with such a name is kept in extendedAttributes like any other attribute.

A key can hold at most 3 multi-batch sessions open at once, and the API at most 7 in total (INGEST_MAX_SESSIONS_PER_CRAWLER, INGEST_MAX_SESSIONS_GLOBAL); a start beyond that returns 429. The built-in worker is only subject to the global limit. A record's extendedAttributes may hold at most 500 keys and 512 KB.

Validation Rules

Field Rule
id (GUID) Valid UUID v4 format, or externalId + idGeneration: "deterministic"
systemId Must exist in Systems table AND be in crawler's allowed systems
displayName Required, max 255 chars
principalType One of: User, ServicePrincipal, ManagedIdentity, WorkloadIdentity, AIAgent, ExternalUser, SharedMailbox
resourceType Free-form text — no enum check. Values currently emitted by crawlers: Group, BusinessRole, Application, AppRole, DelegatedPermission, EntraDirectoryRole, GroupOwnership, ApplicationOwnership, ServicePrincipalOwnership
assignmentType One of exactly: Direct, Indirect, Eligible — ingest rejects any other value. Everything that used to be its own type is modelled differently: ownership is a Direct membership on a GroupOwnership/*Ownership resource; governance is the governed=true flag on the assignment; the former source-detail types (OAuth2Grant, AppRole, AppRoleViaGroup, DirectoryRole, DirectoryRoleEligible) collapse to Direct/Indirect/Eligible with resourceType carrying the detail
origin Optional, on assignments. One of exactly: Automatic (a rule or birthright policy assigned it), Requested (it was requested and approved), Discovered (the source found it on the target system) — ingest rejects any other value. Leave it out when the source does not say; that is not the same as Discovered. Not part of the row key or the reconcile scope, so a grant whose origin changes is the same row
originDetail Optional, on assignments, max 100 chars. The source's own word for the origin (IdentityIQ: Rule, LCM, Aggregation, …), kept so the grouping into origin can be audited
relationshipType One of: Contains, HasAppRole, DelegatesScope, HasApplicationPermission, HasAppOwnership, HasOwnership (GrantsAccessTo is accepted but reserved / not yet emitted)
extendedAttributes Valid JSON object, max 64 KB

OpenAPI / Swagger

The API serves an OpenAPI 3.0 spec and Swagger UI:

  • GET /api/docs — Swagger UI (interactive documentation)
  • GET /api/docs/openapi.json — OpenAPI 3.0 spec file

From this spec, crawlers can auto-generate clients:

# Generate PowerShell client
npx @openapitools/openapi-generator-cli generate \
  -i openapi.json -g powershell -o ./crawler-client-ps

# Generate Python client
npx @openapitools/openapi-generator-cli generate \
  -i openapi.json -g python -o ./crawler-client-py

Observed performance

Measured against the committed load-test dataset (~2.17 M records, ~97 MB of CSV) on a VM with 6 cores / 16 GB RAM:

Phase Records Duration Throughput
Identities 25,000 9 s ~2,800 rows/s
Identity members 76,000 40 s ~1,900 rows/s
Certification decisions 300,000 4 min 12 s (1 batch) ~1,190 rows/s
Resource assignments 1,500,000 ~20 min (20 batches × 75 k) ~1,250 rows/s sustained
Full run ~2.17 M ~30 min ~1,200 rows/s overall

See Scaling & Load Testing for the full analysis, including hardware utilisation (CPU 73 %, memory 87 %, disk 1 % — memory is the limiting factor) and reproduction instructions.


Future Extensions

  • Webhook receiver — source systems push change events
  • NDJSON streaming — for very large datasets
  • Crawler SDK — npm/PyPI/PSGallery package with auth, chunking, retry
  • Crawler templates — wizard in admin UI generates boilerplate
  • Async ingestion — queue-based with job IDs
  • Data quality scoring — completeness and consistency metrics per sync