Observability
Send AIOE's traces and metrics to your OpenTelemetry collector, and read its logs.
The API writes JSON logs to standard output, and can send traces and metrics to any OpenTelemetry collector over OTLP/HTTP: the OpenTelemetry Collector, Grafana Alloy, Datadog, Honeycomb, or Azure Monitor's OTLP ingestion. Telemetry is off until you name a collector.
Turn on traces and metrics
Set these on the API and restart it:
| Variable | Default | What it does |
|---|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT | unset (off) | The collector's base address, for example http://otel-collector:4318. The API sends to /v1/traces and /v1/metrics under it. |
OTEL_EXPORTER_OTLP_HEADERS | none | Extra headers for the collector, as comma-separated key=value pairs, for example an Authorization header. Keep this in a secret. |
OTEL_SERVICE_NAME | aioe-api | The service name on every span and metric. |
AIOE_ENVIRONMENT | NODE_ENV, else development | Sent as deployment.environment.name. The Helm chart sets production; the Azure template sets the env parameter. |
OTEL_METRIC_EXPORT_INTERVAL | 30000 | Milliseconds between metric exports. |
OTEL_LOG_LEVEL | unset | debug prints the OpenTelemetry SDK's own diagnostics, to troubleshoot the export. |
api:
env:
AIOE_ENVIRONMENT: production
OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318Put OTEL_EXPORTER_OTLP_HEADERS in the API's secret rather than in api.env.
What you get
Traces for every HTTP request (except /healthz and /readyz), with spans for Express routing, PostgreSQL queries and Redis commands. Log lines written during a traced request carry its trace ID, so you can move from a log line to its trace.
Metrics of HTTP traffic, plus AIOE's own:
| Metric | Type | Unit | What it counts |
|---|---|---|---|
aioe.ingest.envelopes | counter | envelope | Ingest envelopes accepted from workbenches and nodes. |
aioe.ingest.records | counter | record | Records accepted, by stream. |
aioe.relay.dispatches | counter | request | Remote-control requests dispatched, by outcome. |
aioe.relay.rpc.duration | histogram | ms | Round trip from a remote-control request to the workbench's answer. |
aioe.relay.streams | up-down counter | stream | Workbench relay connections open on this replica. |
aioe.enrolments | counter | enrolment | Workbench enrolments, by outcome. |
aioe.auth.failures | counter | request | Rejected tokens, by kind (person, device, either). |
Good first alerts: aioe.auth.failures rising sharply, aioe.relay.rpc.duration growing, and aioe.relay.streams falling to zero on every replica.
One trace from workbench to platform. AI Workbench can export its own telemetry over OTLP too (the observability.otlpEndpoint setting, see Workbench settings). Point it at the same collector to follow a request from a workbench into AIOE.
Logs
The API logs one JSON object per line to standard output, with a request log for every call except the health checks. Set the level with LOG_LEVEL (info by default; debug, warn and error also work).
Lines worth knowing at start-up:
| Line | Meaning |
|---|---|
connected to postgres | The database is reachable. |
STORE=memory: nothing persists across restarts | You are on the in-memory store. Fine for a trial, never for real use. |
bootstrap tenant ready | Your organisation was created or updated from the AIOE_BOOTSTRAP_* settings. |
relay broker: redis (multi-replica) | Redis is in use. |
AIOE_SIGNING_KEY unset: a signing key was generated | Set the key when you run more than one replica. |
remote control: mesh first, relay fallback | The overlay gateway is configured. |
aioe api listening | The API is serving, with its port and public address. |
Health checks
| Path | Answers | Use it for |
|---|---|---|
/healthz | { "ok": true, "version": "…" } | Liveness, and to see which version is running. |
/readyz | { "ok": true } | Readiness. |
The console answers /healthz on its own port.