Billing & usage
Telemetry
Read recorded gateway usage and spend over the API with the same org key, or bring traces from your existing stack in as telemetry.
For filtered financial breakdowns, CSV exports, and read-only usage analysis, open Spend or read the Spend guide. API-key owners and caller-supplied end-user labels are separate attribution dimensions.
Usage rollups
GET /api/gateway/usage/daily returns a grouped rollup of requests, tokens, and spend. Pass org_id and group_by; an API key reads at scope=org (scope=self needs an end-user session). Spend is reported in spend_nano_usd (billionths of a US dollar).
curl "https://api.experientiallabs.ai/api/gateway/usage/daily?org_id=$ORG_ID&scope=org&group_by=day" \-H "Authorization: Bearer $EXPLABS_API_KEY"
| group_by | One row per |
|---|---|
| day | Calendar day (the row carries day). |
| day_model | Calendar day and model (the row carries day and alias). |
| model | Model slug (the row carries alias). |
| member | Org member or end user (the row carries user_id). |
Per-request events
GET /api/gateway/usage/events is the paginated per-request stream behind the rollup: one row per call, with its model, lane, tokens, and spend. Use it to attribute cost to a specific request or to export raw usage.
curl "https://api.experientiallabs.ai/api/gateway/usage/events?org_id=$ORG_ID&limit=50" \-H "Authorization: Bearer $EXPLABS_API_KEY"
Reading usage in Logs
Logs keeps known token counts, including zero, and labels estimated usage as Estimated. A no data label with a Usage unavailable tooltip means the counter is missing, not zero. Cached input is a subset of input; an unavailable cache count does not mean a cache miss. Hover the token cells for cached input, reasoning, and usage provenance. Older rows may retain counts without a recorded source.
To find recorded cache hits, open Logs and enable Cache hits only beside Errors only in Request history. This shows requests with positive recorded cached input tokens, not necessarily every cache hit. It filters on the server before pagination and works with the other request filters, Load more, and Auto-refresh. The shareable URL uses cache=hit; the request-log API uses cache_hits_only=true. Spend and token totals, charts, and rollups stay unchanged. Rows outside this filter are not classified as cache misses.
Not dispatched means the backend confirmed no provider dispatch and no provider usage. Speed needs known output tokens and positive elapsed time; estimated output produces an Estimated speed. Request evidence covers the request, not each retry in a provider-attempt view. Monetary amounts stay as recorded. A zero charge with missing metering does not prove provider usage was free, and displaying recovered usage does not add a charge.
Bring your own traces
Beyond gateway metering, you can land traces from your existing observability stack as telemetry. This never builds a router or spends credits; it just stores the traces under your organization.
Live pull from a provider
POST /api/orgs/<org_id>/telemetry/traces/pull pulls from a provider you name in transport_kind (one of braintrust, langsmith, langfuse, posthog, mastra, postgres). The credential is a single string, used once and never stored on the row: for Langfuse it is the public_key:secret_key pair; for the token-based providers it is the API key; for postgres it is the connection string.
curl -X POST "https://api.experientiallabs.ai/api/orgs/$ORG_ID/telemetry/traces/pull" \-H "Authorization: Bearer $EXPLABS_API_KEY" \-H "Content-Type: application/json" \-d '{"transport_kind": "langfuse","source_kind": "langfuse","source_label": "prod","credential": "pk-lf-...:sk-lf-..."}'
Upload a file
POST /api/orgs/<org_id>/telemetry/traces/upload reserves an ingest-scoped Storage path after authenticating the xpl_ key. The body is JSON {source_kind, source_label} (one of otel-genai, otlp, langfuse, langsmith, phoenix, braintrust, mastra, posthog, chat-json). The response is a two-hour, path-bound signed upload URL and token, never service credentials. PUT the exact raw file bytes to signed_url, then POST /api/orgs/<org_id>/telemetry/traces/<ingest_id>/finalize to enqueue verification. Finalize returns 202 quickly; the worker checks existence, the 50MB limit, format, exact size, and SHA-256 before projecting. Arize and Phoenix have no live pull yet, so upload them with source_kind=phoenix (or otlp).
TICKET=$(curl -sS -X POST "https://api.experientiallabs.ai/api/orgs/$ORG_ID/telemetry/traces/upload" \-H "Authorization: Bearer $EXPLABS_API_KEY" \-H "Content-Type: application/json" \-d '{"source_kind":"otlp","source_label":"prod-otel-august"}')curl -sS -X PUT "$(echo "$TICKET" | jq -r .signed_url)" \-H "Content-Type: application/octet-stream" \--data-binary @traces.jsonlcurl -sS -X POST "https://api.experientiallabs.ai/api/orgs/$ORG_ID/telemetry/traces/$(echo "$TICKET" | jq -r .ingest_id)/finalize" \-H "Authorization: Bearer $EXPLABS_API_KEY"
GET /api/orgs/<org_id>/telemetry/traceslists the org's landed traces with total_ingests and total_traces, the verify count.
Export calls to Langfuse
The gateway can send every call your keys make to your own Langfuse project as an OpenTelemetry span: input and output, model, provider, token usage, cost, time to first token, and errors, streaming calls included. Connect Langfuse once under Settings, Connections (public key, secret key, and your region's host such as https://us.cloud.langfuse.com), then turn on Live export. Nothing changes in your requests. An agent can do the same with the org's own key: PUT /api/orgs/<org_id>/observation-export with {"enabled": true, "langfuse": {"public_key", "secret_key", "host"}} (omit langfuse to toggle an existing connection; DELETE turns export off).
To keep your application's traces and put each gateway call under the right parent, send the W3C traceparentheader of the span that makes the call. Langfuse's SDKs are built on OpenTelemetry, so the standard propagator writes it for the active observation:
curl "https://api.experientiallabs.ai/v1/chat/completions" \-H "Authorization: Bearer $EXPLABS_API_KEY" \-H "traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" \-H "X-Explabs-Session-Id: session-42" \-H "Content-Type: application/json" \-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Hello"}]}'
| Header | Effect |
|---|---|
| traceparent | W3C trace context. The call becomes a child of this span. |
| X-Explabs-Trace-Name | Names the trace. Omit it when your app already names its trace. |
| X-Explabs-Session-Id | Langfuse session id. |
| X-Explabs-User-Id | Langfuse user id. Defaults to the body's user (or metadata.user_id on Messages). |
| X-Explabs-Observation-Name | Names the observation. Defaults to the operation and model, for example chat gpt-5. |
| X-Explabs-Tags | Your existing request tags, recorded as observation metadata. |
Without a traceparenteach call is its own trace, named after the call. Embeddings appear as embedding observations recording the input and the vector count and size, not the vectors. Cost is what your organization was charged; on your own provider key it is that provider's list price for the tokens, or omitted when the model has no price, so Langfuse can infer it.
provider.zdr) always exports metadata only. Delivery is best effort, like an OpenTelemetry batch exporter: if Langfuse is unreachable or rejects the keys, spans are dropped and your calls are unaffected. Turning on live export turns off broadcast on the same connection, so no call arrives twice. If your application already wraps its OpenAI client with a Langfuse integration, you will see both generations; keep one. Export covers HTTP calls the gateway accepts: a request refused before acceptance (a malformed body or an invalid key) and Responses turns over the WebSocket transport are not exported.See also
How each call is paid for and how to cap spend is in Credits & billing, and every usage endpoint is in the API reference.