Docs
Status
OverviewQuickstartSetup promptsThe core loopAuthenticationOverviewModelsThe waterfallPlansAdding modelsData controlsOpenAI compatibilityEmbeddingsAnthropic APIErrorsIntegrate the gatewayCost APIAccount APICoding agentsCredits & billingSpend & intelligenceSpend APITelemetryBecome a providerProvider guideAPI reference

Get started

  • Overview
  • Quickstart
  • Setup prompts
  • The core loop
  • Authentication

Guides

  • Overview
  • Models
  • The waterfall
  • Plans
  • Adding models
  • Data controls
  • OpenAI compatibility
  • Embeddings
  • Anthropic API
  • Errors

Integrations

  • Integrate the gateway
  • Cost API
  • Account API
  • Coding agents

Billing & usage

  • Credits & billing
  • Spend & intelligence
  • Spend API
  • Telemetry

Providers

  • Become a provider
  • Provider guide

Reference

  • API reference

Billing & usage

Telemetry

Read recorded gateway usage and spend over the API with the same org key, or bring traces from your existing stack in as telemetry.

For filtered financial breakdowns, CSV exports, and read-only usage analysis, open Spend or read the Spend guide. API-key owners and caller-supplied end-user labels are separate attribution dimensions.

Usage rollups

GET /api/gateway/usage/daily returns a grouped rollup of requests, tokens, and spend. Pass org_id and group_by; an API key reads at scope=org (scope=self needs an end-user session). Spend is reported in spend_nano_usd (billionths of a US dollar).

GET /api/gateway/usage/daily
curl "https://api.experientiallabs.ai/api/gateway/usage/daily?org_id=$ORG_ID&scope=org&group_by=day" \
-H "Authorization: Bearer $EXPLABS_API_KEY"
group_byOne row per
dayCalendar day (the row carries day).
day_modelCalendar day and model (the row carries day and alias).
modelModel slug (the row carries alias).
memberOrg member or end user (the row carries user_id).

Per-request events

GET /api/gateway/usage/events is the paginated per-request stream behind the rollup: one row per call, with its model, lane, tokens, and spend. Use it to attribute cost to a specific request or to export raw usage.

GET /api/gateway/usage/events
curl "https://api.experientiallabs.ai/api/gateway/usage/events?org_id=$ORG_ID&limit=50" \
-H "Authorization: Bearer $EXPLABS_API_KEY"

Reading usage in Logs

Logs keeps known token counts, including zero, and labels estimated usage as Estimated. A no data label with a Usage unavailable tooltip means the counter is missing, not zero. Cached input is a subset of input; an unavailable cache count does not mean a cache miss. Hover the token cells for cached input, reasoning, and usage provenance. Older rows may retain counts without a recorded source.

To find recorded cache hits, open Logs and enable Cache hits only beside Errors only in Request history. This shows requests with positive recorded cached input tokens, not necessarily every cache hit. It filters on the server before pagination and works with the other request filters, Load more, and Auto-refresh. The shareable URL uses cache=hit; the request-log API uses cache_hits_only=true. Spend and token totals, charts, and rollups stay unchanged. Rows outside this filter are not classified as cache misses.

Not dispatched means the backend confirmed no provider dispatch and no provider usage. Speed needs known output tokens and positive elapsed time; estimated output produces an Estimated speed. Request evidence covers the request, not each retry in a provider-attempt view. Monetary amounts stay as recorded. A zero charge with missing metering does not prove provider usage was free, and displaying recovered usage does not add a charge.

Mint one key per agent or workload so group_by=member and the event stream break spend out per caller. Humans see the same data at Telemetry and Credits.

Bring your own traces

Beyond gateway metering, you can land traces from your existing observability stack as telemetry. This never builds a router or spends credits; it just stores the traces under your organization.

Live pull from a provider

POST /api/orgs/<org_id>/telemetry/traces/pull pulls from a provider you name in transport_kind (one of braintrust, langsmith, langfuse, posthog, mastra, postgres). The credential is a single string, used once and never stored on the row: for Langfuse it is the public_key:secret_key pair; for the token-based providers it is the API key; for postgres it is the connection string.

POST /api/orgs/{org_id}/telemetry/traces/pull
curl -X POST "https://api.experientiallabs.ai/api/orgs/$ORG_ID/telemetry/traces/pull" \
-H "Authorization: Bearer $EXPLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transport_kind": "langfuse",
"source_kind": "langfuse",
"source_label": "prod",
"credential": "pk-lf-...:sk-lf-..."
}'

Upload a file

POST /api/orgs/<org_id>/telemetry/traces/upload reserves an ingest-scoped Storage path after authenticating the xpl_ key. The body is JSON {source_kind, source_label} (one of otel-genai, otlp, langfuse, langsmith, phoenix, braintrust, mastra, posthog, chat-json). The response is a two-hour, path-bound signed upload URL and token, never service credentials. PUT the exact raw file bytes to signed_url, then POST /api/orgs/<org_id>/telemetry/traces/<ingest_id>/finalize to enqueue verification. Finalize returns 202 quickly; the worker checks existence, the 50MB limit, format, exact size, and SHA-256 before projecting. Arize and Phoenix have no live pull yet, so upload them with source_kind=phoenix (or otlp).

POST /api/orgs/{org_id}/telemetry/traces/upload
TICKET=$(curl -sS -X POST "https://api.experientiallabs.ai/api/orgs/$ORG_ID/telemetry/traces/upload" \
-H "Authorization: Bearer $EXPLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"source_kind":"otlp","source_label":"prod-otel-august"}')
curl -sS -X PUT "$(echo "$TICKET" | jq -r .signed_url)" \
-H "Content-Type: application/octet-stream" \
--data-binary @traces.jsonl
curl -sS -X POST "https://api.experientiallabs.ai/api/orgs/$ORG_ID/telemetry/traces/$(echo "$TICKET" | jq -r .ingest_id)/finalize" \
-H "Authorization: Bearer $EXPLABS_API_KEY"

GET /api/orgs/<org_id>/telemetry/traceslists the org's landed traces with total_ingests and total_traces, the verify count.

Export calls to Langfuse

The gateway can send every call your keys make to your own Langfuse project as an OpenTelemetry span: input and output, model, provider, token usage, cost, time to first token, and errors, streaming calls included. Connect Langfuse once under Settings, Connections (public key, secret key, and your region's host such as https://us.cloud.langfuse.com), then turn on Live export. Nothing changes in your requests. An agent can do the same with the org's own key: PUT /api/orgs/<org_id>/observation-export with {"enabled": true, "langfuse": {"public_key", "secret_key", "host"}} (omit langfuse to toggle an existing connection; DELETE turns export off).

To keep your application's traces and put each gateway call under the right parent, send the W3C traceparentheader of the span that makes the call. Langfuse's SDKs are built on OpenTelemetry, so the standard propagator writes it for the active observation:

A gateway call nested under your trace
curl "https://api.experientiallabs.ai/v1/chat/completions" \
-H "Authorization: Bearer $EXPLABS_API_KEY" \
-H "traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" \
-H "X-Explabs-Session-Id: session-42" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Hello"}]}'
HeaderEffect
traceparentW3C trace context. The call becomes a child of this span.
X-Explabs-Trace-NameNames the trace. Omit it when your app already names its trace.
X-Explabs-Session-IdLangfuse session id.
X-Explabs-User-IdLangfuse user id. Defaults to the body's user (or metadata.user_id on Messages).
X-Explabs-Observation-NameNames the observation. Defaults to the operation and model, for example chat gpt-5.
X-Explabs-TagsYour existing request tags, recorded as observation metadata.

Without a traceparenteach call is its own trace, named after the call. Embeddings appear as embedding observations recording the input and the vector count and size, not the vectors. Cost is what your organization was charged; on your own provider key it is that provider's list price for the tokens, or omitted when the model has no price, so Langfuse can infer it.

Prompts and responses are sent from memory and never stored by the gateway for this. Turn off Include prompts and responses to export metadata only. A request routed under zero data retention (the organization setting or provider.zdr) always exports metadata only. Delivery is best effort, like an OpenTelemetry batch exporter: if Langfuse is unreachable or rejects the keys, spans are dropped and your calls are unaffected. Turning on live export turns off broadcast on the same connection, so no call arrives twice. If your application already wraps its OpenAI client with a Langfuse integration, you will see both generations; keep one. Export covers HTTP calls the gateway accepts: a request refused before acceptance (a malformed body or an invalid key) and Responses turns over the WebSocket transport are not exported.

See also

How each call is paid for and how to cap spend is in Credits & billing, and every usage endpoint is in the API reference.

PreviousSpend APINextBecome a provider