Docs
Status
OverviewQuickstartSetup promptsThe core loopAuthenticationOverviewModelsThe waterfallPlansAdding modelsData controlsOpenAI compatibilityEmbeddingsAnthropic APIErrorsIntegrate the gatewayCost APIAccount APICoding agentsCredits & billingSpend & intelligenceSpend APITelemetryBecome a providerProvider guideAPI reference

Get started

  • Overview
  • Quickstart
  • Setup prompts
  • The core loop
  • Authentication

Guides

  • Overview
  • Models
  • The waterfall
  • Plans
  • Adding models
  • Data controls
  • OpenAI compatibility
  • Embeddings
  • Anthropic API
  • Errors

Integrations

  • Integrate the gateway
  • Cost API
  • Account API
  • Coding agents

Billing & usage

  • Credits & billing
  • Spend & intelligence
  • Spend API
  • Telemetry

Providers

  • Become a provider
  • Provider guide

Reference

  • API reference

Billing & usage

Spend API

Exact, scoped usage reports for integrations and agents. The dashboard reads the same reporting stores, money definitions and snapshot contracts.

Base URL and authority

The canonical reporting base is https://api.experientiallabs.ai/api/orgs/{org_id}/spend. Send a provisioning xpl_ key as Authorization: Bearer <key>; find its organization with GET /api/whoami. Keep management credentials on your server. See Account API for creating the first provisioning key.

  • Every canonical Spend route, including handle reads and comparison cancellation, requires provisioning strength for API keys. An inference key receives 403; another organization receives 404.
  • API keys use scope: "org" only. Personal scope: "self" requires a signed-in human member and is bound by the server, never by a supplied user ID or actor header.
  • The trusted backend retains USER-strength member reads. The dashboard retains its existing workspace-admin or platform-admin gate; this API does not widen its browser routes or make them cookie-free key endpoints.
  • GET /api/v1/key remains the presented inference key's own usage/limit summary. This is not a claim that every inference-key read is key-scoped: the existing Cost API activity and settled usage export retain their documented org scope, with separate restrictions for plan-bound keys.

Ordinary analytics and prepared reports are free and make no model call. They are separate from credit-billed Intelligence.

Acquire once, then poll the same load

  1. Send POST /analytics with an unpinned query and refresh: false. Omit as_of and snapshot_id. Add comparison: "previous_elapsed" only when you need the prior period.
  2. On status: "loading", persist id. The response has no report, build, or candidate snapshot. Poll GET /analytics/{load_id}, not another POST.
  3. On ready, use report.query as your next request's input. A fast-path ready response can have id: null and build: null; it needs no polling.
  4. On failed, read error_code. A transport timeout is an unknown outcome, not permission to repost. If you already have an id, resume its GET. Without one, stop and reconcile before deliberately starting new work.
Start analytics with curl
# EXPLABS_API_KEY must be a provisioning key; ORG_ID comes from /api/whoami.
curl --fail-with-body --connect-timeout 5 --max-time 20 \
"https://api.experientiallabs.ai/api/orgs/$ORG_ID/spend/analytics" \
-H "Authorization: Bearer $EXPLABS_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"query": {
"start": "2026-09-01T00:00:00Z",
"end": "2026-09-08T00:00:00Z",
"time_zone": "UTC",
"scope": "org",
"basis": "requests",
"group_by": "model",
"grain": "day",
"filters": [],
"tags": [],
"sort": "spend",
"direction": "desc",
"offset": 0,
"limit": 200
},
"refresh": false,
"comparison": "previous_elapsed"
}'
The server acquisition deadline is 60 seconds from issuance. The dashboard observes for up to 70 seconds to allow delivery; neither polling nor reconnecting extends the server deadline. Those are not per-HTTP-request timeouts or a guarantee of completion. The bounded example below uses a 70-second observation budget and separate network timeouts. There is no promised Retry-After header.

refresh: false can reuse a completed result with its real cutoff. Use refresh: true only for an explicit refresh; keep the prior same-slice complete report visible while waiting. Do not combine its totals with a new chart.

One snapshot for every projection

C = report.coverage.projected_through is the certified data cutoff. F = report.query.as_of is the read fence, also echoed as coverage.as_of. They are different timestamps; C can precede F. A ready result has complete: true, pending_requests: 0, a non-null C, and a registered snapshot_id. This certifies projection coverage, not that every provider supplied known usage.

Keep the returned as_of and snapshot_id together across query, view, series, facets, activity, export and compatible build handles. A client-authored timestamp alone is not a registered snapshot. Never replace F with your clock or label it as C.

For another table page, send the same frozen query to POST /query with offset: next_offset. Stop at null. Sorting and pagination do not change the population. For a different grouping, POST /view reads the report and selected-dimension series at that same registered snapshot. A valid report may accompany series.status: "unavailable"; do not render an old chart beside it.

A build id belongs to one grouping, grain and predicate set. Never reuse a model-grouped build's report or series as a key-grouped report. A regrouped table pages through ordinary /query with the same F and token, or uses a separately prepared build whose query matches. Expired or privacy-invalidated snapshots fail closed; explicit refresh replaces the full envelope.

Query language and limits

Unpinned SpendQuery
{
"start": "2026-09-01T00:00:00Z",
"end": "2026-09-08T00:00:00Z",
"time_zone": "UTC",
"scope": "org",
"basis": "requests",
"group_by": "model",
"grain": "day",
"filters": [],
"tags": [],
"sort": "spend",
"direction": "desc",
"offset": 0,
"limit": 200
}
ContractMeaning
start / end / time_zoneTimezone-aware ISO timestamps; start inclusive, end exclusive, settlement-time reporting. IANA timezone defaults UTC. Range at most 3,660 days. Hourly range at most seven days.
scope / basis / grainorg (default) or self; requests (default) or attempts; hour, day (default), month. Hours are UTC buckets; days/months follow the selected calendar.
group_by / tag_keymodel (default), key, member, end_user, provider, funding, surface, outcome, error, app, prompt, tag, source. Exactly one tag_key is required only for tag grouping. app is the calling agent or application (claude_code, codex, opencode, hermes, ...) classified from request headers; null means unidentified.
filtersAt most 12 unique dimensions, each {dimension,values,missing,exclude?}. Up to 50 unique values; key/member values are UUIDs. Tag predicates belong in tags instead.
tagsAt most 16 distinct keys, each {key,values,missing,exclude?}. Up to 50 values. Exact case-sensitive values, never prompt text. See the request-tag guide.
sort / directionspend (default), requests, tokens, errors, cache_rate, name; desc (default) or asc. Spend ranks all-in value. Chart leaders always rank spend descending, independent of table sort.
offset / limitoffset 0..100000 (default 0); limit 1..200 (default 50). Follow returned next_offset; totals cover the whole slice, not just that page.
as_of / snapshot_idReuse the server-returned fence and token. Analytics/latest reject pinned input. View requires both. Series needs as_of; durable consistency requires the registered token too.

Values within a predicate use OR; different predicates use AND. No predicate means all values. An empty included set with missing: false matches nothing. missing: true includes the unattributed bucket; exclude: true negates the whole value/missing set. Missing attribution remains distinct from the chart's Other remainder.

Charts contain at most 368 buckets and seven globally ranked entities plus exact Other, not a sample or the current table page. Series entries contain value, label, other, figures and bucket-aligned cells; dimension series also distinguishes kind: value | missing | other. Other is not a selectable identifier.

CSV supports at most 10,000 groups and 16 MiB, with nine decimal places in dollar columns and formula-safe text. Prepared output is bounded at 10,000 groups, 100,000 cells and 32 MiB of logical output; results expire after 24 hours. Preparation does not remove Activity or facet limits. No path silently shortens dates, samples totals or declares an incomplete export complete.

The eleven exact report figures

ContractMeaning
countInteger. Settled customer requests, or terminal provider attempts under basis=attempts. Not a successful-completion count.
rejected_countInteger. Requests with zero attempts and failed or expired_before_dispatch status. Cancellations are not rejections. Always zero on the attempts basis.
errorsInteger. Non-completed outcomes, including cancellations, incomplete responses, unknown interruptions and rejections. Not only provider errors; do not add rejected_count to it.
input_tokensInteger. Input volume, including cached input.
output_tokensInteger. Output volume, including reasoning.
cached_input_tokensInteger. Subset of input_tokens, not additional tokens.
reasoning_tokensInteger. Subset of output_tokens, not additional tokens.
paid_nano_usdDecimal integer string. Settled platform credit charges.
byok_nano_usdDecimal integer string. BYOK usage estimated at list rates, not a platform debit or provider invoice.
free_nano_usdDecimal integer string. Free promotional usage at list value, not a credit debit.
unknown_usage_countInteger. Usage whose provider accounting is unknown. A nonzero value makes the measured money provisional, even when projection coverage is complete.

The eleven fields occur on totals, each group, each bucket and each series cell. Money uses decimal integer strings: 1 USD = 1,000,000,000 nano-USD. Use Python integers or JavaScript BigInt, not floating-point addition. Report figures do not serialize a fourth spend_nano_usd field: derive all-in Spend as paid_nano_usd + byok_nano_usd + free_nano_usd. Only paid is charged credits.

Unknown usage is not free usage or a measured zero. Failed calls can incur charges; completed calls can charge zero. Pending holds are not finalized Spend. Request and attempt counts are alternative bases, never additive. Pre-admission protocol rejections remain outside this projection (includes_pre_admission_rejections: false); accepted failures before dispatch are counted requests.

Prompt groups hash system/developer content and tool declarations, not full requests, sessions or all user messages. Facet labels and Activity are content-free attribution, not proof of the end user's identity.

Endpoint reference

Every path in this table is relative to https://api.experientiallabs.ai/api/orgs/{org_id}/spend. Wrappers matter: /analytics takes {query,...}, while /query takes the query itself.

ContractMeaning
POST /analytics{query, refresh?:false, comparison?:"previous_elapsed"}. Unpinned query only. Returns {id,status,report,build,error_code,comparison?}; status is loading, ready or failed.
GET /analytics/{load_id}Resume the same load. Query parameters: sort, direction, offset, limit, comparison=previous_elapsed (optional). Same analytics envelope; preserve the original id.
DELETE /analytics/{load_id}/comparison204 with no body. Cancels optional comparison work, not the primary report. There is no primary-load DELETE endpoint.
POST /latest{query, comparison?:"previous_elapsed"}, unpinned only. No new primary work. Without comparison: {build,report} or null. With comparison: a ready analytics envelope or null. A request error is not a cache miss.
POST /snapshot{query}. Returns a server-frozen SpendQuery. Retain as_of and snapshot_id when present; timestamp-only mode is not a registered visibility fence.
POST /querySpendQuery directly, not {query}. Returns SpendReport: query, totals, buckets, groups, filter_labels, group_count, next_offset, coverage.
POST /viewSpendQuery directly, with registered as_of and snapshot_id. Returns {report,series:{status:"ready",data:SpendDimensionSeries}} or {report,series:{status:"unavailable",reason:"error"}}. Never starts analytics or a build.
POST /seriesFrozen SpendQuery directly. Returns {query,coverage,totals,buckets,series}; the selected group_by drives seven spend-ranked leaders plus exact Other.
POST /model-seriesFrozen SpendQuery directly. Returns the model-specific {query,coverage,totals,buckets,series} projection. Use /series for other dimensions.
POST /facets{query,kind,dimension?,tag_key?,search?,limit?,offset?}. Returns {query,kind,dimension,tag_key,search,options,next_offset,coverage}.
POST /activity{query,cursor?,errors_only?:false,include_features?:false}. Returns {query,coverage,rows,next_cursor}; cursor is {time,id}.
POST /exportSpendQuery directly. Complete grouped CSV, not JSON or request-level rows. Includes exact dollar columns, query, cutoff and measure definitions.
POST /builds{query}. Explicit free report preparation. Returns SpendBuild; unpinned input freezes at admission, pinned input retains its registered snapshot.
GET /builds/{build_id}SpendBuild: id, status, query, created_at, updated_at, expires_at, progress, error_code. Status: queued, running, completed, failed, cancelled or expired.
GET /builds/{build_id}/reportCompleted SpendReport. Query parameters: sort, direction, offset, limit. Does not change the build's grouping or filters.
GET /builds/{build_id}/seriesCompleted SpendDimensionSeries for this build's grouping, globally ranked by spend.
GET /builds/{build_id}/exportCompleted grouped CSV for this build's query. Same bounded export contract as POST /export.

Facet kinds are dimension (requires a non-tag dimension), tag_keys (neither dimension nor tag_key), or tag_values (requires tag_key). Search is at most 200 characters; limit and offset follow the query's 1..200 and 0..100000 bounds. Follow its own next_offset, retaining the frozen query. Discovery omits only its own target predicate so selected values do not hide alternatives.

Activity rows contain id, request_id, created_at, dimensions, tags, figures, detail. Pass its returned next_cursor unchanged with the same query. include_features: true adds source-discriminated Intelligence debit rows: source: "intelligence", request_id: null, zero API counts/tokens and only a paid money leg. Report-link metadata is member-authorized. Captured bodies are a separate member-authorized surface; a Spend key gains no prompt access.

Optional previous-period comparison

Opt in with comparison: "previous_elapsed" when starting analytics and on its GET polls. The primary report is usable while comparison.status is pending. A ready comparison has exact totals, start, end, as_of and projected_through. Pending or unavailable comparisons have no totals: do not substitute zero or calculate a percentage.

The observed current interval ends at min(query.end, C). Its equally long predecessor ends at query.start, with millisecond derivation milliseconds_v1. Comparison retains the primary C and F, not a separately refreshed report. Unavailable reasons are not_prepared, coverage, capacity, deadline, cancelled, invalidated, error. DELETE of the load's comparison leaves the primary report intact.

Bounded Python polling and pagination

Install httpx; set EXPLABS_API_KEY to a provisioning key and ORG_ID to its organization. The checkpoint stores the response, never the credential. Keep it private because reports contain attribution. A later invocation resumes the same id, never automatically repeats an uncertain POST. To start a different report, explicitly choose a new checkpoint after reviewing any unresolved admission.

Python: acquire, resume, page
import json
import os
import tempfile
import time
from pathlib import Path
import httpx
base = "https://api.experientiallabs.ai/api/orgs/" + os.environ["ORG_ID"] + "/spend"
headers = {"Authorization": "Bearer " + os.environ["EXPLABS_API_KEY"]}
checkpoint = Path("spend-checkpoint.json")
query = {
"start": "2026-09-01T00:00:00Z",
"end": "2026-09-08T00:00:00Z",
"time_zone": "UTC",
"scope": "org",
"basis": "requests",
"group_by": "model",
"grain": "day",
"filters": [],
"tags": [],
"sort": "spend",
"direction": "desc",
"offset": 0,
"limit": 200
}
# Run one polling process per checkpoint. No credential is stored.
record = {"base": base, "query": query, "started_at": time.time(),
"result": {"status": "admitting"}}
try:
fd = os.open(checkpoint, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
except FileExistsError:
resumed = True
record = json.loads(checkpoint.read_text(encoding="utf-8"))
if record["base"] != base or record["query"] != query:
raise RuntimeError("Checkpoint belongs to another org/query. Choose a new path.")
else:
resumed = False
with os.fdopen(fd, "w", encoding="utf-8") as stream:
json.dump(record, stream)
stream.flush()
os.fsync(stream.fileno())
def save(result):
record["result"] = result
with tempfile.NamedTemporaryFile(mode="w", encoding="utf-8", dir=checkpoint.parent,
prefix=".spend-", delete=False) as stream:
temporary = Path(stream.name) # tempfile creates it with mode 0600.
json.dump(record, stream)
stream.flush()
os.fsync(stream.fileno())
try:
os.replace(temporary, checkpoint)
finally:
temporary.unlink(missing_ok=True)
with httpx.Client(headers=headers, timeout=15.0, verify=True,
trust_env=False, follow_redirects=False) as client:
observe_until = time.monotonic() + max(0, record["started_at"] + 70 - time.time())
result = record["result"]
if resumed:
if result["status"] == "admitting":
raise RuntimeError("Admission outcome unknown. Reconcile; do not repost.")
# Reauthorize before displaying any cached attribution, even after 70s.
previous = result.get("report")
if result.get("id"):
page = {} if previous is None else {
name: previous["query"][name] for name in ("sort", "direction", "offset", "limit")
}
response = client.get(base + "/analytics/" + result["id"], params=page)
response.raise_for_status()
result = response.json()
else:
response = client.post(base + "/query", json=result["report"]["query"])
response.raise_for_status()
result["report"] = response.json()
if previous is not None and result["status"] == "ready":
assert result["report"]["coverage"] == previous["coverage"]
for field in ("as_of", "snapshot_id"):
assert result["report"]["query"][field] == previous["query"][field]
save(result)
else:
# The exclusive checkpoint claim is durable before admission.
response = client.post(base + "/analytics", json={"query": query, "refresh": False})
response.raise_for_status()
result = response.json()
save(result)
for poll in range(35):
if result["status"] != "loading":
break
remaining = observe_until - time.monotonic()
if remaining <= 2:
break
time.sleep(2)
response = client.get(
base + "/analytics/" + result["id"],
timeout=min(15.0, max(0.1, observe_until - time.monotonic())),
)
response.raise_for_status()
result = response.json()
save(result)
if result["status"] != "ready":
raise RuntimeError("No report yet; keep checkpoint. " + str(result.get("error_code")))
report = result["report"]
frozen = report["query"]
coverage = report["coverage"]
assert frozen["snapshot_id"] and frozen["as_of"]
assert coverage["complete"] and coverage["pending_requests"] == 0
assert coverage["projected_through"] is not None
assert coverage["as_of"] == frozen["as_of"]
totals = report["totals"]
all_in = sum(int(totals[k]) for k in ("paid_nano_usd", "byok_nano_usd", "free_nano_usd"))
print("All-in nano-USD:", all_in, "unknown usage:", totals["unknown_usage_count"])
page_until = time.monotonic() + 120
for page_number in range(100):
for group in report["groups"]:
print(json.dumps(group))
offset = report["next_offset"]
if offset is None:
break
if page_number == 99 or time.monotonic() >= page_until:
raise RuntimeError("Page budget reached; resume the checkpoint, same snapshot.")
# Ordinary query pagination: never borrow a differently grouped build id.
response = client.post(
base + "/query", json={**frozen, "offset": offset},
timeout=min(15.0, max(0.1, page_until - time.monotonic())),
)
response.raise_for_status()
report = response.json()
assert report["query"]["as_of"] == frozen["as_of"]
assert report["query"]["snapshot_id"] == frozen["snapshot_id"]
assert report["coverage"] == coverage
result["report"] = report
save(result)

The polling loop checks elapsed time and request count; HTTPX timeouts bound network phases, not an absolute process wall clock. Pagination also stops at a page and time budget. Checkpointed pages may be printed again after an interruption: deduplicate by frozen query and group value if persisting rows, and read totals once rather than adding repeated page totals. This script does not retry any POST on an error.

Lightweight key and identity summaries

GET /api/orgs/{org_id}/usage/by-key and GET /api/orgs/{org_id}/usage/by-identity are bounded live usage summaries for provisioning keys or existing member sessions. They share the API and dashboard reader, not a separate browser total. window=24h|7d|30d defaults to 7d.

ContractMeaning
by-key envelope{report_kind:"live_usage_summary",window,keys:[{api_key_id,key_label,models,totals,last_used_at}]}. Each model has model, request_count, error_count, input_tokens, output_tokens and the money fields below.
by-identity envelope{report_kind:"live_usage_summary",window,identities:[{identity_id,display_name,active,keys:[{api_key_id,key_label,request_count}],totals,last_used_at}]}. Null identity preserves traffic whose key no longer resolves.
totalsrequest_count, error_count, input_tokens, output_tokens, cost_usd, estimated_cost_usd, free_usd, paid_nano_usd, byok_nano_usd, free_nano_usd, spend_nano_usd.
moneyThe four nano-USD fields are exact decimal integer strings on model rows and key/identity totals. spend_nano_usd = paid + byok + free. Existing numeric cost_usd (paid), estimated_cost_usd (BYOK) and free_usd remain for display.
limitsOne consistent bounded rollup result per read, at most 10,000 key/model cells, not 10,000 keys, and a 32 MiB source-cell JSON budget; final response formatting has separate overhead. Transport overhead can reach the source-read bound earlier. No partial success or offset stitching; HTTP 422 usage_summary_too_large directs callers to Spend analytics.
report_kind: "live_usage_summary" is not a certified Spend report. There is no registered snapshot or coverage proof, and unknown usage is not reported, not zero. Do not synthesize complete: true or unknown_usage_count: 0. For frozen reconciliation, detailed attribution, arbitrary dates or large results, use Spend analytics.

Failures and recovery

Business errors return {error,code?}; inspect HTTP status and code. Schema errors can use FastAPI's detail array instead. Do not assume every error has a code or Retry-After. An unavailable comparison or chart is not a failed primary report, and an HTTP 200 analytics envelope still requires checking its status.

ContractMeaning
400 / 422Invalid query, unsupported wrapper, timezone or range. invalid_spend_query is a business code; schema validation uses {detail:[{loc,msg,type}]}. Fix the input, not the retry interval.
401 / 403 / 404Missing/invalid credential; inference key lacks provisioning authority; foreign org or inaccessible handle. Stop and clear retained private results. Never treat these as empty usage.
409 snapshot_expired / snapshot_invalidatedThe old financial or attribution envelope is no longer readable. Explicitly refresh, replacing every projection together, never silently mixing snapshots.
409 report_incompleteProjection is not complete for an export. Acquire a ready analytics report; do not download partial totals as final.
422 report_too_large / series_too_large / export_too_largePrepare an exact report where enabled or narrow filters/dates. Monthly grain solves bucket limits, not cardinality or source-work limits.
429 analytics_capacity / build_capacityNew work is at capacity. Preserve a previous valid result; retry admission only after an explicit decision, not an automatic POST loop.
503 analytics_disabled / builds_unavailable / reporting_initializing / report_timeoutEnvironment readiness or bounded read failure. Disabled preparation is not empty data. Retain an issued id and recover by GET; narrow an interactive query when appropriate.
404 build_not_found / analytics_not_foundUnknown or inaccessible handle. Check authority and the saved id; do not claim zero usage.
410 build_expired; 409 build_invalidated / build_not_readyExpired or invalidated results require explicit refresh. A not-ready build should be polled at its existing id.
status=failed + error_codeAn analytics response can succeed at HTTP level but fail its work, for example analytics_deadline. No report is present. Stop the load; a new load is an explicit action.
422 usage_summary_too_largeThe lightweight summary exceeded its cell/byte bound. Use POST /analytics, not repeated summary reads or client-side offset assembly.

Related: Spend dashboard and request tags, per-request Cost API, and machine-readable API guide.

PreviousSpend & intelligenceNextTelemetry