Docs
Status
OverviewQuickstartSetup promptsThe core loopAuthenticationOverviewModelsThe waterfallPlansAdding modelsData controlsOpenAI compatibilityEmbeddingsAnthropic APIErrorsIntegrate the gatewayCost APIAccount APICoding agentsCredits & billingSpend & intelligenceSpend APITelemetryBecome a providerProvider guideAPI reference

Get started

  • Overview
  • Quickstart
  • Setup prompts
  • The core loop
  • Authentication

Guides

  • Overview
  • Models
  • The waterfall
  • Plans
  • Adding models
  • Data controls
  • OpenAI compatibility
  • Embeddings
  • Anthropic API
  • Errors

Integrations

  • Integrate the gateway
  • Cost API
  • Account API
  • Coding agents

Billing & usage

  • Credits & billing
  • Spend & intelligence
  • Spend API
  • Telemetry

Providers

  • Become a provider
  • Provider guide

Reference

  • API reference

Guides

Embeddings

Turn text into vectors for semantic search, retrieval, and similarity with the OpenAI embeddings API. Use the same organization key as your other gateway calls.

Endpoint and authentication

Send POST https://api.experientiallabs.ai/v1/embeddings with Authorization: Bearer <key> and Content-Type: application/json. Use an inference-enabled xpl_ key from API Keys. Keep the key in an environment variable, not in source control.

The canonical OpenAI SDK base URL is https://api.experientiallabs.ai/v1. The unified integration base also serves POST https://api.experientiallabs.ai/api/v1/embeddings. A client that appends /embeddings to the bare host reaches the same relay; these are not separate serving or caching implementations.

Set your organization key
export EXPLABS_API_KEY="xpl_..." # replace with your actual key

Choose an embedding model

Check GET https://api.experientiallabs.ai/v1/models with your key for actual availability before sending a request. The launch targets are direct OpenAI text-embedding-3-small, text-embedding-3-large, and text-embedding-ada-002; a model must be activated and available to your key to be callable. These examples use text-embedding-3-small.

Discover available models
curl "https://api.experientiallabs.ai/v1/models" \
-H "Authorization: Bearer $EXPLABS_API_KEY"

The model catalog carries current rates and model details. Input-token limits, batch size, and supported dimensions depend on the selected model and provider. Split oversized inputs or batches rather than assuming every OpenAI-compatible model has the same limits.

Create embeddings

Run python -m pip install openai for the Python example, or npm install openai for JavaScript. Both read the key you exported above. A single string or a flat token-ID array embeds one item; a list of strings or token-ID arrays embeds a batch. Match each result to its input using data[].index.

POST /v1/embeddings
curl "https://api.experientiallabs.ai/v1/embeddings" \
-H "Authorization: Bearer $EXPLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-small",
"input": "A small example for semantic search.",
"encoding_format": "float"
}'
The official OpenAI Python SDK requests base64 when you omit encoding_format, then decodes each vector into floats before returning it. That default works here. Explicitly set encoding_format="float" for JSON vectors on the wire, or encoding_format="base64"to receive the encoded strings without the SDK's automatic decoding.

Supported request fields

FieldContract
modelRequired. An embedding model slug available to your key. A chat model is not an embedding model.
inputRequired. A nonempty string; a nonempty list of nonempty strings; a nonempty array of nonnegative integer token IDs (one vector); or a nonempty list of nonempty token-ID arrays (a batch). Do not mix text and token arrays in one batch. Empty strings or arrays, negative IDs, booleans, and floats are rejected.
dimensionsOptional positive integer supported by the selected model. OpenAI text-embedding-3 models support reduced dimensions; text-embedding-ada-002 does not. Omit it to use the model's default size.
encoding_formatOptional. "float" returns arrays of numbers; "base64" returns base64-encoded vectors. Omit it for float on the raw HTTP API.
streamOptional. Omit it or send false for non-streaming compatibility only. true is unsupported; all other values are invalid. The response is always one JSON body.
userOptional string, at most 1,024 characters. Gateway-only end-user attribution; not forwarded to the upstream provider.
Token IDs must match the chosen model's tokenizer. The gateway forwards them without decoding or automatic conversion between tokenizers. Upstream models enforce vocabulary and input limits; gateway acceptance does not guarantee token-array support for every custom provider.

Unknown parameters are rejected, not silently ignored. There is no streaming: omit stream or send stream: false for non-streaming compatibility only. stream: true is unsupported. Chat options, tool calls, and prompt-cache controls do not apply to this endpoint.

For the upstream contract, see the OpenAI embeddings reference. The examples below use the tokenizer for the selected OpenAI model.

Send token IDs with tiktoken

Install python -m pip install openai tiktoken and use the key exported above. Choose the tokenizer from the model name, not from the gateway hostname. This example sends two token arrays in one request. To embed just one item, pass input=token_batches[0] instead; a flat array produces one vector, not one per token.

Tokenize for the chosen model
import os
import tiktoken
from openai import OpenAI
model = "text-embedding-3-small"
encoding = tiktoken.encoding_for_model(model)
documents = ["A small example for semantic search.", "A second document."]
token_batches = [encoding.encode(text) for text in documents]
client = OpenAI(
base_url="https://api.experientiallabs.ai/v1",
api_key=os.environ["EXPLABS_API_KEY"],
max_retries=0,
)
result = client.embeddings.create(model=model, input=token_batches)
for item in result.data:
print(item.index, len(item.embedding))

See the tiktoken documentation. Do not reuse these IDs with a different model unless its tokenizer matches. For a custom model, follow that model's tokenizer and provider documentation rather than assuming it uses tiktoken.

LangChain OpenAIEmbeddings

Install python -m pip install langchain-openai in your application. Use the standard OpenAIEmbeddings client with your gateway base URL and key. Its default length-checking path tokenizes with tiktoken and sends token-array batches for this OpenAI model; leave that behavior enabled. No custom client or tokenizer workaround is needed for this setup.

Embed documents and a query
import os
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(
model="text-embedding-3-small",
base_url="https://api.experientiallabs.ai/v1",
api_key=os.environ["EXPLABS_API_KEY"],
max_retries=0,
)
vectors = embeddings.embed_documents(
["A small example for semantic search.", "A second document."]
)
query_vector = embeddings.embed_query("What is semantic search?")
print(len(vectors), len(vectors[0]), len(query_vector))

LangChain may split long documents and combine their chunk vectors. Provider limits still apply to each request. See the OpenAIEmbeddings reference for that client's behavior. A custom provider may need a different tokenizer or input format; this example does not establish support for every OpenAI-compatible endpoint.

Response and usage

The response has object: "list", a model, and a data list. Each item has object: "embedding", a zero-based index, and an embedding containing either floats or a base64 string. The response arrives as one JSON body, never SSE.

Billing uses provider-reported input tokens: usage.prompt_tokens equals usage.total_tokens. Embeddings generate no output tokens; the gateway records zero output tokens internally. Vector length is not an output-token charge. Platform-funded calls draw credits at the active catalog input rate; on BYOK, your provider bills you directly. See Credits & billing and your usage history for settled spend.

Your own Azure deployments

Embeddings serve on your own Azure OpenAI deployments. Add each Azure resource as an account under the Azure OpenAI provider and map the model (for example text-embedding-3-large) to that resource's deployment name. On the model's page, order the Azure accounts and, if you want a last resort, your OpenAI key. A request tries the first account and fails over down that list on a provider error or exhausted quota; a rate limit fails over as the model's failover policy decides (see Waterfall). Your list never falls through to platform-funded credit.

Caching and retries

There is no embeddings response cache and no caller-keyed replay. Repeating a request sends it upstream again and can incur another charge, even with the same Idempotency-Key. The gateway does not honor that header on embeddings.

Completion prompt caching does not cache embedding vectors. Internal router vector reuse is not a cache for this endpoint either. Store and reuse vectors in your application if you want to avoid repeated provider work. The SDK examples disable automatic retries with max_retries=0 (Python) or maxRetries: 0 (JavaScript), so retrying is an explicit choice. After a timeout, the provider may already have done billable work.

Data controls and content guardrails

Embeddings do not run identity content-guardrail classifiers. A policy on your API key's identity does not inspect, block, or redact embedding input or vectors. Those checks apply to chat completions, not this endpoint; validate sensitive input in your application before sending it.

Content inspection is separate from provider routing and retention. Read Data controls for the organization's storage and provider-policy settings.

Errors and recovery

  • 400 invalid_parameter: fix the field named by error.param, such as empty input, mixed text and token arrays, a boolean or float token ID, an invalid dimension, or a non-boolean stream value.
  • 400 unsupported_parameter: remove a field this endpoint does not accept, such as chat parameters, or turn off stream: true.
  • 400 unsupported_capability: the selected model does not serve embeddings. Choose an embedding model from the model listing.
  • 401: check the key. For model access, credit, or rate refusals, follow the returned error message. Provider limits and failures can also prevent a valid request from completing.

Read the error reference for the common envelope and recovery steps. Fix validation errors before retrying, and remember that an embeddings retry is a new billable request, not an idempotent replay.

PreviousOpenAI compatibilityNextAnthropic API