API reference
Keep your SDK.
Change the base URL.
Base URL: https://api.embeddings.app/v1. Authenticate with your embeddings.app workspace API key. Keep production keys in server-side code.
中文語料,直接送出。
JavaScript / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.EMBEDDINGS_API_KEY,
baseURL: "https://api.embeddings.app/v1",
});
const result = await client.embeddings.create({
model: "bge-m3",
input: "The quick brown fox jumps over the lazy dog.",
encoding_format: "float",
});
console.log(result.data[0].embedding.length); // 1024Pass encoding_format: "float" explicitly, including with OpenAI SDK v6. Use "base64" for little-endian float32 vectors.
Python
from openai import OpenAI
import os
client = OpenAI(api_key=os.environ["EMBEDDINGS_API_KEY"],
base_url="https://api.embeddings.app/v1")
result = client.embeddings.create(model="bge-m3",
input="A document to search.",
encoding_format="float")
print(len(result.data[0].embedding)) # 1024Raw HTTP and batches
curl https://api.embeddings.app/v1/embeddings \
-H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"bge-m3","input":["Workspace pulse","收件箱"],"encoding_format":"float"}'model and input are required. Input is a non-empty string or an array of 1–150 strings. Token arrays and streaming are rejected.
Optional fields: dimensions, encoding_format and input_type (document, the default, or query). Requests are limited to 256 KiB; each model has a per-input byte limit.
The response contains object: "list", model and ordered data. meter records exact billed UTF-8 bytes, cached inputs, charge and balance in integer micro-USD. usage appears only when the provider reports token counts; counts are not invented for Workers AI.
Inputs are Unicode NFC-normalized. Identical model, dimensions, purpose and text reuse persistent vectors. Successful duplicates are free, including duplicates within a batch.
Document search
Upload up to 60,000 UTF-8 bytes of plain text. The original is private in R2; deterministic chunks use BGE-M3 and a 1,024-dimensional cosine Vectorize index.
curl https://api.embeddings.app/v1/documents \
-H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"guide.txt","text":"Our product helps teams search their documents."}'A successful ingest returns a mutation ID and index_status: "pending". Vectorize visibility is asynchronous; retry search after the index catches up.
curl https://api.embeddings.app/v1/search \
-H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query":"How do teams find documents?","top_k":5}'Search returns scored chunks and source metadata from your workspace namespace. Pass those sources to your preferred answer model. Model execution for generated answers is outside this API.
Endpoints
| Endpoint | Result |
|---|---|
| GET /v1/models | Configured models and service prices |
| POST /v1/embeddings | Ordered vectors and usage meter |
| POST /v1/documents | Private document ingestion |
| POST /v1/search | Workspace-scoped nearest chunks |
| GET /v1/usage | Balance and recent payload-free operations |
| POST /v1/billing/checkout | Hosted checkout when billing is configured |
Errors and spending limits
400: invalid input or unsupported options. 401: invalid or revoked key. 402: insufficient prepaid credit. 429: workspace or service limit. 503: missing configuration or unavailable index.
502 / 409: an uncertain provider outcome is held for reconciliation. Credit was reserved before dispatch, and the service will not automatically retry a possibly billed operation. Contact support@gridheap.com with the operation ID from your usage history.
Initial hard limits: 120 workspace requests/minute, 2 million workspace input bytes/day, 200 provider calls/day and 1 million new input bytes/day across the service. Operators may configure the service limits. There are no scheduled backfills, recurring jobs or automatic top-ups.