embeddings.app

API reference

Keep your SDK.
Change the base URL.

Base URL: https://api.embeddings.app/v1. Authenticate with your embeddings.app workspace API key. Keep production keys in server-side code.

中文語料,直接送出。

JavaScript / TypeScript

import OpenAI from "openai";
const client = new OpenAI({
  apiKey: process.env.EMBEDDINGS_API_KEY,
  baseURL: "https://api.embeddings.app/v1",
});
const result = await client.embeddings.create({
  model: "bge-m3",
  input: "The quick brown fox jumps over the lazy dog.",
  encoding_format: "float",
});
console.log(result.data[0].embedding.length); // 1024

Pass encoding_format: "float" explicitly, including with OpenAI SDK v6. Use "base64" for little-endian float32 vectors.

Python

from openai import OpenAI
import os
client = OpenAI(api_key=os.environ["EMBEDDINGS_API_KEY"],
                base_url="https://api.embeddings.app/v1")
result = client.embeddings.create(model="bge-m3",
                                 input="A document to search.",
                                 encoding_format="float")
print(len(result.data[0].embedding)) # 1024

Raw HTTP and batches

curl https://api.embeddings.app/v1/embeddings \
  -H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"bge-m3","input":["Workspace pulse","收件箱"],"encoding_format":"float"}'

model and input are required. Input is a non-empty string or an array of 1–150 strings. Token arrays and streaming are rejected.

Optional fields: dimensions, encoding_format and input_type (document, the default, or query). Requests are limited to 256 KiB; each model has a per-input byte limit.

The response contains object: "list", model and ordered data. meter records exact billed UTF-8 bytes, cached inputs, charge and balance in integer micro-USD. usage appears only when the provider reports token counts; counts are not invented for Workers AI.

Inputs are Unicode NFC-normalized. Identical model, dimensions, purpose and text reuse persistent vectors. Successful duplicates are free, including duplicates within a batch.

Document search

Upload up to 60,000 UTF-8 bytes of plain text. The original is private in R2; deterministic chunks use BGE-M3 and a 1,024-dimensional cosine Vectorize index.

curl https://api.embeddings.app/v1/documents \
  -H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name":"guide.txt","text":"Our product helps teams search their documents."}'

A successful ingest returns a mutation ID and index_status: "pending". Vectorize visibility is asynchronous; retry search after the index catches up.

curl https://api.embeddings.app/v1/search \
  -H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"How do teams find documents?","top_k":5}'

Search returns scored chunks and source metadata from your workspace namespace. Pass those sources to your preferred answer model. Model execution for generated answers is outside this API.

Endpoints

EndpointResult
GET /v1/modelsConfigured models and service prices
POST /v1/embeddingsOrdered vectors and usage meter
POST /v1/documentsPrivate document ingestion
POST /v1/searchWorkspace-scoped nearest chunks
GET /v1/usageBalance and recent payload-free operations
POST /v1/billing/checkoutHosted checkout when billing is configured

Errors and spending limits

400: invalid input or unsupported options. 401: invalid or revoked key. 402: insufficient prepaid credit. 429: workspace or service limit. 503: missing configuration or unavailable index.

502 / 409: an uncertain provider outcome is held for reconciliation. Credit was reserved before dispatch, and the service will not automatically retry a possibly billed operation. Contact support@gridheap.com with the operation ID from your usage history.

Initial hard limits: 120 workspace requests/minute, 2 million workspace input bytes/day, 200 provider calls/day and 1 million new input bytes/day across the service. Operators may configure the service limits. There are no scheduled backfills, recurring jobs or automatic top-ups.