embeddings.app

Embeddings, with less plumbing

From text to vectors.
One API. Your choice of model.

Connect your retrieval pipeline to direct embedding providers through one OpenAI-compatible endpoint. Prepaid usage, reusable vectors, and document search on Cloudflare.

Managed BGE-M3 · 1,024 dimensions · 中文語料,直接送出。

Use the client you already know.

Read the API docs
curl https://api.embeddings.app/v1/embeddings \
  -H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"bge-m3","input":["Workspace pulse","收件箱"],"encoding_format":"float"}'

A single string or up to 150 inputs. One vector per item, in the same order.

Your retrieval pipeline

Generate. Store. Find.

01

Bring your text

Upload a bounded plain-text document. The private original lives in R2.

02

Generate once

Workers AI embeds deterministic chunks. A shared cache reuses successful vectors.

03

Search your corpus

Vectorize retrieves matching chunks in your workspace. Feed the sources to your own answer model.

Stable model and vector size throughout retrieval. Index updates are asynchronous.

A model change, not an integration project.

ModelProviderDimensions$/1M input bytesStatus
bge-m3cloudflare1,024$0.10Live
text-embedding-3-smallopenai1,536 configurable$0.40Needs credentials
text-embedding-3-largeopenai3,072 configurable$0.80Needs credentials
voyage-3.5voyage1,024 configurable$0.60Needs credentials
embed-v4.0cohere1,536 configurable$0.80Needs credentials
jina-embeddings-v3jina1,024 configurable$0.60Needs credentials
gemini-embedding-001google3,072 configurable$0.80Needs credentials
qwen3-embedding-0-6bdatabricks1,024 configurable$0.40Needs credentials
gte-large-endatabricks1,024$0.80Needs credentials

BGE-M3 is the initial managed model. Other adapters are implemented and become callable when their credentials are configured. Inspect the catalogue.

Pay for new work

Only unique uncached inputs spend prepaid credit. See the charge and remaining balance in each response.

A budget that stops calls

Requests reserve credit before dispatch. Daily byte and provider-call caps bound spending; no automatic retries or top-ups.

Providers stay behind the API

Your application holds one workspace key. Provider credentials stay on the server.

Start with your own corpus.

Prepaid access is available by request. Tell us what you are building and which models you need.