Embeddings, with less plumbing
From text to vectors.
One API. Your choice of model.
Connect your retrieval pipeline to direct embedding providers through one OpenAI-compatible endpoint. Prepaid usage, reusable vectors, and document search on Cloudflare.
Managed BGE-M3 · 1,024 dimensions · 中文語料,直接送出。
Use the client you already know.
Read the API docscurl https://api.embeddings.app/v1/embeddings \
-H "Authorization: Bearer $EMBEDDINGS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"bge-m3","input":["Workspace pulse","收件箱"],"encoding_format":"float"}'A single string or up to 150 inputs. One vector per item, in the same order.
Your retrieval pipeline
Generate. Store. Find.
Bring your text
Upload a bounded plain-text document. The private original lives in R2.
Generate once
Workers AI embeds deterministic chunks. A shared cache reuses successful vectors.
Search your corpus
Vectorize retrieves matching chunks in your workspace. Feed the sources to your own answer model.
Stable model and vector size throughout retrieval. Index updates are asynchronous.
A model change, not an integration project.
| Model | Provider | Dimensions | $/1M input bytes | Status |
|---|---|---|---|---|
bge-m3 | cloudflare | 1,024 | $0.10 | Live |
text-embedding-3-small | openai | 1,536 configurable | $0.40 | Needs credentials |
text-embedding-3-large | openai | 3,072 configurable | $0.80 | Needs credentials |
voyage-3.5 | voyage | 1,024 configurable | $0.60 | Needs credentials |
embed-v4.0 | cohere | 1,536 configurable | $0.80 | Needs credentials |
jina-embeddings-v3 | jina | 1,024 configurable | $0.60 | Needs credentials |
gemini-embedding-001 | 3,072 configurable | $0.80 | Needs credentials | |
qwen3-embedding-0-6b | databricks | 1,024 configurable | $0.40 | Needs credentials |
gte-large-en | databricks | 1,024 | $0.80 | Needs credentials |
BGE-M3 is the initial managed model. Other adapters are implemented and become callable when their credentials are configured. Inspect the catalogue.
Pay for new work
Only unique uncached inputs spend prepaid credit. See the charge and remaining balance in each response.
A budget that stops calls
Requests reserve credit before dispatch. Daily byte and provider-call caps bound spending; no automatic retries or top-ups.
Providers stay behind the API
Your application holds one workspace key. Provider credentials stay on the server.
Start with your own corpus.
Prepaid access is available by request. Tell us what you are building and which models you need.