AI & Machine Learning
AI Batch Infer
Workers AI async batch embeddings with queueRequest and poll
Submit a batch of Workers AI embedding inferences with queueRequest: true, then poll by request_id via the Asynchronous Batch API. Without an AI binding, routes return demo queued/poll responses.
Features
- Queue embedding batches (
POST /batch) - Poll results by
request_id(GET /batch/:requestId) - Configurable model via
MODELvar (default@cf/baai/bge-m3)
API Reference
POST /batch
Body
{
"texts": ["first phrase", "second phrase"]
}| Field | Required | Description |
|---|---|---|
texts | Yes | 1–20 non-empty strings, max 2000 chars each |
Example Request
curl -X POST "https://your-worker.workers.dev/batch" \
-H "Content-Type: application/json" \
-d '{"texts":["Cloudflare Workers","edge compute"]}'Success Response
{
"status": "queued",
"model": "@cf/baai/bge-m3",
"request_id": "000-000-000",
"mode": "live"
}Error Codes
400- Invalid body (INVALID_BODY)502- Batch submit failure (BATCH_ERROR)
GET /batch/:requestId
Poll status/results for a queued batch.
Example Request
curl "https://your-worker.workers.dev/batch/000-000-000"While processing, the API may return queued or running. When complete, the response includes per-item results.
Error Codes
400- Invalid id (INVALID_REQUEST_ID)502- Poll failure (BATCH_ERROR)
Use Cases
- Offline embedding of large document sets
- Batch summarization / classification without interactive latency
- Avoid capacity errors by queueing inference for eventual completion
Limitations
- Total batch payload must stay under ~10 MB (Workers AI limit)
- Needs a batch-capable model and account AI access for live mode
- Demo mode returns placeholder embeddings
Use in your project
Copy these files into an existing Worker. Prefer Deployment to try the full experiment first. Source: apps/experiments/ai-batch-infer.
DependenciesNoneBindingsWorkers AI (AI binding)PlatformWorkers AI
import { DEFAULT_MODEL, MAX_REQUEST_ID_LENGTH, MAX_TEXT_LENGTH, MAX_TEXTS, MIN_TEXTS,} from "../constants/defaults";import type { BatchPollResponse, QueuedBatchResponse } from "../types/batch";import type { AIRunBinding } from "../types/env";export function resolveModel(model: string | undefined): string { const trimmed = model?.trim(); return trimmed || DEFAULT_MODEL;}export function hasAIBinding(ai: unknown): ai is AIRunBinding { return ( typeof ai === "object" && ai !== null && "run" in ai && typeof (ai as AIRunBinding).run === "function" );}export function validateTexts(body: unknown): string[] | null { if (!body || typeof body !== "object") return null; const texts = (body as { texts?: unknown }).texts; if (!Array.isArray(texts)) return null; if (texts.length < MIN_TEXTS || texts.length > MAX_TEXTS) return null; const validated: string[] = []; for (const item of texts) { if (typeof item !== "string") return null; const trimmed = item.trim(); if (!trimmed || trimmed.length > MAX_TEXT_LENGTH) return null; validated.push(trimmed); } return validated;}export function validateRequestId(input: string | undefined): string | null { if (!input || typeof input !== "string") return null; const trimmed = input.trim(); if (!trimmed || trimmed.length > MAX_REQUEST_ID_LENGTH) return null; if (!/^[a-zA-Z0-9._-]+$/.test(trimmed)) return null; return trimmed;}export function demoQueuedResponse(model: string): QueuedBatchResponse { return { status: "queued", model, request_id: `demo-${crypto.randomUUID()}`, mode: "demo", note: "AI binding unavailable; returning a demo queued response", };}export function demoPollResponse(requestId: string, model: string): BatchPollResponse { return { status: "completed", model, request_id: requestId, mode: "demo", note: "AI binding unavailable; returning demo poll results", responses: [ { id: 0, success: true, result: { data: [0.1, 0.2, 0.3], }, }, ], };}function isQueuedResponse(value: unknown): value is QueuedBatchResponse { if (!value || typeof value !== "object") return false; const obj = value as Record<string, unknown>; return ( obj.status === "queued" && typeof obj.request_id === "string" && typeof obj.model === "string" );}export async function submitBatch( ai: AIRunBinding, model: string, texts: string[]): Promise<QueuedBatchResponse> { const result = await ai.run( model, { requests: texts.map((text) => ({ text })) }, { queueRequest: true } ); if (!isQueuedResponse(result)) { throw new Error("Unexpected batch queue response shape"); } return { ...result, mode: "live" };}export async function pollBatch( ai: AIRunBinding, model: string, requestId: string): Promise<BatchPollResponse> { const result = await ai.run(model, { request_id: requestId }); if (!result || typeof result !== "object") { throw new Error("Unexpected batch poll response shape"); } return { ...(result as BatchPollResponse), mode: "live" };}Deployment
Confirm AI binding
wrangler.json declares ai.binding: "AI" and vars.MODEL (default @cf/baai/bge-m3).
Test your deployment
curl -X POST "https://your-worker.workers.dev/batch" \
-H "Content-Type: application/json" \
-d '{"texts":["hello","world"]}'Local Development
cd apps/experiments/ai-batch-infer
npm install
npm run devWorkers AI batch typically needs remote AI / account access.
Configuration
- Workers AI binding
AI - Var
MODEL— batch embedding model id
Cloudflare Features Used
- Workers - Edge compute runtime
- Workers AI - Inference at the edge
- Asynchronous Batch API -
queueRequest+ poll byrequest_id