This site is not affiliated with or endorsed by Cloudflare, Inc. It simply showcases experiments built using Cloudflare services.
Cloudflare Experiments
AI & Machine Learning

AI Batch Infer

Workers AI async batch embeddings with queueRequest and poll

Submit a batch of Workers AI embedding inferences with queueRequest: true, then poll by request_id via the Asynchronous Batch API. Without an AI binding, routes return demo queued/poll responses.

Features

  • Queue embedding batches (POST /batch)
  • Poll results by request_id (GET /batch/:requestId)
  • Configurable model via MODEL var (default @cf/baai/bge-m3)

API Reference

POST /batch

Body

{
  "texts": ["first phrase", "second phrase"]
}
FieldRequiredDescription
textsYes1–20 non-empty strings, max 2000 chars each

Example Request

curl -X POST "https://your-worker.workers.dev/batch" \
  -H "Content-Type: application/json" \
  -d '{"texts":["Cloudflare Workers","edge compute"]}'

Success Response

{
  "status": "queued",
  "model": "@cf/baai/bge-m3",
  "request_id": "000-000-000",
  "mode": "live"
}

Error Codes

  • 400 - Invalid body (INVALID_BODY)
  • 502 - Batch submit failure (BATCH_ERROR)

GET /batch/:requestId

Poll status/results for a queued batch.

Example Request

curl "https://your-worker.workers.dev/batch/000-000-000"

While processing, the API may return queued or running. When complete, the response includes per-item results.

Error Codes

  • 400 - Invalid id (INVALID_REQUEST_ID)
  • 502 - Poll failure (BATCH_ERROR)

Use Cases

  • Offline embedding of large document sets
  • Batch summarization / classification without interactive latency
  • Avoid capacity errors by queueing inference for eventual completion

Limitations

  • Total batch payload must stay under ~10 MB (Workers AI limit)
  • Needs a batch-capable model and account AI access for live mode
  • Demo mode returns placeholder embeddings

Use in your project

Copy these files into an existing Worker. Prefer Deployment to try the full experiment first. Source: apps/experiments/ai-batch-infer.

DependenciesNoneBindingsWorkers AI (AI binding)PlatformWorkers AI
import {  DEFAULT_MODEL,  MAX_REQUEST_ID_LENGTH,  MAX_TEXT_LENGTH,  MAX_TEXTS,  MIN_TEXTS,} from "../constants/defaults";import type { BatchPollResponse, QueuedBatchResponse } from "../types/batch";import type { AIRunBinding } from "../types/env";export function resolveModel(model: string | undefined): string {  const trimmed = model?.trim();  return trimmed || DEFAULT_MODEL;}export function hasAIBinding(ai: unknown): ai is AIRunBinding {  return (    typeof ai === "object" &&    ai !== null &&    "run" in ai &&    typeof (ai as AIRunBinding).run === "function"  );}export function validateTexts(body: unknown): string[] | null {  if (!body || typeof body !== "object") return null;  const texts = (body as { texts?: unknown }).texts;  if (!Array.isArray(texts)) return null;  if (texts.length < MIN_TEXTS || texts.length > MAX_TEXTS) return null;  const validated: string[] = [];  for (const item of texts) {    if (typeof item !== "string") return null;    const trimmed = item.trim();    if (!trimmed || trimmed.length > MAX_TEXT_LENGTH) return null;    validated.push(trimmed);  }  return validated;}export function validateRequestId(input: string | undefined): string | null {  if (!input || typeof input !== "string") return null;  const trimmed = input.trim();  if (!trimmed || trimmed.length > MAX_REQUEST_ID_LENGTH) return null;  if (!/^[a-zA-Z0-9._-]+$/.test(trimmed)) return null;  return trimmed;}export function demoQueuedResponse(model: string): QueuedBatchResponse {  return {    status: "queued",    model,    request_id: `demo-${crypto.randomUUID()}`,    mode: "demo",    note: "AI binding unavailable; returning a demo queued response",  };}export function demoPollResponse(requestId: string, model: string): BatchPollResponse {  return {    status: "completed",    model,    request_id: requestId,    mode: "demo",    note: "AI binding unavailable; returning demo poll results",    responses: [      {        id: 0,        success: true,        result: {          data: [0.1, 0.2, 0.3],        },      },    ],  };}function isQueuedResponse(value: unknown): value is QueuedBatchResponse {  if (!value || typeof value !== "object") return false;  const obj = value as Record<string, unknown>;  return (    obj.status === "queued" && typeof obj.request_id === "string" && typeof obj.model === "string"  );}export async function submitBatch(  ai: AIRunBinding,  model: string,  texts: string[]): Promise<QueuedBatchResponse> {  const result = await ai.run(    model,    { requests: texts.map((text) => ({ text })) },    { queueRequest: true }  );  if (!isQueuedResponse(result)) {    throw new Error("Unexpected batch queue response shape");  }  return { ...result, mode: "live" };}export async function pollBatch(  ai: AIRunBinding,  model: string,  requestId: string): Promise<BatchPollResponse> {  const result = await ai.run(model, { request_id: requestId });  if (!result || typeof result !== "object") {    throw new Error("Unexpected batch poll response shape");  }  return { ...(result as BatchPollResponse), mode: "live" };}

Deployment

Click the deploy button

Deploy to Cloudflare Workers

Confirm AI binding

wrangler.json declares ai.binding: "AI" and vars.MODEL (default @cf/baai/bge-m3).

Test your deployment

curl -X POST "https://your-worker.workers.dev/batch" \
  -H "Content-Type: application/json" \
  -d '{"texts":["hello","world"]}'

Local Development

cd apps/experiments/ai-batch-infer
npm install
npm run dev

Workers AI batch typically needs remote AI / account access.

Configuration

  • Workers AI binding AI
  • Var MODEL — batch embedding model id

Cloudflare Features Used

On this page