Network & Monitoring
Broken Link Checker
Extract links from a page and report HTTP status codes from the edge
Fetch a webpage, extract a[href] links, and probe each URL's HTTP status from Cloudflare's edge.
Features
- Extracts absolute and same-origin relative links from HTML
- Probes up to 25 links by default (max 50) with HEAD, falling back to GET
- Runs probes concurrently (5 at a time) with per-link timeouts
- Stateless; no bindings required
API Reference
GET /check
Prop
Type
Example Request
curl "https://your-worker.workers.dev/check?url=https://www.cloudflare.com&limit=10"Success Response
url string
Final page URL after redirects
checked number
Number of links probed
broken number
Count of non-OK responses
links array
Per-link { href, statusCode, ok, error? }
{
"url": "https://www.cloudflare.com/",
"checked": 10,
"broken": 1,
"links": [
{ "href": "https://www.cloudflare.com/", "statusCode": 200, "ok": true },
{ "href": "https://www.cloudflare.com/missing", "statusCode": 404, "ok": false }
]
}Error Codes
400- Missing or invalidurl(INVALID_URL)400- Response was not HTML (NOT_HTML)502- Page fetch failed (FETCH_ERROR)
Use Cases
- Find dead links before publishing content
- Spot-check navigation and footer URLs
- Complement link extraction tools with live status checks
- Build simple site QA workflows on Workers
Limitations
- Does not execute JavaScript; only links present in raw HTML
- Caps probes (default 25) to stay under Worker time limits
- Some origins block HEAD or bots; status codes may reflect bot filtering
Use in your project
Copy these files into an existing Worker. Prefer Deployment to try the full experiment first. Source: apps/experiments/broken-link-checker.
DependenciesNoneBindingsNonePlatformFetch API
import { PAGE_FETCH_TIMEOUT_MS, USER_AGENT } from "../constants/defaults";import type { CheckResponse } from "../types/check";import { extractLinks } from "./extract";import { probeLinks } from "./probe";export type PageFetchResult = | { ok: true; url: string; html: string } | { ok: false; code: "FETCH_ERROR" | "NOT_HTML"; error: string };export async function fetchPageHtml(url: string): Promise<PageFetchResult> { const controller = new AbortController(); const timeoutId = setTimeout(() => controller.abort(), PAGE_FETCH_TIMEOUT_MS); try { const res = await fetch(url, { method: "GET", redirect: "follow", signal: controller.signal, headers: { "User-Agent": USER_AGENT, Accept: "text/html,*/*" }, }); clearTimeout(timeoutId); if (!res.ok) { return { ok: false, code: "FETCH_ERROR", error: `HTTP ${res.status}` }; } const contentType = res.headers.get("content-type") ?? ""; if (!contentType.includes("text/html") && !contentType.includes("application/xhtml")) { return { ok: false, code: "NOT_HTML", error: `Expected HTML, got content-type: ${contentType || "unknown"}`, }; } return { ok: true, url: res.url || url, html: await res.text() }; } catch (e) { clearTimeout(timeoutId); return { ok: false, code: "FETCH_ERROR", error: e instanceof Error ? e.message : "Unknown error", }; }}export async function checkBrokenLinks( pageUrl: string, html: string, limit: number): Promise<CheckResponse> { const hrefs = extractLinks(html, pageUrl).slice(0, limit); const links = await probeLinks(hrefs); const broken = links.filter((l) => !l.ok).length; return { url: pageUrl, checked: links.length, broken, links, };}Deployment
Deploy
Follow the deployment wizard. No bindings or secrets required.
Test your deployment
curl "https://your-worker.workers.dev/check?url=https://example.com"Local Development
cd apps/experiments/broken-link-checker
npm install
npm run devcurl "http://localhost:8787/check?url=https://www.cloudflare.com"