> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hadoseo.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Cloudflare Worker Middleware: Call Hado SEO From Your Own Edge

> Run your own Cloudflare Worker in front of your app that pre-filters bot traffic and fetches prerendered HTML from Hado SEO. No DNS proxy required — keep your existing infrastructure.

## Self-Hosted Middleware: Prerender Without Changing Your DNS

Most Hado SEO setups point your DNS at our proxy. If you'd rather **keep your own
infrastructure** — your CDN, your Worker, your routing — you can run a thin
Cloudflare Worker in front of your site that pre-filters bot traffic and calls the
Hado SEO **render endpoint** for prerendered HTML.

<Info>
  **Advanced / self-hosted.** This guide is for teams that want to integrate at
  the edge themselves. If you just want SEO with zero code, use the
  [DNS setup](/guides/dns-setup) instead.
</Info>

### How it works

Your Worker is a **thin, untrusted client**. All the heavy lifting — authentication,
domain authorization, authoritative bot verification, caching, and rendering — lives
in the Hado SEO service. Your Worker only has to:

1. Decide if a request *looks* like a bot (a cheap User-Agent check).
2. For bots, ask the Hado SEO render endpoint for prerendered HTML.
3. Serve that HTML on `200`, or serve your app normally on `204` (passthrough).

<Note>
  Your bot check can be loose. Hado SEO **re-verifies every bot authoritatively**
  (User-Agent + published IP ranges), so a human or spoofed crawler that slips
  through simply gets a `204` and is served your normal app. False positives only
  cost one subrequest.
</Note>

## Prerequisites

* Your site's domain is on **Cloudflare** (the zone is active and serving your
  traffic) — the Worker runs on your zone, in front of your existing app.
* **Node 18+** locally, for the `wrangler` CLI. (Everything below can also be
  done in the Cloudflare dashboard UI; the CLI path is shown because it's
  exactly reproducible.)

## Setup

<Steps>
  <Step title="Add your domain in Hado SEO">
    Sign up at [hadoseo.com/auth](https://hadoseo.com/auth). In onboarding,
    when asked **"How are you connecting?"**, choose **Cloudflare Worker**.
    Enter your custom domain — that's all; unlike the DNS path, no app URL is
    needed. This registers **domain ownership**: the render endpoint only
    serves domains your account owns.
  </Step>

  <Step title="Copy your API key">
    The onboarding success screen issues an API key (`hado_sk_...`) and shows
    it **once** — copy it now. If you already closed it, create a new key in
    **Settings → API Keys**. See [Authentication](/api-reference/authentication).
  </Step>

  <Step title="Copy your render endpoint">
    From the same screen (or the dashboard's domain page):

    ```
    https://serve.hadoseo.com/v1/render
    ```
  </Step>

  <Step title="Create the Worker project">
    The fastest path is cloning the ready-made template — Worker, config, and
    deploy scripts included:

    ```bash theme={null}
    git clone https://github.com/hado-seo/hado-cloudflare-worker.git my-site-seo-proxy
    cd my-site-seo-proxy && npm install
    ```

    Or build it by hand — a fresh directory with three files is all it takes:

    ```bash theme={null}
    mkdir my-site-seo-proxy && cd my-site-seo-proxy
    npm init -y && npm install -D wrangler
    ```

    Save [the Worker below](#the-worker) as `src/worker.js`, and add:

    ```toml wrangler.toml theme={null}
    name = "my-site-seo-proxy"
    main = "src/worker.js"
    compatibility_date = "2025-05-05"

    # Route the Worker in front of your site — every hostname that serves
    # pages. Include www (or any other subdomain) if it serves traffic.
    # Keep routes ABOVE the [vars] section: in TOML, keys below a [table]
    # header belong to that table, and routes inside [vars] attach nothing.
    routes = [
      { pattern = "example.com/*", zone_name = "example.com" },
      { pattern = "www.example.com/*", zone_name = "example.com" }
    ]

    [vars]
    HADO_RENDER_ENDPOINT = "https://serve.hadoseo.com/v1/render"
    ```

    Replace `example.com` with your domain in both `pattern`s and `zone_name`.
  </Step>

  <Step title="Set your API key as a secret">
    ```bash theme={null}
    npx wrangler login                    # first time only
    npx wrangler secret put HADO_API_KEY  # paste hado_sk_... when prompted
    ```

    The key is a **secret**, never a `[vars]` entry — vars are visible in the
    dashboard and in `wrangler.toml`.
  </Step>

  <Step title="Deploy">
    ```bash theme={null}
    npx wrangler deploy
    ```

    Deploying attaches the routes: from this moment every request to your
    domain passes through the Worker. Humans are completely unaffected (one
    header check, then your app) — there is no cutover moment to schedule.
  </Step>

  <Step title="Verify">
    Three checks, in order:

    ```bash theme={null}
    # 1. Humans still get your app untouched.
    curl -sI https://example.com/ | head -3

    # 2. A bot UA gets prerendered HTML. The FIRST hit of a page may return
    #    your app shell (the render finishes in the background); run it twice.
    curl -s -A "GPTBot/1.0" https://example.com/ | head -40
    curl -s -A "GPTBot/1.0" https://example.com/ | head -40   # → full HTML
    ```

    3. Open your Hado SEO dashboard → **Analytics**: the test crawls appear
       within a few minutes. Then run the
       [SEO Bot Crawler Test](https://hadoseo.com/free-seo-bot-crawler-test)
       against a page to confirm real crawlers receive fully rendered HTML.
  </Step>
</Steps>

## The Worker

This is the same `src/worker.js` shipped in the
[hado-cloudflare-worker template](https://github.com/hado-seo/hado-cloudflare-worker) —
clone that instead of copy-pasting if you want the wrangler config and deploy
scripts along with it.

```js src/worker.js theme={null}
// A loose pre-filter derived from Hado's bot registry and known-good crawler
// list. Hado SEO re-verifies bots authoritatively, so a false positive only
// costs one subrequest — but every token here must be BOT-specific: a token
// that also appears in human in-app-browser UAs (e.g. "pinterest",
// "snapchat") would make real visitors wait on the render call.
const BOT_UA = new RegExp(
  [
    // Search engines
    "googlebot|googleother|storebot-google|google-inspectiontool|adsbot-google",
    "adidxbot|mediapartners|apis-google|feedfetcher-google|google-read-aloud",
    "bingbot|yandex|baiduspider|duckduckbot|slurp|seznambot|sogou|exabot",
    "petalbot|seekport|bravebot|applebot|amazonbot",
    // AI crawlers, assistants, and user-initiated agent fetches
    "gptbot|oai-searchbot|chatgpt-user|chatgpt-agent",
    "claudebot|claude-user|claude-searchbot|anthropic-ai",
    "perplexitybot|perplexity-user|mistralai-user|duckassistbot|amzn-user",
    "google-extended|google-cloudvertexbot|google-notebooklm",
    "gemini-deep-research|googleagent-mariner|meta-externalfetcher",
    "ccbot|bytespider|tiktokspider|youbot|cohere|diffbot|omgili|ai2bot",
    "timpibot|tavilybot",
    // Social & messaging link previews
    "facebookexternalhit|facebookbot|facebot|meta-externalagent|twitterbot",
    "linkedinbot|pinterestbot|slackbot|slack-imgproxy|discordbot|whatsapp",
    "telegrambot|redditbot|quora-bot|bitlybot|skypeuripreview",
    "googlemessages|google-chat-link-preview|bufferlinkpreviewbot",
    "metadatascraper",
    // SEO & monitoring tools
    "ahrefsbot|semrushbot|siteauditbot|dataforseobot|rsiteauditor|sebot-wa",
    "screaming frog|mj12bot|dotbot|rogerbot|blexbot|barkrowler|serpstatbot",
    "seokicks|gtmetrix|uptimebot|ia_archiver|archive\\.org_bot",
  ].join("|"),
  "i",
);

// How long to wait for a cold render before giving up and serving your app.
const RENDER_TIMEOUT_MS = 15000;

export default {
  async fetch(request, env, ctx) {
    const ua = request.headers.get("user-agent") || "";
    const isBot = BOT_UA.test(ua);

    // Humans: record AI referrals (ChatGPT, Perplexity, …) with a
    // fire-and-forget beacon, then serve your app untouched.
    if (request.method === "GET" && !isBot) {
      const referer = request.headers.get("referer");
      const dest = request.headers.get("sec-fetch-dest") || "document";
      if (referer && dest === "document") {
        ctx.waitUntil(
          fetch(new URL("/v1/event/referral", env.HADO_RENDER_ENDPOINT), {
            method: "POST",
            headers: {
              authorization: `Bearer ${env.HADO_API_KEY}`,
              "content-type": "application/json",
              "user-agent": ua, // the ORIGINAL visitor UA
            },
            body: JSON.stringify({ url: request.url, referer }),
          }).catch(() => {}),
        );
      }
      return fetch(request);
    }

    // Only prerender bot GET navigations. Everything else → your app, untouched.
    if (request.method !== "GET") {
      return fetch(request);
    }

    try {
      const endpoint = new URL(env.HADO_RENDER_ENDPOINT);
      endpoint.searchParams.set("url", request.url);

      const controller = new AbortController();
      const timer = setTimeout(() => controller.abort(), RENDER_TIMEOUT_MS);

      const res = await fetch(endpoint, {
        signal: controller.signal,
        headers: {
          authorization: `Bearer ${env.HADO_API_KEY}`,
          // Forward the ORIGINAL bot identity so Hado can verify it.
          "user-agent": ua,
          // The visitor IP MUST travel in x-hadoseo-client-ip: on a
          // worker-to-worker hop, transport-level headers like
          // cf-connecting-ip are rewritten to YOUR Worker's egress IP,
          // and real crawlers would be rejected as spoofed.
          "x-hadoseo-client-ip": request.headers.get("cf-connecting-ip") || "",
          // Analytics-only? Uncomment to record the crawl without prerendering:
          // "x-hadoseo-mode": "passthrough",
        },
      }).finally(() => clearTimeout(timer));

      // 200 → prerendered HTML; 404 → a soft-404 snapshot (your app renders
      // "not found" content for this route). Serve BOTH with their real
      // status — swallowing the 404 would show crawlers a 200 for a page
      // that doesn't exist (a soft-404 penalty). 404 is safe to pass
      // through: Hado's caller errors are only ever 400/401/403/429.
      if (res.status === 200 || res.status === 404) {
        return new Response(res.body, {
          status: res.status,
          headers: {
            "content-type": "text/html; charset=utf-8",
            // Reuse Hado's cache directives so your edge can cache repeat hits.
            "cache-control":
              res.headers.get("cache-control") || "public, max-age=0",
          },
        });
      }

      // 301/302/307/308 → a routing rule you configured in the Hado
      // dashboard. Return the redirect to the bot verbatim.
      const location = res.headers.get("location");
      if (res.status >= 301 && res.status <= 308 && location) {
        return new Response(null, { status: res.status, headers: { location } });
      }

      // A blocked path's noindex/nofollow directives ride the 204 —
      // copy them onto your app's response as X-Robots-Tag.
      const robots = res.headers.get("x-hadoseo-robots");
      if (robots) {
        const originRes = await fetch(request);
        const withRobots = new Response(originRes.body, originRes);
        withRobots.headers.set("x-robots-tag", robots);
        return withRobots;
      }

      // Plain 204 (human/spoofed/over-limit) or any 4xx → your app, untouched.
    } catch (err) {
      // Network error or timeout → fail open.
    }

    return fetch(request);
  },
};
```

## What the Worker sends

`GET {HADO_RENDER_ENDPOINT}?url=<absolute request URL>` with:

| Header | Required | Purpose |
| - | - | - |
| `Authorization: Bearer <key>` | ✅ | Your API key (or `x-api-key: <key>`). |
| `user-agent` | ✅ | The **original** bot UA — Hado verifies it. |
| `x-hadoseo-client-ip` | ✅ | The **original** client IP (copy your `cf-connecting-ip` into it) — used for authoritative IP verification. It must travel in this dedicated header: on a worker-to-worker hop the transport-level `cf-connecting-ip` is rewritten to your Worker's own egress IP, and real crawlers would look spoofed. |
| `x-hadoseo-mode: passthrough` | — | Optional. Analytics-only mode (see below). |

The Worker also fires a fire-and-forget `POST {HADO_RENDER_ENDPOINT origin}/v1/event/referral`
beacon for **human page navigations that carry a Referer**, with body
`{"url": "<landing URL>", "referer": "<referer>"}` and the same auth. This is
how visits arriving from AI assistants (ChatGPT, Perplexity, Claude, …) show up
in your AI-referral analytics — no IP is ever sent on this path. It always
returns `204`; failures are silently ignored and never affect the visitor.

<Warning>
  Always forward the **visitor's** `user-agent`, and copy the visitor's IP into
  `x-hadoseo-client-ip` — never your Worker's own values. Hado SEO verifies bots
  by matching the UA against published crawler IP ranges — if the IP that
  arrives is your Worker's egress, real crawlers will be rejected as spoofed
  and served a `204`.
</Warning>

## How to handle the response

| Status | Meaning | What your Worker does |
| - | - | - |
| `200` | Prerendered HTML (cache hit, or a fresh render). | Return the body to the bot. |
| `404` | Soft-404 snapshot — the route renders "not found" content. | Return the body **with the 404 status** so crawlers see the truth (never swallow it into a 200). |
| `301`–`308` | A [routing rule](/dashboard/routing-rules) you configured (redirect). | Return the redirect (status + `Location`) to the bot verbatim. |
| `204` | Passthrough — not a verified bot, a blocked or proxy-ruled path, over quota, or a render that wasn't ready in time. | Serve your app: `fetch(request)`. If the response carries `x-hadoseo-robots` (a blocked path's noindex/nofollow), copy it onto your response as `X-Robots-Tag`. |
| `400` | Missing/invalid `url`. | Fall through; check the URL you send. |
| `401` | Missing/invalid API key. | Fall through; fix your key/secret. |
| `403` | The URL's host isn't owned by your key's account. | Fall through; add the domain in your dashboard. |
| `429` | Rate limit exceeded. | Fall through; retry later or upgrade your plan. |

The service **fails open**: on any internal error or timeout it returns `204`
rather than a `5xx`, so your site never breaks. Your Worker should mirror that —
fall through to `fetch(request)` on anything that isn't a `200`, a `404`
snapshot, or a redirect.

<Note>
  **Routing rules & blocked paths.** Redirect rules from your dashboard are
  returned as real `30x` responses so bots see them. Proxy rules and blocked
  paths come back as `204` — your own edge serves those paths (Hado never
  prerenders them), and blocked-path robots directives arrive in
  `x-hadoseo-robots` for you to forward. Rule changes take effect within about
  a minute (the service caches your domain config briefly).
</Note>

### Redirects for human visitors

The Worker only routes **bot** traffic through Hado SEO, so your dashboard's
redirect rules apply to crawlers — which is what consolidates link equity and
keeps search engines pointed at the right URLs. Human visitors go straight to
your app, untouched, and that's deliberate: it keeps the Worker from adding
any latency to real users.

For humans, configure the same redirects where they're free and instant on
your own zone: Cloudflare **Redirect Rules** (or **Bulk Redirects** for long
lists) under *Rules* in your Cloudflare dashboard, or in your app's own
routing. When you move a page, set the redirect in both places — your zone
rule covers visitors, your Hado rule guarantees crawlers see the `301` even
before caches turn over.

## Analytics-only (passthrough mode)

Only want **crawl and AI-agent analytics**, without serving prerendered HTML? Send
the `x-hadoseo-mode: passthrough` header on every request. Hado SEO records who
crawled what (feeding your [crawl analytics](/dashboard/analytics)) and returns
`204`, so your app is always served as-is.

```js theme={null}
headers: {
  authorization: `Bearer ${env.HADO_API_KEY}`,
  "user-agent": ua,
  "x-hadoseo-client-ip": request.headers.get("cf-connecting-ip") || "",
  "x-hadoseo-mode": "passthrough",
},
```

Because passthrough never renders, it uses no render quota — only your per-minute
rate limit applies.

Passthrough mode still populates the **snapshot timeline**: when a bot crawls a
page, Hado fetches that page from your origin (at most once per URL per day)
and archives what it served. These appear in the page drawer with an **origin**
label — "what your origin served near crawl time" — alongside your crawl
analytics, with no extra setup and no load beyond one background fetch per
page per day.

## Caching repeat bot hits (optional)

Every `200` comes back with `Cache-Control: public, s-maxage=… , stale-while-revalidate=…`.
If you serve the Worker response from Cloudflare's cache, repeat bot hits for the
same URL are served from your edge and never reach Hado SEO. The simplest approach
is to store the render in the Worker's `caches.default`, keyed by the request URL,
and revalidate in the background with `ctx.waitUntil`.

<Card title="How Hado SEO Works" icon="gears" href="/how-it-works">
  Understand caching, bot verification, and the render pipeline behind the endpoint.
</Card>
