Cloudflare AI Updates & Release Notes

Follow

169 updates curated from 1 source by the Releasebot Team. Last updated: Oct 2, 2026

Get this feed:
  • Oct 2, 2026
    • Date parsed from source:
      Oct 2, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Introducing Web Search API

    Cloudflare AI launches Web Search API in beta, letting AI agents and apps search the live internet and ground responses in current information. It supports Ceramic.ai, Exa, and Linkup, runs through AI Gateway, and lets users bring their own provider key.

    Web Search API is now available in beta. Web Search API lets your AI agents and applications search the Internet and ground their responses in live information, instead of guessing URLs or relying on a model's training cutoff.

    At launch, you can choose between three search providers: Ceramic.ai, Exa, and Linkup. All three support Zero Data Retention for requests made through Cloudflare, and all have committed to Cloudflare's verified bot crawling standards.

    Web Search API runs through AI Gateway, so search requests appear in your gateway logs and are billed to your AI Gateway credits at each provider's list API price, with no additional markup. You can also bring your own provider API key.

    Call Web Search API with the REST API:

    curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/ \
      --request POST \
      --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
      --header "Content-Type: application/json" \
      --data '{
        "query": "What are some fun things to do in Salt Lake City as fall approaches?",
        "provider": "ceramic",
        "limit": 5,
        "options": { "gateway": { "id": "default" } }
      }'
    

    Or from a Worker with the AI binding:

    const response = await env.AI.websearch({
      gatewayId: "default",
      query: "What are some fun things to do in Salt Lake City as fall approaches?",
      provider: "exa",
      limit: 5,
    });
    const results = await response.json();
    

    To get started, refer to How to use Web Search API.

    Original source
  • Oct 2, 2026
    • Date parsed from source:
      Oct 2, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Run the Pi Durable harness on Cloudflare with the Agents SDK

    Cloudflare AI adds first-class Pi harness support in the Agents SDK, letting developers build long-running agents with Pi 1.0, Pi Durable, and PiHarness for durable, persisted work that can survive interruptions, restarts, and network issues.

    The Agents SDK now provides first-class support for building agents using the Pi harness.

    You can build long-running agents using the combination of Pi 1.0, Pi Durable, and the new PiHarness class that the Cloudflare Agents SDK provides, ensuring your agent's work is durably persisted, even if interrupted mid-turn.

    Built with Earendil, this integration is our first step toward first-class support for third-party agent harnesses on Cloudflare.

    Beta

    PiHarness is in beta. Pi Durable is a new, experimental package, and the PiHarness API will likely change as Pi Durable matures.

    PiHarness is a new "Lifecycle capability" provided by the Cloudflare Agents SDK. Pi Durable provides the agent harness and the Lifecycle is responsible for keeping the agent running in the Durable Object. The Lifecycle is a core concept in the Agents SDK ensuring that long-running work can run in a Durable Object, surviving restarts, crashes, and network issues. We will share more on Lifecycle capabilities in the near future.

    Install

    npm i agents@latest @earendil-works/pi-durable @earendil-works/pi-ai
    

    Both Pi packages are optional peer dependencies of agents, so you only install them if you use the harness.

    Use it in an Agent

    Creating a Pi agent requires configuring the Pi Harness with a model, skills, and tools, then registering the PiHarness with the Agent class.

    Add tools with extensions

    Both tools and system prompt sections are provided to the Pi Harness via extensions.

    For more information on creating and configuring extensions, refer to Extensions.

    Learn more

    • Pi harness documentation
    • Pi harness extensions
    • pi-ai model provider
    • Pi harness example, with WebSockets, a browser UI, and a @cloudflare/computer Workspace for the model's tools
    • Lifecycle
    • Pi Durable announcement from Earendil
    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Cloudflare and hundreds of other software products.

    Create account
  • Oct 1, 2026
    • Date parsed from source:
      Oct 1, 2026
    • First seen by Releasebot:
      Oct 1, 2026
    • Modified by Releasebot:
      Oct 2, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Introducing Clef: Cloudflare's first open-source decision models, now on Workers AI

    Cloudflare AI launches Clef and Clef-flash, its first Workers AI decision models, bringing fast structured decisions, vision support, and drop-in compatibility with Jev. It also adds an RL fine-tuning service and open-sources the model weights on Hugging Face.

    Meet @cf/cloudflare/clef and @cf/cloudflare/clef-flash, the first models trained by the Cloudflare Workers AI team, available on Workers AI today.

    Clef is a decision model, in the same family as Typesafe's Jev. Instead of generating text, it reads an input state and a set of typed questions, then returns a probability for every allowed answer. Your agent gets a structured decision it can act on immediately, for example: route the ticket, block the request, or escalate to a human. There is no free-form output to parse and no reasoning tokens to wait for.

    Both models are hosted on Workers AI as Clef and Clef-flash. We are also open-sourcing the weights under the Apache 2.0 license on Hugging Face: Clef and Clef-flash. Read the launch blog post for the full story, including how we trained them.

    We are also launching a reinforcement learning (RL) fine-tuning service to help you tune Clef for your own workloads. Sign up to work with us as a design partner.

    Built for the hot path

    Clef is designed to be fast so decisions come back in milliseconds. Across our 43 benchmark runs, we achieved speeds where Clef is 2.5x faster than Jev at the median, and Clef-flash 13x faster.

    Hosting on Workers AI adds to that speed. Requests run on GPUs across Cloudflare's network, running close to your users, so the network round trip stays short. You can put Clef directly in the request path of your agent, then hand off to an LLM on Workers AI to take action.

    Leading the benchmarks

    Across 10 decision benchmarks, a Clef model scores highest on 7, ahead of Jev and other open decision models. A few highlights:

    • BFCL (case exact): Clef-flash 98.76
    • BANKING77 (macro-F1): Clef 94.20
    • CLINC150+OOS (macro-F1): Clef 97.43
    • Home appliances (case exact): Clef-flash 97.73

    On Typesafe's own workflow evals, Clef beats Jev in 3 of 4 areas: invoice processing, customer service, and security incidents. The full results are on the Hugging Face model card.

    Drop-in compatible with Jev

    Clef follows the System One API, so you can switch an existing Jev integration to Clef by changing the endpoint and model. Ask up to 64 questions per request, in three types:

    • noul: A yes/no question. Returns the probability that the answer is yes.
    • choice: Pick one option from a set you define. Returns the chosen option, a probability per option, and a confidence value.
    • score: Rate against an ordered rubric. Returns a probability-weighted score and a probability per level.

    Example JavaScript usage provided.

    What you can build with decision models

    • Support triage: Decide whether a ticket is urgent and which team owns it, then route it without a human in the loop.
    • Threat intelligence: Classify a website by category. Paired with Browser Run, Clef fetched, rendered, and classified a domain in 2.2 seconds, compared to 4.7 seconds for gpt-oss-120b in the same workflow.
    • Trust and safety: Score user submissions against your own policy rubric and act on the probability.
    • Agent guardrails: Let an agent check "should I take this action?" in tens of milliseconds before calling a tool.
    • Visual classification: Pass up to four images alongside the state. Unlike text-only decision models, Clef has a vision encoder.

    Get started

    Use Clef through the Workers AI binding (env.AI.run()) or the REST API at /ai/run. You can also use AI Gateway with these endpoints.

    For more information, refer to the Clef model page, the Clef-flash model page, and pricing.

    Original source
  • Oct 1, 2026
    • Date parsed from source:
      Oct 1, 2026
    • First seen by Releasebot:
      Oct 1, 2026
    • Modified by Releasebot:
      Oct 2, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    The best way to do MCP auth just got better: Workers OAuth Provider goes v1, with a new split API and full support for MCP 2026-07-28

    Cloudflare AI releases workers-oauth-provider v1 with a split OAuth API for MCP servers, letting authorization and resource servers run in separate Workers. It adds MCP spec support, step-up authorization, sliding refresh expiry, resumable KV cleanup, and migration help.

    @cloudflare/workers-oauth-provider is now v1, with a new split API. One Worker acts as the authorization server: it signs users in and issues tokens. Your MCP server acts as the resource server, and can run in another Worker. It validates each token with the authorization server over a Service Binding, without crossing the public Internet.

    • It supports the MCP 2026-07-28 authorization specification, including Client ID Metadata Documents and issuer identification. It still works with older clients, including those using Dynamic Client Registration.
    • insufficientScope() gives you step-up authorization in one line.
    • The authorization server and the resource server can run in different Workers with different WAF and rate limiting rules.
    • One authorization server can issue tokens for many MCP servers.
    • A migration skill ships in the npm package so a coding agent can perform the upgrade.

    The split API

    Example JavaScript code provided for the split API usage.

    Other updates and helpers

    • Consent page and upstream sign-in helpers implement the MCP confused deputy protections.
    • Sliding refresh token expiry with refreshTokenIdleTTL.
    • Resumable KV cleanup with purgeExpiredData().
    • An internal reason on every error passed to onError.

    Upgrade with the migration skill

    Instructions and code snippets provided.

    OAuthResourceServer publishes the RFC 9728 protected resource metadata that MCP clients use to find your authorization server. It answers requests without a token with a 401 challenge that points to that metadata. It also rejects tokens issued for any other resource.

    You can still use OAuthProvider as both the authorization server and the MCP server. For most 0.x deployments, the only required change is to add resourceMetadata: { resource }.

    Example wrangler.jsonc configuration snippet provided.

    Original source
  • Oct 1, 2026
    • Date parsed from source:
      Oct 1, 2026
    • First seen by Releasebot:
      Oct 1, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    AI Search is generally available

    Cloudflare AI releases AI Search generally available with usage-based pricing, default hybrid search, included Workers AI embeddings and reranking, multimodal image support, OCR for scanned PDFs, larger file limits, and optional source type inference for new instances.

    AI Search is now generally available

    Usage-based billing begins on November 1, 2026, with included monthly ingestion, storage, semantic query, and full-text query usage. Cloudflare will send a reminder email the week before billing begins.

    Refer to Limits & pricing for rates and included usage.

    Hybrid search is on by default

    New AI Search instances use hybrid search by default. Hybrid search combines semantic vector retrieval with full-text matching. You can choose a different index method when you create an instance.

    Refer to Hybrid search for details.

    Workers AI embeddings and reranking are included

    Workers AI embedding and reranking calls made by AI Search are included in AI Search pricing. These calls no longer appear on your Workers AI bill or in your AI Gateway logs. Generation, query rewriting, and external providers continue to use your account and gateway.

    Refer to Limits & pricing for details.

    Multimodal model and image support

    AI Search supports the @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2 multimodal embedding models. Search and chat requests can include images through the REST API and public endpoint.

    Refer to Supported models for the full list of embedding models.

    OCR availability and increased file limits

    Optical character recognition (OCR) is available on every account for scanned PDFs. Plain-text or code files and PDFs with OCR enabled can be up to 10 MiB. PDFs without OCR and other supported formats remain limited to 4 MiB.

    Refer to Data source for file limits and Limits & pricing for OCR pricing.

    Source type inference

    When you create an AI Search instance, the type field is optional. AI Search infers a website source from an HTTP or HTTPS URL, or an R2 source from an existing bucket name.

    Refer to Data source for details.

    Original source
  • Similar to Cloudflare AI with recent updates:

  • Sep 30, 2026
    • Date parsed from source:
      Sep 30, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    • Modified by Releasebot:
      Oct 1, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Sandbox SDK 1.0: control every sandbox from your own Durable Object

    Cloudflare AI releases Sandbox SDK 1.0, giving Durable Objects direct control over sandbox containers with custom images, snapshots, terminals, streamed commands, preview hosting, outbound request handling, and agent workflows, while keeping Sandbox SDK 0.x supported for migrations.

    Sandbox SDK 1.0 is available

    Your own Durable Object class now controls each sandbox container directly, through the Durable Object container API on this.ctx.container.

    With 1.0, your class can:

    • Choose the image and instance size each time it starts a sandbox. One class can run sandboxes on different images, and a deploy does not restart sandboxes that are running.
    • Save the files of a sandbox as a snapshot , in public beta, and start the same sandbox or a new one from it.
    • Decide when each sandbox stops, for example when a task finishes, when its user goes idle, or after it saves a snapshot.
    • Run commands with streamed input and output, send them signals, and open terminals .
    • Serve previews from ports in the sandbox, with your own hostnames and authentication.
    • Handle outbound requests for each hostname in Worker code, so credentials and bindings stay in your Worker.
    • Expose only the methods that you want callers to use.
    • Run an agent in the same Durable Object, and give the model a tool that runs commands in the sandbox.

    If you use Sandbox SDK 0.x

    Your 0.x applications keep running, and @cloudflare/sandbox 0.x stays on npm. Sandbox SDK 0.x receives bug and security fixes until 2026-12-31, and its documentation stays at Sandbox SDK 0.x .

    When you are ready, Migrate from Sandbox SDK 0.x shows the 1.0 code for each 0.x feature, including preview URLs, tunnels, background processes, terminals, backups, and the code interpreter. The guide keeps the preview URLs and named tunnels that your 0.x application created working. You can move every sandbox in one deploy, or run a 1.0 class next to your 0.x class and move sandboxes one at a time. The deploy that moves an existing class to the new policy is one-way, so the guide shows how to rehearse it first.

    The Sandboxes documentation also covers Dynamic Workers, for untrusted code in JavaScript, Python, or WebAssembly.

    Original source
  • Sep 30, 2026
    • Date parsed from source:
      Sep 30, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Pay for AI inference with Machine Payments

    Cloudflare AI adds Machine Payments in beta for AI Gateway, letting eligible /ai/run inference requests use x402 to pay from a stablecoin wallet instead of a prepaid credit balance. It supports select open models and requires U.S.-based customers with a credit card on file.

    AI Gateway now supports Machine Payments in beta. With Machine Payments, clients can use the x402 protocol to pay for eligible inference requests directly from a stablecoin wallet instead of maintaining a prepaid credit balance.

    Machine Payments is available for the /ai/run endpoint with select open models. To request x402 payment, authenticate with a Cloudflare API token and include the Cloudflare-specific Payment-Method: x402 header:

    curl -iX POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run" \
    --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
    --header "Payment-Method: x402" \
    --header "Content-Type: application/json" \
    --data '{
      "model": "z-ai/glm-4.7-flash",
      "input": {
        "messages": [
          {
            "role": "user",
            "content": "What is Cloudflare?"
          }
        ]
      }
    }'
    

    An x402-compatible client handles the payment challenge, signs an authorization from the client's wallet, and retries the request. Machine Payments currently requires customers to be based in the United States and have a credit card on file.

    For prerequisites, eligible models, and transaction details, refer to Machine Payments (x402).

    Original source
  • Sep 30, 2026
    • Date parsed from source:
      Sep 30, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Monetization Gateway closed beta

    Cloudflare AI adds Monetization Gateway in closed beta, letting sellers charge agents for access to APIs, MCP tools, sites, and datasets while using x402 to handle payment authorization in the HTTP request flow.

    Monetization Gateway is now available in closed beta. Sellers can use it to charge agents for access to APIs, Model Context Protocol (MCP) tools, sites, and datasets.

    Sellers (domain owners) define which requests require payment, the cost, and where the payment should be sent. Buyers receive the payment instructions, sign an authorization, and receive the resource after the payment has been settled. The Monetization Gateway uses the x402 protocol to handle payment authorization within the HTTP request flow.

    To learn more, request access in the Cloudflare dashboard ↗︎, review the Monetization Gateway documentation, or read the blog ↗︎.

    Original source
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Sep 29, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Identify model overuse and potential savings with User Insights

    Cloudflare AI adds AI Gateway User Insights to help customers understand AI traffic, group conversations by task, track turns, and compare model fit with cost and latency. The new Potential Savings view also spotlights requests that may use faster or cheaper models at no extra cost.

    AI Gateway User Insights now gives you more context about the traffic flowing through your gateway. It shows what users and agents are doing with AI, and where a selected model may be more capable than a task requires.

    On the analysis side, User Insights groups conversations by task, tracks conversation turns, and helps you compare model fit with cost and latency.

    The Potential Savings view highlights requests that may work with faster or less expensive models without compromising output quality. These are the same signals that Cloudflare's Auto Router uses to select a model based on task and cost.

    These new insights are available to all AI Gateway customers at no additional cost. For more information, refer to User Insights.

    Original source
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Sep 29, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Connect multiple clients to one Browser Run session

    Cloudflare AI adds concurrent Browser Run sessions, letting multiple Workers connect to the same browser at once. Each puppeteer.connect() gets its own CDP connection, with separate browser contexts to keep pages, cookies, and storage isolated while reducing cold starts and browser count.

    Browser Run sessions now accept multiple concurrent connections. Before, a session accepted only one connection at a time, and other Workers had to wait until that connection closed. Now multiple Workers can connect to the same browser at the same time.

    Each puppeteer.connect() call opens its own Chrome DevTools Protocol (CDP) connection. Create a separate browser context for each request to keep its pages, cookies, and storage apart from other clients.

    Sharing sessions means fewer new browsers to launch, less cold-start time, and fewer concurrent browsers counted against your limits.

    Concurrent connections require @cloudflare/puppeteer version 1.1.0 or later.

    Refer to Reuse sessions for a full example.

    Original source
  • Sep 28, 2026
    • Date parsed from source:
      Sep 28, 2026
    • First seen by Releasebot:
      Sep 28, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Browser Run adds WebMCP to Kitesurf and moves to document.modelContext

    Cloudflare AI expands WebMCP support to Kitesurf sessions and updates tool access for DevTools, AI agents, and CDP clients.

    WebMCP now works in Kitesurf sessions as well as Lab sessions. Both backends use the document.modelContext API from the WebMCP Community Group draft. Lab sessions no longer expose navigator.modelContextTesting.

    To list and run page tools:

    • Chrome DevTools: Use the Application > WebMCP panel in the live view of a Lab session or in the Kitesurf playground.
    • AI agents: Start Chrome DevTools MCP with the --category-experimental-webmcp flag to add the list_webmcp_tools and execute_webmcp_tool tools.
    • CDP clients: Use the WebMCP CDP domain.
    Original source
  • Sep 25, 2026
    • Date parsed from source:
      Sep 25, 2026
    • First seen by Releasebot:
      Sep 25, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Subscribe to Browser Run crawl events

    Cloudflare AI adds Browser Run crawl lifecycle events to Queues for progress tracking and downstream processing without polling.

    Browser Run crawl jobs can publish lifecycle events to Cloudflare Queues. Subscribe to started, updated, and finished events to track progress or trigger downstream processing without polling.

    To create an account-level subscription, run the following command:

    npx wrangler queues subscription create <QUEUE_NAME> --source browserRun --events crawl.started,crawl.updated,crawl.finished
    

    For payload examples, refer to the Browser Run event schemas.

    Original source
  • Sep 21, 2026
    • Date parsed from source:
      Sep 21, 2026
    • First seen by Releasebot:
      Sep 22, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Browser Run adds session and DevTools methods to browser bindings

    Cloudflare AI adds typed Browser Run bindings for session management and DevTools operations, with new methods for acquiring sessions, connecting clients, Live View URLs, target control, and outboundByHost routing through other Workers.

    Browser Run browser bindings

    now provide typed methods for session management and DevTools operations. You can acquire a session, connect a browser client, create Live View URLs, manage targets, and close sessions without constructing HTTP requests.

    The new acquire() and launch() methods also accept outboundByHost. This lets you route requests for selected hostnames through another Worker, including a Worker that adds authentication or reaches a private service.

    Use connectSession(sessionId) when you need to acquire and connect in separate steps. The method returns a session-pinned webSocket Fetcher for a CDP client.

    The binding also includes session methods for Live View, active sessions, session history, limits, session details, and cleanup. The nested devtools binding provides typed methods for browser version information, protocol descriptions, and target operations such as listing, creating, activating, and closing targets.

    Refer to the Browser binding API documentation for method signatures and the outbound Worker feature guide for routing examples.

    Original source
  • Sep 18, 2026
    • Date parsed from source:
      Sep 18, 2026
    • First seen by Releasebot:
      Sep 19, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Inspect logs, network requests, and DOM in Session Recordings

    Cloudflare AI adds Browser Run Session Recordings with an Inspect panel, bringing Logs, Network, and DOM views to help debug sessions without reproducing them. It also supports HAR downloads, API access to raw network data, and multi-tab inspection from the Cloudflare dashboard.

    Browser Run Session Recordings now include an Inspect panel, giving you more context to understand what happened during a browser session without having to reproduce it.

    The Logs tab lets you search captured console output and filter messages by level. The Network tab shows each request's method, status, headers, payload, response, and timing waterfall, with the option to download the session's network activity as a HAR file.

    You can also retrieve recorded network activity via API as raw JSON or a HAR file for use in your own debugging and analysis workflows.

    The DOM tab provides an expandable view of the page structure at the end of the recording and lets you copy the reconstructed HTML. For sessions with multiple browser tabs, the Inspect panel updates to show data for the tab selected in the recording viewer.

    To get started, enable recording when launching a browser session. After the session closes, open Browser Run > Runs in the Cloudflare dashboard and select the recording icon next to the session.

    Refer to the Session recording documentation for setup instructions and current limits.

    Original source
  • Sep 17, 2026
    • Date parsed from source:
      Sep 17, 2026
    • First seen by Releasebot:
      Sep 17, 2026
    • Modified by Releasebot:
      Sep 22, 2026
    Cloudflare logo

    Cloudflare AI by Cloudflare

    Reject busy synchronous inference requests

    Cloudflare AI adds rejectIfBusy for synchronous Workers AI requests to fail fast when capacity is unavailable.

    The rejectIfBusy option lets synchronous Workers AI inference requests fail when capacity is unavailable. Use it when your application should not wait in a capacity queue.

    Pass the option as the third argument to the Workers AI binding:

    For the native REST API, add the option to the request body:

    Refer to Reject busy requests for OpenAI-compatible usage and error behavior.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official product update announcements from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.