AI Gateway Updates & Release Notes

Follow

25 updates curated from 1 source by the Releasebot Team. Last updated: Oct 2, 2026

Get this feed:
  • Oct 2, 2026
    • Date parsed from source:
      Oct 2, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway, Web Search API - Introducing Web Search API

    AI Gateway adds Web Search API in beta, letting AI agents and apps search the web with live information through providers like Ceramic.ai, Exa, and Linkup, with logging, billing through AI Gateway credits, and support for BYO provider keys.

    Web Search API

    Web Search API is now available in beta. Web Search API lets your AI agents and applications search the Internet and ground their responses in live information, instead of guessing URLs or relying on a model's training cutoff.

    At launch, you can choose between three search providers: Ceramic.ai, Exa, and Linkup. All three support Zero Data Retention for requests made through Cloudflare, and all have committed to Cloudflare's verified bot crawling standards.

    Web Search API runs through AI Gateway, so search requests appear in your gateway logs and are billed to your AI Gateway credits at each provider's list API price, with no additional markup. You can also bring your own provider API key.

    Call Web Search API with the REST API:

    curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/ \
    --request POST \
    --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
    --header "Content-Type: application/json" \
    --data '{
    "query": "What are some fun things to do in Salt Lake City as fall approaches?",
    "provider": "ceramic",
    "limit": 5,
    "options": { "gateway": { "id": "default" } }
    }'
    

    Or from a Worker with the AI binding:

    const response = await env.AI.websearch({
    gatewayId: "default",
    query: "What are some fun things to do in Salt Lake City as fall approaches?",
    provider: "exa",
    limit: 5,
    });
    const results = await response.json();
    

    To get started, refer to How to use Web Search API.

    Original source
  • Sep 30, 2026
    • Date parsed from source:
      Sep 30, 2026
    • First seen by Releasebot:
      Oct 1, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Pay for AI inference with Machine Payments

    AI Gateway adds Machine Payments in beta, letting clients use the x402 protocol to pay for eligible inference requests from a stablecoin wallet instead of a prepaid credit balance. It supports select open models on /ai/run and requires a Cloudflare API token plus x402 payment header.

    AI Gateway now supports Machine Payments in beta. With Machine Payments, clients can use the x402 protocol to pay for eligible inference requests directly from a stablecoin wallet instead of maintaining a prepaid credit balance.

    Machine Payments is available for the /ai/run endpoint with select open models. To request x402 payment, authenticate with a Cloudflare API token and include the Cloudflare-specific Payment-Method: x402 header:

    curl -iX POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run" \
    --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
    --header "Payment-Method: x402" \
    --header "Content-Type: application/json" \
    --data '{
    "model": "z-ai/glm-4.7-flash",
    "input": {
    "messages": [
    {
    "role": "user",
    "content": "What is Cloudflare?"
    }
    ]
    }
    }'
    

    An x402-compatible client handles the payment challenge, signs an authorization from the client's wallet, and retries the request. Machine Payments currently requires customers to be based in the United States and have a credit card on file.

    For prerequisites, eligible models, and transaction details, refer to Machine Payments (x402).

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Cloudflare and hundreds of other software products.

    Create account
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Identify model overuse and potential savings with User Insights

    AI Gateway adds User Insights to give customers richer context on AI traffic, grouping conversations by task, tracking turns, and comparing model fit with cost and latency. It also highlights potential savings with faster or less expensive models at no extra cost.

    AI Gateway User Insights now gives you more context about the traffic flowing through your gateway. It shows what users and agents are doing with AI, and where a selected model may be more capable than a task requires.

    On the analysis side, User Insights groups conversations by task, tracks conversation turns, and helps you compare model fit with cost and latency.

    The Potential Savings view highlights requests that may work with faster or less expensive models without compromising output quality. These are the same signals that Cloudflare's Auto Router uses to select a model based on task and cost.

    These new insights are available to all AI Gateway customers at no additional cost. For more information, refer to User Insights.

    Original source
  • Sep 14, 2026
    • Date parsed from source:
      Sep 14, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Prevent Unified Billing fallback for BYOK third-party providers

    AI Gateway adds a setting to require provider credentials for third-party requests, preventing fallback to Unified Billing with Cloudflare-managed credentials. It also supports request-level controls for BYOK enforcement while leaving Workers AI billing unchanged.

    AI Gateway can now require credentials for third-party provider requests.

    Credentials must accompany the request or be stored on the gateway. This setting prevents fallback to Unified Billing with Cloudflare-managed credentials.

    Turn on Require provider credentials in your gateway settings. To use the API, set byok_only to true in the request body of a PUT request to update the gateway:

    {
      "byok_only": true
    }
    

    To require provider credentials for one third-party request, set the cf-aig-no-wholesale header to true. This header cannot relax the gateway setting.

    Requests without applicable credentials then return an HTTP 400 response. Workers AI requests remain allowed, and the setting does not change their configured billing mode.

    For configuration details and request-level controls, refer to Prevent Unified Billing fallback for BYOK third-party providers.

    Original source
  • Sep 9, 2026
    • Date parsed from source:
      Sep 9, 2026
    • First seen by Releasebot:
      Sep 10, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - AI Gateway custom costs support cache tokens

    AI Gateway adds cache-read and cache-write token rates to custom costs, improving billing accuracy for negotiated cache pricing across providers and avoiding double-counting.

    AI Gateway custom costs now support cache-read and cache-write token rates. This lets custom cost metrics reflect negotiated cache pricing across providers.

    Add per_cache_read_token or per_cache_write_token to the cf-aig-custom-cost header:

    {
    "per_token_in": 0.000001,
    "per_token_out": 0.000002,
    "per_cache_read_token": 0.0000001,
    "per_cache_write_token": 0.0000005
    }
    

    Cache-token pricing activates when either cache rate is present. An omitted cache rate defaults to per_token_in. If both cache rates are omitted, AI Gateway preserves the existing input and output calculation.

    Providers can include cache tokens within input tokens or report them separately. AI Gateway automatically accounts for these differences and prevents double-counting.

    For more information, refer to Custom costs.

    Original source
  • Similar to AI Gateway with recent updates:

  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - AI Gateway consolidates monthly usage invoice line items and standardizes model names

    AI Gateway updates monthly usage invoices with a single total cost per model and standardizes model names across invoices and logs for clearer billing. Input and output token line items are no longer shown on usage invoices, while credit purchase invoices are unchanged.

    AI Gateway monthly usage invoices

    AI Gateway monthly usage invoices, issued at the beginning of each month for the previous month's usage, now show a single total cost for each model. These invoices no longer break out input and output token quantities and unit prices into separate line items. This change does not apply to invoices for AI Gateway credit purchases.

    For example, an invoice that previously included these separate line items:

    anthropic claude-haiku-4-5-20251001 Input Tokens: 40,000 tokens at $0.000001 ($0.04)

    anthropic claude-haiku-4-5-20251001 Output Tokens: 24,000 tokens at $0.000005 ($0.12)

    The updated invoice includes one line item: anthropic/claude-haiku-4.5: $0.16.

    AI Gateway has also standardized model names across invoices and logs. Model variants that previously appeared with provider-specific version suffixes now use a consistent provider/model identifier.

    For more information, refer to the Unified Billing documentation and AI Gateway logging documentation.

    Original source
  • Aug 19, 2026
    • Date parsed from source:
      Aug 19, 2026
    • First seen by Releasebot:
      Aug 20, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Get 50% off GPT-5.6 Sol through AI Gateway

    AI Gateway adds GPT-5.6 Sol support with a limited-time 50% discount for Unified Billing users. Customers can send requests to openai/gpt-5.6-sol with automatic promotional pricing through September 18, 2026, with no promo code required.

    GPT-5.6 Sol is available through AI Gateway, and for a limited time you can use it at 50% off. If you are already using AI Gateway, point to the openai/gpt-5.6-sol model and the discounted pricing applies automatically — no promo code needed.

    The promotion is available for Unified Billing users only (not Bring Your Own Keys). Load credits onto AI Gateway and start sending requests to openai/gpt-5.6-sol.

    Discounted pricing during the promotion:

    Usage

    Promotional price

    Standard price

    Input

    $2.50 per 1M tokens

    $5 per 1M tokens

    Output

    $15 per 1M tokens

    $30 per 1M tokens

    Cache read

    $0.25 per 1M tokens

    $0.50 per 1M tokens

    The promotion runs through September 18, 2026. After that date, GPT-5.6 Sol requests return to standard pricing.

    For more details, refer to the Unified Billing documentation and the GPT-5.6 Sol model page.

    Original source
  • Aug 7, 2026
    • Date parsed from source:
      Aug 7, 2026
    • First seen by Releasebot:
      Aug 7, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway, Workers AI - Workers AI and AI Gateway unify model access and billing

    AI Gateway now provides a unified path for Workers AI and supported third-party models, with shared AI bindings and REST APIs plus observability, logging, caching, security, and billing controls. It also adds unified billing with prepaid credits and higher rate limits for frontier models.

    Unified entrypoints and observability

    Workers AI and AI Gateway now provide a unified path for accessing models and managing inference traffic. Use the same AI binding and REST API to call models hosted on Workers AI or by supported third-party providers, with AI Gateway providing observability, logging, caching, security, and billing controls.

    The AI binding supports both Workers AI and third-party models through env.AI.run(). The REST API provides shared /ai/ endpoints with Cloudflare authentication across providers.

    Route a Workers AI request through AI Gateway by specifying a gateway ID. Use default to automatically create a gateway on the first authenticated request, or specify an existing gateway to separate applications and workloads:

    const response = await env.AI.run(
    "@cf/zai-org/glm-5.2",
    {
    messages: [{ role: "user", content: "What is the capital of France?" }],
    },
    {
    gateway: { id: "default" },
    },
    );
    const response = await env.AI.run(
    "@cf/zai-org/glm-5.2",
    {
    messages: [{ role: "user", content: "What is the capital of France?" }],
    },
    {
    gateway: { id: "default" },
    },
    );
    

    Requests routed through AI Gateway can be logged and included in analytics for request volume, errors, latency, token usage, and costs. You can also configure controls such as caching, rate limiting, and request retries on the gateway.

    Unified billing and higher rate limits

    You can now use prepaid AI Gateway credits to pay for Workers AI inference. This provides one credit balance for Workers AI and supported third-party model providers. To use credits for Workers AI, set the gateway's Workers AI billing setting to Unified billing. Workers AI requests routed through that gateway deduct from your credit balance in real time.

    Prepaid credits also provide access to the following Workers AI frontier models without requiring the Workers Paid plan. Each frontier Workers AI model has a rate limit of 50 requests per minute per account, per model when billed with AI Gateway credits, compared to 20 requests per minute through standard Workers AI billing:

    • @cf/moonshotai/kimi-k2.6
    • @cf/moonshotai/kimi-k2.7-code
    • @cf/zai-org/glm-5.2

    These limits are designed for typical agentic and coding workloads, where requests to frontier models can take longer to complete.

    For details, refer to Workers AI limits, Workers AI pricing, Unified Billing, and the AI Gateway model catalog.

    Original source
  • Aug 5, 2026
    • Date parsed from source:
      Aug 5, 2026
    • First seen by Releasebot:
      Aug 6, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Track AI spend and catch anomalous usage with User Insights

    AI Gateway adds User Insights, giving teams a single dashboard for AI spend visibility and abnormal usage detection. It shows organization-wide cost, requests, tokens and adoption, plus user-level drilldowns, with no extra setup and no additional cost.

    AI Gateway now includes User Insights

    AI Gateway now includes User Insights, a dashboard that gives you two things at once: clear visibility into how much your organization spends on AI, and a security signal that surfaces users whose usage suddenly looks abnormal. It works on the traffic already flowing through your gateway, so there is no additional setup.

    On the spend side, User Insights shows organization-wide totals for cost, requests, tokens, and adoption, and lets you drill into an individual user to see their spend, top models and providers, cache hit rate, and more. To attribute usage to individual users, add a user identifier with custom metadata or put your gateway behind Cloudflare Access.

    On the security side, User Insights baselines each user's normal usage from their 95th percentile (p95) session cost over the last 30 days, then flags sessions that exceed both that baseline and an organization-level threshold. A sudden jump above a user's own pattern is often the first sign of a compromised credential or a misbehaving agent, so you can investigate before it shows up on your bill.

    User Insights is available to all AI Gateway customers at no additional cost.

    Original source
  • Aug 5, 2026
    • Date parsed from source:
      Aug 5, 2026
    • First seen by Releasebot:
      Aug 5, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway, Access - Identity-aware controls are now available in AI Gateway

    AI Gateway now integrates with Cloudflare Access, adding endpoint protection and identity-aware controls for authenticated users. It can use Access identity in logs, analytics, routing, and spend controls, with verified user IDs added to request metadata as cf.user_id.

    AI Gateway now integrates with Cloudflare Access, giving you two new capabilities:

    • Protect your gateway endpoint. Put your AI Gateway behind Access so you can set policies that control who is allowed to call a specific gateway's endpoint.
    • Identity-aware controls. When traffic reaches AI Gateway through an Access-protected custom domain, AI Gateway can use the authenticated user's Access identity in logs, analytics, routing, and spend controls.

    With identity-aware controls, you can set spend limits by authenticated user, control which gateways different users can access, filter logs by user, and build policies without passing user IDs from the client application. AI Gateway adds the verified Access user ID to request metadata as cf.user_id.

    For setup instructions, refer to Cloudflare Access.

    Original source
  • Jun 12, 2026
    • Date parsed from source:
      Jun 12, 2026
    • First seen by Releasebot:
      Jul 18, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - View the user agent of requests in AI Gateway logs

    AI Gateway adds user agent logging and filtering to help identify which SDK, library, or app sent each request.

    AI Gateway logs now capture the user agent of the client that made each request, making it easier to identify which SDK, library, or application sent the traffic flowing through your gateway. For example, you can tell apart requests coming from openai-python versus a custom application or a Cloudflare Worker.

    The user agent appears alongside the other details in each log entry, and you can filter logs by user agent (equals, does not equal, or contains) in the dashboard.

    For more information, refer to Logging.

    Original source
  • Jun 5, 2026
    • Date parsed from source:
      Jun 5, 2026
    • First seen by Releasebot:
      Jul 18, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Control AI costs with spend limits

    AI Gateway adds spend limits that track dollar-based budget usage and block requests once the cap is reached. The new controls can be scoped by model, provider, or custom metadata, with fixed or sliding time windows for flexible cost control across Unified Billing and BYOK requests.

    AI Gateway now supports spend limits — cost-based budgets that track cumulative dollar spend and block requests when the budget is exceeded. Unlike rate limiting, which caps the number of requests, spend limits track actual cost based on token usage and model pricing.

    You can scope limits by model, provider, or custom metadata dimensions. For example, give each user a $200/day budget, cap total gateway spend at $10,000/day, or limit a specific model to $50/day per user. Each rule uses a configurable time window with fixed or sliding enforcement.

    Spend limits work with both Unified Billing and BYOK requests for models with known pricing.

    For more details, refer to the Spend limits documentation.

    Original source
  • May 21, 2026
    • Date parsed from source:
      May 21, 2026
    • First seen by Releasebot:
      Jul 18, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Call any AI model through AI Gateway's new REST API

    AI Gateway now supports the AI REST API on api.cloudflare.com, giving users one unified way to call models from OpenAI, Anthropic, Google, or Workers AI. It adds multiple compatible endpoints, automatic logging, caching, rate limiting, guardrails, and Unified Billing for third-party models.

    AI Gateway now uses the AI REST API on api.cloudflare.com.

    You can call any model — whether from OpenAI, Anthropic, Google, or hosted on Workers AI — through one unified API, using the same endpoints and authentication regardless of provider. Four endpoints are available:

    • POST /ai/run — universal endpoint for all models and modalities
    • POST /ai/v1/chat/completions — OpenAI SDK compatible
    • POST /ai/v1/responses — OpenAI Responses API compatible
    • POST /ai/v1/messages — Anthropic SDK compatible
    curl -X POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/v1/chat/completions" \
    --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
    --header "Content-Type: application/json" \
    --data '{
    "model": "openai/gpt-5.5",
    "messages": [{"role": "user", "content": "What is Cloudflare?"}]
    }'
    

    All AI Gateway features — logging, caching, rate limiting, and guardrails — are applied automatically. Third-party models are billed through Unified Billing, so you do not need to manage separate provider API keys.

    Third-party model requests are routed through your account's default gateway, which is created automatically on first use. To route requests through a specific gateway, add the cf-aig-gateway-id header.

    If you are already calling Workers AI models through the existing REST API, that path (/ai/run/@cf/{model}) continues to work. To call Workers AI models through AI Gateway, use the @cf/ model prefix (for example, @cf/moonshotai/kimi-k2.6) and include the cf-aig-gateway-id header to specify which gateway to route through.

    For more details and examples, refer to the REST API documentation.

    Original source
  • Apr 2, 2026
    • Date parsed from source:
      Apr 2, 2026
    • First seen by Releasebot:
      Jul 18, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Automatically retry on upstream provider failures on AI Gateway

    AI Gateway now supports automatic gateway-level retries, helping requests recover from upstream errors without client-side changes. It adds configurable retry counts, delays, and backoff strategies, with per-request headers able to override the defaults.

    AI Gateway now supports automatic retries at the gateway level. When an upstream provider returns an error, your gateway retries the request based on the retry policy you configure, without requiring any client-side changes.

    You can configure the retry count (up to 5 attempts), the delay between retries (from 100ms to 5 seconds), and the backoff strategy (Constant, Linear, or Exponential). These defaults apply to all requests through the gateway, and per-request headers can override them.

    This is particularly useful when you do not control the client making the request and cannot implement retry logic on the caller side. For more complex failover scenarios — such as failing across different providers — use Dynamic Routing.

    For more information, refer to Manage gateways.

    Original source
  • Mar 17, 2026
    • Date parsed from source:
      Mar 17, 2026
    • First seen by Releasebot:
      Jul 18, 2026
    Cloudflare logo

    AI Gateway by Cloudflare

    AI Gateway - Log AI Gateway request metadata without storing payloads

    AI Gateway adds support for the cf-aig-collect-log-payload header, giving users control over whether request and response bodies are stored in logs while preserving usage metadata like token counts, model, provider, cost, status code, and duration.

    AI Gateway now supports the cf-aig-collect-log-payload header, which controls whether request and response bodies are stored in logs. By default, this header is set to true and payloads are stored alongside metadata. Set this header to false to skip payload storage while still logging metadata such as token counts, model, provider, status code, cost, and duration.

    This is useful when you need usage metrics but do not want to persist sensitive prompt or response data.

    curl https://gateway.ai.cloudflare.com/v1/$ACCOUNT_ID/$GATEWAY_ID/openai/chat/completions \
    --header "Authorization: Bearer $TOKEN" \
    --header 'Content-Type: application/json' \
    --header 'cf-aig-collect-log-payload: false' \
    --data '{
    "model": "gpt-4o-mini",
    "messages": [
    {
    "role": "user",
    "content": "What is the email address and phone number of user123?"
    }
    ]
    }'
    

    For more information, refer to Logging.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official product update announcements from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.