AI Models Release Notes

Release notes for leading AI models, APIs and AI platforms

Get this feed:

Products (16)

Latest AI Models Updates

  • Sep 2, 2026
    • Date parsed from source:
      Sep 2, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Anthropic logo

    Claude by Anthropic

    Building commerce agents with Claude

    Claude launches a commerce agent blueprint with reference shopping and merchant agents, guardrails, live demos, and a Claude Code plugin to help teams build AI-powered shopping experiences fast across retail, travel, telecom, and ticketing.

    What's in the blueprint

    Many of the world’s largest retailers, marketplaces, e-commerce platforms, and travel companies use Claude to build agents that make shopping easier. Enterprise customers like Shopify, Priceline, and others have agents that let consumers use AI to search for what they want in plain language, find it, compare it, and buy it.

    Today, we're launching a blueprint to help build commerce agents on Claude. It contains the harnesses, patterns, and guardrails an engineering team needs to get a commerce agent running in days, with reference implementations of a shopping agent and a merchant agent for retail, travel, telecom, and ticketing platforms. It also includes a Claude Code plugin to get you started.

    The code deploys where you already build with Claude, including the Claude API, Amazon Bedrock, Microsoft Foundry, or Google Cloud Vertex AI. You can also work with our solutions and ecosystem partners such as Accenture, Mastercard, and Visa, who are working with us to enable clients and merchant communities to leverage the blueprints.

    It’s available today, with live demos for each vertical and an engineering deep-dive on how it was built, just in time for holiday season planning.

    The repository contains complete, working implementations of a shopping agent and merchant agent that can be built using the Messages API, Agent SDK, or Claude Managed Agents (beta). You can see them running in a self-guided demo before writing any code, and then work with Claude Code to customize them to your catalogs, policies, brand, and more.

    The shopping agent

    The shopping agent lives inside your app or website. The blueprint includes the integration points for catalog, cart, checkout, customer preferences, and order history, and leaves payment to you, whether that is your existing checkout or an agentic payments provider.

    A customer can say “I need a tent, sleeping bag, and stove for a weekend trip with two kids,” and the agent can take it from there. Here’s what it can do:

    • Search the catalog and assemble the right set of items, including multi-item requests.
    • Remember the customer's preferences and tailor what it suggests.
    • Show products, comparisons, and the cart right in the conversation, not just as text.
    • Build the cart and hand it to checkout.
    • Answer customer service questions in the same conversation, like where an order is, how to return or exchange an item, and what the refund policy says, instead of sending the customer to a support page.

    The agent features guardrails designed to constrain prices and products to actual catalog data, and avoids manipulative upsell patterns. In the repository, these are skills and tools for catalog search, multi-item planning, deep research, personalization, customer care, and in-conversation UI.

    The merchant agent

    The merchant agent supports the people running the store. A user can ask “what should we discount to clear last season’s inventory?” and get an answer based on their own data. Here’s what it can do:

    • Answer questions about sales performance like what's selling and what isn't.
    • Track inventory and proactively flag problems, like an item about to sell out before a promotion starts.
    • Recommend pricing and promotions based on the store's own sales history.
    • Draft marketing campaigns to move the products that need moving.

    When the agent proactively suggests a change, a person approves it before anything goes live, meaning users get the final say while their agent watches the store. In the repository, these capabilities ship as skills for sales analytics, catalog and inventory management, marketing and promotions, and in-portal UI such as charts and dashboards.

    Trusted across the industry

    Companies that serve shoppers, travelers, subscribers, and merchants build and run agents on Claude. Here's what they have to say about building commerce agents with Claude:

    “AI will fundamentally reshape commerce, but trust must remain at the center of every transaction. Merchants are telling us they want more control over how AI engages their customers. Our collaboration with Anthropic on their commerce blueprint helps bring together the intelligence of Claude with the trust, security, and global reach of the Visa network, empowering merchants to deliver better customer experiences while maintaining the relationships that drive their businesses forward.”
    Jack Forestell, Chief Product and Strategy Officer

    “Trust is the currency of commerce, and it is even more critical in the agentic era. With Anthropic's commerce blueprint, we're helping merchants build their own agents with Claude to drive their growth. By combining AI innovation with trusted payments and commerce infrastructure, we're helping connect consumers, merchants and AI agents securely, seamlessly and at scale.”
    Sherri Haymond, Executive Vice President, Global Head of Digital Commercialization

    “Commerce agents are quickly becoming a critical capability for organizations seeking to deliver the personalized, intelligent customer experiences that today’s consumers expect. Our latest research revealed that 85% are now open to collaboration with an AI agent and nearly three in four would trust a personal AI agent more than their best friend to make a purchase on their behalf. This is more than a shift in how people shop. Agentic commerce is rewriting the rules of brand value – fundamentally shaping what gets purchased, when, where and by whom. Anthropic’s commerce agent blueprint provides a proven starting point that can help organizations accelerate deployment and build differentiated experiences that increase customer satisfaction, loyalty, and growth. Combined with Accenture's deep retail and consumer goods expertise, we can help clients move from concept to production and realize value from agentic AI faster.”
    Kath Gramling, Global Consumer Goods, Retail and Travel lead

    “A trip is one of the most complex things a person buys: flights, hotels, cars, and dozens of options to weigh against each other. Penny, our AI assistant, navigates all of that in one conversation and surfaces the best options and best value. We built the latest generation of Penny on Claude because that kind of reasoning is exactly what Claude models are good at.”
    Cobus Kok, Vice President, AI Experiences

    “Millions of consumers, businesses and accountants run their finances on Intuit. Working with partners like Anthropic, we're building highly personalized experiences that provide customers with a clear understanding of what's shifting in their business and why, so they can take action with complete confidence. We are creating a financial system of intelligence by combining frontier AI reasoning, including Claude, with our proprietary data, capabilities, intelligence, and human expertise that powers the next level of prosperity for our customers.”
    Chris Kasten, Intuit’s Chief Architect and SVP of Engineering, Platform and Development Xceleration Group

    “We want our merchants to be everywhere customers are shopping, and increasingly that means a conversation with an agent. We're building on Anthropic's blueprints with a reference storefront implementation that connects them to a merchant's store through Catalog, UCP and Shop Sign-in. Merchants can use Claude to build agents that help customers find products, check out, and answer questions about their orders.”
    Vanessa Lee, VP Product

    “Commerce and stunning customer experiences require personalization, and every brand on Klaviyo sits on more customer data and decisions than any team could act on by hand. Claude closes that gap, turning consumer preferences and performance data into the insights, campaigns and personalization that drive revenue. That's why we keep building with Anthropic: agents do complex analysis, design and decision making, and businesses can focus on delighting customers.”
    Andrew Bialecki, Founder and Co-CEO

    “Wix’s mission has always been to make complex technology simple and accessible for our users. For merchants, that means providing powerful commerce capabilities without adding operational complexity, and agents are a natural next step. Our engineers had a working commerce agent taking prompts within fifteen minutes, and the pilot showed the potential of combining Anthropic’s AI capabilities with Wix’s commerce platform and deep expertise in commerce for SMB.”
    Dror Zalika, Head of Commerce at Wix

    “Our engineers had the blueprint running with no blockers; the setup worked exactly as documented. The practices it bakes in, from tool iteration limits to prompt caching, are the ones we recognized from building Zomato's own agent. Teams standing up their first agent on Claude will skip weeks of trial and error.”
    Akhil Bansal, Senior Engineering Manager

    “Our engineers had both commerce agents from Anthropic's blueprint running locally in well under an hour, with live conversations working on the first attempt. We ran the Claude Code workflow twice and got two different architectures back, each designed to what we'd asked for. For a team starting from scratch, that turns days of agent scaffolding into hours.”
    Ashley Nader, Staff Product Manager

    “Much of what we build at Square is about giving sellers time back, and agents are a big leap in our ability to do that. We're building agentic tools that watch sales, labor, and inventory and come back with real next steps, not just an answer, while keeping sellers in control. Trust is the hardest part of that work, and Claude helps us meet a high standard.”
    Willem Avé, Head of Product

    Getting started

    The blueprint is available today. Contact our sales team to learn more, schedule a demonstration, or discuss how to implement for your organization.

    1. Fork the repository at github.com/anthropics/commerce-agents.
    2. Read the engineering deep-dive at claude.com/blog/the-anatomy-of-effective-commerce-agents.
    3. See the vertical demos and request a working session at claude.com/solutions/commerce.

    Register for our webinar to see the deep dive where we'll share live walkthroughs, demos, and cover how commerce builders can get the most out of Claude.

    Original source
  • Sep 2, 2026
    • Date parsed from source:
      Sep 2, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Google logo

    Gemini by Google

    Proactive cyber defense for governments and enterprises

    Gemini launches the Fairwind Program, giving trusted governments and partners early access to advanced cyber defense tools that can find, verify, and fix vulnerabilities at scale with Gemini 3.8 Flash Cyber and CodeMender.

    Today, we’re launching our Fairwind Program, a limited access program for governments and trusted partners to use our most advanced cyber defense capabilities.

    Defenders wanting to use advanced AI have faced a difficult dilemma: adopt enormous frontier models that could be expensive to deploy and difficult to control across enterprise codebases, or turn to smaller open-weight models that might struggle with complex vulnerability remediation and require teams to build their own tooling and infrastructure from scratch. Until now.

    Today, we’re launching our Fairwind Program to bring the best of Google’s AI and cyber defense capabilities to a trusted group of Google Cloud customers, government agencies, and cybersecurity partners, to help them proactively solve cyber risks at scale. As a first step, the Fairwind Program will give defenders access to powerful and advanced Gemini models to help them autonomously find and fix vulnerabilities, protecting critical infrastructure, public services, and national security.

    Finding and autonomously fixing vulnerabilities

    The Fairwind Program offerings bring together our most advanced cyber model, Gemini 3.8 Flash Cyber, with our CodeMender harness, to help defenders find, verify, and fix vulnerabilities at agentic scale. Spotting weaknesses creates awareness and fear; autonomously finding and fixing vulnerabilities delivers security.

    CodeMender with Gemini 3.8 Flash Cyber delivers the specialized reasoning to write and validate code fixes, at a fraction of the operating cost of traditional frontier models. Instead of taking weeks to manually fix vulnerabilities, defenders can now generate verified, deployment-ready patches in minutes — within an organization’s secure cloud environment.

    Scaling frontline defense

    Providing early access to these powerful cyber capabilities gives trusted defenders a vital adaptation window to harden their systems before bad actors have a chance to exploit new capabilities. We’re staging initial access to government and enterprise partners most critical to society’s resilience:

    • Governments and national cyber authorities: Hardening public-sector networks and citizen services against targeted intrusions.
    • Critical infrastructure operators: Protecting essential services across healthcare, telecommunications, energy, and financial networks from operational disruption.
    • Core technology platforms: Securing widespread software foundations to uplift digital security for millions of downstream users at once.

    To ensure these powerful AI capabilities are used responsibly, participating organizations agree to strict operational standards, including limiting access to employees within their internal cybersecurity, incident response, or penetration testing teams and deploying protections like multi-factor authentication.

    We have more than 650 participating partners globally, including:

    What our Fairwind Program partners are saying

    Trusted defenders and industry leaders are already putting Gemini 3.8 Flash Cyber into practice:

    The Fairwind Program will evolve alongside our partners and users' needs. We will adapt our product offerings and expand partner access, collaborating closely with industry, governments, and open-weight community leaders to strike the right balance between open access and robust security.

    While we are prioritizing Gemini 3.8 Flash Cyber access for customers in the Fairwind Program, any Google Cloud customer can proactively secure their code by using CodeMender with publicly available models hosted on Gemini Enterprise Agent Platform, in combination with industry-leading solutions offered through AI Threat Defense.

    Making an ecosystem-scale impact on cyber defense

    Years of Google’s pioneering zero-trust architecture, advanced AI defenses, and built-in security allow us to protect billions of accounts daily – keeping more people and organizations safe online than anyone else. The Fairwind Program builds on this experience and is part of our broader commitment to global cyber resilience across the entire digital ecosystem, including helping to fortify grassroots cyber defense.

    Through Google.org, our latest commitment brings our total cybersecurity funding to more than $100 million globally. We’re pleased to release our 2026 Google.org US Cybersecurity Impact Report, which details $36 million in funding for 35 cyber clinics to date, providing free, hands-on security support to over 1,250 hospitals, public school districts, and municipal utilities in the U.S.

    Providing a security advantage

    The defender’s edge comes from shrinking the time between detecting a flaw and patching it. Through Google’s Fairwind Program, government and enterprise partners gain autonomous tools to repair systems faster and at scale, keeping them one step ahead of agentic-speed threats.

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Anthropic and hundreds of other software products.

    Create account
  • Sep 2, 2026
    • Date parsed from source:
      Sep 2, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Google logo

    Gemini by Google

    Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

    Gemini releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, bringing faster, lower-cost reasoning, coding and cybersecurity capabilities for agentic workflows. The update adds stronger vulnerability discovery, automated patching, and broader availability across developer, enterprise and consumer surfaces.

    Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity.

    Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7. Gemini 3.8 introduces 2 variants:

    • Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains. It is available at the same introductory price1 as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.
    • Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program.

    While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity.

    Gemini 3.8 Flash: built for long-horizon coding and autonomous agents

    Gemini 3.8 Flash delivers substantial gains from 3.7 Flash, often approaching the performance of higher-cost frontier models.

    On DeepSWE v1.1 (Long-Horizon Software Engineering) 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost.

    Additionally, 3.8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains.

    In quantitative and professional fields that require advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models in benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. 3.8 Flash also achieves a 54.9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields.

    These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.

    For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.

    Gemini 3.8 Flash Cyber: expert cyber performance

    Gemini 3.8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program, provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration.

    Autonomous vulnerability discovery

    On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models.

    To better capture real-world defensive needs which are not limited to just C/C++ codebases like in CyberGym, we also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%.

    Automated patching

    With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation.

    CWE-Bench, run by Collinear, is a challenging external benchmark for patching capabilities. On this benchmark, Gemini 3.8 Flash Cyber is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost.

    Real-world impact: securing Google’s code

    We’re already using Gemini 3.8 Flash Cyber to secure code across Google. For example:

    • The Chrome Security team found that 3.8 Flash Cyber produced 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models that are much larger.
    • Wiz found that Gemini 3.8 Flash Cyber achieves +7.5-9.7% higher recall on their internal penetration testing benchmark for a 2.3-5.2x lower cost compared to other leading frontier models.
    • Google’s Cloud Vulnerability Research team leveraged the 3.8 Flash Cyber model to find a critical foundational vulnerability in less than 2 hours, a vulnerability for which research and discovery usually takes months.

    What our Fairwind Program partners are saying

    [Quotes from partners shown as images]

    Built with safety in mind

    3.8 Flash ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, as per our Frontier Safety Framework. 3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities.

    Gemini 3.8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan, protecting Gemini model users from prompt-injection related malicious attacks.

    Gemini 3.8 Flash and Cyber: get started today

    • Developers: Build with 3.8 Flash and explore agent-first workflows in Google Antigravity or start building today in the Gemini API via Google AI Studio and Android Studio, or generate UIs in Stitch. Get started with our developer docs.
    • Enterprises: Access 3.8 Flash in Gemini Enterprise.
    • Consumers: 3.8 Flash is available to Google AI Pro and Ultra subscribers across the Gemini app, AI Mode in Google Search and Gemini in Google Sheets.
    • Cyber: Through our new Fairwind Program, we’re providing trusted government authorities, as well as critical infrastructure operators and software maintainers with prioritized access to Gemini 3.8 Flash Cyber. Apply for access.
    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Google logo

    Gemini by Google

    The latest AI news we announced in August 2026

    Gemini releases a major August AI roundup featuring Gemini 3.7 Flash, Gemini 3.5 Transcribe, new Pixel 11 devices, Gemini in Chrome on Android, and productivity upgrades in Gemini Live, plus expanded video, music, weather, and climate AI tools.

    Here’s a recap of some of our biggest AI updates from August, including the launch of Gemini 3.7 Flash, Gemini 3.5 Transcribe, and the all-new Pixel 11 series of devices.

    For more than 20 years, we’ve invested in machine learning and AI research, tools, and infrastructure to build products that make everyday life better for more people. Teams across Google are working on ways to unlock AI’s benefits in fields as wide-ranging as healthcare, crisis response, and education. To keep you posted on our progress, we're doing a regular roundup of Google's most recent AI news.

    Here’s a look back at some of our AI announcements from August.

    In August, we continued advancing AI responsibly — making it faster, more accessible, and truly practical for everyone, from software developers and students to creatives and scientists. We’re bringing intelligence directly to where people work and live, with powerful new hardware designed for Gemini in the Pixel 11 series, cost-efficient developer models like Gemini 3.7 Flash, the rollout of Gemini in Chrome on Android, and hands-free voice tools across Google Workspace and Gemini Live. As the Gemini app officially crossed 1 billion monthly users, we also expanded our creative and scientific footprint, introducing studio-quality video and music generation alongside open-source AI models that predict weather patterns and tackle climate challenges. Overall, August marked a shift toward AI that isn't just powerful in theory, but useful in everyday reality.

    Build better agents at a lower cost with Gemini 3.7 Flash.

    We released Gemini 3.7 Flash as our most intelligent workhorse model yet for coding and agents. It arrived just three weeks after our launch of 3.6 Flash, delivering substantial improvements across software engineering, knowledge work, and web development workflows — with an introductory price of half the original 3.6 Flash cost per million tokens.

    Try the new Pixel 11 series, designed for Gemini Intelligence.

    At Made by Google 2026, we unveiled Pixel 11, Pixel 11 Pro, Pixel 11 Pro XL, and Pixel 11 Pro Fold. These devices come with major camera upgrades, enhanced durability, and our fastest, most powerful chip, Google Tensor G6, that runs the latest Gemini Nano model. They’re also designed for Gemini Intelligence to deliver time-saving, personal help. See all the announcements from Made by Google 2026.

    Start the semester with one year of Gemini, on us.

    We’re offering one year of a Google AI plan free of charge for eligible college students around the world — plus new and enhanced study tools — so students can make the most of this school year. We’ve also got tools to help teachers and students as they head back to school, including a dedicated student hub, new teacher-led tools, and SAT prep in Gemini.

    Level up your learning with Search.

    To help you start the semester with confidence, we've added new AI-powered learning features in Search — all built to be safe by design. With these updates, Search can help you grasp complex concepts through interactive visuals, generate practice quizzes for exams like the SAT, ACT, GRE, and LSAT, learn step-by-step with Lens, and stay organized with notebooks.

    Get more intelligent transcription with Gemini 3.5 Transcribe.

    Our latest speech-to-text model delivers precise, intelligent real-time transcription for developer workflows like voice agents, live captioning, and post-call analytics. Unlike conventional models that struggle with noise and jargon, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, context-aware understanding.

    Get more done with new productivity features in Gemini Live.

    Gemini Live is moving beyond conversation to handle complex tasks on your behalf. With new features like Personal Intelligence, Daily Brief, Spark, and hands-free inbox management, you can easily talk through your day and delegate your to-dos without missing a beat.

    See how more than 1 billion people are using the Gemini app every month.

    The Gemini app officially surpassed 1 billion monthly users, making it the fastest-growing product in Google’s history. To mark the milestone, we shared some usage insights, such as: 63% percent of users now talk directly to Gemini, including more “voice only” users — with busy parents 43% more likely to use it for everyday tasks. Gemini now generates 150 million+ images every day, and small businesses are power users, relying on Gemini's all-in-one image, video, and audio creation to craft marketing materials.

    Generate videos with more control using Gemini Omni 1.1 Flash.

    We introduced Gemini Omni 1.1 Flash to bring people even more precision and control for generating videos. The new capabilities deliver studio-quality video production — including scene extension, first-and-last-frame interpolation, crisp 4K upscaling, and faster prototyping. Omni 1.1 Flash is now available in Google Flow, Google AI Studio, the Gemini Enterprise Agent Platform, and the Gemini app.

    Explore how Gemma is offline everywhere from outer space to underwater.

    We released Gemma to help developers build responsible, innovative AI applications anywhere. Over one billion downloads later, Gemma supports environments from phones and edge infrastructure to space. You can explore some of the ways people are using it — from researching interspecies communication to driving medical breakthroughs — and share, discover, and collaborate in our new community repository.

    Learn about Operation Blue Skies, a project to reduce aviation climate impact with AI.

    Our AI-powered forecasts already help flight crews and air traffic controllers adjust routes to avoid forming contrails — all within normal flight operations. Now, we’re partnering with the UK Government and aviation leaders to expand this technology across the North Atlantic, helping airlines reduce aviation’s climate impact on a global scale.

    See how WeatherNext 2 demonstrated a massive leap forward in predicting cyclones.

    In a Nature paper, our researchers showed that WeatherNext 2 predicts cyclone track, intensity, and wind structure with state-of-the-art accuracy — delivering a decade of meteorological progress in one model. Now, we’re open-sourcing WeatherNext 2 to the research community to help build global climate resilience.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    OpenAI logo

    ChatGPT by OpenAI

    September 1, 2026

    ChatGPT adds Healthcare Public Data for eligible Clinicians users in the United States, bringing nine read-only apps for searching public healthcare sources like biomedical research, clinical trials, medication information, Medicare data, and provider records.

    Healthcare Public Data in ChatGPT for Clinicians

    Eligible ChatGPT for Clinicians users in the United States can now use Healthcare Public Data in ChatGPT. The plugin brings together nine apps for searching public healthcare sources, including biomedical research, clinical trials, medication information, Medicare data, and provider records.

    To get started, install the plugin from the Plugin directory, then connect the apps you want to use. These apps are read-only and do not access patient charts. Do not include protected health information in searches sent to public sources.

    For details, see: Using Healthcare Public Data in ChatGPT and Codex.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Anthropic logo

    Claude by Anthropic

    September 1, 2026

    Claude launches Fable 5.1 and Mythos 5.1, its most advanced models for coding and knowledge work.

    Claude Fable 5.1 and Claude Mythos 5.1 launch

    We just launched Claude Fable 5.1 and Claude Mythos 5.1, the world’s most advanced models for coding and knowledge work. For more information, see our blog post: Claude Fable 5.1 and Mythos 5.1.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    OpenAI logo

    OpenAI

    Healthcare organizations can now connect EHR and additional industry data to ChatGPT

    OpenAI adds a new Epic EHR integration and Healthcare Public Data plugin for ChatGPT for Healthcare, bringing authorized patient context, official healthcare datasets, and governed workflows into one secure workspace for clinical, research, and administrative teams.

    Bring ChatGPT and EHR context together

    A new Epic integration and Healthcare Public Data plugin help teams review authorized patient context and work with structured information from official sources in ChatGPT for Healthcare.

    Healthcare organizations need AI that works across the systems and information central to care and operations. Patient context, medical evidence, public healthcare data, and organizational knowledge often live in different places. Connecting these sources in a governed workspace helps teams find the right information, understand it in context, and put it to work across the business.

    Today, we’re introducing a new electronic health record integration that brings authorized patient context from Epic into ChatGPT for Healthcare, along with the Healthcare Public Data plugin for direct, structured access to official healthcare datasets like PubMed, DailyMed, and CMS Coverage. Together, these capabilities bring ChatGPT closer to the systems and sources healthcare teams trust, while supporting the controls and compliance healthcare work requires.

    Healthcare organizations can now connect Epic environments to ChatGPT. Instead of searching across appointment notes, laboratory results, medications, and specialist documentation, clinicians can ask:

    • What has changed since this patient’s last visit?
    • Which recent lab results should I review before today’s appointment?
    • Have there been medication changes or new specialist recommendations?
    • What follow-ups, referrals, or unresolved issues should I be aware of?

    ChatGPT brings together relevant information from the authorized patient record, summarizes important developments, and points back to supporting chart information.

    The integration supports two complementary experiences:

    • EHR context in ChatGPT: Bring authorized patient information from a supported EHR into ChatGPT to review patient history, identify changes, and prepare for appointments.
    • ChatGPT in the EHR workflow: In supported deployments, ChatGPT can be integrated directly into an EHR layout, enabling AI assisted workflows without leaving the patient chart.

    “As a pilot partner, we’re exploring how the new EHR integration with ChatGPT for Healthcare can help clinical teams understand what has changed and what matters most across a complex patient record. By bringing relevant information together more quickly and comprehensively, the technology has the potential to reduce time spent synthesizing data and give clinicians more time with patients. We’re also engaging frontline teams to validate these capabilities in practice and help shape where they can add the most value.”
    —Suresh Gunasekaran, President and CEO, UCSF Health

    Designed to complement existing EHR workflows, these experiences help clinicians review authorized patient information alongside other trusted sources in ChatGPT.

    Work across nine official healthcare sources with one new plugin

    Patient context is one part of the information healthcare teams rely on. They also draw on current research and official information about medications, coverage, clinical trials, and providers.

    ChatGPT for Healthcare already helps teams answer clinical questions and synthesize medical research with trusted clinical search across thousands of medical sources. Building on that foundation, the Healthcare Public Data plugin brings together dedicated connectors to nine official public healthcare sources, including ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed, and PubMed. Teams can work with specific records, fields, identifiers, and versions while focusing the task on these authoritative sources. This makes it easier to compare and verify precise information such as trial eligibility criteria, medication identifiers, coverage policy versions, and provider records without searching each source separately.

    A research team could use ClinicalTrials.gov to identify actively recruiting trials and compare eligibility criteria, while a pharmacy team could use DailyMed to confirm the latest label and warnings for a medication. A population health team planning a diabetes-prevention program could bring relevant research, active trials, and Medicare coverage information together in a source-backed view, helping program leaders evaluate options and identify open questions.

    Evaluated on real healthcare work

    Putting connected healthcare context to work depends on models that can interpret it accurately. OpenAI partners with hundreds of physicians across 60 countries, 49 languages, and 26 medical specialties to help us define, measure, and improve health responses in ChatGPT. To date, these physicians have reviewed more than 700,000 model responses across examples that reflect real-world healthcare questions. Their feedback improves model behavior and strengthens healthcare-specific tools.

    To understand how ChatGPT performs when working with connected EHR context, physicians evaluated responses across 27 clinical use cases, including pre-visit review, clinical timelines, medication review, and handoff summaries. Across 4,363 ratings, physicians rated 99.1% of responses safe across all use cases.

    In a separate two-round evaluation, physicians reviewed hundreds of ChatGPT responses to nuanced clinical questions based on large U.S. healthcare datasets. For each of the five connected data sources tested, more than 93% of responses were rated as having “good” or better accuracy.

    These evaluations measure how ChatGPT works with healthcare context, from surfacing relevant information to citing supporting evidence and preparing work that clinicians and staff can review. We continue to invest in dedicated training and evaluation for healthcare to further improve ChatGPT’s performance and reliability.

    Put healthcare and business information to work

    The same governed workspace that helps teams find and understand healthcare information in context can also help them use it across research, operations, business, and technology. Clinical and business teams can use ChatGPT Work to turn information into reports, analyses, presentations, and plans for review. Technical teams can use Codex to build and improve software that supports care delivery and business operations.

    Plugins for Microsoft SharePoint, Google Drive, Salesforce, Slack, and other enterprise systems expand the approved business context available in ChatGPT while preserving existing permissions.

    “At AdventHealth, the value of AI starts with our people. By putting tools like ChatGPT Work, Codex, and connected business data in the hands of our team members, AI can reduce routine work and help practical innovations move more quickly into action. The goal is simple: give caregivers and teams more time for the human connection that enables our connected, whole-person care.”
    —Robert Purinton, Chief AI Officer at AdventHealth

    Get started with ChatGPT for Healthcare

    ChatGPT for Healthcare gives clinical, research, and administrative teams a governed workspace for using AI. It combines healthcare-specific capabilities with enterprise controls such as role-based access, single sign-on, and audit logs. With an applicable Business Associate Agreement, customers can use ChatGPT Work, Codex, apps, and plugins in the same workspace to support HIPAA-compliant workflows.

    • ChatGPT for Healthcare customers: Ask your workspace administrator to enable the EHR integration and Healthcare Public Data plugin.
    • ChatGPT Enterprise customers: Contact your OpenAI account team to confirm eligibility and the right configuration for your Regulated Workspace.
    • Individual clinicians: Eligible U.S. ChatGPT for Clinicians users can install the Healthcare Public Data plugin. The EHR integration is not available for individual accounts.

    Healthcare organizations can start with the capabilities that fit their needs and expand over time, bringing more of the information central to care and operations into one governed workspace.

    Ready to get started with ChatGPT for Healthcare? Contact our sales team⁠ to find the right solution for your organization.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Anthropic logo

    Anthropic

    Claude Fable 5.1 and Mythos 5.1

    Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1, bringing stronger coding, knowledge work, and scientific research capabilities with lower pricing, improved safeguards, and new enterprise privacy options. Fable 5.1 is generally available, while Mythos 5.1 is limited to trusted access programs.

    We’re introducing Claude Fable 5.1 and Claude Mythos 5.1

    They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.

    Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.

    Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.

    Price

    Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.

    Data retention

    Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.

    Safeguards

    We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon.

    A new performance frontier

    Claude Fable 5.1 sets a new standard on coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves similar or better results than Fable 5 at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)

    Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying.

    Here, you can see how Fable 5.1 compares across various benchmarks:

    Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us:

    “In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”
    — Craig Falls, Head of Quantitative Research, Jane Street Capital

    Scientific research

    We tested Claude Fable 5.1 and Claude Mythos 5.1’s scientific research capabilities across a wide range of domains. What we found—which includes the early examples we share below—adds to the evidence that AI models will soon make important contributions to scientific discovery.

    Molecular design

    Many modern medicines work by binding to targets within the body to block, activate, or deliver something to them. High-affinity binders are necessary for drugs to work at lower doses; designing one is the first step in the development process for many common drug modalities. To see how well Claude Mythos 5.1 could do at this task, we gave the model access to open-source protein design and folding tools and sent its designs to two external organizations for experimental validation. Mythos 5.1 proved able to design very high-affinity binders. On three targets, its binding affinities were 10 times higher than the best designs submitted to Adaptyv Bio’s protein design competitions. Its hit rate (that is, the number of designs that were viable binders) was the strongest we’ve measured to date: it reached nearly 50% across 12 targets. (Hit rates of 10-15% are typical in protein design today.)

    Computational analysis and modeling

    Claude Fable 5.1 trained a neural network to create a new, high-resolution elevation map of a third of the planet Venus. Its work was based on radar images taken by NASA’s Magellan mission more than 30 years ago and a map that already existed for one-fifth of the planet. Claude’s new map now reveals details down to two to three kilometers, rather than 10 to 20, and shows heights up to 25% more accurately than before.

    We’re releasing this map under a Creative Commons license in advance of upcoming NASA VERITAS and ESA EnVision missions, in the hope it might help them determine which geologic features to target for future observation.

    Computational biology

    In computational biology, it’s common to run task-specific machine learning models on GPUs. The speed of these models is therefore a bottleneck to research progress. Mythos 5.1 provided one solution to this problem: by writing custom GPU kernels and caching their intermediate results, it sped up seven open-source deep learning models by up to 2.5 times (with identical outputs).

    The benefits of such speed-ups accumulate quickly. In any given experiment, biologists might run these models thousands of times (for example, testing every possible mutation near every human gene). On analyses like these, the optimized models cut estimated GPU costs by 30 to 60%. This kind of optimization would normally take a team of performance engineers weeks, and is often unaffordable for academic labs. Mythos 5.1 was able to do it in just days, using the publicly available source code alone. We plan to open-source these optimizations soon.

    As our models’ scientific capabilities improve, our investment in scientific progress is also growing. Last week, we previewed the Model Hardware Standard, which allows Claude to directly and safely operate laboratory equipment. We’ve also recently expanded our support for scientists through our AI for Science program, which provides free credits to researchers working on high-impact scientific projects, and we are offering steeply discounted usage through our new Claude Team plan for scientists.

    Safety, security, and alignment

    AI models’ agentic capabilities have become much more powerful over the past two years. But as we’ve documented, greater autonomy comes with new risks. Work on safety, security, and alignment needs to advance at the same pace as AI capabilities. Yesterday, we published a report describing how we are improving our own alignment and security efforts.

    Prior to releasing Claude Fable 5.1 and Claude Mythos 5.1, we (and, in some cases, external researchers) subjected the models to extensive testing for risks across many areas. We describe these efforts in full in our System Card; below is a brief summary.

    Chemical and biological risks

    We tested the extent to which Claude Mythos 5.1 could help create chemical or biological weapons. This involved expert red-teaming, automated evaluations, and a tabletop exercise that paired PhD-level biologists with AI experts, testing whether the models could match human specialists’ performance. Mythos 5.1’s capabilities are greater than those of Mythos 5. However, our evaluations indicate that it still falls short of the next risk tier defined in our Responsible Scaling Policy. We are therefore deploying Mythos 5.1 with the same safeguards that we applied to Mythos 5, which restrict access to research biology capabilities.

    Cyber risks

    We ran a suite of evaluations to assess the cyber capabilities of Claude Mythos 5.1 (with cybersecurity safeguards off). Overall, the model demonstrates the strongest cyber capabilities of any model we’ve released, though it still falls within the lower category of risk in our Frontier Compliance Framework. We also performed extensive stress-testing of our cybersecurity safeguards for Fable 5.1: as well as our own dynamic evaluation of their robustness, we commissioned external testing from two organizations, along with automated testing by Gray Swan. As with Fable 5 and Opus 5, we have not found evidence of a critical-severity jailbreak for these safeguards.

    Agentic safety

    We ran evaluations of how Claude Mythos 5.1 responds to malicious requests and prompt injections (adversarial instructions hidden within content processed by AI models). It refused malicious agentic coding and computer use requests at a comparable rate to Mythos 5, Sonnet 5, and Opus 5, and it is our most robust model to date on an external prompt injection benchmark.

    Alignment

    We tested the model’s behavior through static and interactive behavioral evaluations, analyses of its internal thinking using natural language autoencoders, misalignment-related capability evaluations, a review of our training data, and analyses of our internal pilot use. We also received reports from external testing.

    Our automated behavioral audit found that Claude Mythos 5.1 is better aligned across most metrics than its predecessor, Mythos 5. The model is significantly less likely than Mythos 5 to try to access resources outside of its test environment when assigned an otherwise impossible task. It is also less likely than Mythos 5 to use motivated reasoning to justify its actions (for instance, by reasoning that the situation is a simulation or evaluation), and it is less likely to ignore explicit constraints in pursuit of users’ goals. From our review of its training data, Mythos 5.1 both attempts and succeeds at reward hacking (or cheating) at a lower overall rate than Mythos 5.

    Though generally our alignment evaluations showed improvements, our testing found the model can still sometimes bypass approvals and auto-mode classifiers (as we discuss in more detail in our System Card). There are also limitations to the coverage provided by our alignment assessment. Currently, our automated behavioral audit provides less visibility into very long-context work and multi-agent settings. We also have less coverage of impossible tasks (which can elicit more abnormal and misaligned behavior) than we’d like, although we’ve recently made improvements in this domain and are working hard to continue doing so.

    We have also improved our safeguards so that they allow our models to be more useful without compromising on safety. We describe these changes below.

    Automated safeguards for enterprises

    Enterprise Frontier Safeguards (EFS) allows us to detect and respond to misuse of our models while still providing our enterprise customers the privacy of a zero data retention agreement. With EFS, customers store their data on their own cloud infrastructure, rather than on Anthropic’s systems; any human review is, by default, done by the customer themselves, rather than Anthropic. We developed EFS in close collaboration with more than 100 customers across industries like financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, and with our cloud partners at Amazon Web Services, Google Cloud, and Microsoft Azure.

    EFS will be supported on Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google’s Agent Platform, and Microsoft Foundry. It’s rolling out in phases, starting this fall. As noted above, customers who are eligible for EFS can use Fable 5.1 (and Fable 5) with zero data retention until EFS is ready. You can read more about EFS here; to request access, please complete this form.

    More precise safeguards for biology and cybersecurity

    In the past few months, we’ve made progress in making our safeguards for Fable 5.1 more precise: ensuring that they’re less likely to flag benign content (like queries about medical issues or cyberdefenders using the model to make their systems safer), but still ensuring they provide robust protection against genuine threats.

    As we recently shared, our latest biology safeguards for Fable 5.1 and Fable 5 fire 85% less often for benign requests related to elementary biology and medical questions (relative to those that launched with Fable 5). However, queries related to research and development in the life sciences will still be directed to our Opus models. We’re making the model’s life sciences capabilities available to professionals through an access program for Claude Mythos 5.1 that we’ve developed in partnership with the US government, which we discuss below.

    With Fable 5.1, we’re updating our cybersecurity safeguards to be more precise. We’re also now allowing Fable 5.1 to be used for identifying software vulnerabilities—that is, to conduct the kind of defensive work that improves software security. As a result of these changes, Claude Code users can expect an average of around 60% fewer interventions per session from our cyber safeguards, relative to the previous safeguards on Fable 5. Our safeguards do, however, still redirect several kinds of dual-use cybersecurity tasks (tasks that might have helpful or harmful applications) to our Opus models. This includes penetration testing, exploit generation, and binary-based vulnerability scanning.

    Anti-distillation mechanisms

    Distillation is a method used to extract the capabilities of advanced models. It is often employed on an industrial scale, using thousands of fake accounts. Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards. Fable 5.1 comes with strengthened mechanisms to make distillation attacks harder. For example, it is no longer possible for new API accounts (those created from today onwards) to manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking. This closes off a common, publicly documented distillation technique, which allowed distillers to illicitly extract Claude’s thinking. We’re rolling out the change gradually, to minimize disruption: existing accounts are not currently affected by this change, though it will apply to all users with future model releases. A small number of customers’ custom integrations will then be affected. Our Help Center article explains more about this change and the adjustments that developers can make.

    Trusted access for Claude Mythos 5.1

    Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations whose work is affected by the cybersecurity and life sciences restrictions outlined above. It will be available through two trusted access programs:

    • Cyber Verification Program: The CVP currently provides access to certain Opus and Sonnet-class models with reduced cyber safeguards for defensive security work. In the near future, this program will also include access to Claude Mythos-class models. To apply to join the CVP, click here.
    • Life Sciences Verification Program: The LSVP is designed so that life sciences professionals can use Claude Mythos 5.1 with safeguards designed for professional research and development activities (while all other safeguards remain in place). In partnership with the US government, we have enrolled our first participants, and we plan to expand access to this program to the broader life science community.

    In addition to these trusted access programs, Claude Security, our product that scans codebases for vulnerabilities and suggests patches for human review, is now also powered by Claude Mythos 5.1.

    Compliance with the EU AI Act

    In July 2026, Anthropic (along with 190 other signatories, including several other major AI model providers) signed the EU AI Act’s Code of Practice on Transparency of AI-Generated Content.

    This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their organization, or their conversations with Claude.

    The Act also required us to provide a way for users to tell whether a text likely contains the watermark. We are thus rolling out a detection API in private preview. It is currently available to eligible organizations as required under EU law (such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups). It is also available for enterprises who are similarly obligated to verify watermarking for their own compliance with the Act. We plan to expand access to the detection API over time. You can register interest in access here.

    Cost and availability

    Claude Fable 5.1 is available today on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started with claude-fable-5-1 on the Claude API.

    As mentioned above, we have reduced the price of Fable 5.1’s cache reads (where the model reuses context it has already processed) wherever usage is billed by token, such as on our API. Cache reads now cost 75% less, or $0.25 per million tokens.

    This change leads to a substantial reduction in the overall cost of running the model. For typical workloads, costs are reduced by around 25% relative to Fable 5. For complex coding and highly agentic tasks, the savings could be up to around 45%. The graph below illustrates why this change makes such a big difference:

    Indexed cost of running the same workloads on Fable 5 and Fable 5.1, at usage-based pricing measured at default effort over four weeks of actual usage in August 2026. Typical workload covers Fable usage across Claude Enterprise, Claude Code, and the API. Highly agentic workload covers context-heavy, tool-heavy work, where cache reads make up most of the cost.

    Fable 5.1’s pricing is otherwise the same as Fable 5’s: $10 per million input tokens and $50 per million output tokens. In parallel, we’re continuing our work to bring many of the improvements of Fable 5.1 to the rest of our model family.

    As discussed above, Claude Mythos 5.1 is available to vetted cyberdefenders and life scientists. Currently, it is only available to a set of US organizations, though we’re coordinating with the US government to expand access to a broader set of domestic and international partners as quickly as possible. To register interest in access to Claude Mythos 5.1 for cyberdefense through the CVP, see here.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Anthropic logo

    Anthropic

    Developing Enterprise Frontier Safeguards with our customers

    Anthropic introduces Enterprise Frontier Safeguards, a new opt-in security and privacy solution that pairs zero data retention with automated misuse detection in customer-controlled cloud storage, with no Anthropic human review required and phased rollout starting later this fall.

    Today we’re announcing Enterprise Frontier Safeguards (EFS), a solution that combines the privacy of zero data retention (ZDR) with state-of-the-art safeguards for detecting misuse. EFS works by storing data in cloud infrastructure controlled by the customer, not Anthropic. EFS will be rolling out to customers in phases, starting later this fall. To make the transition smooth, eligible customers will receive ZDR on Fable 5 and Fable 5.1 until EFS is ready.

    We developed EFS in close collaboration with more than 100 customers in industries like financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, and with our cloud partners at Amazon Web Services, Google Cloud, and Microsoft Azure.

    EFS will be supported on Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google’s Agent Platform, and Microsoft Foundry.

    Solving the dilemma of frontier security

    Mythos-class models, like Claude Fable 5.1, represent a major increase in intelligence and agentic capabilities. However, with that increase comes the potential for both misuse and autonomous misbehavior.

    Over the last few months, we’ve seen substantial evidence of attempted misuse of AI models. These range from typical forms of abuse, such as fraud, to sophisticated cyberattacks, which can include agents autonomously engaging in destructive behavior. Some of these instances involve theft or misappropriation of enterprise customers’ credentials, which are difficult to detect without the ability to monitor traffic and detect abnormal behavior.

    Furthermore, because the most sophisticated misuse can involve many tasks spread across multiple sessions and accounts, it is not sufficient to run automated analysis on each interaction separately and then instantaneously discard the data. Effective detection requires storing data for a meaningful period of time so that it can be correlated across time and accounts.

    For this reason, we introduced 30-day data retention starting with Fable 5. This policy was not motivated by a desire to train on enterprise data: Anthropic has never trained on enterprise data without explicit permission, and never will.

    The enterprises we worked with generally understood the safety and security value of data retention, but many–especially in regulated industries–found it difficult to use models with data retention. We therefore sat down with customers to design a solution that could provide the best of both worlds: the privacy of ZDR and the safety allowed by monitoring across time and accounts.

    Designed with our customers

    We built Enterprise Frontier Safeguards with feedback from the experts who will use it every day: security, product, compliance, and delivery teams. One of the groups we worked with was the Analysis and Resilience Center for Systemic Risk (ARC), whose members include the chief information security officers of the largest US banks, including Goldman Sachs, Morgan Stanley, Citi, Bank of America, and Wells Fargo.

    We also worked with leaders at companies such as Comcast, KPMG, Mastercard, Salesforce, and Visa, to make sure the design held up across industries. Our conversations spanned a quarter of the Fortune 100, every US global systemically important bank, and virtually every regulated industry.

    Here is what we heard from this wide range of customers, and what we built into EFS to address these common concerns:

    On monitoring

    Enterprises have long applied monitoring for insider risk, and now want help upleveling monitoring for agents. Their concerns were about Anthropic’s automated monitoring systems meeting their regulatory standards.

    With EFS, customers control how data gets reviewed. When monitoring detects a pattern that needs attention, those signals are sent directly to customers so they can review what the automated systems detected.

    On data storage

    It’s a lot of work for enterprises to add another “trusted data vendor” for a number of reasons. They need to notify all of their customers who these vendors are and update contracts. They also have internal requirements for safely storing and auditing data, given its high level of sensitivity. Because of these concerns, we architected EFS so that customers have the ability to store data on their existing cloud infrastructure.

    In EFS, customers can control their data storage and management. Customers want the ability to have their data live in infrastructure they control, under their own encryption keys, access policies, and audit logging. Activity data used for monitoring can be stored in the customer’s own cloud account (such as Amazon S3, Azure Blob Storage, or Google Cloud Storage).

    On automated and human review

    Even as automated review is becoming more effective, a person looking at a flag still adds value by confirming real misuse and clearing false positives. But what we heard from many customers, especially those in regulated industries, is that the person doing that review needs to be one of their own. Many operate under rules that tightly govern who may see certain information—privileged legal material, non-public information, drug-safety reports. Their teams are already trained and cleared for that work.

    EFS has automated safety monitoring, no Anthropic human review required. Customers want protection against cyberattacks, and appreciate that these can be difficult to detect if they unfold across many sessions and accounts. With EFS, automated systems analyze a rolling window of traffic for signals of serious misuse, including attempts to develop offensive cyber or biological capabilities and signs of stolen or leaked credentials. Those flags go directly to the customer and their people take it from there – no human review by Anthropic employees is required.

    How EFS works

    These controls are designed to work the same way whether you access Claude directly from Anthropic or through a cloud partner. Customers on Amazon Web Services, Google Cloud, and Microsoft Azure will get equivalent controls, with their activity data stored in their own cloud account, in the environment they already trust. We’re also working to support third-party offerings that serve customers that are eligible for Enterprise Frontier Safeguards.

    Customer-owned storage, Customer-Managed Encryption Keys, and fully automated review are each opt-in, so you enable the ones your organization needs. None of them change model behavior, API pricing, or rate limits.

    Anthropic doesn’t charge for Enterprise Frontier Safeguards. If customers elect to store their data in their cloud account, their cloud provider bills them for that storage, as well as reads, writes, and data egress fees, the same way it bills any other resource.

    Getting started

    Enterprise Frontier Safeguards will roll out to customers in phases, with the goal of making it broadly available later this fall. To request access to Enterprise Frontier Safeguards, please complete this form.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Google logo

    Gemini by Google

    Introducing agentic video understanding with Gemini

    Gemini launches agentic video understanding across its latest Flash models, cutting token use and costs while improving video analysis quality. The new capability works for video uploads and YouTube videos in the Gemini API, Google AI Studio, and Gemini Enterprise Agent Platform.

    Our new agentic feature for video analysis cuts token consumption by up to 88%, reduces costs by up to 66%, and boosts quality by up to 7%.

    Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.

    The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

    Benchmarks

    Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. Across standard video analysis benchmarks, Gemini models with agentic video understanding reduce analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%.

    These efficiency gains are especially pronounced on long-form video (from 10-minute how-to guides to 90-minute lectures and multi-hour recordings), where static processing forces developers to choose between high token costs or techniques that drop critical details.

    Activating agentic video understanding drops token consumption by up to 88% and boosts accuracy by up to 7% with Gemini 3.7 Flash.

    While these gains span all three supported models, Gemini 3.7 Flash with agentic understanding offers the best possible quality overall and the best combination of quality and cost efficiency, putting it at the accuracy-to-cost pareto frontier among tested models for video understanding.

    Using agentic video understanding places Gemini 3.7 Flash at the accuracy-to-cost pareto frontier for video analysis.

    How it works

    Instead of static processing where the model ingests media streams at a fixed frame rate, agentic video understanding enables Gemini to take an active, goal-directed role in determining what to watch, at what speed, and through which modality (frames, audio, or transcript), fetching only the moments and signals needed. While developers could previously do this manually, with agentic video understanding, Gemini can accomplish it through an agentic loop, invoking an internal tool to load the relevant part of the video file, significantly reducing development overheads.

    Capabilities and use cases

    Agentic video understanding transforms how developers can process long-form video content across a variety of demanding applications.

    • Sub-second moment retrieval: Pinpoint split-second state changes and tight cut boundaries that are easily missed at 1 FPS, making precise automated video editing possible.
    • Long-form needle-in-a-haystack search: Answer complex queries across multi-hour videos without consuming millions of tokens.
    • Anomaly detection: Resample interesting time windows at higher FPS to inspect rapid motion and subtle visual artifacts.
    • Counting action & object: Accurately track repeated physical movements and distinct objects over time.

    Real-world results

    Many of our early access partners saw strong performance while testing with agentic video understanding. Here’s what they have to say:

    Getting started

    Agentic video understanding is available via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform, launching across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. It uses standard Gemini API token pricing with no additional feature fee.

    To enable it, simply set processing to "agentic" in the API configuration. Read our developer guide to get more insights into the feature and how to get started.

    We are also bringing the efficiency and quality improvements of agentic video understanding to billions of users across Google products. The feature will roll out to all users in the Gemini app across Flash and Flash-Lite models soon. And in the coming months, agentic video understanding will also power YouTube's ‘Ask YouTube’ feature on the video watch page, leveraging Gemini to deliver higher-quality answers grounded in the visuals.

    Acknowledgement for their contribution to this work:
    Sergi Caelles, Filip Pavetić, Ahmet Iscen, Suhas Yogin, and the Agentic Vision team.

    Original source