OpenAI Updates & Release Notes
135 updates curated from 192 sources by the Releasebot Team. Last updated: Jul 10, 2026
- Jul 10, 2026
- Date parsed from source:Jul 10, 2026
- First seen by Releasebot:Jul 10, 2026
ChatGPT Work: Take on your most ambitious tasks
OpenAI launches ChatGPT Work, powered by GPT-5.6, to turn scattered notes and drafts into finished work.
Powered by GPT‑5.6, ChatGPT Work brings together context from your team’s tools to turn scattered notes, drafts, and ideas into finished work — and keeps projects moving while you stay in control.
Available to all plans on desktop today, and rolling out to Plus, Pro, Business, Enterprise, and Edu on web and mobile over the next few days.
https://openai.com/chatgpt-work/
Original source - Jul 9, 2026
- Date parsed from source:Jul 9, 2026
- First seen by Releasebot:Jul 10, 2026
GPT-5.6 System Card
OpenAI releases GPT-5.6, a new three-model family with Sol as the flagship, Terra as the lower-cost option, and Luna as the fastest, most cost-efficient model. The launch pairs stronger cybersecurity and biology safeguards with extensive safety testing and ongoing deployment monitoring.
GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet—are built to deliver these models safely and at scale, around the world.
Under our Preparedness Framework, we are treating Sol, Terra and Luna as High capability in both Cybersecurity and Biological and Chemical risk. None of them reach our High threshold in AI Self-Improvement. We have implemented a tailored set of safeguards, adapted to each model’s capability profile, to sufficiently minimize the associated risks.
This system card is a detailed report of the work we did to understand and mitigate GPT-5.6’s safety risks before deployment. The six most important things to know are that:
These models are a meaningful step up in cybersecurity capability, but they do not reach our risk framework’s highest level (Critical). GPT-5.6 Sol and Terra can find vulnerabilities and pieces of exploits, but in cybersecurity testing they were unable to carry out autonomous, end-to-end attacks against hardened targets. Separate evaluations examined misaligned behavior in agentic coding tasks and found GPT-5.6 shows a greater tendency than GPT-5.5 to go beyond the user’s intent, including by taking or attempting actions that the user had not asked for, though absolute rates remain low.
We are taking a more conservative approach as we continue to strengthen the system against adaptive attacks. Compared with previous models, our GPT‑5.6 Sol cyber safeguards block roughly ten times more potentially harmful activity. Because these measures can create friction for benign users, we provide an option in ChatGPT and Codex to easily retry prompts on lower-capability models, and we will continue reducing the impact of our safeguards on benign users while maintaining a high robustness bar. This reflects our iterative deployment approach: starting conservatively and improving based on what we learn from real-world use.
To make these models safe, we added new technology to a safety stack that is more than the sum of its parts. The models are trained to be safe, Sol and Terra are served with newly added activation classifiers focused on sensitive domains that watch the model and can intervene to stop unsafe answers during generation, and certain conversations are scanned so unsafe outputs are blocked in real time if they cross safety boundaries. We also have automated safety systems that look for unsafe patterns across conversations that would not be clear from any single moment.
Severe harm requires a chain of successful steps, and our safeguards place barriers throughout that chain. Based on our threat modelling in cybersecurity and biology, we’ve designed our safety stack so that even if an attacker does complete one step on the path to harm, safeguards will still stop the model from allowing severe harm. We also have programs in place so that when GPT-5.6 models are broadly available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders.
Our safeguard testing was more intensive than for any earlier release, and we are continuing to test during deployment. Expert humans and external testers used a diverse set of approaches to find gaps. We’ve also dedicated over 700,000 A100e GPU hours to automatically find universal jailbreaks, and we will run automated red teaming continuously during deployment. As jailbreaks are reported, we reproduce, mitigate and retest for them so that gaps are addressed.
Providing broad access, particularly for cybersecurity capabilities, will have important safety benefits. Our testing suggests that GPT-5.6 is better at finding and fixing cyber vulnerabilities than at exploiting those vulnerabilities in real attacks. That gives defenders an opportunity to harden systems before cybersecurity weaknesses are exploited—an opportunity that may narrow as offensive capabilities improve. Our safeguards therefore focus on making malicious use at scale harder, while still enabling the day-to-day work of securing systems.
The card further details model data and training, model safety including disallowed content evaluations, vision capabilities, avoidance of accidental data-destructive actions, user confirmations during computer use, robustness evaluations including jailbreaks and prompt injection, health benchmarks including HealthBench and dynamic mental health benchmarks, hallucination performance, alignment including forecasting misaligned behavior, chain of thought evaluations, metagaming, bias evaluations, and preparedness assessments.
Preparedness Framework designates all three GPT-5.6 models as High capability in Biological and Chemical and Cybersecurity domains, but below High in AI Self-Improvement. The card describes extensive capabilities assessments, including biological and chemical capabilities (e.g., virology troubleshooting, protein design), cybersecurity capabilities (e.g., capture the flag challenges, vulnerability identification and exploitation), and AI self-improvement capabilities (e.g., debugging internal research, kernel optimization, training loop improvements).
Safeguards include layered defenses such as model safety training, real-time monitoring with activation classifiers and safety reasoners, automated and third-party red-teaming for jailbreaks, actor-level enforcement, trust-based access programs for biology and cybersecurity, and security controls to protect intellectual property and model weights.
The system card emphasizes ongoing evaluation, iterative improvements, and a commitment to balancing capability deployment with safety and risk mitigation.
Original source All of your release notes in one feed
Join Releasebot and get updates from OpenAI and hundreds of other software products.
- Jul 9, 2026
- Date parsed from source:Jul 9, 2026
- First seen by Releasebot:Jul 9, 2026
GPT‑5.6 is now the preferred model in Microsoft 365 Copilot
OpenAI adds GPT‑5.6 to Microsoft 365 Copilot, making it the new preferred model in Word, Excel, PowerPoint, Chat and Cowork. The update brings stronger AI help for drafting, analysis, presentations and collaboration with more useful work from every token.
Today, OpenAI announced GPT‑5.6, which will become the new preferred model in Microsoft 365 Copilot—in Word, Excel, PowerPoint, Chat and Cowork. For Microsoft 365 customers, the update brings OpenAI's latest flagship model series into productivity tools people use every day, helping them create, analyze, and collaborate with more capable AI assistance across workstreams. GPT‑5.6 is OpenAI’s latest flagship model series, which delivers more useful work from every token, with stronger performance per dollar and on demand capability for the most complex tasks.
With GPT‑5.6, Microsoft 365 users will be able to create higher-quality work products with less effort across the apps they already rely on:
- In Word, GPT‑5.6 can help people draft, edit, and refine documents with fewer rounds of prompting.
- In Excel, it can support deeper analysis while using tokens more efficiently, helping users move faster from data to insights.
- In PowerPoint, it can help turn early ideas into more polished, visually compelling presentations with less manual guidance.
- In Cowork, it can help users complete complex, cross-functional work and produce higher-quality outputs with less manual coordination.
“We can’t wait for customers to see what GPT‑5.6 in Microsoft 365 will do, enabling them to work even more effectively with AI in the tools they use every day,” said Nitin Agrawal, President, Copilot & Agents Core, Microsoft. “Using Copilot powered by OpenAI’s latest model, customers will be able to produce more polished outputs in Word, Excel, PowerPoint, Cowork, and Copilot Chat, whether they are drafting documents, analyzing data, creating presentations, or collaborating across teams. We’re excited to continue building with OpenAI to bring more powerful AI experiences to people and organizations around the world.”
“Microsoft 365 is where millions of people write, analyze, create, and collaborate every day,” said Nikunj Handa, Head of API Product, OpenAI. “By bringing GPT‑5.6 to Microsoft 365 Copilot through the OpenAI API, we're helping organizations get more useful work from every token, and more value from AI in the tools they already use.”
In addition to serving the models natively, Microsoft will also access OpenAI models directly through the API to bring GPT‑5.6 to Microsoft 365 customers.
Our partnership with Microsoft has always been about bringing the benefits of advanced AI to more individuals and organizations, and we’re excited to continue building on that shared commitment.
Original source - Jul 9, 2026
- Date parsed from source:Jul 9, 2026
- First seen by Releasebot:Jul 9, 2026
GPT‑5.6: Frontier intelligence that scales with your ambition
OpenAI launches GPT‑5.6 family with Sol, Terra, and Luna for general availability, bringing stronger coding, knowledge work, design, cybersecurity, and science performance plus new ultra multi-agent acceleration and updated API tooling.
Efficient by default, maximum performance on demand
We’re launching the GPT‑5.6 family of models for general availability following our limited preview: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.
GPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost. The result is stronger performance per dollar: more successful work for the same spend, or comparable results at a lower total cost. We also introduce a new way to accelerate the most demanding work: ultra is our highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster. Stronger computer use and design judgment make GPT‑5.6 Sol our most polished collaborator yet, helping it inspect, refine, and deliver ready-to-use results.
We trained GPT‑5.6 to get more useful work from every token. On Agents’ Last Exam (opens in a new window), an evaluation of long-running professional workflows across 55 fields, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable: GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost. On the Artificial Analysis Intelligence Index (opens in a new window), a broad measure of intelligence spanning agentic work, coding, scientific reasoning, and general capabilities, GPT‑5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost.
GPT‑5.6 launches with our most robust safeguards to date, designed to be resilient against determined and adaptive misuse without broadly limiting legitimate work. Before general availability, we put the models and safeguards through our most extensive evaluation period yet, combining human red teaming with large-scale automated testing. During the preview, we worked closely with expert organizations and with trusted partners to pressure-test defenses and strengthen safeguards before broader launch. The resulting system layers protections trained into the model with real-time checks, monitoring, and access calibrated to trust and risk.
A leap forward in design
GPT‑5.6 Sol is our best coding model yet. On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one-third less. That advantage extends across the family: Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost. It also sets new state-of-the-art results on Terminal‑Bench 2.1 and DeepSWE, which test complex command-line workflows and long-horizon engineering in real codebases.
GPT‑5.6 can write and run lightweight programs that coordinate tools, process intermediate results, monitor progress, and choose the next action as work unfolds. This lets tool-heavy tasks advance with fewer tokens, fewer model round trips, and less guidance. Instead of requiring developers to script every step or passing every tool response back through the model, Programmatic Tool Calling (opens in a new window) in the Responses API can filter large amounts of intermediate data, retain only what matters, and adapt its workflow along the way.
For problems that reward a greater investment of time and compute, GPT‑5.6 can push beyond this efficient default. max gives GPT‑5.6 even more time than xhigh to reason and explore alternatives, run checks, and revise its approach. ultra goes further by coordinating four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks. The charts below compare ultra’s default four-agent setup with a one-agent baseline across BrowseComp, SEC-Bench Pro, and Terminal-Bench 2.1; BrowseComp and SEC-Bench Pro also show 16-agent configurations. Across all three evaluations, adding parallel agents shifts the score-latency frontier upward and to the left, reaching stronger results in less time. In the API, developers can build ultra-like experiences using the multi-agent beta in the Responses API.
A leap forward in design
GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back.
GPT‑5.6’s frontend capabilities also turn natural-language requests into polished, interactive explanations and visualizations within ChatGPT Work.
End-to-end knowledge work
GPT‑5.6 delivers better results for professional tasks. It takes messy context from your documents and everyday workflows like Slack, Notion, Microsoft 365, and Google Drive, and converts it into expert-level, shareable artifacts.
GPT‑5.6’s strength on knowledge work shows up in evaluations spanning long-horizon professional analysis, browsing, tool use, and computer use. GPT‑5.6 Sol sets new state-of-the-art results on BrowseComp at 92.2% and OSWorld 2.0 at 62.6%; on OSWorld, it surpasses Opus 4.8 while using 85% fewer output tokens. Here, the performance-per-dollar gains extend across the GPT‑5.6 family. Luna nearly matches GPT‑5.5’s peak performance at less than half the estimated cost, while Terra surpasses it at a lower cost.
GPT‑5.6 Sol improves quality in presentations, documents, and spreadsheets, producing outputs that are more polished and accurate. It can create fully editable presentations from scratch, translating a prompt and source material into a coherent visual narrative with strong layouts, hierarchy, and design.
The improvement is especially pronounced when following templates and reference decks. GPT‑5.6 can infer a deck’s design system—layouts, typography, spacing, colors, and recurring content patterns, including rules embedded in the Slide Master—and apply those conventions consistently to new material. In this example, when asked to update numbers based on a reference file, the GPT‑5.5 output is missing key components from the master slide, while GPT‑5.6 follows the reference structure more faithfully.
GPT‑5.6 also creates more visually refined documents and spreadsheets. It follows complex reference formats more faithfully, which is important for repeatable knowledge work activities. It handles equations and financial models with greater precision, and makes better use of typography, spacing, hierarchy, and page or worksheet layout.
Early customers testing GPT‑5.6 saw improvements to knowledge work outputs across domains.
Pushing the frontier on cyber and science
GPT‑5.6 is our strongest cybersecurity model yet, achieving frontier performance with significantly fewer tokens. On ExploitBench2, which measures progress from reaching vulnerable code through arbitrary code execution, it scores 73.5% versus GPT‑5.5’s 47.9% at a comparable output-token budget. On ExploitGym3, which asks agents to turn real-world vulnerabilities into working exploits, it almost doubles GPT‑5.5’s peak pass rate, from 15.1% to 24.9% under the two-hour cap; with six hours, it reaches 33.7%. On SEC-Bench Pro, which tests proof-of-concept generation on complex software, it scores 71.2% versus GPT‑5.5’s 45.8% at an improved latency.
GPT‑5.6 supports important defensive tasks such as secure code review, patching, threat modeling, and blue teaming. Qualified individuals and organizations in OpenAI Daybreak’s Trusted Access for Cyber program can access more of its defensive capability through more precise safeguards for verified work in authorized environments, including vulnerability triage and validation, malware analysis, detection engineering, and patch validation.
Individuals can verify their identity and request trusted access (opens in a new window), and organizations can apply for their teams. Individual members will need to enable Advanced Account Security (opens in a new window) with hardware-backed passkeys by September 1 to retain access to our most cyber-capable frontier models; those who do not will return to default access. Users who do not already have hardware-backed passkeys can receive preferred pricing (opens in a new window) from our partner, Yubico. We are also taking additional steps to restrict access to high-risk entities and in high-risk jurisdictions.
GPT‑5.6 Sol also shows broad gains across scientific research. On life sciences evaluations, GPT‑5.6 demonstrates Pareto improvements over GPT‑5.5 on real-world biology, life science research workflows, and chemistry.
GPT‑5.6 accelerates OpenAI
GPT‑5.6 is our strongest model yet for accelerating AI research. Inside OpenAI, researchers use it across the development loop: diagnosing failures, optimizing training systems, running experiments, and interpreting results. We already saw that acceleration and stronger adoption during the internal testing period of GPT‑5.6, as average daily output tokens per active researcher were more than twice the highest level observed for GPT‑5.5.
This way of working is quickly becoming standard. Over the past six months, the share of research compute devoted to internal coding inference grew 100-fold, while internal agentic token usage increased approximately 22-fold. These adoption metrics do not measure research progress on their own, but they show how rapidly AI assistance is increasing for research and across other teams like sales, marketing, user ops, finance, and more.
To measure this capability directly, we developed an internal suite of evaluations based on real AI research tasks, including debugging research systems, optimizing kernels and training recipes, running machine-learning experiments, and improving another model.
Scaling safety and security with capability
As model capabilities increase, we strengthen our safety stack so advanced intelligence can remain broadly useful while applying greater scrutiny to the highest-risk uses. For GPT‑5.6, we built our most robust safety system to date, calibrated to each model’s capabilities and powered by more compute than ever before.
The GPT‑5.6 models are more capable than our earlier models in both biology and cybersecurity but do not cross the Critical threshold in either category. In cybersecurity, our testing suggests GPT‑5.6 is better at finding and fixing vulnerabilities than at reliably carrying out autonomous, end-to-end attacks against hardened targets—giving defenders an opportunity to strengthen systems before weaknesses are exploited. In biology, our testing suggests GPT‑5.6 can support legitimate research but does not provide the end-to-end capability needed to create, engineer, or synthesize a highly dangerous novel threat.
Both domains are inherently dual-use. In cybersecurity, the same capabilities that could help an attacker exploit a vulnerability can help a defender find it, reproduce it, and build a reliable fix. Overblocking therefore creates a security risk of its own. It can prevent defenders from testing systems and deploying patches while malicious actors continue using other models, including increasingly capable open-source models, as well as established tools. Effective safeguards account for the context and likely consequences of a request, preserving legitimate defensive work while applying stronger controls where the evidence indicates a serious risk of harm.
GPT‑5.6’s safeguards are layered for greater accuracy and redundancy, and designed to adapt quickly as new attacks emerge. Protections trained into the model work alongside real-time checks, continuous monitoring, and account-level enforcement, to help the system remain safe even when a particular layer does not work as intended. In many systems, classifier flags alone decide what to block, relying on lower intelligence models that are harder to change in order to prevent harm. Our approach adds a reasoning monitor that reviews the conversation to determine if there is a potential for harm. This design is intended to enable defensive work while blocking serious misuse, with the most sensitive capabilities reserved for verified users through Trusted Access. Because some protections use test-time reasoning, we can rapidly update them to close gaps without retraining classifiers from scratch.
We are taking a more conservative approach as we continue to strengthen the system against adaptive attacks. Compared with previous models, our GPT‑5.6 Sol cyber safeguards block roughly ten times more potentially harmful activity. Because these measures can create friction for benign use, we provide an option in ChatGPT and Codex to easily retry prompts on lower-capability models, and we will continue reducing the impact of our safeguards on benign use while maintaining a high robustness bar. This reflects our iterative deployment approach: starting conservatively and improving based on what we learn from real-world use.
Before general availability, we ran our most intensive safety evaluations to date, including extensive red teaming, robust capability and safeguard testing with external experts, and approximately 700,000 A100e GPU hours of black-box automated red teaming. This enabled us to systematically probe likely weak points, surface jailbreaks, and help us strengthen the system before launch.
There is no such thing as perfect security, and our work to secure increasingly capable models continues. New weaknesses will be discovered, as will new jailbreaks that circumvent existing safeguards. Each new generation of model will also create new avenues for attack and misuse. We build for that reality through layered safeguards, continuous monitoring, rapid remediation, and collaboration across the defensive community. For GPT‑5.6, we have paired our existing security (opens in a new window) and biology bug bounty programs with a new rapid-remediation process and our strongest monitoring effort to date. Findings from researchers, monitoring, and real-world misuse will feed into new evaluations and stronger safeguards on an ongoing basis.
Read more about our safeguards in the updated GPT‑5.6 system card (opens in a new window).
Availability and pricing
GPT‑5.6 spans three model tiers: Sol, our flagship; Terra, a lower-cost model with performance competitive with GPT‑5.5; and Luna, our fastest and most affordable model. The number identifies the generation, while Sol, Terra, and Luna are durable capability tiers that can advance on their own cadence.
GPT‑5.6 is available starting today across ChatGPT, Codex, and the OpenAI API. The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours.
- Chat: Plus, Pro, Business, and Enterprise users access GPT‑5.6 Sol through medium and higher effort settings. Pro and Enterprise users can also select GPT‑5.6 Sol Pro for the highest-quality results on complex tasks.
- ChatGPT Work and Codex: Free and Go users access GPT‑5.6 Terra. Plus, Pro, Business, and Enterprise users can choose among GPT‑5.6 Sol, Terra, and Luna and set an effort level for each. max is available to all users with access to GPT‑5.6 in ChatGPT Work and Codex and can be toggled on in settings. In ChatGPT Work, ultra is available to Pro and Enterprise users. In Codex, it is available to Plus and higher plans.
- API: Developers can access Sol, Terra, and Luna through the OpenAI API. In the Responses API, Programmatic Tool Calling lets GPT‑5.6 write and run programs in-memory that coordinate tools and process intermediate results, making it Zero Data Retention (ZDR) compatible. Multi-agent, initially available in beta, lets GPT‑5.6 run concurrent subagents and synthesize their work in a single request.
GPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints (opens in a new window) and a 30-minute minimum cache life. For GPT‑5.6 and later models, cache writes are billed at 1.25x the model’s uncached input rate, while cache reads continue to receive the 90% cached-input discount.
Original source - Jul 9, 2026
- Date parsed from source:Jul 9, 2026
- First seen by Releasebot:Jul 9, 2026
ChatGPT is now a partner for your most ambitious work
OpenAI introduces ChatGPT Work, an agent that helps users tackle bigger projects across apps and workflows, create docs, slides, sheets, and web apps, and keep work moving with scheduled tasks, desktop browser tools, and Codex-powered editing and review.
Introducing ChatGPT Work
Introducing ChatGPT Work, an agent in ChatGPT that helps you take on more ambitious tasks. It can gather information across your apps and workflows to create finished materials like sheets, slides, docs, and web apps, and stay with complex projects for hours by breaking them into smaller steps and completing them independently.
With Codex technology built-in, ChatGPT can now move beyond answering questions to getting real work done across web, mobile, and desktop. More than 5 million people use Codex every week. Although it began as a coding agent for developers, more than 1 million people now use it for work outside software development, showing how its capabilities can support a wider range of tasks.
To better manage these tasks, ChatGPT Work is powered by our latest frontier model, GPT‑5.6, which is also rolling out today. GPT‑5.6 makes ChatGPT state of the art at reasoning through multi-step tasks and creating materials that follow your templates and reference files.
The best way to learn how to use ChatGPT Work is to give it a task you already know well: analyze a month-end budget variance, turn source materials into a marketing campaign brief, or prepare for a sales meeting. You can follow its progress, answer questions, change direction, and approve important actions.
You can even ask ChatGPT Work to take on entire workflows with a single request. For example, it can turn customer research into a campaign brief, use that brief to create marketing assets, and adapt those assets for different markets while carrying context through every step.
Even when you’re away from your computer or phone, ChatGPT Work can keep projects moving forward with Scheduled Tasks. For example, it can independently turn new messages from Microsoft Teams and Slack into updated docs or slides, then share important changes with your team.
In early testing, we’ve seen how ChatGPT Work expands what users can do:
Used ChatGPT Work to build a repeatable system for reviewing thousands of leads each month. It traced customer touchpoints across Zapier’s CRM, email, and other tools, found where follow-ups broke down, and generated a weekly executive dashboard that highlighted missed pipeline and revealed seven figures in potential sales. — Angela Ferrante, Head of Enterprise Marketing at Zapier
Nearly 100% of teams inside OpenAI, including finance and sales, now use ChatGPT Work and Codex to move faster, take on harder tasks, and spend more time with customers.
- In sales, ChatGPT Work turned a discovery conversation into a tailored proof of concept for a mission-critical problem within 24 hours—a process that normally takes weeks. ChatGPT structured the notes, routed the request to a solutions architect, and collaborated with the technical team, freeing the lead to focus on the customer and serve as a high-value consultative partner.
- In finance, ChatGPT Work reduced month-end close and forecasting from days to hours by helping teams find source data, move it into Excel or Sheets, reconcile it, create slides, and verify the results. This lets the finance team spend more time understanding what changed in the forecast, explaining why it changed, and advising leaders on what the company should do next.
On web and mobile, ChatGPT Work is rolling out today now for Pro, Enterprise, and Edu plans. It will roll out to Plus and Business plans over the next few days. In the ChatGPT desktop app, Chat, Work, and Codex are available on every plan, including Free, and is available globally to download on Windows and Mac.
Work anywhere. Go further on desktop.
ChatGPT Work is designed to keep tasks moving forward wherever you are. You can ask it to start a task from your phone, review a draft on the go, or check the status of a longer-running workflow between meetings. When you return to your desk, you can pick up the same work on the web.
For an even more powerful experience, the ChatGPT desktop app now goes further. On desktop, ChatGPT can use your local files and apps to get work done. For web-based work, ChatGPT’s new built-in browser lets it bring in websites, tools, and online files, giving you one place to move work forward.
Starting today, the Codex app is merging with the new ChatGPT desktop app. Codex remains the same powerful coding agent for developers and technical professionals, now with new capabilities across core workflows, including inline editing within diffs, pull request review in the side panel, faster computer use (powered by GPT‑5.6), and support for multiple repositories in a single project.
Create slides, sheets, docs, and Sites from your apps and workflows
To get started with ChatGPT Work, connect the tools and context where your work already happens, using plugins.
Plugins connect ChatGPT to apps and systems like Slack and Microsoft Teams, Google Drive and SharePoint, email, calendars, CRMs, project trackers, and other internal tools. ChatGPT will automatically know when to reference a plugin based on your prompt, but you can also direct ChatGPT to pull context from a specific app by typing “@” followed by the app name in your prompt. The new unified plugins directory brings plugins into one place, and ChatGPT can suggest relevant ones during your conversations.
Once your apps and tools are connected, ChatGPT can understand what you’re trying to do, pull in information from relevant sources, create documents, decks, and analyses, and keep refining drafts in the background while you stay in control.
We’re also introducing Sites in ChatGPT in public beta. With Sites, you can turn your work or ideas into an interactive site or web app and share it with your team or publicly through a URL. Sites are useful when you want to create things like live dashboards, project trackers, launch calendars, prototypes, internal portals, and interactive reports. You can test the Sites you build right inside ChatGPT and bring fresh web context into your project, too. ChatGPT can also update them as the underlying information changes.
Delegate repetitive tasks to focus on work that matters
ChatGPT Work can help take repetitive tasks off your plate, freeing you up for more impactful work.
Scheduled Tasks let you ask ChatGPT to perform an action once, repeat it on a schedule or when an event occurs, or monitor for changes over time.
Scheduled Tasks can use your connected apps and browser to:
- Review new Slack updates each week and refresh a recurring meeting agenda.
- Check websites and dashboards each morning, summarize what changed, and send a report.
- Monitor new customer feedback and turn recurring themes into prioritized product ideas.
- Update a presentation when new feedback arrives by email.
You remain in control of how ChatGPT works with you. You decide what it can access, when it should check in, and when it needs your approval before taking action. You can review progress and steer the work as priorities change.
Get work done faster across the web and your desktop apps
On desktop, ChatGPT now includes a built-in browser to help you gather information online, use web-based tools, and refine web-based work in one place.
You can ask ChatGPT to research a market, compare sources, pull information from websites, or open and refine files from Google Workspace and Microsoft 365 inside the app. It can use the browser to bring in fresh context, take steps across web pages, and keep the work moving while you review and guide the result.
On desktop, Computer Use lets ChatGPT use your computer on your behalf to execute tasks in the background across your apps, tools, and browser—clicking, typing, and moving files where they need to go. You can use it for a one-time task or as part of a Scheduled Task when recurring work includes steps on your computer.
We are also updating our Chrome extension to make it possible to use ChatGPT directly in Chrome’s sidebar. These capabilities build on what we learned from Atlas and from the users who helped us understand how agentic tools can make browser-based work more useful. We’ll begin sunsetting the standalone Atlas browser, and will share information with users about how to transition to ChatGPT.
From goals to real outcomes across every team
ChatGPT Work is powerful out of the box for all kinds of work: it can support complex workflows using the apps, files, and tools each team chooses to connect. Below are real use cases we’ve seen across teams:
Sales:
Sellers can keep a live command center current as account activity changes, without rebuilding account plans by hand. ChatGPT Work can synthesize new signals, update next steps, and help sellers spend more time moving deals forward.
[Example prompt: Create an automation that monitors for new account activity and updates this site every day at 8am]
Security and governance for organizations
Your organization remains in control even as ChatGPT takes on more substantive work for your teams. ChatGPT is built on the security, privacy, compliance, and workspace management foundation of ChatGPT Enterprise. Enterprise and Edu admins can centrally manage who has access, what company context ChatGPT can use, which tools it can connect to, and what actions it can take. The Compliance API provides visibility into ChatGPT Work conversations and actions at scale to support enterprise oversight.
Controls are tailored to each environment. On web, admins can manage access to plugins and connected tools, configure browser use and network access for cloud environments, and restrict sensitive actions in connected systems. On desktop, ChatGPT Work builds on Codex’s enterprise governance model and admin controls, bringing enterprise safeguards to work involving local files, apps, browsers, and tools—including policies for managing agent network access.
Auto-review adds another layer of protection by using our most advanced models to review important actions involving connected tools and APIs before they happen, helping prevent unauthorized sharing of sensitive information. During adversarial red teaming, auto-review blocked 100% of attempts to extract protected data, including attacks the reviewing model had not seen during training.
Availability and pricing
ChatGPT Work starts rolling out today on web and mobile, beginning with Pro, Enterprise, and Edu users and expanding to Plus and Business users over the next few days. The updated ChatGPT desktop app is available globally today for Mac and Windows, with Chat, Work, and Codex available to users on every plan, including Free.
If you already use the Codex app, you can update it as usual and it will become the new ChatGPT desktop app. Developers can make Codex the default view when they open the desktop app and choose the Codex logo as the app icon. Desktop Codex projects remain accessible on the go through the ChatGPT mobile app. The existing version of the ChatGPT desktop app will be renamed ChatGPT Classic.
ChatGPT Work is designed for longer, more involved work than a typical chat request, so usage works differently. Usage varies with the amount of work required, and more complex tasks may use more of your plan’s included usage. ChatGPT Work follows the same usage structure as Codex.
ChatGPT Enterprise and Edu admins can also set spend controls in the Admin Console to manage ChatGPT Work usage as adoption grows. Admins can support high-impact work across teams without raising limits broadly by setting workspace-level defaults, configuring group limits, creating individual overrides for people who need more capacity, and reviewing requests for additional credits with user-submitted project details and rationale.
What’s next
This is the first step towards a broader vision for ChatGPT—where intelligence goes beyond answering questions to helping everyone turn their biggest ideas into reality.
Original source Similar to OpenAI with recent updates:
- ChatGPT updates189 release notes · Latest Jul 9, 2026
- Claude updates114 release notes · Latest Jul 9, 2026
- Codex updates197 release notes · Latest Jul 9, 2026
- OpenAI Models updates48 release notes · Latest Jul 9, 2026
- Anthropic updates52 release notes · Latest Jul 9, 2026
- Claude Code updates391 release notes · Latest Jul 11, 2026
- Jul 9, 2026
- Date parsed from source:Jul 9, 2026
- First seen by Releasebot:Jul 9, 2026
Introducing GPT-5.6 series: Sol, Terra and Luna. Coming July 9
OpenAI previews GPT-5.6 Sol, a new flagship model for developers and enterprises, alongside Terra and Luna. It brings stronger frontier reasoning, long-horizon agentic work, new max reasoning effort, and ultra mode for faster complex tasks, with public launch planned for July 9.
UPDATE
GPT-5.6 Sol, along with Terra and Luna, will launch publicly on Thursday, July 9.
Today OpenAI is previewing GPT-5.6 Sol, its newest flagship model for developers and enterprises, alongside Terra and Luna.
Sol is built for frontier reasoning and long-horizon agentic work; Terra is a balanced everyday model with GPT-5.5-competitive performance at 2x lower cost; and Luna is the fastest, most affordable member of the family.
Breaking new ground
GPT-5.6 Sol advances coding, scientific reasoning, long-horizon planning, and agentic workflows, while improving reliability and efficiency across demanding real-world tasks. Sol establishes new high-water marks across some of OpenAI’s most challenging evaluations:
- Coding: GPT-5.6 Sol sets a new SOTA on Terminal-Bench 2.1, which tests command-line workflows requiring planning, iteration, and tool coordination.
- Cybersecurity: GPT-5.6 Sol shifts the efficiency frontier for vulnerability research and controlled exploitation tasks, achieving competitive results on ExploitBench while using roughly one-third of the output tokens compared with another leading frontier system.
- Biology: On SecureBio evaluations, GPT-5.6 reached top reported scores including 53.5% on the Virology Capabilities Test, 60.0% on Molecular Biology, 68.4% on Human Pathogen Capabilities, and 68.3% on World-Class Bio, about 9 percentage points above GPT-5.5.
These gains are especially important for developers building agents that need to reason over large codebases, operate tools, debug multi-step failures, conduct research, or assist with defensive security work.
What’s new
GPT-5.6 introduces a new max reasoning effort, giving Sol more time to reason deeply on difficult tasks. It also introduces ultra mode, which goes beyond a single-agent setup by using subagents to accelerate complex work.
Availability
During the preview, GPT‑5.6 models will initially be available through the API and Codex to a select group of trusted partners and organizations. OpenAI plans to make them more broadly available to people using ChatGPT, Codex, and the API soon.
Learn more
- Read the full announcement: Previewing GPT‑5.6 Sol: a next-generation model
- Read the GPT-5.6 Preview System Card: GPT-5.6 Preview System Card - OpenAI Deployment Safety Hub
- Jul 8, 2026
- Date parsed from source:Jul 8, 2026
- First seen by Releasebot:Jul 9, 2026
GPT-Live System Card
OpenAI introduces GPT-Live-1 and GPT-Live-1 mini, new full-duplex voice models that make AI conversations feel more natural and intelligent. The launch brings default voice models for paid and free users, plus stronger voice safety safeguards and broad red-teaming results.
GPT-Live-1 and GPT-Live-1 mini are a new generation of voice models designed to make conversations with AI feel more natural and intelligent.
These models — enabled by research advances — are full-duplex, meaning they can listen and respond continuously instead of waiting for a clearly defined turn to end. This means they can follow pauses, interruptions, and changes in pace, and decide in the moment whether to respond or keep listening.
GPT-Live-1 will be the default voice model for paid users, while GPT-Live-1 mini will be the default model for free users.
The most important things to know about our safety work for this launch are that:
- We trained these new models to respond safely, using the same infrastructure we rely on for training our flagship models. They can also delegate more complex work to our other models, and when they do, the resulting work will reflect the safety training of the underlying model that is doing that work. The evaluations in this card describe how the new GPT-Live-1 models perform with delegation, matching the deployment context.
- These models have system-level safety integrations that are designed to be on par with the existing safety stack for text models, while adopting some new safeguards specifically for the new voice modality: inputs and generated outputs are checked as the conversation unfolds; when potentially unsafe content is detected, the system can steer or interrupt the response, play a spoken safety message, provide support resources in text, or, in higher-risk cases, end the voice conversation.
- As part of our safety work for this launch, we built new evaluations that focus specifically on the distinctive ways that people use voice models, as opposed to text-based chats, and on observations from real-world use of the existing Advanced Voice Mode. We report those results below.
- The same monitoring, review, and enforcement infrastructure that we use for our text models also applies to the GPT-Live models, enabling us to measure prevalence, detect abuse, and enforce our safety policies.
Like OpenAI’s other models, GPT-Live-1 and GPT-Live-1 mini were trained on diverse datasets, including information that is publicly available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate. Our data processing pipeline includes rigorous filtering to maintain data quality and mitigate potential risks. We use advanced data filtering processes to reduce personal information from training data. We also employ safety classifiers to help prevent or reduce the use of harmful or sensitive content, including explicit materials such as sexual content involving a minor.
Note that comparison values from previously launched models are from the latest versions of those models, so may vary slightly from values published at launch for those models.
GPT-Live-1 and GPT-Live-1 mini are intended to be used in accordance with OpenAI’s Usage Policies, Service Terms, and Terms of Use. These policies apply universally to OpenAI services and are designed to ensure safe and responsible usage of AI technology. You can review OpenAI’s Usage Policies at openai.com/policies/usage-policies/.
If you need assistance with respect to GPT-Live-1 and GPT-Live-1 mini, you can find further information on OpenAI’s website (openai.com), or you can contact OpenAI Support by opening the chat bubble icon displayed at the bottom-right of help.openai.com.
A list of the languages that ChatGPT currently supports can be found here.
Model Safety
3.1 Voice-Native Evaluations for Disallowed Content
3.1.1 Voice-Native Evaluations: Production PromptsIn these newly developed evaluations, we use real audio examples from users who have chosen to share their voice interactions to help improve our models. Before these examples are used, they are processed through our privacy and eligibility safeguards, including checks for user permissions and deletion/opt-out status, filtering of ineligible data, and steps to reduce personal information through PII scrubbing and de-identification. We then transcribe the audio, generate the model’s response, and evaluate that response for safety.
We compare the new GPT-Live models to their respective predecessors, the models that power Advanced Voice Mode (AVM).
These evaluations are not prevalence weighted, meaning they do not reflect rates of safety performance we see in real usage. Instead, these evaluations were meant to be difficult. For each category below, the evaluation is built around cases in which the existing AVM models were not yet giving ideal responses.
Table 1: Voice-Native Evaluations: Production Prompts
Category | AVM | GPT-Live-1 | AVM mini | GPT-Live-1 mini
Sexual | 0.96 | 0.97 | 0.97 | 0.95
Illicit behavior | 0.74 | 0.97 | 0.60 | 0.94
Mental health | 0.90 | 0.90 | 0.78 | 0.84
Personal data | 0.96 | 0.96 | 0.95 | 0.95
Emotional reliance | 0.88 | 0.82 | 0.78 | 0.78
Self-harm | 0.89 | 0.96 | 0.81 | 0.92We observe that the GPT-Live models generally provide equal or better safety performance across these adversarially selected prompts than AVM models. GPT-Live-1 shows a slight regression on emotional reliance from 0.88 to 0.82, and GPT-Live-1 mini shows a slight regression on sexual content from 0.97 to 0.95. Note that neither of these are statistically significant.
3.1.2 Voice-Native Evaluations: Synthetic PromptsThe below evaluations are similar to the ones above, except that for these we synthetically generate audio prompts to target challenging edge cases. The text for these prompts is generated from safety policies and related guidance, targeting specific safety categories, policy boundaries, and difficult cases. This approach allows us to deliberately cover a broad range of scenarios, including rare or hard-to-sample safety-relevant situations. We then convert the prompts into speech and use them as audio inputs, allowing us to assess whether models apply the intended safety behavior when safety-relevant content is presented in spoken form. These evaluations focus more intensively on the areas where we believe our existing models are least likely to give ideal responses. As a result, the numbers in the table below are likewise not a guide to safety performance across all production traffic.
Table 2: Voice-Native Evaluations: Synthetic Prompts
Category | AVM | GPT-Live-1 | AVM mini | GPT-Live-1 mini
Sexual | 0.76 | 0.97 | 0.85 | 0.95
Illicit behavior | 0.63 | 0.97 | 0.63 | 0.97
Mental health | 0.57 | 0.84 | 0.47 | 0.81
Personal data | 0.90 | 0.97 | 0.79 | 0.97
Emotional reliance | 0.72 | 0.91 | 0.72 | 0.89
Self-harm | 0.72 | 0.98 | 0.71 | 0.96
Hate | 0.87 | 1.00 | 0.82 | 0.98
Gore | 0.61 | 0.97 | 0.80 | 0.96We observe that the GPT-Live models uniformly provide equal or better safety performance across these synthetic, adversarially selected prompts than AVM models.
The synthetic evaluation is designed to assess different risk surfaces than the production evaluation, so performance may vary between them. The synthetic set tests targeted, policy-grounded adversarial scenarios, including rare cases that may be underrepresented in production data. Strong performance on this set indicates that safety training is transferring for clear, intentionally constructed risks, but may not translate to production behavior. The production set reflects real-world user behavior and often includes more ambiguous or borderline context, longer interaction histories, and persistent attempts to steer the model toward unsafe outputs. Failures in the production set can be subtler and less severe, but still important. We therefore view the two benchmarks as complementary.
Red Teaming
OpenAI worked with a team of internal and external red teamers across languages to stress test the models’ safety training with no system level mitigations. We started very early in the process to baseline performance with no model safety training, and ultimately completed two additional rounds of testing as we continued to improve safeguards. Red teaming spanned multiple categories including child-coded voice, impersonation, speaker identification, sensitive train identification, self-harm, emotional reliance, scams and manipulation, and audio-specific perturbations. Early findings were used to prioritize risk areas to focus mitigation efforts on (e.g. sexual content, emotional reliance, and self harm) while validating some other areas that were policy-compliant by default (e.g. voice cloning and impersonations). Consequent follow-up rounds validated that the mitigations we built to address the identified issues are robust.
Preparedness Framework
The Safety Advisory Group reviewed this launch and determined that neither GPT-Live-1 nor GPT-Live-1 mini, when operating without delegation, could plausibly be considered High in any of our Preparedness Framework’s Tracked Categories – Biological and Chemical Risk, AI Self-Improvement, or Cybersecurity.
- For the Biological and Chemical domain, where GPT-Live models can use our highly capable flagship models, the GPT-Live experience inherits the safeguards of those underlying models. In addition, we’ve built automated monitors that may interrupt and end the call when potentially harmful conversations are detected, or degrade user experience for repeated abuse. We also take actor level enforcement actions where necessary.
- For the Cybersecurity domain, delegated work will likewise receive the safeguards associated with the model to which work is delegated. In addition, cybersecurity risk from the GPT-Live models themselves is highly constrained at launch because these models lack broad access to tools independently of the models to which they delegate, and do not have code execution capability. We will reassess the cybersecurity safeguards posture before enabling additional tools.
- AI Self-Improvement capability evals were not run, as GPT-Live-1 and GPT-Live-1 mini are less capable than GPT-5.5 Thinking across several intelligence evaluations.
- Jul 8, 2026
- Date parsed from source:Jul 8, 2026
- First seen by Releasebot:Jul 8, 2026
Introducing GPT‑Live
OpenAI launches GPT-Live, a new ChatGPT Voice experience with fuller-duplex conversation, smarter answers, better listening, and visual cards. It rolls out globally to ChatGPT users, with GPT-Live-1 and GPT-Live-1 mini powering voice across plans.
Entering a new era of human-AI interaction
We’re launching GPT‑Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.
GPT‑Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT‑Live can show it’s paying attention with phrases like “mhmm” or “yeah”, engage in quick back-and-forth, or just stay quiet when you need a moment to think. The result is a voice experience that is refreshingly easy to talk to.
GPT‑Live is also our smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the result back into the conversation when it’s ready. While it works, GPT‑Live can keep talking with you and maintain the flow of conversation. At launch, GPT‑Live will use GPT‑5.5 in the background. As we release new frontier models, we’ll continuously update the model used by GPT‑Live.
These advances power a new ChatGPT Voice experience that is more intelligent and natural to use. Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work.
We’re beginning to roll out two versions of GPT‑Live – GPT‑Live‑1 and GPT‑Live‑1 mini – to ChatGPT users globally today. We also plan to bring them to the API soon, and developers and enterprises can sign up to be notified using this form.
Previous approaches
Older generations of voice AI systems brought us closer to that vision, but with important tradeoffs.
Cascaded voice systems
Cascaded voice systems rely on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe your speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled us to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted.
Turn-based voice models
Turn-based voice models like ChatGPT Advanced Voice Mode processed and generated audio within a single model, reducing latency and making conversations smoother — but they still operated through discrete turns. The model had to wait for the user to stop speaking before responding, resulting in rigid back-and-forth. In addition, because turn detection is based on silence, even a brief pause or background noise could be mistaken for the end of turn — causing the model to interrupt at unnatural times.
Our new approach
GPT‑Live addresses these limitations through two architectural changes.
Continuous interaction
First, we built GPT‑Live for continuous interaction using a full-duplex architecture. Instead of processing a sequence of separate messages, GPT‑Live continuously processes input while generating output. The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool.
This allows the model to engage in more natural back-and-forth, maintain a better sense of time, and even perform live translation.
Delegation for deeper work
Second, we decoupled GPT‑Live — which handles continuous interaction — from deeper work. When a question requires search, reasoning, or more agentic capabilities, GPT‑Live can delegate the task to another model like GPT‑5.5. This allows it to keep the conversation going, even as it handles multiple tasks in the background.
This architectural change also allows GPT‑Live to continuously use the latest models and agents, combining frontier intelligence with natural interaction.
Evaluations
We built new human evaluations to measure pleasantness and the flow of conversation. In these head-to-head comparisons, GPT‑Live‑1 and GPT‑Live‑1 mini are strongly preferred over Advanced Voice Mode in matched 5–10 minute conversations that measure overall preference, turn-taking, interruptions, conversational flow, and how natural each interaction felt.
- GPQA: GPT‑Live‑1 substantially outperforms Advanced Voice Mode on GPQA, which tests expert-level scientific reasoning across biology, chemistry, and physics.
- BrowseComp: GPT‑Live‑1 shows strong gains over Advanced Voice Mode on BrowseComp, which tests agentic web search and the ability to find difficult-to-locate information.
- τ³-Voice Telecom (internal variant)**: GPT‑Live‑1 outperforms Advanced Voice Mode on τ³-Voice Telecom, which tests voice agents on realistic, multi-turn telecom support tasks.
A new ChatGPT Voice experience
Each week, more than 150 million people talk to ChatGPT using features like Voice and Dictation. They use it to get hands-free everyday help, to practice languages, tell bedtime stories, or just chat during their commute.
Starting today, when you tap the Voice button to talk with ChatGPT, you’ll get an improved experience powered by GPT‑Live—with more natural conversations, smarter answers, better listening, and visual responses.
More natural conversations
Talking with ChatGPT should now feel much more like a real conversation. You can interrupt with a question, pause to gather your thoughts, or ask ChatGPT to slow down. It naturally acknowledges what you’re saying with phrases like “mhmm” or “got it,” so you know it’s following along. We’ve also remastered the nine distinct voices in ChatGPT for GPT‑Live.
Smarter answers
ChatGPT Voice can now draw on our latest frontier models, giving you smarter answers when you need them. You can also choose the level of reasoning that fits your needs: Instant for fast responses, or Medium and High when you want ChatGPT to spend more time thinking.
Better listening
If you take a moment to think, ChatGPT Voice now waits instead of jumping in and interrupting. If you ask it to stay quiet and listen, it will. And when there’s background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on your voice instead of getting distracted.
Visual answers at a glance
Some answers are more useful when you can see them. While you’re talking, ChatGPT can now show rich visual cards for topics like weather, stocks, sports, and more. Voice also continues to support search, memory, images, and file uploads.
The result is a ChatGPT Voice experience that feels more natural, more capable, and more useful in everyday life.
Safety designed for voice
GPT‑Live was designed to be safe by default. It builds on the safety advances from our latest models while adding dedicated safety training across key risk areas and new safeguards designed specifically for voice.
Expanded safety testing
To better reflect how people use voice in real-life settings, we began by expanding our safety testing to include new audio-native evaluations. We also created synthetic evaluations that use generated audio to focus more intensively on key safety areas, drawing on what we learned from Advanced Voice Mode. Those areas include self-harm, psychosis and mania, emotional reliance on AI, violence, and sexual content. Internal experts also red-teamed the model for risks unique to voice.
In our testing, GPT‑Live performed comparably to or better than Advanced Voice Mode across nearly all of the areas we evaluated. You can read more about our testing and safeguards in the GPT‑Live system card (opens in a new window).
Built-in safeguards
Because voice conversations unfold in real time, we also built safeguards that can act while the model is speaking. When the system detects potentially unsafe output, it can steer the model toward a safer response, surface additional safety messaging or resources, or end the voice conversation in higher-risk cases. For conversations involving self-harm, we adapted ChatGPT’s support flows for voice, including offering expert-vetted crisis helpline support.
We designed additional protections to support teen users, and trained age-appropriate behavior directly into the model to reduce the risk of inappropriate responses. Parents can choose whether their teen can use ChatGPT Voice through Parental Controls, and linked parents may be notified in higher-risk situations involving signs of potential self-harm or suicidal intent.
Learning from real-world use
We’re also rolling out longer-term measurement and post-launch monitoring focused on emotional reliance to continue improving our understanding and refining safeguards. Building on our previous research into affective use and emotional well-being, this will help us identify emerging patterns and improve how the system responds in emotionally sensitive interactions.
Finally, GPT‑Live is designed for conversation, not voice impersonation. It uses a set of predefined voices in ChatGPT, with safeguards to prevent it from imitating a real person’s voice.
We’re committed to supporting safety and well-being and will keep strengthening these protections as we learn from real-world use.
Availability & limitations
GPT‑Live is rolling out now to ChatGPT users globally across iOS, Android, and ChatGPT.com (opens in a new window). GPT‑Live‑1 will become the default model powering ChatGPT Voice for Go, Plus, and Pro users, and GPT‑Live‑1 mini will become the default for Free users. More availability details can be found in our Help Center (opens in a new window).
We’ve optimized GPT‑Live for some of the most popular languages in ChatGPT. For certain languages, the model may have a non-native accent or gaps in fluency. We’re actively working to improve the experience across languages.
At launch, GPT‑Live will not support voice with video or screen sharing in ChatGPT, but we’re working to introduce these capabilities soon. You can still access legacy versions of ChatGPT Voice, including Standard and Advanced Voice Mode, where these features are available.
Original source - Jul 6, 2026
- Date parsed from source:Jul 6, 2026
- First seen by Releasebot:Jul 7, 2026
New Realtime models on the API: gpt-realtime-2.1 and gpt-realtime-2.1-mini
OpenAI releases gpt-realtime-2.1 and gpt-realtime-2.1-mini for low-latency voice and multimodal experiences, with at least 25% lower p95 latency across Realtime voice models, improved recognition and noise handling, and stronger reasoning, tool use, and instruction following.
We’ve released two new Realtime models for building low-latency voice and multimodal experiences: gpt-realtime-2.1 and gpt-realtime-2.1-mini.
With this release, we’ve also reduced p95 latency by at least 25% across Realtime voice models through improved caching.
gpt-realtime-2.1 updates GPT-Realtime-2 with improved alphanumeric recognition, silence and noise handling, and interruption behavior. It supports speech-to-speech interactions with configurable reasoning effort, instruction following, and tool use for more complex voice-agent workflows.
gpt-realtime-2.1-mini is a mini reasoning model for faster, lower-cost realtime voice interactions.
A quick way to choose:
- Use gpt-realtime-2.1 when you want the strongest realtime reasoning, tool use, instruction following, and voice-agent behavior.
- Use gpt-realtime-2.1-mini when you want a faster, more cost-efficient option for realtime voice experiences.
Pricing
Model Text input Text cached input Text output Audio input Audio cached input Audio output Image input Image cached input
gpt-realtime-2.1 $4.00 $0.40 $24.00 $32.00 $0.40 $64.00 $5.00 $0.50
gpt-realtime-2.1-mini $0.60 $0.06 $2.40 $10.00 $0.30 $20.00 $0.80 $0.08Try them in the Playground and let us know what you build with these models and how they perform in your realtime voice workflows. Share your feedback, questions, and examples in the thread.
Original source - Jun 30, 2026
- Date parsed from source:Jun 30, 2026
- First seen by Releasebot:Jul 1, 2026
Introducing GeneBench-Pro
OpenAI introduces GeneBench-Pro, a research-level benchmark for judging AI agents in computational biology. It expands GeneBench with harder, more realistic synthetic tasks, open-sources representative questions, and reports strong model results on scientific reasoning under uncertainty.
A research-level benchmark measuring how AI agents navigate ambiguity and make consequential judgments in computational biology.
Scientific data rarely arrive with instructions. Researchers must decide whether a pattern reflects biology or noise, whether the data can support the question being asked, and how each result should change what they do next. AI agents are increasingly capable of executing complex analyses, but real scientific research also depends not simply on recalling facts or following a predefined workflow but also on making these higher-order judgments.
Today, we’re introducing GeneBench-Pro—a challenging, research-level benchmark for testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires. It expands on GeneBench to cover harder, more realistic tasks across genomics, quantitative biology, and translational medicine, capturing the complexity, iterative nature, and ambiguity of scientific research in computational biology.
To date, there have been few convincing assessments of the system-level judgment calls that make real-world computational research difficult. These include handling ambiguity, revising assumptions, choosing the correct analysis path, and knowing when a result is decision-ready. Because these skills are difficult to formalize, they are also difficult to assess rigorously, even as weaknesses in them increasingly constrain overall AI performance.
GeneBench-Pro is designed to precisely measure these higher-level capabilities. Within GeneBench-Pro, we define “research taste” as the chains of judgment calls that shape an analysis: which questions the data can support, how early diagnostics should change the model or estimand, and when an initial plan needs to be revised. Each GeneBench-Pro problem gives the model a realistic and messy dataset, brief experimental context, and a target estimand tied to a downstream decision. To answer correctly, the model must explore the data, choose an appropriate analytical approach, engage in an iterative process of experimentation, and supply a final answer.
Dataset construction
In biology, the cost of data generation (e.g., genome sequencing) has fallen dramatically, and some researchers now argue that the limiting factor is no longer sample collection but downstream computation and analysis. GeneBench-Pro is built to assess progress in addressing that bottleneck, with 129 questions covering a broad range of computational biology settings and methods.
GeneBench-Pro is also designed to avoid common benchmark failures. Many long-horizon biology benchmarks construct multi-step questions around messy historical datasets, where there may be no single correct path through the analysis. An agent might choose one defensible cutoff, while another might choose a different but equally defensible option, reflecting the arbitrary choices made by the benchmark creator more than any fundamental differences in model performance. The reverse can also happen: if a problem is too numerically insensitive, an agent can make fundamental errors in an analysis and still produce a passing result.
To avoid these failure modes, each GeneBench-Pro problem is built synthetically: we know the full causal structure and directly simulate the data-generating process. That enables us to tune the complexity of each problem, ensure that reasonable differences in subjective analytical choices still produce accepted numerical results, and verify (through ablation studies) that plausible but incorrect analyses fail. We then audit problem drafts through detailed trace analyses to check for information leakage and unintended solution pathways. This gives us confidence that getting the right answer depends on choosing the correct analytic pathway and not on exploiting a shortcut or matching an arbitrary author preference.
We sent 82 of the 129 GeneBench-Pro questions to external domain experts, including graduate students, postdoctoral researchers, industry scientists, and professors. Reviewers assessed each problem’s realism, whether the target answer was identifiable, and whether the methods and estimators were appropriate. Feedback was used to improve problems.
Evaluation and grading
Each GeneBench-Pro problem is a self-contained scientific analysis. Agents receive access to an isolated workspace with a short prompt, data files, and a standard bioinformatics stack including Python, scientific computing libraries, and basic genomics packages like PLINK 2.0 (although the problems do not require domain-specific tooling).
Because we control the full data-generation process, we can grade correctness deterministically against known targets, avoiding model-choice variability and verbosity effects found in standard rubric-based evaluation.
Each problem also comes with rich metadata, including the intended analysis structure, attached data files, a detailed multi-page case study, and expert review outcomes. We are fully open-sourcing 10 representative GeneBench-Pro questions on Hugging Face, with an interactive web interface for browsing them. Finally, we will provide a 50-question subset to Artificial Analysis for independent, third-party benchmarking in the near future.
Results
Our strongest model, GPT‑5.6 Sol, attains a pass rate of 28.7% at the highest reasoning level (31.5% with Pro mode enabled). That is a sharp increase from when we began building the original GeneBench; at that time, our best frontier model, GPT‑5, scored below 5%. Progress on this benchmark suggests that frontier models are improving quickly, even on less tangible, systems-level scientific reasoning. At the current pace, this benchmark may be saturated by the end of the year.
The results also show the impact of scaling test-time compute. At the lowest reasoning level, GPT‑5.6 Sol only achieves a single-digit passrate. At the highest reasoning level, GPT‑5.6 Sol solves nearly six times as many questions as GPT‑5.2 does while using about two-thirds as many tokens.
Comparisons across model families suggest that GPT models are among the strongest systems at high-level scientific reasoning under quantitative uncertainty. The performance gap between GPT‑5.6, GPT‑5.5 and leading open-source models such as GLM 5.2 is significantly larger than we would expect when extrapolating from coding benchmarks, indicating that open-source models are more specialized for coding than for broader reasoning ability.
We used frontier GPT models to evaluate and harden problems during development. As such, we suspected GeneBench-Pro might be biased against GPT models relative to other model families. However, competitor models at best matched the performance of the corresponding GPT model at the time of release, and tended to fall short considerably.
These evaluation results—as high as 31.5% on GPT‑5.6 Sol (Pro)—are striking given the difficulty of the GeneBench-Pro questions. In a survey, our reviewers estimated that a typical GeneBench-Pro problem would take a human expert around 20–40 hours to complete. At a conservative $200 per hour, that puts the human labor cost of a single problem in the thousands of dollars. Current AI agents are still too unreliable to replace human experts, but the cost gap is large, with inference costs at only several dollars per problem. That means even partial automation at current capabilities could create meaningful economic and scientific value.
Still, the fact that frontier models still solve fewer than a third of these problems shows that there is substantial room for improvement. Models can make partial progress on challenging problems, but they struggle to close the inferential loop. This failure pattern mirrors the contrast between human experts and novices. Experts use their experience to frame the problem and adapt their approach, while novices make observations but struggle to integrate them into the broader context of the problem.
Achieving near-perfect performance will require evaluations that both reliably measure progress and identify where models still fail. Benchmarks like GeneBench-Pro can help to turn a vague capability deficiency into something we can diagnose and improve.
If agents can reliably automate this class of analysis, they could significantly accelerate scientific discovery. Human genetic evidence is already central to target prioritization and translational follow-up, because mechanisms with genetic support are much more likely to lead to approved treatments.
Meanwhile, sequencing costs have plummeted, and biobank-scale datasets now link molecular, phenotypic, and health-record information at unprecedented breadth. The limiting factor is shifting from data generation to turning the information into actionable insights. Models that can consistently perform analyses now handled by teams of human experts could transform industrial research by accelerating hypothesis triage, target follow-up, and the iteration cycle between data generation and decision-making.
GeneBench-Pro represents an initial effort to evaluate the more abstract skills involved in good scientific judgment possessed by experienced. These skills allow them to intuit and identify the most promising initial analyses, iterate and revise their thinking when data contradict initial assumptions, and arrive at conclusions upon which downstream clinical, academic, or business decisions may depend.
We anticipate that as model capabilities advance, benchmarks that probe model abilities at these higher levels of abstraction will become increasingly useful, beyond those that simply test book knowledge or the ability to execute routine analyses.
Original source - Jun 26, 2026
- Date parsed from source:Jun 26, 2026
- First seen by Releasebot:Jul 10, 2026
Introducing GPT-5.6 series: Sol, Terra and Luna. Coming July 9 10am PT
OpenAI previews GPT-5.6 Sol with Terra and Luna, bringing a new flagship model family for developers and enterprises. Sol targets frontier reasoning and long-horizon agentic work, while Terra and Luna emphasize lower cost and speed. New max reasoning effort and ultra mode are also introduced.
UPDATE
GPT-5.6 Sol, along with Terra and Luna, will launch publicly on Thursday, July 9.
Today OpenAI is previewing GPT-5.6 Sol, its newest flagship model for developers and enterprises, alongside Terra and Luna.
Sol is built for frontier reasoning and long-horizon agentic work; Terra is a balanced everyday model with GPT-5.5-competitive performance at 2x lower cost; and Luna is the fastest, most affordable member of the family.
Breaking new ground
GPT-5.6 Sol advances coding, scientific reasoning, long-horizon planning, and agentic workflows, while improving reliability and efficiency across demanding real-world tasks. Sol establishes new high-water marks across some of OpenAI’s most challenging evaluations:
- Coding: GPT-5.6 Sol sets a new SOTA on Terminal-Bench 2.1, which tests command-line workflows requiring planning, iteration, and tool coordination.
- Cybersecurity: GPT-5.6 Sol shifts the efficiency frontier for vulnerability research and controlled exploitation tasks, achieving competitive results on ExploitBench while using roughly one-third of the output tokens compared with another leading frontier system.
- Biology: On SecureBio evaluations, GPT-5.6 reached top reported scores including 53.5% on the Virology Capabilities Test, 60.0% on Molecular Biology, 68.4% on Human Pathogen Capabilities, and 68.3% on World-Class Bio, about 9 percentage points above GPT-5.5.
These gains are especially important for developers building agents that need to reason over large codebases, operate tools, debug multi-step failures, conduct research, or assist with defensive security work.
What’s new
GPT-5.6 introduces a new max reasoning effort, giving Sol more time to reason deeply on difficult tasks. It also introduces ultra mode, which goes beyond a single-agent setup by using subagents to accelerate complex work.
Availability
During the preview, GPT‑5.6 models will initially be available through the API and Codex to a select group of trusted partners and organizations. OpenAI plans to make them more broadly available to people using ChatGPT, Codex, and the API soon.
Learn more
- Read the full announcement: Previewing GPT‑5.6 Sol: a next-generation model
- Read the GPT-5.6 Preview System Card: GPT-5.6 Preview System Card - OpenAI Deployment Safety Hub
- Jun 26, 2026
- Date parsed from source:Jun 26, 2026
- First seen by Releasebot:Jun 27, 2026
Introducing GPT-5.6 series: Sol, Terra and Luna
OpenAI previews GPT-5.6 Sol, a new flagship model for developers and enterprises, with Terra and Luna joining the family. The release highlights stronger coding, scientific reasoning, long-horizon planning, and agentic workflows, plus new max reasoning effort and ultra mode for complex work.
Today OpenAI is previewing GPT-5.6 Sol, its newest flagship model for developers and enterprises, alongside Terra and Luna.
Sol is built for frontier reasoning and long-horizon agentic work; Terra is a balanced everyday model with GPT-5.5-competitive performance at 2x lower cost; and Luna is the fastest, most affordable member of the family.
Breaking new ground
GPT-5.6 Sol advances coding, scientific reasoning, long-horizon planning, and agentic workflows, while improving reliability and efficiency across demanding real-world tasks. Sol establishes new high-water marks across some of OpenAI’s most challenging evaluations:
- Coding: GPT-5.6 Sol sets a new SOTA on Terminal-Bench 2.1, which tests command-line workflows requiring planning, iteration, and tool coordination.
- Cybersecurity: GPT-5.6 Sol shifts the efficiency frontier for vulnerability research and controlled exploitation tasks, achieving competitive results on ExploitBench while using roughly one-third of the output tokens compared with another leading frontier system.
- Biology: On SecureBio evaluations, GPT-5.6 reached top reported scores including 53.5% on the Virology Capabilities Test, 60.0% on Molecular Biology, 68.4% on Human Pathogen Capabilities, and 68.3% on World-Class Bio, about 9 percentage points above GPT-5.5.
These gains are especially important for developers building agents that need to reason over large codebases, operate tools, debug multi-step failures, conduct research, or assist with defensive security work.
What’s new
GPT-5.6 introduces a new max reasoning effort, giving Sol more time to reason deeply on difficult tasks. It also introduces ultra mode, which goes beyond a single-agent setup by using subagents to accelerate complex work.
Availability
During the preview, GPT‑5.6 models will initially be available through the API and Codex to a select group of trusted partners and organizations. OpenAI plans to make them more broadly available to people using ChatGPT, Codex, and the API soon.
Learn more
- Read the full announcement: Previewing GPT‑5.6 Sol: a next-generation model
- Read the GPT-5.6 Preview System Card: GPT-5.6 Preview System Card - OpenAI Deployment Safety Hub
- Jun 26, 2026
- Date parsed from source:Jun 26, 2026
- First seen by Releasebot:Jun 26, 2026
Previewing GPT-5.6 Sol: a next-generation model
OpenAI begins a limited preview of GPT-5.6, introducing Sol, Terra and Luna with stronger reasoning, coding, biology and cybersecurity performance, a new ultra mode, tougher safety protections, updated pricing, and more predictable prompt caching ahead of broader availability.
Capabilities
We're beginning a limited preview of the GPT‑5.6 series: Sol, our flagship model; Terra, a balanced model for everyday work; and Luna, a fast and affordable model. Terra has competitive performance to GPT‑5.5 while being 2x cheaper and Luna brings strong capability at our lowest cost.
GPT‑5.6 Sol launches with our most robust safety stack to date. We strengthened protections for higher-risk activity, sensitive cyber requests, and repeated misuse, and spent multiple weeks finding weaknesses, pressure-testing our system, and hardening it against real-world attacks.
We believe in broad access, and we plan to make GPT‑5.6 Sol, Terra, and Luna generally available in the coming weeks. As part of our ongoing engagement with the U.S. government, we previewed our plans and the models’ capabilities ahead of today’s launch. At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly. During this preview, we will continue testing and coordinating closely with partners as we work toward broader availability. We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them. We are taking this short-term step because we believe it is the strongest path to broader availability in the coming weeks, while we work with the Administration to develop the cyber Executive Order framework and a repeatable process for future model releases.
GPT‑5.6 Sol is our strongest model yet. To give a preview of model performance, we share a set of evaluations highlighting improved agentic capabilities in coding, biology, and cybersecurity, with additional safety and preparedness evaluations available in our system card (opens in a new window). We will share an expanded suite of evaluation results when we make the model broadly available.
With GPT‑5.6, we’re introducing a new max reasoning effort to give Sol the most time to reason deeply. Additionally, we’re introducing a new ultra mode that goes beyond the capabilities of a single agent by leveraging subagents to accelerate complex work.
For coding workflows, GPT‑5.6 Sol sets a new state of the art on Terminal‑Bench 2.1, which tests command-line workflows requiring planning, iteration, and tool coordination.
GPT‑5.6 Sol also shows broad improvements in biology workflows. On GeneBench v1, which evaluates long-horizon genomics and quantitative-biology analyses, it achieves stronger results than GPT‑5.5 while using fewer tokens.
GPT‑5.6 Sol is our most capable model yet for cybersecurity. It shifts the performance-efficiency frontier for long-horizon security tasks including vulnerability research and exploitation. On ExploitBench², GPT‑5.6 Sol is competitive with Mythos Preview using only ~1/3 of the output tokens. On ExploitGym (opens in a new window)³, a benchmark created by UC Berkeley researchers in collaboration with OpenAI and other frontier labs, GPT‑5.6 Sol, Terra, and Luna models all demonstrate strong improvements in cyber capabilities as we increase reasoning.
Stronger cyber capabilities with stronger safeguards
We developed GPT‑5.6 Sol, Terra and Luna with our most robust safeguards to date, with configurations matched to each model’s capabilities. As the model becomes more capable, we design safeguards to increasingly hold up to real-world adversarial pressure while preserving access to legitimate work such as code review, vulnerability research, patch development, debugging, security education, and defensive testing. Our goal is to make prohibited offensive activity more difficult, uncertain, and detectable without unnecessarily limiting those beneficial uses. Based on our assessment of the model and safeguards, we expect substantial benefit for legitimate defensive work, while meaningfully constraining prohibited offensive use.
GPT‑5.6 Sol is better at helping people find and fix vulnerabilities than reliably carrying out end-to-end attacks. As these capabilities continue to advance, our priority is to make sure they reach and benefit defenders, who can use these tools to find weaknesses, develop patches, and strengthen systems more broadly.
GPT‑5.6 Sol does not cross the Cyber Critical threshold under our Preparedness Framework. In evaluations involving Chromium and Firefox, it identified bugs and exploitation primitives—the building blocks of an exploit—but did not autonomously produce a functional full-chain exploit under the conditions tested. Still, benchmark thresholds cannot capture every way a model may be used or combined with other tools. That uncertainty, along with the model’s broader step change in capabilities, is why we are pairing the model’s increased capabilities with stronger safeguards and a phased release. We share more details about our safeguards in the GPT‑5.6 Preview system card (opens in a new window).
A layered safeguard stack
No single safeguard is sufficient against determined or adaptive misuse. Across the GPT‑5.6 preview, we use layered safeguards, with exact configurations varying across models, and pressure-test them for real-world attacks. These include protections trained into the model, real-time checks during generation, account-level signals, differentiated access, monitoring, enforcement, and continued testing.
GPT‑5.6 is trained to refuse prohibited cyber assistance, including when users attempt to disguise their intent or jailbreak the model. These model-level safeguards establish the first boundary around what the model should and should not help with.
Real-time cyber and biology misuse classifiers provide another layer by evaluating output as it is generated. For higher risk cases, if they detect a potential violation, the generation may be paused while a larger reasoning model reviews the conversation and its context. If the output is assessed as disallowed, it is withheld before it reaches the user.
Flagged activity can also trigger account-level review across relevant conversations and risk signals, consistent with our terms and policies around content retention and review. Looking beyond a single conversation helps our systems distinguish persistent malicious behavior from legitimate dual-use security work, where similar technical concepts may appear in very different contexts.
Together, these layers make the overall approach more robust than any one safeguard on its own. Model behavior reduces the likelihood of harmful responses, real-time systems can intervene during generation, account-level review can identify broader patterns, and differentiated access preserves important defensive work without making the most sensitive capabilities broadly available by default.
Especially during the preview, users may encounter safeguards that block or refuse some requests. Other requests may take longer because generation is paused for additional review. Safeguards may occasionally intervene on legitimate work, particularly in dual-use areas where defensive and offensive activity can initially look similar.
That is part of what the preview is designed to test. We want to understand not only whether the safeguards constrain misuse, but whether legitimate users can still complete normal work reliably and efficiently. Feedback during the preview will help us reduce unnecessary blocks and delays, improve how the safeguards interpret context, and create a smoother experience before wider release.
We are also working with enterprise customers on longer-term approaches—including privacy-preserving detection, customer-operated safety controls, and access calibrated to the risk of a customer, user, or workload—to advance safety while supporting enterprise privacy requirements.
Improving robustness with automated red-teaming
Safeguards also need to remain effective when attackers adapt their tactics. A protection that works only on a fixed set of known attacks is not robust enough for a frontier model.
That’s why we are applying more intelligence and compute than ever before to safety, using our own models to find weaknesses and improve safeguards faster. We dedicated over 700,000 A100-equivalent GPU hours to automated red teaming aimed at finding universal jailbreaks: attacks that can work across many prompts or contexts, not just one narrow setting. Focusing on these harder, more general attacks let us test the safeguards beyond a fixed set of known failures. It also lets us explore far more attack patterns than human testing alone could cover, identify failure patterns earlier, and shorten the path from finding a weakness to addressing it.
In addition to automated red-teaming, we worked with third-party testers to conduct extensive human expert red teaming, which will continue in the preview period. Human red-teaming complements the automated work by testing safeguards against creative experts trying to misuse the model in ways our systems might not anticipate.
No evaluation can represent every product configuration, multi-step attack, or real-world workflow. We therefore maintain a rapid-response process to reproduce, assess, prioritize, and remediate newly discovered jailbreaks, then add them to our ongoing evaluations so we can test against similar failures in the future.
Availability and pricing
During the preview, GPT‑5.6 models will initially be available through the API and Codex to a select group of trusted partners and organizations. We plan to make them more broadly available to people using ChatGPT, Codex, and the API soon.
In this new naming system introduced with GPT‑5.6, the number identifies a model’s generation, while Sol, Terra, and Luna identify durable capability tiers that can advance on their own cadence. Together, the family gives people and developers clearer choices across intelligence, speed, and cost.
GPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life. For GPT‑5.6 and later models, cache writes are billed at 1.25x the model’s uncached input rate, while cache reads continue to receive the 90% cached-input discount.
We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.
We’re excited to continue learning from this preview period, and to bring GPT‑5.6 Sol, Terra and Luna to more people soon.
Original source - Jun 26, 2026
- Date parsed from source:Jun 26, 2026
- First seen by Releasebot:Jun 26, 2026
GPT-5.6 Preview System Card
OpenAI launches GPT-5.6 as a new model family with Sol, Terra and Luna, pairing stronger cybersecurity capability with its most robust safety stack yet. The release starts as a limited preview for trusted partners, with broader general availability planned in the coming weeks.
GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet—are built to deliver these models safely and at scale, around the world.
We believe in broad access, and we plan to make GPT-5.6 Sol, Terra, and Luna generally available in the coming weeks. As part of our ongoing engagement with the U.S. government, we previewed our plans and the models’ capabilities ahead of today’s launch. At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly. During this preview, we will continue testing and coordinating closely with partners as we work toward broader availability.
Under our Preparedness Framework, we are treating Sol, Terra and Luna as High capability in both Cybersecurity and Biological and Chemical risk. None of them reach our High threshold in AI Self-Improvement. We have implemented a tailored set of safeguards, adapted to each model’s capability profile, to sufficiently minimize the associated risks.
This system card is a detailed report of the work we did to understand and mitigate GPT-5.6’s safety risks before deployment. The five most important things to know are that:
These models are a meaningful step up in cybersecurity capability, but they do not reach our risk framework’s highest level (Critical). GPT-5.6 Sol and Terra can find vulnerabilities and pieces of exploits, but in cybersecurity testing they were unable to carry out autonomous, end-to-end attacks against hardened targets. Separate evaluations examined misaligned behavior in agentic coding tasks and found GPT-5.6 shows a greater tendency than GPT-5.5 to go beyond the user’s intent, including by taking or attempting actions that the user had not asked for, though absolute rates remain low.
To make these models safe, we added new technology to a safety stack that is more than the sum of its parts. The models are trained to be safe, Sol and Terra are served with newly added activation classifiers focused on sensitive domains that watch the model and can intervene to stop unsafe answers during generation, and certain conversations are scanned so unsafe outputs are blocked in real time if they cross safety boundaries. We also have automated safety systems that look for unsafe patterns across conversations that would not be clear from any single moment.
Severe harm requires a chain of successful steps, and our safeguards place barriers throughout that chain. Based on our threat modelling in cybersecurity and biology, we’ve designed our safety stack so that even if an attacker does complete one step on the path to harm, safeguards will still stop the model from allowing severe harm. We also have programs in place so that when GPT-5.6 models are broadly available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders.
Our safeguard testing has already been more intensive than for any earlier release, and we are continuing to test during the preview period. Expert humans and external testers used a diverse set of approaches to find gaps. We’ve also dedicated over 700,000 A100e GPU hours to automatically find universal jailbreaks, and we will run automated red teaming continuously during deployment. As jailbreaks are reported, we reproduce, mitigate and retest for them so that gaps are addressed.
Providing broad access, particularly for cybersecurity capabilities, will have important safety benefits. Our testing suggests that GPT-5.6 is better at finding and fixing cyber vulnerabilities than at exploiting those vulnerabilities in real attacks. That gives defenders an opportunity to harden systems before cybersecurity weaknesses are exploited—an opportunity that may narrow as offensive capabilities improve. Our safeguards therefore focus on making malicious use at scale harder, while still enabling the day-to-day work of securing systems.
In this card, we show how performance changes with reasoning effort—the amount of thinking a model uses to work through a problem. Rather than report a single score, we show a curve across different levels of effort. This gives a fuller picture of what the model can do and how much effort it takes to get there.
Note that we are continually iterating on our models. Comparison values from previously-launched models are from recent snapshots of those models, and may vary slightly from values published in previous cards.
We plan to publish an updated version of this system card when making the GPT-5.6 family of models generally available.
[The system card continues with detailed sections on Model Data and Training, Model Safety, Robustness Evaluations, Health, Alignment, Bias Evaluations, Preparedness, Safeguards, and References, providing comprehensive evaluation results, safety measures, threat modeling, and performance benchmarks for the GPT-5.6 family of models. It includes extensive data on cybersecurity and biological capabilities, safety training, real-time safeguards, automated red-teaming, actor-level enforcement, trust-based access programs, and security controls. The document also discusses the evaluation methodologies, results on various benchmarks, and ongoing research and monitoring efforts to ensure safe deployment and use of these advanced AI models.]
Original source - Jun 24, 2026
- Date parsed from source:Jun 24, 2026
- First seen by Releasebot:Jun 24, 2026
OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI introduces Jalapeño, its first Intelligence Processor, a custom LLM inference accelerator built with Broadcom to make AI faster, more reliable, and more accessible. Early testing points to strong performance per watt as OpenAI expands its full-stack infrastructure strategy.
OpenAI and Broadcom (NASDAQ: AVGO) today unveiled Jalapeño, OpenAI’s first Intelligence Processor: an accelerator architected around OpenAI’s vision for the future of LLM inference, and the first AI accelerator in a multi-generation compute platform the companies are building together to make advanced AI faster, more reliable, and more accessible to more people.
Jalapeño was delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom President and CEO Hock Tan and President Charlie Kawwas, marking an important step in OpenAI’s strategy to build the full stack behind its models and products.
OpenAI designed the chip from scratch around its deep understanding of LLM fundamentals, informed by its roadmap of models, kernels, serving systems, and product needs, with partners Broadcom and Celestica, helping industrialize the platform through chip implementation, board, rack system integration, high-performance networking, and scalable production systems. Jalapeño is designed with flexibility to work with all LLMs guided by OpenAI’s insights into the inference needs of current and future AI models across the industry. Engineering samples of the Jalapeño chip are running ML workloads in the lab at production target frequency and power, including GPT‑5.3‑Codex‑Spark.
While OpenAI is still measuring final performance, early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art. A detailed technical report on performance will be presented in the coming months. The architecture reduces data movement and balances compute, memory, and networking resources to achieve realized utilization much closer to theoretical peak performance. Broadcom’s silicon implementation and networking technologies, including Tomahawk networking silicon, help bring the platform to large-scale production.
“The world is moving to a compute-powered economy,” said Greg Brockman, President and Co-Founder of OpenAI. “Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant, resulting in AI which is faster, more reliable, more affordable for people and businesses, and can be used to solve more important problems. By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access.”
“Jalapeño was designed from the ground up for LLM inference using detailed insights from our close collaboration with OpenAI researchers,” said Richard Ho, who leads OpenAI’s hardware program. “We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits.”
“Our collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure required for the next decade of AI,” said Hock Tan, President and CEO, Broadcom. “This is just the beginning of a multi-generation roadmap. By co-developing our industry-leading silicon directly with OpenAI, we are enabling the deployment of gigawatt scale data centers with Microsoft and other partners beginning in 2026.”
Designed to be the best inference platform for LLMs
Jalapeño is a blank-slate design for modern LLM inference, not a general-purpose accelerator adapted from earlier AI workloads. It is informed by the systems OpenAI runs every day across ChatGPT, Codex, the API, and future agentic products, while also being designed for current and future LLMs across the industry. The goal is to combine the power and throughput of today’s leading AI accelerators with latency closer to the fastest specialized inference systems, making Jalapeño well suited for interactive LLM products at scale.
That is the full-stack advantage. OpenAI is not only developing frontier models or building products on top of them; it is designing the infrastructure underneath them: chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience. Because OpenAI operates across the stack, each layer can be optimized around the same goal: making its models faster, more reliable, and more affordable for users.
Jalapeño strengthens the flywheel behind OpenAI’s progress. Better infrastructure drives compute efficiency. Greater compute efficiency enables better training and serving, ultimately powering more capable AI models. Better models become better products for people, developers, and businesses. Better products drive more usage, more customers, and more revenue, which lets OpenAI reinvest in the next generation of infrastructure. Over time, that cycle helps make intelligence more capable, more reliable, and less expensive for everyone.
Nine-month tape-out, accelerated by OpenAI models
Jalapeño was co-developed from initial design to manufacturing tape-out in just nine months, and the custom AI accelerator program represents what we believe to be the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors. That speed reflects deep software-hardware co-development with OpenAI’s engineering teams, Broadcom’s silicon implementation expertise, and the use of OpenAI models to accelerate parts of the design and optimization process.
The same models served to users are helping improve the infrastructure used to run future models. If AI can help engineers design better chips faster, it can lower the cost of compute across the industry and help democratize access to advanced AI.
Building a multi-generation platform with partners
Jalapeño is the first step in a multi-generation compute platform designed for initial deployment by the end of 2026 and expanding in the years ahead, combining OpenAI-designed accelerators with Broadcom silicon implementation, networking, and connectivity technologies; and Celestica’s board, rack, and system expertise.
Making advanced AI more broadly available
The point of this work is simple: inference is where AI reaches people. Every improvement in cost, speed, and reliability can show up as a faster ChatGPT answer, a Codex task that can take more steps with less waiting, an API product that is cheaper to build, or more dependable access when demand is high.
Democratizing AI means making advanced models available, dependable, and affordable enough for more people to use every day. Jalapeño helps OpenAI turn more of its infrastructure into useful intelligence for students, developers, small businesses, researchers, enterprises, and anyone trying to learn, create, or solve hard problems.
Original source
Curated by the Releasebot team
Releasebot is an aggregator of official product update announcements from hundreds of software vendors and thousands of sources.
Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.