Cursor Release Notes

Follow

137 release notes curated from 140 sources by the Releasebot Team. Last updated: Sep 25, 2026

Get this feed:
  • Sep 23, 2026
    • Date parsed from source:
      Sep 23, 2026
    • First seen by Releasebot:
      Sep 25, 2026
    Cursor logo

    Cursor

    Improved token efficiency for longer agent runs

    Cursor improves its agent harness to cut token costs by 7% while keeping quality high. It trims system prompts, loads tools dynamically, boosts cache reuse, compresses file reads, and tunes subagent behavior for longer, more efficient runs.

    As agents have matured and learned to tackle more ambitious tasks, token spend has shifted. Agents now work for longer and carry more context from one step to the next, making the way we assemble and manage that context increasingly important.

    Notes: System & tool defs includes compaction summaries. User text includes manually attached skills. Skills & plugins includes skill descriptions, MCP tool descriptions, and rules that go in static context.

    Over the past few months we've responded to this shift by improving the efficiency of Cursor's agent harness. The harness gives us direct control over how each request is assembled, how context is reused, and when work is divided across agents. Changes across each of these layers reduced token costs for users by 7% without reducing agent quality.

    Trimming the system prompt

    Every agent turn includes context supplied by Cursor before the model begins working. This includes the system prompt and definitions for the tools the agent can use. Because this context is included throughout a conversation, it had become one of the largest sources of spend that we fully control.

    When models were less capable, we had to spell out instructions for tool usage, task management, and code-change workflows. We also had to guard against strange behaviors like extremely long hash dumps, binary output, and emojis.

    As models improved, much of that direction became unnecessary. Instead of long lists of "DO NOT do this," "You must," or "Important" instructions, we could simply define how a tool behaves and models would generally comply. This was true across model families, allowing us to trim roughly 66% of our system prompt.

    Over time, we continue to add and remove instructions as new models require new guidance, which then flows into the training of future models. Leveraging A/B tests on a large user base is crucial to effectively optimizing the harness for real traffic. While evals can be a fast and useful proxy, they often represent "hard" problems and don't properly reflect the true distribution of user requests.

    Loading tools only when needed

    The system prompt is only one part of the context Cursor supplies on every turn. Another is tool definitions, which had grown dramatically over the course of the year as we added more powerful capabilities to the Cursor agent, including background shell monitoring, cloud subagents, and more reliable access to web content. Most of these tools are important, but each is needed in fewer than 20% of conversations.

    That created an opportunity to improve efficiency by keeping tools available without including their full definitions in every request. We'd solved a similar problem earlier this year when we moved MCP tools into dynamic context, loading them only when needed. This reduced total tokens by 46.9% across sessions that called an MCP tool.

    We have now applied the same technique to our own built-in tools.

    To decide which tools to keep in static context, we A/B tested several configurations based on how often each tool was used and whether models needed to see it from the start. We tracked token usage, cost, latency, tool-call errors, and overall agent usage to make sure the savings did not degrade quality.

    Ultimately, we kept the high-frequency tools for reading, searching, editing, and using the shell in static context. We also retained ask_question, which some models tended to hallucinate calls for, and tools that are crucial to specific product flows, such as create_plan in Plan Mode. The remaining tools now load when the agent needs them.

    Improving cache reuse

    After reducing the amount of static context in each request, we improved how effectively repeated context could be cached across turns.

    Every agent turn resends a long request containing tools, system instructions, setup, and the conversation so far. Much of the beginning stays the same from one turn to the next, while the conversation at the end continues to grow.

    Prompt caching allows the model provider to reuse that unchanged prefix. However, caching configurability can vary by provider. Before GPT-5.6, the cache boundary was determined automatically based on the latest request. Even though tools and system instructions rarely changed, they were not cleanly marked as reusable on their own.

    Since GPT-5.6, the OpenAI API allows clients to mark explicit cache breakpoints alongside its default implicit caching. We now place breakpoints after stable layers of the request and before the growing conversation, allowing later turns to reuse more of the unchanged prefix.

    Breakpoints only help if the prefix itself stays stable, so we also tightened what sits at the front of each request. We did this by reserving tools and system instructions for content that rarely changes, and by moving more variable setup past the cache boundaries into our "phantom user message." This holds user- and request-specific context like skills, subagents, and environment info.

    These changes reduced the rate of cold cache misses by 20%.

    Compressing file reads

    Another large source of token spend is the context an agent adds as it works, much of which comes from reading files.

    Cursor's agent reads files through a Read tool, which traditionally numbered every line because models are not good at counting lines on their own and need to cite specific sections for the user.

    A single line number uses only around three to five tokens, but when an agent reads tens of thousands of lines during a session, numbering every one adds a meaningful amount of context.

    We reduced that overhead by including line numbers only on every tenth line. This is still frequent enough for models to cite code properly, and the change reduced cache-read tokens by 1.6% with no reduction in quality.

    Using subagents strategically

    Longer agent runs create more opportunities to delegate work to subagents. This can reduce token spend because each subagent typically starts with a fresh context window rather than carrying the parent agent's full conversation. Once it reports its results, the parent can continue without carrying the subagent's full working context.

    This kind of context isolation between agents and subagents does carry a coordination tax, though, because agents that do not share context can duplicate work or pursue tasks that are no longer necessary.

    We made two changes to capture the efficiency benefits without adding unnecessary coordination. First, we removed instructions that strongly encouraged agents to use subagents for codebase exploration. As subagents became more prevalent in training data and researchers incorporated them into post-training, models learned this pattern natively. Removing the extra prompting produced more balanced subagent usage.

    We also tightened how subagents select models. Cursor can spawn subagents using any of our available models, which makes it possible to shore up blind spots across models or pair an expensive planning model with a cheaper one for implementation. We updated the tool arguments so agents choose a different model only when directed by the user or the harness.

    Continuing to improve harness efficiency

    We'll continue measuring how context accumulates across longer runs and testing where the harness can reduce repeated processing without affecting agent quality. Over time, we expect this will allow token use to grow far more slowly than the amount of work agents can complete. We've also taken these learnings to Grok Bot, where we're working to optimize its unique harness so that users can accomplish the most work at the lowest cost.

    Original source
  • Sep 23, 2026
    • Date parsed from source:
      Sep 23, 2026
    • First seen by Releasebot:
      Sep 25, 2026
    Cursor logo

    Cursor

    Bots for the last mile: Rollouts, Security Review

    Cursor ships software development bots that help teams get safe, reliable code into production faster. Rollouts watches changes from PR to production and flags regressions, while Security Reviewer finds and fixes security issues in codebases. Both are available today.

    Today we're releasing software development bots that help you get safe, reliable code into production faster.

    • Rollouts watches a change from PR to production, flags regressions, and acts to restore a healthy state.
    • Security Reviewer finds and fixes security issues in your codebase.

    Why we built these

    Writing code is no longer the slow part. What hasn't sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.

    That work is hard, repetitive, and needs a lot of context. Sounds like the type of work that bots should be doing!

    We've said we're building toward self-driving codebases. Rollouts and Security Review autonomously handle common tasks every software team wishes they had more expertise and time to do.

    Rollouts

    Rollouts tracks a change from the moment the PR opens to the moment you're confident it works in production.

    Connect source control, your deploy system, and your telemetry (Datadog, Grafana, Honeycomb, or wherever your metrics and traces live). Before merge, Rollouts reads the diff and writes a monitoring plan: the risks it sees, the effects the change is supposed to have, and the places your instrumentation can't tell you whether it worked. Edit the plan if it's missing something.

    After deploy, Rollouts compares the signals in the plan against the pre-deploy baseline. When it finds a regression, it tells you which change it suspects and what it plans to do. Depending on how you've configured it, that's a ping to the author, a paused progressive rollout, or a revert PR waiting for approval.

    Three things it does well today:

    • It catches regressions confined to one endpoint in one region, before a global alert would fire
    • It tells intended effects apart from regressions, so a deliberate spike doesn't page anyone
    • It flags missing instrumentation before merge, which is the most common reason a bad change goes unnoticed

    Coming soon: feature flag integration so Rollouts can ramp and unramp traffic directly, and awareness of release trains and deploy freezes.

    Security Reviewer

    Security Reviewer runs on every PR, reads the change in the context of the whole codebase, and reports vulnerabilities with an explanation and a proposed fix.

    Static analysis pattern-matches: it flags every string concatenation near a SQL call and misses the authorization check that stopped running after a refactor. Security Review reads code the way a security engineer does: where does user input enter, where does it end up, what does it pass through on the way.

    Out of the box it looks for:

    • Injection across SQL, command, template, and LDAP surfaces
    • Missing or broken authentication and authorization on new and changed routes
    • Secrets and credentials committed to source
    • Unsafe deserialization and unvalidated redirects
    • Dependency changes that pull in known vulnerabilities
    • Insecure defaults in infrastructure and config

    Each finding has a severity, an attack path, and a one-click fix.

    Get started

    Rollouts and Security Reviewer are available today on Teams and Enterprise plans. Enable either bot from the automations tab to get started.

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Cursor and hundreds of other software products.

    Create account
  • Sep 23, 2026
    • Date parsed from source:
      Sep 23, 2026
    • First seen by Releasebot:
      Sep 25, 2026
    Cursor logo

    Cursor

    Rollouts and Security Review

    Cursor launches two new bots for shipping code faster and safer. Rollouts monitors pull requests through deployment and reports health by environment, while Security Review flags exploitable bugs in every PR. Both are available now on Teams and Enterprise plans.

    Rollouts

    Today we're launching two Cursor bots for the last mile of shipping code. Rollouts watches every change as it deploys and reports its health per environment. Security Review reports exploitable bugs on every pull request.

    Both are available today on Teams and Enterprise plans.

    Rollouts

    Rollouts attaches a monitor to every pull request and watches the change as it deploys, reporting change health per environment: verified healthy, regression detected, or inconclusive. It's the Cursor version of Firetiger Change Monitors, rebuilt with the Bot Development Kit.

    Enable it from the dashboard and connect source control, your deploy system, and your telemetry provider. Rollouts starts watching on the next pull request.

    Monitoring plans

    When a pull request opens, Rollouts reads the diff and the systems it touches, then writes a monitoring plan as a PR comment. The plan lists the risks it identified, the effect the change is meant to have, the signals it will check, and any gaps in instrumentation that would make the change hard to verify. Edit the plan in the PR and Rollouts uses your version.

    Deploy tracking

    Rollouts wakes on deploy events for the change's commit and runs the plan against your logs, metrics, and traces. It tracks each environment separately, so a change can be verified in staging and still flagged in production. Rollouts checks the change's intended effect alongside error and latency signals, and reports back on the PR when it reaches a verdict.

    Regressions

    When Rollouts detects a regression, it names the change it suspects and notifies the author. Depending on configuration, it can also open a revert PR for review or hand the finding to a cloud agent for a fix. Rollouts does not merge or roll back on its own today.

    Integrations

    Rollouts connects to Origin or GitHub for source control, to your continuous delivery system for deploy events, and to Datadog and other telemetry providers for signals. Feature flag integration is coming soon.

    Security Review

    Security Review is available today. It reads every pull request in the context of the codebase and posts one review comment reporting exploitable bugs. Style and quality stay with Bugbot.

    Enable it from the dashboard for the repositories you want reviewed. Draft PRs are skipped.

    What it reports

    Security Review looks for injection across SQL, command, and template surfaces, along with authentication and authorization bypasses, including checks that a refactor stopped running. It also flags secrets and credentials committed to source, SSRF and unvalidated redirects, unsafe deserialization, and dependency changes that introduce known vulnerabilities. It traces where user input enters and what it passes through.

    Findings

    Each finding carries a severity, the attack path, and a proposed fix. Dismiss one with a reason and Security Review won't raise it again on that PR.

    Team rules

    Add rules for your codebase, such as which client external calls must go through or which tables are never queried from a request handler, and Security Review enforces them on every PR.

    Get started

    Rollouts and Security Reviewer are available today on Teams and Enterprise plans. Enable either bot from the automations tab.

    For the next 10 days, we're including usage credits so teams can try Rollouts on real changes. Teams and Enterprise customers receive credits for roughly 50 and 500 changes, respectively.

    Original source
  • Sep 10, 2026
    • Date parsed from source:
      Sep 10, 2026
    • First seen by Releasebot:
      Sep 11, 2026
    Cursor logo

    Cursor

    Cursor Projects

    Cursor launches Projects, a new beta workspace for larger coding efforts that keeps context over months, delegates work to agents, runs in the cloud, and can react to signals like Slack, schedules, and PRs while rolling out to all users today.

    Projects in Cursor

    Today we're launching Projects in Cursor. Projects lets you take on larger bodies of work, such as a feature, a migration, or a full app. It maintains context over months of work, delegates tasks to thousands of subagents, and performs recurring work without being prompted.

    Access Projects from the left-hand nav. The coordinator agent in a project doesn't write code itself; it plans the work, delegates it to agents that implement it, and brings the finished work back to you to check. Coordinators create and manage agents on your behalf, running as many in parallel as the work needs.

    A Project runs on its own computer in the cloud, so closing your laptop doesn't stop it. When something needs testing on your machine, the coordinator spins up a local agent to run it there.

    You shouldn't have to onboard an agent every time you start a task. Each Project maintains a set of files that sync across every cloud and local machine its agents use. Agents add research and artifacts, along with what they learn about the codebase and how you prefer work to be done. If one agent figures out how to test a service, for example, every future agent can use those instructions. The shared context grows with the Project, making the coordinator more effective over time.

    Tell the coordinator agent to watch a Slack channel, run on a schedule, or follow all your PRs. This way it can take action based on signals it detects, without waiting for your prompt.

    Connect Slack and point it at a bug-report channel, and it starts delegating each time a bug comes in.

    Projects are available in beta and rolling out to all users starting today.

    Original source
  • Sep 10, 2026
    • Date parsed from source:
      Sep 10, 2026
    • First seen by Releasebot:
      Sep 10, 2026
    Cursor logo

    Cursor

    Introducing Projects

    Cursor launches Projects, a beta feature for tackling larger bodies of work with a coordinator agent that delegates to thousands of subagents, keeps shared context over time, and can run on schedules or respond to signals like Slack and PRs.

    Direct thousands of agents through one coordinator

    Today we're launching Projects in Cursor. Projects lets you take on larger bodies of work, such as a feature, a migration, or a full app. It maintains context over months of work, delegates tasks to thousands of subagents, and performs recurring work without being prompted.

    In February we outlined our vision for a third era of software development, where fleets of agents take on entire bodies of work. Projects is the concrete implementation of that vision. By moving up a level of abstraction, it frees developers from managing agents and lets them direct the work itself.

    At Cursor, we've been using Projects for several months, doing work such as running migrations of a few hundred PRs, keeping our design system consistent, and shipping Projects itself. We've found it to be a substantial productivity multiplier: new users merge 30% more PRs while users who primarily use Projects merge six times as many.

    You oversee a Project by chatting with its coordinator agent. The coordinator doesn't write code itself but directs other agents that do. Because it delegates rather than executes, it is never blocked and is always responsive to direction.

    There are three core capabilities that make Projects possible:

    • Cloud by default, local when needed. A Project runs on its own computer, so closing your laptop doesn't stop it. This lets a Project run more subagents in parallel than your laptop could support. When something needs testing on your machine, the coordinator spins up a local agent to run it there.

    • Shared context. You shouldn't have to onboard an agent every time you start a task. Each Project maintains a set of files that sync across every cloud and local machine its agents use. Agents add research and artifacts, along with what they learn about the codebase and how you prefer work to be done. If one agent figures out how to test a service, for example, every future agent can use those instructions. This context grows with the Project, making the coordinator more effective over time.

    • Subscriptions. The coordinator can watch a Slack channel, run on a schedule, or follow all your PRs, fixing CI and acting when they open or merge. This way it can take action based on signals it detects, without waiting for you to prompt it.

    How we use Projects at Cursor

    Three patterns cover most of what our engineers do with Projects.

    Feature work

    Most engineers create a Project for a substantial body of work. A feature usually starts with agents researching the system and recording what they learn as shared context. The coordinator then creates a plan and sends agents to implement and test different parts of it in parallel.

    With each turn of feedback, the Project learns your architecture and preferences. When the feature is ready to try, the coordinator can start an agent on your computer and run it locally. After it ships, the same Project can monitor logs and handle bug reports with the full context behind the original decisions.

    Migrations

    Projects are especially useful for migrations that are easy to start and difficult to finish. At Cursor, we've used them to adopt new frameworks and replace styling systems across hundreds of PRs.

    You work with the coordinator to establish a safe approach, then it applies that approach incrementally across the codebase. Early on, you review each PR closely. As the fixes hold up, you review less, and the coordinator keeps working through the migration on its own.

    Gardening

    Projects are great for handling work that never really ends, such as maintaining code quality or watching for regressions. You can tell the coordinator to follow new PRs, listen for bug reports in Slack, or run on a schedule, and it acts whenever new work appears.

    One engineer on our team runs a design-system Project this way. At first, the engineer reviewed each fix and corrected the ones it got wrong. Now the coordinator scans every new PR, extracts components that belong in the design system, and adds a lint rule whenever it sees the same mistake twice. The Project is on track to touch 20 to 100 PRs a day, so the coordinator organizes the work and the engineer checks in where attention is needed.

    Get started with Projects

    Projects are available in beta and rolling out to all users starting today. Start a Project from the left hand nav, describe what you want built, and the coordinator takes it from there. It works best on work that will outlive a single chat, whether that's a feature with several PRs, a migration, or a job you want handled while you're away.

    Original source
  • Similar to Cursor with recent updates:

  • Sep 2, 2026
    • Date parsed from source:
      Sep 2, 2026
    • First seen by Releasebot:
      Sep 3, 2026
    Cursor logo

    Cursor

    Self-hosted machines

    Cursor adds self-hosted machines for local tool execution, keeping code, builds, and secrets inside your network. It also brings dynamic machine pools, support for cloud agents on existing infrastructure, and computer use on Linux and Mac for browser and desktop control.

    Cursor supports self-hosted machines, which let you keep tool execution entirely in your own network.

    Your codebase, build outputs, and secrets all stay on internal machines running in your infrastructure, while the agent handles tool calls locally.

    Dynamic pool scheduling

    My Machines connects a single laptop or VM to your account for personal workflows.

    Team pools are named queues of workers for a team or enterprise. Capacity can grow as requests arrive and shrink when workers disconnect, so your self-hosted machines can scale with demand. Pools are not tied to one repository: name the pool, and any available worker can claim the request.

    Pools can also hibernate idle machines, then restore within a reconnect window when a follow-up arrives, so you don't keep expensive capacity warm just for the next prompt.

    Run on your sandboxes

    Cloud agents can now execute on infrastructure you already use, including from AWS Lambda, Coder, Cloudflare, Daytona, Modal, Namespace, Vercel, and E2B.

    Computer use on Linux and Mac

    Self-hosted workers now support computer use on Linux and Mac. With the right desktop packages, an agent can click, type, take screenshots, and drive the browser. You can watch its desktop or take control from Cursor.

    Original source
  • Sep 2, 2026
    • Date parsed from source:
      Sep 2, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Cursor logo

    Cursor

    Run cloud agents on machines you manage

    Cursor adds Self-Hosted Machines for cloud agents, letting teams run agents on their own infrastructure while Cursor still manages planning and orchestration. It also expands pool scaling, sandbox provider support, and Linux computer use for more flexible, secure agent workflows.

    Cursor cloud agents can execute on dynamically scheduled pools of machines inside your network. You manage the underlying infrastructure, while agents are still started and managed from Cursor.

    This gives teams more control over where agents execute and what infrastructure they use. Agents can work next to internal services and source control, run on custom hardware, or use operating systems and build pipelines that are difficult to package as a Cloud Agent build.

    Cloud agents now create more than 60% of the pull requests we merge internally and are taking on a growing share of software work at many of the largest enterprises we work with. As their role expands, the machines they run on matter more too. These new capabilities make it practical for teams to provide and manage that infrastructure at scale.

    With Lambda MicroVMs as the compute layer for Cursor Cloud Agents, developers can run AI-powered coding agents in their own AWS account. Each machine launches near-instantly from a snapshot, suspends when idle, and resumes with full state. Your coding agents benefit from Lambda's fast startup, strong isolation, and zero fleet management, while Cursor orchestrates the work.

    Control where agents execute

    Cursor-hosted environments remain the default for cloud agents. Each session runs on a dedicated VM inside the Cursor cloud, with its dependencies installed and its own network controls. Per-agent isolation, secret redaction, egress controls, and signed commits meet the security requirements of most teams.

    Teams generally use Self-Hosted Machines when:

    • Agent tool execution needs to happen inside their network, with direct access to source control, internal services, and code repositories.
    • Agents require custom hardware, such as GPUs or Macs for iOS development, or infrastructure such as Kubernetes, sandboxes, or managed VMs.
    • Their operating system or build pipeline is difficult to package as a Cloud Agent build.

    With Self-Hosted Machines, only the execution environment moves while the agent loop, inference, and planning remain in the Cursor cloud. Tool outputs flow back to Cursor for inference and may contain code, and agent transcripts may be processed and stored by Cursor. Teams can continue to access cloud agents from the desktop app, cursor.com, mobile, Slack, GitHub, and Linear.

    Workers connect your infrastructure to the Cursor agent loop

    With Self-Hosted Machines, tool execution moves from a Cursor-hosted VM to a machine in your environment. That machine holds the working copy of the repository, edits files, and runs commands. A worker connects it to the rest of the agent system.

    To register a machine, run a worker by installing the Cursor CLI and running

    agent worker start
    

    This opens a long-lived outbound HTTPS connection to the Cursor cloud. When a session begins, Cursor's agent harness handles inference and planning, then sends tool calls to a dedicated worker for execution. The worker returns the results for the next round of inference. Cursor never initiates a connection into your network.

    Workers can be configured in two ways.

    1. My Machines. This configuration connects a single laptop or VM to your account and is best suited to personal workflows.
    2. Pools. A pool is a named queue of workers that can serve a team or enterprise. Capacity increases as requests arrive and decreases after workers disconnect, letting your existing cloud infrastructure scale with developer demand.

    Developers should have the flexibility to run coding agents on the platform that best supports their workflow, and companies should not have to compromise on control of where agents run and what they can access. The future of development will be built on powerful agents, running in secure, isolated environments.

    Cloud agents adapt to your infrastructure

    Worker pools can now scale in response to queued requests and serve work from any repository. We have also added support for several sandbox providers and computer use on Linux alongside Mac.

    Pools scale with demand and serve any repository

    Demand for cloud agents often arrives in bursts and Self-Hosted Machines pools adjust to those bursts automatically. This happens through a controller which watches the request queue and uses a spawn script supplied by the team to start machines as needed.

    If a pool has an available worker, that worker claims the request. Otherwise, the request waits until more capacity becomes available, so teams do not have to decide how many machines to leave running.

    Teams can set an idle timeout for each worker connection. Once it expires, the machine can reset and re-enter the pool. Teams can also preserve its workspace in case the agent receives a follow-up.

    Self-Hosted Machines put teams in control of where Cursor agents run, and Vercel Sandbox makes it effortless. Every task gets an isolated sandbox on demand, no fleet to manage, and nothing sitting idle.

    Leaving a machine running while its agent is idle can be expensive. But if the machine is released, the agent may need several minutes to reconstruct its workspace when a follow-up arrives. With hibernation, teams can snapshot and stop an idle machine instead. If a follow-up arrives within the reconnect window, the snapshot is restored and a worker starts with the same ID. Otherwise, the request can move to a new machine.

    Pools are not tied to individual repositories. A request only needs to identify the pool, and any available worker can claim it. This lets one pool serve many repositories.

    Workers run across supported sandbox providers

    Self-Hosted Machines does not require building a custom sandbox layer from scratch. We partner with AWS Lambda, Cloudflare, Coder, Daytona, E2B, Modal, Namespace, and Vercel, allowing workers to be started and orchestrated wherever a team's sandboxes already run.

    Cursor Self-Hosted Machines on Modal gives each Cloud Agent session a Modal Sandbox, so you can hand it a machine tailor-made for its task.

    Agents control browsers on Linux and Mac

    Linux workers now support computer use alongside Macs. With the required computer use dependencies installed, including Chrome or Chromium, an agent can click, take screenshots, and control the browser. You can watch its desktop or take control directly from Cursor.

    You can't build iOS or macOS apps without a Mac. Namespace Devboxes spin up a real Mac for each Cursor Cloud Agent, which can now perform that work on Apple silicon.

    Bring cloud agents into your environment

    Teams have spent years shaping their infrastructure around how they build software. Self-Hosted Machines lets cloud agents fit more naturally into it, and we're excited to see how far teams take them.

    To connect a machine or configure a pool, get started in the docs.

    Original source
  • Aug 27, 2026
    • Date parsed from source:
      Aug 27, 2026
    • First seen by Releasebot:
      Aug 27, 2026
    Cursor logo

    Cursor

    Start from scratch, without a repo

    Cursor adds a smoother Cloud Agents workflow that starts without a connected GitHub or other SCM provider, then lets users save work to a Cursor Origin repo, preview live in the browser, and publish to Vercel with a live URL.

    Cloud Agents no longer require a connected GitHub or other third-party SCM provider to get started. Prompt from the get-go, then save your work to a Cursor Origin repo.

    Get started

    Select Start from scratch in the repo picker, then prompt the agent. In the background, Cursor creates an Origin repo for you.

    Turn it into a real repo, whenever you want

    Once the agent's build is to your liking, click the Create repo button to save your work in an Origin repo. Choose a custom name or pick one of the suggested ones, then set the visibility to private or internal. You'll get a fully scaffolded Origin repo, ready to share or keep building on. Navigate to the Codebase tab to find your repo and access your project.

    A live preview, right in the browser

    Cursor now port-forwards your cloud agent's live environment straight to your browser, so you can preview it and use tools like design mode.

    Publish your work

    Connect a Vercel account and hit publish to get a live URL for what you built.

    A Vercel account is required to use the publish feature.

    Get started today.

    Original source
  • Aug 19, 2026
    • Date parsed from source:
      Aug 19, 2026
    • First seen by Releasebot:
      Aug 19, 2026
    Cursor logo

    Cursor

    Cloud Agents and Cursor Harness Improvements

    Cursor adds always-on cloud agents that can pick up work from events, hold long-lived goals, and keep running through complex sessions. It also introduces subscriptions for PRs and Slack, custom modes, subagents on their own machines, /goal, and smoother steering while agents work.

    We're continuing to improve cloud agents and the Cursor harness so always-on agents can operate as a system, building and shipping software on their own without the need for intervention at each loop.
    With this release, cloud agents can automatically pick up work in response to events, hold a goal until it's met, and stay on course through long-running sessions.

    Subscriptions

    Cursor can now monitor your PRs, watch a Slack thread, or run scheduled tasks. Cursor Agent subscribes to an event source (a thread or conversation) and wakes when something happens. Subscriptions are available for cloud agents only, for now.
    Cloud agents automatically subscribe to PRs they create and drive them to completion, fixing CI and addressing bot comments. In Slack, ask
    @cursor check back in an hour and keep going until that feedback is in.

    Custom modes

    Use any skill as a Custom Mode: a skill that stays pinned in the chat. Custom modes keep agents focused on a skill - you can think about it like "always on" skills.
    From /, pick a skill and press
    (Mac) or Alt+Enter (Windows), or choose Use as Mode.

    Subagents on their own machines

    Subagents can now run on their own virtual machines. Each gets an isolated copy of the project with clean context in its own cloud environment.
    Have subagents test the parent agent's changes in fresh environments or swarm independent fixes without collisions. Try
    run a swarm of subagents to test my app for bugs, each in its own environment.

    /goal

    Use /goal to give the agent a long-lived objective to work towards until it's fully complete.
    Try /goal fix all flaky tests and make CI green in a new chat. Pair it with a custom mode to follow a playbook, or /loop for recurring check-ins.

    Steering improvements

    You can now send a message to steer the agent while it's working without interruption. Follow-ups wait for the next tool call instead of cutting the agent off mid-action.
    Type a follow-up and hit Send now, or press
    twice.

    Original source
  • Aug 18, 2026
    • Date parsed from source:
      Aug 18, 2026
    • First seen by Releasebot:
      Aug 18, 2026
    Cursor logo

    Cursor

    Git at any scale

    Cursor introduces Origin and its Continuity Git storage system, aimed at making repository hosting more reliable, scalable, and consistent. It replaces heavy replica coordination with an S3-backed write-ahead log, local NVMe repositories, and linearizable pushes to support faster reads and smoother operations at scale.

    Hosting Git repositories at scale is a nightmare

    When Linus Torvalds designed the first version of the information manager from hell (that's actually the tagline for Git, look it up), he had a very specific use case in mind: his own. He wanted to replace BitKeeper, the distributed version control system that was being used to develop the Linux Kernel. Of course, the replacement had to be distributed too. The Kernel is an unusual software project; it is extremely decentralized, with many different maintainers for its many different subsystems. A distributed version control system is a natural fit for this workflow.

    Twenty years later, Git has become an industry standard, but the truth is that its distributed nature is more of a hindrance than an advantage. The average open-source software project doesn't operate with a decentralized workflow. The average company definitely doesn't. They use the many advantages of the distributed model (such as being able to work offline, delay pushes, etc) but they very much rely on a centralized host. And hosting a Git repository, it turns out, is an incredibly hard thing to do.

    What's hard about Git?

    The challenge in hosting Git repositories at scale is inherent in the design of Git itself: a distributed version control system means that all instances of a repository are identical. There's nothing special about the repository on a Git server that doesn't apply to a repository on a developer's laptop. Although at first it may appear that this makes hosting Git repositories straightforward (simply put an HTTP daemon in front of an on-disk copy of a repository and you've got a Git server going!), there are many hard scalability and reliability challenges that make this quite the opposite.

    In a normal Git repository, your code and metadata (files, commits, trees) are compressed and stored in packfiles — a simple binary serialization format which is convenient to deal with on a local machine, but not ideal to manage at scale on a server. Packfiles are the fundamental building block of Git storage and Git networking. When you push or fetch data from a repository, it's transferred as a packfile.

    This is how Git works by design, but it would be fair to think that it needn't be that way. After all, you do not control the Git client (at least not without annoying your users and adding a lot of friction), but within the walls of your own server, you can do anything you want. Nothing ties you to using packfiles — Linus is not going to come over and check. The only restriction is that you do need to receive and send packfiles over the network for all Git operations.

    Over the years, companies that tried hosting Git repositories at scale noticed that this packfile-based design was a major limitation on both availability and scalability. Packfiles are large binary files that must exist on a filesystem for Git to access them. The simple approach of having an HTTP server in front of a repository on disk has a very low ceiling. Ideally you'd want the repository to exist on many disks and many machines (this lets you run many Git operations in parallel, and keeps your repository available when a server crashes). But how do you do that?

    There are broadly three possible approaches to accomplish this, in increasing order of complexity: distribute the filesystem, distribute the packfiles, or distribute Git itself.

    Git without packfiles

    Git is a content-addressable data store. All objects in a Git repository (blobs, trees, commits, etc) are keyed by the SHA-1 of their contents. This is something that intuitively maps very well to a distributed key-value store (the key is the SHA-1; the value is the actual object), and could provide a clean way to scale out the storage of a repository. But this actually doesn't work.

    Here's the issue: the actual layout of a Git repository is a directed acyclic graph (DAG for short). You can look up any object via its SHA, but to perform even the most trivial operation in the repo, you must actually walk the DAG step by step.

    If you want to do an operation like listing the recent changes in a repository, you must process its commits. When you process a commit, you get a pointer to the root of its tree. From that tree, you get pointers to each file and each subtree. From the original commit, you get a pointer to its parent (the one that comes before it in the history). Crucially, at every step of this walk, you don't know the value of the next pointer until you fetch the previous one. If every fetch requires a round trip to a distributed store, things become very expensive very fast.

    This approach to distributing Git at the object level has been tried before, many times, and it often fails at scale. The most promising implementation was attempted by my former mentor Shawn Pearce when he was working on the version control systems team at Google. His approach was storing the objects in a distributed hash table. This was only possible thanks to JGit, a custom Git implementation in Java. Like any good ol' Java library, JGit provides enough interfaces and factories and interface factories to abstract all the details of a normal Git repository, including replacing its on-disk packfiles with a DHT. Although the system worked and results were good enough for normal Git operations, the limitations of the Git protocol (which again, require packfiles to be sent over the network regardless of how you store data on the server) made the git clone performance bad enough to discard the design altogether.

    GitHub and filesystems

    A couple years after Git started to escape its Linux Kernel bubble, a scrappy startup was born in San Francisco. GitHub was founded in 2008 as a social coding platform with a very prescient tagline, "Git repository hosting: no longer a pain in the ass." I'm not joking here either, look it up. There was, all the way back in 2008, a broad consensus that despite (or perhaps because of) Git's distributed design, you actually needed a centralized way to host Git repositories to make them user-friendly, and doing this was very painful. GitHub was set on changing that.

    Its platform started as (and mostly still is) a Rails monolith. The very first versions were running off a single, albeit beefy, machine, with a Ruby server and copies of the repositories on disk next to it. Scaling a Rails app is easy: deploy more instances of it. But in this particular case, since Git is involved, they quickly ran into the recurring question we're trying to solve here: If the Rails app needs to access the Git repositories on disk, how do you deploy more copies of them?

    Being a thrifty bunch of misfits, the early systems engineers at GitHub tried the simplest approach that could possibly fix their scaling problems. The thinking was that, if they focused on distributing the filesystem (instead of packfiles, or Git itself), they could keep the Rails app unchanged and spend their time shipping more features for the ever-growing user base, instead of doing weird stuff with Git. Very pragmatic. It didn't work.

    The team attempted many approaches to a distributed filesystem for Git data: the most obvious one, using NFS to store all repositories on a centralized server, was quickly discarded. The default implementation of Git makes a lot of assumptions about filesystem semantics (locking, tearing, reading, syncing...) that ensure decent performance on the local filesystem of a slow developer laptop, but pay no attention to how they behave over a networked filesystem. It was slow, and it was buggy.

    Further attempts were made with (frankly, in retrospect, horrific) technologies that replicated the filesystem at the block level. A short-lived deployment with GFS. A longer-lived deployment based on DRBD. They all hit a wall. They were terrible to operate day to day, and they didn't make up for it with good performance. It all boils down to the design of packfiles on disk.

    We've already seen how Git's graph-like data structures make round-trips prohibitively expensive. Unfortunately, a very similar principle also applies to the underlying data on-disk. There is no correlation between the layout of objects in the DAG and the way they're placed in a packfile. The key heuristic used when generating packfiles is minimizing their size; objects are placed randomly throughout the pack, they are compressed, and crucially they're rarely stored whole. Most objects are stored as a delta on top of another object in the same packfile. Reading an individual object, after following the many logical hops in the graph data structure, also involves following physical hops in the on-disk format.

    This kind of random walk across gigabytes of data, which must happen for every single Git operation performed on a repository, just doesn't play well with a networked filesystem (whether it replicates at the file or at the block level). The only way this works without slowing down to a crawl is if you can cache the whole file locally. But with hundreds of thousands of repositories in the same filesystem, caching is not an option.

    Eventually, the systems engineers at GitHub bit the bullet and gave up distributing the filesystem. They started developing an RPC system so that repositories could live on dedicated fileservers, and updated the Rails app to do all operations remotely. This provided a good chunk of horizontal scalability, but didn't fix their availability, nor the performance for the busiest repositories. After all, every repository was still stored only on a single machine.

    Spokes and Consistency

    Spokes was originally developed at GitHub around 2013, and it has since become an industry standard. Most Git hosting services use a variant of the Spokes approach (application-level replication for Git repositories) in their architecture. The main reason Spokes has worked well for many years is that it made three fundamental choices that, over time, have been proven to be optimal:

    • It doesn't distribute Git itself; it works at the packfile level.
    • It stores all data as actual Git repositories on local NVMe disks.
    • It replicates the Git data, but keeps all copies consistently in sync.

    Because of the random read patterns across packfiles we've just discussed, storing plain Git repositories on NVMe drives is basically a requirement to ensure all basic Git operations remain fast. They also keep clones efficient because you don't have to transform the data into what the Git client expects. They also let you focus on building a product on top of Git, as opposed to maintaining a fork of Git yourself that can operate on your weird repositories.

    Keeping all the copies of the data consistently in sync is also, crucially, very good. This is something you find out the hard way, but the Git client really doesn't play well with eventual consistency. If your local Git client pushes a commit and then fails to read it immediately after a fetch, that's bad news. Git finds that very confusing. If you run your CI pipeline across a hundred runners and three of them don't find the commit they're supposed to test after cloning your repository, that's bad news. It's also a very poor user experience.

    Working with an eventually consistent view of a Git repository has a lot of sharp edges, whether it's on the client or in the backend. Hence, Spokes pays a very high complexity cost to ensure the system is always fully consistent. Let's see exactly what this means.

    Spokes is a consensus-based distributed system. It works by storing several copies of your Git repository on different servers. Whenever you push new data, an orchestrator fans out your push so that every instance of your repository receives a copy. The "fan-out" is synchronized with a classic consensus algorithm called 3PC (three-phase commit) so that a push is only accepted if a majority of the nodes acknowledge it.

    Before we can talk more about the way Spokes uses 3PC, we need to understand how a Git push works. A Git push has two components: a packfile and a reference transaction. The packfile, which we've already talked about, contains the objects you're pushing to the repository (blobs, trees, and commits with your changes). The transaction is what actually publishes your changes to the repository by updating one or more references (e.g. the branch you're working on) to point to the commits you've just pushed.

    This separation is very convenient here, because a pushed commit is not visible ("reachable" in Git parlance) until the reference that points to it has been updated. This means we can implement consensus for our pushes by fanning out the packfiles to all hosts simultaneously (we don't need to synchronize here) and then doing three-phase commit with the reference transaction, which is much smaller and faster to synchronize than the packfile. Git itself has support for preparing reference transactions: it can acquire a lock on the reference, verify that the existing value is what's expected, and then hold the lock until it receives a commit or an abort command for the transaction.

    Spokes distributes packs and then performs a three-phase commit for each push’s transaction. You can increase the number of replicas and the latency in this simulator to see how it affects commit throughput.

    With this design, we ensure that every push is fully synchronized across all the replicas. Reads (fetches, clones) can then be safely routed to any single replica, because every replica is always up to date.

    This is essentially how Spokes works, and it has been working quite well for the past 13 years. Of course, Spokes is not perfect — no system is. In 2026, the way people use Git repositories has changed drastically, and we have learned many important lessons about building distributed systems along the way. Time and experience have shown which of Spokes's choices turned out to be optimal, and which did not.

    One flaw that has turned out to be critical is the constrained horizontal scalability of 3PC. When Spokes was initially released, three replicas per repository was the sweet spot. You could serve your average repository from three copies with capacity to spare, with enough redundancy to keep accepting pushes even if one machine went down.

    In 2026, things look very different. The average repository for an enterprise company is now a massive monorepo. Three replicas are not enough to serve the traffic for such repos, particularly when it comes to CI. Of course, nothing stops Spokes from running with more than three replicas, except the dreaded tail at scale. Three-phase commit maps very elegantly to the Git transaction model, but as a consensus algorithm, it has fundamental limitations: the latency of every step is bound by the slowest of all the servers in the cluster. The more replicas you add to a cluster, the worse push throughput gets.

    This scalability constraint also applies the other way. When agents work with Git repositories at scale, they often operate outside of a monorepo by creating vast numbers of small repositories, many of them throwaway, and most of them barely touched. Spokes struggles here because it still requires three replicas for every one of these repositories. Three mostly idle replicas, which cannot be trimmed down because then the system wouldn't be fully consistent and data loss would be possible. With three-phase commit, the floor is always too high, and the ceiling too low.

    Another flaw, impossible to see up front, but painfully obvious after having suffered through it, is that Spokes can be rough to operate at scale. Because the repositories on disk are always the source of truth for consensus, every copy of every repository is very important. You have to treat repositories as pets, not cattle.

    This means, for starters, that you need to know exactly where every repository is. This adds a dependency (and a potential availability issue) on an external database that must keep a very large routing table mapping every repository to every machine where it's replicated. Every repository must also be checksummed, and its checksums constantly updated in that table, to ensure the repository remains valid on disk. As soon as something bad happens to the repository (and trust me, bad things happen all the time — Git can be very finicky in practice), you must detect it and schedule a repair job to bring it back to a healthy state. And you must do it very quickly! Because, again, the repositories on disk are the source of truth. A corrupted copy is as bad as a missing one. If two of the three copies are corrupt, the system can no longer accept pushes: there's no quorum.

    Continuity

    Continuity (Cnt for short) is the Git storage system we've developed at Cursor, with a very clear approach: learning from everything that Spokes did well, and fixing the things that, after many years, we now know are problems.

    Cnt is a simple system (a system cannot be easy to operate if it is not simple). The core primitive behind it is a write-ahead log, which we store in S3-compatible object storage. In production, we run directly on S3, but we designed it so it can be deployed on any cloud.

    When a repository receives a push, we store the push as a WAL entry in S3. We never acknowledge a push until it has been fully persisted. Each push is stored as a separate object; we write the pushed packfile to disk and upload it to S3 simultaneously. Uploading a WAL entry, however, does not publish it. A push is only visible once we successfully prepare its reference transaction on a local copy of the repository and record a pointer to the WAL entry in the WAL index file, which is its own object in the store. This forces all pushes to be linearizable.

    We try not to do one single S3 write per push, because in busy repositories, this puts a hard cap on push throughput based on the latency of the S3 PUT operation. With a carefully tuned batching implementation, and with the only requirement of having to synchronize the reference transaction with a single local repository instead of a quorum of replicas, we have a system that can ingest pushes as fast as our disk allows.

    The local copy of the repository is, of course, a normal Git repository stored on a very fast NVMe drive. We do the same thing that Spokes does because I think Spokes got that exactly right. It allows us to reuse all the amazing OSS work of the Git community, including the upstream Git client and its many performance optimizations. It lets us focus on shipping new features, instead of doing weird stuff with Git.

    Consensus

    We've seen that one thing that makes a Spokes cluster hard to operate is that it's very important to keep track of the location of every repository on each server. Cnt does this very differently. Where does every repository live? The answer is "anywhere". It doesn't matter! We treat repositories like a warm cache on disk, but the source of truth is always the write-ahead log in S3. The system is stateless, and there are no routing tables (and no relational database to operate — hashtag blessed). If a repository is missing from the local disk when accessed on a host, we just materialize it from the WAL. We can do this very efficiently, but of course we don't want to do this all the time, because it'd be wasteful. In production, we use rendezvous hashing to map a repository ID to the list of nodes where we expect it to be. All the state we require to route repositories is the repository ID and the current set of healthy nodes in a cluster. But if this state gets out of sync (e.g., a node becomes unhealthy), that's perfectly fine too. We'll just materialize the repository on whichever node comes next.

    What about consensus? Elections? Which server is the primary for a given repository? It also doesn't matter! There's no state and no consensus here. Any server can be the primary. All updates to the write-ahead log are synchronized with an atomic compare-and-swap (CAS) operation on S3, so it's always safe for any instance of a repository to receive a push. Again, just like with routing, letting an arbitrary server act as the primary isn't the most efficient thing (it leads to CAS retries, which can delay pushes), so in practice we always choose the same server as the primary, the first one in the ranked list from rendezvous hashing. But in the corner cases — when there's a deploy, a failover, a network blip — we just don't care exactly which server is the primary. The system is designed to always be correct when degraded, and always fast when healthy.

    Replication

    Having a write-ahead log in S3 opens a world of possibilities when it comes to scale. We can have literally any number of replicas, because the scalability of S3 is unmatched and all the replicas catch up directly from there. We perform optimistic replication by sending gossip UDP packets around our cluster. The packets contain all the required metadata for each replica to catch up directly from S3 after every push. "That is insane," I hear you mumble from behind your screen across time and space. "UDP is not a reliable transport." Of course it isn't. Nothing is reliable in a distributed system! The wire is not reliable, the routing is not reliable, and the topology is not reliable either. But it's OK: it doesn't matter. Each replica knows the ETag of the last version of the WAL index it's caught up with. When you perform a read operation on a replica, we do a conditional GET to S3 with the ETag we expect. A 304 response with no body (conveniently, an almost instant operation — less than 10ms on average because it's a metadata-only S3 operation) means we're up to date and we can serve the fetch or the clone straight away. A 200 response comes with the newest version of the WAL index, which we use to catch up before serving the read.

    It doesn't matter if the replication UDP packet is lost, or if it arrives at the wrong server because the topology shifted. All reads on all replicas are fully consistent, because they're verified against the source of truth, which is S3. The system is designed to always be correct when degraded, and always fast when healthy.

    The implications of this are twofold. First, because the system is always consistent, building infrastructure on top of it is trivial. We (our agents, our web interface, our clients) always see a globally consistent view of the repository. And because the system scales in both directions, every repository gets just the right number of replicas. A large monorepo can be deployed across hundreds of replicas to serve all the load from its CI jobs. Millions of tiny repositories created by agents can be served with one replica each; we don't need more than one to ensure availability, because S3 is the source of truth. In fact, an idle repository doesn't even need that: when a replica hasn't received traffic for a while, we garbage collect it from the node's disk, and simply materialize it again from the WAL the next time a fetch comes in.

    Compaction

    Write-ahead logs require periodic compaction. You cannot let the log grow unbounded: a full restore replays every entry, so the more entries, the more expensive it becomes.

    Coincidentally, a normal Git repository also requires periodic compaction, even though Git is not based on a WAL. We've seen that the fundamental unit of storage in a Git repository is the packfile. Each time you push to a remote copy of a repository, or fetch into your local copy, you create a new packfile. This doesn't scale indefinitely: each packfile has its own attached index, which allows Git to efficiently look up the objects it contains, but this lookup is only efficient on a per-packfile basis. If you're looking for a specific object, and your repository has 100 packfiles, you'll need to open the index for each one of them and look up the object until you find it in one of the packfiles. An efficient operation is not efficient if it must be performed hundreds or thousands of times.

    Modern Git has gotten very good at working around this; it now supports multi-pack indexes and incremental geometric compaction. But eventually you must bite the bullet and repack your Git repository on disk. Historically, this has been a constant availability issue for systems like Spokes, because repacking is a very CPU-heavy operation, even when done incrementally, and it must be performed on all the replicas of the system. Accidentally triggering a maintenance operation on two or more Spokes nodes for the same repository will easily cause the repository to fail over.

    Here, we amortize the cost of compaction. Only the primary does compactions, and the result of the compaction applies to both the on-disk repository and the WAL. Since all replicas follow the WAL, they also follow the compaction events. Replicas don't repack; they simply download the already-compacted packs from S3, trading bandwidth for CPU.

    Scale

    Replication and compaction are the two key factors that determine how well a Git storage system behaves under load. As we’ve just seen, they’re intrinsically linked: the more pushes per second a repository ingests, the more read performance degrades, because the packfiles of every push must be compacted for Git operations to remain efficient. If you replicate these pushes, the compaction must be either replicated or performed independently on each replica.

    Continuity’s WAL-first design offers fully consistent horizontal scalability: you can deploy an arbitrary number of replicas, and the throughput for read-only Git operations grows linearly with them. Because all replicas in the cluster are fully consistent, this allows us to scale the Git protocol (clones, fetches) and all the RPC operations that Origin performs on top of repositories (web UI interactions, the REST API, all our agentic interfaces, etc.)

    We have run synthetic stress tests with up to 100 replicas and seen consistent linear scaling for reads, without any regressions in push throughput.

    The push throughput of a cluster depends on the latency at which we can update our WAL on S3. Using S3 Standard, we can sustain up to 120 pushes/s while compacting and replicating the compacted data to all other nodes. We have also deployed high-performance clusters on S3 Express One Zone, which has much lower latency for PUT operations. There, we can ingest more than 300 pushes/s, and we are effectively bottlenecked by the speed at which Git can compact the on-disk data. We’re working on innovative ways to lay out this data on disk to reduce the impact of compaction: our goal is to continue optimizing the speed at which a Git repository can ingest code without relaxing our hard durability and consistency guarantees.

    WAL as truth

    S3 is a great piece of technology. The whole concept of blob storage that was pioneered with the S3 API has turned out to be a very powerful building block for large data storage systems, and this most definitely also applies to hosting Git repositories. The design presented here is novel on many ways, but it's not the first one to store packfiles as blobs. Azure DevOps (Microsoft's own competitor to Microsoft's own GitHub) has a very successful Git storage system that stores packfiles in blob storage and their references in a relational database (MS SQL Server). There are many trade-offs to a system like this. A relational database scales well with large reference transactions. But then you have to operate a relational database. We have a strong belief that the consistency of Git data is more important than any other consideration. This is what really tipped the scales for us into designing a WAL-based system that doesn't depend on external databases.

    There are many things that can go wrong with a Git repository in production. Data corruption at rest, bugs during repacking, races during pushes. It's one big collection of corner cases. Most of these have been ironed out in Git upstream. But not all of them. No system is without bugs, not even those that are OSS and widely deployed. Our consistency model ensures that we keep track of every fundamental operation that happens to a repository. We never acknowledge a push until it has been fully persisted to the WAL. We linearize all pushes. Every view of every repository we access is always fully consistent.

    Since every push is in the WAL, we can look at every state a repository has ever been in. We have full provenance data for all pushes, and also for all repacks. We can rewind and fast-forward every replica. We don't have to synchronize any state with any external database, whether it's a database that only stores references, or a database that stores all object data. When (not if) we hit a bug in Git, we can pinpoint exactly what happened and revert it. And besides the bugs that already exist in Git, we introduce very few new ones, because throughout all this, all Git operations are performed on a normal Git repository on disk, using off-the-shelf tooling.

    Origin

    We are acutely aware of how important it is to host somebody's source code. I think everybody who reads and understands this blog post is just as aware of it. A company can grind to a halt if its developers cannot push or pull from its Git repositories. The productivity cost of five minutes of downtime in your CI system is hard to quantify in dollars, but it is, by any measure, a humongous amount.

    Agents have fundamentally changed the way we work with software, and in many ways they've made this situation worse. More code, more PRs, more CI runs. Version control is at the core of all of this, and it is possibly the hardest thing to change overnight.

    We've faced these difficulties internally at Cursor for many months now, and we've put considerable thought and care into building a platform that solves them for us and that can hopefully solve them for our customers too. Our focus right now is on providing the smoothest possible off-ramp into more reliability, more performance and more scale, and making the migration as painless as possible.

    Origin is not an experiment; it is the result of many decades of experience building these same systems, from people who deeply understand the magnitude of the challenges involved. We have an engineering and operational philosophy that has been proven to work, and a strong commitment to continue evolving it as the landscape of version control evolves.

    We're hoping you'll place your trust in us and our platform.

    Original source
  • Aug 17, 2026
    • Date parsed from source:
      Aug 17, 2026
    • First seen by Releasebot:
      Aug 17, 2026
    Cursor logo

    Cursor

    Origin Code Hosting

    Cursor launches Origin in early beta on paid plans, bringing hosted codebases, synced GitHub repos, pull requests, and code browsing into one place. It also adds agent-powered repo actions and app integrations for Vercel, Depot, and Buildkite.

    Cursor can now host your code.

    Origin begins rolling out today in early beta on all paid plans. We're starting with the essentials, designed for agent scale: repos, pull requests, code browsing, and GitHub sync. Agent-native features ship soon.

    Origin Repos

    The new Codebase tab is home for Origin repos.

    Click +New to create a new repo and name it. Once you do, a page shows you how to install the CLI, with commands for how to clone a repo or push a local project. Push, and your code is hosted on Origin.

    Name your codebase when you create your first repo. That name becomes part of every repo's URL: cursor.com/codebase/acme-corp.

    Bring your GitHub repos

    Your GitHub repos can sit alongside the ones Cursor hosts. Connect GitHub to Cursor, pick your org, and you'll see the repos you can sync. Select one and Cursor pulls it in. You choose what gets synced and can disconnect a repo at any time. Anyone with read or write access to a synced repo can view it in Cursor too.

    Synced repos update in real time. Browse, search, and pull from the copy in Origin. Pushes keep going to GitHub, which stays the source of truth for anything started there. Icons next to each repo name tell you which ones Cursor hosts and which came from GitHub.

    Pull requests

    Every repo has pull requests. Open one to see the timeline, commits, checks, and files changed. Review the diff, leave comments, and merge.

    Pull requests on synced repos sync both ways: comment in Cursor and it posts to GitHub, react or reply on GitHub and it shows up in Cursor within seconds. Got a review assigned to you on GitHub? Review and merge it from Cursor.

    Agents in every repo

    Your code, PRs, and agents are now in the same place. Ask Cursor questions about code you're browsing. It can answer, make changes, update PRs, or push a branch.

    App extensions for Cursor repos

    We're building an app ecosystem so your whole stack works seamlessly with Origin. Integrations with Vercel, Depot, and Buildkite are already available, with more coming soon.

    Connect Vercel from a repo's Apps tab and every PR gets a preview deployment where you can test and make comments. Merge, and it ships to production. For CI, connect Depot or Buildkite. Both run your existing GitHub Actions workflows and Buildkite also runs its native pipelines.

    Settings

    Every repo has settings. Check sync status for GitHub repos, manage who has access, and see which apps are connected.

    Origin is rolling out in early beta to all paid plan users starting today, except enterprise orgs whose admins opt out. Name your codebase and create your first repo.

    Learn more in our docs or get started today.

    Original source
  • Aug 13, 2026
    • Date parsed from source:
      Aug 13, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Cursor logo

    Cursor

    Cloud agents start 3x faster with builds

    Cursor introduces builds for Cloud Agents, preparing ready-to-use development environments in the background so agents start up to 3x faster and keep running from the last successful build when environments break. It also adds clearer build logs, status, and debugging in the dashboard.

    Agents are only as capable as the environments they run in. Fast, reliable development environments allow agents to take ambitious, long-running tasks from start to finish.

    Until now, every cloud session began with extensive setup: boot a machine, clone the repositories, and run the install script. On a large, complex repo, this just-in-time boot could take several minutes before the agent started executing.

    Today we're introducing builds: ready-to-use copies of your development environment that Cursor prepares continuously in the background, at no additional cost. When you kick off an agent, it starts in a ready environment so you get a response up to 3x faster.

    And when a bad commit or dependency update breaks your environment, agents keep using the last successful build. Your work continues uninterrupted while you debug in the background.

    Faster boot times

    A build is a copy of your development environment that Cursor prepares in the background. By default, Cursor runs a new build every hour. Instead of setting up the environment from scratch each session, agents boot into a ready version: repos cloned, dependencies installed, and the install script fully executed.

    When a build succeeds, it becomes the environment that future agents start from. Cursor keeps warm copies ready with new agents forking a live machine instead of restoring one from disk. This allows sessions to start almost instantly instead of keeping the next agent waiting.

    With environment setup already complete, agents get to real work much faster. At Cursor, our internal environments now boot 10x faster and time to first token is 3x faster.

    Our customers are seeing the same:

    We kick off more than 2,000 automated agent runs a week without any manual prompting. With builds, every run boots quickly into an environment we know is good and broken builds never take down the agent fleet. Our largest, most complex repos now start in just a few seconds.

    That combination of speed and reliability is what lets us hand more of our engineering work to agents that run entirely on their own.

    Blair McAlpine
    Senior Engineer, Faire

    More resilient agents

    Cloud agents always start from the latest successful build. If a dependency bump breaks your install script or a Docker build fails, that build never becomes active and you're notified of the issue. New and existing sessions keep running safely while you debug the environment in the background, either manually or with an agent.

    Better observability and easier debugging

    You can now inspect each build directly in your Cloud Agents dashboard, with:

    • A Builds tab for each environment, with type, status, start time, and versioning
    • Build details with logs and the exact commit SHAs the build captured
    • A record that ties each agent run to exactly the build it used
    • A configurable threshold for a build's git state so agents don't start too far behind your default branch

    Agents can also inspect and manage builds using the built-in Cursor Cloud MCP.

    Get started with builds today

    For an existing environment, open it in the Cloud Agents dashboard, go to the Builds tab, and click Enable Builds. Or click Run setup agent first to test the migration and review any proposed config changes.

    Because builds work by using filesystem snapshots, there are a few things worth checking at this stage:

    • Update your install command to cover anything that can be prepared ahead of time, like dependencies
    • If install needs credentials for private registries, use team or environment secrets. User secrets stay out of builds and are added when the agent starts.
    • The start command still runs when you first prompt an agent. Use it for services that must be fresh when the session begins, like bringing up Docker containers or other long-running processes

    On August 17th, all new and existing environments will use builds by default, with no additional cost to you.

    Learn more in our docs.

    Original source
  • Aug 13, 2026
    • Date parsed from source:
      Aug 13, 2026
    • First seen by Releasebot:
      Aug 13, 2026
    Cursor logo

    Cursor

    Cloud Agents Start 3x Faster with Builds

    Cursor introduces Builds for Cloud Agents, preparing ready-to-use development environments in the background so agents start faster and more reliably. It keeps successful builds warm, helps recover from broken environments, and adds build history and debugging tools in the dashboard.

    Agents do their best work when they start in a ready environment: repos cloned, dependencies installed, and your install script already run.

    This release introduces builds: ready-to-use copies of your development environment that Cursor prepares in the background. Agents boot into a ready environment instead of setting up from scratch each session. Builds are included with Cloud Agents at no additional cost.

    Faster starts

    Cursor runs a new build of your environment regularly. When a build succeeds, it becomes the environment future agents start from. Cursor keeps warm copies ready so the next agent does not wait. Internally, our environments now boot 10x faster, with 3x faster time to first token.

    Use your install command for anything that can be prepared ahead of time. The start command still runs when you first prompt an agent and should focus on services that need to be fresh in the session.

    More resilient agent runs

    When a bad commit or dependency update breaks your environment, agents keep using the last successful build. The broken build never becomes active, you are notified of the issue, and your agents keep working while you debug in the background.

    Build history and debugging

    Each environment has a Builds tab in the Cloud Agents dashboard. You can inspect build status, logs, commit SHAs, and which build each agent run used. Agents can also inspect and manage builds using built-in tools.

    Getting started

    New environments use Builds automatically. For an existing environment, open it in the Cloud Agents dashboard, go to the Builds tab, and click Enable Builds. Or click Run setup agent first to test the migration and review any proposed config changes.

    You can also trigger a Build manually, debug a failing Build with an agent, and control how stale Builds refresh with a configurable threshold.

    Learn more in our announcement post and docs.

    Original source
  • Aug 12, 2026
    • Date parsed from source:
      Aug 12, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Cursor logo

    Cursor

    Introducing Grok 4.6

    Cursor adds Grok 4.6, a new model focused on long-running agents, complex coding, and interactive visual work. It brings stronger multi-step task handling, improved self-checking, and available today in Cursor with 2x included usage for the first week.

    Today we are releasing Grok 4.6 together with SpaceXAI.

    Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.

    Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.

    Grok 4.6 is available today in Cursor and Grok Build. We’re offering 2x included usage inside Cursor and Grok Build for the first week.

    Training Grok 4.6

    Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. This produced a stronger foundation for the SFT and RL stages that followed.

    We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work. We filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior.

    Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more.

    Turning ambitious ideas into working projects

    We tested Grok 4.6 on projects designed to stretch its range and ability to sustain work over many steps. We found the model is especially strong at turning a broad product idea into a working first version. It can research unfamiliar domains, structure the application, implement the core interactions, and continue refining the result through several rounds of feedback.

    On longer trajectories, we also started to see more self-testing and verification, with the model checking its own work before moving on.

    Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. This has made it especially useful for projects where the fastest route to a good result was to begin with something substantial and then iterate in the loop.

    Safety and capabilities

    Grok 4.6's safeguards have been improved and calibrated in line with the model's capabilities.

    Our safety stack is designed to maximize utility and security across legitimate use cases, allowing Grok 4.6 to be helpful and safe in domains such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research.

    Our safeguard evaluation work reflects Grok 4.6’s expanded capabilities, with our widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, as well as extensive post-deployment third-party testing.

    Get started with Grok 4.6

    Grok 4.6 is available today in Cursor and Grok Build. It's also available in the SpaceXAI API and through partners including OpenRouter, Vercel, and Cloudflare.

    Pricing starts at $2 per million input tokens and $6 per million output tokens. A fast variant is available at twice the price.

    We’re including 2x usage inside Cursor and Grok Build for the first week.

    Original source
  • Aug 6, 2026
    • Date parsed from source:
      Aug 6, 2026
    • First seen by Releasebot:
      Aug 6, 2026
    Cursor logo

    Cursor

    How Cursor Router chooses the right model for the task

    Cursor launches Cursor Router with Auto Intelligence and Auto Balance, improving model routing to boost user satisfaction while lowering costs. The system now adapts from production traffic, adds Opus 5 to the mix, and keeps moving closer to the model frontier.

    On July 22, we launched Cursor Router with two new configurations, Auto Intelligence and Auto Balance

    Since then, we have continued improving both modes as new models have arrived and our routing system has learned from more production traffic.

    Today, Auto Intelligence delivers above Fable-level user satisfaction at 68% lower cost, a further 18% reduction since its launch. Auto Balance outperforms Opus 4.8 at 41% lower cost, a further 8% reduction over the same period, while further increasing user satisfaction by 3%.

    We're working towards a Cursor Router that improves alongside the model frontier. This post explains how the current system works.

    Cursor Router increases satisfaction and cuts cost vs. frontier models

    Numbers in graph reported relative to Opus 4.8

    A data-driven approach to routing

    Cursor Router is built around the idea that model selection should be learned from how models perform on real developer work, rather than inferred from benchmark scores.

    The router makes each decision using signals from the current turn and recent conversation state. These include structured features such as the task category, along with recent tool calls and the broader context of the work.

    From there, routing happens in two parts.

    First, we need to decide whether a turn is simple enough for a price-efficient model. Compass, our complexity predictor, makes this decision.

    Second, if the turn is more demanding, we need to decide which frontier model is most likely to perform well on that kind of work. To make that decision, we classify the turn using a taxonomy of tasks, domains, and modifiers learned from real developer traffic.

    Building a dataset

    To develop the routing system, we first needed a dataset that reflected the conditions it would encounter in production. We built it from live Cursor traffic so it would preserve the actual mix of developer tasks, the context surrounding each turn, and the effects of switching between models.

    As always, we respected users' privacy mode and data retention settings throughout this process.

    The dataset contains hundreds of thousands of turns sampled across a range of models. Each datapoint includes the conversation signals available to the router, along with two outcomes we use to compare routing choices.

    1. Performance. We infer performance from what the user does next. Moving on to the next task is a strong positive signal, while correcting the agent is a strong negative one.
    2. Cost. We calculate cost from API pricing and token usage for that turn. Because the data comes from live traffic, it also captures costs that benchmarks often miss, including cache misses caused by switching models.

    Predicting complexity with Compass

    Compass estimates the complexity of each turn by predicting whether the user will be satisfied with Cursor's response. We train it on the performance signal mentioned above.

    We use the resulting prediction as a proxy for complexity. This works because users rarely ask for corrections after simple tasks, like making a commit, while they're more likely to make follow-up requests when the work is more complex.

    We evaluated Compass online and confirmed that its scores are strong predictors of user satisfaction. Turns that Compass rated as most likely to succeed received a positive performance signal 96% of the time, while turns it rated as least likely to succeed received one 71% of the time.

    In practice, Compass assigns each turn a continuous complexity score between 0 and 1. We set a threshold within that range to determine which turns stay on a price-efficient model and which are upgraded to a frontier model. Lower thresholds keep more traffic on the price-efficient model, while higher thresholds upgrade more often.

    Learning model strengths

    After Compass tells us when a turn is complex enough to justify using a frontier model, the next question is which frontier model to use.

    To answer it, we built a taxonomy from real developer traffic that describes each turn across three dimensions:

    • Domains identify where the work happens: backend, database schemas, frontend
    • Tasks identify what the developer wants done: fixing bugs, running commands, writing tests
    • Modifiers capture characteristics that cut across domains and tasks, but may change which model performs best: bounded edits, product questions, visual-heavy changes

    We then compare how different models perform across those categories. We found that no model dominates every kind of work, and each has categories where it outperforms:

    • Grok offers strong value across broad, routine work. Its low inference cost made it especially effective for categories such as Git commands and general database operations.
    • Sol performs especially well on planning and codebase comprehension. It also delivered strong results across several implementation tasks at a lower cost than other frontier models.
    • Opus performs well on execution-heavy work. It showed particular strengths in devops, database queries, and performance optimization.
    • Fable excels at debugging and visual implementation. Its quality gains were most valuable on complex tasks where they justified its higher cost.

    Cursor Router uses those differences to match each turn to the model best suited to it.

    Combining into an algorithm

    Compass and the taxonomy play complementary roles. Compass estimates the model-agnostic complexity of the turn and compares that score with a routing threshold. Depending on where the score falls, we either send the turn to Grok, given its low inference cost, or use the taxonomy to identify the frontier model with the strongest observed performance on that kind of work.

    When Compass does send a turn to the taxonomy router, model selection follows two rules:

    1. Only route when performance is clearly better. A candidate model becomes eligible only when its observed performance on that task label clears a one-sided 75% uplift threshold against the price-efficient model. Roughly, this means we need 75% confidence that the improvement is real.
    2. Choose the best mix within the budget. From the eligible candidates, the optimizer chooses the traffic-weighted combination expected to deliver the largest performance gain while keeping the average cost per turn within the mode's budget.

    Together, the Compass threshold and the task router's cost budget define each mode's position on the cost-performance curve. Auto Balance keeps more traffic on the price-efficient path and gives the task router a smaller budget. Auto Intelligence gives the task router more room to select frontier models when the expected performance gain justifies the cost.

    Evaluating performance in production

    We evaluated our routing policies in two stages. First, we used cross-validation to tune the Compass thresholds and optimization budgets without overfitting to a particular split. We then evaluated the selected policies on a held-out test set that had not been used during training.

    This gives us a more reliable estimate of how each policy should perform on new traffic. It helps us eliminate weak candidates and compare expected cost and performance before deployment. But offline analysis still cannot fully capture how a policy will behave in production, and benchmarks are limited for the same reason. Live developer traffic remains the most representative test.

    Offline evaluation surfaces candidate policies to test online

    *Cost and performance relative to Opus 4.8

    We then tested the policies on live traffic, where we could measure user satisfaction and the actual cost of each turn under production conditions. This captures effects that are difficult to model offline, including token usage, caching, and the cost of switching between models.

    Before launch, we tested both modes on live traffic and found that each improved the cost-performance tradeoff relative to individual frontier models. Auto Balance delivered higher satisfaction than Opus 4.8 at lower cost, while Auto Intelligence approached Fable-level satisfaction at a much lower cost.

    We have since repeated this process as the routing system and available models have improved, moving both modes further beyond the cost-performance frontier.

    Keeping pace with the model frontier

    Since launching Cursor Router, we've added Opus 5 to the routing mix and improved Compass's predictions. That gives the router both a stronger set of models to choose from and a better signal for deciding when each one is worth using.

    Over time, we want the router to become more adaptive by predicting each model's expected quality and cost, learning from production outcomes, and updating continuously. As the system improves, Cursor users will be able to benefit from frontier models where they're needed most, without paying frontier-model prices on every turn.

    Read more in our docs.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.