Feature Flagging and Experimentation Release Notes
Release notes for feature flagging, A/B testing and experimentation platforms
Products (17)
Latest Feature Flagging and Experimentation Updates
- Sep 2, 2026
- Date parsed from source:Sep 2, 2026
- First seen by Releasebot:Sep 2, 2026
v2.269.0
Flagsmith adds cohort sync improvements, including CSV re-synchronisation, sync state visibility, managed segment updates, and a generic cohort sync webhook. It also allows uppercase trait keys, adds flag dependencies in evaluation context, and includes several bug fixes and dependency updates.
2.269.0 (2026-09-02)
Features
- allow uppercase characters in trait keys (#8403) (4d39c39)
- cohort segment detail view with CSV re-synchronisation (#8387) (7cb90aa)
- cohort synchronisation keys and provider connect modals (#8420) (9dd6141)
- cohorts: expose sync state and allow updating the managed segment (#8386) (bbc4bd1)
- generic cohort sync webhook (#8380) (0f0e09a)
- ingestion: Deny non-HTTPS traffic to S3 buckets (#8411) (8044b6d)
- sdk: Add flag dependencies to the evaluation context (#8396) (58eb348)
Bug Fixes
- 500 creating a cohort sync key with a master API key (#8390) (17243dc)
- Amplitude list creation response shape (#8402) (e7ffc35)
- hide Try it out until the project has 3+ flags (#8419) (50e0309)
- project admin feature import permissions (#8421) (cf77245)
- queue membership count refreshes for cohort segments (#8438) (37b5bdf)
- return only listId from Amplitude list creation (#8400) (e252a36)
- return the Amplitude list ID at every depth their parsers read (#8399) (ab19267)
- usage: separate the plan meter from the filters (#8357) (93a6d02)
Dependency Updates
- Bump flagsmith to 6.2.3 (#8423) (798b675)
- Bump flagsmith-flag-engine to 11.0.0 (#8410) (c3cfb47)
CI
- frontend-deploy: announce a deploy when it starts (#8398) (894bfa4)
- frontend-deploy: report every deploy outcome in one Slack message (#8371) (cb66498)
Refactoring
- forms: adopt FieldLabel and FieldError for hand-rolled labels (#8117) (cf98e07)
- styles: diff view and code theme on shared tokens (#8069) (76da16c)
- types: name the trait value shape (#8366) (bce2bf4)
- Sep 1, 2026
- Date parsed from source:Sep 1, 2026
- First seen by Releasebot:Sep 2, 2026
Version 4.41.3
ABTasty fixes session cookie encoding and segment value formatting for more secure browsing and accurate data handling.
AI Generated Release Note
Bug Fixes
- Session Cookie Encoding: Improved encoding of special characters within session cookies to ensure secure browsing experience.
- Segment Value Formatting: Corrected the formatting of segment hit values, ensuring accurate data representation and processing.
All of your release notes in one feed
Join Releasebot and get updates from Flagsmith and hundreds of other software products.
- Aug 27, 2026
- Date parsed from source:Aug 27, 2026
- First seen by Releasebot:Aug 27, 2026
v2.268.0
Flagsmith adds experimental warehouse auto-connect for new orgs and flag-driven plan gating for free orgs.
2.268.0 (2026-08-27)
Features
- api: auto-connect Flagsmith warehouse for new orgs via experimental_flags value (#8377) (27dbc25)
- fe: flag-driven warehouse plan gating for free orgs (#8379) (2be8a90)
- Aug 27, 2026
- Date parsed from source:Aug 27, 2026
- First seen by Releasebot:Aug 27, 2026
Unlock more learning with every experiment
Growthbook launches Learnings to turn experiment results into a living knowledge library, helping teams capture evidence-backed insights from experiments and research, scope them by project or tag, and share them with people and AI agents through APIs, MCP, and GrowthBook Skills.
Running an experiment gives you an answer to a question. Running thousands of experiments gives you a lot of answers, but also something much more valuable: a body of evidence about how your product, users, and business actually behave.
The problem is that this knowledge is surprisingly difficult to use in practice.
A PM working on onboarding might not know that another team tested a similar idea six months ago. An engineer building a new checkout flow might not know which patterns have consistently helped conversion on other parts of the product. Even when people remember that relevant experiments exist, reading through dozens or hundreds of them to find the important patterns is rarely practical.
Today, we are launching Learnings in GrowthBook to help solve this problem.
Learnings let teams capture what they have learned across experiments, user research, and other sources of evidence, then make that knowledge available to both people and AI agents when they are building something new.
Learnings: Turn experiment learnings into a living knowledge library.
Experiments produce more than winners
Experiments produce more than winners
The most obvious output of an experiment is a decision. Those decisions create incremental gains, and over time those gains compound.
But there is another output from experimentation that is easier to overlook: knowledge. Every experiment produces it, including the ones that lose or move nothing at all.
You might learn that:
- Showing pricing earlier in the funnel consistently improves qualified conversions.
- Simplifying onboarding helps new users but hurts activation for experienced users.
- Social proof matters on acquisition pages but has little effect inside the product.
- Asking users to configure everything upfront creates friction, while progressive configuration performs better.
None of these conclusions necessarily come from a single experiment. They emerge after five, twenty, or a hundred experiments, once somebody notices the pattern.
This is where an experimentation program becomes more powerful than a sequence of isolated A/B tests: individual tests give you answers, but the program gives you a model of how your users behave. You are gradually identifying patterns and putting your learnings to work for future growth.
Learnings compound your experimentation program
Learnings compound your experimentation program
If an experiment improves conversion by 2%, that improvement can continue generating value for as long as the change remains in the product (novelty effects aside). That’s real compounding, and it’s the return most programs measure.
But that win only compounds one thing: a metric. A learning acts on something different: the quality of the next decision.
Imagine your team learns through repeated experiments that users perform better when complex actions are introduced progressively rather than all at once. That insight might influence your next onboarding flow. Then your settings experience. Then a new AI feature. Then the way an agent designs a workflow six months later.
It doesn’t stay with one team either. An insight can travel to whoever picks your onboarding feature or settings work next. Your experimentation program can widen, since everyone starts from organizational knowledge.
The value is not limited to the experiment that produced the learning. It changes the starting point of future work. Instead of beginning every project from first principles, your team starts with a set of evidence-backed assumptions about what tends to work (and what doesn’t).
Learnings in GrowthBook
Learnings in GrowthBook
GrowthBook Learnings are designed to capture this organizational knowledge explicitly.
A learning can reference evidence from multiple experiments, rather than being tied to the result of a single test. That matters because many useful conclusions only become visible across a collection of experiments.
The New Learning form in GrowthBook, with fields for title, description, status, tags, projects, and supporting and contradicting experiments.
Every learning can cite the experiments that support it and the ones that don't, then be scoped to the projects and tags where the pattern actually held
You can use Learnings to document things like:
- Patterns that repeatedly improve a metric
- Approaches that consistently fail
- Differences between user segments
- Design principles supported by experimentation
- Unexpected behaviors observed across multiple tests
- Areas where the evidence is contradictory or uncertain
The same change isn’t universally applicable, though. Streamlining a flow can lift conversion in one part of your product and lower it in another, and the right amount of friction depends on what the user is trying to do. Learnings can be scoped to specific projects or tags, so that patterns that have been tested for your self-serve signup aren’t applied to your enterprise onboarding flow.
And learnings do not have to come exclusively from experiments either. You can capture qualitative findings from user research, customer interviews, support conversations, or other sources and combine them with quantitative evidence.
The goal is not to turn every observation into an immutable rule.
It is to give your organization a shared, evidence-backed memory. Each learning can be updated if new evidence comes in, and also have a specific status for when the learning is not verified yet, or if it’s no longer relevant. Statuses for learnings are entirely customizable. Once captured this way, that memory becomes something both people and AI agents can draw on, which raises the question of how much of it to hand an agent at once, and in what form.
The context window problem for organizations
The context window problem for organizations
There is a useful analogy to working with AI coding tools.
If you give an agent every line of code your company has ever written, you have technically given it more information. But you have not necessarily given it better context.
The useful question is:
what does the agent actually need to know to make this decision well?
Organizations have the same problem. After thousands of experiments, nobody should need to read thousands of experiment reports before starting a project. They need the relevant conclusions.
Learnings act as a compressed context layer over your experimentation history. The underlying experiments are still there as evidence, but people and agents can work from the higher-level patterns that those experiments have established.
For example:
Learning:
Users are more likely to complete complex setup flows when advanced configuration is deferred until after initial success.Evidence:
Seven onboarding experiments across three product areas.A PM planning a new onboarding experience can start with that knowledge instead of rediscovering it. An engineer can incorporate it while designing the implementation. The same applies to agents, even more strongly. An agent without access to your evidence will still produce a confident answer, but it will be generic and maybe an approach you’ve already disproved. Point your agents at your Learnings, and they start from what you already know.
Learnings are available through GrowthBook's APIs and MCP support, just like your experiments and other experimentation data. Through GrowthBook Skills, you can instruct agents to retrieve relevant Learnings before they design or implement something.
Instead of giving agents generic product-development best practices, you can give them exactly what they need to know to make this decision well: your company’s own compressed, decision-relevant knowledge, grounded in real outcomes and updated as new evidence comes in.
AI can help find the patterns humans miss
AI can help find the patterns humans miss
As experimentation programs grow, manually identifying patterns becomes harder if not impossible. A team running ten experiments a year can probably remember most of them. A company running thousands cannot. GrowthBook can use AI to analyze your experiment history and surface commonalities across results.
GrowthBook's Find Learnings panel proposing a learning about friction reduction in conversion flows, citing three supporting experiments.
GrowthBook proposes candidate learnings from your experiment history, with the supporting evidence and a suggested next step attached. You decide which ones to keep.
Perhaps a certain type of messaging consistently works for new users but not existing customers. Maybe several unrelated experiments show that reducing perceived commitment improves activation. Maybe a UI pattern that teams keep proposing has actually failed in four different areas of the product.
These patterns are easy to miss when every experiment is analyzed independently.
AI makes it possible to search across a much larger body of evidence, while Learnings give teams a place to review, refine, and preserve the conclusions that matter.
Good experiment hygiene becomes even more valuable
Good experiment hygiene becomes even more valuable
There is an important prerequisite.
AI cannot infer very much from an experiment called "Homepage test 7" with no hypothesis, description, or conclusion.
The better your experiment documentation, the more useful your accumulated knowledge becomes.
This is one reason GrowthBook has increasingly invested in experiment quality and hygiene, including checklists and workflows that encourage teams to document the hypothesis, context, results, and conclusions of an experiment. Good documentation has always made individual experiments easier to understand.
With Learnings, it also makes the entire experimentation history more valuable. Every well-documented experiment becomes another piece of evidence that can contribute to future decisions.
From experimentation history to organizational memory
From experimentation history to organizational memory
The long-term value of experimentation is not just making better decisions today. It is making every future decision from a stronger starting point.
Experiments that win improve the product. All your experiments improve the ideas, assumptions, and decisions that come next, increasingly including the ones your agents make. That’s the part that really compounds. Firmer footing leads to better experiments, which produce better evidence, which produces better judgment. Over time, your experimentation program is not only improving your product, it’s improving your organization’s ability to build one.
Experiments should not disappear into a results archive once a decision has been made. The best ones should keep teaching you.
Original source - August 2026
- No date parsed from source.
- First seen by Releasebot:Aug 27, 2026
Go AI SDK reference
LaunchDarkly releases a Go AI SDK reference for AgentControl, adding completion, agent, and judge config modes plus new tracker methods for AI metrics, feedback, and graph traversal. It also replaces the older Config and Tracker APIs and supports resumption across processes.
This topic documents how to get started with the Go AI SDK, and links to reference information on all of the supported features.
The Go AI SDK is designed for use with AgentControl. It is in a pre-1.0 release and the API may change based on feedback. You can follow development or contribute on GitHub.
This version replaces the previous Config and Tracker API
If your codebase calls aiClient.Config() or tracker.TrackRequest(), those methods have been deprecated. This version introduces separate completion, agent, and judge config modes, along with a new set of tracker methods. Review this reference before you upgrade.
SDK quick links
LaunchDarkly’s SDKs are open source. In addition to this reference guide, we provide source, API reference documentation, and sample applications:
Get started
LaunchDarkly AI SDKs interact with AgentControl configs. Configs are the LaunchDarkly resources that manage model configurations and messages for your generative AI applications.
Try the Quickstart
This reference guide describes working specifically with the Go AI SDK. For a complete introduction to LaunchDarkly AI SDKs and how they interact with configs, read Quickstart for AgentControl.
You can use the Go AI SDK to customize your config based on the context that you provide. This means both the messages and the model evaluation in your generative AI application are specific to each end user, at runtime. You can also use the AI SDK to record metrics from your AI model generation, including duration and tokens, and to evaluate model output with judges.
Follow these instructions to start using the Go AI SDK in your application.
Install the SDK
First, install the AI SDK as a dependency in your application. How you do this depends on what dependency management system you are using:
- If you are using the standard Go modules system, import the SDK packages in your code and go build will automatically download them. The SDK and its dependencies are modules.
- Otherwise, use the go get command and specify the SDK version, such as go get github.com/launchdarkly/go-server-sdk-ai.
The Go AI SDK is built on the Go SDK, so install that as well.
Here is how:
import ( ld "github.com/launchdarkly/go-server-sdk/v7" "github.com/launchdarkly/go-server-sdk-ai/ldai" )Initialize the client
After you install and import the SDK, create a single, shared instance of LDClient. Then, use it to initialize the AI client. The AI client is how you interact with configs. Specify the SDK key to authorize your application to connect to a particular environment within LaunchDarkly.
The Go SDK uses an SDK key
The Go SDK uses an SDK key. Keys are specific to each project and environment. They are available on the SDK keys page under Settings. To learn more about key types, read Keys.
Here is how:
client, _ := ld.MakeClient("YOUR_SDK_KEY", 5*time.Second) aiClient, err := ldai.NewClient(client) if err != nil { // Client couldn't be created }This example assumes you have imported the LaunchDarkly SDK package as ld, as shown above.
Best practices for error handling
The second return type in these code samples (_) represents an error in case the LaunchDarkly client does not initialize. Consider naming the return value and using it with proper error handling.
Configure the context
Next, configure the context that will use the config, that is, the context that will encounter generated AI content in your application. The context attributes determine which variation of the config LaunchDarkly serves to the end user, based on the targeting rules in your config. If you are using template variables in the messages in your config’s variations, the context attributes also fill in values for the template variables.
Here is how:
context := ldcontext.NewBuilder("example-context-key"). Kind("user"). Name("Sandy Smith"). SetString("email", "[email protected]"). SetValue("groups", ldvalue.ArrayOf(ldvalue.String("Acme"), ldvalue.String("Global Health Services"))). Build()Customize a config
Then, use one of the config retrieval methods to customize a config. Customization means that any variables you include in the messages when you define the config variation have their values set to the context attributes and variables you pass in. The AI SDK provides three config modes, completion, agent, and judge. You set the mode for a particular config when you create it in the LaunchDarkly UI.
The customization process within the AI SDK is similar to evaluating flags in one of LaunchDarkly client-side, server-side, or edge SDKs, in that the SDK completes the customization without a separate network call. If it cannot perform the evaluation or LaunchDarkly is unreachable, it returns the fallback value you provide. For example, you might use an empty, disabled default as a fallback value, or a fully configured default. Either way, you should make sure to check for this case and handle it appropriately in your application.
All three config modes share a common set of methods through an embedded base: Key(), Enabled(), Model(), ModelName(), Provider(), ProviderName(), and CreateTracker().
Customize configs in completion mode
In completion mode, each variation in your config includes a single set of roles and messages used to prompt your generative AI model. Use CompletionConfig to customize the config.
The CompletionConfig method takes a config key, a context, a fallback value, and optional variables. It performs the evaluation, then returns an AICompletionConfig object with the customized messages and model configuration.
Here is how:
fallbackValue := ldai.NewAICompletionConfigDefault(). WithEnabled(true). WithModelName("claude-sonnet-4-6") config := aiClient.CompletionConfig( "example-config-key", context, fallbackValue, map[string]interface{}{"exampleCustomVariable": "exampleCustomValue"}, )After you call CompletionConfig, you can pass the customized messages directly to your AI provider. To learn more, read Customizing AgentControl configs.
Customize configs in agent mode
In agent mode, use AgentConfig or AgentConfigs to customize the config. The AgentConfig method customizes a single agent config. The AgentConfigs method customizes a batch of them. Agent configs add Instructions() and JudgeConfiguration() on top of the shared base methods.
Here is how:
fallbackValue := ldai.NewAIAgentConfigDefault(). WithEnabled(true). WithModelName("claude-sonnet-4-6") agent := aiClient.AgentConfig("example-agent-key", context, fallbackValue, variables) instructions := agent.Instructions()Customize model parameters
Every default value supports builder-style setters that customize the model, provider, and underlying parameters before evaluation:
fallbackValue := ldai.NewAICompletionConfigDefault(). WithEnabled(true). WithModelName("claude-sonnet-4-6"). WithProviderName("anthropic"). WithModelParam("temperature", 0.7). WithCustomModelParam("top_k", 40). WithTool(myTool)Disable a config by default
Provide a fallback value with Enabled set to false so the client falls back to your default behavior if the flag targeting rules do not enable the config, or if LaunchDarkly is unreachable.
Use template configs
CompletionConfigTemplate, AgentConfigTemplate, and JudgeConfigTemplate skip Mustache interpolation and do not accept a variables parameter. Use a template method if you plan to interpolate message content yourself, or if a config has no placeholders to fill.
Evaluate input and output pairs with a judge
Use JudgeConfig to retrieve a judge config. Judge configs add Messages() and EvaluationMetricKey() to the shared base methods.
Judges are constructed directly, not through the client
Unlike completion mode and agent mode, the AI SDK does not expose a client-level method to create a judge. Construct one directly from the judge subpackage, and implement the Provider interface yourself so the judge can call your model.
Here is how:
judgeConfig := aiClient.JudgeConfig("example-judge-key", context, fallbackValue, variables) type myProvider struct{} func (p *myProvider) InvokeStructuredModel( messages []datamodel.Message, schema map[string]interface{}, ) (judge.StructuredResponse, error) { // Call your AI provider and return a structured response. } j, err := judge.New(judgeConfig, tracker, &myProvider{}, "example-judge-key", loggers)Call Evaluate to score a single input and output pair, or EvaluateMessages to evaluate a full message list against a response. samplingRate is a value between 0.0 and 1.0. If Evaluate skips the call because of sampling, or because the judge config is empty, it returns nil, nil rather than an error:
result, err := j.Evaluate(input, output, samplingRate)Recording the judge response is your responsibility, not the judge’s. Call tracker.TrackJudgeResponse yourself after Evaluate or EvaluateMessages returns. The judge does not call it for you.
Call provider, record metrics from AI model generation
TrackDuration, TrackSuccess, TrackTimeToFirstToken, and TrackTokens are at-most-once, which means LaunchDarkly records only the first call for each. TrackSuccess and TrackError are also mutually exclusive, the first one you call wins. TrackTokens replaces the deprecated TrackUsage method, use TrackTokens in new code.
Here is how:
if config.Enabled() { tracker, err := config.CreateTracker() // Make a request to a provider using details from config. // For example, you can pass model parameters (config.ModelParam) or messages (config.Messages). tracker.TrackSuccess() tracker.TrackDuration(elapsed) tracker.TrackTokens(tokenUsage) } else { // Application path to take when the config is disabled. }Each tracker call shares a single runId, a UUIDv4 created when you create the tracker. The runId correlates every event you record on that tracker as one AI run.
Alternatively, you can use TrackDurationOf to wrap a function call and measure its wall-clock duration automatically. For applications that require streaming, use the package-level TrackMetricsOf function to wrap an operation that returns a value and an error. Pass the tracker to TrackMetricsOf along with a function that extracts an AIMetrics value from the operation result. TrackMetricsOf records the operation’s success or error, duration, and tokens, and is the recommended replacement for the deprecated TrackRequest method.
TrackFeedback is multi-fire. Call it as many times as you receive feedback, for example once per user reaction. TrackJudgeResponse is also multi-fire. Use it to record a judge evaluation result.
Go names this tracker method differently than Java and .NET
Go names this method TrackJudgeResponse. The Java and .NET AI SDKs name the equivalent method TrackJudgeResult.
Use GetSummary to read the metrics recorded on a tracker so far, ResumptionToken to get a token for reconstructing the tracker later, and GetTrackData to read the raw track data.
To learn more, read Tracking AI metrics.
Build and traverse agent graphs
An agent graph links multiple agent configs together into nodes and edges. Retrieve a graph definition with AgentGraph:
graph := aiClient.AgentGraph("example-graph-key", context, variables)AgentGraphDefinition exposes Enabled(), RootNode(), GetNode(key), GetChildNodes(key), GetParentNodes(key), TerminalNodes(), and CreateTracker(), which returns a *GraphTracker.
Each AgentGraphNode exposes Key(), Config(), Edges(), and IsTerminal(). Each GraphEdge exposes the target node’s Key() and a Handoff() map of values passed along that edge.
Traverse and ReverseTraverse walk the graph in topological order. Traverse visits a node only after every reachable predecessor of that node has been visited, starting from the root. ReverseTraverse visits a node only after every reachable descendant has been visited, ending at the root. Both orderings are cycle-safe and deterministic, and give each callback its own dependency-scoped context. LaunchDarkly never mutates the initialContext value you pass in:
graph.Traverse(func(node *ldai.AgentGraphNode, context map[string]interface{}) interface{} { // Run the node. return result }, initialContext)GraphTracker records graph-level and edge-level metrics separately. Graph-level methods, such as TrackInvocationSuccess, TrackInvocationFailure, TrackDuration, TrackTotalTokens, and TrackPath, are at-most-once. TrackInvocationSuccess and TrackInvocationFailure are mutually exclusive, so the first call wins, and TrackDuration ignores non-finite values. Edge-level methods, such as TrackHandoffSuccess, TrackHandoffFailure, and TrackRedirect, are multi-fire.
Use GetSummary to read a GraphMetricSummary, and ResumptionToken to get a token for reconstructing the tracker later.
Resume tracking across processes
Both trackers support resumption, so you can start tracking an AI run in one process and continue it in another.
For a completion, agent, or judge tracker, call CreateTracker on the client with a saved token. For a graph tracker, call CreateGraphTracker on the client, or the package-level TrackerGraphFromResumptionToken function directly:
tracker, err := aiClient.CreateTracker(token, context) graphTracker, err := aiClient.CreateGraphTracker(token, context)Reconstructing a tracker from a resumption token preserves the original runId, so at-most-once guards still apply across the resumed run.
Supported features
This SDK supports the following features:
- Anonymous contexts
- Context configuration
- Customizing AgentControl configs
- Private attributes
- Tracking AI metrics
- August 2026
- No date parsed from source.
- First seen by Releasebot:Aug 27, 2026
AgentControl
LaunchDarkly introduces AgentControl, an add-on for managing LLM configs outside application code so teams can customize prompts, test variations, run judges, and roll out AI changes safely with targeting, monitoring, and experiments.
AgentControl
The topics in this category explain how to use LaunchDarkly AgentControl to manage your configs. You can use AgentControl to customize, test, and roll out new large language models (LLMs) in your generative AI applications.
An AgentControl config is a single resource that you create in LaunchDarkly to control how your application uses large language models. It lets teams manage prompts, instructions, and model settings outside of application code so they can iterate, experiment, and release changes more safely without redeploying. To learn how to create one, read Create configs.
Choose a configuration mode
When you create a config, you select a configuration mode that defines how the model behaves in your application.
AgentControl supports two modes:
- Completion mode: Configure prompts using messages and roles for single-step model responses. To learn more, read Create and manage config variations.
- Agent mode: Configure multi-step workflows using structured instructions. Agent mode does not create a separate resource. To learn more, read Agents.
You can attach judges to both completion-mode and agent-mode config variations directly in the LaunchDarkly UI. You can also invoke a judge programmatically using the AI SDK.
Both modes use the same config resource and support variations, targeting rules, monitoring, experimentation, and lifecycle management.
Both completion mode and agent mode can integrate with external tools or APIs. Tool usage depends on how your application and SDK are implemented, not on the selected configuration mode. Agent mode enables structured, multi-step workflows. You can integrate external tools in either mode.
With AgentControl, you can:
- Manage model configuration outside of your application code so you can update prompts and settings at runtime without deploying changes.
- Upgrade to new model versions and roll out changes gradually and safely.
- Add new model providers and progressively shift production traffic between them.
- Compare variations to determine which performs better based on cost, latency, satisfaction, or other metrics.
- Run experiments to measure the impact of generative AI features on end-user behavior.
AgentControl supports advanced use cases such as retrieval-augmented generation, integration with external tools or APIs, and evaluation in production. You can:
- Track which knowledge base or vector index is active for a given model or audience.
- Experiment with different chunking strategies, retrieval sources, or prompt and instruction structures.
- Evaluate outputs using side-by-side comparisons or online evaluations with judges in completion mode or agent mode, or invoke a judge programmatically using the AI SDK.
- Build guardrails into runtime configuration using targeting rules to block risky generations or switch to fallback behavior.
- Apply different safety filters by user type, geography, or application context.
- Use live metrics, including satisfaction and quality signals you define, to guide rollouts.
These capabilities let you evaluate model behavior in production, run targeted experiments, and adopt new models safely without being locked into a single provider or manual workflow.
If you use an AI agent to create and manage configs, you can use LaunchDarkly agent skills to help AI coding agents execute common tasks safely and consistently.
Availability
AgentControl is an add-on feature. Access depends on your organization’s LaunchDarkly plan. If AgentControl does not appear in your project, your organization may not have access to it.
To enable AgentControl for your organization, contact your LaunchDarkly account team. They can confirm eligibility and assist with activation.
For information about pricing, visit the LaunchDarkly pricing page or contact your LaunchDarkly account team.
How AgentControl works
Every config contains one or more variations. Each variation defines model settings with messages for completion mode or instructions for agent mode. You define targeting rules to control which variation LaunchDarkly serves to a given context.
In your application, you use one of LaunchDarkly’s AI SDKs to evaluate a config for a given context. The LaunchDarkly SDK evaluates targeting rules and selects a variation. The AI SDK plug-in then uses that variation to return the resolved configuration, including model settings and messages or instructions.
As part of this evaluation, the AI SDK resolves any variables in your prompts using context attributes and additional variables you provide. This enables you to tailor prompts and model settings for each context at runtime. When you update prompts, instructions, or model configuration in LaunchDarkly, those changes take effect immediately without requiring you to redeploy your application.
LaunchDarkly does not invoke model providers on your behalf. Your application is responsible for calling the model provider directly using its own credentials and the configuration returned by the AI SDK. LaunchDarkly does not proxy or independently invoke model providers.
After your application calls the model provider, use the AI SDK to track AI metrics such as generation count, token usage, latency, errors, and evaluation scores. LaunchDarkly aggregates these metrics and displays them on the Monitoring tab.
The topics in this category explain how to create configs and variations, update targeting rules, monitor related metrics, and incorporate AgentControl into your application.
Additional resources
In this section:
Set up AgentControl configs
- Quickstart for AgentControl
- Create configs
- Create and manage config variations
- Create and manage AI model configurations
- Tools
- Prompt snippets
Config evaluations
- Playgrounds
- Offline evaluations
- Datasets
- Online evaluations
- Judges
- Run experiments with AgentControl
Agents
- Agents
- Agent graphs
Deliver and monitor configs
- Config targeting
- Monitor config performance
- Understand AI impact with AI Insights
- Manually instrument LLM spans
Manage AgentControl configs
- Manage AgentControl configs
- Compare config variation versions
- AgentControl and information privacy
In our guides:
- Managing AI model configuration outside of code
- Using targeting to manage AI model usage by tier
In our SDK documentation:
- .NET AI SDK reference
- Go AI SDK reference
- Node.js (server-side) AI SDK reference
- Python AI SDK reference
- Ruby AI SDK reference
- August 2026
- No date parsed from source.
- First seen by Releasebot:Aug 27, 2026
LLM observability
LaunchDarkly adds LLM observability and conversation views to capture spans, stitch agent runs into readable transcripts, and help teams monitor latency, tokens, costs, errors, and model outputs across traces and configs.
How LLM observability works
This topic explains how LaunchDarkly captures and displays large language model (LLM) spans, and how it groups related LLM spans into conversations. You can use LaunchDarkly LLM observability features to monitor the performance of your models in production and diagnose problems with them.
LLM observability helps your team:
- Optimize LLM latency by monitoring token usage and request duration
- Investigate provider errors
- Compare model outputs by reviewing prompt and response pairs across environments
- Read what an agent did, to determine the turn where it went wrong
- Compare cost and latency across different agent runs
- Analyze downstream impact by connecting LLM spans with session or error data
How LLM observability works
When your application calls an LLM provider:
- Instrumentation in the LaunchDarkly observability SDK captures telemetry about the model request.
- The SDK exports LLM telemetry as span attributes in OpenTelemetry traces.
- LaunchDarkly records the LLM spans for display on the Traces page.
Each span includes the detailed information you need to evaluate model behavior across environments, such as the model name, prompt and response content, token usage, request duration, and provider information.
LaunchDarkly marks LLM spans with a green indicator labeled “LLM” in the traces view.
About LLM conversations
A single agent run rarely fits into one trace. A typical agent run involves a variety of activities that generate multiple traces over a period of time, such as:
- Answering a question
- Calling one or more tools to perform tasks or gather information
- Waiting for a person to reply, or for a tool call to return data
- Repeating these actions minutes or hours later after new information becomes available
Because each of these activities arrives as a separate trace, reading a single, logical conversation with an LLM requires organizing and connecting the different activities span-by-span.
LaunchDarkly automatically stitches together all spans that share a common conversation identifier. It orders the messages and tool calls into a single, readable transcript. LaunchDarkly also rolls up key attributes such as the duration, token usage, models, and errors across the entire agent run.
You can use LaunchDarkly LLM conversations to browse and display full agent runs as they occurred in response to user prompts and tool responses. Conversations also work as a starting point to dive into span information at any turn in the conversation to learn details about problems or errors that occurred when interacting with an agent.
Set up LLM observability
To set up LLM observability, you configure a LaunchDarkly observability SDK in your application to instrument the generative AI attributes LaunchDarkly reads on each span. The steps vary by SDK and by LLM provider.
To learn more about instrumenting your application so LLM spans and conversations render correctly, read Instrumenting LLM applications.
Associate traces with AgentControl
LLM observability captures spans for any instrumented model call. When you use LaunchDarkly AgentControl together with the LaunchDarkly Observability plugin and spans are successfully exported, LaunchDarkly associates traces with the AgentControl config that generated them.
The AI SDK annotates the root span with the underlying feature flag key for the evaluated config. LaunchDarkly uses this annotation to link related spans to the correct config when spans are exported through the Observability SDK.
With this integration, you can:
- Filter traces by config key or variation
- Correlate model behavior with variations and targeting
- Investigate latency, errors, and quality signals in context
If you do not use the LaunchDarkly Observability SDK, you can still associate traces with AgentControl by adding span attributes in your tracing pipeline. For example, attach the AgentControl config key and evaluated variation to the root span when your application calls the model provider.
View and analyze LLM spans
LaunchDarkly displays LLM observability data in two places:
- The Monitoring tab on a config, when spans relate to that config
- The global Traces page
Each view serves a different purpose.
View traces for a specific config
To view trace data for a config:
- Click Agents. The AgentControl menu appears.
- Click Configs.
- Open the Monitoring tab.
If LaunchDarkly links spans to that config, it displays them in this panel.
Why "No traces detected" appears
If the config page shows No traces detected, LaunchDarkly has not linked any spans to that config.
LaunchDarkly evaluates configs using a context. To associate spans with a config:
- The config must evaluate with a valid context.
- The LLM call must occur within the same request flow.
If your application calls a model without a LaunchDarkly context, or outside the config evaluation flow, LaunchDarkly records the span on the global Traces page but does not associate it with the config. To learn more about how LaunchDarkly makes this association, read Associate traces with AgentControl.
Show all LLM spans
To explore all captured LLM spans, open the Telemetry section and navigate to the Traces list.
The Traces page lists all captured spans, including LLM spans that may not relate to a specific config. LaunchDarkly marks LLM spans with a green indicator. Select a span to view detailed model telemetry.
Use the search bar, filters, and time range selector to analyze spans. The trace detail panel shows a timeline of generation steps and related spans, provider and model metadata such as latency and token usage, prompt and response content, and any provider errors or exceptions.
Filter on LLM span attributes
LaunchDarkly builds LLM observability and conversation views from the OpenTelemetry generative AI semantic conventions. It normalizes several common non-standard formats when it receives them, including the OpenLLMetry llm.* attributes and Claude Code telemetry. Search and display use the normalized names, so filter on the names below rather than the names your instrumentation emits.
LLM spans include the following attributes:
- gen_ai.request.model: The model that handled the request.
- gen_ai.provider.name: The provider that handled the request.
- gen_ai.usage.input_tokens: The number of input tokens processed.
- gen_ai.usage.output_tokens: The number of tokens in the output.
- gen_ai.prompt.0.content: The input prompt text. Subsequent messages use increasing indexes.
- gen_ai.completion.0.content: The generated response. Subsequent responses use increasing indexes.
- duration: The total latency of the span.
- service.name: The name of the emitting service.
Use these attributes to filter, search, and investigate model behavior. For example, to find slow requests to a specific model:
Example search query:
gen_ai.request.model=gpt-4 AND duration>2sInstrumentation that follows the current semantic conventions reports message content as gen_ai.input.messages and gen_ai.output.messages instead of the indexed attributes above. For the full set of attributes LaunchDarkly reads and how to set them, read Instrumenting LLM applications.
Read LLM conversations
The identifier LaunchDarkly groups on is the gen_ai.conversation.id span attribute. Every span that carries the same value joins the same conversation.
Conversation terminology
Understanding these terms will help you read the conversation view:
- Conversation: One complete interaction, identified by a conversation ID set on every related span. A conversation can span many traces.
- Trace: One connected unit of work within a conversation, such as a single agent invocation. LaunchDarkly displays trace boundaries in the conversation transcript, but they are a background detail rather than the main structure.
- Turn: One message or one tool call. Turns are the unit you read, navigate, and select in the conversation view.
- Mission: One question from a person, combined with everything the agent did to answer the question. Missions start at each user turn, so each one covers a single request and its follow-through.
- Evaluation: A quality score attached to a turn or to a whole conversation, such as the relevance or safety rating produced by a model-graded evaluator.
Find a conversation
Navigate to “Monitor,” click Traces, and select the Conversations tab.
The list displays one row per conversation, with the following columns:
- Conversation: The conversation ID your application set.
- Last activity: When the most recent span in the conversation started.
- Duration: Elapsed time from the first span to the end of the last span.
- Traces: How many distinct traces the conversation spans.
- Spans: Total spans carrying the conversation ID.
- Model: Every distinct model used in the conversation.
- Provider: Every distinct provider used in the conversation.
- Input tokens, Output tokens: Token counts summed across the conversation.
- Errors: How many spans reported an error status.
Sort by any column, and add or remove columns with the column picker. The list covers the last seven days.
You can also open a conversation from any span inside it. Select a span on the traces page, then use the scope control to switch from the single trace to the whole conversation.
Read the transcript
The conversation view has two tabs:
- Messages displays a continuous transcript of user messages, assistant replies, and tool calls in the order they happened, grouped into missions. Use this tab to follow what happened.
- Waterfall displays the conversation’s spans on a timeline. Use this tab to investigate timing and nesting.
Click the span icon on a turn to display raw data for the turn in a span details drawer.
Each turn displays the metadata available for it, which may include the model, token counts, latency, cost, and any evaluation scores. Tool calls display the tool name with a summary of the arguments and result, and expand to the full payload.
Turns whose timing LaunchDarkly could not determine do not display a time offset. This happens when a message is recovered from a later snapshot of the conversation history rather than observed directly. To learn more, read Emit message content.
Read the summary
The conversation header summarizes the whole thread:
Metric - How LaunchDarkly calculates it
- Duration: Elapsed time from the earliest span to the end of the latest span.
- Traces: Distinct traces. Hover to read the per-trace breakdown.
- Steps: Number of turns in the transcript, after LaunchDarkly removes duplicate observations of the same message.
- LLM: Spans identified as model calls or agent invocations.
- Tools: Tool execution spans. Hover to read the per-tool breakdown.
- Total tokens: Input, output, cache read, and cache write tokens combined. Hover to read the breakdown of tokens.
- Models: Distinct models. Hover to read the per-model call count.
- Estimated cost: Token counts priced using your configured model costs.
Why some totals differ between views
The Conversations list and the conversation header count different things.
The header adds prompt-caching tokens to its total, because cache reads often make up most of a cached workload’s context. The list reports only input and output tokens. For a conversation that uses prompt caching, expect the header total to be substantially larger than the list total.
The header’s Steps count removes duplicate observations of the same message, while LLM and Tools count spans. An agent runtime that emits several spans for one logical tool call raises the Tools count without changing Steps.
Read evaluations
If your application emits evaluation results, LaunchDarkly displays them in two places:
- On a turn, as a badge with the evaluation name and score. Scores between 0 and 1 display as fractions, and other numeric scores display on a 0 to 10 scale. Names that suggest a risk measure, such as toxicity or hallucination, use low scores to indicate positive results.
- On the conversation, as an overall score in the header for when an evaluator scored the conversation as a whole.
Attach evaluation results to the span where you want them to appear. If you attach a result to a separate evaluation span, it appears on that span rather than on the turn it describes.
Troubleshooting
When a conversation does not render as expected, the cause is usually a missing or misplaced attribute in your instrumentation. The following table maps common symptoms to their likely causes:
- The conversation does not appear in the list: gen_ai.conversation.id is missing, or set under a different key. It must be on every span in the conversation.
- The conversation appears, but the transcript is empty: No span carries gen_ai.input.messages, gen_ai.output.messages, or gen_ai.system_instructions.
- A tool renders before the message that requested it: The model call’s output has no matching tool_call part.
- Every turn appears at the start of the conversation: tool_call parts are on the envelope span rather than the per-call spans.
- The transcript opens with the agent replying to nothing: The opening question was never recorded.
- The person’s question appears as agent output: Message text is accumulated in a shared buffer rather than keyed per message.
- Tokens appear in the conversation but the list displays zero: Token usage uses the older prompt_tokens and completion_tokens names.
- Estimated cost is blank: At least one model call is missing a model, a provider, or a token count, or uses a model with no configured cost.
- Turns display no timing: Those messages were recovered from a later history snapshot rather than observed directly.
- Evaluation badges do not appear: Evaluation events are attached to a separate evaluation span rather than to the span being scored.
- The LLM Agent default dashboard reports less activity than expected: Inference spans carry HTTP or database attributes and are classified by those instead.
- A failed run appears as an empty conversation: The run failed before its first model call, so the error is recorded on a span that contributes no turns.
Privacy and data handling
LLM spans and conversations display prompt and response text, tool arguments, and tool results. This content can include personally identifiable information (PII) depending on your application.
Review your organization’s data-handling policies before you enable LLM observability or emit message content. Redact sensitive values in your instrumentation rather than after LaunchDarkly receives them.
Original source - Aug 26, 2026
- Date parsed from source:Aug 26, 2026
- First seen by Releasebot:Aug 26, 2026
v2.267.0
Flagsmith adds usage breakdowns by request type or SDK and fixes segment change request bypasses in 2.267.0.
2.267.0 (2026-08-26)
Features
- usage: break usage down by request type or SDK (#8343) (2fc223a)
Bug Fixes
- segments: Segment Change Requests bypassable via segment deletion (#8358) (56b6579)
Dependency Updates
- api: update dependency flagsmith-private to >=0.13.0,<1 (#8384) (bed0b2f)
- api: Upgrade pytest to 9.0.3 and swap in pytest-lazy-fixtures (#8375) (0091e0c)
- Aug 26, 2026
- Date parsed from source:Aug 26, 2026
- First seen by Releasebot:Aug 26, 2026
v2.266.0
Flagsmith adds cohort CSV sync, Amplitude and Mixpanel cohort sync, CSV segment creation, and cohort summaries on segments. It also improves onboarding, SCIM deactivated user support, billing usage visibility, and includes several bug fixes and refactoring work.
2.266.0 (2026-08-26)
Features
- cohort CSV sync endpoint and cohort summary on segments (#8352) (b9b32ba)
- cohorts: Amplitude cohort sync endpoints and sync keys (#8290) (cc20443)
- create segment from CSV drawer (#8283) (0102a0c)
- Mixpanel cohort sync webhook (#8338) (c25dc8c)
- onboarding: gradual rollout quest screen (#8286) (2c313e0)
- SCIM: Support deactivated user membership (#8370) (c0b2a31)
- usage: Show usage against the plan limit for the current billing period (#8320) (763938b)
- wire CSV segments to the cohorts API (#8295) (d0485cc)
Bug Fixes
- deep-link: feature value renders as [object Object] (#8341) (ac56c47)
- identify user before sending warehouse test event (#8356) (3ff27da)
- identities: label the rows in the identity override selector (#8372) (0be48ea)
- identities: multivariate override editor hides the identity's override value (#8279) (2e5ea89)
- keep experiment refresh state until requests settle (#8351) (0f81888)
- treat stale experiment refresh requests as settled (#8347) (66b630a)
Refactoring
- featureStateToValue: extract into a Flux-free module (#8344) (0212201)
- Aug 26, 2026
- Date parsed from source:Aug 26, 2026
- First seen by Releasebot:Aug 26, 2026
@vercel/[email protected]
Flags SDK adds a metricEnvironment client option to associate evaluation metrics with an environment in ingestion.
Minor Changes
#453 cc8c266 Thanks @luismeyer! - Add a metricEnvironment client option for associating evaluation metrics with an environment when sending them to the ingestion endpoint.
Original source