Datadog Release Notes

Follow

32 release notes curated from 47 sources by the Releasebot Team. Last updated: Jul 17, 2026

Get this feed:

Datadog Products

  • Jul 17, 2026
    • Date parsed from source:
      Jul 17, 2026
    • First seen by Releasebot:
      Jul 17, 2026
    Datadog logo

    Datadog

    Answer any cost question faster with the Cloud Cost skill in Bits Chat

    Datadog adds the Cloud Cost skill in Bits Chat, bringing conversational cloud cost analysis, anomaly investigation, and root cause insights that connect cost and observability data to help teams track budgets and understand AI spend faster.

    Managing cloud, AI, and SaaS costs means answering a steady stream of questions from finance, leadership, and engineering teams.

    What changed? Which team owns the spend? Was an increase expected? Are we still on track against the budget? When each answer requires moving between dashboards, filtering cost data by team or service, or manually correlating billing data with observability data, it can slow down investigations while costs continue to rise.

    The Cloud Cost skill in Bits Chat brings Cloud Cost Management analysis into a conversational workflow. You can ask cost questions in plain language, investigate cost anomalies, and get answers grounded in both cost and observability data within minutes. In this post, we’ll show how you can:

    • Ask Bits Chat anything about your costs
    • Spot Anthropic cost spikes before they grow
    • Find the root cause of a cost change

    Ask Bits Chat anything about your costs

    The Cloud Cost skill gives FinOps practitioners and engineers one place to ask cost questions without needing to know which dashboard to open or how to construct a query. You can get started with Bits Chat in the Datadog navigation bar, or you can start an investigation directly from a cost anomaly. From there, Bits Chat analyzes the relevant cost data and responds with a concise summary, then asks what you want to explore next.

    Because the Cloud Cost skill works across many cost categories, including cloud, SaaS, AI, custom, and Datadog, teams can use the same workflow for any cost question. FinOps teams can use it to track budgets, identify cost owners, and investigate anomalies. Engineers can use it to understand the cost impact of their own services and workloads, and identify optimization opportunities. For example, you can ask Bits to investigate why Anthropic costs increased last week for team:web-store, identify which teams are driving the highest OpenAI spend this month, or show total AI cost by provider for the last 30 days.

    Bits Chat can investigate cost monitor alerts, cost anomalies, and cost changes; identify teams, services, accounts, regions, or resources that are driving spend; and compare actual or forecasted spend against budgets. It can also correlate cost changes with observability metrics such as CPU, memory, request volume, and storage size. This gives teams the technical context they need to understand whether a cost change came from a usage increase, a configuration change, or another source.

    Spot Anthropic cost spikes before they grow

    AI usage can become expensive quickly, especially when costs are spread across providers, models, teams, users, and services. AI Costs in Cloud Cost Management gives FinOps and engineering teams a unified location for analyzing AI spend across providers such as OpenAI, Anthropic, Amazon Bedrock, Google Gemini, Vertex AI, GitHub Copilot, and Cursor. Datadog normalizes this spend so teams can understand which providers, models, users, and API keys are contributing to cost changes.

    This level of granularity helps teams detect AI cost anomalies before they cause larger budget issues. For example, Datadog can surface an unexpected Anthropic cost spike that might otherwise go unnoticed until the next billing review. But identifying a spike is only half the problem—understanding what caused it is another. Without a connected cost investigation workflow, a team would need to pull billing exports, cross-reference logs, and search for the service or team that caused the change.

    With the Cloud Cost skill, the team can start the investigation in one click. When the team clicks “Investigate” directly from the anomaly, Bits automatically pulls in the relevant context and kicks off a root cause analysis. This lets the team move from detection to investigation without leaving the AI Costs page or having to do any manual work.

    Find the root cause of a cost change

    Once an investigation starts, Bits Chat performs root cause analysis by using both cost data and observability data from across Datadog. For a cost anomaly investigation, the initial analysis typically includes a daily cost chart for the baseline and investigation periods so you can see exactly when a change started. It also summarizes the total dollar change, percentage change, and projected annual impact when applicable.

    Bits Chat also provides rate-versus-usage context, which helps teams determine whether spend increased because of a pricing change or because a team consumed more of a resource. For AI spend, this distinction is especially important. An Anthropic spike might come from more requests, larger prompts, higher output token usage, a model change, or a combination of these factors. By correlating cost with observability metrics, Bits can help connect the billing change to the system behavior behind it.

    In the preceding example screenshot, Bits identifies that the Anthropic cost spike is driven by two teams: Customer Support (57%) and AI Platform (37%). Together, these teams account for the full $2,895.49 increase—a 188.15% jump—with a projected annual impact of $62,167.87. Instead of knowing only that Anthropic costs went up, the engineer can see exactly which teams are driving it and by how much, which makes it immediately clear who to loop in.

    From there, the team can keep drilling into the investigation. Bits can find the top services, accounts, regions, resources, or tags driving the change. It can also compare actual and forecasted spend against related budgets, identify optimization opportunities, and create a Datadog Notebook that captures the investigation for the owning team. This helps preserve the full context, so the team does not need to reconstruct the analysis later from Slack threads or separate reports.

    Get started with the Cloud Cost skill in Bits Chat

    The Cloud Cost skill in Bits Chat helps FinOps and engineering teams investigate cost changes, budget risks, and AI spend in the same place where they already monitor their systems. By grounding answers in both cost and observability data, Bits Chat can help teams move from a broad cost question to an actionable explanation in minutes.

    To get started, make sure you’ve configured Cloud Cost Management for the cost sources you want to analyze. Users also need Bits Chat access and Cloud Cost Management permissions for the data they ask about. If you want Bits to create or update investigation records, you can also grant the relevant Notebooks permissions. Learn more in the Cloud Cost skill documentation, Bits Chat documentation, and AI Costs in Cloud Cost Management documentation.

    If you don’t already have a Datadog account, you can sign up for a 14-day free trial to start investigating your cloud costs with Datadog.

    Original source
  • Jul 16, 2026
    • Date parsed from source:
      Jul 16, 2026
    • First seen by Releasebot:
      Jul 16, 2026
    Datadog logo

    Datadog

    Datadog Expands UK Data Hosting Capabilities on AWS Europe (London) Region

    Datadog launches its products and services on the AWS Europe (London) Region, giving customers more options to keep observability and security data in the UK with lower latency, stronger governance, and improved support for compliance and operational resilience.

    Local data storage capacity gives Datadog customers greater flexibility in how they meet UK data residency, governance, compliance and security requirements

    London, UK — 16 July 2026 — Datadog, Inc. (NASDAQ: DDOG), the leading AI-powered observability and security platform for cloud applications, announced it has launched its products and services on the Amazon Web Services (AWS) Europe (London) Region.

    This gives Datadog’s customers and partners additional options to store observability and security data within the UK, with lower latency and streamlined incident response for greater operational resilience. It also supports customers as they navigate evolving data governance and compliance considerations, while maintaining unified visibility across cloud, hybrid, and AI environments.

    The launch is particularly relevant for organisations in regulated sectors, including financial services, healthcare, government and higher education, where in-region data residency and operational continuity are key considerations. By keeping observability and security data closer to the workloads being monitored, Datadog helps organisations investigate incidents faster, improving system reliability, and strengthening security operations.

    “UK enterprises scaling cloud and AI workloads increasingly want the option to keep observability and security data in-region to support their data governance and operational resilience,” said Yanbing Li, Chief Product Officer at Datadog. “AI is accelerating the complexity of modern systems, creating more data, dependencies, and potential points of failure. This milestone helps customers keep observability and security data close to their workloads while maintaining the resilience and governance needed to operate at scale. With end-to-end visibility across infrastructure, applications, security, and AI, Datadog gives customers the complete picture they need to maintain operational control.”

    Steve Barrett, VP EMEA at Datadog added, “Cloud adoption is now the norm for UK organisations, and AI is adding a new layer of operational complexity. As systems become more distributed, observability becomes increasingly important to maintaining resilience, security and control. Launching on the AWS Europe (London) Region gives customers the option to keep their observability and security data in-region, helping them strengthen governance while continuing to scale with confidence.”

    For more information, visit:

    https://docs.datadoghq.com/getting_started/site/

    About Datadog

    Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high-growth AI leaders, Datadog enables businesses to move faster with clarity and confidence.

    Forward-Looking Statements

    This press release may include certain “forward-looking statements” within the meaning of Section 27A of the Securities Act of 1933, as amended, or the Securities Act, and Section 21E of the Securities Exchange Act of 1934, as amended including statements on the benefits of new products and features. These forward-looking statements reflect our current views about our plans, intentions, expectations, strategies and prospects, which are based on the information currently available to us and on assumptions we have made. Actual results may differ materially from those

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Datadog and hundreds of other software products.

    Create account
  • Jul 15, 2026
    • Date parsed from source:
      Jul 15, 2026
    • First seen by Releasebot:
      Jul 15, 2026
    Datadog logo

    Datadog Agent by Datadog

    7.81.1

    Datadog Agent adds reflector-based Kubernetes event collection and upgrades builds to Go 1.26.5.

    Agent

    Prelude

    Released on: 2026-07-15

    • Please refer to the 7.81.1 tag on integrations-core for the list of changes on the Core Checks

    New Features

    • Add a new reflector-based Kubernetes event collection path, enabled via event_collection_mode: watch.

    Enhancement Notes

    • Agents are now built with Go 1.26.5.

    Datadog Cluster Agent

    Prelude

    Released on: 2026-07-15 Pinned to datadog-agent v7.81.1: CHANGELOG.

    Original source
  • Jul 8, 2026
    • Date parsed from source:
      Jul 8, 2026
    • First seen by Releasebot:
      Jul 10, 2026
    • Modified by Releasebot:
      Jul 18, 2026
    Datadog logo

    Datadog Agent by Datadog

    7.81.0

    Datadog Agent ships broader observability and scaling updates, including default Datadog v3 metrics forwarding, new GPU monitoring metrics, APM remote configuration improvements, and in-place vertical scaling as the default autoscaling strategy, plus Linux process autoconfig and log pipeline enhancements.

    Agent

    Prelude

    Released on: 2026-07-08

    • Please refer to the 7.81.0 tag on integrations-core for the list of changes on the Core Checks

    Metrics are now forwarded to the new Datadog v3 API by default (/api/intake/metrics/v3/series). The v3 API payload format is more compact, reducing outbound bandwidth from Agents to Datadog.

    If you configure additional_endpoints to forward to a non-Datadog endpoint, you will likely need to disable v3 for this endpoint. Otherwise you will see 404s. This can be done via:

    use_v3_api:
      series:
        endpoints: "<additional_endpoint>": false
    

    Example:

    additional_endpoints:
      "https://non-datadog-endpoint": apikey2
    

    will need:

    use_v3_api:
      series:
        endpoints:
          "https://non-datadog-endpoint": false
    

    Metrics sent to Observability Pipelines Worker continue to use the v2 API by default.

    To keep using v2 endpoint, set use_v3_api.series.enabled: "false" (global) or use_v3_api.series.endpoints: { "": "false" } (per-endpoint; shown above).

    Upgrade Notes

    • The DDOT feature gate exporter.datadogexporter.metricremappingdisabled has been removed and replaced with exporter.datadogexporter.DisableAllMetricRemapping.
    • Removed the agent status py subcommand (which wasn't officially supported)
    • On Linux, the agent process manager systemd units were renamed from datadog-agent-procmgrd.service / datadog-agent-procmgrd-exp.service to datadog-agent-procmgr.service / datadog-agent-procmgr-exp.service. The dd-procmgrd binary and its paths are unchanged.

    On upgrade, the installer stops and removes the legacy procmgrd-suffixed unit files so only one process manager daemon binds the socket. Update any custom automation that referenced the old unit names.

    • Upgrade OpenTelemetry Collector dependencies from v0.152.0 to v0.153.0 (core v1.58.0 to v1.59.0).

    See the full upstream changelogs: collector-contrib v0.153.0, collector core v0.153.0.

    • Upgrade OpenTelemetry Collector dependencies from v0.153.0 to v0.154.0 (core v1.59.0 to v1.60.0).

    See the full upstream changelogs: collector-contrib v0.154.0, collector core v0.154.0.

    New Features

    • In-place vertical scaling is enabled as the default strategy for workload autoscaling.
    • New metrics for GPU memory have been added to the GPU Monitoring product:
      • gpu.memory.utilization: Ratio of used memory compared to total memory.
    • Add passthrough entry for genresources EVP intake track.
    • This change adds two new metric points for the GPU Monitoring product:
      • gpu.pci.link.speed.current: Current usable bandwidth for the PCI link in bytes per second
      • gpu.pci.link.speed.max: Max usable bandwidth for the PCI link in bytes per second
    • Add a new ReportIssue method to the Python bridge to report issues to Agent Health Platform
    • APM: The trace-agent can now receive span tag equivalence and peer tag mapping updates over Remote Configuration and apply them at runtime, without an agent restart. The feature is opt-in via the new remote_configuration.apm_semantics.enabled setting (default false). Stats aggregation picks up the updated peer-tag keys on the next span processed. If the backend removes or untargets a previously-applied payload, the trace-agent reverts to the mappings it ships with. Existing deployments see no behavior change with default settings.
    • APM: remote_configuration.agent_config.enabled is now a settable configuration entry that controls the trace-agent's Remote Configuration subscription for agent-config updates (such as runtime log-level overrides) independently from remote_configuration.apm_sampling.enabled. When the user has explicitly set apm_sampling.enabled but not agent_config.enabled, the trace-agent mirrors the former into the latter so existing configurations continue to behave exactly as before.
    • On Linux, when the DDOT extension is installed with the Datadog Agent, DDOT is now managed by dd-procmgrd through processes.d/datadog-agent-ddot.yaml instead of relying on the legacy datadog-agent-ddot systemd unit. Uninstalling the extension removes that config file. To roll back to the legacy behavior manually, remove processes.d/datadog-agent-ddot.yaml and restart datadog-agent.
    • Add Go stack trace aggregation to the auto multi-line log pipeline. When auto multi-line aggregation is enabled (logs_config.auto_multi_line_detection), multi-line Go crash dumps (panic:, fatal error:, runtime: errors, signal crashes, and unexpected faults) are automatically detected and combined into a single log entry using a streaming state-machine parser.
    • gpu: all gpu.nvlink.* metrics now have a nvlink_port tag and are emitted per-port. We provide GPU-level alternatives for certain metrics such as gpu.nvlink.throughput.data.rx/tx.total
    • Enable instrumentation_crd_controller.enabled and a new autodiscovery provider will schedule checks derived from DatadogInstrumentation custom resources deployed in the Kubernetes cluster.
    • Parses and collects kubernetes.pod.cpu.requests, kubernetes.pod.memory.requests, kubernetes.pod.cpu.limits, and kubernetes.pod.memory.limits.
    • Process Autodiscovery is now enabled by default on Linux through the process autoconfig feature. It can be disabled with DD_AUTOCONFIG_EXCLUDE_FEATURES=process.
    • Register process_manager.enabled in the Agent configuration schema (pkg/config/schema/core_schema.yaml), set its default in pkg/config/setup, and document it in config_template.yaml. On Windows, this option controls whether the core Agent starts dd-procmgr-service. On Linux, dd-procmgrd is started by systemd; this setting is ignored there.

    Enhancement Notes

    • Use compensated floating point summation to accurately calculate the sum and average aggregates of histograms for inputs where magnitudes significantly vary.
    • Scale .sum, .avg, and .count aggregates by the exact 1/SampleRate to avoid undercount of these aggregates for sample rates whose reciprocal is not an integer (e.g. @0.21).
    • The macOS battery check now adds a power_state:battery_critical tag to the system.battery.power_state metric when the operating system reports a degraded battery.
    • Update the SNMP traps database with new MIB additions, including PANZURA-TRAP-MIB.
    • Updated the ntp check to support the default location of systemd-timesyncd (/etc/systemd/timesyncd.conf). The check now parses NTP= and FallbackNTP= keys in addition to the existing chrony/ntp.conf server / pool / peer directives.
    • On startup the Datadog Agent will now validate its configuration against the schema and report any violations through the Agent Health pipeline.
    • APM stats now mask additional metric tag values that exceed the value length or per-bucket cardinality limits.
    • APM : The enable_otlp_container_tags_v2 behavior is now enabled by default. Container tags on OTLP traces are now extracted using the infraattributes processor instead of calling the tagger directly, reducing redundant work and outgoing traffic. To opt out, set disable_otlp_container_tags_v2 in apm_config.features.
    • Agents are now built with Go 1.26.4.
    • CWS: Add support for monitoring the socket system call, enabling detection rules based on socket creation events (domain, type, protocol).
    • The comp/dataobs/queryactions component now supports an optional schedule field on Data Observability monitor queries. The field accepts a standard 5-field cron expression (e.g. "20 * * * *" for 20 minutes past every hour) and enables wall-clock-aligned scheduling in place of the fixed interval_seconds cadence. When both schedule and interval_seconds are set on the same query, schedule takes precedence and interval_seconds is ignored. At least one of the two fields must be set; the agent rejects Remote Configuration payloads containing queries where neither field is provided or where the cron expression is syntactically invalid.
    • Expanded the functionality of the experimental fentry-based network connection tracer. This tracer remains experimental and disabled by default.
    • gpu: add new PCI link width metrics gpu.pci.link.width.{current,max} and add degraded PCI link metrics gpu.pci.link.{width,speed}.degraded.
    • gpu: add gpu.nvlink.errors.fec.{none,light,heavy} metrics to easily group error thresholds
    • The in-place vertical autoscaler throttles disruptive resizes to at most 15% of a workload's replicas per reconcile, configurable via autoscaling.workload.in_place_vertical_scaling.disruption_tolerance_percent.
    • use the /healthz route to check and validate kubelet connection, instead of the deprecated /spec route.
    • The agent automatically detects Kueue-related labels in pods and adds them as kueue_local_queue and kueue_cluster_queue tags.
    • Network Config Management: Adds support for Cisco ASA firewalls by adding a new profile for these. Previously, Cisco ASA was not supported and would be unmonitored by the NCM integration.
    • Extended the ntp check's systemd-timesyncd discovery to also read drop-in files under /etc/systemd/timesyncd.conf.d/, /run/systemd/timesyncd.conf.d/, /usr/local/lib/systemd/timesyncd.conf.d/, /usr/lib/systemd/timesyncd.conf.d/. This covers hosts where NTP= is set by cloud-init or another tool that writes a drop-in instead ...

    Read more

    Assets 2

    Original source
  • Jul 8, 2026
    • Date parsed from source:
      Jul 8, 2026
    • First seen by Releasebot:
      Jul 8, 2026
    Datadog logo

    Datadog

    Monitor your .NET MAUI apps with Datadog RUM

    Datadog adds an official .NET MAUI SDK for Real User Monitoring, giving teams a supported way to monitor crashes, errors, network activity, and Session Replay from a single NuGet package while viewing MAUI app data alongside other mobile telemetry in Datadog.

    Detect and troubleshoot crashes in .NET MAUI apps

    As .NET Multi-platform App UI (MAUI) becomes the default cross-platform UI framework in the Microsoft ecosystem, many teams are standardizing on it to build mobile applications for iOS and Android. However, observability has not kept pace with the shift in adoption. Developers often rely on unsupported community bindings or maintain their own wrappers around native iOS and Android SDKs, which introduces instability and ongoing maintenance. The retirement of Microsoft Visual Studio App Center has also left many teams without a clear path for monitoring crashes, errors, and user activity in production.

    Datadog Real User Monitoring (RUM) provides an official .NET MAUI SDK that enables teams to instrument applications by using a single supported NuGet package. The SDK supports many core RUM features, including built-in crash reporting, error tracking, and network monitoring, along with Session Replay. All telemetry data flows into existing Datadog views, so you can analyze .NET MAUI applications alongside other mobile apps built with native iOS and Android or other cross-platform frameworks such as React Native, Flutter, and Kotlin Multiplatform.

    In this post, we’ll explore how you can use the SDK to:

    • Detect and troubleshoot crashes in .NET MAUI apps
    • Understand user sessions and performance across .NET MAUI apps

    Mobile crashes in production often require significant effort to reproduce and diagnose, especially when stack traces are incomplete or difficult to interpret. Community-maintained bindings can compound the problem by introducing gaps in instrumentation or becoming incompatible with new SDK releases.

    The Datadog .NET MAUI SDK automatically captures crashes and errors across managed and native layers of an application. Coverage includes .NET exceptions and platform-specific issues such as iOS App Hangs and Android Application Not Responding (ANR) events. Automatic instrumentation enables teams to begin collecting telemetry data without adding custom logic or maintaining separate integrations.

    Datadog deobfuscates stack traces through its symbol upload process, which uses PDB files to restore readable method names and source context. Engineers can investigate issues with clear, actionable information instead of working with obfuscated output.

    For example, consider a scenario where a team releases a new version of a .NET MAUI app and begins receiving crash reports shortly afterward. An engineer can navigate to Datadog Error Tracking, identify a spike in crashes tied to the new release, and review the associated stack traces. Correlating the crash with recent code changes helps the team isolate the faulty method and begin remediation.

    Crash and error data from .NET MAUI apps appears alongside data from other supported mobile apps in the same Datadog views, including the RUM Explorer, Error Tracking, and session details pages. The shared interface enables teams to apply existing workflows without introducing additional tools.

    Understand user sessions and performance across .NET MAUI apps

    Understanding how users interact with an application and how services respond plays a key role in maintaining performance and reliability. Limited instrumentation often leaves teams without visibility into network latency, failed requests, and the sequence of user actions that led to issues.

    The .NET MAUI SDK automatically tracks network requests made through common libraries, such as HttpClient. Automatic collection provides out-of-the-box visibility into request latency and error rates, in addition to automated correlation with downstream backend dependencies through distributed APM traces.

    View and action tracking give developers control over how user activity is captured and organized in session data. Developers can use views to track the screens that users navigate to. With actions, developers can capture user interactions that align with specific workflows.

    Session, network, and interaction data all appear in the RUM Explorer and session detail pages. Teams can use the unified dataset to analyze performance and identify patterns across iOS, Android, and other supported platforms.

    Get started with .NET MAUI monitoring in Datadog

    The Datadog .NET MAUI SDK provides a supported path to instrument cross-platform mobile apps without relying on community-maintained bindings or custom wrappers. Automatic crash reporting, network monitoring, and manual instrumentation capabilities give teams visibility into application health and user behavior. To learn more, read our .NET MAUI SDK documentation.

    If you’re new to Datadog, you can sign up for a 14-day free trial to start monitoring your .NET MAUI apps.

    Original source
  • Similar to Datadog with recent updates:

  • Jul 8, 2026
    • Date parsed from source:
      Jul 8, 2026
    • First seen by Releasebot:
      Jul 8, 2026
    Datadog logo

    Datadog

    Protect AWS Strands Agents with Datadog AI Guard

    Datadog adds AI Guard support for AWS Strands Agents, bringing inline protection for prompts, responses, tool calls, and tool results. The plugin helps teams monitor, block, and investigate unsafe agent behavior with Datadog traces and policy controls.

    AI agents can reason through tasks, call tools, and adapt their next steps based on intermediate results. That flexibility is useful for building agentic applications, but it also creates security risk at runtime: A prompt injection attempt can change the agent’s instructions, a malicious request can try to exfiltrate sensitive data, and an unsafe tool call can lead to an action that the application owner did not intend.

    Datadog AI Guard now works with AWS Strands Agents through a Strands plugin that evaluates prompts, model responses, and tool interactions as the agent runs. By using the Strands native hook system, AI Guard can monitor or block unsafe behavior in the agent loop without requiring teams to scatter security checks throughout application code. In this post, we’ll show how to:

    • Monitor the Strands agent loop
    • Evaluate prompts, responses, and tool calls inline
    • Configure enforcement without changing agent code
    • Investigate AI Guard evaluations in Datadog

    Monitor the Strands agent loop

    Strands Agents takes a model-driven approach to orchestration. The model reasons through the task, chooses tools, builds context from previous steps, and decides when it has enough information to respond. This design helps teams build agents that can handle open-ended workflows, but it also means the application’s behavior can change with each user request and model decision.

    Traditional application security controls are not always designed for this type of runtime behavior. An agent’s risk can depend on the sequence of prompts, intermediate tool results, and tool calls that led to a particular action. The attack surface is a dynamic sequence of model decisions that changes with every interaction. If teams add custom checks directly into each part of that workflow, those checks can make the agent harder to audit and maintain. Teams also need to redeploy the application whenever they change those checks.

    The AI Guard plugin for Strands Agents gives teams a central place to evaluate agent behavior as the agent runs. The plugin assesses every interaction in the context of the full agent session to catch multistep attacks that become harmful only after several tool calls. It registers callbacks on Strands life cycle events, sends relevant content to AI Guard for evaluation, and applies the configured response before the agent loop continues. This makes AI Guard part of the agent’s execution path rather than a separate review layer after the fact.

    Evaluate prompts, responses, and tool calls inline

    AI Guard evaluates the parts of an agent session where risk most often appears: user input, assistant output, tool invocations, and tool results. This inline approach helps teams inspect agent behavior in context, including the relationship between a tool call and the preceding steps that produced it.

    The AIGuardStrandsPlugin registers callbacks for four Strands hook events:

    HOOK EVENT WHAT AI GUARD EVALUATES RESPONSE WHEN BLOCKED BeforeModelCallEvent User prompts, excluding tool results Raises AIGuardAbortError AfterModelCallEvent Assistant text content Raises AIGuardAbortError BeforeToolCallEvent Pending tool call and conversation context Cancels the tool with a descriptive message AfterToolCallEvent Tool result and conversation context Replaces the tool result content

    These checks help protect against common risks in production agent workflows. Prompt protection evaluates user prompts and model responses for attacks such as prompt injection and jailbreaking. Tool protection analyzes tool calls, arguments, intent, and surrounding context to help determine whether an invocation should continue. Sensitive data protection detects personally identifiable information (PII), secrets, and other sensitive content in LLM inputs and outputs.

    The AI Guard plugin also avoids duplicate evaluations. Tool results that AI Guard evaluates during AfterToolCallEvent are excluded from the next BeforeModelCallEvent scan, which prevents the same content from being evaluated twice. If the AI Guard API is unreachable because of a network error, the plugin logs the failure at the debug level and allows the agent to continue.

    Configure enforcement without changing agent code

    AI Guard supports a monitor mode that lets teams observe evaluations before they start blocking traffic. This is useful when you are first adding AI Guard to an agent, tuning policy behavior, or evaluating how detections map to real production traffic. After your team has reviewed the results, you can switch a service to blocking mode from the AI Guard settings.

    You can also adjust detection sensitivity to make policies stricter or more lenient for a given service. For tool-specific controls, AI Guard supports a tool denylist that proactively blocks selected tools from being used by the agent. This gives security and platform teams a way to reduce risk for sensitive actions, such as file operations, administrative APIs, or payment-related tools.

    Because these controls are managed in Datadog, teams can update policies without editing the agent or redeploying the application. This is especially helpful when multiple teams own different agents across environments. Security teams can start in monitor mode, review how agents behave in practice, and then tighten controls as policies mature.

    Investigate AI Guard evaluations in Datadog

    After you add AI Guard to a Strands agent, evaluations appear in Datadog as spans that show whether an interaction was safe or unsafe. For unsafe interactions, Datadog includes the attack category, such as prompt injection, data exfiltration, tool misuse, or jailbreaking. These spans link to Datadog APM and Agent Observability traces, so teams can move from a flagged evaluation to the broader agent session that produced it.

    This trace context is important because agent attacks often unfold over several steps. A single prompt, tool call, or response might look harmless in isolation, but the full session can reveal how an instruction changed the agent’s behavior or how a tool result influenced the next model call. By reviewing the span side panel, teams can inspect previous prompts, tool calls, and related context as part of the same investigation. Every evaluation decision is a traceable event linked to the full LLM trace, giving teams the audit trail they need to demonstrate that their agent behaved within policy.

    AI Guard also provides an aggregate view of evaluations in the Signals tab. Instead of requiring teams to review raw evaluation logs one by one, Datadog groups and ranks events that warrant investigation. This helps security and platform teams prioritize high-signal activity across agents and environments.

    Protect agentic workflows in production

    AI Guard for AWS Strands Agents helps teams evaluate agent behavior at the points where runtime risk appears: prompts, responses, tool calls, and tool results. By using the Strands hook system, the plugin brings AI Guard into the agent loop while keeping policy configuration separate from application logic. The result is a centralized way to observe, govern, and block unsafe agent behavior while preserving the trace context that teams need for investigation.

    To get started, read the Datadog AI Guard documentation and the AWS AI Guard Strands plugin documentation. If you’re interested in trying AI Guard, sign up for the AI Guard Limited Availability Program.

    If you’re new to Datadog, sign up for a 14-day free trial.

    Original source
  • Jul 6, 2026
    • Date parsed from source:
      Jul 6, 2026
    • First seen by Releasebot:
      Jul 7, 2026
    Datadog logo

    Datadog

    Reduce SAST false positives with agentic evaluation and Bits Memories

    Datadog adds agentic evaluation and Bits Memories to Static Code Analysis, helping Bits AI triage SAST findings with repository-wide evidence and team-specific security knowledge to reduce false positives and improve explanations.

    Automatically triage SAST findings with full repository context by using agentic evaluation

    Static application security testing (SAST) tools are intentionally conservative. Traditional scanners identify code that appears exploitable and flag the snippet for review, even when protections elsewhere in the application prevent exploitation. Although that approach helps teams catch vulnerabilities, it also creates false positives that consume developer time, slow remediation efforts, and make future alerts easier to dismiss.

    As part of Datadog Static Code Analysis in Datadog Code Security, Bits AI already helps teams prioritize findings by assessing whether a finding is likely to be a true positive or a false positive and providing a short explanation. Many findings, however, can’t be evaluated from the flagged file alone. A function might appear vulnerable until you discover that every caller validates its inputs or that an authorization check runs elsewhere in the request path. Other findings depend on team-specific conventions that aren’t visible in the code at all.

    To help teams distinguish real vulnerabilities from false positives when evidence exists outside the flagged file, Datadog Static Code Analysis includes agentic evaluation and Bits Memories. Agentic evaluation brings repository-wide analysis to findings, and Bits Memories incorporates your organization’s knowledge into false positive assessments. In this post, we’ll explore how these capabilities help you:

    • Automatically triage findings with full repository context
    • Capture your team’s security conventions
    • Combine repository-wide evidence with organizational knowledge

    For example, consider a server-side request forgery (SSRF) finding generated during a scan. The following code shows the function that triggered the finding:

    func FetchURL(ctx context.Context, target string) (*http.Response, error) {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, target, nil)
        if err != nil {
            return nil, err
        }
        return http.DefaultClient.Do(req)
    }
    

    Viewed in isolation, the function appears risky because it accepts a user-supplied URL and passes it directly to an HTTP client. However, the flagged function doesn’t tell the whole story. To understand whether the finding represents a genuine vulnerability, you need to know how the function is used in the application. The following code shows one of the function’s callers:

    func HandlePreview(w http.ResponseWriter, r *http.Request) {
        target := r.URL.Query().Get("url")
        if !allowlist.Valid(target) {
            http.Error(w, "blocked", http.StatusBadRequest)
            return
        }
        FetchURL(r.Context(), target)
    }
    

    The caller in this example validates the URL against an allowlist before invoking the function. Bits follows those relationships automatically with agentic evaluation, inspecting callers and related code paths before assessing the finding. If every reachable caller validates the input, the evidence points toward a false positive. If a caller passes user input directly to the function, the evidence indicates a genuine vulnerability.

    Repository context also improves the explanations that accompany findings. Instead of providing only generic rationale, Bits references the specific validators, wrappers, or callers that informed its assessment. Security teams can review the reasoning and understand why a finding was classified as a likely true positive or false positive.

    Agentic evaluation operates within controlled boundaries. Bits receives read-only access to repository contents and can inspect code, search files, and gather evidence. It cannot modify files, execute code, or access resources outside the repository. Bits treats what it reads as untrusted evidence, not instructions. All evaluations also pass through Datadog AI Guard, which provides defense against prompt injection attempts.

    Capture your team’s security conventions with Bits Memories

    Exploring a repository helps when the missing context lives in code, but some triage decisions depend on organizational knowledge that isn’t visible in the repository. One team might treat an outbound URL as safe after it passes through a central allowlist, while another team might require validation closer to the call site. General-purpose models don’t know those conventions, which makes it difficult to assess findings accurately.

    Bits Memories brings organizational context for more accurate assessments. Memories are scoped per rule in your organization and apply across all of your repositories. A memory for one rule has no effect on how Bits evaluates other rules. Within that scope, Bits draws on two sources of information. The first source is your organization’s history of false positive reports for that specific rule. When you repeatedly identify a recurring pattern as a false positive, Bits uses that history as additional context during future evaluations.

    The second source of information for Bits Memories is custom context that your team provides in free-text notes. You can document framework-specific behavior, internal validation libraries, authorization patterns, review guidance, and other information that isn’t visible in the repository. A custom note might explain that all requests pass through a shared authorization layer, or that a proprietary validation function implements a company-wide security standard.

    The memories that you create through false positive reports and custom context serve as evidence, not a suppression list. A past report or custom note does not silence every future finding for that rule. Bits evaluates each finding and determines whether the prior context applies. If the pattern matches previous assessments, the memory contributes additional evidence. If the pattern differs from past examples, Bits relies on current evidence instead of the memory.

    Scans also reflect the latest organizational knowledge. When teams submit false positive reports or update custom context, subsequent evaluations incorporate that new data rather than relying on outdated verdicts or information.

    Combine repository-wide evidence with your team’s knowledge

    Agentic evaluation and Bits Memories are designed to complement each other. When used in tandem, they give Bits a more complete picture of a finding to help it reach an accurate verdict.

    For example, consider the SSRF finding that we highlighted previously. Agentic evaluation can determine that every caller validates the URL before invoking the flagged function. Bits Memories can add historical context showing that your security team has repeatedly classified findings as false positives when that same validation pattern is present. The combination of evidence provides a more conclusive assessment than either source alone.

    Human reviewers don’t assess findings in a vacuum. They trace data paths, consult framework conventions, and draw on institutional memory. By combining repository-wide evidence with organizational knowledge, Bits can bring that same depth of reasoning to findings.

    Get started with agentic evaluation and Bits Memories

    Agentic evaluation and Bits Memories make SAST triage more accurate by extending Bits’s view beyond a single file to the full repository and your team’s accumulated knowledge. Together, these capabilities help reduce false positives, provide more trustworthy explanations, and direct developer attention toward findings that deserve investigation. To learn more, read the documentation about AI enhancements in Static Code Analysis and the blog post about our open source AI-native SAST.

    If you’re new to Datadog, you can sign up for a 14-day free trial to get started with Datadog Code Security and Static Code Analysis.

    Original source
  • Jul 6, 2026
    • Date parsed from source:
      Jul 6, 2026
    • First seen by Releasebot:
      Jul 6, 2026
    Datadog logo

    Datadog

    Monitor watchOS and visionOS apps with Datadog RUM

    Datadog adds RUM support for watchOS and visionOS, bringing crash reporting, error tracking, and session-level observability to Apple Watch and Apple Vision Pro apps. It works in all regions, including GovCloud, with fully deobfuscated stack traces and no separate SDK.

    Apple’s platform ecosystem is evolving as developers build production applications for watchOS and visionOS. Whether it’s a fitness app on Apple Watch or an immersive spatial computing experience on Apple Vision Pro, these platforms have moved beyond the experimental phase to support real users. Despite this growth in adoption, teams lack visibility into how their apps behave on these devices. Unlike iOS, where mature observability tooling exists, watchOS and visionOS have remained largely unobserved: Crashes go undiagnosed, errors surface without context, and sessions pass without insight into user experience.

    Datadog Real User Monitoring (RUM) fills the visibility gap by supporting watchOS and visionOS in all regions, including GovCloud. Datadog has extended the existing dd-sdk-ios package to compile, run, and be fully tested on both platforms, so no separate SDK is required.

    In this post, we’ll explore how RUM brings Apple Watch and Apple Vision Pro apps the same crash reporting, error tracking, and session-level observability that iOS teams rely on.

    Monitor crashes and errors in watchOS and visionOS apps with deobfuscated stack traces

    When issues occur on watchOS and visionOS apps, teams often have limited information to work with. Determining the source of a crash, understanding the impact of a runtime error, and reconstructing the user journey that led to a failure can require piecing together information from multiple sources. Datadog RUM brings those signals together in a single view through crash reporting and error tracking with fully deobfuscated stack traces, as well as session tracking:

    • Crash reporting automatically captures crashes on both platforms. This capability is especially valuable on watchOS, where constrained hardware resources and a different application life cycle can make failures more difficult to reproduce and investigate. Rather than working from a string of memory addresses, you get a human-readable trace that points to the exact line of code that failed.
    • Error tracking identifies runtime errors across WatchKit and SwiftUI watchOS apps, as well as immersive and windowed visionOS experiences. By correlating these errors with application context, you can triage and resolve issues more quickly.
    • Session tracking shows how users interact with Apple Watch and Apple Vision Pro apps. Session-level data helps you understand real usage patterns, identify friction points, and measure the impact of releases.

    Deobfuscation is powered by Datadog’s symbolication platform, which automatically collects watchOS and visionOS system symbols. The process to upload your application’s dSYM files is the same as for iOS, so crash reports and runtime errors include readable stack traces with no additional configuration beyond what iOS teams already do.

    Start monitoring your watchOS and visionOS apps with Datadog RUM

    Datadog RUM for watchOS and visionOS gives you the same level of visibility into your Apple Watch and Apple Vision Pro apps that you get for iOS apps. You can monitor crashes, track errors, and analyze user sessions by using the same workflows that you already have in place. Stack traces are fully deobfuscated, and you don’t need to adopt or maintain a separate SDK. For setup instructions, read the RUM documentation for Apple platform monitoring.

    If you don’t already have a Datadog account, you can sign up for a 14-day free trial to get started monitoring your watchOS and visionOS apps.

    Original source
  • Jul 1, 2026
    • Date parsed from source:
      Jul 1, 2026
    • First seen by Releasebot:
      Jul 3, 2026
    • Modified by Releasebot:
      Jul 18, 2026
    Datadog logo

    Datadog Agent by Datadog

    7.80.4

    Datadog Agent releases Prelude with more Linux SSI installation traces and a Datadog Cluster Agent pinned to v7.80.4.

    Agent

    Prelude

    Released on: 2026-07-01

    • Please refer to the 7.80.4 tag on integrations-core for the list of changes on the Core Checks

    Bug Fixes

    • Add more traces during SSI installation on Linux host

    Datadog Cluster Agent

    Prelude

    Released on: 2026-07-01 Pinned to datadog-agent v7.80.4: CHANGELOG.

    Assets 2

    Original source
  • Jul 1, 2026
    • Date parsed from source:
      Jul 1, 2026
    • First seen by Releasebot:
      Jul 3, 2026
    Datadog logo

    Datadog

    Accelerate investigations with AI in Datadog Incident Response

    Datadog introduces AI-powered Incident Response capabilities that help teams investigate root causes, catch up on active incidents with chat summaries, and automatically capture bridge call discussions in a unified incident timeline for faster, more coordinated response.

    Diagnose root causes with Bits Investigation as an active AI responder

    Engineering teams spend much of their incident response time investigating the problem and coordinating the response. Both tasks become harder when telemetry data lives in one place, deployment history is stored in another, and conversations unfold across chat channels and incident bridges. Responders often spend the first part of an incident rebuilding context before they can begin testing hypotheses and working toward resolution.

    Datadog Incident Response includes three new AI-powered capabilities that help teams investigate incidents and coordinate responses more effectively without leaving their existing workflows. In this post, we’ll explore how you can:

    • Diagnose root causes with Bits Investigation as an active AI responder
    • Catch up on active incidents with AI-generated chat summaries
    • Capture bridge discussions in a unified incident timeline

    AI can help responders investigate incidents, but the quality of an investigation depends on the context available to the AI. When telemetry data, deployment history, ownership information, and incident activity are spread across different systems, AI tools often have to fill gaps with assumptions. Those assumptions can send responders down the wrong path and prolong an incident.

    Bits Investigation joins active incidents as a responder alongside the human response team and analyzes the same incident context that the team uses. For example, Bits can reason over a latency graph shared in a Slack thread, a summary generated from an incident bridge call, and a runbook that is posted in a linked Confluence repository. Access to that context enables Bits to develop hypotheses based on the available evidence rather than making assumptions.

    Responders can start an investigation by Bits from the Datadog web app or by using the @Datadog investigate command in an incident channel. They can also provide additional context as part of the request. For example, a responder can direct the investigation toward a specific service, deployment window, or suspected dependency issue. Bits uses that information as a starting point while continuing to analyze related telemetry data and incident context.

    As the investigation progresses, Bits publishes updates in a chat thread while it works. Bits then shares a final summary that includes root cause findings and recommended next steps. Team members can follow the investigation without switching to a separate AI interface or monitoring another application.

    Bits also becomes part of the incident record. The agent appears in the incident’s responder list, and the full investigation remains available within the incident overview page. During postmortems and audits, teams can review the investigation alongside the rest of the incident timeline.

    Catch up on active incidents with AI-generated chat summaries

    Many responders join incidents after remediation efforts are already in progress. Catching up often requires reading hundreds of messages, reviewing timeline events, and piecing together context from several conversations. That process can delay a responder’s ability to contribute.

    AI-generated incident summaries help responders get current context quickly. By using Datadog Incident Management integrations with tools such as Slack, Google Chat, and Microsoft Teams, responders can request a summary of an incident’s current state and recent activity. The summary is delivered as an ephemeral response that is visible only to the requester. As a result, the requester gets current context without creating additional channel noise or interrupting teammates who are actively working on remediation.

    Because the summary draws from the same incident context that powers other Incident Management workflows, the output reflects activity already captured in the incident channel and timeline. Responders receive a synthesized view of the investigation, ongoing remediation work, and key developments without needing to reconstruct the timeline manually.

    Teams can use AI-generated summaries throughout the incident life cycle. New responders can get up to speed quickly, incident commanders can review the latest state before communicating with stakeholders, and engineers who are returning to an incident can refresh their understanding.

    Capture bridge discussions in a unified incident timeline

    Critical decisions happen during incident bridge calls. Engineers discuss hypotheses, evaluate remediation options, and agree on next steps in real time. Without an automated way to capture those conversations, teams risk failing to document important context or must assign someone to record decisions and updates. Manual note-taking is an inefficient use of engineering time, requiring a skilled responder to divert their attention from resolving the incident.

    Incident Management integrates with video conferencing tools such as Zoom, Slack, Microsoft Teams, and Google Meet, automatically capturing incident bridge discussions and generating AI-powered summaries throughout meetings. After a meeting begins, a Datadog Transcriber joins the call and starts capturing discussion context.

    During active bridge calls, Datadog publishes AI-generated summaries approximately every 10 minutes to the incident timeline and the incident channel. Responders who join late or leave temporarily can review recent discussion points without interrupting the meeting.

    When the bridge call ends, Datadog automatically generates and publishes a post-meeting summary. The AI understands that the video call is happening in the context of an incident and tailors the summary accordingly, as opposed to using generic summarization. Key decisions, remediation plans, and discussion outcomes become part of the same incident record that already contains responder actions, status changes, automation activity, and telemetry data.

    You can also control which incidents receive summaries. Configuration options enable teams to scope summarization by severity, visibility, and tags. Private incidents are excluded from summarization by default.

    Get started with AI in Datadog Incident Response

    Datadog Incident Response combines AI-powered investigation, incident summarization, and meeting note-taking capabilities to help teams respond to incidents more effectively. Bits Investigation analyzes incident context alongside human responders, AI-generated summaries help engineers catch up on active incidents, and Incident Management collects decisions from incident bridge discussions and adds them to the incident record. Together, these capabilities unify telemetry data, investigation findings, chat activity, and meeting discussions in a shared context to help teams spend less time gathering information, coordinating handoffs, and documenting activity.

    To learn more about AI capabilities in Incident Response, check out the Incident AI documentation and Bits Investigation documentation.

    If you don’t already have a Datadog account, you can sign up for a 14-day free trial to start using Incident Response.

    Original source
  • Jun 30, 2026
    • Date parsed from source:
      Jun 30, 2026
    • First seen by Releasebot:
      Jul 1, 2026
    Datadog logo

    Datadog

    Debug and evaluate your AI app from your coding agent with Datadog Agent Observability

    Datadog adds Agent Skills for Agent Observability, bringing coding agents into the loop for trace classification, root-cause analysis, eval bootstrapping, experiment analysis, and end-to-end investigation workflows through MCP Server and Pup CLI.

    Coding agents like Claude Code, Cursor, and Codex CLI handle the coding parts of building an AI application well. The harder work comes after: understanding why a response went wrong, building eval sets that reflect real production behavior, and keeping up with an application that changes faster than any one-off script can. Teams spend 60–80% of their time on evaluation and error analysis, and much of that work needs to be redone every time the stack shifts.

    Datadog Agent Observability already captures the telemetry data needed to answer those questions. It traces every prompt and response and runs online evaluations over them. To make that telemetry data usable from inside your coding agent, we’ve built two foundations. The Agent Observability toolset in the Datadog MCP Server gives agents structured access to Agent Observability data. The Pup CLI, a command-line interface into much of Datadog’s API surface. On top of these foundations, we’re shipping a set of Agent Skills that package common AI engineering tasks into single commands. Drop them into your agent’s skills directory, and your coding agent can classify sessions, debug production failures, and evaluate new versions of your application against real traffic.

    In this post, we’ll show you how to:

    • Give your coding agent access to your Agent Observability data
    • Turn production traces into evaluation datasets directly in your coding agent
    • Analyze experiments and view the results in Datadog Notebooks
    • Take an investigation from traces to a coding-agent generated fix

    Give your coding agent access to your Agent Observability data

    While you’re in your coding agent, the data you most need to evaluate is already in Datadog. That includes your production LLM traces, evaluation results, and experiment metrics. The skills below all rely on the same foundation that pulls that data into context. Some teams prefer MCP for this and others prefer a CLI, so we support both.

    The Datadog MCP Server gives your agent native tool calls for searching traces, walking span trees, pulling experiment summaries, and writing findings into a Datadog Notebook. Two toolsets matter here: The llmobs toolset covers trace and experiment access, and the core toolset covers cross-product utilities like Notebook export. MCP is the right fit when you want your agent to reason about Datadog data alongside other tasks, and it works in any MCP-compatible client.

    Pup CLI drives the same Datadog API surface from the shell. A single command like pup llmobs traces search --ml-app task-cruncher --has-error returns trace IDs without leaving the terminal. Pup CLI is the right fit for scripting, CI runs, and agents whose workflows suit pipes better than tool calls. Both foundations are useful on their own, but the skills below package the most common eval workflows on top of them so you don’t end up scripting the same investigation repeatedly.

    To make the skills concrete, we’ll thread a single example through all of them. Task Cruncher is a conversational task-management assistant, similar to tools such as Linear or Jira, whose users have drifted from simple “create this task” requests toward multi-project coordination questions. The agent struggles with these types of complex queries, and since they’re not well represented in existing offline evaluations, the team has no signal that anything is wrong.

    Turn production traces into evaluation datasets directly in your coding agent

    The first three skills operate directly on production traces.

    1. agent-observability-session-classify: Assigns a binary thumbs-up or thumbs-down directly on traces, filling gaps where direct user feedback is missing.
    2. agent-observability-trace-rca: Runs an initial error analysis on production traces, with the ability to output results to a Datadog Notebook.
    3. agent-observability-eval-bootstrap: Bootstraps potential evaluators from production traces.

    agent-observability-session-classify

    The Task Cruncher team collects thumbs up/thumbs down feedback on every session, but only a few percent of customers click it. If customers are unhappy, they typically leave.

    The agent-observability-session-classify skill helps close this gap using a technique called weak labeling. When customers don’t provide explicit feedback on a session, this skill attempts to derive a binary label using existing Datadog data including the trace itself, Real User Monitoring (RUM), and Audit Trail. For each session, it outputs a thumbs up/thumbs down label and reasoning about the failure mode.

    agent-observability-trace-rca

    The next skill accepts traces and user feedback and performs an initial error analysis. Feedback can come from our end-user feedback feature (thumbs-up or thumbs-down), online evals, or the session-classify skill above. The skill runs on traces across a given time range in an ML application, sampling if necessary. It then outputs a report. If it runs with your codebase in context, it can also suggest specific fixes for your agent to implement.

    Beyond those suggested fixes, the skill helps you and your team understand errors in your traces. We’ve found the best format for this is a Datadog Notebook, which can be shared with the team and further modified with additional analysis. The skill supports native exports.

    agent-observability-eval-bootstrap

    The third skill operates on production traces to an improved eval set. The Task Cruncher team’s problem traces back to customers asking multi-project queries the agent wasn’t built to handle. This skill bootstraps new evaluators, either in our Python SDK or as a JSON struct you can import into an external evaluation framework.

    Analyze experiments and view the results in Datadog Notebooks

    Those three skills all operate on production traces and other online data. Switching over to the offline experiments side, we’ve got two skills to help set up and analyze experiments.

    1. agent-observability-experiment-py-bootstrap: Bootstraps a new experiment using our Python SDK, automating the manual steps in our existing setup.
    2. agent-observability-experiment-analyzer: Conducts a first-pass analysis of experiment results. Similar to the trace RCA skill above, it produces useful output for your coding agent to act on and a human-readable report that can be exported to a Notebook.

    agent-observability-experiment-py-bootstrap

    This Python-only skill creates an experiment using our Python experiments SDK. It asks a few questions, then covers environment setup, dataset creation, a placeholder experiment task your coding agent can fill in, and two or three evaluators in the requested style.

    Once created, you can run the experiment in Agent Observability’s experiments tool.

    agent-observability-experiment-analyzer

    In the Task Cruncher example, the team had existing experiments. However, they didn’t contain the necessary data to catch the problem they’re running into. Whether you’re like the Task Cruncher team and have an existing experiment that needs updating, or you’re creating your experiment from scratch like above, the agent-observability-experiment-analyzer skill can help you make sense of the results. The skill is run on a specific experiment ID. It first asks you which experiment metric, or metrics, you’d like to do an analysis on. Here, we select just the answer_quality_judge, which clearly has the lowest score of all evaluation metrics and our most interest in improving.

    Like the trace analysis skill, the experiment analysis skill produces a report you can export to a Notebook, along with actionable follow-ups for your coding agent. You choose which ones to implement, or you can examine the Notebook to decide what to do next.

    The Notebook contains the key findings, a summary of potential follow-ups, and supporting telemetry data.

    Take an investigation from traces to a coding-agent generated fix

    These skills feed into the agent-observability-eval-pipeline skill, an end-to-end pipeline.

    agent-observability-eval-pipeline goes from raw traces to a functioning experiment setup and analysis, with a few questions along the way. The skill walks you through the whole lifecycle: analyzing traces, constructing useful evaluators, generating a dataset and experiment, and analyzing the experiment to make meaningful code changes.

    This higher-order skill walks through the stages one at a time, introducing platform concepts that you may not have encountered before. Because the lifecycle can quickly get overcrowded with context, the stage-based approach keeps you focused on a single objective at a time. You can move between phases as you see fit, but you will always be working within one phase, which keeps the session focused.

    The pipeline runs in six phases, each with the same structure: a banner naming what’s being produced, a one-paragraph explanation of why it matters, the action (a sub-skill call or a small executable step), and a checkpoint that waits for your confirmation. The skill also links the generated artifacts so you can follow each step in the UI.

    Getting started with Agent Observability

    Pick whichever foundation matches how you work. MCP gives your agent native tool calls into Datadog data and is the right starting point if you already work in Claude Code, Cursor, Codex CLI, or Gemini CLI. Pup CLI gives you a shell-driven CLI for the same data plus the broader Datadog product surface, and is the better fit if you script a lot or run things in CI. You can use both for full coverage, then layer the skills on top.

    To use the MCP Server, add it from inside Claude Code.

    See the Datadog MCP Server setup docs for Cursor, Codex CLI, Gemini CLI, and other MCP-compatible clients.

    To use Pup, install and authenticate it:

    • go install github.com/datadog-labs/pup@latest
    • export PATH="$HOME/go/bin:$PATH"
    • pup auth login

    Tokens last about an hour. Run pup auth refresh if a command returns a 401. The full Pup CLI documentation covers the rest of the command surface.

    Install the skills by copying any of the skill directories from datadog-labs/agent-skills into your agent’s skills directory. For Claude, the install would look like this:

    • git clone https://github.com/datadog-labs/agent-skills
    • cp -r agent-skills/dd-llmo/traces-to-evals ~/.claude/skills/

    Agent Observability already contains the traces, evaluations, and experiment data needed to improve AI applications. By combining the MCP Server, Pup CLI, and Agent Skills, teams can move from investigation to evaluation and remediation without leaving their coding environment.

    Once your foundation is in place and the skills are installed, your coding agent can run these workflows against your production data. To learn more about the underlying telemetry data, see the Agent Observability documentation. If you’re not already a Datadog customer, sign up for a 14-day free trial to start instrumenting and collecting traces.

    Original source
  • Jun 24, 2026
    • Date parsed from source:
      Jun 24, 2026
    • First seen by Releasebot:
      Jul 3, 2026
    Datadog logo

    Datadog Agent by Datadog

    7.80.3

    Datadog Agent ships Prelude with Go 1.25.11, fixes for workload autoscaling consistency, Private Action Runner self-enrollment behind proxies, and zlib-compression handling for v3beta metrics intake shadow payloads.

    Agent

    Prelude

    Released on: 2026-06-24

    • Please refer to the 7.80.3 tag on integrations-core for the list of changes on the Core Checks

    Enhancement Notes

    • Agents are now built with Go 1.25.11.

    Bug Fixes

    • Workload autoscaling: fixed a bug where, when running the Cluster Agent in high-availability mode (multiple replicas), the burstable mode of a DatadogPodAutoscaler could leave the CPU limit in place on a random subset of pods. The CPU-limit removal is now re-derived from the autoscaler spec in the admission controller, so every replica applies it consistently regardless of which one handles the admission request.
    • Fix Private Action Runner self-enrollment failing silently on hosts with no direct internet access when a proxy is configured in datadog.yaml. Enrollment requests now respect the agent proxy settings (proxy.https, proxy.http, and no_proxy).
    • Disable v3beta metrics intake shadow payloads when zlib compression is used.

    Datadog Cluster Agent

    Prelude

    Released on: 2026-06-24 Pinned to datadog-agent v7.80.3: CHANGELOG.

    Original source
  • Jun 24, 2026
    • Date parsed from source:
      Jun 24, 2026
    • First seen by Releasebot:
      Jun 26, 2026
    Datadog logo

    Datadog

    Automatically enrich security logs with MITRE ATT&CK context before they reach your SIEM

    Datadog adds MITRE ATT&CK Enrichment Packs to Observability Pipelines, automatically tagging security logs with ATT&CK tactics and techniques across Okta, Palo Alto, FortiGate, AWS WAF, CloudTrail, and Windows for faster investigations and detections.

    To detect and investigate threats, security teams need to collect telemetry data from identity providers, cloud platforms, web application firewalls, and endpoints. But these diverse sources describe the same tactics, techniques, and procedures (TTPs) differently according to their own vendor-specific language. For example, a failed Windows logon appears as an event ID, while an Okta account lockout appears as an identity event. A firewall, meanwhile, may represent a similar attack through a completely different log format. Because of this incompatibility, analysts often spend valuable time translating vendor-specific events into a common security context before they can investigate or respond. This step adds overhead and slows security investigations.

    Observability Pipelines addresses this challenge with MITRE ATT&CK Enrichment Packs: preconfigured mappings that automatically tag security events with the relevant ATT&CK tactics and techniques. MITRE ATT&CK Enrichment Packs enrich your logs as they move through your pipeline, before they reach your SIEM, data lake, or archive. That means ATT&CK context is already there when the log lands, ready for detections, dashboards, investigations, and reporting.

    In this post, we’ll explore how these packs help teams:

    • Automatically enrich logs with MITRE ATT&CK context
    • Investigate security incidents across every source

    Automatically enrich logs with MITRE ATT&CK context

    MITRE ATT&CK gives security teams a shared framework for describing how attackers operate, from initial access through exfiltration. For telemetry pipelines, that context can help teams enrich logs and events with relevant tactics and techniques before the data reaches downstream security tools, making it easier to prioritize the signals that matter for detection, threat hunting, and incident response.

    Observability Pipelines now brings this context directly into your logs through MITRE ATT&CK Enrichment Packs. The initial release includes packs for:

    • Okta (identity and access): Tags authentication events that signal account abuse, including logins, MFA tampering, MFA fatigue, impersonation, API token creation, and account lockouts
    • Palo Alto (network and perimeter): Tags firewall activity including exploit attempts, command-and-control traffic, malware transfers, denial-of-service, VPN access, and admin brute force
    • FortiGate (network and perimeter): Tags the same firewall behaviors as Palo Alto and also maps data-loss events to exfiltration techniques
    • AWS WAF (web): Tags web-layer attacks including exploit attempts, brute force, bot activity, anonymous proxy traffic, and credential stuffing
    • CloudTrail (cloud): Tags cloud control-plane activity including console logins, IAM changes, defense evasion, and cloud infrastructure reconnaissance
    • Windows (endpoint): Tags endpoint-only behaviors like PowerShell execution, scheduled task creation, service installation, and event log clearing

    Teams can browse and add packs directly from Observability Pipelines. Each pack comes preconfigured with mapping logic, so there’s nothing to build from scratch.

    For example, suppose you’re a security engineer using Okta logs in Splunk to detect identity attacks. On its own, an event such as user.session.impersonation.grant means something only to analysts who already know Okta’s event taxonomy. After you add the Okta Pack, that same event arrives tagged as the MITRE ATT&CK tactic “Privilege Escalation” and the technique “Use Alternate Authentication Material (T1550).” Detection rules can then target MITRE ATT&CK fields rather than vendor-specific event names.

    Once the pack is added, you can validate the enrichment logic against production log samples by using Live Capture. Live Capture shows exactly how an event is transformed as it moves through the pipeline. For example, the raw Okta event below enters on the left and exits on the right with its MITRE tags added, along with a security:true flag and a source field.

    Investigate security incidents across every source

    Security teams often spend as much time normalizing data as investigating threats. When the same attacker behavior appears in different formats across systems, teams can end up maintaining separate detection logic for each source.

    MITRE ATT&CK Enrichment Packs apply one enrichment model across all supported sources. Identity events, network activity, cloud control-plane changes, and endpoint behaviors all arrive with the same fields, so teams write detections once and build dashboards on the same fields across every source. Inside each pack, processors match specific events and apply the corresponding MITRE ATT&CK mappings automatically, so analysts receive events already labeled with attacker intent. Each enriched event also carries a flag marking it as security-relevant.

    That standardization is what makes investigations fast. Say you’re investigating a possible account takeover. Without MITRE tags, you’d query each source in its own syntax and manually stitch together the timeline. With the tags applied in the pipeline, you can filter on @mitre.tactic:Privilege Escalation and see those events side by side, already labeled, from the moment they arrive. From there, you can narrow down to a specific technique or widen to @security:true to surface every security-relevant event at once.

    Because the tagging happens in the pipeline, before logs leave your environment, every destination gets the same context automatically. You can route the tagged events to the SIEM or data lake of your choice, including Splunk, Microsoft Sentinel, Datadog Cloud SIEM, or a data lake like Databricks or ClickHouse. The MITRE fields are already there when the log lands, and detection rules fire right away, without anything needing to be looked up.

    Start enriching your security logs today

    By tagging logs with MITRE ATT&CK context, you normalize security logs automatically, investigate activity across sources with a shared taxonomy, and speed up detections.

    MITRE ATT&CK Enrichment Packs are included with Observability Pipelines at no additional cost and are available in all regions and environments outside GovCloud. To get started, open the Packs gallery in Observability Pipelines and add the pack that matches your log source. To learn more, check out the Observability Pipelines documentation, or read our blog on how Observability Pipelines can enrich logs with additional context on-stream. And if you’re not yet a Datadog customer, sign up for a 14-day free trial.

    Original source
  • Jun 22, 2026
    • Date parsed from source:
      Jun 22, 2026
    • First seen by Releasebot:
      Jun 26, 2026
    Datadog logo

    Datadog

    Using Evaluation Frameworks with Agent Observability

    Datadog adds native support for DeepEval and Pydantic Evals in Agent Observability, letting teams run, compare, and monitor LLM evaluations in Datadog Experiments with trace-linked results, regression visibility, and continuous production traffic scoring.

    AI teams have invested heavily in evaluation frameworks, yet getting those frameworks beyond local experimentation remains challenging.

    Teams using open source libraries like DeepEval and Pydantic Evals gain flexibility and research-grounded metrics, but operationalizing those evaluations still requires brittle custom integration code that doesn’t scale. SaaS eval platforms often prioritize convenience, which can come at the cost of flexibility when teams need to port or extend their metric definitions over time. The result is that even mature teams with carefully tuned, task-specific evaluators end up with siloed artifacts: evals that work in a notebook, break in CI, and vanish entirely in production monitoring.

    In this post, we explain how Datadog Agent Observability addresses this gap by letting teams run their existing DeepEval evaluations natively within Datadog Agent Observability Experiments. Datadog also supports Pydantic Evals, a code-first evaluation framework that provides its own dataset, evaluator, and LLM-as-a-judge primitives, for teams that prefer it or already use it alongside Pydantic AI. The examples in this post use DeepEval, but the same patterns also apply to Pydantic Evals. Together, these integrations give teams a single place to define, run, and monitor evaluation quality across every stage of development and deployment.

    We’ll cover:

    • Why framework portability matters for LLM evals
    • How to set up experiments with Datadog Agent Observability
    • How to connect eval scores to production traces
    • How to run LLM evaluations continuously on production traffic

    Why framework portability matters for LLM evals

    Evaluations are an engineering asset, not a platform feature. A team that has built a suite of DeepEval evaluations has accumulated organizational knowledge about what “good” looks like for their application. That knowledge is encoded in the rubrics, thresholds, and human validation behind every G-Eval judge, RAG faithfulness metric, and custom evaluator in the suite. Rewriting those evaluators to conform to a platform’s proprietary metric definitions means discarding that investment rather than simply porting it.

    Datadog Agent Observability doesn’t replace the open source eval ecosystem but wraps around it. You define what to measure and how to measure it, using the frameworks you already trust. The platform handles operationalization. It runs those evaluations at scale across hundreds or thousands of examples and tracks results over time to surface regressions. It also monitors token usage and cost across runs, and connects offline eval scores to production traces so you can verify that improvements in your Experiments environment actually translate to better user experiences. The open source scaffolding stays intact. The platform provides infrastructure for continuous eval runs, trace-linked regression visibility, and verification that offline improvements hold in production.

    Set up experiments with Datadog Agent Observability

    Before running experiments, enable Agent Observability in your Datadog account and install the required libraries. The example below uses ddtrace 4.8 or later and works with any version of DeepEval:

    pip install ddtrace deepeval pydantic
    

    Then enable Agent Observability instrumentation in your application:

    from ddtrace.llmobs import LLMObs
    LLMObs.enable(
        ml_app = "your-llm-app",
        api_key = "<YOUR_DD_API_KEY>",
        app_key = "<YOUR_DD_APP_KEY>",
        site = "<YOUR_DD_SITE>",
    )
    

    Step 1: Define your dataset

    A dataset is a collection of inputs and expected outputs. The inputs are passed directly to your task function, whether that is a RAG pipeline, an agent, or any other LLM application, which produces an actual output. The experiment then compares that actual output against the expected output you provide to score each example. All you need to define a dataset are a name, a version, and a list of those input and expected output pairs.

    from ddtrace.llmobs import LLMObs
    dataset = LLMObs.create_dataset(
        dataset_name = "rag-customer-support-v1",
        description = "Example dataset containing customer support examples",
        records =[
            {
                "input_data" : {
                    "question" : "How do I reset my password?"
                },
                "expected_output" : {
                    "answer" : "Click 'Forgot Password' on the login page..."
                },
                "metadata" : {
                    "difficulty" : "easy"
                }
            },
            {
                "input_data" : {
                    "question" : "What's your refund policy?"
                },
                "expected_output" : {
                    "answer" : "We offer 30-day refunds for..."
                },
                "metadata" : {
                    "difficulty" : "easy"
                }
            },
        ],
    )
    

    Step 2: Configure your DeepEval or Pydantic evaluator

    Existing DeepEval metrics like G-Eval judges, RAG faithfulness metrics, and custom LLM-as-a-judge implementations can be used without modification.

    from deepeval.metrics import GEval
    from deepeval.test_case import LLMTestCaseParams
    helpfulness_evaluator = GEval(
        name = "Helpfulness",
        criteria = "Determine whether the response directly answers the user's question with actionable steps.",
        evaluation_steps =[
            "Check whether the content of the 'actual output' contradict the content of the 'expected output'",
            "You should also heavily penalize omission of detail",
            "Vague language, or contradicting OPINIONS, are not OK",
            "The user's question should be answered by the 'actual output'"
        ],
        evaluation_params =[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
        async_mode = True,
    )
    

    Setting async_mode=True runs evaluations concurrently across the dataset. For a dataset of 100 examples, this can significantly reduce the time evaluations take to run.

    Step 3: Define your task and run the experiment

    The task function takes an input from your dataset and returns an output, which is typically a call to your LLM application or RAG pipeline.

    from ddtrace.llmobs import LLMObs
    def my_rag_task(input_data):
        question = input_data["question"]
        response = your_rag_pipeline(question)
        return {"answer": response}
    
    experiment = LLMObs.Experiment(
        name = "rag-customer-support-baseline",
        dataset = dataset,
        task = my_rag_task,
        evaluators =[helpfulness_evaluator]
    )
    
    experiment.run()
    

    When experiment.run() is called, Datadog executes the task function across every example in the dataset, runs the DeepEval metrics in parallel, and uploads results to the Experiments UI for analysis.

    Analyze experiment results in Datadog

    Once an experiment completes, Datadog makes results available in the Datadog Agent Observability Experiments UI. You can select any prior experiment run as a baseline and view side-by-side comparisons of eval scores, latency, token usage, and cost. If switching to a different model improved helpfulness scores by 12% but introduced a 3× latency increase, the same view shows both changes without cross-referencing separate tools.

    For any low-scoring example, you can drill into the full trace to see the exact prompt sent to the model, the completion, the eval score, and evaluator reasoning. This visibility reduces the need to reproduce failures locally or reconstruct context from logs after the fact.

    Connect eval scores to production traces

    Eval scores in isolation have a practical ceiling. A helpfulness score that drops from 0.82 to 0.74 between runs raises questions about what caused the drop. Answering it requires knowing which examples regressed, what changed in the prompt or model output, and whether the issue originated in retrieval or generation. It also requires understanding how the regression correlates with latency or token usage.

    Without observability, this means manually correlating data from an eval framework and a separate logging system. Engineers have to copy trace IDs, cross-reference timestamps, and piece together context that should already be connected.

    Running DeepEval metrics with Datadog automatically links every eval score to the trace, prompt, and token count that produced it. Regressions are clickable, explorable, and reproducible within the same Datadog platform used to monitor the rest of your application.

    Run LLM evals continuously on production traffic

    Most teams treat evals as a pre-deployment gate where a batch job in CI that produces a pass or fail decision before a change ships. These evals can catch regressions before they reach users, but they do not surface issues that emerge in production as traffic patterns, user inputs, or upstream dependencies change over time.

    With Datadog, evals can run continuously on sampled production traffic alongside offline experiment workflows. The same evaluators used during development can score live completions, and the results feed into the same dashboards and alerting infrastructure used for the rest of the application stack. Teams can catch quality regressions as they happen rather than learning about them from user feedback.

    Get started with Datadog Agent Observability

    Datadog Agent Observability lets teams run DeepEval and Pydantic Evals evaluations natively within Datadog Experiments without needing to rewrite existing evaluators or adopting proprietary metric definitions. By connecting offline eval scores to production traces, teams can catch quality regressions at every stage of development and deployment, not just at the pre-deployment gate. As LLM applications grow more complex, continuous evaluation against live traffic becomes as essential as any other part of the observability stack. To learn more, check out the Agent Observability documentation.

    If you don’t have a Datadog account, you can sign up for a 14-day free trial to get started with Agent Observability.

    Original source
  • Jun 17, 2026
    • Date parsed from source:
      Jun 17, 2026
    • First seen by Releasebot:
      Jul 3, 2026
    Datadog logo

    Datadog Agent by Datadog

    7.80.2

    Datadog Agent releases a security-focused update with CIS Docker compliance rule tuning, a Cluster Agent AppSec admission mutator fix, improved container log recovery for idle streams, and OTel metrics sampling changes. Datadog Cluster Agent also fixes AKS webhook reconciliation issues.

    Agent

    Prelude

    Released on: 2026-06-17

    • Please refer to the 7.80.2 tag on integrations-core for the list of changes on the Core Checks

    Enhancement Notes

    • Compliance: CIS Docker rules (scope: docker) are no longer evaluated on Kubernetes nodes where the kubelet's CRI runtime is not Docker (e.g. containerd, CRI-O), avoiding false positives on GKE Container-Optimized OS which ships dockerd alongside containerd. The runtime is read from the kubelet's --container-runtime-endpoint flag or the containerRuntimeEndpoint field of its --config YAML; if it cannot be determined the rules continue to evaluate.

    Security Notes

    • Fixed a confused-deputy vulnerability in the Cluster Agent's AppSec ingress-nginx admission mutator where the pod's --configmap=/ argument was trusted verbatim, allowing a user with pod-create permission in one namespace to make the Cluster Agent service account create or update ConfigMaps and add labels and annotations in arbitrary namespaces. The mutator now requires the portion to match the pod's own namespace (or use the $(POD_NAMESPACE) downward-API substitution) and skips mutation otherwise, emitting a warning event on the pod. The vulnerability affected Cluster Agent releases starting from 7.78.0.

    Bug Fixes

    • Fix an issue where container log collection could stop for an individual container without recovering and without any error in the Agent logs. When a container's log stream was idle longer than logs_config.docker_client_read_timeout, the read timeout could cause the underlying Docker connection to close in a way that the tailer treated as a permanent shutdown, silently stopping log collection for that container until it was recreated or the Agent was restarted. The tailer now reconnects in this case, and only stops when the Agent is intentionally shutting down. Low-volume containers (for example, services that log only periodically) were the most affected.
    • OTel Agent: Disable v3 series API shadow sampling, which is incompatible with the zlib compression the OTel Agent forces for the metrics intake.

    Datadog Cluster Agent

    Prelude

    Released on: 2026-06-17 Pinned to datadog-agent v7.80.2: CHANGELOG.

    Bug Fixes

    • Fixed an issue where the admission controller connectivity probe webhook did not include the AKS selector requirements when admission_controller.add_aks_selectors was enabled, which could cause repeated webhook reconciliation conflicts on AKS.
    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.