Pinecone Release Notes

Follow

40 release notes curated from 30 sources by the Releasebot Team. Last updated: Oct 2, 2026

Get this feed:
  • September 2026
    • Date parsed from source:
      Sep 28, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Pinecone logo

    Pinecone

    Released Go SDK v7.0.0

    Pinecone releases v7.0.0 of the Go SDK with the latest stable API, schema-based document indexes, and the Documents API for full-text search. It also includes migration-breaking changes, updated module path, and removal of pod-based index creation.

    Released v7.0.0 of the Pinecone Go SDK. This version uses the latest stable API version, 2026-07, and adds support for schema-based document indexes and the Documents API, which brings full-text search to the Go SDK. The v6 release line continues to target 2026-04.

    Warning

    Before upgrading to v7.0.0, update all relevant code to account for the following breaking changes. See the v7 migration guide for full details.

    • The module path is now github.com/pinecone-io/go-pinecone/v7.
    • Pod-based index creation is removed (CreatePodIndex, CreatePodIndexRequest). Existing pod-based indexes keep working.
    • Enum constants are prefixed with their type name. For example, Aws is now CloudAWS, and Cosine is now IndexMetricCosine.
    • ListImports takes a *ListImportsRequest, and QueryByVectorIdRequest.SparseValues is removed.
    • SourceCollection and Schema on the legacy create requests now return an error. ConfigureIndexParams.Embed is replaced by ConfigureIndexParams.Schema.
    • Backup field types changed, and some invalid requests now fail client-side.
    Original source
  • September 2026
    • No date parsed from source.
    • First seen by Releasebot:
      Sep 22, 2026
    Pinecone logo

    Pinecone

    Pinecone BYOC: Trusted AI Knowledge in the Customer Cloud

    Pinecone announces general availability of Bring Your Own Cloud on AWS, Google Cloud, and Azure, letting enterprises run its AI knowledge platform inside their own cloud accounts while keeping customer data in their boundary and preserving the managed Pinecone experience.

    Today, we are announcing the general availability of Pinecone Bring Your Own Cloud (BYOC) on AWS, Google Cloud, and Azure, bringing Pinecone’s trusted AI knowledge platform to where enterprise data needs to live.

    AI becomes transformative when it works with a company’s proprietary knowledge. Customer context, policies, and operational history allows its agents to make decisions and carry out work using expertise the business has built over years.

    Organizations have spent years controlling where sensitive knowledge lives and who can reach it. Providing access to it typically meant managing knowledge infrastructure ranging from inference, document parsing, and vector databases. Platform teams shouldered the burden of tuning and maintaining the system, including keeping retrieval quality and performance stable across a diverse set of AI workloads.

    With BYOC, customer data and the knowledge derived from it remain in the customer’s account, while Pinecone manages the platform operations. This means teams can bring sensitive AI workloads to production without taking on the complexity of operating knowledge infrastructure themselves. The APIs and interfaces remain the same as the managed service, providing organizations with the flexibility to select the right deployment model for each workload based on its security, connectivity, and operational requirements.

    Keeping proprietary knowledge inside the customer cloud

    Pinecone’s platform architecture separates the systems that manage the service from those that store and process customer data.

    • Control Plane: Handles management operations such as resource lifecycle, authentication, and service health. It does not store or process customer content or request payloads.
    • Data Plane: Stores, processes, and serves customer data and knowledge. AI agents and applications connect directly to this for read and write operations. The only data shared with Pinecone are anonymized operational metrics and traces for monitoring and support.

    With BYOC, the data plane runs inside the customer's selected cloud account and region, including those beyond where Pinecone's standard service is available. Vectors, documents, metadata, and request payloads remain within the customer-controlled boundary.

    Zero-access BYOC model

    Pinecone does not require SSH, VPN, inbound network access, or a standing cross-account IAM role to manage the service. Upgrades, scaling actions, and maintenance work are retrieved using an outbound call from the Pinecone control plane and executed locally.

    This pull-based mechanism allows Pinecone to manage the database without a persistent access path into the customer environment. Additionally, BYOC works alongside SSO, RBAC, SCIM + SAML, audit logging, encryption, and private-networking controls available with Pinecone's Enterprise plan so customers can have complete confidence in ensuring their proprietary knowledge is secure.

    Keep the managed Pinecone experience

    In addition to Pinecone handling upgrades, scaling, maintenance, and service health monitoring, customers retain access to Pinecone’s support and engineering teams for troubleshooting, incident response, and ongoing operational guidance.

    Teams use the same Pinecone APIs, SDKs, and control plane workflows across the BYOC and standard deployments. This means each workload can use the deployment model that fits its data governance and access requirements without creating a separate development path.

    Toyota brings manufacturing knowledge to AI within its environment

    Toyota Motor North America (TMNA) was one of Pinecone’s first BYOC customers to deploy it in production. TMNA used Pinecone to ground AI applications with decades of proprietary manufacturing knowledge while keeping that knowledge secure inside Toyota’s environment.

    “Decades of manufacturing know-how live in our documentation, and that institutional knowledge is one of the most valuable assets we have. It also happens to be complex — highly structured engineering data sitting alongside unstructured process documents, across a lot of formats and a lot of different access patterns. Pinecone BYOC runs inside our own environment, so that knowledge never leaves our boundary and is served only to models we’ve already vetted. It handles that complexity at the scale our operations demand, with the enterprise security and governance controls our teams require. A critical requirement for how our team can use AI with confidence.”

    — Kordel France, Head of AI Engineering, Toyota Motor North America

    Bringing trusted AI knowledge to more environments

    Our vision is to have Pinecone deployable wherever enterprise knowledge is. BYOC extends Pinecone’s trusted AI knowledge platform to customer-controlled cloud environments today, and our work continues beyond BYOC.

    We are developing a fully self-managed option for air-gapped and highly restricted networks where both the control plane and data plane will run inside the customer environment. Reach out if you're interested in shaping the security and deployment requirements of a self-managed Pinecone offering.

    Get started

    Talk with your Pinecone account team to review your requirements and plan your BYOC deployment, or contact us to get connected with us.

    For more information, the Pinecone BYOC product page and the BYOC documentation cover the operating model and the architecture in detail.

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Pinecone and hundreds of other software products.

    Create account
  • Sep 17, 2026
    • Date parsed from source:
      Sep 17, 2026
    • First seen by Releasebot:
      Sep 17, 2026
    Pinecone logo

    Pinecone

    VQ-bench: a Composable Vector Quantization Framework

    Pinecone launches VQ-bench, a new open-source benchmark for vector quantization with a public website, GitHub repo, and research paper. It standardizes fair, reproducible comparisons across popular quantizers and highlights early results on recall, reconstruction error, and encoding speed.

    Quantizers

    Before a vector database can search vectors, it has to store them. But storing high-dimensional vectors at full precision is quite expensive. Vector quantization (VQ) reduces the number of bits needed to store a vector, making it a critical part of maintaining a vector database.

    Because VQ is so important (to both vector databases and LLMs), many research papers are published on the topic every year. Pinecone has been using quantization since its first prototypes. But we can always do better, so we set out to survey and benchmark newer results. We were pretty overwhelmed by just how many quantizers are out there. To make matters worse, every paper seemed to evaluate performance differently, measuring different metrics on different datasets and optimizing for different hardware. We were unable to find any systematic attempt to evaluate the leading methods against one another.

    Of course, faithfully implementing dozens of quantizers from scratch comes with its own challenges. Luckily, as we dug deeper into the literature, we began to notice a pattern. Many published quantizers are actually just slight variations of existing ones. In fact, most of them are built from a relatively small set of primitive operations. That gave us an idea: what if we published an open-source library of these core primitives, where building a quantizer was as easy as writing a recipe of which primitives to use and in what order? Then, we would be able to evaluate all of these quantizers in a fair and reproducible way. It would also make it easier to experiment with new variations of existing quantizers or invent new ones altogether.

    This was the start of the VQ-bench project. With this post, we're excited to share VQ-bench with the public, including:

    • A public website with a running benchmark of popular quantizers
    • A GitHub repo where you can contribute your own quantizers and primitives
    • A paper on VQ-bench (presented at VecDB@VLDB 2026), along with the talk slides.

    Note that this is just the first iteration of VQ-bench; we encourage feedback, corrections, and contributions, and we will add more quantizers over time.

    Quantizers

    A quantizer is anything that can take a set of vectors, compress them, and recover desired information later on. In VQ-bench, a quantizer must implement four methods:

    • fit: given a sample of vectors (and optionally queries), learn a model
    • encode: given the model and a set of vectors, return per-vector codes
    • reconstruct: given the model and the code for vector x, reconstruct it
    • score: given the model, a query vector q, and the code for x, estimate the dot-product score ⟨q, x⟩

    Primitives

    Quantizers are rarely built from scratch. In the literature, they are assembled from a small set of basic operations, which VQ-bench formalizes as primitives. A primitive implements the same four methods as any other quantizer, plus two more that specify exactly how it hands data to the next stage:

    • apply: given the model, transform the vectors into what the next stage should see
    • apply_queries: given the model, transform the queries into what the next stage should see

    A primitive's reconstruct and score methods also take as input the next stage's reconstruction and score estimate, respectively. That makes six methods in total. The extra two are the chaining contract: they are what let primitives be composed, which is the subject of the next section.

    VQ-bench implements three groups of primitives:

    • Conditioners: transform the data and pass it downstream (Center, Normalize, PCA, RandomRotate, ...).
    • Rounders: cast each vector to a finite codebook, passing the residual downstream (CastUint, CastAngular, CastNormal, KMeans, ...).
    • Splitters: split the vectors and quantize each part with its own chain of primitives (Segment).

    Pipelines

    A pipeline is a special type of quantizer given by composing two or more primitives in a chain. Compressing a vector walks it forward through the chain, and recovering a vector (or its score) walks it backward.

    • The forward pass: fit and encode follow the same path. At each stage, they perform that stage's job (learning the model / computing the codes). Then, they call apply to transform the vectors to the next stage and recurse. At the end, fit concatenates each stage's model and encode concatenates each stage's codes.
    • The backward pass: reconstruct starts at the last stage. Each stage above it folds its own contribution back in (e.g., adding back the mean, undoing a rotation, etc.) until the first stage has an approximation of the original vector.
    • score works the same way, except every stage needs the query as it saw the data. So, it begins by walking just the query forward with apply_queries. Then, it performs the backward pass on the score.

    A quantizer does not have to be a pipeline. Anything that implements the four methods qualifies, and the interface leaves room for methods that are built some other way. But most published quantizers can be expressed as pipelines of primitives, which is what makes the decomposition worth building on.

    For example, E-RaBitQ is a popular quantizer (which we found to be quite performant in our experiments). The E-RaBitQ pipeline consists of four primitives:

    1. Center: subtract the average dataset vector from each vector
    2. Normalize: scale each vector to unit norm
    3. Random Rotation: apply a random orthogonal (or random Hadamard) rotation to each vector
    4. Angular Cast: snap each vector to a b-bit integer grid by rounding to the nearest grid point in angle.

    A diagram of this pipeline and table for the primitive functions are given below.

    Experimental Results

    We evaluated a suite of 14 quantizers on 5 datasets from VIBE. Each dataset consists of vectors to encode and queries to score. Below, we present some results for two of the datasets: ArXiv (1,344,643 vectors in 768 dimensions) and Yahoo (677,305 vectors in 384 dimensions). You can view the full results on the website.

    Reconstruction error

    Reconstruction MSE is the traditional metric for VQ, and it's important for applications like LLM weight compression. To measure it, we sample 1000 random dataset vectors x. A quantizer reconstructs x̂ and we measure the average value of ∥x̂ − x∥².

    Recall

    For vector databases, a more relevant metric is recall, specifically for reranking. To measure it, we take each query and compute the 1000 dataset vectors of maximum dot-product. A quantizer estimates these 1000 scores, and we measure what fraction of the estimated top-10 were contained in the true top-10 (averaging this fraction over all queries).

    Encode time

    We also measure how long it takes to encode the entire dataset. Note that encoding is done in chunks and accelerated via multithreading. These results were obtained on an Apple M2 Pro with 16GB RAM using 6 threads.

    Discussion

    Overall, we can see some clear trends. PQ and OPQ consistently have the lowest reconstruction MSE. EDEN and E-RaBitQ are comparable in terms of recall, especially at higher bit budgets. EDEN is also much faster to encode than PQ, OPQ, and E-RaBitQ, making it a good candidate for most quantization applications.

    Contribute

    We built VQ-bench to be extended, and the repo takes two kinds of contributions.

    Got a new quantizer? Usually just a few lines of code. The E-RaBitQ pipeline above is four primitives in a list, and many published quantizers are a similar reordering of primitives the library already ships.

    Got a new primitive? Implement the six methods above and it composes with every other primitive in the catalog. Every pipeline can use it, including the ones nobody has written yet.

    Either way, you get the evaluation harness. A short config runs your method over the whole suite, measured exactly the way every other method is measured: recall@k, reconstruction and score error, bias, softmax KL and total variation, size in bits per dimension, and encode, score, and reconstruction cost. Both lists keep growing as we add datasets and metrics. We refresh the published benchmark on a regular cadence, and new methods are folded in then.

    We also want corrections. If we implemented your quantizer wrong, or we missed a method worth including, open an issue and tell us.

    Original source
  • September 2026
    • Date parsed from source:
      Sep 15, 2026
    • First seen by Releasebot:
      Sep 17, 2026
    Pinecone logo

    Pinecone

    Pod-to-serverless migration supports up to 500 million records

    Pinecone now supports migrating pod-based indexes to serverless with up to 500 million records.

    Migrating a pod-based index to serverless now supports indexes with up to 500 million records, up from 50 million. If you need to migrate a larger index, contact Support before you start.

    Original source
  • September 2026
    • Date parsed from source:
      Sep 15, 2026
    • First seen by Releasebot:
      Sep 17, 2026
    Pinecone logo

    Pinecone

    Gemini 3.5 Flash now available for Assistant chat

    Pinecone adds support for Google’s Gemini 3.5 Flash model in Pinecone Assistant.

    Pinecone Assistant now supports Google’s Gemini 3.5 Flash model. To use this model, set model: "gemini-3.5-flash" in your chat requests. For more information, see Choose a model.

    Original source
  • Similar to Pinecone with recent updates:

  • September 2026
    • Date parsed from source:
      Sep 15, 2026
    • First seen by Releasebot:
      Sep 17, 2026
    Pinecone logo

    Pinecone

    Gemini 2.5 Pro and o4-mini deprecations for Assistant

    Pinecone Assistant now routes Gemini 2.5 Pro to Gemini 3.5 Flash and o4-mini to GPT-5 at the same price.

    Google has deprecated the Gemini 2.5 Pro model, and OpenAI has deprecated o4-mini. Pinecone Assistant automatically routes all chat requests that specify gemini-2.5-pro to Gemini 3.5 Flash, and all chat requests that specify o4-mini to GPT-5, at the same price. No code changes are required. To update your code to explicitly use the newer models, set model: "gemini-3.5-flash" or model: "gpt-5" in your chat requests. For more information, see Choose a model.

    Original source
  • September 2026
    • Date parsed from source:
      Sep 13, 2026
    • First seen by Releasebot:
      Sep 14, 2026
    Pinecone logo

    Pinecone

    Higher namespace limit on the Enterprise plan

    Pinecone raises the Enterprise namespace limit for serverless indexes to 1,000,000, with higher limits available on request.

    The Enterprise limit for namespaces per serverless index is now 1,000,000. Pinecone can accommodate more for specific use cases. Contact Support to request a higher limit.

    Original source
  • Sep 9, 2026
    • Date parsed from source:
      Sep 9, 2026
    • First seen by Releasebot:
      Sep 9, 2026
    Pinecone logo

    Pinecone

    Full-Text Search is Now Generally Available In Pinecone Database

    Pinecone introduces generally available full-text search in Pinecone Database, adding BM25 keyword ranking and text-match filters alongside dense vectors in one index. It helps retrieval, RAG, recommendations, and agents find exact identifiers and phrases without a separate search system.

    If you run search, recommendations, RAG, or agents, the queries hitting your retrieval system come in more than one kind. Semantic search was built for when you need to search by meaning. The embedding captures what the user meant, and a question finds the right data even when its wording differs from the text.

    Identifiers and exact phrases are a different kind of query. SKUs, part numbers, error codes, case numbers, and order IDs are literal strings. The literal string is the query, and embeddings rarely match literal strings. In a best case scenario, an embedding model splits PROD-001 into subword tokens (if the token was present in the embedding model's training data) and places it beside every other SKU. So, a search for the part number a customer provided, PROD-001 for this example, returns PROD-002 and PROD-003 alongside the correct part. The result is confidently close when the user needs exactly correct. Nothing throws an error. The agent answers from the wrong documentation, and nothing in the response indicates that there's a problem at all.

    Matching exact words has usually meant running a node-based search system. These engines are built for the job, and they are good at matching exact words, but running one is its own project that your team owns: nodes to size, capacity to plan as the corpus grows, shards, memory, merge cycles to tune, and version upgrades to run. The engine is good at the job, but keeping it running requires your team's time and effort.

    After months of private preview validation with our user partners, today, we're announcing that full-text search is generally available in Pinecone Database. It adds BM25 keyword ranking and text-match filters to the same index as your dense vectors, so agents get the exact order number or error code embeddings miss, with no separate system to tune and maintain.

    TL;DR

    If your search, recommendation, RAG, or agent workload mixes questions with identifiers and exact phrases, Full-Text Search gives you:

    • Exactly correct instead of confidently close on the queries embeddings get wrong: SKUs, error codes, case numbers, order IDs, and exact quoted phrases.
    • BM25 keyword ranking across multiple text fields per index, with Lucene query syntax, fuzzy matching, and tokenization and stemming in 18 languages.
    • Text-match filters that restrict the results of a semantic search to the documents matching the filter. One query both retrieves based on the identifier and ranks the rest by meaning.
    • One index and one schema for text fields, dense vectors, and sparse vectors, queried through the Documents API, with no cluster to tune, patch, or upgrade.
    • Usage-based capacity by default, metered in the same read units and write units as vector indexes. Dedicated Read Nodes are available for provisioned read capacity, and BYOC for running inside your own cloud.

    The queries vector embeddings get wrong

    Semantic search works on the queries it was designed for, where meaning is the signal. A literal string is a different test, and a vector-only index fails it the same way every time.

    • Lookalikes rank as well as the right record. Every SKU with the same pattern shares most of its subword tokens with PROD-001, and the ranking has no way to tell which one the customer meant.
    • The failure is silent. A vector search always returns its nearest neighbors. There is no empty result to check for and no error to catch. A wrong answer looks the same as a right one until a customer notices.
    • The alternative is a node-based lexical engine. A lexical engine handles the exact word queries well, but it comes with nodes your team sizes, tunes, and upgrades.

    Those failures have a business cost:

    • Wrong actions that look right. An agent that retrieves the lookalike ships the wrong part or applies the wrong policy, and each one erodes users' trust.
    • An ongoing cost of ownership. Someone sizes the nodes, plans capacity as the corpus grows, and runs the upgrades, and that work continues for as long as the system is in production.
    • Slower delivery. Engineers spend their time maintaining search infrastructure instead of building the product.

    Full-text search is built for the queries where the answer depends on a literal string, and it runs in the same index as the semantic search that handles all of your other queries.

    Where exact-match precision matters

    Identifier-heavy retrieval.

    SKUs, part numbers, error codes, and case numbers are where embedding models fail most visibly, and where full-text search shines. BM25 ranks the term itself, and a text-match filter returns nothing when no document carries the identifier. So, a miss reads as a miss instead of a lookalike.

    Recommendations with hard constraints.

    A customer who names a specific model or a compatible part number expects every recommendation to honor it, and embeddings alone can't do that (e.g., accessories for PROD-001 land next to the ones for PROD-002). A text-match filter pins the candidate set to items whose text carries the literal term, and the dense vector query ranks within it, so similarity never overrides a constraint the user stated. This is also true for content recommendations keyed to a ticker symbol, regulation number, or a streaming video ID.

    RAG over enterprise documents.

    A conceptual question needs semantic search. A query for a specific clause or an error code requires the exact term, and only lexical search reliably returns it. With full-text search and semantic search in the same index, the RAG system retrieves the clause the question depends on instead of a passage that resembles it.

    Agent tool use over structured data.

    Agents often need to look up data using a value from an earlier step, like an order number or a status string. Full-text search matches the value itself, so every step after that builds on the right data.

    How full-text search works in Pinecone

    Pinecone Database indexes created with the Documents API now run on a document schema. You declare text fields, dense vector fields, and sparse vector fields in one schema, and one index holds all of them.

    You keep:

    • The same Pinecone APIs and SDKs
    • The same index lifecycle
    • The same usage-based capacity model, metered in storage, read units, and write units
    • The same serverless architecture, with no nodes for your team to maintain
    • The same deployment options, including Dedicated Read Nodes and Bring Your Own Cloud

    You add:

    • BM25 keyword ranking across multiple text fields per index
    • Lucene query syntax, including boolean operators and phrase queries
    • Text-match filters that restrict any search, including a semantic one, to documents matching the filter
    • Fuzzy matching with the ~ operator for typo tolerance
    • Tokenization and stemming in 18 languages, plus language-agnostic n-gram tokenization for substring- and prefix-matching

    Search that finds what you mean and exactly what you say

    An agent that retrieves a lookalike record fails quietly. It ships the wrong part or cites the wrong clause, and the error shows up later as a support ticket or a lost customer. Putting a node-based lexical engine behind those queries trades the risk for a cluster your team owns.

    Full-text search, now generally available, closes both gaps inside the vector database you already run. The queries that embeddings get wrong now return the right records, and semantic search keeps handling everything else. One index and one schema sit behind both, with no cluster to tune and maintain.

    Create an index and start using full-text search today, or contact us if you have questions about full-text search or want help sizing an instance.

    Learn more about full-text search

    • Read the docs. The full-text search guide covers schemas, analyzers, query syntax, current limits, and more.
    • Run an example. Searching for Birds with Pinecone Full-Text Search walks through schemas, filters, and query syntax, and the Bird Search demo combines full-text search with multimodal vector search over a Wikipedia bird dataset.
    • Go deeper. Full Text Search: Architecture and Design covers how it's built, for anyone who wants the mechanism behind the API.
    Original source
  • September 2026
    • Date parsed from source:
      Sep 2, 2026
    • First seen by Releasebot:
      Sep 5, 2026
    Pinecone logo

    Pinecone

    General availability: Full-text search and the Documents API

    Pinecone releases generally available full-text search on API version 2026-07, built on the new Documents API. It adds schema-based document indexes with BM25, dense and sparse ranking, richer query options, and metadata-based document updates.

    Full-text search is now generally available on API version 2026-07, built on the new Documents API. Create schema-based document indexes that combine full-text (BM25), dense-vector, and sparse-vector ranking fields in a single index. Query them with BM25 token matching, Lucene query syntax, or vector similarity, narrowed by metadata and text-match filters. You can also fetch, update, and delete documents by metadata filter, not just by ID. This is a change at the REST layer: on 2026-07, POST /indexes is schema-only, so the top-level dimension, metric, vector_type, and spec are replaced by a schema. Existing indexes and SDK create_index calls are unaffected. See Adopt the Documents API to check whether your code is affected and how to move.

    Original source
  • September 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 5, 2026
    Pinecone logo

    Pinecone

    BYOC stores cluster metadata in FoundationDB

    Pinecone adds BYOC deployments that run FoundationDB in your Kubernetes cluster with no managed database to operate.

    BYOC deployments run FoundationDB inside your Kubernetes cluster to store cluster metadata, with no managed database to operate.

    Original source
  • September 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Pinecone logo

    Pinecone

    Bulk import credit is no longer offered

    Pinecone updates Standard and Enterprise import pricing, ending the one-time bulk import credit for new subscriptions.

    New Standard and Enterprise subscriptions no longer receive the one-time $250 bulk import credit (1 TB). Imports from object storage are billed at the standard $0.25/GB rate. Credits already granted stay usable until they expire, 60 days after they were granted. The last balances expire by November 1, 2026. For details, see Understanding cost.

    Original source
  • September 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Pinecone logo

    Pinecone

    Egress metering and monthly egress allowances

    Pinecone adds metered egress for read data, with monthly allowances by plan and billing or read-blocking behavior once limits are reached.

    Pinecone now meters egress, the data returned to you on reads.

    Each plan includes a monthly egress allowance: 1 GB on Starter, 10 GB on Builder, and 100 GB on Standard and Enterprise.

    Egress is metered on reads that return record data, such as query, fetch, list, and search.

    Writes, index statistics, and index management requests aren’t metered.

    Egress applies to indexes that use dedicated read nodes as well as on-demand indexes.

    On Standard and Enterprise, egress beyond the allowance is billed at the rate on the pricing page, and reads keep serving.

    On Starter and Builder, reads that return record data are blocked once the allowance is reached, and the allowance resets at the start of the next billing period.

    Original source
  • August 2026
    • Date parsed from source:
      Aug 25, 2026
    • First seen by Releasebot:
      Aug 26, 2026
    Pinecone logo

    Pinecone

    General availability: Bring your own cloud (BYOC)

    Pinecone adds general availability for Bring Your Own Cloud on Enterprise, keeping data and queries in your environment.

    General availability: Bring your own cloud (BYOC) is now generally available and recommended for production usage on the Enterprise plan, across AWS, GCP, and Azure.

    BYOC runs the Pinecone data plane inside your own cloud account, so your vectors, metadata, and queries stay in your environment.

    An agent in your cluster pulls operations from Pinecone and runs them locally, so Pinecone needs no inbound network access to your infrastructure.

    Original source
  • Aug 6, 2026
    • Date parsed from source:
      Aug 6, 2026
    • First seen by Releasebot:
      Aug 10, 2026
    Pinecone logo

    Pinecone

    General Availability of Pinecone Nexus Proves Knowledge Drives Real Outcomes for Agentic AI

    Pinecone launches Nexus, a generally available knowledge engine that brings enterprise data and workflows into governed, agent-ready knowledge. It promises faster, lower-cost, more accurate agents with native governance and cloud deployment.

    Pinecone Nexus makes agents more accurate, faster, lower cost, and trusted, outperforming agents that use frontier models alone on Sierra’s agentic work benchmark

    NEW YORK, August 6, 2026 / PRNewswire / — Pinecone announced today the general availability of Pinecone Nexus, a knowledge engine that transforms an enterprise's proprietary data and workflows into governed, agent-ready knowledge, delivering it to AI agents in a single call. On its debut on τ-Knowledge, Sierra's open benchmark for the most demanding enterprise knowledge tasks, an agent using Nexus as its knowledge layer posted the top score, outperforming agents built on frontier models from OpenAI, Anthropic, and Google.

    The first era of enterprise AI was built by developers for human users. Retrieval systems like RAG pipelines, vector search, and brute-force agentic search assumed a human in the loop who could read the results, catch the wrong ones, and try again. Agents are now the dominant consumer of enterprise AI. Costs are exploding as they run those same systems autonomously. More than 85% of LLM effort goes to retrieving knowledge from the underlying data, driving accuracy down and latency up on every task.

    Enterprises moving agents into production face another problem: the model is a commodity because every competitor can buy the same one. The only durable advantage is the enterprise's own knowledge and the way its people do the work. Today's agent stacks give both away, reassembling that knowledge on every call and handing it to a model vendor that turns around and competes with the enterprise.

    Pinecone Nexus solves these problems by delivering a knowledge layer purpose-built for agents. Nexus compiles an enterprise's data into governed, domain-specific knowledge that agents query through KnowQL, a declarative query language built for agents. Pre-compiling knowledge lowers token costs by more than 90% over agentic RAG, answers up to 30 times faster, and completes tasks with more than 90% accuracy.

    Nexus deploys in the customer's cloud and runs with zero access, on the models they choose, including open-weight models. Outputs are open and portable with no lock-in. Governance is native: field-level access control, per-field citations, confidence scores, PII-aware ingestion, and lineage back to source. Enterprises keep their own knowledge. They don't hand their competitive moat to model vendors.

    Nexus relies on subject matter experts to shape the knowledge so it fits the business context and workflows the enterprise runs every day. No central ontologies set once and left to decay. Keeping domain experts in the loop extends Pinecone from a developer-first tool into a platform for the line-of-business professionals now driving AI adoption: financial analysts, insurance underwriters, attorneys, account executives, and customer service representatives.

    "Enterprises adopting AI are squeezed from two sides," said Ash Ashutosh, CEO of Pinecone. "Agents burn tokens grinding through raw data, so cost and latency climb while accuracy stays lower than it should be. And every model call risks handing proprietary knowledge to a system that can turn around and compete with you. Nexus puts a knowledge engine in your own cloud, raises accuracy, lowers the total cost of running AI, and keeps your own experts shaping how agents work."

    τ-Knowledge is Sierra's open-source benchmark for agentic customer support work that demands multi-step reasoning, strict policy adherence, and coordinated tool use. The benchmark measures exactly the work for which Nexus is built because it is graded on whether the agent drives the system to the correct end state. The best frontier model on the current leaderboard, GPT-5.5, solves 46.4% of the tasks. With Nexus as the knowledge layer, an agent solved 47.4%, the top score on the benchmark. It also achieved 74% less cost per task compared to an agent using a frontier model without a Nexus knowledge layer.

    "Enterprises running agentic workloads have been hitting a real ceiling on cost, since retrieval and re-orientation can eat up the bulk of token spend before an agent ever reasons”, says Devin Pratt, Research Director at IDC. “Pinecone's approach, compiling proprietary knowledge into a reusable layer instead of re-deriving it on every call, is a sensible response to that problem. It's a promising direction, and one worth watching as more enterprises evaluate precompiled knowledge layers."

    Pinecone Nexus is generally available beginning August 6, 2026, deployed in the customer's own cloud. It sits within the broader Pinecone platform, with the Pinecone Database as its retrieval foundation and Pinecone Marketplace offering production-ready knowledge apps. Learn more at pinecone.io/nexus and pinecone.io/blog/pinecone-nexus-generally-available/.

    About Pinecone

    Pinecone is the trusted AI knowledge company. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 10,000 customers and 1M developers worldwide. Pinecone's mission is to make AI knowledgeable. For more information, visit pinecone.io.

    Media Contact

    Mike Sefanov
    [email protected]
    Sr. Director, Communications

    Original source
  • Aug 6, 2026
    • Date parsed from source:
      Aug 6, 2026
    • First seen by Releasebot:
      Aug 7, 2026
    Pinecone logo

    Pinecone

    The Ceiling Was Never the Model

    Pinecone releases Nexus generally available, bringing compiled governed knowledge for production AI agents in your own cloud. It emphasizes citations, confidence scores, access control, and lower-cost, more reliable answers backed by benchmark and enterprise results.

    Sierra AI built an internal AI agent and named it Pinecone. We're flattered. So we took the benchmark Sierra built and beat it. Name confusion aside, Sierra and Pinecone landed on the same conclusion about production AI: it's the knowledge, not the model.

    Let's clear something up first. Sierra AI, Bret Taylor's AI company, recently announced they had “AI-pilled” the whole company with an internal agent they built and named Pinecone. Great name!

    Here is the part where the joke turns serious. Today, on τ-Knowledge, the benchmark Sierra itself built, an agent using Pinecone Nexus posted the top score, ahead of agents running on frontier models alone. Two companies reached the same conclusion independently: in enterprise production, it is the knowledge, not the model, that decides whether AI works.

    We did not arrive here overnight. We have been building Nexus for about a year. More than 800 organizations signed up for early access, and since Early Access opened in May we have worked closely with over a hundred enterprises, many of them among the largest in the world, running their Data through it. Today Nexus is generally available. This is what we learned, and why the thesis held.

    Start with what those enterprises told us, because it is probably familiar. They moved agents out of the demo and into production, across finance, insurance, legal, retail, and support. The agents stalled. Not because the model was not smart enough. Because of everything the agent had to do before it could be smart.

    Picture a support agent handling a billing dispute. Before it can answer, it reads the ticket, searches the knowledge base, reads what came back, searches again, and re-sends everything it has gathered on the next turn. Most of its time and most of its token budget is spent before it decides anything. Then the retrieval itself betrays it. A vector search returns the top matching chunks of text, stripped of the relationships that connect them, and the agent confidently quotes a refund policy that was revised eighteen months ago. A fluent answer grounded in the wrong version of the policy scores zero, and in production it reaches a real customer.

    Now multiply that by every ticket, every day. The bill does not scale with the price of a token. It scales with retrieval. Blended inference prices fell about 67% year over year, and enterprise AI budgets kept climbing anyway, because one task fans out into dozens of model calls, each re-reading the same documents. More than 85% of an agent's effort goes to fetching knowledge before it reasons. Goldman Sachs projects token consumption to multiply 24x by 2030. In a survey of 306 teams running agents in production, reliability, not model capability, was the top challenge, and 68% cap their agents at ten steps before a human has to step in. A bigger model does not fix a bill that scales with retrieval, or a completion rate capped by what the agent can reach.

    There is a second cost. The model is a commodity, because every competitor can buy the same one. The only durable advantage an enterprise has is its own knowledge and the way its people work. Most agent stacks reassemble that knowledge on every call and hand it to a model vendor. Satya Nadella has made this the center of Microsoft's argument: models are becoming interchangeable, and the moat that lasts is the enterprise's own data, context, and memory, kept under its own control. Alex Karp of Palantir has put it more sharply, warning that frontier labs have oversold their models while quietly absorbing the proprietary edge of the companies paying for them, so enterprises end up paying to lose their advantage. Different companies, same warning from two leaders serving world’s largest enterprises. The moat leaks out one API call at a time.

    That is the problem Nexus was built to remove, and the fix is not a better model. It is moving the knowledge work out of the per-query loop.

    Nexus compiles your data once, ahead of time, into governed, domain-specific knowledge, and agents reuse that compiled layer on every call. The person who understands the work describes it in their own terms: the entities that matter, how they relate, and the shape of the answers the work needs. Nexus turns raw sources into structured knowledge, including the relationships that ordinary retrieval throws away, and resolves conflicts between sources up front, so the layer knows what it knows and flags what is contested. Agents then ask through KnowQL, a query language built for agents. The agent states what it needs and gets back a typed, cited answer in a single call. Compile once, answer every time.

    Go back to the support agent. With a compiled layer it stops grinding through documents and asks for the answer instead. It gets the current policy, with a citation, and because it now has customer-level context it knows the one question only the customer can answer, and when to ask it. We know this because we pointed Nexus at our own support queue on July 17. The share of tickets the agent resolved on its own went from 24.6% to 55.1%. More than half of our tickets now close without a person touching them.

    The same shift shows up on Sierra's benchmark. τ-Knowledge grades an agent on whether it drives the system to the correct end state, and its hardest domains make the agent find and apply the right policy before it acts. On the banking domain, ninety-seven of those tasks, we gave the same frontier models a Nexus layer to query and watched their behavior change.

    The model calls roughly halved, and each one carried less context. GPT-5.2 gained 12% accuracy at 80% lower cost. GPT-5.5 held its accuracy at 77% lower cost. In dollars, that is a task that cost $1.45 falling to $0.53, and the advantage held on 96 or 97 of the 97 tasks, so it is not an average hiding a wide spread. Across the full benchmark, an agent with Nexus posted the top score, 47.4% against the best frontier model's 46.4%, at 74% less cost per task.

    Task completion is the number benchmarks chase. Cost is the number that decides whether an enterprise AI program returns anything. When the same model delivers the same accuracy at a third of the prices, AI finally delivers on the promised enterprise ROI business case. It also drops low enough that, in many cases, smaller and open-weight models clear the bar, which compounds the saving. The AI strategy initiative now becomes the Production AI initiative.

    None of that matters in a regulated industry if the answer cannot be defended, which brings us back to Nadella and Karp. Both are describing a crisis of trust: keep control of your knowledge, and be able to prove where a decision came from. Nexus is built that way. It runs in your own cloud, on the models you choose, with no standing Pinecone access to your data, so your knowledge never leaves your infrastructure. Every field it returns carries a citation and a confidence score. Every answer traces back to the source document and clause it came from. Access control is applied when knowledge is retrieved, not requested in a prompt. That is what lets an agent's answer survive a security review and an auditor, and it is what turns moat preservation from a keynote line into something an enterprise can operate.

    One more thing the hundred-plus enterprises taught us. They validated accuracy fast, usually in the first week. The rest of the work was keeping the layer true as the world moved: new tickets daily, contracts amended, a process doc revised on a Tuesday, and the wiki that disagrees with the contract. Across those engagements they compiled 3.5 million source chunks into nearly 26,000 structured knowledge artifacts, drawn from support tickets, contracts, filings, research papers, and call transcripts. Nexus curates incrementally, so only what changes gets recompiled, and the person who owns the domain keeps control of how the knowledge is shaped. The layer worth building is the one still true in month twelve, not just the one that demos well in week one.

    You do not wait for a better model to build a reliable agent. You give the model better knowledge.

    Pinecone Nexus is generally available today, in your own cloud. Learn more at pinecone.io/nexus.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.