Together AI Release Notes
88 release notes curated from 1 source by the Releasebot Team. Last updated: Jul 30, 2026
- Jul 29, 2026
- Date parsed from source:Jul 29, 2026
- First seen by Releasebot:Jul 30, 2026
July 29, 2026
Together AI deprecates a broad set of fine-tuning models, including Qwen, DeepSeek, Llama, Gemma, Mistral, and Kimi variants, and points users to migration options.
Deprecations
The following models have been deprecated and are no longer available for fine-tuning:
- nvidia/NVIDIA-Nemotron-Nano-9B-v2.
- Qwen/Qwen3-Next-80B-A3B-Instruct.
- Qwen/Qwen3-Next-80B-A3B-Thinking.
- Qwen/Qwen3-0.6B.
- Qwen/Qwen3-0.6B-Base.
- Qwen/Qwen3-1.7B.
- Qwen/Qwen3-1.7B-Base.
- Qwen/Qwen3-4B.
- Qwen/Qwen3-4B-Base.
- Qwen/Qwen3-8B.
- Qwen/Qwen3-8B-Base.
- Qwen/Qwen3-14B.
- Qwen/Qwen3-14B-Base.
- Qwen/Qwen3-32B.
- Qwen/Qwen3-30B-A3B-Base.
- Qwen/Qwen3-30B-A3B.
- Qwen/Qwen3-30B-A3B-Instruct-2507.
- Qwen/Qwen3-235B-A22B.
- Qwen/Qwen3-235B-A22B-Instruct-2507.
- Qwen/Qwen3-Coder-30B-A3B-Instruct.
- Qwen/Qwen3-Coder-480B-A35B-Instruct.
- Qwen/Qwen3-VL-8B-Instruct.
- Qwen/Qwen3-VL-32B-Instruct.
- Qwen/Qwen3-VL-30B-A3B-Instruct.
- Qwen/Qwen3-VL-235B-A22B-Instruct.
- Qwen/Qwen2.5-72B-Instruct.
- Qwen/Qwen2.5-72B.
- Qwen/Qwen2.5-32B-Instruct.
- Qwen/Qwen2.5-32B.
- Qwen/Qwen2.5-14B-Instruct.
- Qwen/Qwen2.5-14B.
- Qwen/Qwen2.5-7B-Instruct.
- Qwen/Qwen2.5-7B.
- Qwen/Qwen2.5-3B-Instruct.
- Qwen/Qwen2.5-3B.
- Qwen/Qwen2.5-1.5B-Instruct.
- Qwen/Qwen2.5-1.5B.
- Qwen/Qwen2-72B-Instruct.
- Qwen/Qwen2-72B.
- Qwen/Qwen2-7B-Instruct.
- Qwen/Qwen2-7B.
- Qwen/Qwen2-1.5B-Instruct.
- Qwen/Qwen2-1.5B.
- moonshotai/Kimi-K2.5.
- moonshotai/Kimi-K2-Thinking.
- moonshotai/Kimi-K2-Instruct-0905.
- moonshotai/Kimi-K2-Instruct.
- moonshotai/Kimi-K2-Base.
- zai-org/GLM-5.
- zai-org/GLM-4.7.
- zai-org/GLM-4.6.
- deepseek-ai/DeepSeek-R1-0528.
- deepseek-ai/DeepSeek-R1.
- deepseek-ai/DeepSeek-V3-0324.
- deepseek-ai/DeepSeek-V3.
- deepseek-ai/DeepSeek-V3.1-Base.
- deepseek-ai/DeepSeek-V3-Base.
- deepseek-ai/DeepSeek-R1-Distill-Llama-70B.
- deepseek-ai/DeepSeek-R1-Distill-Llama-70B-32k.
- deepseek-ai/DeepSeek-R1-Distill-Llama-70B-131k.
- deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.
- deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B.
- meta-llama/Llama-4-Scout-17B-16E.
- meta-llama/Llama-4-Maverick-17B-128E.
- meta-llama/Llama-3.3-70B-32k-Instruct-Reference.
- meta-llama/Llama-3.3-70B-131k-Instruct-Reference.
- meta-llama/Llama-3.2-3B-Instruct.
- meta-llama/Llama-3.2-3B.
- meta-llama/Llama-3.2-1B-Instruct.
- meta-llama/Llama-3.2-1B.
- meta-llama/Meta-Llama-3.1-8B-131k-Instruct-Reference.
- meta-llama/Meta-Llama-3.1-8B-Reference.
- meta-llama/Meta-Llama-3.1-8B-131k-Reference.
- meta-llama/Meta-Llama-3.1-70B-Instruct-Reference.
- meta-llama/Meta-Llama-3.1-70B-32k-Instruct-Reference.
- meta-llama/Meta-Llama-3.1-70B-131k-Instruct-Reference.
- meta-llama/Meta-Llama-3.1-70B-Reference.
- meta-llama/Meta-Llama-3.1-70B-32k-Reference.
- meta-llama/Meta-Llama-3.1-70B-131k-Reference.
- meta-llama/Meta-Llama-3-8B-Instruct.
- google/gemma-3-270m.
- google/gemma-3-270m-it.
- google/gemma-3-1b-it.
- google/gemma-3-1b-pt.
- google/gemma-3-4b-it.
- google/gemma-3-4b-it-VLM.
- google/gemma-3-4b-pt.
- google/gemma-3-12b-it.
- google/gemma-3-12b-it-VLM.
- google/gemma-3-12b-pt.
- google/gemma-3-27b-it.
- google/gemma-3-27b-it-VLM.
- google/gemma-3-27b-pt.
- mistralai/Mixtral-8x7B-v0.1.
- mistralai/Mistral-7B-Instruct-v0.2.
- mistralai/Mistral-7B-v0.1.
- togethercomputer/llama-2-7b-chat.
See Deprecations for migration options.
Original source - Jul 28, 2026
- Date parsed from source:Jul 28, 2026
- First seen by Releasebot:Jul 30, 2026
July 28, 2026
Together AI adds CLI improvements for A/B traffic updates, smarter replica bound handling for deployments, and daily upgrade notices with in-place update prompts. It also deprecates the MiniMaxAI/MiniMax-M2.7 model on serverless.
Improvements
Deprecations
A/B variant percent updates in the CLI
tg beta endpoints update now accepts
--ab-percentto change a variant’s traffic percentage in an existing A/B experiment. The flag takes percentage from or returns it to the control only; other variants stay unchanged. The control must remain at least 1%, and--percentontg beta endpoints abis limited to 1–99.
See Ramp the variant and the endpoints CLI reference.Deploy replica bound inference
tg beta endpoints deploy now infers a missing replica bound:
--min-replicasalone mirrors into the max (including 0 to create a deployment stopped), and--max-replicas 0alone lowers the min to 0. Ontg beta endpoints update, stopping a deployment still requires both--min-replicas 0and--max-replicas 0. Passing a single zero bound is an error.
See the endpoints CLI reference.CLI upgrade notices
The Together CLI now detects when a newer release is available and prints an upgrade notice at most once per day. Interactive sessions offer to run the upgrade in place, using the command that matches your install (uv, pipx, or pip). Set
TOGETHER_DISABLE_VERSION_CHECK=1to turn the check off.
See Get started.Model deprecations
The following models have been deprecated and are no longer available on serverless:
Original source
MiniMaxAI/MiniMax-M2.7.
See Deprecations for migration options. All of your release notes in one feed
Join Releasebot and get updates from Together AI and hundreds of other software products.
- Jul 28, 2026
- Date parsed from source:Jul 28, 2026
- First seen by Releasebot:Jul 28, 2026
- Modified by Releasebot:Jul 29, 2026
July 28, 2026
Together AI deprecates MiniMaxAI/MiniMax-M2.7 on serverless and points users to migration options.
Improvements
Deprecations
The following models have been deprecated and are no longer available on serverless:
MiniMaxAI/MiniMax-M2.7.
See Deprecations for migration options.
Original source - Jul 27, 2026
- Date parsed from source:Jul 27, 2026
- First seen by Releasebot:Jul 28, 2026
July 27, 2026
Together AI adds fine-tuning support for Qwen/Qwen3.6-27B.
Improvements
New models available for fine-tuning
You can now fine-tune the following models:
- Qwen/Qwen3.6-27B.
See Supported models for the full list.
Original source - Jul 24, 2026
- Date parsed from source:Jul 24, 2026
- First seen by Releasebot:Jul 25, 2026
- Modified by Releasebot:Jul 28, 2026
July 24, 2026
Together AI adds realtime transcription to the Python SDK, bringing streaming speech-to-text over WebSocket with reconnects, audio replay, normalized transcript events, and endpoint failover. It also expands fine-tuning comparison metrics filtering in the dashboard.
New releases
Improvements
Python SDK realtime transcription
The Together Python SDK now includes client.beta.realtime.transcription() for streaming speech-to-text over WebSocket. Install with pip install "together[realtime]". The session reconnects with audio replay on transient drops, exposes normalized events such as TranscriptDelta and TranscriptCompleted, and supports application-level failover across endpoints with RealtimeConnectionError (code="no_healthy_workers") and session.pending_audio().
See Streaming transcription.
Fine-tuning comparison metrics filtering
The fine-tuning comparison view now includes the same Metrics filtering control as the single-job Metrics tab. Adjust Sampling rate and Step range, then click Apply to re-fetch metrics for every selected job with matching filters.
See View metrics in the dashboard.
Original source Similar to Together AI with recent updates:
- xAI release notes200 release notes · Latest Jul 31, 2026
- Anthropic release notes732 release notes · Latest Jul 28, 2026
- OpenAI release notes887 release notes · Latest Jul 31, 2026
- Google release notes1766 release notes · Latest Jul 31, 2026
- Notion release notes157 release notes · Latest Jul 31, 2026
- Ubiquiti release notes805 release notes · Latest Jul 31, 2026
- Jul 23, 2026
- Date parsed from source:Jul 23, 2026
- First seen by Releasebot:Jul 24, 2026
- Modified by Releasebot:Jul 28, 2026
July 23, 2026
Together AI adds dedicated container HTTP server mode for OpenAI-compatible endpoints, richer fine-tuning artifact naming in APIs and dashboard, new CLI output for checkpoint registry IDs and training previews, and Prometheus gauges for worker cache utilization and hit rate.
Improvements
Dedicated containers OpenAI-compatible endpoints
Dedicated container inference now supports HTTP server mode for synchronous, OpenAI-compatible endpoints. Run your worker without the --queue flag, wire an OpenAI route with Sprocket, and call it with the OpenAI SDK or plain HTTP, with no Together-specific request shapes.
See Serve an OpenAI-compatible endpoint.
Fine-tuning output object names
Fine-tune retrieve and list-checkpoints responses now include qualified Together model registry names alongside object IDs. On GET /fine-tunes/{id}, use model_object_name and adapter_object_name (LoRA jobs) for the final artifacts in / form. On GET /fine-tunes/{id}/checkpoints, each entry adds object_name with the same naming pattern (including - or -adapter suffixes). Names are resolved on retrieve only, not on list jobs. If the project slug cannot be resolved, the name field falls back to the object ID.
On a completed job in the fine-tuning jobs dashboard, Output model shows the same qualified model_object_name and links to the registry model page.
See Model registry object IDs.
Fine-tuning checkpoint CLI output
tg fine-tuning list-checkpoints now displays registry artifact IDs in the table output. The Registry Artifact column shows object_id@object_revision_id, and a copyable Registry artifacts block prints below the table.
See List checkpoints.
Fine-tuning preview CLI command
tg fine-tuning preview samples rows from an uploaded JSONL training file and shows how a base model tokenizes them before you start a job. The table output highlights masked tokens, trained spans, and truncation; pass --json for the full response.
See Preview.
Worker engine cache metrics
Dedicated endpoint Prometheus metrics now expose worker_engine_kv_cache_utilization and worker_engine_cache_hit_rate gauges. Use them to track KV-cache pressure and prefix-reuse efficiency on deployment workers.
See Monitor endpoints and deployments.
Original source - Jul 22, 2026
- Date parsed from source:Jul 22, 2026
- First seen by Releasebot:Jul 28, 2026
July 22, 2026
Together AI adds a --mode override for beta cluster node remediation approvals, letting users choose how nodes are repaired.
Improvements
Cluster remediation mode override
tg beta clusters remediations approve now accepts a --mode flag to override the recommended repair action when you approve a node remediation. Choose reboot, quick reprovision, migrate to new host, or remove without changing the recommendation in the console first.
See Node repair and the clusters CLI reference.
Original source - Jul 21, 2026
- Date parsed from source:Jul 21, 2026
- First seen by Releasebot:Jul 22, 2026
- Modified by Releasebot:Jul 28, 2026
July 21, 2026
Together AI adds GPU cluster add-on CLI flags, new MoE expert-LoRA targeting for fine-tuning, stricter jig environment variable collision validation, and support for fine-tuning new GLM models.
Improvements
GPU cluster add-on CLI flags
tg beta clusters create and tg beta clusters update now expose flags for Headlamp and Slurm Web cluster add-ons.
Create: Pass --headlamp-addon or --slurm-web-addon to enable an add-on at cluster creation.
Update: Pass --headlamp / --no-headlamp or --slurm-web / --no-slurm-web to toggle add-ons on an existing cluster.
See the clusters CLI reference.
MoE expert-LoRA target modules
Fine-tuning jobs on supported mixture-of-experts models can now target expert feed-forward layers instead of attention projections. Set lora_trainable_modules to w_up,w_gate,w_down to train a compact shared-factor adapter across experts, useful when adapting domain knowledge rather than attention patterns.
See Target MoE expert layers.
Jig environment variable collision validation
jig deploy now rejects configs where the same name appears in both [tool.jig.deploy.environment_variables] and secrets. The error lists each colliding name so you can remove the duplicate or unset the secret before redeploying.
See Secrets.
New models available for fine-tuning
You can now fine-tune the following models:
- zai-org/GLM-5.1.
- zai-org/GLM-5.
See Supported models for the full list.
Original source - Jul 20, 2026
- Date parsed from source:Jul 20, 2026
- First seen by Releasebot:Jul 22, 2026
- Modified by Releasebot:Jul 28, 2026
July 20, 2026
Together AI adds a fine-tuning Metrics dashboard with live training curves, preference-tuning metrics, interactive chart controls, and side-by-side run comparison. It also expands dedicated endpoint model support with deepseek-ai/DeepSeek-V4-Flash.
New releases
New models
Improvements
Fine-tuning metrics dashboard
Fine-tuning jobs now have a Metrics tab in the web dashboard that charts training progress. Metrics stream in live while a job runs and stay available after it completes.
What’s new:
- Live training curves: Watch training loss, learning rate, and gradient norm update as the job trains, without leaving the dashboard.
- Preference-tuning metrics: DPO and other preference jobs add reward accuracy, reward margin, chosen and rejected rewards and log-probabilities, and approximate KL.
- Interactive charts: Zoom and pan across all charts on a shared step axis, switch the x-axis between step and time, toggle linear or logarithmic scales, and sync a hover crosshair across every metric.
- Compare runs: Overlay metrics from multiple jobs to compare fine-tuning runs side by side.
New dedicated endpoint models
The following models are now available for deployment on dedicated endpoints:
- deepseek-ai/DeepSeek-V4-Flash.
- Jul 20, 2026
- Date parsed from source:Jul 20, 2026
- First seen by Releasebot:Jul 17, 2026
- Modified by Releasebot:Jul 21, 2026
July 20, 2026
Together AI adds new serverless model support for thinkingmachines/Inkling and expands dedicated model inference with one-command deployment, traffic splitting, autoscaling, and shadow experiments. It also brings image inputs to evaluations for vision-capable models.
New releases
New models
New serverless models
The following models are now available on serverless:
thinkingmachines/Inkling: 552.8B parameters, 524,288 context length, NVFP4 quantization. Pricing: $1.00 input / $4.05 output / $0.17 cached input (per 1M tokens).Dedicated model inference
Dedicated model inference (DMI) serves a model on reserved GPUs, giving you higher throughput, lower latency, and predictable performance with no hard rate limits. It uses the same inference APIs as serverless, so you can prototype on serverless and deploy on dedicated hardware without changing your application code.
What’s new:
- Deploy in one command: tg beta endpoints deploy creates an endpoint, attaches a deployment, and routes all traffic to it.
- New resource model: Compose models, configs, endpoints, deployments, and traffic splits to control how your model is served. See Concepts.
- A/B tests and shadow experiments: Split live traffic across variants, or mirror traffic to a new deployment without serving its responses.
- Autoscaling and serving controls: Scale on request or GPU metrics, scale to zero on idle, and tune serving behavior per deployment. See Configure autoscaling.
Note
If you’re already using dedicated endpoints, you can migrate to the v2 API by following the migration guide.
New serverless models
The following models are now available on serverless:
thinkingmachines/Inkling: 552.8B parameters, 524,288 context length, NVFP4 quantization. Pricing: $1.00 input / $4.05 output / $0.17 cached input (per 1M tokens).Image inputs for evaluations
You can now evaluate vision-capable models on image datasets. Add an image_data_urls column to your evaluation dataset — a base64 image data URL, or a list of them — and the images are attached to the model and judge requests alongside the text prompt.
See Prepare your dataset for details.
Original source - Jul 17, 2026
- Date parsed from source:Jul 17, 2026
- First seen by Releasebot:Jul 21, 2026
July 17, 2026
Together AI improves privacy settings and endpoint deployment listings for organization admins and API users.
Improvements
Privacy and training opt-in settings
Training opt-in and other privacy toggles now live under Privacy in Organization Settings. Organization admins control data-sharing settings for all traffic sent under the organization’s API keys.
See Privacy and security.
Dedicated endpoint deployment listing
Endpoint get and list responses now include at most the 10 newest deployment summaries per endpoint. Use the deployments list API to retrieve every deployment on an endpoint.
See Manage endpoints and deployments.
Original source - Jul 16, 2026
- Date parsed from source:Jul 16, 2026
- First seen by Releasebot:Jul 22, 2026
July 16, 2026
Together AI introduces dedicated model inference with reserved GPUs, one-command deployment, traffic splitting, autoscaling, and a shared API path from serverless to dedicated hardware. It also adds a new serverless model, image inputs for evaluations, and richer fine-tuning registry IDs.
New releases
New models
Deprecations
Improvements
Dedicated model inference
Dedicated model inference (DMI) serves a model on reserved GPUs, giving you higher throughput, lower latency, and predictable performance with no hard rate limits. It uses the same inference APIs as serverless, so you can prototype on serverless and deploy on dedicated hardware without changing your application code.
What’s new:
- Deploy in one command: tg beta endpoints deploy creates an endpoint, attaches a deployment, and routes all traffic to it.
- New resource model: Compose models, configs, endpoints, deployments, and traffic splits to control how your model is served. See Concepts.
- A/B tests and shadow experiments: Split live traffic across variants, or mirror traffic to a new deployment without serving its responses.
- Autoscaling and serving controls: Scale on request or GPU metrics, scale to zero on idle, and tune serving behavior per deployment. See Configure autoscaling.
Note
If you’re already using dedicated endpoints, you can migrate to the v2 API by following the migration guide.
New serverless models
The following models are now available on serverless:
thinkingmachines/Inkling: 552.8B parameters, 524,288 context length, NVFP4 quantization. Pricing: $1.00 input / $4.05 output / $0.17 cached input (per 1M tokens).Image inputs for evaluations
You can now evaluate vision-capable models on image datasets. Add an image_data_urls column to your evaluation dataset — a base64 image data URL, or a list of them — and the images are attached to the model and judge requests alongside the text prompt.
See Prepare your dataset for details.
Fine-tuning artifact registry IDs
Fine-tune job and checkpoint responses now include Together model registry IDs for completed artifacts. Use model_object_id and model_object_revision_id on completed jobs, adapter_object_id and adapter_object_revision_id on LoRA jobs, and object_id and object_revision_id on list checkpoints entries to reference weights in the model registry. The Together Python SDK (2.24+) and Together TypeScript SDK expose these fields on retrieve and list-checkpoints responses.
Together Python SDK 2.24
Together Python SDK 2.24 is available. OIDC cluster SSH (tg beta clusters ssh) no longer requires a Together API key, and the bastion-to-target SSH hop skips host key prompts for ephemeral cluster nodes. See SSH into a cluster.
Original source - Jul 15, 2026
- Date parsed from source:Jul 15, 2026
- First seen by Releasebot:Jul 21, 2026
July 15, 2026
Together AI adds TogetherLink for connecting popular coding tools to models hosted by Together AI, browser-based SSH for OIDC-enabled Slurm GPU clusters, and a model limits API update that shows whether a model supports full fine-tuning.
New releases
Improvements
TogetherLink
TogetherLink is an open-source launcher that connects Claude Code, Codex, ChatGPT Desktop, Pi Code, and OpenCode to models hosted by Together AI. Install with one command and run your existing tools without hand-editing provider settings.
See Configure Claude Code, Codex, and ChatGPT with Together AI models.
Slurm cluster OIDC SSH
Slurm GPU clusters with OIDC enabled now support browser-based SSH through the Together CLI. Choose OIDC on the cluster details page to sign in without uploading an SSH key. Key-based SSH remains available.
See Cluster management.
Fine-tuning model limits API
The model limits endpoint now returns supports_full_training so you can check whether a model supports full fine-tuning before submitting a job. When the field is false, the model is LoRA-only.
See LoRA vs. full fine-tuning.
Original source - Jul 15, 2026
- Date parsed from source:Jul 15, 2026
- First seen by Releasebot:Jul 16, 2026
July 15, 2026
Together AI adds image inputs for evaluations, letting vision-capable models and judges use image data URLs with prompts.
Image inputs for evaluations
You can now evaluate vision-capable models on image datasets. Add an
image_data_urlscolumn to your evaluation dataset — a base64 image data URL, or a list of them — and the images are attached to the model and judge requests alongside the text prompt.See Prepare your dataset for details.
Original source - Jul 14, 2026
- Date parsed from source:Jul 14, 2026
- First seen by Releasebot:Jul 21, 2026
July 14, 2026
Together AI adds new serverless models and improves GPU cluster health monitoring with clearer alerts and repair guidance.
New models
Improvements
New serverless models
The following models are now available on serverless:
- google/gemma-4-12B-it: 262,144 context length.
GPU cluster health monitoring
Passive health checks now document the full set of monitored failure signals and their recommended repair actions. Slurm node unavailability surfaces as a warning-only alert without an automated repair action.
See Health checks and Node repair.
Original source
Curated by the Releasebot team
Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.
Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.