Together AI Release Notes

Follow

98 release notes curated from 1 source by the Releasebot Team. Last updated: Aug 14, 2026

Get this feed:
  • Aug 13, 2026
    • Date parsed from source:
      Aug 13, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Together AI logo

    Together AI

    August 13, 2026

    Together AI adds new fine-tuning support for zai-org/GLM-5.2 and expands the CLI with endpoint event monitoring plus fine-tuning tools for model limits and tokenized dataset downloads.

    New models available for fine-tuning

    You can now fine-tune the following models:

    zai-org/GLM-5.2.

    See Supported models for the full list.

    Endpoint events in the CLI

    tg beta endpoints events lists a dedicated endpoint’s audit and lifecycle events from the terminal: replica scaling, traffic shifts, status changes, and pauses across every deployment under the endpoint.

    See Monitoring endpoint events and the tg beta endpoints events CLI command.

    Fine-tuning limits and tokenized datasets in the CLI

    Two new fine-tuning commands are available:

    • tg ft model-limits <model> prints a model’s fine-tuning constraints, including sequence-length, batch-size, and LoRA rank limits.
    • tg ft download-tokenized-dataset <ft_id> downloads the tokenized dataset a job trained on, so you can audit exactly what the model saw.

    See the fine-tuning CLI reference.

    Original source
  • Aug 12, 2026
    • Date parsed from source:
      Aug 12, 2026
    • First seen by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    August 12, 2026

    Together AI adds new serverless models with Qwen/Qwen3.8-2.4T-A95B FP4 quantization.

    New serverless models

    The following models are now available on serverless:

    • Qwen/Qwen3.8-2.4T-A95B: FP4 quantization. Pricing: $2.50 input / $6.25 output / $0.50 cached input (per 1M tokens).
    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Together AI and hundreds of other software products.

    Create account
  • Aug 11, 2026
    • Date parsed from source:
      Aug 11, 2026
    • First seen by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    August 11, 2026

    Together AI adds new serverless models and fine-tuning support for Meta Muse-Glimmer-30B and DeepSeek-V4-Flash-0731.

    New serverless models

    The following models are now available on serverless:

    meta-models/Muse-Glimmer-30B: 131,072 context length, FP8 quantization. Pricing: $0.35 input / $1.50 output / $0.04 cached input (per 1M tokens).

    New models available for fine-tuning

    You can now fine-tune the following models:

    deepseek-ai/DeepSeek-V4-Flash-0731.

    See Supported models for the full list.

    Original source
  • Aug 8, 2026
    • Date parsed from source:
      Aug 8, 2026
    • First seen by Releasebot:
      Aug 9, 2026
    • Modified by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    August 8, 2026

    Together AI expands GLM-5.2 serverless to a 512,000-token context length with pricing unchanged.

    Improvements

    Longer context for GLM-5.2

    zai-org/GLM-5.2 on serverless now accepts a 512,000-token context length, up from 262,144. Pricing is unchanged.

    See the GLM-5.2 quickstart.

    Original source
  • Aug 7, 2026
    • Date parsed from source:
      Aug 7, 2026
    • First seen by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    August 7, 2026

    Together AI adds new serverless models and lets DeepSeek-V3.1 LoRA fine-tuning target MoE expert layers.

    New models

    Improvements

    New serverless models

    The following models are now available on serverless:

    • Prism-ML/Ternary-Bonsai-27B: 262,144 context length. Pricing: Free.
    • prunaai/p-image-ideogram: Pricing: from $0.00225 per image.
    • black-forest-labs/FLUX-3: Pricing: $0.17/sec at 720p.
    Expert LoRA for DeepSeek-V3.1

    LoRA fine-tuning jobs on deepseek-ai/DeepSeek-V3.1 can now target the MoE expert layers.

    See Target MoE expert layers for how to enable it.

    Original source
  • Similar to Together AI with recent updates:

  • Aug 5, 2026
    • Date parsed from source:
      Aug 5, 2026
    • First seen by Releasebot:
      Aug 7, 2026
    • Modified by Releasebot:
      Aug 9, 2026
    Together AI logo

    Together AI

    August 5, 2026

    Together AI adds expert-layer LoRA fine-tuning for DeepSeek-V3.1, letting jobs target MoE expert layers.

    Expert LoRA for DeepSeek-V3.1

    LoRA fine-tuning jobs on deepseek-ai/DeepSeek-V3.1 can now target the MoE expert layers.

    See Target MoE expert layers for how to enable it.

    Original source
  • Aug 3, 2026
    • Date parsed from source:
      Aug 3, 2026
    • First seen by Releasebot:
      Aug 5, 2026
    • Modified by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    August 3, 2026

    Together AI adds new serverless DeepSeek-V4-Flash-0731 and enables fine-tuning for DeepSeek-V4-Flash.

    New models

    New serverless models

    The following models are now available on serverless:

    deepseek-ai/DeepSeek-V4-Flash-0731: 1,000,000 context length, FP4 quantization. Pricing: $0.14 input / $0.28 output / $0.03 cached input (per 1M tokens).

    New models available for fine-tuning

    You can now fine-tune the following models:

    deepseek-ai/DeepSeek-V4-Flash.

    See Supported models for the full list.

    Original source
  • Aug 3, 2026
    • Date parsed from source:
      Aug 3, 2026
    • First seen by Releasebot:
      Aug 4, 2026
    Together AI logo

    Together AI

    August 3, 2026

    Together AI adds new models

    New models

    Original source
  • Jul 31, 2026
    • Date parsed from source:
      Jul 31, 2026
    • First seen by Releasebot:
      Aug 4, 2026
    • Modified by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    July 31, 2026

    Together AI adds new serverless models and live model upload progress tracking in the console.

    New models

    Improvements

    New serverless models

    The following models are now available on serverless:

    thinkingmachines/Inkling-Small: 524,288 context length. Pricing: $0.50 input / $1.20 output (per 1M tokens).

    Model upload progress in the console

    While a remote model upload is pending or running, the Models page shows an Uploading badge on the model under My models and floats it to the top of the list. Opening the model shows a live Upload progress event log until the job finishes.

    See Check upload status.

    Original source
  • Jul 29, 2026
    • Date parsed from source:
      Jul 29, 2026
    • First seen by Releasebot:
      Aug 4, 2026
    • Modified by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    July 29, 2026

    Together AI improves Models page visibility filtering and expands project-scoped console support for fine-tuning, files, and evaluations.

    Improvements

    Models page visibility filter

    The Models page now lists Internal-visibility models from every project in your organization under My models, not only from the selected project. A Visibility filter lets you show Internal models, Private models, or both.

    See Upload a fine-tuned model.

    Project scoping in the console

    Fine-tuning, Files, and Evaluations are now available in the Projects UI. Create and manage fine-tuning jobs, uploaded files, and evaluations within a project from the console, not just with project-scoped API keys.

    Original source
  • Jul 29, 2026
    • Date parsed from source:
      Jul 29, 2026
    • First seen by Releasebot:
      Jul 30, 2026
    • Modified by Releasebot:
      Aug 13, 2026
    Together AI logo

    Together AI

    July 29, 2026

    Together AI deprecates a broad set of models for fine-tuning, including many Qwen, DeepSeek, Llama, Kimi, GLM, and other popular options, tightening its supported model lineup.

    Deprecations

    Model deprecations

    The following models have been deprecated and are no longer available for fine-tuning:

    • nvidia/NVIDIA-Nemotron-Nano-9B-v2.
    • Qwen/Qwen3-Next-80B-A3B-Instruct.
    • Qwen/Qwen3-Next-80B-A3B-Thinking.
    • Qwen/Qwen3-0.6B.
    • Qwen/Qwen3-0.6B-Base.
    • Qwen/Qwen3-1.7B.
    • Qwen/Qwen3-1.7B-Base.
    • Qwen/Qwen3-4B.
    • Qwen/Qwen3-4B-Base.
    • Qwen/Qwen3-8B.
    • Qwen/Qwen3-8B-Base.
    • Qwen/Qwen3-14B.
    • Qwen/Qwen3-14B-Base.
    • Qwen/Qwen3-32B.
    • Qwen/Qwen3-30B-A3B-Base.
    • Qwen/Qwen3-30B-A3B.
    • Qwen/Qwen3-30B-A3B-Instruct-2507.
    • Qwen/Qwen3-235B-A22B.
    • Qwen/Qwen3-235B-A22B-Instruct-2507.
    • Qwen/Qwen3-Coder-30B-A3B-Instruct.
    • Qwen/Qwen3-Coder-480B-A35B-Instruct.
    • Qwen/Qwen3-VL-8B-Instruct.
    • Qwen/Qwen3-VL-32B-Instruct.
    • Qwen/Qwen3-VL-30B-A3B-Instruct.
    • Qwen/Qwen3-VL-235B-A22B-Instruct.
    • Qwen/Qwen2.5-72B-Instruct.
    • Qwen/Qwen2.5-72B.
    • Qwen/Qwen2.5-32B-Instruct.
    • Qwen/Qwen2.5-32B.
    • Qwen/Qwen2.5-14B-Instruct.
    • Qwen/Qwen2.5-14B.
    • Qwen/Qwen2.5-7B-Instruct.
    • Qwen/Qwen2.5-7B.
    • Qwen/Qwen2.5-3B-Instruct.
    • Qwen/Qwen2.5-3B.
    • Qwen/Qwen2.5-1.5B-Instruct.
    • Qwen/Qwen2.5-1.5B.
    • Qwen/Qwen2-72B-Instruct.
    • Qwen/Qwen2-72B.
    • Qwen/Qwen2-7B-Instruct.
    • Qwen/Qwen2-7B.
    • Qwen/Qwen2-1.5B-Instruct.
    • Qwen/Qwen2-1.5B.
    • moonshotai/Kimi-K2.5.
    • moonshotai/Kimi-K2-Thinking.
    • moonshotai/Kimi-K2-Instruct-0905.
    • moonshotai/Kimi-K2-Instruct.
    • zai-org/GLM-5.
    • zai-org/GLM-4.7.
    • deepseek-ai/DeepSeek-R1-0528.
    • deepseek-ai/DeepSeek-R1.
    • deepseek-ai/DeepSeek-V3-0324.
    • deepseek-ai/DeepSeek-V3.
    • deepseek-ai/DeepSeek-V3.1-Base.
    • deepseek-ai/DeepSeek-V3-Base.
    • deepseek-ai/DeepSeek-R1-Distill-Llama-70B.
    • deepseek-ai/DeepSeek-R1-Distill-Llama-70B-32k.
    • deepseek-ai/DeepSeek-R1-Distill-Llama-70B-131k.
    • deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.
    • deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B.
    • meta-llama/Llama-4-Scout-17B-16E.
    • meta-llama/Llama-4-Maverick-17B-128E.
    • meta-llama/Llama-3.3-70B-32k-Instruct-Reference.
    • meta-llama/Llama-3.3-70B-131k-Instruct-Reference.
    • meta-llama/Llama-3.2-3B-Instruct.
    • meta-llama/Llama-3.2-3B.
    • meta-llama/Llama-3.2-1B-Instruct.
    • meta-llama/Llama-3.2-1B.
    • meta-llama/Meta-Llama-3.1-8B-131k-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-8B-Reference.
    • meta-llama/Meta-Llama-3.1-8B-131k-Reference.
    • meta-llama/Meta-Llama-3.1-70B-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-70B-32k-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-70B-131k-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-70B-Reference.
    • meta-llama/Meta-Llama-3.1-70B-32k-Reference.
    • meta-llama/Meta-Llama-3.1-405B-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-405B-Reference.
    • meta-llama/Meta-Llama-3.1-405B-10k-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-405B-10k-Reference.
    • meta-llama/Meta-Llama-3.1-405B-8k-Instruct-Reference.
    • meta-llama/Meta-Llama-3.1-405B-8k-Reference.
    • meta-llama/Llama-3-8B-Instruct.
    • Qwen/Qwen2.5-14B.
    • Qwen/Qwen2.5-32B.
    • Qwen/Qwen3-235B-A22B-Instruct-2507-FP8.
    • Qwen/Qwen2-72B.
    • arcee-ai/trinity-mini.
    • BAAI/bge-base-en-v1.5.
    • minimax/speech-2.8-turbo.
    • rime-labs/rime-mist-v3.
    • rime-labs/rime-mist-v3-omni.
    Original source
  • Jul 28, 2026
    • Date parsed from source:
      Jul 28, 2026
    • First seen by Releasebot:
      Jul 30, 2026
    • Modified by Releasebot:
      Aug 4, 2026
    Together AI logo

    Together AI

    July 28, 2026

    Together AI adds CLI improvements for A/B traffic updates, smarter replica-bound handling for deployments, and upgrade notices in the Together CLI. It also brings live dedicated endpoints into evaluations and deprecates MiniMaxAI/MiniMax-M2.7 on serverless.

    Improvements

    Deprecations

    A/B variant percent updates in the CLI

    tg beta endpoints update now accepts --ab-percent to change a variant’s traffic percentage in an existing A/B experiment. The flag takes percentage from or returns it to the control only; other variants stay unchanged. The control must remain at least 1%, and --percent on tg beta endpoints ab is limited to 1–99.

    See Ramp the variant and the endpoints CLI reference.

    Deploy replica bound inference

    tg beta endpoints deploy now infers a missing replica bound: --min-replicas alone mirrors into the max (including 0 to create a deployment stopped), and --max-replicas 0 alone lowers the min to 0. On tg beta endpoints update, stopping a deployment still requires both --min-replicas 0 and --max-replicas 0. Passing a single zero bound is an error.

    See the endpoints CLI reference.

    CLI upgrade notices

    The Together CLI now detects when a newer release is available and prints an upgrade notice at most once per day. Interactive sessions offer to run the upgrade in place, using the command that matches your install (uv, pipx, or pip). Set TOGETHER_DISABLE_VERSION_CHECK=1 to turn the check off.

    See Get started.

    Select dedicated endpoints in evaluations

    In the evaluations console, live dedicated model inference endpoints now appear under My Endpoints in the model picker. Legacy dedicated endpoints appear under My Legacy Endpoints. Only endpoints with at least one live deployment are listed.

    See Supported models.

    Model deprecations

    The following models have been deprecated and are no longer available on serverless:

    MiniMaxAI/MiniMax-M2.7.

    See Deprecations for migration options.

    Original source
  • Jul 28, 2026
    • Date parsed from source:
      Jul 28, 2026
    • First seen by Releasebot:
      Jul 28, 2026
    • Modified by Releasebot:
      Jul 29, 2026
    Together AI logo

    Together AI

    July 28, 2026

    Together AI deprecates MiniMaxAI/MiniMax-M2.7 on serverless and points users to migration options.

    Improvements

    Deprecations

    The following models have been deprecated and are no longer available on serverless:

    MiniMaxAI/MiniMax-M2.7.

    See Deprecations for migration options.

    Original source
  • Jul 27, 2026
    • Date parsed from source:
      Jul 27, 2026
    • First seen by Releasebot:
      Jul 28, 2026
    • Modified by Releasebot:
      Aug 9, 2026
    Together AI logo

    Together AI

    July 27, 2026

    Together AI adds serverless Kimi K3 with 1M context and expands fine-tuning to Qwen3.6-27B.

    New serverless models

    The following models are now available on serverless:

    moonshotai/Kimi-K3: 1,000,000 context length. Pricing: $3.00 input / $15.00 output / $0.30 cached input (per 1M tokens). Supports function calling, structured outputs, and vision inputs.

    See Kimi K3 quickstart.

    New models available for fine-tuning

    You can now fine-tune the following models:

    Qwen/Qwen3.6-27B.

    See Supported models for the full list.

    Original source
  • Jul 24, 2026
    • Date parsed from source:
      Jul 24, 2026
    • First seen by Releasebot:
      Jul 25, 2026
    • Modified by Releasebot:
      Aug 9, 2026
    Together AI logo

    Together AI

    July 24, 2026

    Together AI adds Python SDK realtime transcription over WebSocket and expands fine-tuning comparison metrics filtering, with improved reconnects, normalized transcript events, and failover support for streaming speech-to-text.

    Python SDK realtime transcription

    The Together Python SDK now includes client.beta.realtime.transcription() for streaming speech-to-text over WebSocket. Install with pip install "together[realtime]". The session reconnects with audio replay on transient drops, exposes normalized events such as TranscriptDelta and TranscriptCompleted, and supports application-level failover across endpoints with RealtimeConnectionError (code="no_healthy_workers") and session.pending_audio().

    See Streaming transcription.

    Fine-tuning comparison metrics filtering

    The fine-tuning comparison view now includes the same Metrics filtering control as the single-job Metrics tab. Adjust Sampling rate and Step range, then select Apply to re-fetch metrics for every selected job with matching filters.

    See View metrics in the dashboard.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.