Unsloth Release Notes

Follow

25 release notes curated from 1 source by the Releasebot Team. Last updated: Aug 21, 2026

Get this feed:
  • Aug 20, 2026
    • Date parsed from source:
      Aug 20, 2026
    • First seen by Releasebot:
      Aug 21, 2026
    Unsloth logo

    Unsloth

    Auto compaction + LAN Access

    Unsloth ships v0.1.801-beta with 200+ merged PRs, adding experimental auto compaction for long chats, preview remote and LAN access, faster chat streaming, and support for custom llama.cpp builds. It also releases Unsloth Dynamic v3.0 with new Qwen3.8-27B GGUFs and higher accuracy.

    Thanks for the support for Qwen3.8-27B and Unsloth Desktop last week! For our new v0.1.801-beta release, we merged 200+ PRs to introduce many new features, fixes including:

    Auto Compaction (Experimental) for longer chats beyond context limits

    Remote & LAN Access (Preview) for easy network access without Cloudflare links
    Faster Chat - Improved streaming performance, reduced UI lag, and smoother long conversations.
    Support for custom llama.cpp builds. Toggles for Cache RAM, Mmap, Mlock, Checkpoints, Spe Decoding KV Cache, Vision On / Off
    Unsloth Dynamic v3.0 is released. New Qwen3.8-27B Dynamic v3.0 GGUFs deliver >10% higher top-1 accuracy compared to everyone else. Works with Unsloth.

    Auto compaction (Experimental)

    Long chats can now exceed context limits by moving older turns into a searchable archive.

    Older turns are removed only when needed and remain searchable. Fresh context epochs improve recall without permanently trimming chats. Uses retrieval instead of summarization for better accuracy.

    Remote & LAN access (Preview)

    Access Unsloth from other devices on your network.

    New remote access settings. LAN control, QR codes, and auto-start options. Disabled by default for security.

    Chat improvements

    Faster long chats and improved threading. Projects organize chats, files, and workspaces. Added prompt queueing, shortcuts, edit_file, and better tool support.

    Hardware, inference, API

    Support for custom llama.cpp builds and Intel XPU. More inference controls (cache, mmap, mlock, checkpoints, speculative decoding, vision). Improved GPU validation, VRAM handling, and compatibility. Structured outputs in Responses API. Better llama-server recovery.

    To run or train Qwen3.8, you can download Unsloth Desktop:

    Download for macOS Download for Windows Download for Linux

    Original source
  • Aug 14, 2026
    • Date parsed from source:
      Aug 14, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    • Modified by Releasebot:
      Aug 21, 2026
    Unsloth logo

    Unsloth

    Qwen3.8

    Unsloth adds local support for Qwen3.8-27B and Qwen3.8-2.4T, plus fine-tuning for Qwen3.8-27B. It also brings faster inference, new NVFP4 and GGUF quants, tool calling for connected providers, and Codex login and tools in Chat.

    Qwen3.8-27B and Qwen3.8-2.4T can now be run locally in Unsloth!

    Run on 17GB RAM via Unsloth Dynamic GGUFs. You can also fine-tune Qwen3.8-27B in Unsloth. Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants.

    Other Unsloth updates include:

    • Qwen3.8-27B + extra llama-server arguments allowed + custom VRAM toggle
    • External provider has tool calling + tool support + login with Codex
    • Fast FP8 10x faster MiniMax-H3 inference (3 minutes vs 30)
    • 10% faster inference for GGUFs + Bypass permissions fixed
    • Connected API providers can use their own Search or Unsloth Desktop's built-in Search and tools. Tool results are passed back to the model so it can continue multi-step tasks.
    • Sign in with a Codex subscription and use Codex tools inside Chat.

    To run or train Qwen3.8, you can download Unsloth Desktop:

    Download for macOS Download for Windows Download for Linux

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Unsloth and hundreds of other software products.

    Create account
  • Aug 11, 2026
    • Date parsed from source:
      Aug 11, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    Introducing Unsloth Desktop

    Unsloth launches Unsloth Desktop, an open-source app for running and training models locally on Mac, Windows and Linux. It supports MLX, diffusion, audio, GGUF, private search, RAG, MCP, and faster training with lower VRAM.

    Introducing Unsloth Desktop 🦥 - the first desktop app to run and train models locally.

    Open-source. Runs on Mac, Windows and Linux

    • Supports MLX, diffusion image/video, audio, GGUF
    • Connect Claude Code and Codex to local LLMs
    • 50% more accurate, self-healing tool calls + sandboxed code exec
    • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
    • Train models 2× faster with 70% less VRAM
    • Private web search, deep research, RAG, MCP + exports (NVFP4, GGUF)
    • Use Unsloth’s OpenAI-compatible API and cloud models
    • Securely deploy LLMs remotely and access anywhere

    You can download Unsloth Desktop now:

    Download for macOS Download for Windows Download for Linux

    Original source
  • Jul 29, 2026
    • Date parsed from source:
      Jul 29, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    Kimi K3 + Deep Research + Parallel Chat

    Unsloth adds local Kimi K3 Dynamic GGUF support, parallel chats, and a new Deep Research mode that plans, reads, and cites sources with a local model. It also expands AMD and Intel GPU support, adds DoRA training, and delivers many installer, export, MLX, and inference fixes.

    Kimi K3

    Moonshot AI’s Kimi K3 is a 2.8T-parameter MoE model with 104B active parameters, native vision support and a 1M context window. Kimi K3 is thinking-only and Unsloth supports low, high and max reasoning efforts.

    You can run our Kimi K3 Dynamic GGUFs through Unsloth. Unsloth automatically detects multi-GPU setups and can offload model layers to system memory.

    Kimi K3 is a very large model, so plan your hardware accordingly:

    • UD-IQ1_S is 595GB in disk space.
    • UD-Q4_K_XL is 1.51TB in disk space.
    • For lossless inference, use UD-Q8_K_XL, which is 1.56TB in disk space.

    Read our full Kimi K3 guide and learn more about Unsloth Dynamic 2.0 GGUFs.

    Parallel Chat

    Unsloth can now run multiple conversations at once. Starting a New Chat leaves the previous answer generating, and each active conversation gets its own progress indicator and Stop control.

    • 4 llama-server slots by default, adjustable in the web UI.
    • Tools, uploads, self-healing and agents stay isolated between chats.
    • Stop one chat without interrupting others or restarting the server.
    • Unsloth reduces the slot count when memory is limited.
    • Reloading the model still stops active chats after confirmation.

    Deep Research

    Deep Research turns a local model into a complete research workflow. Give it a question and it will create a plan, search the web, organize the evidence and produce a cited report.

    • Review and edit the plan before research begins.
    • Follow its progress and collected sources while it works.
    • Resume or cancel without losing completed research.
    • Allow or block websites to control source selection.
    • Sources and citations stay saved with each run.
    • Before writing, Unsloth checks for unsupported claims, contradictions and unresolved gaps. Recommendations without direct evidence are marked as testable inferences.

    One run can be active per chat. Deep Research currently works with local models. Optional full-page grounding can read top search results into temporary RAG before synthesis; enable it with UNSLOTH_RESEARCH_AUTO_SCRAPE=1. It remains off by default because it adds scraping time and needs additional context.

    Improved AMD Support

    This update builds on our initial AMD release and setup guide with support for more GPUs and more reliable ROCm inference and training.

    • Improved detection for RDNA2, Radeon, Ryzen, Strix Halo and workstation GPUs.
    • MI50 and Radeon VII support 16-bit LoRA and full fine-tuning on Linux.
    • Fixed 4-bit NaNs, library conflicts and long RDNA startup stalls.
    • Unsupported Windows HIP GPUs can fall back to Vulkan.
    • Vulkan devices show their real names and can be selected individually.
    • Clean Windows installs no longer require Winget or developer tools.

    Training, Models and Platform Updates

    • Intel XPU enables local chat and training on Intel Arc and Data Center GPUs.
    • DoRA is available alongside LoRA and full fine-tuning.
    • Large exports can use all visible GPUs to avoid GPU 0 memory limits.
    • MLX adds streaming datasets, continued pretraining, callbacks and better VLM and LoRA support.
    • Fixed incorrect flash_attention_2 model output.
    • Unsloth and saving utilities now work without bitsandbytes.
    • Fixed .json datasets and Qwen3.5/3.6 MoE and GRPO notebook setup.
    • Hub download folders are easier to find and open.
    • Release notes are available in the Unsloth update popup.
    • unsloth start --as-subagent lets Codex, Claude Code and Pi delegate tasks to a local Unsloth model.
    Original source
  • Jul 20, 2026
    • Date parsed from source:
      Jul 20, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    AMD support is here!

    Unsloth adds local LLM training and inference on AMD GPUs across Windows, WSL and Linux, with faster ROCm builds, broader model support and better GPU detection. The release also improves multi-GPU use, chat restarts, downloads, agents and export reliability.

    Hey everyone!

    This release brings local LLM training and inference to AMD GPUs across Windows, WSL and Linux.

    Starting today, our AMD collaboration, custom Triton kernels, and math algorithms enables you to train and run 500+ models across AMD's Radeon, Instinct, Vulkan, Ryzen and data center GPUs, up to 2× faster with 70% less VRAM and no accuracy loss. Optimized ROCm builds also support GGUF & Safetensors inference.

    July 23 Update

    • Added RDNA2, Gorgon Halo, Vulkan support + fixed AMD installing not detecting GPUs on Strix Halo / other AMD GPUs
    • Better RDNA4, HIP / ROCm failure auto fixing and catching
    • 2x faster unified memory AMD safetensors loading + much faster gradient checkpointing for unified memory devices
    • Added voice dictation / whisper.cpp preliminary support for fast text to speech
    • Fixed rollback environments during installs eating 5GB of disk space - now auto cleans

    Train LLMs Locally on AMD

    Train, run RL, chat with and deploy models locally on AMD GPUs.

    More reliable AMD GPU detection and installation across Windows, WSL and Linux.

    Improved ROCm compatibility for AMD MI300X and MI325X GPUs.

    Remote access Unsloth via unsloth studio --secure through free HTTPS via Cloudflare

    Run Larger Models on Your Hardware

    Use automatic GPU placement or choose exactly which GPUs and model layers to use.

    Move MoE expert layers into system memory to help larger models fit.

    Split models across multiple GPUs or use Tensor Parallelism.

    Save hardware settings separately for each model and quant.

    Faster Chat Restarts and More Reliable Downloads

    Resume long chats without rebuilding the full conversation after an idle model frees its VRAM.

    Stalled Hugging Face XET downloads automatically retry over standard HTTP.

    Existing GGUF files are reused instead of being downloaded again whenever a model loads.

    Better Search, Tools and Agents

    Web search can now read PDF papers, manuals and other PDF results.

    Parallel tool calls, reasoning output and tool retries work more reliably.

    A new opt-in MCP endpoint lets compatible AI clients inspect models and training history, start or stop training, load checkpoints, validate recipes and export GGUFs.

    Enable it with UNSLOTH_STUDIO_ENABLE_MCP=1 and set the required bearer token with UNSLOTH_STUDIO_MCP_TOKEN.

    Training and Export Fixes

    Text-only training with multimodal models no longer truncates long examples before packing.

    Fine-tuned Qwen3.5 and Qwen3.6 MTP models now export correctly to GGUF.

    Fixed a Windows permission error that could stop GGUF exports during the final write step.

    In Case You Missed It

    Our previous Studio release added Dynamic NVFP4 models, deeper personalization, seven new display languages, safer agents and Vulkan GPU acceleration.

    Dynamic NVFP4:

    Unsloth Dynamic NVFP4 keeps accuracy-sensitive layers in FP8 or BF16 while running the rest in W4A4. On NVIDIA Blackwell GPUs, this enables up to 2.5x faster inference, while calibrated FP8 KV caches provide up to 2x longer context.

    Read the Dynamic NVFP4 guide and explore our expanded NVFP4 collection, including Qwen3.6, Qwen3.5, Inkling, GLM-4.7 Flash and Gemma 4.

    Personalize Your Studio:

    Choose between Standard, Classic and Minimal color palettes, each with light and dark modes. You can also customize colors, import fonts and adjust font size, contrast, reduced motion and more.

    Unsloth also includes a Voice settings tab for dictation, custom dictionary settings and read-aloud.

    New Languages:

    Unsloth Studio is now available in French, German, Spanish, Hindi, Arabic, Russian and Korean, alongside Chinese (Simplified), Japanese and Portuguese (Brazil). Browser-language auto-detection is now the default.

    Safer Agents:

    A four-level tool-call permission selector-Ask, Approve for me, Off and Full access-provides finer control over agents. Agent workspace isolation and safer installer checks also reduce the risk of unintended changes.

    Original source
  • Similar to Unsloth with recent updates:

  • Jul 15, 2026
    • Date parsed from source:
      Jul 15, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    Personalization, NVFP4, languages

    Unsloth adds a major customization and performance update with new color palettes, custom fonts, seven more display languages, a Voice settings tab, safer agent tool controls, and Vulkan llama.cpp support for Intel GPUs, plus native Inkling support and expanded Dynamic NVFP4 models.

    Personalize Your Unsloth

    Unsloth is no longer limited to light and dark mode:

    Choose from Standard, Classic, and Minimal palettes, each with light and dark variants. Customize accent, background, and foreground colors. Import your own UI, heading, chat, and code fonts. Adjust font size, contrast, motion, cursor behavior, and font smoothing. Search settings and sync preferences across devices.

    Light and dark modes keep their own customization values.

    Seven New Languages

    Unsloth now supports French, German, Spanish, Hindi, Arabic, Russian, and Korean, alongside Chinese, Japanese, and Portuguese. Browser-language auto-detection is now the default.

    Voice Settings

    A new Voice tab adds controls for dictation, pronunciation dictionaries, and read-aloud. Speech-to-text and text-to-speech support are coming soon; full voice conversations are still in development.

    Safer Agents and Tool Calls

    The old bypass toggle is now a four-level permission selector:

    • Ask: approve every tool call.
    • Approve for me: automatically run read-only actions and pause for potentially unsafe ones.
    • Off: disable tool calls.
    • Full access: allow all calls and disable the sandbox.

    These controls cover terminal, Python, web search, RAG, and MCP tools. Agents also gain isolated workspaces, resumable sessions with --persist, safer remote-installer warnings, live tool-output streaming, and improved web extraction.

    Vulkan and llama.cpp

    The new Vulkan llama.cpp backend gives Intel GPUs GPU-accelerated inference instead of falling back to the CPU. It supports existing VRAM management, automatic context sizing, multi-GPU selection, and layer offloading.

    AMD users can opt in with:

    UNSLOTH_FORCE_VULKAN=1
    

    We also improved update reliability, model-state recovery, and interruption of stalled generations.

    Models and Training

    This release also includes:

    • Native Inkling support and multi-GPU B200 improvements.
    • DeepSeek-V4 eager attention and trainable FP8 grouped experts.
    • Reliable completion-only training through automatic chat-template marker detection.
    • Correct handling of disabled gradient checkpointing.
    • Improved RoPE scaling and context extension.
    • Force-stopping for stuck training runs.
    • Automatic routing for new model architectures and newer Transformers releases.

    Dynamic NVFP4

    We expanded our NVFP4 collection with quantized versions of Qwen3.6, Qwen3.5, Inkling, GLM-4.7 Flash, and Gemma 4.

    Read the Dynamic NVFP4 guide for details.

    We’ve also added native support for Inkling, a 975B-parameter open model with 41B active parameters and up to a 1M-token context window. Licensed under Apache 2.0, it accepts text, images, and audio and generates text.

    Hey guys we got lots of new update for Unsloth, especially customization. Unsloth now yours to personalize: three color palettes plus custom colors and fonts, seven new display languages, and a new Voice settings tab for dictation and read-aloud. Agents get safer with a four-level tool-call permission selector (Ask, Approve for me, Off, Full access) and workspace isolation, and Intel GPUs finally get GPU-accelerated inference through new Vulkan llama.cpp support.

    Original source
  • Jul 7, 2026
    • Date parsed from source:
      Jul 7, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    DeepSeek-V4 + NVFP4

    Unsloth releases a broad upgrade with faster GRPO and MoE training, smarter OpenAI-compatible API serving, richer export options, and expanded language support. It also improves RAG, file chat, installer reliability, model loading, and many training and kernel fixes.

    Unsloth can now export NVFP4, FP8, and imatrix GGUFs after training; act as a llama-swap API system; add Japanese and Brazilian Portuguese support; and includes MLX, safetensors, tool calling, healing support, and more. Unsloth core makes GRPO 1.3x faster, adds HTTP fallback for stalled downloads, improves offline mode, speeds up MoE training by 3-5x, and fixes many bugs. This release series uses unsloth>=2026.7.1.

    DeepSeek-V4-Flash is now supported with Thinking toggles and our improved chat template fixes.

    Smarter OpenAI-Compatible API Serving

    Run one local API endpoint with safer model swapping and better agent-tool recovery.

    API requests can opt into automatic switching between downloaded local GGUFs, while unknown model names safely keep using the current model./v1/models now returns clean model IDs and the local GGUF catalog instead of local .gguf paths.Idle auto-unload can free VRAM after inactivity, and tool-call healing can now be controlled per request.

    Export Improvements

    Exports are more flexible and avoid unnecessary downloads.

    Select multiple export formats at once, including portable FP8/INT8, GGUF LoRA, source-matched exports, imatrix GGUF, and compressed FP8/FP4.Multi-checkpoint exports avoid more repeated base-model downloads.FP8, INT8, and GGUF-LoRA exports now respect trust_remote_code, and GGUF export handles missing quantization settings more reliably.

    RAG and File Chat

    File chat is more useful on real documents.

    RAG attachments can now use whole-document context, with customizable embedding models and Hugging Face search.File chat reads more PDFs and Word documents correctly, including right-to-left text, Indic text, and DOCX tables.Local RAG checks are more reliable behind proxy setups.

    Unsloth Polish and Reliability

    Everyday use should feel smoother and more stable.

    Long training and chat runs are less likely to freeze silently.Compare mode, model switching, model cancellation, Hub browsing, and Hub Discover are more reliable.Project exports, chat exports, settings, guided tours, file dialogs, update screens, and reasoning UI are cleaner and more consistent.

    Installer, Hardware, and Platform Fixes

    Unsloth installs and runs more reliably across platforms.

    macOS installs no longer require CMake or Homebrew when a prebuilt llama.cpp is available, and Apple Silicon support handles paths with spaces and unified memory sizing better.Windows startup, UTF-8 handling, ROCm RAG embedding, and ROCm-on-WSL GPU support were improved.Blackwell GPU prebuilt selection, GGUF fit checks, tensor parallelism for vision/mmproj GGUFs, and local llama.cpp reuse are now more reliable.

    Training, Models, and Kernels

    Training and model loading are more reliable across more setups.

    GRPO now supports sequence packing by default, avoids repeated shared prompts, and handles DDP logit scaling correctly.Full fine-tuning, RL precision settings, gradient checkpointing, DDP RoPE buffers, and MoE LoRA detection were fixed.FP8 quantization/dequantization, Llama 3 RoPE scaling with Transformers v5, PEFT 0.19 LoRA reloads, and fast_generate error messages were improved.

    Original source
  • Jun 18, 2026
    • Date parsed from source:
      Jun 18, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    GLM 5.2 + Hub + 3x longer contexts

    Unsloth adds GLM-5.2 support in Unsloth Studio with all reasoning levels, 3x longer context, forkable and queued chats, a redesigned model hub, secure Cloudflare access, parallel modules, faster logging, and broader GPU and training reliability improvements.

    GLM-5.2 is now supported in Unsloth Studio!

    All reasoning levels supported. 3x longer context lengths are now achievable with our new auto fit algorithm with MTP, allowing longer chats. Bypass permissions mode, forkable chats, queue-able chats, a new hub for model discovery, parallel modules + HTTPS Cloudflare support and more! Use unsloth studio --secure for secure HTTPS global access!

    Better context length algorithm

    As per PR 1 and PR 2, we made Unsloth Studio's determination of memory usage and context length much better, achieving 3x longer context overall:

    Chat Canvas, Forking & Queueing

    Edit assistant messages in place and re-run from any point in the thread. Fork a thread to branch a conversation without losing the original. Temporary (incognito) chats that leave nothing behind. Queue new prompts while a generation is still running instead of waiting. Chat "artifacts" are now canvas, with inline HTML canvas cards that auto-render, a Code view, and DiffusionGemma keeps its raw code visible inline instead of collapsing. Chat search now covers every message and surfaces your own messages first.

    Hub (Redesigned)

    Full-page Hub with a trending feed, search, and custom model paths support. README preview in a split-view feed so you can read before you download. Downloads default to the faster Xet transport, with automatic HTTP fallback if a transfer stalls. New "Load on selection" toggle to set load options before a model loads. Google logo shown for DiffusionGemma and future Gemma derivatives.

    Models & Inference

    DeepSeek-OCR and more vision models now load and run without errors. Fixed fast inference on the latest vLLM (0.22+) so speed-ups work again. Tensor parallelism is more reliable: if the faster MTP path fails, it now recovers on its own instead of crashing. DiffusionGemma now shows the image forming live as it denoises, with accurate speed stats.

    Security & Cloudflare Encrypted Studios

    New --secure Cloudflare-only mode for end-to-end encrypted studios, with server-side tools staying enabled under --secure. Use unsloth studio --secure! Bypass Permissions mode to skip confirmations and disable the tool sandbox when you want it. Auto detect Hugging Face Virus scanning + dangerous files in repos.

    Logging and API

    New API server monitor in Unsloth. Faster API calling and less latency Much better streamlined logs - now with throughput and latency and removed a lot of bloated logs.

    Hardware & Backend

    Better support for Blackwell RTX 50X and 60X GPUs Fix silent downgrading to CPU and not GPU torchao version is now selected from the installed torch. Installer now auto-repairs a broken or CPU-only PyTorch install and warns on silent CPU fallback, across NVIDIA + AMD on Win/Linux/Mac/WSL. Frees the chat model's VRAM when training starts, but only when the GPU is actually tight (no needless reloads otherwise). If llama-server hard-crashes at startup, Unsloth now steps through a recovery ladder instead of just failing.

    Training & General Fixes & Parallel Modules

    MLX training updates. Improved GRPO training reliability with vLLM. Training startup made more reliable, with clearer errors for invalid VLM batches. Unsloth now cleans up leftover backend processes more reliably after crashes, restarts, or interrupted shutdowns. Export, Chat, Training, Recipes are all individualized / compartmentalized! This means you can do all 4 in parallel now! You can chat / do inference while you wait for a training run or an export!

    To update Unsloth or install a new Unsloth Studio, you must use:

    macOS, Linux, WSL:

    curl -fsSL https://unsloth.ai/install.sh | sh
    

    Windows:

    irm https://unsloth.ai/install.ps1 | iex
    
    Original source
  • Jun 12, 2026
    • Date parsed from source:
      Jun 12, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    DiffusionGemma + Gemma 4 MTP

    Unsloth ships major Studio upgrades with support for DiffusionGemma, Gemma 4 MTP and MiniMax-M3, plus faster Gemma 4 runs, audio chat, a new Hub and Chat with Files, easier llama.cpp updates, stronger hardware support, and reliability fixes across chat, APIs, and training.

    Ensure you install the latest v0.1.464-beta or 2026.6.7. DiffusionGemma, Gemma 4 MTP and MiniMax-M3 are all now supported.

    Run and train DiffusionGemma via Unsloth Studio.Gemma 4 MTP is here! Run Gemma 4 ~2x faster with MTP.Audio chat is now supported for Gemma 4 (wav, mp3, m4a, flac, webm).Preserve Think added to Gemma 4.

    Hub + Download Manager (Experimental)

    Added a new Hub page for browsing, downloading, and managing Hugging Face models and datasets.Unsloth can now detect models and datasets already on your machine and show them alongside downloaded assets.Downloaded GGUF models now have direct Run / New Chat actions.

    RAG / Chat with Files (Experimental)

    Added Chat with Files in Unsloth, letting you ask questions over your own documents and knowledge bases.Supports hybrid search, citations, PDF previews, per-thread documents, and a built-in search_knowledge_base tool.

    New Update Button + Hardware Support

    Unsloth now uses constant fresh up to date llama.cpp prebuilts across CUDA, ROCm, Windows, Linux, and macOS.Added an in-app Update llama.cpp button so users can update the local backend without reinstalling Unsloth.Improved Windows / WSL AMD support, Strix Halo ROCm support, Blackwell CUDA selection, and clearer installer messages.

    Local Chat, Tools & API Compatibility

    Local tool calling is more reliable, with better ordering of tool cards, fewer duplicate tool loops, and support for tool use with GGUF vision models.Improved OpenAI-compatible API and Anthropic-compatible API behavior for local Unsloth servers, including better errors, token usage, stop reasons, and Claude Code compatibility.

    Training & Fixes

    Improved MLX support with better model labels, generation speed stats, and fixes for VLM training.Fixed several training and dataset edge cases, including non-writable Hugging Face caches and custom dataset mappings.Added many UI polish fixes across chat, menus, model picker, dark mode, import/export, and settings.

    To update Unsloth or install a new Unsloth Studio, you must use:

    macOS, Linux, WSL:

    curl -fsSL https://unsloth.ai/install.sh | sh
    

    Windows:

    irm https://unsloth.ai/install.ps1 | iex
    
    Original source
  • Jun 3, 2026
    • Date parsed from source:
      Jun 3, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    Gemma 4 12B, New UI, MCP, Projects

    Unsloth releases a broad update with Gemma 4 12B support, MCP improvements, new Projects and Canvas tools, a refreshed chat UI, and CUDA 13.3 runtime updates. It also expands Studio support across Windows, Linux, ROCm, Blackwell, B300, and ARM64.

    This update focuses mainly on Gemma 4 12B, MCP, Projects, Canvas, CUDA 13.3 and the new chat UI. Next week we'll have an even bigger update.

    Gemma 4 12B

    Google releases Gemma 4 12B, a new model that runs locally on 8GB RAM. GGUF / Guide

    Gemma 4 12B Unified supports image, audio and 256K context. Run and train the model via Unsloth Studio.

    MCP

    Remote MCP server support, including custom headers and OAuthLocal command-based MCP server supportMCP can now be turned on from the chat composerBuilt-in presets for common MCP servers

    New Chat UI

    Projects, Canvas, MCP, RAG and Compare controls now live in the plus menuSearch and Code controls are easier to access from the composerMenus, overlays, icons and clickable controls are more consistent across Unsloth

    Projects

    Organize related chats into dedicated project workspacesMove existing chats into projectsCreate and manage projects directly from the sidebar

    Experimental Canvas / Artifacts

    Opens generated HTML in a dedicated canvas panel inside Unsloth StudioSupports interactive outputs, including browser based visualizations and CDN-loaded packagesLets you switch between rendered preview and source code

    Install, Runtime and Hardware

    Windows prebuilt installs no longer require the early CUDA Toolkit checkLinux llama.cpp prebuilts now match the detected runtime cudart majorROCm gfx detection is forwarded into prebuilt selectionBlackwell, B300 and ARM64 Linux support updates

    To update Unsloth or install a new Unsloth Studio, you must use:

    macOS, Linux, WSL:

    curl -fsSL https://unsloth.ai/install.sh | sh
    

    Windows:

    irm https://unsloth.ai/install.ps1 | iex
    

    DO NOT USE unsloth studio update anymore since packaging will not get the latest updates!

    Original source
  • May 31, 2026
    • Date parsed from source:
      May 31, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    CUDA 13.3, Windows, Mac

    Unsloth updates Studio install support with refreshed macOS, Windows, Linux and WSL install paths, re-enabled llama.cpp prebuilt binaries for Apple Silicon and Intel Macs, Windows CUDA 13.3 support, and a CUDA 13.3 fix for gibberish output, while Blackwell binaries remain delayed.

    To update Unsloth or install a new Unsloth Studio, you must use:

    macOS, Linux, WSL:
    curl -fsSL https://unsloth.ai/install.sh | sh
    
    Windows:
    irm https://unsloth.ai/install.ps1 | iex
    

    DO NOT USE unsloth studio update anymore since packaging will not get the latest updates!

    Mac Updates

    Re-enabled llama.cpp prebuilt binaries for Apple Silicon (M1-M4) - Mac OS 14 / 15 / 26 (Tahoe)Apple Silicon Mac OS 13 (Ventura) is source buildIntel (x86_64) for Mac OS 13.3 / 14 / 15 / 26 (Tahoe) uses llama.cpp prebuilt binariesIntel for Max 13.0 - 13.2 is source build

    Windows Updates

    CUDA 13.3 llama.cpp prebuilt binaries now work for WindowsFor CUDA 13.2, CUDA 13.1 and below, Windows devices uses CUDA 12.4 fallback - we'll work on CUDA 13.1 binaries soon.

    CUDA 13.3 Update

    CUDA 13.3 non Linux binaries work. We'll still use CUDA 13.1 for nowCUDA 13.3 solves the CUDA 13.2 gibberish problem - see https://github.com/unslothai/unsloth/issues/4849

    Blackwell GPUs Update

    For now Blackwell will have delayed releases of llama.cpp prebuilt binaries sine CUDA 12.4 does not work - we are working to resolve this soon.

    Original source
  • May 26, 2026
    • Date parsed from source:
      May 26, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    An update before Revamp.

    Unsloth adds broader API calling, image generation and editing, web search, code execution, prompt caching, and support for OpenAI, Anthropic, OpenRouter, local backends, and non-English languages, while also tightening Studio security and offline workflows.

    Hey guys, we're doing one more-ish update before a major revamp which is likely coming this week or next week. Our revamp will change a lot of things, especially with new major features and a lot of design changes.

    NEW: API calling support now with image generation + editing, proper web search, code execution, auto prompt caching. Connect OpenAI, Anthropic and more. Proper support for non-English languages e.g. Japanese, Chinese, Indian etc.

    Many of you may have missed our previous release which only lasted for one day. We introduced:

    • Connect to external inference backends: vLLM, Ollama, llama-server
    • Security improvements
    • Auto MTP speculative decoding for MTP GGUFs; get the best settings customized for your hardware.

    API provider calling & external connections

    You can now connect Unsloth to any API cloud provider (OpenAI, Anthropic, OpenRouter etc.)
    Built-in web search for OpenAI, Anthropic, OpenRouter and Kimi
    Built-in code execution for OpenAI and Anthropic (Anthropic containers persist and are reused across turns)
    Prompt caching is enabled for OpenAI and Anthropic models saving 50 to 90% of costs.
    Image generation + editing
    API key now optional for local providers (llama.cpp / vLLM / Ollama)
    Auto-load models when adding a cloud provider

    Other Unsloth Studio updates

    OpenDocument chat attachments
    o3 reasoning summary payload
    Sending/prompting non-English languages (e.g. Japanese, Chinese) now works properly
    IME composer hardening, RTL dir="auto", long log-line truncation fix
    Tool reasoning trace rendering in UI
    Fully offline support: cached GGUF discovery and offline DNS auto-detect for both inference and training

    Unsloth Studio security improvements

    Authentication rate-limiting, proxy-aware so reverse proxies don't bypass it
    Sandboxed worker with a tightened blocklist (bash, hf upload, NOFILE)
    Path containment so workers can't escape their in-flight tmp dirs
    Strict schema validation across the Unsloth API
    Tightened CSP / security headers (only legitimate favicon hosts allowed)
    Removed the torch.load fallback on training_args.bin so untrusted pickles can never execute on model load
    Hardened Tauri desktop release flow
    Frontend auth: singleflight token refresh, current-password input on changes, working logout, shared 422 helper
    Cancel cleanup now scoped strictly to in-flight tmp dirs so it can never delete user state

    Original source
  • May 19, 2026
    • Date parsed from source:
      May 19, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    MTP + Unsloth Fixes

    Unsloth improves Studio with bug fixes, UI and UX updates, plus faster MTP and better offline mode support.

    Lots of bug fixes, UI, UX fixes to Unsloth! To get the latest updates do:

    macOS, Linux, WSL:

    curl -fsSL https://unsloth.ai/install.sh | sh
    

    Windows:

    irm https://unsloth.ai/install.ps1 | iex
    

    Fixes

    • Fix unsloth studio update not working well
    • Fix getting stuck on reset-password page
    • More offline mode support
    • Improve MTP not being faster on Macs, CPUs and GPUs - now it's much better!
    • Fix Desktop Shortcut not working after update
    • Many many UI UX bug fixes
    Original source
  • May 18, 2026
    • Date parsed from source:
      May 18, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    Qwen3.6 MTP + API Connections

    Unsloth ships a major v0.1.41-beta update with much faster GGUF inference, broader API provider and tool calling support, experimental MLX inference, better non-English language handling, and major training, UI, offline, and security improvements across Unsloth Studio.

    We've got lots of new updates for Unsloth v0.1.41-beta:

    ~2x faster GGUF inference with automatically enabled MTPAPI calling support for OpenAI, Anthropic etc. with auto prompt caching, web search, code execution
    Connect to external inference backends: vLLM, Ollama, llama-server
    Experimental MLX inference
    Proper support for non-English languages
    Security improvements

    MTP speculative decoding support

    1.4 to 2x faster inference!

    Auto MTP speculative decoding for MTP GGUFs; warn when the bundled llama.cpp prebuilt is stale or too old for MTP
    New pre-built llama.cpp binaries for MTP support!

    API provider calling & external connections

    You can now connect Unsloth to any API cloud provider (OpenAI, Anthropic, OpenRouter etc.)
    Built-in web search for OpenAI, Anthropic, OpenRouter and Kimi
    Built-in code execution for OpenAI and Anthropic (Anthropic containers persist and are reused across turns)
    Prompt caching is enabled for OpenAI and Anthropic models saving 50 to 90% of costs.
    API key now optional for local providers (llama.cpp / vLLM / Ollama)
    Auto-load models when adding a cloud provider

    MLX inference (Experimental)

    MLX quants and models now can run locally on your Mac machines!
    We'll be adding thinking, tools and web search soon!

    Other Unsloth Studio updates

    Sending/prompting non-English languages (e.g. Japanese, Chinese) now works properly
    OpenDocument chat attachments
    o3 reasoning summary payload
    IME composer hardening, RTL dir="auto", long log-line truncation fix
    Tool reasoning trace rendering in UI
    Fully offline support: cached GGUF discovery and offline DNS auto-detect for both inference and training
    Lots of UI/UX polish: dark theme refactor, right sidebar redesign, time-of-day sloth mascot, dismissable copyable toasts, larger chat composer, code-execution config polish, composer action pill styling, narrower Discord button

    Training updates

    Gemma attention mask fixes
    Multi Image GRPO
    GRPO hidden-state return experiments
    New Continued Pretraining (CPT) training method as a first-class option
    Gemma-4 MoE LoRA extractor registered to fix grouped_mm contraction crash
    Opt-in fused lm_head + cross-entropy forward, with single-matmul path under UNSLOTH_RETURN_LOGITS=1
    Pass batch size for eval
    Eval/training paths now honour HF_DATASETS_OFFLINE alongside HF_HUB_OFFLINE

    Unsloth Studio security improvements

    Authentication rate-limiting, proxy-aware so reverse proxies don't bypass it
    Sandboxed worker with a tightened blocklist (bash, hf upload, NOFILE)
    Path containment so workers can't escape their in-flight tmp dirs
    Strict schema validation across the Unsloth API
    Tightened CSP / security headers (only legitimate favicon hosts allowed)
    Removed the torch.load fallback on training_args.bin so untrusted pickles can never execute on model load
    Hardened Tauri desktop release flow
    Frontend auth: singleflight token refresh, current-password input on changes, working logout, shared 422 helper
    Cancel cleanup now scoped strictly to in-flight tmp dirs so it can never delete user state

    Original source
  • May 5, 2026
    • Date parsed from source:
      May 5, 2026
    • First seen by Releasebot:
      Aug 14, 2026
    Unsloth logo

    Unsloth

    Unsloth API endpoint

    Unsloth ships a beta bug fix and update focused on more reliable chat and training workflows, with autosaving threads, checkpoint resume support, fixed generation stop behavior, better DPO and VLM GRPO handling, and improved chat history and attachment rendering. It also adds new model support and API endpoint features.

    v0.1.39-beta bug fix May 5th 2026

    Fixes chat history not being shown (existing chat history is not lost) and attachments not attaching correctly. The bug was render-only - use 2026.5.2 or directly call curl -fsSL https://unsloth.ai/install.sh | sh to update

    You can use local LLMs with tools like Claude Code and Codex by connecting them to Unsloth’s API endpoint. This lets you run models like Qwen and Gemma locally, with additional features such as self-healing tool calling, code execution, and web search.

    Using Unsloth as an API inference endpoint is beneficial not only because it is easy to setup and fast, but also because Unsloth provides:

    • Self-healing tool calling, which helps reduce broken or malformed tool calls by 50%
    • Code execution support, allowing Bash and Python execution for more accurate code outputs.
    • Advanced Web search that visits and actually reads webpages to gather in-depth info.
    • Automatic inference settings for GGUF models (temp, top-k etc.)

    New models

    We've also got a handful of new models to run including NVIDIA Nemotron 3 Nano Omni, IBM Granite 4.1 and Mistral 3.5 Medium. We helped Mistral solve some issues with implementation in transformers and GGUFs.

    Unsloth Updates

    • Stopped Unsloth training runs can now resume from checkpoints.
    • Chat threads now autosave and persist more reliably.
    • DPO training hangs in multi-process setups were fixed.
    • VLM GRPO support improved with MROPE updates.
    • Unsloth’s stop button now properly stops generation.
    • Fix chat template disappearing after browser refresh.
    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.