Ollama Release Notes

Follow

135 release notes curated from 1 source by the Releasebot Team. Last updated: Oct 6, 2026

Get this feed:
  • Oct 6, 2026
    • Date parsed from source:
      Oct 6, 2026
    • First seen by Releasebot:
      Oct 6, 2026
    Ollama logo

    Ollama

    v0.40.0-rc6: model: add multimodal embeddings (#18820)

    Ollama adds EmbeddingGemma2Model on MLX with media-aware /api/embed support and a 24-layer bidirectional text encoder.

    Implements the EmbeddingGemma2Model architecture on the MLX runner: 24-layer bidirectional text encoder with PLE, shared gemma4 vision/audio towers, mean-pool + L2 output. /api/embed accepts per-item media via input dicts.

    Original source
  • Oct 6, 2026
    • Date parsed from source:
      Oct 6, 2026
    • First seen by Releasebot:
      Oct 6, 2026
    Ollama logo

    Ollama

    v0.40.0-rc5

    Ollama fixes a patch for the recent mlx update.

    mlx: fix patch for recent mlx update (#18812)

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Ollama and hundreds of other software products.

    Create account
  • Oct 6, 2026
    • Date parsed from source:
      Oct 6, 2026
    • First seen by Releasebot:
      Oct 6, 2026
    Ollama logo

    Ollama

    v0.40.0-rc4: MLX: version bump (#18720)

    Ollama bumps MLX and adds unit test scopes to reduce memory usage.

    MLX: version bump

    add scopes for unit tests to reduce memory usage

    Original source
  • Oct 6, 2026
    • Date parsed from source:
      Oct 6, 2026
    • First seen by Releasebot:
      Sep 25, 2026
    • Modified by Releasebot:
      Oct 6, 2026
    Ollama logo

    Ollama

    v0.40.0

    Ollama adds default MLX support on Apple Silicon and expands model availability with new MLX-compatible models.

    What's Changed

    Models run on MLX on Apple Silicon by default

    In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.

    ollama pull qwen3.8
    ollama run qwen3.8
    

    Additional models include gemma4, qwen3.6 and qwen3.5

    Decision models are now available on MLX as well: Nimble tev1 clef clef-flash

    MLX now has support for an embedding model: embeddinggemma-2

    We will continue testing and enabling additional models.

    Full Changelog: v0.35.1...v0.40.0

    Original source
  • Oct 4, 2026
    • Date parsed from source:
      Oct 4, 2026
    • First seen by Releasebot:
      Oct 6, 2026
    Ollama logo

    Ollama

    v0.40.0-rc3: pull: allow RCs to pull matching min_version (#18790)

    Ollama strips pre-release tags from version strings so RCs can pull models for the same release.

    Strip off pre-release from the version string so RCs can pull models for the same release.

    Original source
  • Similar to Ollama with recent updates:

  • Oct 4, 2026
    • Date parsed from source:
      Oct 4, 2026
    • First seen by Releasebot:
      Oct 4, 2026
    Ollama logo

    Ollama

    v0.40.0-rc1: mlx: match publisher tokenizer semantics (#18779)

    Ollama improves mlx tokenizer compatibility with shared Go/Python reference cases and regression fixes for encoding behavior.

    mlx: match publisher tokenizer semantics

    Honor pretokenizer stage order, split behavior, Unicode boundaries, added-token normalization, and ranked BPE merges. Handle empty added tokens and empty Metaspace input consistently.

    Add shared Go/Python reference cases using published tokenizers, pulling missing models directly and failing on errors, plus focused regressions for configuration precedence, byte fallback, and parallel encoding.

    address comments

    Original source
  • Oct 2, 2026
    • Date parsed from source:
      Oct 2, 2026
    • First seen by Releasebot:
      Oct 4, 2026
    Ollama logo

    Ollama

    v0.35.1-rc2

    Ollama fixes missing build context in CI.

    ci: fix missing build context (#18742)

    Original source
  • Oct 2, 2026
    • Date parsed from source:
      Oct 2, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Ollama logo

    Ollama

    v0.35.1

    Ollama adds support for Cloudflare's Clef and Clef Flash decision models through /v1/systemone, expands web search to ten searches per response, and introduces CAPABILITY declarations in Modelfiles. It also improves decision model handling and updates llama.cpp and MLX.

    Clef decision models

    Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone.

    Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.

    curl http://localhost:11434/v1/systemone -d '{
    "model": "clef-flash",
    "state": "The user took this screenshot.",
    "images": ["<base64-encoded image>"],
    "questions": {
    "has_ollama": {"type": "noul", "instructions": "Does this image contain Ollama?"}
    }
    }'
    
    {
    "model": "clef-flash",
    "answers": {
    "has_ollama": {
    "type": "noul",
    "noul": 0.958
    }
    },
    "usage": {
    "input_tokens": 548,
    "output_tokens": 0
    }
    }
    

    What's Changed

    Models using web search can now perform up to ten searches per response, up from three

    Modelfiles now support CAPABILITY declarations, so model creators can explicitly declare what a model can do. Declarations are preserved when creating from GGUF or safetensors, through model inheritance, and on Modelfile export

    ollama show and the model list now report only decision as the capability for decision models, so clients no longer offer them for general chat, tools, or thinking

    Updated llama.cpp and the MLX engine

    Full Changelog: v0.35.0...v0.35.1

    Original source
  • Oct 1, 2026
    • Date parsed from source:
      Oct 1, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Ollama logo

    Ollama

    v0.35.1-rc1

    Ollama adds clef support for models.

    models: add clef support (#18741)

    Original source
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Oct 2, 2026
    Ollama logo

    Ollama

    v0.35.1-rc0: create: support explicit model capabilities (#18708)

    Ollama adds CAPABILITY declarations to Modelfiles and create requests, with preserved exports and new System One scheduling rules.

    Add CAPABILITY declarations to Modelfiles and an additive capabilities field to create requests. Preserve declarations across GGUF and safetensors creation, inheritance, and Modelfile export.

    Require decision capability before scheduling System One requests instead of matching Qwen architecture/renderer metadata. Retain main's GGUF-only scoring restriction until the separate MLX runtime work lands.

    Extracted from the capability foundation in 36d46a0 on system_one_mlx; MLX scoring and manifest-list changes are intentionally separate.

    Original source
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Ollama logo

    Ollama

    v0.35.1

    Ollama adds ten web searches per response plus MLX and llama.cpp bumps and explicit model capabilities.

    What's Changed

    feat: allow ten web searches per response by @ParthSareen in #18602
    MLX: version bump by @dhiltgen in #18651
    llama.cpp: version bump b11232 by @dhiltgen in #18652
    create: support explicit model capabilities by @dhiltgen in #18708

    Full Changelog

    v0.35.0...v0.35.1-rc0

    Original source
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Ollama logo

    Ollama

    v0.35.0

    Ollama adds decision models via /v1/systemone, based on TypeSafe’s Jev API, with choice, probability and score outputs for ticket triage, routing and classification. It also speeds Settings startup and fixes macOS update, MLX download, and typical_p warning behavior.

    Decision models

    Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API.

    Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.

    Available models:

    • Nimble from Bespoke Labs
    • Tev1 from Together AI
    ollama pull nimble
    

    Send context and one or more questions:

    curl http://localhost:11434/v1/systemone \
    -H 'Content-Type: application/json' \
    -d '{
    "model": "nimble",
    "state": "Our checkout has returned 500 errors since 9am.",
    "questions": {
    "label": {
    "type": "choice",
    "instructions": "Which label fits this ticket?",
    "criteria": {
    "billing": "Payments and refunds",
    "bug": "Software errors",
    "account": "Login and account access"
    }
    }
    }
    }'
    

    Example response:

    {
    "model": "nimble",
    "answers": {
    "label": {
    "type": "choice",
    "choice": "bug",
    "probabilities": {
    "billing": 0.0125,
    "bug": 0.9781,
    "account": 0.0093
    },
    "confidence": 0.8906
    }
    },
    "usage": {
    "input_tokens": 174,
    "output_tokens": 1
    }
    }
    

    The API supports three question types:

    • choice: Select an option and return probabilities for each.
    • noul: Return the probability that a condition is true.
    • score: Return a score across an ordered set of criteria

    What's Changed

    • Settings now opens without waiting for model discovery.
    • Fixed the macOS update menu and icon not reflecting an available update at startup.
    • Fixed stalled MLX model downloads hanging indefinitely.
    • Requests containing the deprecated typical_p parameter now log a warning instead of failing.

    Full Changelog: v0.34.4...v0.35.0

    Original source
  • Sep 28, 2026
    • Date parsed from source:
      Sep 28, 2026
    • First seen by Releasebot:
      Sep 28, 2026
    Ollama logo

    Ollama

    v0.35.0-rc1

    Ollama bounds MLX pull stall retries and lets the watchdog interrupt them.

    mlx: bound pull stall retries and let the watchdog interrupt them (#1…)

    Original source
  • Sep 28, 2026
    • Date parsed from source:
      Sep 28, 2026
    • First seen by Releasebot:
      Sep 28, 2026
    Ollama logo

    Ollama

    v0.35.0-rc0

    Ollama adds System One scoring API.

    feat: add System One scoring API (#18606)

    Original source
  • Sep 25, 2026
    • Date parsed from source:
      Sep 25, 2026
    • First seen by Releasebot:
      Oct 4, 2026
    Ollama logo

    Ollama

    v0.40.0-rc0: llama-server: prepare to remove compatibility patch

    Ollama adds manifest-list storage and smarter runner-aware manifest handling, with show, copy, pull, push and remove now supporting digest selection and child manifest transfers. It also brings lazy GGUF compatibility migration, broader create and import support, and stronger validation and test coverage.

    Add manifest-list storage so runner-specific manifests can coexist under one tag while preserving existing v1 tags as best-effort downgrade anchors. Show/list/copy/remove/pull/push now understand runner and digest selection and transfer referenced child manifests and layers.

    Add lazy local compatibility migration for legacy Ollama GGUFs into llama.cpp-compatible children, covering the patched model families and preserving parser/renderer, templates, projectors, media metadata, split GGUFs, and MTP/draft tensors where applicable.

    Extend create/import/conversion paths for the same compatibility rules, harden manifest-list combine validation, and add focused manifest, migration, transfer, and integration coverage for first-load conversion plus chat/tools/vision/audio/embedding behavior.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.