Moondream Release Notes

Follow

23 release notes curated from 25 sources by the Releasebot Team. Last updated: Sep 30, 2026

Get this feed:
  • Sep 29, 2026
    • Date parsed from source:
      Sep 29, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Moondream logo

    Moondream

    400 tok/s with Qwen3.5

    Moondream releases Photon 2.6 with FP8 support and DFlash speculative decoding, delivering faster Qwen3.5-27B serving on NVIDIA B200 GPUs and outperforming vLLM at tested concurrency levels.

    Photon 2.6 runs Qwen3.5-27B at 411 tokens per second on a single NVIDIA B200, increasing to 825 tok/s with two concurrent requests. It outperforms vLLM at every concurrency we tested, from one to eight requests.

    This release adds FP8 support and DFlash speculative decoding to Photon. Under the hood, we extended our megakernel compiler to handle quantized projections and causal verification of draft tokens. Qwen3.5's recurrent layers made verification particularly interesting: when a draft is rejected, we need to recover the model state without running the expensive projections again.

    Faster on one GPU

    We ran Qwen3.5-27B-FP8 with the same DFlash checkpoint in Photon and vLLM 0.30.0, with thinking enabled in both. Photon is ahead at every concurrency we tested:

    At two concurrent requests, Photon maintains over 400 tok/s per request. The chart divides total throughput by concurrency; the table below shows total throughput.

    Concurrent requests Photon vLLM Speedup 1 411 tok/s 346 tok/s 1.19x 2 825 tok/s 671 tok/s 1.23x 4 1,259 tok/s 1,089 tok/s 1.16x 8 2,206 tok/s 1,989 tok/s 1.11x

    These numbers measure the full serving system, including prefill and host overhead, using the released packages from PyPI.

    The cost of serving at 400 tok/s

    Assume a B200 costs $6.25 per hour, and we keep two requests running at the measured throughput. At 825 tok/s total, that's 2.97 million output tokens per hour, or $2.10 in GPU rental per million output tokens, while maintaining over 400 tok/s per request.

    The cost is comparable to hosted APIs, while per-request speed is much higher than the rates reported in our September 29 snapshot of OpenRouter's Qwen3.5-27B listings: 9 to 22 tok/s across its listed providers, compared with 412 tok/s per request in our two-request benchmark.

    OpenRouter's speeds are P50 measurements of provider traffic; Photon's figure comes from the benchmark above. Provider prices as of September 29, 2026:

    Provider Input / M Output / M Reported P50 tok/s Alibaba $0.195 $1.56 21 SiliconFlow $0.25 $2.00 9 AtlasCloud $0.27 $2.16 13 Phala $0.30 $2.40 10 NovitaAI $0.30 $2.40 12 DeepInfra $0.26 $2.60 22

    Photon's GPU cost falls within that output-price range. The $2.10 estimate allocates the entire GPU rental cost, including time spent processing inputs, to output tokens. It assumes the benchmark workload runs continuously; idle time and operating costs are additional. API providers charge separately for input tokens.

    Adding FP8 to the compiler

    FP8 reduces weight storage from sixteen bits per value to eight, plus block scales. This is especially useful at low concurrency, where reading the weights can take longer than doing the arithmetic.

    We added FP8 projections to the megakernel compiler, including activation quantization for tensor-core computation. The compiler can pipeline weight loading with compute and combine normalization with quantization to avoid writing an intermediate activation back to GPU memory.

    The compiler chooses the tile geometry, too. A tile built for a large batch can leave most of the GPU idle when verifying a short sequence. Smaller tiles expose more parallel work, but change weight reuse, register pressure, and synchronization costs. Our cost model weighs those costs when selecting the execution plan.

    Speculative verification in a megakernel

    Ordinary decoding runs the target model once per new token. DFlash proposes a block of tokens with a smaller draft model, then asks the target to verify them in one causal pass. Keep the accepted prefix, discard the rejected suffix, and draft again.

    One pass through the target's weights can now produce several output tokens. That is a good fit for a megakernel, but it isn't ordinary batching: sixteen positions in one sequence depend on each other. Sixteen independent requests don't.

    The compiler now represents that causal sequence directly. At one concurrent request, Qwen3.5's target verification runs as a compiled megakernel, carrying work across operator boundaries without a CPU dispatch between each operation. Photon handles drafting, acceptance, and state commit around it.

    Handling recurrent state

    Qwen3.5 mixes full attention with Gated DeltaNet, a recurrent architecture. Rejected attention-cache positions can be hidden by shortening the visible history. A recurrent state has already incorporated their values. Changing a length counter won't undo that.

    Our approach is inspired by ReplaySSM, which caches recent inputs and reconstructs recurrent state when needed. For generated verification, we use checkpoint-and-replay: keep the state from before the candidate block, retain the projected inputs, and reconstruct the state at the accepted boundary.

    The compiler keeps Gated DeltaNet's working state in registers across the verification block. When a suffix is rejected, a replay kernel advances the checkpoint through only the accepted positions, using the retained inputs without repeating the large weight projections. Convolution history is committed to the same boundary. If the whole block is accepted, the engine uses the verified state directly.

    Each request owns its draft context and target state. Two concurrent requests can accept different numbers of tokens in the same round; the engine keeps those histories separate while batching the work they can share.

    Try it

    Photon 2.6 is available through the Moondream Python package:

    pip install --upgrade "moondream==2.6.1"
    

    To run Qwen3.5-27B with DFlash on a B200:

    import moondream as md
    from huggingface_hub import snapshot_download
    
    draft = snapshot_download("z-lab/Qwen3.5-27B-DFlash")
    
    with md.photon(
        "Qwen/Qwen3.5-27B-FP8",
        device="cuda",
        draft_model_path=draft,
        max_batch_size=1,
    ) as model:
        result = model.chat(
            [{"role": "user", "content": "Explain how binary search works."}],
            reasoning=True,
            settings={"temperature": 0, "max_tokens": 256},
        )
        print(result["message"])
    

    For concurrent Qwen3.5 requests, set max_batch_size to your desired concurrency and use page_size=64. Both configurations use the same DFlash checkpoint.

    Both FP8 execution and causal verification are implemented in the compiler, rather than a separate Qwen-specific inference path. As we add models, they can use the same optimizations.

    Benchmark setup

    Qwen3.5-27B-FP8 with z-lab/Qwen3.5-27B-DFlash on one NVIDIA B200, against vLLM 0.30.0. Thinking enabled, greedy decoding, identical prompt token IDs and draft checkpoints, prefix caching disabled. Three prompts covering explanation, code, and mathematics, with a 256-token output limit. One warmup repetition followed by three measured repetitions; throughput is total accepted output tokens divided by total request time. Reported tokens include reasoning.

    Original source
  • Sep 24, 2026
    • Date parsed from source:
      Sep 24, 2026
    • First seen by Releasebot:
      Sep 25, 2026
    Moondream logo

    Moondream

    Photon can now speak

    Moondream adds faster speech generation with Photon, cutting p95 startup latency by up to 70% and delivering first audible audio in under 85 ms in tests. It also adds Qwen3-TTS CustomVoice and Kokoro-82M, with streaming support in the same Python package.

    Photon now generates speech with up to 70% lower p95 startup latency than vLLM-Omni. In our Qwen3-TTS tests on H100 and B200, the first audible output arrived in under 85 milliseconds for 95% of requests, at six requests per second. Less waiting for your voice assistant to answer. Faster feedback for the person on the other end.

    Today we're adding Qwen3-TTS CustomVoice, in both 0.6B and 1.7B sizes, and Kokoro-82M. Generate a complete recording or stream audio as it's ready, through the same Moondream Python package you use for transcription and vision.

    We measure the first audible audio, not a chunk of silence. Photon cut that wait by 56–70% across all four comparisons.

    Give it a voice

    Install the latest Moondream Python package:

    pip install --upgrade "moondream>=2.5.0"
    

    Here's a complete Qwen3-TTS example for an NVIDIA GPU:

    import moondream as md
    
    with md.photon("Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice", device="cuda") as voice:
        result = voice.synthesize(
            text="Good morning! What would you like to build today?",
            voice="Ryan",
            language="English",
        )
        pcm = result["audio"]
        sample_rate = result["sample_rate"]
    

    The result is mono, 24 kHz floating-point PCM, ready for your audio output or file writer. Change 1.7B to 0.6B to use the smaller checkpoint. Both accept a voice and language; the 1.7B model also accepts instructions to guide delivery.

    To start playback before the whole recording is ready, set stream=True. Each update contains the next piece of audio, not the recording so far. In this example, play is your application's audio output callback:

    def speak(voice, text, play):
        stream = voice.synthesize(text=text, voice="Ryan", stream=True)
        for update in stream:
            play(update["audio"], update["sample_rate"])
        return stream.result()
    

    stream.result() returns the complete recording. Async applications can use await voice.asynthesize(...) and consume updates with async for.

    A smaller option with Kokoro

    Kokoro has 82 million parameters and runs on either CPU or CUDA. It uses the same synthesis methods, with one difference: you pass phonemes instead of text.

    Text-to-phoneme conversion decides how written words are pronounced. Keeping that step in your application lets you use your own pronunciation dictionary or language frontend, without installing a separate text-processing stack just to run the model. The input is a Unicode string in Kokoro's phoneme vocabulary, not arbitrary IPA or a list of token IDs.

    import moondream as md
    
    with md.photon("hexgrad/Kokoro-82M", device="cpu") as voice:
        result = voice.synthesize(phonemes="həlˈO", voice="af_heart")
        pcm = result["audio"]
    

    Use device="cuda" to run on an NVIDIA GPU. On a Mac, use the CPU path. Kokoro also supports streamed output.

    We're excited to see what you build with speech going in both directions. Give it a try, and let us know how it works in your application.

    Benchmark notes

    Benchmark notes. Qwen3-TTS CustomVoice on one GPU, using English Seed-TTS-Eval prompts and the Ryan voice. Poisson arrivals at six requests per second; 30 seconds of warmup and five minutes of measurement, with matched arrival seeds across engines. Baseline: vLLM-Omni 0.28.0 with tuned chunk-ramp and onset settings. Results report p95 time to first audible audio, not maximum throughput.

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Moondream and hundreds of other software products.

    Create account
  • Sep 22, 2026
    • Date parsed from source:
      Sep 22, 2026
    • First seen by Releasebot:
      Sep 22, 2026
    Moondream logo

    Moondream

    Introducing Moondream Parakeet Redux and Parakeet Ultra

    Moondream adds Parakeet Redux and Parakeet Ultra, two speech-to-text models with 25-language ASR. Redux is smaller and faster for CPUs and Macs, while Ultra improves transcription accuracy on GPUs. Both include automatic speech detection, timestamps, live audio, and file support.

    Making the Moondream VLM smaller and faster

    We've spent a lot of time making the Moondream VLM smaller and faster. A model that fits on your laptop, or costs less to run on a server, opens up more things you can build with it. The same is true for speech.

    Today we're introducing Moondream Parakeet Redux and Moondream Parakeet Ultra, two speech-to-text models based on NVIDIA's Parakeet. Both support automatic speech recognition (ASR) in 25 languages.

    Parakeet Redux compresses the model weights from 1.2 GB to 178 MB, and is designed for fast inference on CPUs and Macs. Parakeet Ultra keeps the full-precision weights and improves transcription accuracy through further training. It's built for running on GPUs.

    Making Parakeet smaller

    Running a model requires repeatedly moving its weights from memory to the processor. On CPUs and Apple Silicon, memory bandwidth and not computation is often the limiter of inference speed. Smaller weights reduce that traffic, so compressing the model can make transcription faster as well as reduce its memory footprint.

    For Parakeet Redux, we reduced the encoder's weights to just three possible values: −1, 0, and +1. This is called ternary quantization. Those weights take much less space to store, and Photon, our inference engine, has specialized kernels that compute directly from the packed representation.

    Here's how fast Parakeet Redux runs in Photon:

    Hardware Audio processed per second Apple M2 CPU, MacBook Air 38 seconds Apple M2 GPU, MacBook Air 43 seconds AMD EPYC 9575F CPU, 8 cores 113 seconds

    On the AMD CPU, Parakeet Redux delivers 2.5 times the throughput of parakeet.cpp, the fastest alternative we measured on the same machine. See the model card for comparisons with other engines and full test details.

    One user has already moved transcription to the CPU to free up GPU memory for LLMs. In their tests, Parakeet Redux on an AMD Ryzen 9 9950X3D CPU beat 8-bit parakeet.cpp on an NVIDIA RTX PRO 6000 GPU for shorter clips and nearly matched it for longer ones.

    Of course, making a model this small comes with a question: how much accuracy do you lose?

    We measure accuracy using word error rate, which counts missed, incorrect, and extra words against a reference transcript. Lower is better.

    On our seven English test sets, average word error rate rises from 6.26% to 6.55%. Across the 25-language FLEURS evaluation, Parakeet Redux actually improves the average, from 11.62% to 10.56%. The biggest tradeoff is background noise, where Parakeet Redux makes more mistakes than the original. We'd use Parakeet Ultra when noise is a concern and a GPU is available.

    Making Parakeet more accurate

    With Parakeet Ultra, we kept the original model's size and focused on improving its transcriptions. Post-training of the model brought down the average error rate in every evaluation group we tested, including English, multilingual speech, business recordings, and speech mixed with background noise.

    Here's how the two models compare with the original, measured by word error rate:

    Evaluation Original Parakeet Parakeet Redux Parakeet Ultra English, seven test sets 6.26% 6.55% 5.80% FLEURS, 25 languages 11.62% 10.56% 9.55% Business speech 6.15% 6.96% 5.79% Background noise 6.72% 9.04% 5.82% Long recordings, 11 TED-LIUM talks 2.71% 2.51% 1.94%

    Parakeet Ultra's multilingual result represents 18% lower average word error rate than the original on that evaluation.

    Knowing where to pause

    Transcribing a long recording means breaking it into smaller pieces. Where you make those cuts matters: splitting in the middle of a word can make it harder to get the transcription right.

    We built voice activity detection (VAD) into both models. It identifies where speech is happening, and Photon uses it to find pauses where it can split the recording into chunks of at most 30 seconds. You can pass in the whole recording without setting up a separate VAD model or cutting the audio yourself.

    On our test of 11 complete TED-LIUM talks, each 10–20 minutes long, Parakeet Redux brings word error rate down from the original's 2.71% to 2.51%. Parakeet Ultra gets it down to 1.94%, a 28% reduction.

    Give them a try

    Both models support files, live audio, and timestamps through the Moondream Python package:

    pip install --upgrade "moondream>=2.4.1"

    Transcribe a recording

    Here's all you need to run Parakeet Redux on your CPU. WAV, MP3, FLAC, and M4A files are supported, among other formats. Long recordings are split automatically.

    import moondream as md
    with md.photon("moondream/parakeet-redux", device="cpu") as speech:
        result = speech.transcribe(audio="meeting.m4a")
    print(result["text"])
    

    For the Apple GPU, use device="mps". The same examples work with Parakeet Ultra on an NVIDIA GPU by changing the model to "moondream/parakeet-ultra" and the device to "cuda". Both models choose the language automatically and preserve punctuation and capitalization.

    Find when each word was spoken

    Use timestamps="word" to get a transcript with sentence segments and word timings. All times are in seconds.

    import moondream as md
    with md.photon("moondream/parakeet-redux", device="cpu") as speech:
        result = speech.transcribe(audio="interview.wav", timestamps="word")
    for segment in result["segments"]:
        print(segment["start"], segment["end"], segment["text"])
        for word in segment["words"]:
            print(word["start"], word["end"], word["word"])
    

    Use timestamps="segment" for sentence timings alone, or "none" if you only need the text.

    Read the transcript as it develops

    You don't have to wait for a whole recording to finish. With stream=True, Photon returns updated transcripts as it works through the audio:

    import moondream as md
    with md.photon("moondream/parakeet-redux", device="cpu") as speech:
        updates = speech.transcribe(audio="meeting.m4a", timestamps="segment", stream=True)
        for update in updates:
            # Each update replaces the previous transcript; don't append it.
            print(update["text"])
        final = updates.result()
    print("Final transcript:", final["text"])
    

    Transcribe live audio

    You can also feed audio in as it arrives from a microphone or a network stream. Pass an asynchronous iterator of mono audio chunks to atranscribe, along with their sample rate. Each chunk should be a nonempty, one-dimensional NumPy array or CPU Torch tensor containing raw PCM samples.

    This function connects your application's audio source to the transcriber:

    import moondream as md
    async def transcribe_live(audio_chunks, sample_rate):
        with md.photon("moondream/parakeet-redux", device="cpu") as speech:
            updates = await speech.atranscribe(audio=audio_chunks, sample_rate=sample_rate, timestamps="word", stream=True)
            async for update in updates:
                print(update["text"])
            final = await updates.aresult()
        return final
    

    Call it with await transcribe_live(audio_chunks, sample_rate=48_000) for a source producing 48 kHz audio. End the iterator when the recording stops to receive the final transcript. Live previews begin after four seconds of audio and are scheduled every two seconds after that; earlier text can change as more context arrives. As with files, each update is a replacement snapshot.

    Work with part of a file or audio already in memory

    You can select a time range without editing the source file, or pass encoded audio bytes directly:

    from pathlib import Path
    import moondream as md
    with md.photon("moondream/parakeet-redux", device="cpu") as speech:
        # Transcribe the second minute of a recording.
        clip = speech.transcribe(audio="interview.mp3", clip_start_seconds=60, clip_end_seconds=120)
        print(clip["text"])
        # Encoded audio bytes, such as a file uploaded to your application.
        result = speech.transcribe(audio=Path("speech.wav").read_bytes())
        print(result["text"])
    

    For raw audio already in memory, pass a one-dimensional mono NumPy array or CPU Torch tensor as audio and supply sample_rate, for example speech.transcribe(audio=samples, sample_rate=48_000). Encoded files and bytes carry their own sample rate. Clip ranges apply to recorded audio, not live streams.

    Weights are available under the CC-BY-4.0 license, same as NVIDIA's original model. The Parakeet Redux and Parakeet Ultra model cards have more detailed benchmarks and examples. Give them a try, and let us know what you're building.

    Happy Moondreamin'.

    Original source
  • Sep 9, 2026
    • Date parsed from source:
      Sep 9, 2026
    • First seen by Releasebot:
      Sep 10, 2026
    Moondream logo

    Moondream

    Photon 2.2: Five New NVIDIA GPUs, Ampere to Blackwell

    Moondream releases Photon 2.2 with support for five more NVIDIA GPUs, expanding coverage to seven GPUs across Ampere, Ada, Hopper, and Blackwell. It also brings faster speech, lower latency, and higher throughput for Whisper, Qwen3-ASR, Parakeet, and Moondream models.

    Photon 2.2 adds five NVIDIA GPUs: A10/A10G, A100, RTX 3090, L4, and RTX PRO 6000 Blackwell. With H100 and B200, Photon now runs on seven GPUs across four architectures: Ampere, Ada, Hopper, and Blackwell.

    We benchmarked speech on L4, A10/A10G, and A100 at concurrency 1, 2, 4, and 8. Photon beats the reference engine in 35 of 36 configurations and ties in one.

    • Parakeet TDT 0.6B v3: 3.1x to 4.8x the throughput of NVIDIA NeMo
    • Qwen3-ASR 0.6B: 1.1x to 2.2x the throughput of qwen-asr on vLLM
    • Qwen3-ASR 1.7B: 1.0x to 1.2x the throughput of qwen-asr on vLLM

    Vision, language, and Whisper numbers on the new GPUs will follow.

    Also in 2.2

    • Whisper transcription now runs on L4 and RTX 3090.
    • Lower transcription latency and less host overhead for Qwen3-ASR and Parakeet. The largest gains are at concurrency 1.
    • Lower latency and higher throughput for Moondream and Qwen models.

    Speech benchmarks

    Parakeet TDT 0.6B v3: Photon throughput vs. NVIDIA NeMo

    Higher is better. 1.0x = NeMo.

    Photon throughput divided by reference-engine throughput. Higher is better. C is concurrent requests.

    Reference engines: NVIDIA NeMo for Parakeet, qwen-asr on vLLM for Qwen3-ASR.

    Parakeet gains are large on all three GPUs. Qwen3-ASR gains are smaller. On smaller GPUs, some workloads already saturate compute or memory bandwidth, which leaves less launch and scheduling overhead for Photon to remove.

    The compiler keeps compounding

    Photon compiles each model into one GPU program, a megakernel. All model execution happens inside it.

    When we started, getting one megakernel to compile at all was hard. Making it beat other engines was harder. But each compiler improvement applies to every model and every GPU it targets, so coverage compounds:

    • Photon 2.0, August 3: Qwen, Gemma, and Moondream on H100. That took months.
    • Photon 2.1, August 26: B200, plus Whisper, Qwen3-ASR, and Parakeet. Three weeks.
    • Photon 2.2, September 9: five more GPUs. Two weeks.

    Try Photon 2.2

    pip install --upgrade moondream
    

    Photon is free. Local inference needs no API key.

    Have a model, GPU, or realtime workload that needs to be faster? Email [email protected] with the model, the GPU, and the latency or throughput you need. We will compile it on Photon and send you numbers.

    Original source
  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 3, 2026
    Moondream logo

    Moondream

    Photon 2.1: Speech Recognition on H100 and B200, Up to 3.1× Faster

    Moondream adds streaming speech recognition to Photon with Whisper, Qwen3-ASR, and Parakeet, plus full NVIDIA B200 support across its model catalog. The release brings faster live transcription, segment timestamps, and broader Blackwell readiness for latency-sensitive audio apps.

    Photon can now hear: streaming ASR for Whisper, Qwen3-ASR, and Parakeet, plus B200 support across every Photon model

    Photon can now hear.

    Photon 2.1 adds automatic speech recognition for Whisper large-v3-turbo, Qwen3-ASR 0.6B and 1.7B, and Parakeet TDT 0.6B v3, with performance beating competitors by up to 3.1×.
    It also adds NVIDIA B200 support across every Photon model. Under the hood, this release extends Photon's compiler and megakernel architecture to speech recognition and NVIDIA Blackwell.

    Speech recognition

    Photon now transcribes both long-form audio files and live PCM streams. Transcripts stream back as the audio arrives, making Photon suitable for live agents, meeting transcription, voice interfaces, and other latency-sensitive applications.

    Progressive results include segment timestamps. Depending on the model, Photon also supports word timestamps, language detection, and prompts for domain-specific vocabulary.

    Photon is built for latency-sensitive workloads. We tested all four ASR models at concurrency 1 and 8 on both H100 and B200. Photon won all 16 configurations. Speedups ranged from 1.2× for Qwen3-ASR 1.7B against Qwen's qwen-asr stack on H100 at concurrency 1, to 3.1× for Whisper against vLLM on B200 at concurrency 1. Full results are on the benchmarks page.

    Built for B200

    Photon 2.1 brings the full model catalog to NVIDIA B200. Across every tested model and batch sizes 1, 2, 4, and 8, Photon beat vLLM and SGLang in 51 of 52 matched tests, by up to 2.8×.

    The work behind B200 support also lays the foundation for more Blackwell GPUs.

    Why it's faster

    Photon is built differently from other inference engines. It uses a custom compiler that generates optimized GPU megakernels for each model and chip. Megakernels fuse all GPU operations into a single call, instead of dispatching long chains of separate kernels. This reduces launch overhead and data movement, especially at the low batch counts common in live audio.

    How to get it

    Photon 2.1 is available today. Install or upgrade with pip install --upgrade moondream. Docs are at docs.moondream.ai. Happy Moondreaming.

    Original source
  • Similar to Moondream with recent updates:

  • Aug 3, 2026
    • Date parsed from source:
      Aug 3, 2026
    • First seen by Releasebot:
      Aug 5, 2026
    Moondream logo

    Moondream

    Photon 2.0: Inference engine for Physical AI

    Moondream launches Photon 2.0, a new inference compiler with megakernels for Moondream, Qwen, and Gemma on NVIDIA H100. It promises faster throughput, quicker cold starts, and a production-ready path for physical AI workloads.

    Compiled megakernels for Moondream, Qwen, and Gemma that outperform vLLM and SGLang on NVIDIA H100

    Modern AI inference infrastructure is built for chat. Engines like vLLM and SGLang are really good at queuing prompts, in large batches, and maximize tokens per second across GPU fleets. That's exactly what agentic workloads need.

    Physical AI is different. Robots, cameras, industrial systems, and computer-use agents continuously ingest images, audio, and video. They run with low concurrency, and will sacrifice throughput for quicker responses. They worry about strict latency deadlines, limited memory and several models sharing a GPU. Serving live perception is a different problem. It needs a different inference engine.

    Today, we’re releasing Photon 2.0, our next step toward building the inference stack for physical AI. Photon began as the fastest inference engine for Moondream. Photon 2.0 runs Moondream, Qwen 3.5, and Gemma 4 on NVIDIA H100. In matched benchmarks, Photon outperformed vLLM and SGLang in every throughput test and brought every model online faster.

    We built Photon 2.0 for a different purpose, and we took a very different approach. Let's explain why.

    Hand-tuning doesn't scale

    Modern inference stacks are built from large libraries of hand-written and hand-tuned GPU kernels. The results are remarkable. With contributions from hundreds of developers, vLLM and SGLang have pushed open inference forward at an extraordinary pace. But for Physical AI, the combinations are enormous:

    Every model × every chip × every deployment objective

    Physical AI solutions often use numerous models, and run on a wide set of chips (edge, on-prem, cloud). The objectives often are different too. One deployment needs maximum throughput,another needs the lowest possible p99 latency, or another needs to fit three specific models into limited memory, or with a strict power budget. Writing efficient GPU inference code is hard, and hand-optimizing all of these combinations is practically impossible. Thankfully there's a better approach: a compiler.

    We took the knowledge we had from building Photon for Moondream and turned it into a generalized inference compiler. We added support for Qwen and Gemma to prove the compiler generalized, and the results were amazing. This approach is exciting because it opens the door to so many possibilities. This launch is just the first step.

    Megakernels

    A key technique behind Photon 2.0's performance is called a 'Megakernel'. We all know inference happens on a GPU, but in practice, it's typically done by having the CPU manage a large, tightly coordinated set of GPU pgorams that execute the inference path. There's a lot of 'chattiness' between the CPU and GPU, with inefficient memory bandwith between the two.

    Our inference compiler produces a single 'megakernel': one large set of instructions that runs the entire inference on the GPU alone. This reduces launch and synchronization overhead. More importantly, it gives the compiler visibility across operator boundaries, where many of the best optimization opportunities live.

    The compiler is winning

    Photon 2.0 launches with support for:

    • Moondream 2 and 3
    • Qwen3.5/Qwen3.6 0.8B, 2B, 4B, and 9B
    • Gemma 4 E2B and E4B

    The initial hardware supported is NVIDIA H100. We tested Photon, vLLM, and SGLang using the ChartQA benchmark. Photon came out ahead on every batch count (1, 2, 4 and 8 tested).

    These numbers measure the full serving system, not isolated kernels, with each engine given the same models and the same request streams. The result we care most about: one compiler produced these speedups across Moondream, Qwen, and Gemma, without per-model hand-tuning.

    Models ready sooner

    Another area where Physical AI workloads differ: cold-starts. This matters during autoscaling, process recovery, software updates, development, and power cycling. Systems that run outside of large, permanently warm GPU clusters. In our matched cold-start tests, Photon beat vLLM and SGLang on all models supported (note: SGLang doesn't support Moondream).

    Available today

    Photon 2.0 is ready for production workloads today. It's available through the Moondream package:

    pip install moondream
    

    For example:

    import moondream as md
    from PIL import Image
    model = md.photon("Qwen/Qwen3.5-4B")
    
    # Ask a question
    image = Image.open("photo.jpg")
    answer = model.query(image, "What's in this image?")["answer"]
    print("Answer:", answer)
    

    Check out the docs. The Photon inference engine is Apache 2.0, and the megakernels are free to run. Our compiler itself is proprietary.

    Roadmap

    Photon 2.0 currently supports a narrow set of models, chips, and deployment objectives. We are already expanding that matrix, with new updates every week. Follow @moondreamai on X for the latest, or reach out with specific requests. We are especially looking for design partners building:

    • Live video understanding
    • Robotics
    • Industrial perception
    • Computer-use agents

    Inference is becoming a compiler problem

    Hand-written kernels aren't going away. Assembly never disappeared either. But assembly stopped being the way most software gets built. Inference is entering its compiler era.

    Original source
  • Jul 7, 2026
    • Date parsed from source:
      Jul 7, 2026
    • First seen by Releasebot:
      Jul 8, 2026
    Moondream logo

    Moondream

    Moondream 3.1: Beyond Benchmarks

    Moondream launches 3.1 with stronger benchmark results, faster real-time vision performance, a new Cloudflare Workers AI partnership, simpler licensing, and lower cloud pricing. It also expands support across Lens, Photon, Hugging Face, and Moondream Cloud for easier deployment.

    Today we're excited to announce the launch of Moondream 3.1 and a partnership with Cloudflare. The benchmarks on this new model are strong. Many are best-in-class, and we'll get to them. But your use cases aren't benchmarks. This launch is about making Moondream the best model for what you need.

    In this post:

    • Benchmarks
    • Our partnership with Cloudflare
    • New license
    • New cloud pricing

    But first, obligatory benchmarks

    All figures below are like-for-like against the same baselines under the same settings: Qwen3.5 9B, SAM 3, LocateAnything, and Gemma 4 12B. Detection metrics are [email protected]. The bold value marks the best score in each chart. An "n/a" means the model cannot do that task. SAM 3 and LocateAnything are detection/segmentation specialists, not general VLMs. Setup details in the benchmark methodology footer.

    State of the art on COCO and ODinW-13, on DOTA-v2, on ChartQA and PixMo-Count. But look at the dense detection numbers. Crowded shelves, crowds, bins of parts: that's where 3.1 pulls the furthest ahead, both over other models and over the last Moondream. It's the biggest jump in this release. Full methodology in the footer. The scores are good. How we got them is the better story.

    What benchmarks don't measure

    A benchmark is a fixed set of questions and images that someone else assembled, scored against answers they chose in advance. That makes it good for one thing: comparing models on equal footing, which is why we publish ours.

    But your vision tasks never look like the benchmarks. Every customer who builds on Moondream brings a task specific to them, and they're delightfully different, even weird:

    A factory points a camera at a weld seam and asks: is this weld clean? The answer that matters isn't a paragraph. It's "No. Porosity at top edge."

    A warehouse runs
    detect: crushed pallets
    across aisle footage and gets a bounding box on the one pallet to pull before it takes a rack down.

    A broadcast team runs
    point: the ball
    on a live feed and gets a coordinate that tracks the ball and the player on it, frame after frame, so the camera operator doesn't have to.

    There is no WeldBench. So with 3.1 we spent less effort chasing benchmark numbers and more on the workflow that adapts Moondream to a specific task.

    How we made it

    No special research pipeline produced these numbers. We used
    Lens
    , our fine-tuning product: the same hosted API customers use to tune Moondream with SFT and RL.

    We treated each benchmark like a customer task. We used Lens to improve Moondream 3.1 on each task, using the same API our customers do. These fine-tuned LoRAs achieved state-of-the-art results on the key benchmarks listed above.

    Then one extra step, the only one a customer doesn't need: on-policy distillation, which folds what every LoRA learned back into the base weights. That's why the downloaded model scores at the top with no tuning at all.

    The distilled base lands just behind the individual LoRAs, so if you only care about one task you can stop early: train the LoRA on Lens and ship it.

    These numbers came out of the same workflow we ship to customers. We just ran it on public benchmarks instead of private data, and every step is available to you in Lens.

    The right answer, on time

    On a live feed, latency is a hard limit. A camera watching a checkout line, a weld seam, or a loading dock can't wait several seconds for each result, so speed counts as much as accuracy. Here's how Moondream compares on throughput against the models benchmarked above:

    The vertical axis averages each model's scores on the benchmarks above. The horizontal axis is throughput: average requests (inferences) per second, measured across our full benchmark suite on the same setup. You want the top-right corner: high score, high throughput. Moondream is the only model there, at 34 requests per second. That's seven times faster than the next fastest model here, and over twenty times faster than the closest model on accuracy.

    Part of that is the architecture: 2B active parameters means less work per frame. Part of it is
    Photon
    , our free inference engine, which squeezes real-time vision out of whatever hardware you point it at. Together they put a state-of-the-art vision model on a live feed within reach, without standing up a dedicated infrastructure project.

    Announcing our partnership with Cloudflare

    Moondream already runs on-prem, on desktops, and in our cloud (Fal too!). Today we're adding one more option, and it's the second announcement of the day: we're partnering with Cloudflare to bring Moondream 3.1 to
    Workers AI
    , their serverless global inference platform, as
    @cf/moondream/moondream3.1-9B-A2B
    .

    Vision workloads are large and latency-sensitive. Frames arrive constantly, and every millisecond of round trip to a distant data center eats the speed advantage of running a small model in the first place. So take the fastest model (Moondream) and run it on the network with the most edge presence (Cloudflare). Requests run close to your users and your cameras, and responses stream: in Cloudflare's testing, first tokens came back in roughly 20 to 30 ms, with point at ~145 ms and detect at ~160 ms end to end on a simple image.

    That's fast enough to call the model inline while handling a request instead of pushing work to a background queue: moderate an upload before it's stored, drive a live overlay from a video frame, pull fields from a document during a form submission, or let an agent read a screenshot and pick its next step in a single turn.

    Workers AI supports query, caption, point, and detect, and Cloudflare offers Moondream at the same price as Moondream Cloud: $0.30 per million input tokens and $1.00 per million output tokens.

    New license

    We've adopted a new license for Moondream 3.1: the
    Moondream Model License
    . It's the same underlying theme as before, but hopefully simpler to understand. The weights are open and source-available: use them commercially, self-host them, fine-tune, quantize, and redistribute them, no strings on any of that. The single restriction is that you can't turn around and offer general-purpose Moondream inference or fine-tuning as a hosted service to others. A narrow, domain-specific app built on Moondream is fine. In short: build whatever you want on it, just don't resell Moondream itself as an API.

    New cloud pricing

    Moondream's inference efficiency keeps going up, and we're passing the savings on to you. Cloud pricing for Moondream 3.1 drops from $2.50 per million output tokens to $1.00 per million. Input tokens stay at $0.30 per million. The Batch API keeps its 50% discount on top of that.

    Ready when you are

    Moondream 3.1 is live everywhere, today. Try it in the
    playground
    , download the weights on
    Hugging Face
    , make your first call on
    Moondream Cloud
    , or run it at the edge on
    Workers AI
    .

    And it launches with full support across the platform:
    Lens
    can fine-tune it on your task right now, and
    Photon
    runs it on whatever hardware you have, from a rack of H100s to the laptop in your bag.

    Moondream 3.1. State of the art on the benchmarks. State of the art on your tasks.

    Benchmark methodology

    Models were evaluated on a single H100 SXM5 with a batch size of 16. The highest accuracy and reasoning settings were chosen for each model. LocateAnything was evaluated in slow mode, maximizing accuracy. vLLM was used to evaluate LocateAnything, Qwen3.5, and Gemma 4. Photon was used to evaluate Moondream. Hugging Face Transformers was used to evaluate SAM 3.

    Original source
  • Jun 8, 2026
    • Date parsed from source:
      Jun 8, 2026
    • First seen by Releasebot:
      Jun 9, 2026
    Moondream logo

    Moondream

    Photon is now free

    Moondream ships a faster Photon release that is now completely free for local use, with no API key required. It boosts inference on Windows, Mac, and NVIDIA GPUs, speeds up finetunes on more hardware, and fixes a small accuracy issue on older cards.

    In software there is an old adage: good, fast, cheap, pick two.

    For Photon 1.3.0, we did not get the memo. This release is faster, it fixes bugs, and Photon is now completely free. We picked all three for you.

    Starting with version 1.3.0, running Moondream locally with Photon is totally free. No API key required, just download and start calling it. You will still need to pass your MOONDREAM_API_KEY if you want to run a finetuned model, or if you want telemetry into your inference activity. These are still free, the API key is just there to link it to your account.

    If you have been waiting to try Moondream in production, or wondering about inference costs at the edge or on-prem, now is a perfect time to get started.

    Moondream is faster

    Photon runs on Windows, Mac, and NVIDIA GPUs. This release makes Moondream faster across all of them.

    With NVIDIA GPUs, the biggest gains are seen on older cards. On an A100, you can expect roughly 25 to 44 percent more throughput on standard queries, and up to about 70 percent more when the model reasons step by step. Answers also come back sooner, with latency down about 30%. A10 cards see similar gains, roughly 30 to 45 percent. Jetson Thor is up to 50% faster at low batch sizes. Newer hardware is faster as well. Check out our performance results for more details.

    What this means in practice: the same GPU now handles more images per second and returns each answer sooner. Photon with Moondream does more per machine, so you can size down your hardware or push more throughput on what you have.

    Decoding is also faster on Apple Silicon Macs, so local development on a laptop feels quicker.

    Finetunes run faster, on more hardware

    Lens, our finetune service, lets you customize Moondream for your own tasks. Want to teach Moondream to detect defective welds, or determine your own brand compliance — the possibilities are endless. This release makes these finetunes run much faster and available on far more machines.

    The overhead of using a finetune dropped sharply. A large finetune that previously added about 140 milliseconds per request now adds under 1 millisecond. You only pay for the size of the finetune you actually use, not the maximum the engine supports.

    Finetunes are now supported on Apple Silicon and Windows, in addition to NVIDIA (Windows and Mac could not run them at all before). If your team works on Macs or Windows machines, you can now use customized Moondream models there.

    An accuracy fix for older GPUs

    We fixed a small accuracy issue that affected a few older GPUs, including the A100, A10, and RTX 3090. On these cards, the math that prepared data for the model was rounding the wrong way and pushing values slightly low. Output is now slightly more accurate on these cards. The effect was minor, and newer GPUs were never affected.

    How to get it

    To install, pip install moondream. Docs are at docs.moondream.ai. Happy Moondreaming.

    Original source
  • Jun 4, 2026
    • Date parsed from source:
      Jun 4, 2026
    • First seen by Releasebot:
      Jun 5, 2026
    Moondream logo

    Moondream

    Popping the GPU Bubble

    Moondream's Photon inference engine now achieves near-realtime VLM inference, with up to 35% higher decode throughput by hiding GPU bubbles through pipelined decoding, ping-pong slots, safer constrained decoding, and cleaner request teardown.

    Photon, Moondream's inference engine, achieves near-realtime VLM inference (~33ms on NVIDIA B200). This is a peek into how it delivers up to 35% higher decode throughput by optimizing how the GPU works.

    The bubble

    How do you make an AI model run as fast as possible? This is a question we obsess over at Moondream HQ. The GPU handles all the math involved in model inference, so at first glance it doesn't seem like there's much to it: just tell it what to do and wait for the answer. But if you start looking at how it actually works under the hood, you find that the GPU often sits idle, not for lack of work, but because the CPU hasn't told it what to do next yet. This phenomenon is called a GPU bubble.

    When a typical AI model generates text, it produces one token at a time (a token is a chunk of text, roughly a few characters). Each token depends on the tokens before it, a property called autoregressive, so generation is sequential. You can't compute the third token before you have the second. This decode loop involves a round trip between the CPU and GPU. The GPU does most of the heavy lifting to run the actual model, performing billions of arithmetic operations to produce the next token. But there's also a surprising amount of work done by the CPU. It selects which requests to run next, sets up the metadata the GPU needs for them, picks the actual token out of the model's output and records it, and more.

    The challenge is that one token's worth of GPU work is small, while the CPU housekeeping is a fixed cost paid on every trip. If the GPU has to wait for that housekeeping before it can start the next token, it sits idle for part of every loop. This is why we get GPU bubbles.

    In this post we're going to dive into how Photon hides these bubbles using a technique called pipelined decoding. The idea is to overlap the two kinds of work: we start GPU work on the next token while the CPU is still finishing the last one.

    Here's the shape of the problem.

    In the blocking version (top), every step is a baton pass. The CPU plans and launches a forward, the GPU runs it, then the CPU synchronizes, waits for the results to land, commits them, and only then starts planning the next step. This is because the plan depends on the token we select. For example, if the model indicates it has finished answering, then we need to schedule a new pending request from our queue. The GPU sits idle waiting for the CPU to finish its commit-plan-launch work.

    The fix is to pipeline the loop. Launch the next forward while the current step's token is still coming back and being committed. That's the pipelined version (bottom): the forwards run back-to-back, and the CPU work is overlapped underneath them.

    The reason we can is that the token we just sampled doesn't have to leave the GPU. The next forward reads it straight from GPU memory as its input. We still want a copy on the CPU eventually, to detokenize it, stream it, and decide whether the request is done, but that is bookkeeping we can do a moment later, in the background, while the next forward already runs. Not waiting on that copy is the move that removes the bubble.

    Making it safe requires three things, that we cover in the rest of this post: keeping step buffers from colliding (ping-pong slots), getting the sampling order right for constrained decoding (forward now, sample later), and cleaning up after a request finishes (zombies).

    Mechanism 1: ping-pong slots

    To run a decode step, the GPU needs a working set of buffers: a place to stage the input (the last generated token and its position in the sequence), a place for the model to write its output (the logits, one score per word in the vocabulary), a place to land the sampled token, and some bookkeeping the attention kernel needs to find each sequence's cached keys and values (its KV cache). We keep pinned (page-locked) host buffers on both ends, so the copies on and off the GPU run as background DMA (direct memory access) transfers instead of blocking the CPU.

    These buffers are allocated once and reused on every step. We work hard to avoid performing GPU memory allocations at runtime, because they can cause device synchronization and introduce bubbles. Fixed buffer addresses are also needed for capturing the decode step once as a CUDA graph and replaying it, reducing kernel launch overhead. We call this bundle a DecodeSlot.

    This works, but introduces a blocker for pipelining. The buffers stay in use until the step is done, so we cannot start the next step until the current one finishes. To overlap two steps, the second step needs its own working set, otherwise it can overwrite the results of the first step before the CPU has read them. So we keep two slots and alternate between them, ping-pong style.

    One thing to note about launch: we don't execute kernels the instant we issue a launch from CPU. Instead, we enqueue them onto a stream -- an ordered queue that the GPU drains in order. Work on the same stream runs sequentially, while work on separate streams can overlap. Both slots put their forwards onto the same compute stream. The slots are not for GPU parallelism. They only exist so the CPU can process one slot's results while the GPU runs the other slot's forward.

    The forwards all share that one compute stream, but the copies do not. Each step's device-to-host copy, the one that brings the sampled token back for bookkeeping, goes on a separate copy stream, so it can run while the GPU is busy with the next forward. That is what lets us not wait for it. We anchor the copy to an event recorded the instant the step's outputs are written, so it waits on exactly that step's work and nothing queued behind it.

    A slot only becomes free once its results have been read, not just once the GPU is done with it. Its pinned host buffer is the landing site for a copy that may still be in flight, so handing the slot to a new step too early would overwrite a copy mid-transfer, creating a hard-to-debug corruption bug. So the slot stays reserved through the commit that reads it, and is released only once that commit has finished.

    Mechanism 2: forward now, sample later

    The next forward can run ahead because it doesn't depend on anything the CPU does with the last token. But two things about the next step do depend on the last step's committed result. One is which sequences are still in the batch: if a request just finished, it shouldn't be in the next forward. That is the next section (zombies). The other is what tokens the next step is even allowed to sample, and that one is this section.

    It comes from constrained decoding. Moondream's spatial skills return structured output instead of free text: point returns a coordinate, detect returns boxes, segment returns an outline. We get those from the same decode loop by restricting which tokens the model may produce at each step: we force the scores (the logits) of the disallowed ones to negative infinity before we sample. A point step has to emit a coordinate, a detect request walks an x, y, size cycle, and so on. Which tokens are allowed, the mask, depends on what has been produced so far, so the mask for step t+1 depends on the token we sampled at t.

    The dependency is in sampling, not in the forward.

    Each scheduler tick goes through three phases: launch, commit, and finalize:

    1. Launch the forward for t+1. It doesn't depend on the mask, so it goes immediately.
    2. Commit step t: wait on the in-flight copy and advance the request's decode state. That is needed to decide the mask for t+1.
    3. Finalize sampling for t+1: with the state current, build the mask and sample.

    Sampling t+1 lands after committing t because the commit is what makes t+1's mask correct. We call this "commit-before-finalize" ordering. The GPU runs the t+1 forward through steps 2 and 3, so the commit disappears from the critical path.

    For plain text there is no mask, so forward and sampling can both run a step ahead. For constrained sequences the forward still runs ahead, but sampling waits on the previous commit, which caps how far ahead we get with no special-casing. One loop handles both.

    Mechanism 3: zombies: finalize early, release late

    Back in forward now, sample later we flagged two ways the next step depends on the last step's committed result. The sampling mask was one. Batch membership is the other, and it takes a bit of care to handle right.

    To launch step t+1 we first decide its batch, which sequences are in it, and we do that before committing step t. So what happens when a sequence hits its stop token at t, but is already baked into t+1's forward? You can't un-launch GPU work. The sequence is finished, yet still physically present in a batch that's executing.

    Photon calls these zombies, and instead of bolting on cancellation logic, it lets the behavior emerge from two per-sequence fields:

    • finalized: True after the sequence has hit EOS or its length cap.
    • inflight_refs: the number of in-flight steps that still reference this sequence (0, 1, or 2).

    When step t commits and detects EOS, the sequence is marked finalized and its result is emitted — but it isn't torn down, because inflight_refs is still nonzero (step t+1 references it). At step t+1's commit, the sequence is already finalized, so the commit is skipped: no token is appended, no state mutates. The zombie was harmlessly along for the ride — it occupied its slot and wrote some KV that nobody will read. Only when inflight_refs finally hits 0 are its KV pages and LoRA slot released.

    This finalize-early, release-late dance is a small amount of refcounting that replaces what would otherwise be a thicket of "cancel this row mid-flight" special cases.

    Prefill rides the same pipeline

    So far this has all been about decode steps, but a real serving loop is constantly doing two different kinds of work: prefill (processing a new request's prompt + image, the expensive one-shot forward over many tokens) and decode (one token at a time for everyone already running).

    Photon doesn't separate them. A prefill is just another kind="prefill" launch in the same two-slot pipeline. Because the pipeline only cares that a slot is free, not what kind of work last used it, a prefill forward can be launched into one slot while a decode step from the other slot is still being committed, and vice versa. The expensive prefill forward runs on the GPU while the CPU commits decode results; the next decode forward runs while the CPU finishes admitting the just-prefilled request. The same commit ordering (and the same inflight_refs bookkeeping) keeps everything correct across the two kinds, so none of the zombie or constrained-decode logic needs a special case for "what if a prefill is in flight."

    This matters most when outputs are short. A request that emits three tokens spends almost all of its life in prefill and admission, not decode, so a workload of many short requests is really a stream of prefills with a little decode sprinkled in. Sharing one pipeline is what lets that stream overlap its own CPU bookkeeping instead of serializing prefill behind decode and back again.

    A cost model for the bubble

    How much should pipelining actually buy you? You can predict it from the parts of a decode step, and then check the prediction against measurement.

    A decode step is three pieces of work:

    • forward: the heavy GPU matmuls. At decode this is memory-bandwidth bound: every token streams the whole weight set through the cores, so it has a floor near weight_bytes / memory_bandwidth. It shrinks as memory gets faster or as the model gets smaller.
    • sampling: turning the scores into a committed token: the constrained-decode mask, the argmax/sample, the spatial (grounding) decode, and the device→host copy of the result. All GPU work.
    • bookkeeping: the CPU around it. Choose the next batch (plan), launch the graph (launch), commit the previous step (commit).

    A blocking loop runs the three in series, so the GPU sits idle through the bookkeeping — that idle is the bubble. Pipelining slides the bookkeeping of one step underneath the forward + sampling of the next, so the period collapses toward forward + sampling and the bubble disappears. Measured per step, pipelined, that's exactly what we see — the GPU is busy for essentially the whole period (steady-state medians, moondream2, ms):

    forward + sampling ≈ period; the leftover GPU idle is under 0.05 ms. So what was hiding it worth? It comes down to a tug-of-war between two things — how much of a step you manage to tuck away, against a small penalty for running ahead:

    speedup = T_block / T_pipe × (1 − z)
    bubble hidden zombie tax

    Two symbols, two ideas. The first term is the win, and it's the whole GPU-speed story: how long a step takes blocking (T_block) over how long it takes pipelined (T_pipe) — i.e. how much faster the step runs once the bookkeeping is tucked underneath it.

    The second, z, is the price of running ahead — the zombie tax from Mechanism 3. Launch step t+1 before committing t, and a sequence that just finished still has a forward in flight: a wasted step. On a single stream that's one wasted forward for every L tokens the request generated, so about 1% at L ≈ 110. Pack a batch, though, and it nearly vanishes — the zombie is just one more row in a step that's already paying full price to stream the weights, so it rides along almost free. The tax bites hardest at one stream and fades exactly where throughput lives, which is why predicting it needs both L and the batch size.

    Here's that step, measured both ways — blocking idles each step while the CPU commits the last token and re-launches; pipelining runs that work (and the async mask upload) underneath the forward, so the forwards never stop:

    Now put real numbers in it. Measure each piece on its own — the two step times and L — and the model's prediction should land on what the benchmark actually delivers (depth-1 blocking vs depth-2 pipelined, nothing else changed):

    Three things to read out of it:

    1. The win grows with GPU speed. Same workload, +12% on a 3090 but +35% on a B200 at 32 streams. The bookkeeping is GPU-speed-independent, so as the forward shrinks — faster memory, or a smaller model — the bubble is a bigger share of the step. Pipelining is insurance against the GPU getting faster, which for us is the same thing as the model getting smaller.
    2. The zombie tax is real but small, and it amortizes. At one stream the zombie is a whole wasted forward — about 1% at L≈110. At batch it's one extra row in a step that's memory-bound on the weights, not the row count, so it costs almost nothing: at 32 streams the 3090's observed +11.6% lands right on the no-zombie per-step ratio. The tax bites at a single stream and fades exactly where throughput lives. (The B200's 32-stream row sits a few points under prediction for a duller reason — at ~4 ms/step the whole run is under half a second, so prefill and the end-of-run batch ramp-down are a visible slice of the wall.)
    3. It only pays once the bubble is actually hideable. (This is how we caught a bug, in fact: the pipelined numbers came out at blocking speed, traced to an accidental synchronous copy while building the constrained-decode mask. Moving it to the copy stream was worth +11% on the 3090 and +34% on the B200.)

    It's never just one thing

    That's the whole technique: ping-pong slots so two steps don't collide, a forward/sampling split so even constrained decoding can run ahead, and a little zombie refcounting so finished requests tear down cleanly. The GPU stops waiting on the CPU, and you get back anywhere from a few percent to a third; more the faster your accelerator/model is.

    But Photon isn't fast because of this one technique, or any single technique. It's fast because dozens of these details compound across the serving stack: how we resize and tile images on the way in, the kernels that run the model, the scheduler ordering here, and the synchronization points we remove from the hot path. No one piece is the whole story; the stack gets fast when enough of them line up.

    We'll keep writing these up, one corner of the stack at a time. Follow us on Twitter so you don't miss the next one. And keep an eye out for Photon 2.0, coming soon: we can't share details yet, but it's a big one.

    Original source
  • May 1, 2026
    • Date parsed from source:
      May 1, 2026
    • First seen by Releasebot:
      May 2, 2026
    Moondream logo

    Moondream

    Photon 1.2.0: Faster Inference, Now on Mac, Windows, Blackwell, and Jetson Thor

    Moondream ships Photon 1.2.0 with faster local vision AI and broader native hardware support across Apple Silicon, Windows x86_64, NVIDIA Blackwell, Jetson Thor, and existing GPUs. It lowers latency, boosts throughput, and makes on-device deployment easier.

    Production vision AI that runs everywhere — now faster, on more hardware.

    Moondream's mission is simple: production vision AI that runs everywhere. Most cloud VLMs take seconds to respond, which doesn't work for systems that need to be fast and often run on-device or at the edge. So we built the full stack ourselves: our own models, a fine-tune service (Lens), and an inference engine (Photon). Today's Photon update in the Moondream 1.2.0 release makes it faster still.

    Getting Started

    To install:

    pip install moondream
    

    Then run locally by setting

    local=True
    

    :

    import moondream as md
    from PIL import Image
    
    model = md.vl(api_key="YOUR_API_KEY", local=True)
    image = Image.open("photo.jpg")
    
    print(model.caption(image)["caption"])
    

    That local=True flag is the important part. It tells Moondream to run inference on your machine using Photon instead of sending the request to the hosted API. With Photon 1.2.0, local Moondream inference now supports:

    Platform | What's new
    Apple Silicon | Native inference on M-series Macs
    Windows x86_64 | Native CUDA inference (no WSL required) or Linux containers
    NVIDIA Blackwell | Support for B200 and RTX PRO 6000
    NVIDIA Jetson Thor | Edge inference on JetPack 7 / CUDA 13
    Existing NVIDIA GPUs | Faster prefill, MoE, dispatch, and tail latency

    The result: Moondream is now easier to deploy across laptops, workstations, edge devices, and production GPU servers.

    Why This Release Matters

    Production vision AI depends on more than model quality. It needs to:
    • be strong fast enough for real applications.
    • run on the hardware teams already use.
    • work outside a single cloud or GPU environment.
    • be simple enough to install and ship.

    Photon 1.2.0 improves Moondream across all of those dimensions. It expands native hardware support, reduces setup complexity, improves single-request latency, and increases throughput on both new and existing GPUs.

    That matters for applications like:

    Use case | What Photon improves
    Interactive image apps | Faster answers from a single request
    Production APIs | Higher request throughput
    Robotics and inspection | Local inference without cloud round trips
    Desktop tools | Native Mac and Windows support
    Edge devices | Vision AI where network latency or privacy matters
    Private workflows | Images can stay on-device

    Native Moondream Inference on Apple Silicon

    Photon now runs on Apple M-series Macs starting with macOS 13 Ventura and Python 3.12. Photon uses native Metal kernels across the decode path, including paged attention, rotary embeddings, KV cache management, MoE routing, sampling, and layer norm. KV cache sizing is automatically tuned to the Mac's unified memory.

    Reference performance on ChartQA, batch size 4, direct mode:

    Hardware | Moondream 2 | Moondream 3
    MacBook Pro, M5 Max, 48 GB | 7.26 requests/sec | 4.58 requests/sec
    Mac mini, M2, 24 GB | 0.79 requests/sec | 0.55 requests/sec
    Mac mini, M4 base, 16 GB | 0.84 requests/sec | —

    Apple Silicon support makes local Moondream development much more practical: demos, prototypes, desktop apps, and privacy-sensitive workflows can run directly on a Mac.

    Native Windows Support

    Photon now supports native Windows x86_64 inference. This is not a Linux wrapper. Photon's kernel-loading runtime has been rebuilt to support Windows directly, including MSVC compatibility, Windows DLL loading semantics, and cross-platform library naming across the kernel stack. Windows systems now run the same CUDA kernels as Linux x86_64.

    Low-Latency and High-Throughput Inference on Blackwell

    Photon 1.2.0 adds support for NVIDIA Blackwell, including B200 data-center GPUs and RTX PRO 6000 workstation GPUs.

    B200 is now the fastest hardware Photon supports:

    Hardware | Model | Single-request latency | Batch 64 throughput
    NVIDIA B200 | Moondream 2 | ~23 ms | 93.61 requests/sec
    NVIDIA B200 | Moondream 3 | ~30 ms | 71.27 requests/sec

    The single-request latency is derived from batch size 1 performance. This is the number that matters when an application needs an answer immediately. The batch 64 number shows high-volume throughput. This is the number that matters when a system is serving many requests at once. At batch size 64, B200 is:

    Model | Speedup vs. H100
    Moondream 2 | 1.49× faster
    Moondream 3 | 1.23× faster

    Photon also supports RTX PRO 6000, which reaches 39.3 requests/sec on Moondream 2 and 39.7 requests/sec on Moondream 3 at batch size 64. Under the hood, this release includes Blackwell-specific MoE kernels and dedicated Blackwell flash-attention kernels for both decode and prefill. The practical result is lower latency for interactive workloads and higher throughput for production serving.

    Edge Inference on Jetson Thor

    Photon now also supports NVIDIA Jetson AGX Thor 64 GB on JetPack 7.

    This brings Moondream to a new class of edge deployments: robotics, inspection systems, kiosks, vehicles, cameras, and embedded vision products where cloud inference may add latency, cost, or privacy concerns.

    Reference performance:

    Hardware | Model | Single-request latency | Batch 64 throughput
    Jetson AGX Thor | Moondream 2 | ~152 ms | 14.53 requests/sec
    Jetson AGX Thor | Moondream 3 | ~147 ms | 12.05 requests/sec

    That means Moondream can run locally on Jetson Thor and return vision-language answers in well under a second.

    Photon also now ships a multi-CUDA Linux aarch64 wheel. The same install works across Jetson Thor, Jetson Orin, and GH200 systems. Photon selects the correct CUDA build automatically: CUDA 13 for Thor on JetPack 7, CUDA 12 for Jetson Orin and GH200 systems on JetPack 6.

    Faster on Existing NVIDIA GPUs

    Photon 1.2.0 also improves performance on existing NVIDIA hardware, including L40S, RTX 4090, Jetson Orin, A100, A10/A10G, L4, and RTX 6000.

    The main improvements are:

    Improvement | Impact
    Faster FP8 prefill on Ada and Jetson Orin | Better performance for FP8 KV cache deployments
    New native paged flash-attention kernels | Faster prefill and decode paths
    Faster MoE inference | Better Moondream 3 performance across GPUs
    Lower per-call dispatch overhead | Faster batch 1 and small-batch inference
    More consistent tail latency | More predictable application performance

    Small-batch performance is especially important in real applications. When a user asks a question about an image, batch size 1 latency determines how fast the answer comes back. Photon 1.2.0 reduces overhead in that path while also improving throughput for larger production batches.

    Conclusion

    Photon 1.2.0 expands where Moondream can be deployed and improves how fast it responds. Full benchmark details, including additional batch sizes and chain-of-thought mode results, are available in PERFORMANCE.md.

    With Moondream, you don't have to compromise. You can get sophisticated visual reasoning at near-realtime speeds, and it runs everywhere. Got a production-level vision challenge? Contact us, we'd love to talk.

    Original source
  • Apr 16, 2026
    • Date parsed from source:
      Apr 16, 2026
    • First seen by Releasebot:
      Apr 17, 2026
    Moondream logo

    Moondream

    Lens: Moondream's Finetune Service

    Moondream launches Lens, a fine-tuning product for production-ready vision AI. It helps users improve model accuracy with simple pay-as-you-go RL and supervised fine-tuning, then run the tuned model in Cloud or locally with Photon.

    Lens in Action: PTZOptics

    VLMs are a big leap for vision AI. They reason about vision at a higher level, and they're much easier to use than the previous generation of models. But they've been hard to put into production for three reasons. First, they're slow, and production systems often need real-time decisions. Second, they struggle to run locally, which production systems often require for security, reliability, or cost. Third, they suffer from a "last-mile" problem: the VLM looks promising in the lab, but in the real world, accuracy falls short.

    Moondream is a different kind of VLM. It's purpose-built for production systems, and unsurprisingly, we've been tackling these three problems. First, we built small models that are lightweight enough to run everywhere. Then we launched Photon, our inference engine that achieves 20ms inference time on an H100. And today, we're happy to announce the launch of Lens, our fine-tuning product that solves the "last-mile" problem.

    We've been working with a partner, PTZOptics, who make network-attached remote controlled cameras. In many cases, customers want the camera to act as if it had a smart camera operator controlling it: following the action of a soccer game, zooming in and out at crucial times in a presentation, or detecting anomalies for security or operational reasons.

    With Moondream, this is now a reality. You can have the camera track complicated things ("the person in the red shirt"), take inventory of what's shown, or get alerted when actions occur ("someone's hurt"). And with Lens, you can teach Moondream new skills, or tune it when the accuracy is lacking.

    Simple API, Pay-as-You-Go

    Lens is a simple API that provides fine-tuning, through both reinforcement learning and supervised fine-tuning. There's no hardware to set up or binaries to worry about. And it's simple pay-as-you-go. We've seen great results with as few as a dozen images.

    As soon as you're done improving the model, you can invoke it immediately through our Cloud, or run it locally with Photon. It's the easiest and fastest way to go from fine-tune to production.

    See the Difference

    Here are a few examples of Lens fine-tunes across very different domains. In each case, the base model struggled, and the fine-tuned model nailed it.

    Broadcast Sports: Detecting the Ball Handler

    We fine-tuned Moondream to detect the player with the ball in NBA broadcast footage. The base model returned dozens of false positives (red boxes). After fine-tuning with RL, it finds just the ball handler (green box). F1 jumped from 28% to 79%, and false positives dropped from 61 to 2.

    Before fine-tuning
    After fine-tuning

    Training took 54 minutes and cost $16.89.
    See the full interactive example →

    Geolocation: Country Identification from Street View

    We trained Moondream to identify countries from street-view imagery. The base model guesses the wrong continent entirely. After fine-tuning with just 25 images per country, it reads road markings, signage, and landscape cues correctly, beating GPT-5.4's 69.8% accuracy with 71.1%.

    Before fine-tuning
    After fine-tuning

    See the full interactive example →

    Medical Imaging: Glaucoma Staging

    We fine-tuned Moondream to classify retinal images by glaucoma severity. The base model defaulted to "early" for nearly every image. The fine-tuned model distinguishes severity correctly, performing 2x better than GPT-5.4.

    Before fine-tuning
    After fine-tuning

    Training took 47 minutes and cost $15.68.
    See the full interactive example →

    You can explore all of our fine-tune examples at moondream.ai/p/lens.

    Need Help Getting Started?

    For customers that are new to vision AI and fine-tuning, we offer help. Our dedicated production team can work with you to deliver a fine-tuned model for your specific use case. Contact [email protected] for more info.

    What We've Been Building Toward

    This is what we've been building toward. VLMs are finally in a form factor that fits the real world.

    Try Lens today, and if you're at NAB next week, come see the PTZOptics demo live. Questions? Drop us a line at [email protected].

    Original source
  • Mar 25, 2026
    • Date parsed from source:
      Mar 25, 2026
    • First seen by Releasebot:
      Mar 26, 2026
    Moondream logo

    Moondream

    Photon: Real-Time Vision AI Is Finally Here

    Moondream introduces Photon, a faster production inference engine for vision AI, delivering 2x speed over similar vLLM setups and over 60 inferences per second on H100s. It is built for real-time image and video analysis across edge to data center.

    The era of production vision AI isn't coming. It's here.

    Vision Language Models (VLMs) changed the game. Instead of building custom CV pipelines for every task, you can now just prompt a model about an image in plain language. That alone made vision AI easier and cheaper to adopt. But VLMs also unlocked something deeper: visual reasoning that simply wasn't possible before. Problems that were out of reach for traditional AI systems are now solvable, and almost anyone can afford to try.

    The result has been an explosion of new vision AI applications. Manufacturing defect detection. Broadcast video analysis. Retail inventory and loss prevention. What used to be research-grade problems are now powering a new wave of startups, and Moondream is at the center of many of them.

    But there's a gap between what VLMs can do and what they can do fast enough to matter.

    Most people's experience with a VLM looks like this: you ask it a question about an image, wait a few seconds (sometimes tens of seconds), and get an answer back. The answers are often impressive. The wait is often a dealbreaker. When you're processing live video, running a manufacturing line, or making real-time decisions, a few seconds of latency kills the use case entirely.

    We heard this over and over from customers. They wanted everything Moondream offers: the accuracy, the grounding, the ease of use. But they needed it faster than any VLM had delivered before.

    Photon is our answer.

    Why We Could Build This

    Photon isn't just fast inference code. The real advantage is that we own the entire stack. We design the model. We design the inference engine. We design the fine-tuning platform and the deployment tools.

    We made architectural decisions at model design time to optimize for the hardware we actually deploy on. We knew which GPU operations would matter on which chips, and shaped the model around that. You can't retrofit those decisions onto an existing model. Compared to similar-sized models on vLLM, Photon is 2x faster. You don't get that from better kernels alone.

    On an H100, Photon delivers over 60 inferences per second. That's frame-by-frame video processing on server-class hardware. On edge devices, including older ones limited by supply chain realities, we still deliver meaningful throughput.

    Here's what matters in practice: production vision AI systems rarely run just one inference per image. You're often analyzing the same frame in multiple ways. Photon gives you the headroom to do that.

    What This Changes

    Live broadcasting with real-time moderation. Manufacturing lines running at full speed with frame-by-frame defect detection. Security systems that keep pace with camera feeds. These were theoretically possible before. Now they're operationally viable.

    Speed also affects cost. When you run inference faster on a GPU, each inference gets cheaper. Photon supports operation batching, which lets you trade slightly higher per-inference latency for much better total throughput. The result is that real-time image and video analysis can fit much tighter budgets than before.

    Moondream Is Now Production-Ready, End to End

    Moondream is becoming a complete stack for production vision AI. The model works well at grounding tasks across industries. Lens, our upcoming fine-tuning platform, makes it easy and cheap to improve accuracy on your specific use case. And now Photon gives you a best-in-class inference engine that runs everywhere, edge to data center, with full support.

    Getting started takes minutes, not days:

    pip install moondream
    
    import moondream as md
    from PIL import Image
    
    model = md.vl(api_key="YOUR_API_KEY", local=True)
    image = Image.open("photo.jpg")
    
    print(model.caption(image))
    # => {"caption": "A golden retriever sitting on a park bench, looking ..."}
    

    Moondream is free to download and run however you want. Photon is for teams that need faster, production-ready performance. See pricing for details and the documentation to get started.

    What's Next

    Lens, our fine-tuning product, is launching soon. More hardware support for Photon is on the way. As both products mature, they'll integrate more tightly so you can fine-tune on your data and deploy through Photon in a single step.

    We're going to stay focused on making Moondream the best production-ready VLM. Faster. Less memory. Lower cost. Running everywhere.

    Original source
  • Mar 10, 2026
    • Date parsed from source:
      Mar 10, 2026
    • First seen by Releasebot:
      Mar 11, 2026
    Moondream logo

    Moondream

    Moondream Segmenting Update: Better Masks, Better Benchmarks, 40% Faster

    Moondream announces a faster, higher accuracy segmenting upgrade now live in Moondream Cloud. The new model boosts RefCOCO benchmarks, inscribes native SVG masks, and enhances referring expressions with faster latency. Local inference will follow later in the week, making it a clear product release.

    We introduced segmenting as a Moondream skill in September 2025 with Moondream 3 Preview. It launched with state-of-the-art scores on segmenting benchmarks. Despite the launch of several segmenting vision models since then, Moondream remains top dog.

    Today, we're excited to announce that we've raised the bar even further with an improvement now live on Moondream Cloud. This new version produces better segmenting results, achieves better benchmark scores, and does it 40% faster than before.

    Examples

    • Car closest to the top of lombard street
    • castle-like building
    • Headlights
    • man wearing blue shirt and jeans, standing near the left railing of the bridge, looking down
    • Runner in the lead with the longest hair
    • Transamerica Pyramid
    • Waldo wearing the number 25317
    • White 911

    Moondream Segmenting Recap

    Put simply, what makes Moondream segmenting different is that it:

    • produces native SVG masks (vectors, not bitmasks)
    • state of the art on segmentation benchmarks
    • offers quick inference speeds, even if raw speed is not the only thing we optimize for
    • supports deep, native referring capabilities such as "the person touching the door"

    Benchmark Improvements

    This latest segmentation model delivers a significant leap in performance across all major referring expression benchmarks. On RefCOCO+ Val, which tests attribute-based reasoning without positional cues, we achieve 79.1 mIoU, a 4.4-point improvement over the previous state-of-the-art (which was also Moondream!). RefCOCOg Val, which evaluates complex natural language descriptions, sees similar gains at 80.7 mIoU. We also report 88.2 mIoU on RefCOCO-M, our high-fidelity benchmark with pixel-accurate masks, underscoring that these gains translate to real-world precision, not just benchmark optimization.

    | Metric: mIoU | RefCOCO Val | RefCOCO+ Val | RefCOCOg Val | RefCOCO-m |
    | Old | 81.8 | 74.7 | 76.4 | 86.9 |
    | New | 83.2 (+1.4) | 79.1 (+4.4) | 80.7 (+4.3) | 88.2 (+1.3) |

    How We Compare

    Most segmentation-capable VLMs are either accurate or fast, but not both. Large multimodal models with bolted-on segmentation decoders can handle complex queries, but they are slow and expensive to run at scale. Lightweight models are fast, but they choke on anything beyond simple noun phrases. Moondream closes this gap: state-of-the-art accuracy at speeds that make latency-sensitive and high-throughput applications practical.

    Moondream vs. SAM 3

    SAM 3 can segment generic concepts like "car" or "person", but it can't natively resolve referring expressions. For prompts like "the person touching the door" or "laundry on the floor," you need to pair it with a larger reasoning model that adds 10s of seconds of latency and drives up cost. Moondream handles complex prompts natively, returns crisp higher-quality SVG masks, at a 5x lower price point.

    Conclusion

    This update is live now on Moondream Cloud. If you're already using segmentation, you get better quality and lower latency immediately. Later this week, we'll also be releasing the model for local inference, along with a technical whitepaper for those who want to go deeper. Learn more about Moondream's segmentation skill at /skills/segment.

    Original source
  • Dec 19, 2025
    • Date parsed from source:
      Dec 19, 2025
    • First seen by Releasebot:
      Dec 20, 2025
    Moondream logo

    Moondream

    We added Moondream 3 Preview support to Moondream Station

    Moondream Station launches Mac with Moondream 3 Preview, delivering native MLX performance and quantized models for Apple Silicon. The one click installer runs on Mac Windows and Linux and showcases snappy inference, with 35+ tokens per second on an M1 Max.

    Moondream Station

    Moondream Station is our free on-prem client for Moondream. A one-click (or one command) installer makes it a snap to get Moondream running on your Mac, PC, or Linux box instantly. Today we're happy to announce Moondream 3 Preview support on Mac. Try it out for yourself (works on mac, windows, linux):

    pip install moondream-station

    Built for Apple Silicon

    To get the most out of Apple Silicon, we built Mac inference to be fully MLX native and added quantized Moondream 3 support. The result is snappy performance. You'll need a Mac with at least 16GB of memory. On an M1 Max with 64GB, we are seeing over 35 tokens per second. Here's a demo of how it works on that M1 Max:

    Moondream uses dedicated grounding tokens, so any x or y coordinate only requires one token. This means inferences for grounded skills like point or detect feel near instantaneous.

    Using the API

    Once Moondream Station is running, you can connect to it using our Python client:

    # pip install moondream
    import moondream as md
    from PIL import Image
    
    # Connect to Moondream Station
    model = md.vl(endpoint="http://localhost:2020/v1")
    
    # Load an image
    image = Image.open("path/to/image.jpg")
    
    # Ask a question
    answer = model.query(image, "What's in this image?")["answer"]
    print("Answer:", answer)
    

    What's next

    We are planning to make more improvements to Moondream Station over the next few weeks. If you have ideas or requests, reach out on Discord.

    Happy holidays from the Moondream team.

    Original source
  • Oct 17, 2025
    • Date parsed from source:
      Oct 17, 2025
    • First seen by Releasebot:
      Dec 6, 2025
    Moondream logo

    Moondream

    Announcing Moondream Cloud

    Moondream launches Moondream Cloud, a hosted vision AI platform built on Moondream 3 Preview. It emphasizes speed and cost, with pay‑as‑you‑go pricing, $5 free credits, and enterprise options including on‑prem and compliance. A clear release highlighting a new product rollout.

    Fast, cheap, smart. Pick three.

    We're excited to launch Moondream Cloud, a hosted version of Moondream that makes it easy to build cutting-edge vision applications.

    When choosing vision AI tech, three things matter most: intelligence, speed, and cost. With our recent launch of Moondream 3 Preview, our model already delivers top-tier intelligence, reaching SOTA on visual reasoning and grounding tasks, outperforming top frontier models. Our Moondream Cloud release focuses on the other two: speed and cost.

    Pricing

    Moondream Cloud is pay-as-you-go. No subscriptions, no commitments, just load up credits and you're done. To help you start building right away, you get $5 in free monthly credits too (no credit card required!).

    Our pricing is token based: Moondream 3 Preview costs $0.30 per million input tokens, and $2.50 per million output tokens. These token rates are simlar to Gemini 2.5 Flash and GPT-5 Mini. But token pricing doesn't tell the full story. Moondream uses a custom SuperBPE tokenizer that means we generate 21% fewer tokens for the same output text. We have dedicated grounding tokens that represent points with two tokens and object bounding boxes with three tokens, where competing models have to use tens of tokens. And we represent images of all resolutions with 729 tokens, leading to significant savings on prefill.

    We simulated a workload where each of the three examples below are processed once a minute, for 30 days. To do this for all three images would cost:

    Comparisons

    We compared Moondream Cloud with Gemini Flash 2.5 and GPT-5 Mini. Both are vision-capable and similarly priced. (We skipped Claude Haiku 4.5 because its vision capabilities were significantly behind on the tasks we evaluated.)

    Example 1: Pointing

    Average runtime: Moondream 3 (Preview) 1.52 seconds, Gemini 2.5 Flash 3.02 seconds, GPT-5 Mini 27.58 seconds
    Input tokens: Moondream 737, Gemini 1,352, GPT-5 Mini 419
    Output tokens: Moondream 25, Gemini 241, GPT-5 Mini 1,372
    Monthly cost (1 RPM): Moondream $12, Gemini $35, GPT-5 Mini $123

    In this example, Moondream is cheaper because we use both fewer input and fewer output tokens. We require fewer tokens both because we encode the image efficiently (compared to Gemini 2.5 Flash), and because we don't need a complicated text prompt to get the model to output just the list of 2D points. On the outputs, Moondream benefits from having dedicated grounding tokens, requiring only two tokens per point. The result is that Moondream is significantly cheaper to run.

    Example 2: Object detection

    Average runtime: Moondream 4.56 seconds, Gemini 7.69 seconds, GPT-5 Mini 52.88 seconds
    Input tokens: Moondream 737, Gemini 1,839, GPT-5 Mini 1,849
    Output tokens: Moondream 103, Gemini 1,524, GPT-5 Mini 3,271
    Monthly cost (1 RPM): Moondream $21, Gemini $170, GPT-5 Mini $302

    Again, Moondream is more efficient because our grounding tokens mean we only emit three tokens per bounding box -- two tokens encoding the position of the middle of the box, and one token encoding both the height and width. Like before you'll notice we're also significantly more accurate.

    Example 3: OCR

    Average runtime: Moondream 3.92 seconds, Gemini 3.44 seconds, GPT-5 Mini 18.47 seconds
    Input tokens: Moondream 743, Gemini 1,395, GPT-5 Mini 1,812
    Output tokens: Moondream 414, Gemini 533, GPT-5 Mini 528
    Monthly cost (1 RPM): Moondream $54, Gemini $75, GPT-5 Mini $65

    This one is more evenly matched, since we're emitting normal text output. But Moondream still wins on cost because of more efficient image encoding, and more efficient output tokenization (using our custom tokenizer).

    Throughput and Data Privacy

    On the free tier, we allow up to two requests per second. When you hold $10 or more in paid credits, we increase that to 10 requests per second. We never train on your data, and no data is persisted after returning responses.

    We also offer enterprise plans with:

    • On-prem inference (run Moondream in your own infrastructure)
    • Compliance options (e.g. HIPAA)
    • Dedicated consulting and support
    • Volume-based pricing

    Reach out at [email protected] to discuss your needs.

    Conclusion

    Moondream exists for one reason: to power the next wave of vision AI agents. Our new 9B parameter mixture-of-experts Moondream 3 (Preview) model combines the speed of a 2B model with state-of-the-art visual reasoning and grounding, with no compromises. And now, with Moondream Cloud, using it as simple as it gets. Fast, cheap, smart -- pick three.

    Go to the cloud console to grab an API key, then check out our documentation to get started!

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.