From 9269e48b27e079af7082208e696206d93ea8e592 Mon Sep 17 00:00:00 2001 From: Vikhyat Korrapati Date: Tue, 1 Sep 2026 13:40:54 -0700 Subject: [PATCH 1/2] Document Photon 2.1.1 model support --- docs/overview.md | 5 +-- docs/quickstart.md | 2 +- docs/running-locally.md | 25 +++++++++----- docs/transcription.md | 76 ++++++++++++++++++++++++++++------------- 4 files changed, 74 insertions(+), 34 deletions(-) diff --git a/docs/overview.md b/docs/overview.md index e9ca195..4548215 100644 --- a/docs/overview.md +++ b/docs/overview.md @@ -9,8 +9,9 @@ title: "Overview" Moondream is an open-weight family of Vision Language Models (VLMs) built for powerful, efficient visual reasoning. The Moondream 3 family uses a mixture-of-experts architecture with grounded visual reasoning, a 32k context window, and native support for multiple vision skills—like pointing, counting, and object detection—all designed with a deployment-friendly ethos. For local inference, [Photon](/running-locally) exposes Moondream, Qwen, Gemma, -and Whisper models through the Moondream Python package. For hosted inference, -choose a model through the [Cloud API quickstart](./quickstart). +and speech models including Whisper, Qwen3-ASR, and Parakeet through the +Moondream Python package. For hosted inference, choose a model through the +[Cloud API quickstart](./quickstart). ### Key stats diff --git a/docs/quickstart.md b/docs/quickstart.md index eba1080..64a1ea5 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -344,7 +344,7 @@ console.log(result.caption); Want to run models on your own hardware instead of using the Cloud API? Start with the [Photon local inference guide](/running-locally), or use -[Whisper speech transcription](/transcription) on a supported NVIDIA GPU. +[Photon speech transcription](/transcription) on a supported NVIDIA GPU. ## Next Steps diff --git a/docs/running-locally.md b/docs/running-locally.md index 44f23ee..1d4ad9e 100644 --- a/docs/running-locally.md +++ b/docs/running-locally.md @@ -9,9 +9,9 @@ description: High-performance local inference with Photon on NVIDIA GPUs and App Photon is Moondream's high-performance local inference engine for NVIDIA GPUs (Linux x86_64 / aarch64 or Windows AMD64) and Apple Silicon Macs. It supports -Moondream, Qwen, and Gemma vision-language models, plus Whisper speech -transcription on NVIDIA GPUs. Photon provides custom CUDA and Metal kernels, -automatic batching, paged KV caching, and prefix caching. +Moondream, Qwen, and Gemma vision-language models, plus Whisper, Qwen3-ASR, and +Parakeet speech transcription on NVIDIA GPUs. Photon provides custom CUDA and +Metal kernels, automatic batching, paged KV caching, and prefix caching. ## Requirements @@ -96,7 +96,13 @@ qwen = md.photon("Qwen/Qwen3.5-4B") gemma = md.photon("google/gemma-4-E2B-it") # Whisper large-v3-turbo (NVIDIA GPU) -speech = md.photon("openai/whisper-large-v3-turbo") +whisper = md.photon("openai/whisper-large-v3-turbo") + +# Qwen3-ASR 0.6B (NVIDIA GPU) +qwen_asr = md.photon("Qwen/Qwen3-ASR-0.6B") + +# Parakeet TDT 0.6B v3 (NVIDIA GPU) +parakeet = md.photon("nvidia/parakeet-tdt-0.6b-v3") ``` | Family | Supported models | @@ -104,17 +110,20 @@ speech = md.photon("openai/whisper-large-v3-turbo") | Moondream | Moondream 2, Moondream 3 Preview, Moondream 3.1 9B A2B | | Qwen 3.5 | 0.8B, 2B, 4B, 9B, 27B, and 35B-A3B; base variants where published | | Qwen 3.6 | 27B and 35B-A3B; BF16 and FP8 checkpoints | -| Gemma 4 | E2B, E4B, and 31B base and instruction variants | +| Gemma 4 | E2B, E4B, 26B-A4B, and 31B base and instruction variants | | Whisper | Whisper large-v3-turbo transcription and English translation | +| Qwen3-ASR | 0.6B and 1.7B transcription with optional forced-alignment word timestamps | +| Parakeet TDT | 0.6B v3 transcription with native word timestamps | Use `md.photon_models()` to list the exact identifiers registered by the installed release. Models expose `model.tasks` and `model.supports(task)` so applications can discover their capabilities; not every model implements every Moondream-specific skill. -Whisper currently runs through Photon's CUDA path on supported NVIDIA GPUs. -See [Speech Transcription](/transcription) for files, progressive results, live -PCM, translation, and timestamps. +Speech models currently run through Photon's CUDA path on supported NVIDIA +GPUs. See [Speech Transcription](/transcription) for model-specific language, +prompt, translation, and timestamp support, plus files, progressive results, +and live PCM. Model weights are automatically downloaded from Hugging Face on first run and cached locally. diff --git a/docs/transcription.md b/docs/transcription.md index 0a0a678..a17699e 100644 --- a/docs/transcription.md +++ b/docs/transcription.md @@ -2,26 +2,47 @@ id: transcription slug: transcription title: Speech Transcription -description: Transcribe and translate files or live PCM with Photon and Whisper +description: Transcribe files or live PCM with Photon's Whisper, Qwen3-ASR, and Parakeet models --- # Speech Transcription -Photon serves `openai/whisper-large-v3-turbo` through the Moondream Python -package on supported NVIDIA GPUs. It handles long-form files, progressive -transcripts, live mono PCM, English translation, and segment or word -timestamps. +Photon serves four speech-to-text checkpoints through the same model-bound +`transcribe` interface on supported NVIDIA GPUs: + +- `openai/whisper-large-v3-turbo` +- `Qwen/Qwen3-ASR-0.6B` +- `Qwen/Qwen3-ASR-1.7B` +- `nvidia/parakeet-tdt-0.6b-v3` + +All four handle long-form files, progressive transcripts, live mono PCM, and +segment timestamps. Each also offers word timestamps within model-specific +limits. Language, prompting, translation, and timestamp behavior varies by +model. ## Installation -Install Moondream 2.1 or newer: +Install Moondream 2.1.1 or newer: ```bash -pip install --upgrade "moondream>=2.1" +pip install --upgrade "moondream>=2.1.1" ``` -Photon downloads the Whisper checkpoint from Hugging Face on first use and -caches it locally. Local Photon models do not require a Moondream API key. +Photon downloads the selected checkpoint from Hugging Face on first use and +caches it locally. Qwen3-ASR also downloads its forced aligner when word +timestamps are requested. Local Photon models do not require a Moondream API +key. + +## Choose a model + +| Capability | Whisper large-v3-turbo | Qwen3-ASR | Parakeet TDT | +|------------|------------------------|-----------|--------------| +| Language | Automatic or forced | Automatic or forced | Automatic; code not reported | +| English translation | Yes | No | No | +| Initial prompt | Yes | Yes | No | +| Segment timestamps | Yes | Yes | Yes | +| Word timestamps | Alignment | Forced alignment | Native token durations | +| Sampling controls | Temperature fallback | Temperature and top-p | Deterministic | ## Transcribe a file @@ -31,7 +52,7 @@ from pathlib import Path import moondream as md -with md.photon("openai/whisper-large-v3-turbo") as speech: +with md.photon("Qwen/Qwen3-ASR-0.6B") as speech: result = speech.transcribe( audio=Path("meeting.m4a"), timestamps="word", @@ -66,9 +87,10 @@ with md.photon("openai/whisper-large-v3-turbo") as speech: ## Translate to English -Set `task="translate"` to translate non-English speech into English. Use -`language` when the source language is known, or omit it for automatic language -detection. +Whisper alone supports English translation. Set `task="translate"` to translate +non-English speech into English. Use `language` when the source language is +known, or omit it for automatic language detection. Qwen3-ASR and Parakeet +reject translation requests. ```python from pathlib import Path @@ -167,21 +189,28 @@ async def final_transcript(audio_chunks): ## Options -Arguments are passed directly to the selected Photon model: +Arguments are passed directly to the selected Photon model. Unsupported +model-specific options fail with a clear error instead of being ignored: | Argument | Default | Description | |----------|---------|-------------| | `audio` | Required | Encoded audio, raw PCM, or an asynchronous live PCM iterator | | `sample_rate` | None | Required for raw or live PCM; do not set for encoded audio | -| `language` | Automatic | Source language code such as `"en"` or `"es"` | -| `task` | `"transcribe"` | Use `"translate"` for English translation | -| `timestamps` | `"segment"` | `"none"`, `"segment"`, or `"word"` | -| `initial_prompt` | None | Text context for names or domain-specific vocabulary | -| `condition_on_previous_text` | `True` | Carry bounded text context between long-form windows | +| `language` | Automatic | Source language code such as `"en"` or `"es"`; Whisper and Qwen3-ASR only | +| `task` | `"transcribe"` | Use `"translate"` for English translation with Whisper only | +| `timestamps` | `"segment"` | `"none"`, `"segment"`, or `"word"`; implementation varies by model | +| `initial_prompt` | None | Text context for names or vocabulary; Whisper and Qwen3-ASR only | +| `condition_on_previous_text` | `True` | Carry bounded context; Whisper and Qwen3-ASR only | | `clip_start_seconds` | `0.0` | Start offset for an encoded source | | `clip_end_seconds` | Source end | Exclusive end offset for an encoded source | | `stream` | `False` | Return progressive transcript snapshots | -| `settings` | None | Sampling settings such as `temperature` and `max_tokens` | +| `settings` | None | Model-specific decode bounds and sampling settings | + +Qwen3-ASR word timestamps use forced alignment for Chinese, Cantonese, +English, German, Spanish, French, Italian, Portuguese, Russian, Korean, and +Japanese. Use segment or no timestamps for its other transcription languages. +Parakeet detects language implicitly and returns `language: null`; it does not +accept language forcing, prompts, translation, temperature, or top-p. Live PCM does not accept clip ranges. For encoded bytes and binary streams, Photon snapshots at most 64 MiB before incremental decoding. Use an asynchronous @@ -193,13 +222,14 @@ sessions are limited to 24 hours. The final dictionary includes: - `text`: the complete transcript or English translation; -- `language`: the detected or requested source language; +- `language`: the detected or requested source language, or `None` for Parakeet; - `task`: `"transcribe"` or `"translate"`; - `segments`: timestamped segments, with `words` when requested; - source, clip, and decode diagnostic fields. -Word entries include `word`, `start`, `end`, and `probability`. All timestamps -are in seconds relative to the source audio. +Word entries include `word`, `start`, and `end`, plus `probability` when the +selected model reports it. All timestamps are in seconds relative to the source +audio. For vision-language models and supported hardware, see [Run Moondream Locally](/running-locally). From 45e5e96e401be93ffbe7725cc5c6f4306e862385 Mon Sep 17 00:00:00 2001 From: Vikhyat Korrapati Date: Tue, 1 Sep 2026 13:48:52 -0700 Subject: [PATCH 2/2] Clarify Photon speech-to-text terminology --- docs/overview.md | 4 ++-- docs/running-locally.md | 15 ++++++++------- 2 files changed, 10 insertions(+), 9 deletions(-) diff --git a/docs/overview.md b/docs/overview.md index 4548215..3b0f7ad 100644 --- a/docs/overview.md +++ b/docs/overview.md @@ -9,8 +9,8 @@ title: "Overview" Moondream is an open-weight family of Vision Language Models (VLMs) built for powerful, efficient visual reasoning. The Moondream 3 family uses a mixture-of-experts architecture with grounded visual reasoning, a 32k context window, and native support for multiple vision skills—like pointing, counting, and object detection—all designed with a deployment-friendly ethos. For local inference, [Photon](/running-locally) exposes Moondream, Qwen, Gemma, -and speech models including Whisper, Qwen3-ASR, and Parakeet through the -Moondream Python package. For hosted inference, choose a model through the +and speech-to-text models including Whisper, Qwen3-ASR, and Parakeet through +the Moondream Python package. For hosted inference, choose a model through the [Cloud API quickstart](./quickstart). ### Key stats diff --git a/docs/running-locally.md b/docs/running-locally.md index 1d4ad9e..255519b 100644 --- a/docs/running-locally.md +++ b/docs/running-locally.md @@ -9,9 +9,10 @@ description: High-performance local inference with Photon on NVIDIA GPUs and App Photon is Moondream's high-performance local inference engine for NVIDIA GPUs (Linux x86_64 / aarch64 or Windows AMD64) and Apple Silicon Macs. It supports -Moondream, Qwen, and Gemma vision-language models, plus Whisper, Qwen3-ASR, and -Parakeet speech transcription on NVIDIA GPUs. Photon provides custom CUDA and -Metal kernels, automatic batching, paged KV caching, and prefix caching. +Moondream, Qwen, and Gemma vision-language models, plus automatic speech +recognition (ASR) with Whisper, Qwen3-ASR, and Parakeet on NVIDIA GPUs. Photon +provides custom CUDA and Metal kernels, automatic batching, paged KV caching, +and prefix caching. ## Requirements @@ -120,10 +121,10 @@ installed release. Models expose `model.tasks` and `model.supports(task)` so applications can discover their capabilities; not every model implements every Moondream-specific skill. -Speech models currently run through Photon's CUDA path on supported NVIDIA -GPUs. See [Speech Transcription](/transcription) for model-specific language, -prompt, translation, and timestamp support, plus files, progressive results, -and live PCM. +Speech-to-text models currently run through Photon's CUDA path on supported +NVIDIA GPUs. See [Speech Transcription](/transcription) for model-specific +language, prompt, translation, and timestamp support, plus files, progressive +results, and live PCM. Model weights are automatically downloaded from Hugging Face on first run and cached locally.