Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions docs/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,9 @@ title: "Overview"
Moondream is an open-weight family of Vision Language Models (VLMs) built for powerful, efficient visual reasoning. The Moondream 3 family uses a mixture-of-experts architecture with grounded visual reasoning, a 32k context window, and native support for multiple vision skills—like pointing, counting, and object detection—all designed with a deployment-friendly ethos.

For local inference, [Photon](/running-locally) exposes Moondream, Qwen, Gemma,
and Whisper models through the Moondream Python package. For hosted inference,
choose a model through the [Cloud API quickstart](./quickstart).
and speech-to-text models including Whisper, Qwen3-ASR, and Parakeet through
the Moondream Python package. For hosted inference, choose a model through the
[Cloud API quickstart](./quickstart).

### Key stats

Expand Down
2 changes: 1 addition & 1 deletion docs/quickstart.md
Original file line number Diff line number Diff line change
Expand Up @@ -344,7 +344,7 @@ console.log(result.caption);

Want to run models on your own hardware instead of using the Cloud API? Start
with the [Photon local inference guide](/running-locally), or use
[Whisper speech transcription](/transcription) on a supported NVIDIA GPU.
[Photon speech transcription](/transcription) on a supported NVIDIA GPU.

## Next Steps

Expand Down
26 changes: 18 additions & 8 deletions docs/running-locally.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,10 @@ description: High-performance local inference with Photon on NVIDIA GPUs and App

Photon is Moondream's high-performance local inference engine for NVIDIA GPUs
(Linux x86_64 / aarch64 or Windows AMD64) and Apple Silicon Macs. It supports
Moondream, Qwen, and Gemma vision-language models, plus Whisper speech
transcription on NVIDIA GPUs. Photon provides custom CUDA and Metal kernels,
automatic batching, paged KV caching, and prefix caching.
Moondream, Qwen, and Gemma vision-language models, plus automatic speech
recognition (ASR) with Whisper, Qwen3-ASR, and Parakeet on NVIDIA GPUs. Photon
provides custom CUDA and Metal kernels, automatic batching, paged KV caching,
and prefix caching.

## Requirements

Expand Down Expand Up @@ -96,25 +97,34 @@ qwen = md.photon("Qwen/Qwen3.5-4B")
gemma = md.photon("google/gemma-4-E2B-it")

# Whisper large-v3-turbo (NVIDIA GPU)
speech = md.photon("openai/whisper-large-v3-turbo")
whisper = md.photon("openai/whisper-large-v3-turbo")

# Qwen3-ASR 0.6B (NVIDIA GPU)
qwen_asr = md.photon("Qwen/Qwen3-ASR-0.6B")

# Parakeet TDT 0.6B v3 (NVIDIA GPU)
parakeet = md.photon("nvidia/parakeet-tdt-0.6b-v3")
```

| Family | Supported models |
|--------|----------------|
| Moondream | Moondream 2, Moondream 3 Preview, Moondream 3.1 9B A2B |
| Qwen 3.5 | 0.8B, 2B, 4B, 9B, 27B, and 35B-A3B; base variants where published |
| Qwen 3.6 | 27B and 35B-A3B; BF16 and FP8 checkpoints |
| Gemma 4 | E2B, E4B, and 31B base and instruction variants |
| Gemma 4 | E2B, E4B, 26B-A4B, and 31B base and instruction variants |
| Whisper | Whisper large-v3-turbo transcription and English translation |
| Qwen3-ASR | 0.6B and 1.7B transcription with optional forced-alignment word timestamps |
| Parakeet TDT | 0.6B v3 transcription with native word timestamps |

Use `md.photon_models()` to list the exact identifiers registered by the
installed release. Models expose `model.tasks` and `model.supports(task)` so
applications can discover their capabilities; not every model implements every
Moondream-specific skill.

Whisper currently runs through Photon's CUDA path on supported NVIDIA GPUs.
See [Speech Transcription](/transcription) for files, progressive results, live
PCM, translation, and timestamps.
Speech-to-text models currently run through Photon's CUDA path on supported
NVIDIA GPUs. See [Speech Transcription](/transcription) for model-specific
language, prompt, translation, and timestamp support, plus files, progressive
results, and live PCM.

Model weights are automatically downloaded from Hugging Face on first run and cached locally.

Expand Down
76 changes: 53 additions & 23 deletions docs/transcription.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,26 +2,47 @@
id: transcription
slug: transcription
title: Speech Transcription
description: Transcribe and translate files or live PCM with Photon and Whisper
description: Transcribe files or live PCM with Photon's Whisper, Qwen3-ASR, and Parakeet models
---

# Speech Transcription

Photon serves `openai/whisper-large-v3-turbo` through the Moondream Python
package on supported NVIDIA GPUs. It handles long-form files, progressive
transcripts, live mono PCM, English translation, and segment or word
timestamps.
Photon serves four speech-to-text checkpoints through the same model-bound
`transcribe` interface on supported NVIDIA GPUs:

- `openai/whisper-large-v3-turbo`
- `Qwen/Qwen3-ASR-0.6B`
- `Qwen/Qwen3-ASR-1.7B`
- `nvidia/parakeet-tdt-0.6b-v3`

All four handle long-form files, progressive transcripts, live mono PCM, and
segment timestamps. Each also offers word timestamps within model-specific
limits. Language, prompting, translation, and timestamp behavior varies by
model.

## Installation

Install Moondream 2.1 or newer:
Install Moondream 2.1.1 or newer:

```bash
pip install --upgrade "moondream>=2.1"
pip install --upgrade "moondream>=2.1.1"
```

Photon downloads the Whisper checkpoint from Hugging Face on first use and
caches it locally. Local Photon models do not require a Moondream API key.
Photon downloads the selected checkpoint from Hugging Face on first use and
caches it locally. Qwen3-ASR also downloads its forced aligner when word
timestamps are requested. Local Photon models do not require a Moondream API
key.

## Choose a model

| Capability | Whisper large-v3-turbo | Qwen3-ASR | Parakeet TDT |
|------------|------------------------|-----------|--------------|
| Language | Automatic or forced | Automatic or forced | Automatic; code not reported |
| English translation | Yes | No | No |
| Initial prompt | Yes | Yes | No |
| Segment timestamps | Yes | Yes | Yes |
| Word timestamps | Alignment | Forced alignment | Native token durations |
| Sampling controls | Temperature fallback | Temperature and top-p | Deterministic |

## Transcribe a file

Expand All @@ -31,7 +52,7 @@ from pathlib import Path
import moondream as md


with md.photon("openai/whisper-large-v3-turbo") as speech:
with md.photon("Qwen/Qwen3-ASR-0.6B") as speech:
result = speech.transcribe(
audio=Path("meeting.m4a"),
timestamps="word",
Expand Down Expand Up @@ -66,9 +87,10 @@ with md.photon("openai/whisper-large-v3-turbo") as speech:

## Translate to English

Set `task="translate"` to translate non-English speech into English. Use
`language` when the source language is known, or omit it for automatic language
detection.
Whisper alone supports English translation. Set `task="translate"` to translate
non-English speech into English. Use `language` when the source language is
known, or omit it for automatic language detection. Qwen3-ASR and Parakeet
reject translation requests.

```python
from pathlib import Path
Expand Down Expand Up @@ -167,21 +189,28 @@ async def final_transcript(audio_chunks):

## Options

Arguments are passed directly to the selected Photon model:
Arguments are passed directly to the selected Photon model. Unsupported
model-specific options fail with a clear error instead of being ignored:

| Argument | Default | Description |
|----------|---------|-------------|
| `audio` | Required | Encoded audio, raw PCM, or an asynchronous live PCM iterator |
| `sample_rate` | None | Required for raw or live PCM; do not set for encoded audio |
| `language` | Automatic | Source language code such as `"en"` or `"es"` |
| `task` | `"transcribe"` | Use `"translate"` for English translation |
| `timestamps` | `"segment"` | `"none"`, `"segment"`, or `"word"` |
| `initial_prompt` | None | Text context for names or domain-specific vocabulary |
| `condition_on_previous_text` | `True` | Carry bounded text context between long-form windows |
| `language` | Automatic | Source language code such as `"en"` or `"es"`; Whisper and Qwen3-ASR only |
| `task` | `"transcribe"` | Use `"translate"` for English translation with Whisper only |
| `timestamps` | `"segment"` | `"none"`, `"segment"`, or `"word"`; implementation varies by model |
| `initial_prompt` | None | Text context for names or vocabulary; Whisper and Qwen3-ASR only |
| `condition_on_previous_text` | `True` | Carry bounded context; Whisper and Qwen3-ASR only |
| `clip_start_seconds` | `0.0` | Start offset for an encoded source |
| `clip_end_seconds` | Source end | Exclusive end offset for an encoded source |
| `stream` | `False` | Return progressive transcript snapshots |
| `settings` | None | Sampling settings such as `temperature` and `max_tokens` |
| `settings` | None | Model-specific decode bounds and sampling settings |

Qwen3-ASR word timestamps use forced alignment for Chinese, Cantonese,
English, German, Spanish, French, Italian, Portuguese, Russian, Korean, and
Japanese. Use segment or no timestamps for its other transcription languages.
Parakeet detects language implicitly and returns `language: null`; it does not
accept language forcing, prompts, translation, temperature, or top-p.

Live PCM does not accept clip ranges. For encoded bytes and binary streams,
Photon snapshots at most 64 MiB before incremental decoding. Use an asynchronous
Expand All @@ -193,13 +222,14 @@ sessions are limited to 24 hours.
The final dictionary includes:

- `text`: the complete transcript or English translation;
- `language`: the detected or requested source language;
- `language`: the detected or requested source language, or `None` for Parakeet;
- `task`: `"transcribe"` or `"translate"`;
- `segments`: timestamped segments, with `words` when requested;
- source, clip, and decode diagnostic fields.

Word entries include `word`, `start`, `end`, and `probability`. All timestamps
are in seconds relative to the source audio.
Word entries include `word`, `start`, and `end`, plus `probability` when the
selected model reports it. All timestamps are in seconds relative to the source
audio.

For vision-language models and supported hardware, see
[Run Moondream Locally](/running-locally).
Loading