Process extension workspace for MOSS-SoundEffect v2.0 using the upstream runtime from OpenMOSS/MOSS-TTS and weights from OpenMOSS-Team/MOSS-SoundEffect-v2.0.
This repository is an MIT-licensed Modly integration wrapper. MOSS-SoundEffect v2.0 itself, including the upstream runtime and model weights, belongs to OpenMOSS/OpenMOSS-Team. See NOTICE and the upstream model/runtime links before redistributing weights or using the model commercially.
- Accepts a text prompt in English or Chinese.
- Runs the upstream
MossSoundEffectPipeline. - Writes a
.wavartifact underWorkflows/MOSS-SoundEffect/inside the workflowworkspaceDir. - Emits Modly process-runner JSONL messages:
progress,log,done,error.
- Bucket:
process-extension - Setup seam: root
setup.py - Runtime seam: root
moss_soundeffect_process.py - Model ownership: Modly global model assets under
models/moss-soundeffect-v2-process-extension/generate-soundeffect - Output contract:
done.result.filePathpoints to a workspace.wavaudio artifact;done.result.textcontains JSON metadata fallback
This extension expects Modly builds with audio as a first-class ArtifactKind. For older Modly builds that only support image, text, mesh, and scene, the process still emits JSON metadata in done.result.text and the generated .wav in done.result.filePath, but the node manifest may need to be downgraded to output: "text" until audio support is available.
This extension now resolves a host/backend lane before any pip install happens and fails fast on unsupported matrices.
Implemented matrix:
- Linux
aarch64+ NVIDIASM >= 121+ Python3.12: use local wheelhouse lanelinux-aarch64/local-wheelhouse-sm121whendependencies/wheelhouse/pytorch-sm121-cu130-aarch64-cp312/WHEELHOUSE.jsonexists - Linux
x86_64oraarch64+ NVIDIASM >= 121without a validated local wheelhouse: fail fast. Public wheels are not assumed compatible with this host class unless explicitly verified. - Linux
x86_64oraarch64+ NVIDIASM == 120: prefercu130when driver evidence is new enough; otherwise fallback to experimentalcu128 - Linux
x86_64oraarch64+ NVIDIASM < 120:cu128lane preserved to avoid breaking established users - Windows
x86_64+ NVIDIASM == 120+ Python3.11or3.12: validatedcu128lane, kept experimental while coverage expands beyond the reported RTX 5060 Ti host - Windows + NVIDIA
SM >= 121:cu130as experimental only - Windows
x86_64+ Python3.11+ pre-SM 8.0NVIDIA GPUs: setup and model loading can be validated withcu126, but practical generation is not supported for this GPU class. Treat this lane as setup/import probe only.
Fail-fast / unsupported:
- macOS / MPS
- CPU-only
- Windows lanes below
SM 121for practical inference unless explicitly validated on the target GPU class (SM 120is the current exception) - Any host where setup cannot resolve an evidence-backed lane
Critical rule: legacy cu128 wheels are blocked for SM 12.1+ host classes. Those hosts require a validated CUDA 13 wheelhouse or another explicitly verified lane.
For Linux aarch64 + cp312 + SM 12.1+, the preferred path is a local wheelhouse built from source with TORCH_CUDA_ARCH_LIST="12.1a". Public indexes are not assumed to cover that host class unless explicitly verified.
- Install/apply the extension from this repository/workspace.
- Run extension setup from Modly so Electron executes root
setup.py. - Let extension setup download model assets into Modly's global
models/moss-soundeffect-v2-process-extension/generate-soundeffect, or place an equivalent Hugging Face snapshot there before running. - Run the workflow node with a text prompt.
setup.py is designed for Modly's injected JSON context and will:
- use
ext_dirandpython_exefrom the setup payload when provided - create an extension-owned
venv/ - apply lane-specific Python support. The local SM121 wheelhouse lane currently targets Python
3.12; Windows SM120 accepts Python3.11/3.12; the Windows pre-SM80 probe lane can use Modly's Python3.11but is not considered viable for practical inference. - resolve the host lane from
gpu_sm,cuda_version,python_exe,ext_dir, local platform facts, and a lightnvidia-smiprobe when available - emit the detected host facts plus selected lane before dependency installation
- fail fast before installation when the host matrix is unsupported
- install lane-specific torch packages instead of hardcoding a single
cu128stack - prefer a local
wheelhouselane for Linuxaarch64+cp312+SM 12.1+whenWHEELHOUSE.jsonis present underdependencies/wheelhouse/pytorch-sm121-cu130-aarch64-cp312 - install the upstream Python runtime package from public GitHub
- run
pip check - run a post-install probe for
torch,torch.version.cuda, CUDA availability, device capability, torch arch coverage, andMossSoundEffectPipeline - reuse a previously ready runtime when
.modly/setup-ready.jsonstill matches the current setup schema, selected lane, and runtime requirements hash - download the Hugging Face model snapshot into the extension-owned logical model directory only when required sentinels are missing, unless forced explicitly
- validate required model sentinels after download
- write
.modly/setup-ready.json
Setup now records setup_schema_version, runtime_requirements_hash, and model_assets_signature in .modly/setup-ready.json.
- If the previous sentinel status is
readyorruntime_ready_model_missing, the selected lane is unchanged, the runtime hash still matches, and the extensionvenvstill exists, setup skips torch/runtime reinstalls and runs the post-install probe against the existing environment. - If the probe fails after a skipped install, setup falls back to the normal repair path and reinstalls the runtime for the resolved lane.
- If required model sentinels already exist at the logical model root, setup skips
snapshot_downloadand emits amodel-assets-skippedstatus event. - If the previous sentinel status is
error, or if the schema/hash/lane changed, setup does not trust the old state and reinstalls as needed.
Force flags:
force_reinstall=true: ignore reusable runtime state and reinstall dependencies for the resolved laneforce_model_download=true: force a model snapshot refresh even when sentinels already existskip_model_download=trueordownload_model_assets=false: keep runtime setup but do not download missing model assetsskip_probe=true: skip the post-install health probe when explicitly needed
Important: Modly's UI checks model readiness under the global modelsDir, not inside the extension folder. For this extension, setup provisions the same owner-scoped global model directory that the UI checks.
- expected model root:
models/moss-soundeffect-v2-process-extension/generate-soundeffect - default setup behavior: download
OpenMOSS-Team/MOSS-SoundEffect-v2.0into that root - optional setup escape hatch: pass
skip_model_download=true/download_model_assets=falseonly for dry setup or manually pre-seeded assets - runtime behavior: fail closed if sentinels are missing; runtime generation never downloads weights
This is deliberate. Downloading from the process runtime during generation creates ambiguous failures and bad UX. Setup owns model asset provisioning and records downloads_started, missing_model_files, and readiness status in .modly/setup-ready.json.
The node manifest intentionally does not expose hf_repo/download_check as a UI model-download button until Modly's process-extension ownership support is available in the running app. Otherwise the current UI can show a false Download state for process extensions even when setup has already provisioned the global model directory.
model_index.jsontransformer/diffusion_pytorch_model.safetensorsvae/vae_128d_48k.pthscheduler/scheduler_config.jsontokenizer/tokenizer.jsontext_encoder/model.safetensors.index.json- at least one
text_encoder/model-*.safetensorsshard
prompt: text prompt; English or Chineseseconds: default10, max30num_inference_steps: default100cfg_scale: default4.0sigma_shift: default5.0seed: default0,-1means randomtorch_dtype:bfloat16,float16,float32disable_compile: defaulttrue; if enabled, runtime setsTORCHDYNAMO_DISABLE=1andTORCHINDUCTOR_DISABLE=1before torch-heavy imports. This is now an explicit runtime choice again, not something silently forced by an override env path.output_name: optional custom filename stem; sanitized and forced to.wav
SM 12.1+: use a validated local wheelhouse or another explicitly verified CUDA 13 lane. Do not rely on legacy CUDA 12.8 wheels for this GPU class.- Pre-
SM 8.0NVIDIA GPUs: do not usebfloat16; the runtime rejects it and recommendsfloat16. Even withfloat16, this GPU class is not considered viable for useful MOSS-SoundEffect v2 generation based on current tests. - For any unvalidated GPU class, start with a short smoke test before attempting long/high-step generation.
- output directory:
Workflows/MOSS-SoundEffect/ - output filename: sanitized from
output_nameor prompt text - output type: WAV, 48 kHz
- manifest workflow output:
audio done.result.text: JSON metadata containing the WAV path and generation parametersdone.result.filePath: workspace-relative WAV artifact path
Absolute output paths and traversal are rejected.
- Static review of
manifest.json,setup.py, and runtime path safety - Modly setup run on Linux + NVIDIA CUDA with Python 3.12
- Confirm
.modly/setup-ready.jsonexists - Confirm logical model directory contains all sentinels
- Smoke test with a short prompt and low-risk parameters
- Confirm
done.result.filePathresolves to a real workspace WAV file
- Native Modly audio output handling requires a build that includes the audio ArtifactKind/workflow/preview changes.
- For older Modly builds without audio support, temporarily change the manifest node to
output: "text"and rely ondone.result.textmetadata plusdone.result.filePath. - First run may be slow due to
torch.compile/ Triton graph compilation. disable_compile=trueis still a valid escape hatch for fragile compile stacks, but it is NOT a substitute for the correct setup lane.- If the host is
SM 12.1+, setup must land on a validated CUDA 13-compatible lane. Legacycu128is intentionally rejected for that host class. - Source builds for Triton/PyTorch stay OUTSIDE automatic Modly setup. This repository now includes
tools/build-sm121-wheelhouse/with manual, reproducible helper scripts and aWHEELHOUSE.example.jsontemplate, but setup only consumes an already prepared local wheelhouse. - Windows pre-
SM 8.0NVIDIA GPUs can complete setup/import probes in some cases, but they are not supported practical inference targets for MOSS-SoundEffect v2.
manifest.json: Modly extension contractsetup.py: setup entrypoint wrappermoss_soundeffect_process.py: process runner entrypoint wrappermoss_soundeffect_ext/setup_runtime.py: setup implementationmoss_soundeffect_ext/host_compat.py: host detection and lane resolutionmoss_soundeffect_ext/process_runtime.py: process JSONL implementationmoss_soundeffect_ext/common.py: shared constants and safety helperstools/build-sm121-wheelhouse/: manual docs/scripts to build a local SM121 wheelhouse without making setup.py compile PyTorch or Triton