From 508dfe0d0bfdcfd130e48fa2ea06637b81e587ee Mon Sep 17 00:00:00 2001 From: NewCommer00 Date: Wed, 15 Apr 2026 23:49:46 +0800 Subject: [PATCH] feat(pitd, f0, docs): improve PITD and add hybrid F0 backend MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit BREAKING CHANGE: PITD scaler default changed from 2.0 to 1.0; PITD results prior to v0.9.0 are unreliable feat(f0): add hybrid F0 backend with fallback; improve stability and reduce discontinuities fix(pitd): fix overly flat PITD curves (issue #21); special thanks to @ma0shu for helping identify and diagnose this critical bug ❤ docs(readme): add v0.9.0+ warning; document scaler change; update troubleshooting --- README.en.md | 71 +++++-- README.md | 42 ++++- .../expressive_config.json" | 8 +- .../expressive_config.json" | 4 +- .../expressive_config.json" | 6 +- expressions/dyn.py | 25 ++- expressions/pitd.py | 75 +++++--- expressions/tenc.py | 29 ++- expressive_gui.py | 23 ++- locales/app.pot | 136 +++++++------ locales/en/LC_MESSAGES/app.po | 138 ++++++++------ locales/zh_CN/LC_MESSAGES/app.po | 136 +++++++------ tests/test_seqtool.py | 93 +++++++++ tests/test_wavtool.py | 178 ++++++++++++++++-- utils/seqtool.py | 58 +++++- utils/wavtool.py | 84 ++++++++- 16 files changed, 811 insertions(+), 295 deletions(-) diff --git a/README.en.md b/README.en.md index 79a1693..1e8d5f2 100644 --- a/README.en.md +++ b/README.en.md @@ -7,6 +7,15 @@

+> [!WARNING] +> 🚨 **Please Read Before Downloading** 🚨 +> +> **It is strongly recommended to use v0.9.0 or later**. Earlier versions of the **PITD** expression parameter processing algorithm contain a [critical flaw](https://github.com/NewComer00/expressive/releases/tag/v0.9.0) that may result in **incorrect pitch curve generation**. To download the latest version, please visit the [Releases page](https://github.com/NewComer00/expressive/releases). +> +> For users migrating from an older version to `v0.9.0` or later, note that the default value of the **PITD Scaler** is now `1.0` (previously `2.0`). If you have an old configuration file, please set the **PITD Scaler** to `1.0`. +> +> **🎵 Thank you for using Expressive🎵** + # Expressive **Expressive** is a [DiffSinger](https://github.com/openvpi/diffsinger) expression parameter importer developed for [OpenUtau](https://github.com/stakira/OpenUtau). It aims to extract expression parameters from real human vocals and import them into the appropriate tracks of your project. @@ -21,18 +30,17 @@ The current version supports importing the following expression parameters: | **Working with OpenUtau** | **Data Viewer** | |:---:|:---:| -| | | +| | | -> - *OpenUtau version from [keirokeer/OpenUtau-DiffSinger-Lunai](https://github.com/keirokeer/OpenUtau-DiffSinger-Lunai)* -> - *Singer model from [yousa-ling-official-production/yousa-ling-diffsinger-v1](https://github.com/yousa-ling-official-production/yousa-ling-diffsinger-v1)* +> - *Example from [`examples/明天会更好`](examples/明天会更好). Click to view details.* > [!TIP] >
> 👉 Click to expand the full voiced demo video 👈 > ->

+>

>

> >
@@ -47,6 +55,8 @@ By default, this application uses [rmvpe-onnx](https://github.com/newcomer00/rmv The [swift-f0](https://github.com/lars76/swift-f0) and [CREPE](https://github.com/marl/crepe) pitch extraction backends are also available. The former runs on CPU only and is the fastest option, though its accuracy is modest. The latter is a classic algorithm in the field and runs more slowly. In a CUDA environment, the CREPE backend will automatically enable GPU acceleration. +There is also a newly added experimental **hybrid** backend available. The hybrid backend combines the prediction results of rmvpe-onnx and swift-f0, primarily using the pitch extraction results from rmvpe-onnx. In voiced segments of the audio, if the confidence of rmvpe-onnx is low and the confidence of swift-f0 is high, the result from swift-f0 is used for correction, improving the overall accuracy of pitch extraction. + > \* On Windows, TensorFlow 2.10 is the last version that supports GPU acceleration, and Python 3.10 is the highest Python version supported by its `.whl` files. ## 📌 Use Case @@ -297,24 +307,53 @@ Relaunching the application should restore normal functionality, and this issue #### Future Plan The NiceGUI framework has begun improving its drag-and-drop support and should resolve this in a future release. -### PITD expression curve is overall too flat +--- + +### PITD expression curve is overly flat #### Symptom -The extracted PITD expression curve is too flat, with almost no significant variation overall. Pitch changes in the reference vocal are not reflected in the expression curve. -#### Possible Cause -The two confidence thresholds in the PITD extractor are set **too high**, causing many pitch changes to be discarded. +The extracted PITD expression curve is too flat, with almost no significant variation. Pitch changes in the reference vocal are not properly reflected in the curve. -#### Solution -First try using the best-performing rmvpe-onnx backend (with default confidence thresholds). If the issue persists, try lowering both confidence thresholds. In general, the **Utau vocal** is relatively clean, so it is advisable to first adjust the confidence threshold for the **Reference vocal**. +#### Possible Causes + +1. In versions earlier than **v0.9.0**, there is an issue in the conversion between pitch and PITD values, which can cause the curve to appear overly flat. +2. The two confidence thresholds in the PITD extractor are set **too high**, causing many pitch changes to be discarded. + You can observe missing segments in the original pitch curve in [`expressive-viewer`](#iewer). + +#### Solutions -### PITD expression curve has sudden jumps or spikes at certain positions +1. Please upgrade to **v0.9.0 or later**. +2. First try using the best-performing **rmvpe-onnx** or **hybrid** backend (with default confidence thresholds). + If the issue persists, try lowering both confidence thresholds. You can use the pitch confidence curve in [`expressive-viewer`](#viewer) as a reference when tuning. + In general, the **Utau vocal** is relatively clean, so it is recommended to adjust the confidence threshold for the **reference vocal** first. + +#### Future Plans +Incorporate semantic information into the PITD expression extraction algorithm. + +--- + +### PITD expression curve has sudden jumps or spikes #### Symptom -The PITD expression curve changes too rapidly at certain positions, with very large jumps or spikes that clearly do not match natural vocal behavior. -#### Possible Cause -The two confidence thresholds in the PITD extractor are set **too low**, causing erroneous detection results to be accepted. +The PITD expression curve changes too abruptly at certain positions, with large jumps or spikes that do not match natural vocal behavior. -#### Solution -First try using the best-performing rmvpe-onnx backend (with default confidence thresholds). If the issue persists, try increasing both confidence thresholds. In general, the **Utau vocal** is relatively clean, so it is advisable to first adjust the confidence threshold for the **Reference vocal**. +#### Possible Causes + +1. In versions earlier than **v0.9.0**, there is an issue in the conversion between pitch and PITD values. +2. There is noise in the reference audio around the corresponding timestamps. + You can observe abnormal spikes in the original pitch curve in [`expressive-viewer`](#viewer) +3. The two confidence thresholds in the PITD extractor are set **too low**, causing incorrect detections to be accepted. + This may also appear as spikes in the pitch curve in [`expressive-viewer`](#viewer). + +#### Solutions + +1. Please upgrade to **v0.9.0 or later**. +2. Try denoising the reference audio using tools such as [UVR](https://github.com/Anjok07/ultimatevocalremovergui) or [MSST](https://github.com/SUC-DriverOld/MSST-WebUI). +3. First try using the best-performing **rmvpe-onnx** or **hybrid** backend (with default confidence thresholds). + If the issue persists, try increasing both confidence thresholds. You can use the pitch confidence curve in [`expressive-viewer`](#viewer) to guide your adjustments. + In general, the **Utau vocal** is relatively clean, so it is recommended to adjust the confidence threshold for the **reference vocal** first. + +#### Future Plans +Incorporate semantic information into the PITD expression extraction algorithm. diff --git a/README.md b/README.md index 0b08bda..9103852 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,15 @@

+> [!WARNING] +> 🚨 **下载前请注意** 🚨 +> +> **强烈建议您使用 `v0.9.0` 及以上版本**。早先版本的 **PITD** 表情参数处理算法存在[严重缺陷](https://github.com/NewComer00/expressive/releases/tag/v0.9.0),会导致**音高曲线绘制错误**。下载最新版本请前往 [Releases 页面](https://github.com/NewComer00/expressive/releases)。 +> +> 对于从旧版本迁移到 `v0.9.0` 及以上版本的用户,新版本中 **PITD 缩放因子(Scaler)的默认值为 `1.0`**,不再是原来的 `2.0`。若您有旧版本的配置文件,请将 **PITD 缩放因子(Scaler)设置为 `1.0`**。 +> +> **🎵 感谢您使用 Expressive🎵** + # Expressive **Expressive** 是一个为 [OpenUtau](https://github.com/stakira/OpenUtau) 开发的 [DiffSinger](https://github.com/openvpi/diffsinger) 表情参数导入工具,旨在从真实人声中提取表情参数,并导入至工程的相应轨道。 @@ -21,18 +30,17 @@ | **工作流程** | **数据可视化** | |:---:|:---:| -| | | +| | | -> - *OpenUtau 版本来自 [keirokeer/OpenUtau-DiffSinger-Lunai](https://github.com/keirokeer/OpenUtau-DiffSinger-Lunai)* -> - *歌手模型来自 [yousa-ling-official-production/yousa-ling-diffsinger-v1](https://github.com/yousa-ling-official-production/yousa-ling-diffsinger-v1)* +> - *示例来自 [`examples/明天会更好`](examples/明天会更好),点击查看详情信息* > [!TIP] >
> 👉 点击展开完整有声演示视频 👈 > ->

+>

>

> >
@@ -47,6 +55,8 @@ 应用也提供了 [swift-f0](https://github.com/lars76/swift-f0) 与 [CREPE](https://github.com/marl/crepe) 音高提取后端。前者仅依赖 CPU,效果一般,但速度最快。后者是业内的经典算法,速度较慢。在 CUDA 环境下,CREPE 后端会自动启用 GPU 加速。 +应用还新增了一个实验性的 **hybrid** 后端。该后端融合了 rmvpe-onnx 与 swift-f0 的预测结果,以 rmvpe-onnx 的音高提取结果为主,在音频有声段中,如果 rmvpe-onnx 的置信度较低且 swift-f0 的置信度较高,则采用 swift-f0 的结果进行修正,从而提升整体音高提取的准确性。 + > \* 在 Windows 平台下,TensorFlow 2.10 是最后一个支持 GPU 加速的版本,Python 3.10 是它的 `.whl` 文件支持的最高 Python 版本。 ## 📌 使用场景 @@ -302,16 +312,25 @@ graph TB; #### 未来计划 NiceGUI 框架已经开始着手改进文件拖拽支持,应该在未来的版本中能够解决此问题。 +--- + ### PITD 表情曲线整体变化过于平缓 #### 问题现象 提取出的 PITD 表情曲线过于平缓,整体上几乎没有大的起伏,参考人声中的音高变化并没有反映到表情曲线上。 #### 可能原因 -PITD 表情提取器中,两个置信度阈值设置**过高**,许多音高变化没有被采信。 +1. 在早于 `v0.9.0` 的版本中,PITD 表情曲线取值与音高之间的换算有问题,会导致 PITD 表情曲线整体非常平缓。 +2. PITD 表情提取器中,两个置信度阈值设置**过高**,许多音高变化没有被采信。您可以在 [`expressive-viewer`](#可视化工具viewer) 中观察到,原始的音高曲线中有很多不该出现的缺失部分。 #### 解决方案 -请先尝试使用效果最好的 rmvpe-onnx 后端(默认置信度阈值)。若问题仍在,尝试降低两个置信度阈值。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。 +1. 请下载安装 `v0.9.0` 及之后的版本。 +2. 请先尝试使用效果最好的 rmvpe-onnx 或 hybrid 后端(默认置信度阈值)。若问题仍在,尝试降低两个置信度阈值。您可以参考 [`expressive-viewer`](#可视化工具viewer) 的音高置信度曲线来辅助调整。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。 + +#### 未来计划 +为 PITD 表情提取算法引入语义信息。 + +--- ### PITD 表情曲线在某些位置变化过快,出现跳跃或毛刺 @@ -319,7 +338,14 @@ PITD 表情提取器中,两个置信度阈值设置**过高**,许多音高 PITD 表情曲线在某些位置变化过快,出现非常大的跳跃或毛刺,明显不符合人声的变化规律。 #### 可能原因 -PITD 表情提取器中,两个置信度阈值设置**过低**,错误的识别结果被采信。 +1. 在早于 `v0.9.0` 的版本中,PITD 表情曲线取值与音高之间的换算有问题。 +2. 参考音频的对应时间戳附近有噪声。您可以在 [`expressive-viewer`](#可视化工具viewer) 中观察到,原始的音高曲线中有很多不该出现的尖刺。 +3. PITD 表情提取器中,两个置信度阈值设置**过低**,错误的识别结果被采信。您可以在 [`expressive-viewer`](#可视化工具viewer) 中观察到,原始的音高曲线中有很多不该出现的尖刺。 #### 解决方案 -请先尝试使用效果最好的 rmvpe-onnx 后端(默认置信度阈值)。若问题仍在,尝试增加两个置信度阈值。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。 +1. 请下载安装 `v0.9.0` 及之后的版本。 +2. 可使用 [UVR](https://github.com/Anjok07/ultimatevocalremovergui) 、[MSST](https://github.com/SUC-DriverOld/MSST-WebUI) 等工具对参考音频去噪声(denoise)。 +3. 请先尝试使用效果最好的 rmvpe-onnx 或 hybrid 后端(默认置信度阈值)。若问题仍在,尝试增加两个置信度阈值。您可以参考 [`expressive-viewer`](#可视化工具viewer) 的音高置信度曲线来辅助调整。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。 + +#### 未来计划 +为 PITD 表情提取算法引入语义信息。 diff --git "a/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json" "b/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json" index 2f6e244..1dd2b38 100644 --- "a/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json" +++ "b/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json" @@ -18,13 +18,13 @@ }, "pitd": { "selected": true, - "backend": "rmvpe-onnx", + "backend": "hybrid", "confidence_utau": null, "confidence_ref": null, "align_radius": 1, "semitone_shift": 0, - "smoothness": 4, - "scaler": 2.2 + "smoothness": 2, + "scaler": 1.0 }, "tenc": { "selected": true, @@ -32,7 +32,7 @@ "align_radius": 1, "smoothness": 6, "scaler": 1.0, - "bias": 10 + "bias": 15 } } } diff --git "a/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json" "b/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json" index cd5bb2a..813ec20 100644 --- "a/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json" +++ "b/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json" @@ -18,13 +18,13 @@ }, "pitd": { "selected": true, - "backend": "rmvpe-onnx", + "backend": "hybrid", "confidence_utau": null, "confidence_ref": null, "align_radius": 1, "semitone_shift": 0, "smoothness": 2, - "scaler": 2.0 + "scaler": 1.0 }, "tenc": { "selected": true, diff --git "a/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json" "b/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json" index 73a4124..f7418ae 100644 --- "a/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json" +++ "b/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json" @@ -18,13 +18,13 @@ }, "pitd": { "selected": true, - "backend": "rmvpe-onnx", + "backend": "hybrid", "confidence_utau": null, "confidence_ref": null, "align_radius": 1, "semitone_shift": 0, "smoothness": 2, - "scaler": 2.0 + "scaler": 1.0 }, "tenc": { "selected": true, @@ -32,7 +32,7 @@ "align_radius": 1, "smoothness": 6, "scaler": 1.2, - "bias": 10 + "bias": 15 } } } diff --git a/expressions/dyn.py b/expressions/dyn.py index cbb331a..3df9cd8 100644 --- a/expressions/dyn.py +++ b/expressions/dyn.py @@ -10,6 +10,7 @@ register_expression ) from utils.seqtool import ( + seq_spline_smoothing, unify_sequence_time, align_sequence_tick, gaussian_filter1d_with_nan, @@ -24,10 +25,11 @@ class DynLoader(ExpressionLoader): expression_name = "dyn" expression_info = _l("Dynamics (curve)") args = SimpleNamespace( - trim_silence = Args(name="trim_silence", type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501 - align_radius = Args(name="align_radius", type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501 - smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501 - scaler = Args(name="scaler" , type=float, default=1.5 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501 + trim_silence = Args(name="trim_silence" , type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501 + align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501 + smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501 + scaler = Args(name="scaler" , type=float, default=1.5 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501 + spline_smoothing = Args(name="spline_smoothing", type=bool , default=True, help=_l("Perform **spline smoothing** on the final expression curve for extra smoothness")), # noqa: E501 ) plots = SimpleNamespace( expression = Plot(tag=expression_info , title=expression_info , x_label=_l("Tick") , y_label=expression_name, legends=[expression_name] ), # noqa: E501 @@ -37,10 +39,11 @@ class DynLoader(ExpressionLoader): def get_expression( self, - trim_silence = args.trim_silence.default, - align_radius = args.align_radius.default, - smoothness = args.smoothness .default, - scaler = args.scaler .default, + trim_silence = args.trim_silence .default, + align_radius = args.align_radius .default, + smoothness = args.smoothness .default, + scaler = args.scaler .default, + spline_smoothing = args.spline_smoothing.default, ): self.logger.info(_("Extracting expression...")) @@ -68,6 +71,12 @@ def get_expression( # Generate expression curve dyn_val = get_experssion_dynamics(time_aligned_ref_rms, smoothness, scaler) + if spline_smoothing: + # Final spline smoothing of the expression curve + # NOTE: All NaN positions except the leading and trailing ones will be interpolated + # Only preserving NaN at the head/tail to avoid edge artifacts of spline smoothing + dyn_val = seq_spline_smoothing(dyn_tick, dyn_val, nan_policy='preserve_head_tail') + # Collect plots self.collect_plot(self.plots.expression, (dyn_tick, dyn_val)) self.collect_plot(self.plots.raw_rms, (ref_time, ref_rms), (utau_time, utau_rms)) diff --git a/expressions/pitd.py b/expressions/pitd.py index fc6bd82..5ae40dd 100644 --- a/expressions/pitd.py +++ b/expressions/pitd.py @@ -12,6 +12,7 @@ ) from utils.i18n import _, _l, _lf from utils.seqtool import ( + seq_spline_smoothing, unify_sequence_time, align_sequence_tick, gaussian_filter1d_with_nan, @@ -29,17 +30,19 @@ class PitdLoader(ExpressionLoader): "rmvpe-onnx": _l("finest accuracy, fast, CPU only (ONNX Runtime)"), "swift-f0": _l("fair accuracy, fastest, CPU only (ONNX Runtime)"), "crepe": _l("good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)"), + "hybrid": _l("based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)"), } - confidence_utau_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.95, "crepe": 0.80} - confidence_ref_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.93, "crepe": 0.60} + confidence_utau_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.95, "crepe": 0.80, "hybrid": 0.03} + confidence_ref_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.93, "crepe": 0.60, "hybrid": 0.03} args = SimpleNamespace( - backend = Args(name="backend" , type=str , default="rmvpe-onnx", choices=list(backend_choices.keys()), help=_lf("**F0 detection backend** for extracting pitch from WAV files. Available options:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.backend_choices.items()]))), # noqa: E501 - confidence_utau = Args(name="confidence_utau", type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **UTAU** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_utau_recommended.items()]))), # noqa: E501 - confidence_ref = Args(name="confidence_ref" , type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **reference** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_ref_recommended.items()]))), # noqa: E501 - align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501 - semitone_shift = Args(name="semitone_shift" , type=int , default=None, help=_l("**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is an octave higher than the reference WAV, set to 12; if lower, set to -12. Omit to enable automatic shift estimation")), # noqa: E501 - smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501 - scaler = Args(name="scaler" , type=float, default=2.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501 + backend = Args(name="backend" , type=str , default="rmvpe-onnx", choices=list(backend_choices.keys()), help=_lf("**F0 detection backend** for extracting pitch from WAV files. Available options:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.backend_choices.items()]))), # noqa: E501 + confidence_utau = Args(name="confidence_utau" , type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **UTAU** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_utau_recommended.items()]))), # noqa: E501 + confidence_ref = Args(name="confidence_ref" , type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **reference** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_ref_recommended.items()]))), # noqa: E501 + align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501 + semitone_shift = Args(name="semitone_shift" , type=int , default=None, help=_l("**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is an octave higher than the reference WAV, set to 12; if lower, set to -12. Omit to enable automatic shift estimation")), # noqa: E501 + smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501 + scaler = Args(name="scaler" , type=float, default=1.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501 + spline_smoothing = Args(name="spline_smoothing", type=bool , default=True, help=_l("Perform **spline smoothing** on the final expression curve for extra smoothness")), # noqa: E501 ) plots = SimpleNamespace( expression = Plot(tag=expression_info , title=expression_info , x_label=_l("Tick") , y_label=expression_name , legends=[expression_name] ), # noqa: E501 @@ -50,13 +53,14 @@ class PitdLoader(ExpressionLoader): def get_expression( self, - backend = args.backend .default, - confidence_utau = args.confidence_utau.default, - confidence_ref = args.confidence_ref .default, - align_radius = args.align_radius .default, - semitone_shift = args.semitone_shift .default, - smoothness = args.smoothness .default, - scaler = args.scaler .default, + backend = args.backend .default, + confidence_utau = args.confidence_utau .default, + confidence_ref = args.confidence_ref .default, + align_radius = args.align_radius .default, + semitone_shift = args.semitone_shift .default, + smoothness = args.smoothness .default, + scaler = args.scaler .default, + spline_smoothing = args.spline_smoothing.default, ): self.logger.info(_("Extracting expression...")) @@ -94,16 +98,22 @@ def get_expression( time_aligned_ref_pitch, unified_utau_pitch, semitone_shift=semitone_shift, - smoothness=smoothness, ) # Calculate pitch delta for USTX pitch editing pitd_val = get_pitch_delta( time_pitch_aligned_ref_pitch, unified_utau_pitch, + smoothness=smoothness, scaler=scaler, ) + if spline_smoothing: + # Final spline smoothing of the expression curve + # NOTE: All NaN positions except the leading and trailing ones will be interpolated + # Only preserving NaN at the head/tail to avoid edge artifacts of spline smoothing + pitd_val = seq_spline_smoothing(pitd_tick, pitd_val, nan_policy='preserve_head_tail') + # Collect plots self.collect_plot(self.plots.expression, (pitd_tick, pitd_val)) self.collect_plot(self.plots.confidence, (ref_time, ref_confidence), (utau_time, utau_confidence)) @@ -167,7 +177,7 @@ def get_wav_features(wav_path, backend="rmvpe-onnx", confidence_threshold=0.8, c return wav_time, wav_pitch, wav_confidence, wav_features -def align_sequence_pitch(query, reference, semitone_shift=None, smoothness=0): +def align_sequence_pitch(query, reference, semitone_shift=None): """Align pitch sequences by shifting in semitones and applying smoothing. Args: @@ -175,7 +185,6 @@ def align_sequence_pitch(query, reference, semitone_shift=None, smoothness=0): reference (numpy.ndarray): Target reference pitch values. semitone_shift (int, optional): Semitones to shift the query pitch. If None, estimated automatically. - smoothness (int, optional): Smoothing sigma. Defaults to 0. Returns: tuple: (pitch_aligned_query, semitone_shift) @@ -189,22 +198,30 @@ def align_sequence_pitch(query, reference, semitone_shift=None, smoothness=0): ) print(_("Estimated Semitone-shift: {}").format(semitone_shift)) - pitch_aligned_query = gaussian_filter1d_with_nan( - query * np.exp2(semitone_shift / 12), - sigma=smoothness, - ) + pitch_aligned_query = query * np.exp2(semitone_shift / 12) return pitch_aligned_query, semitone_shift -def get_pitch_delta(query, reference, scaler=2.5): +def get_pitch_delta(query, reference, smoothness=2, scaler=1.0): """Calculate the scaled pitch difference between two sequences. + PITD is expressed in cents (100 cents = 1 semitone). + The renderer applies PITD on top of the base pitch, so the delta + must be in the same unit OpenUtau expects for the PITD expression. + Args: - query (numpy.ndarray): Pitch values from the query sequence. - reference (numpy.ndarray): Pitch values from the reference sequence. - scaler (float, optional): Scaling factor. Defaults to 2.5. + query (numpy.ndarray): Pitch values from the query sequence. + reference (numpy.ndarray): Pitch values from the reference sequence. + smoothness (int, optional): Smoothing sigma. Defaults to 2. + scaler (float, optional): Scaling factor. Defaults to 1.0. Returns: - numpy.ndarray: Scaled pitch difference values. + numpy.ndarray: Scaled pitch difference in cents, preserving NaN for unvoiced frames. """ - return scaler * (query - reference) + voiced = (query > 0) & (reference > 0) + + delta = np.full_like(query, fill_value=np.nan) + delta[voiced] = 1200.0 * np.log2(query[voiced] / reference[voiced]) + + delta = gaussian_filter1d_with_nan(delta, sigma=smoothness) + return scaler * delta diff --git a/expressions/tenc.py b/expressions/tenc.py index 730d56e..9078a49 100644 --- a/expressions/tenc.py +++ b/expressions/tenc.py @@ -10,6 +10,7 @@ register_expression ) from utils.seqtool import ( + seq_spline_smoothing, unify_sequence_time, align_sequence_tick, gaussian_filter1d_with_nan, @@ -24,11 +25,12 @@ class TencLoader(ExpressionLoader): expression_name = "tenc" expression_info = _l("Tension (curve)") args = SimpleNamespace( - trim_silence = Args(name="trim_silence", type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501 - align_radius = Args(name="align_radius", type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501 - smoothness = Args(name="smoothness" , type=int , default=6 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501 - scaler = Args(name="scaler" , type=float, default=1.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501 - bias = Args(name="bias" , type=int , default=10 , help=_l("**Bias** offset added to the expression curve. Positive values shift the curve upward; negative values shift it downward")), # noqa: E501 + trim_silence = Args(name="trim_silence" , type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501 + align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501 + smoothness = Args(name="smoothness" , type=int , default=6 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501 + scaler = Args(name="scaler" , type=float, default=1.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501 + bias = Args(name="bias" , type=int , default=10 , help=_l("**Bias** offset added to the expression curve. Positive values shift the curve upward; negative values shift it downward")), # noqa: E501 + spline_smoothing = Args(name="spline_smoothing", type=bool , default=True, help=_l("Perform **spline smoothing** on the final expression curve for extra smoothness")), # noqa: E501 ) plots = SimpleNamespace( expression = Plot(tag=expression_info , title=expression_info , x_label=_l("Tick") , y_label=expression_name, legends=[expression_name] ), # noqa: E501 @@ -38,11 +40,12 @@ class TencLoader(ExpressionLoader): def get_expression( self, - trim_silence = args.trim_silence.default, - align_radius = args.align_radius.default, - smoothness = args.smoothness .default, - scaler = args.scaler .default, - bias = args.bias .default, + trim_silence = args.trim_silence .default, + align_radius = args.align_radius .default, + smoothness = args.smoothness .default, + scaler = args.scaler .default, + bias = args.bias .default, + spline_smoothing = args.spline_smoothing.default, ): self.logger.info(_("Extracting expression...")) @@ -70,6 +73,12 @@ def get_expression( # Generate expression curve tenc_val = get_experssion_tension(time_aligned_ref_rms, smoothness, scaler, bias) + if spline_smoothing: + # Final spline smoothing of the expression curve + # NOTE: All NaN positions except the leading and trailing ones will be interpolated + # Only preserving NaN at the head/tail to avoid edge artifacts of spline smoothing + tenc_val = seq_spline_smoothing(tenc_tick, tenc_val, nan_policy='preserve_head_tail') + # Collect plots self.collect_plot(self.plots.expression, (tenc_tick, tenc_val)) self.collect_plot(self.plots.raw_rms, (ref_time, ref_rms), (utau_time, utau_rms)) diff --git a/expressive_gui.py b/expressive_gui.py index 58b4416..3d8a392 100644 --- a/expressive_gui.py +++ b/expressive_gui.py @@ -543,12 +543,15 @@ async def update_placeholders(): with ui.card().classes("w-full").bind_visibility_from( state["expressions"]["dyn"], "selected" ): - with ui.row().classes("w-full"): - ui.label(dyn_info).classes("text-lg font-bold") - ui.space() + ui.label(dyn_info).classes("text-lg font-bold") + + with ui.grid(columns=2).classes("w-full"): ui.switch(_("Trim Silence")).bind_value( state["expressions"]["dyn"], "trim_silence", ).tooltip_md(dyn_args.trim_silence.help) + ui.switch(_("Spline Smoothing")).bind_value( + state["expressions"]["dyn"], "spline_smoothing", + ).tooltip_md(dyn_args.spline_smoothing.help) with ui.grid(columns=3).classes("w-full"): ui.number(label=_("Align Radius"), min=1, format="%d").bind_value( @@ -574,6 +577,11 @@ async def update_placeholders(): ): ui.label(pitd_info).classes("text-lg font-bold") + with ui.grid(columns=2).classes("w-full"): + ui.switch(_("Spline Smoothing")).bind_value( + state["expressions"]["pitd"], "spline_smoothing", + ).tooltip_md(pitd_args.spline_smoothing.help) + with ui.grid(columns=3).classes("w-full"): def on_backend_change(e): nonlocal ui_confidence_utau, ui_confidence_ref @@ -633,12 +641,15 @@ def on_backend_change(e): with ui.card().classes("w-full").bind_visibility_from( state["expressions"]["tenc"], "selected" ): - with ui.row().classes("w-full"): - ui.label(tenc_info).classes("text-lg font-bold") - ui.space() + ui.label(tenc_info).classes("text-lg font-bold") + + with ui.grid(columns=2).classes("w-full"): ui.switch(_("Trim Silence")).bind_value( state["expressions"]["tenc"], "trim_silence", ).tooltip_md(tenc_args.trim_silence.help) + ui.switch(_("Spline Smoothing")).bind_value( + state["expressions"]["tenc"], "spline_smoothing", + ).tooltip_md(tenc_args.spline_smoothing.help) with ui.grid(columns=3).classes("w-full"): ui.number(label=_("Align Radius"), min=1, format="%d").bind_value( diff --git a/locales/app.pot b/locales/app.pot index e9c0b29..12d076e 100644 --- a/locales/app.pot +++ b/locales/app.pot @@ -8,7 +8,7 @@ msgid "" msgstr "" "Project-Id-Version: expressive VERSION\n" "Report-Msgid-Bugs-To: EMAIL@ADDRESS\n" -"POT-Creation-Date: 2026-04-11 19:51+0800\n" +"POT-Creation-Date: 2026-04-15 21:12+0800\n" "PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n" "Last-Translator: FULL NAME \n" "Language-Team: LANGUAGE \n" @@ -148,59 +148,63 @@ msgstr "" msgid "Expression Selection" msgstr "" -#: expressive_gui.py:549 expressive_gui.py:639 +#: expressive_gui.py:549 expressive_gui.py:647 msgid "Trim Silence" msgstr "" -#: expressive_gui.py:554 expressive_gui.py:609 expressive_gui.py:644 +#: expressive_gui.py:552 expressive_gui.py:581 expressive_gui.py:650 +msgid "Spline Smoothing" +msgstr "" + +#: expressive_gui.py:557 expressive_gui.py:617 expressive_gui.py:655 msgid "Align Radius" msgstr "" -#: expressive_gui.py:559 expressive_gui.py:620 expressive_gui.py:649 +#: expressive_gui.py:562 expressive_gui.py:628 expressive_gui.py:660 msgid "Smoothness" msgstr "" -#: expressive_gui.py:564 expressive_gui.py:625 expressive_gui.py:654 +#: expressive_gui.py:567 expressive_gui.py:633 expressive_gui.py:665 msgid "Scaler" msgstr "" -#: expressive_gui.py:593 +#: expressive_gui.py:601 msgid "UTAU Confidence" msgstr "" -#: expressive_gui.py:599 +#: expressive_gui.py:607 msgid "Reference Confidence" msgstr "" -#: expressive_gui.py:605 +#: expressive_gui.py:613 msgid "Backend" msgstr "" -#: expressive_gui.py:614 +#: expressive_gui.py:622 msgid "Semitone Shift" msgstr "" -#: expressive_gui.py:615 +#: expressive_gui.py:623 msgid "Auto Estimation" msgstr "" -#: expressive_gui.py:659 +#: expressive_gui.py:670 msgid "Bias" msgstr "" -#: expressive_gui.py:667 +#: expressive_gui.py:678 msgid "Import Config" msgstr "" -#: expressive_gui.py:673 +#: expressive_gui.py:684 msgid "Export Config" msgstr "" -#: expressive_gui.py:684 +#: expressive_gui.py:695 msgid "Process" msgstr "" -#: expressive_gui.py:726 +#: expressive_gui.py:737 msgid "Migrate expressions from real singers to DiffSingers (GUI)" msgstr "" @@ -279,102 +283,112 @@ msgstr "" msgid "Expression result is empty. Skipping USTX update." msgstr "" -#: expressions/dyn.py:25 +#: expressions/dyn.py:26 msgid "Dynamics (curve)" msgstr "" -#: expressions/dyn.py:27 expressions/tenc.py:27 +#: expressions/dyn.py:28 expressions/tenc.py:28 msgid "" "**Trim silence** from the leading and trailing edges of the audio before " "extracting expression" msgstr "" -#: expressions/dyn.py:28 expressions/pitd.py:39 expressions/tenc.py:28 +#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29 msgid "" "**Radius** for the FastDTW alignment algorithm; larger values allow more " "flexible alignment but increase computation time" msgstr "" -#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29 +#: expressions/dyn.py:30 expressions/pitd.py:43 expressions/tenc.py:30 msgid "" "Controls the **smoothness** of the expression curve using Gaussian " "filtering. Higher values produce smoother curves but may lose fine detail" msgstr "" -#: expressions/dyn.py:30 expressions/pitd.py:42 expressions/tenc.py:30 +#: expressions/dyn.py:31 expressions/pitd.py:44 expressions/tenc.py:31 msgid "" "**Scaling factor** applied to the expression curve. Values >1 amplify the" " expression, =1 keeps original intensity, <1 reduces it" msgstr "" -#: expressions/dyn.py:33 expressions/dyn.py:35 expressions/pitd.py:45 -#: expressions/pitd.py:48 expressions/tenc.py:34 expressions/tenc.py:36 +#: expressions/dyn.py:32 expressions/pitd.py:45 expressions/tenc.py:33 +msgid "" +"Perform **spline smoothing** on the final expression curve for extra " +"smoothness" +msgstr "" + +#: expressions/dyn.py:35 expressions/dyn.py:37 expressions/pitd.py:48 +#: expressions/pitd.py:51 expressions/tenc.py:36 expressions/tenc.py:38 msgid "Tick" msgstr "" -#: expressions/dyn.py:34 expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/tenc.py:37 msgid "raw_rms" msgstr "" -#: expressions/dyn.py:34 expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/tenc.py:37 msgid "Raw RMS" msgstr "" -#: expressions/dyn.py:34 expressions/pitd.py:46 expressions/pitd.py:47 -#: expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/pitd.py:49 expressions/pitd.py:50 +#: expressions/tenc.py:37 msgid "Time (s)" msgstr "" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "RMS" msgstr "" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46 -#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49 +#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "Reference" msgstr "" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46 -#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49 +#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "UTAU" msgstr "" -#: expressions/dyn.py:35 expressions/tenc.py:36 +#: expressions/dyn.py:37 expressions/tenc.py:38 msgid "aligned_rms" msgstr "" -#: expressions/dyn.py:35 expressions/tenc.py:36 +#: expressions/dyn.py:37 expressions/tenc.py:38 msgid "Aligned RMS" msgstr "" -#: expressions/dyn.py:45 expressions/pitd.py:61 expressions/tenc.py:47 +#: expressions/dyn.py:48 expressions/pitd.py:65 expressions/tenc.py:50 msgid "Extracting expression..." msgstr "" -#: expressions/dyn.py:77 expressions/pitd.py:114 expressions/tenc.py:79 +#: expressions/dyn.py:86 expressions/pitd.py:124 expressions/tenc.py:88 msgid "Expression extraction complete." msgstr "" -#: expressions/pitd.py:27 +#: expressions/pitd.py:28 msgid "Pitch Deviation (curve)" msgstr "" -#: expressions/pitd.py:29 +#: expressions/pitd.py:30 msgid "finest accuracy, fast, CPU only (ONNX Runtime)" msgstr "" -#: expressions/pitd.py:30 +#: expressions/pitd.py:31 msgid "fair accuracy, fastest, CPU only (ONNX Runtime)" msgstr "" -#: expressions/pitd.py:31 +#: expressions/pitd.py:32 msgid "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)" msgstr "" -#: expressions/pitd.py:36 +#: expressions/pitd.py:33 +msgid "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)" +msgstr "" + +#: expressions/pitd.py:38 #, python-format msgid "" "**F0 detection backend** for extracting pitch from WAV files. Available " @@ -384,7 +398,7 @@ msgid "" "\n" msgstr "" -#: expressions/pitd.py:37 +#: expressions/pitd.py:39 #, python-format msgid "" "Minimum **confidence level** for keeping detected pitch values in the " @@ -395,7 +409,7 @@ msgid "" "\n" msgstr "" -#: expressions/pitd.py:38 +#: expressions/pitd.py:40 #, python-format msgid "" "Minimum **confidence level** for keeping detected pitch values in the " @@ -406,55 +420,55 @@ msgid "" "\n" msgstr "" -#: expressions/pitd.py:40 +#: expressions/pitd.py:42 msgid "" "**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is" " an octave higher than the reference WAV, set to 12; if lower, set to " "-12. Omit to enable automatic shift estimation" msgstr "" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "confidence" msgstr "" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "Pitch Extraction Confidence" msgstr "" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "Confidence" msgstr "" -#: expressions/pitd.py:47 +#: expressions/pitd.py:50 msgid "raw_pitch" msgstr "" -#: expressions/pitd.py:47 +#: expressions/pitd.py:50 msgid "Raw Pitch" msgstr "" -#: expressions/pitd.py:47 expressions/pitd.py:48 +#: expressions/pitd.py:50 expressions/pitd.py:51 msgid "Pitch (Hz)" msgstr "" -#: expressions/pitd.py:48 +#: expressions/pitd.py:51 msgid "aligned_pitch" msgstr "" -#: expressions/pitd.py:48 +#: expressions/pitd.py:51 msgid "Aligned Pitch" msgstr "" -#: expressions/pitd.py:190 +#: expressions/pitd.py:199 #, python-brace-format msgid "Estimated Semitone-shift: {}" msgstr "" -#: expressions/tenc.py:25 +#: expressions/tenc.py:26 msgid "Tension (curve)" msgstr "" -#: expressions/tenc.py:31 +#: expressions/tenc.py:32 msgid "" "**Bias** offset added to the expression curve. Positive values shift the " "curve upward; negative values shift it downward" @@ -472,22 +486,22 @@ msgstr "" msgid "Zoom" msgstr "" -#: utils/wavtool.py:87 +#: utils/wavtool.py:88 #, python-brace-format msgid "Loading F0 data from cache file: '{}'" msgstr "" -#: utils/wavtool.py:130 +#: utils/wavtool.py:133 #, python-brace-format msgid "F0 data saved to cache file: '{}'" msgstr "" -#: utils/wavtool.py:304 +#: utils/wavtool.py:380 #, python-brace-format msgid "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" msgstr "" -#: utils/wavtool.py:310 +#: utils/wavtool.py:386 #, python-brace-format msgid "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" msgstr "" diff --git a/locales/en/LC_MESSAGES/app.po b/locales/en/LC_MESSAGES/app.po index 7c47ad1..2621c73 100644 --- a/locales/en/LC_MESSAGES/app.po +++ b/locales/en/LC_MESSAGES/app.po @@ -7,7 +7,7 @@ msgid "" msgstr "" "Project-Id-Version: expressive\n" "Report-Msgid-Bugs-To: https://github.com/NewComer00/expressive/issues\n" -"POT-Creation-Date: 2026-04-11 19:51+0800\n" +"POT-Creation-Date: 2026-04-15 21:12+0800\n" "PO-Revision-Date: 2026-03-01 20:21+0800\n" "Last-Translator: NewComer00\n" "Language: en\n" @@ -151,59 +151,63 @@ msgstr "Track Number" msgid "Expression Selection" msgstr "Expression Selection" -#: expressive_gui.py:549 expressive_gui.py:639 +#: expressive_gui.py:549 expressive_gui.py:647 msgid "Trim Silence" msgstr "Trim Silence" -#: expressive_gui.py:554 expressive_gui.py:609 expressive_gui.py:644 +#: expressive_gui.py:552 expressive_gui.py:581 expressive_gui.py:650 +msgid "Spline Smoothing" +msgstr "Spline Smoothing" + +#: expressive_gui.py:557 expressive_gui.py:617 expressive_gui.py:655 msgid "Align Radius" msgstr "Align Radius" -#: expressive_gui.py:559 expressive_gui.py:620 expressive_gui.py:649 +#: expressive_gui.py:562 expressive_gui.py:628 expressive_gui.py:660 msgid "Smoothness" msgstr "Smoothness" -#: expressive_gui.py:564 expressive_gui.py:625 expressive_gui.py:654 +#: expressive_gui.py:567 expressive_gui.py:633 expressive_gui.py:665 msgid "Scaler" msgstr "Scaler" -#: expressive_gui.py:593 +#: expressive_gui.py:601 msgid "UTAU Confidence" msgstr "UTAU Confidence" -#: expressive_gui.py:599 +#: expressive_gui.py:607 msgid "Reference Confidence" msgstr "Reference Confidence" -#: expressive_gui.py:605 +#: expressive_gui.py:613 msgid "Backend" msgstr "Backend" -#: expressive_gui.py:614 +#: expressive_gui.py:622 msgid "Semitone Shift" msgstr "Semitone Shift" -#: expressive_gui.py:615 +#: expressive_gui.py:623 msgid "Auto Estimation" msgstr "Auto Estimation" -#: expressive_gui.py:659 +#: expressive_gui.py:670 msgid "Bias" msgstr "Bias" -#: expressive_gui.py:667 +#: expressive_gui.py:678 msgid "Import Config" msgstr "Import Config" -#: expressive_gui.py:673 +#: expressive_gui.py:684 msgid "Export Config" msgstr "Export Config" -#: expressive_gui.py:684 +#: expressive_gui.py:695 msgid "Process" msgstr "Process" -#: expressive_gui.py:726 +#: expressive_gui.py:737 msgid "Migrate expressions from real singers to DiffSingers (GUI)" msgstr "Migrate expressions from real singers to DiffSingers (GUI)" @@ -290,11 +294,11 @@ msgstr "Expression written to USTX file: '{}'" msgid "Expression result is empty. Skipping USTX update." msgstr "Expression result is empty. Skipping USTX update." -#: expressions/dyn.py:25 +#: expressions/dyn.py:26 msgid "Dynamics (curve)" msgstr "Dynamics (curve)" -#: expressions/dyn.py:27 expressions/tenc.py:27 +#: expressions/dyn.py:28 expressions/tenc.py:28 msgid "" "**Trim silence** from the leading and trailing edges of the audio before " "extracting expression" @@ -302,7 +306,7 @@ msgstr "" "**Trim silence** from the leading and trailing edges of the audio before " "extracting expression" -#: expressions/dyn.py:28 expressions/pitd.py:39 expressions/tenc.py:28 +#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29 msgid "" "**Radius** for the FastDTW alignment algorithm; larger values allow more " "flexible alignment but increase computation time" @@ -310,7 +314,7 @@ msgstr "" "**Radius** for the FastDTW alignment algorithm; larger values allow more " "flexible alignment but increase computation time" -#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29 +#: expressions/dyn.py:30 expressions/pitd.py:43 expressions/tenc.py:30 msgid "" "Controls the **smoothness** of the expression curve using Gaussian " "filtering. Higher values produce smoother curves but may lose fine detail" @@ -318,7 +322,7 @@ msgstr "" "Controls the **smoothness** of the expression curve using Gaussian " "filtering. Higher values produce smoother curves but may lose fine detail" -#: expressions/dyn.py:30 expressions/pitd.py:42 expressions/tenc.py:30 +#: expressions/dyn.py:31 expressions/pitd.py:44 expressions/tenc.py:31 msgid "" "**Scaling factor** applied to the expression curve. Values >1 amplify the" " expression, =1 keeps original intensity, <1 reduces it" @@ -326,74 +330,86 @@ msgstr "" "**Scaling factor** applied to the expression curve. Values >1 amplify the" " expression, =1 keeps original intensity, <1 reduces it" -#: expressions/dyn.py:33 expressions/dyn.py:35 expressions/pitd.py:45 -#: expressions/pitd.py:48 expressions/tenc.py:34 expressions/tenc.py:36 +#: expressions/dyn.py:32 expressions/pitd.py:45 expressions/tenc.py:33 +msgid "" +"Perform **spline smoothing** on the final expression curve for extra " +"smoothness" +msgstr "" +"Perform **spline smoothing** on the final expression curve for extra " +"smoothness" + +#: expressions/dyn.py:35 expressions/dyn.py:37 expressions/pitd.py:48 +#: expressions/pitd.py:51 expressions/tenc.py:36 expressions/tenc.py:38 msgid "Tick" msgstr "Tick" -#: expressions/dyn.py:34 expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/tenc.py:37 msgid "raw_rms" msgstr "raw_rms" -#: expressions/dyn.py:34 expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/tenc.py:37 msgid "Raw RMS" msgstr "Raw RMS" -#: expressions/dyn.py:34 expressions/pitd.py:46 expressions/pitd.py:47 -#: expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/pitd.py:49 expressions/pitd.py:50 +#: expressions/tenc.py:37 msgid "Time (s)" msgstr "Time (s)" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "RMS" msgstr "RMS" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46 -#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49 +#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "Reference" msgstr "Reference" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46 -#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49 +#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "UTAU" msgstr "UTAU" -#: expressions/dyn.py:35 expressions/tenc.py:36 +#: expressions/dyn.py:37 expressions/tenc.py:38 msgid "aligned_rms" msgstr "aligned_rms" -#: expressions/dyn.py:35 expressions/tenc.py:36 +#: expressions/dyn.py:37 expressions/tenc.py:38 msgid "Aligned RMS" msgstr "Aligned RMS" -#: expressions/dyn.py:45 expressions/pitd.py:61 expressions/tenc.py:47 +#: expressions/dyn.py:48 expressions/pitd.py:65 expressions/tenc.py:50 msgid "Extracting expression..." msgstr "Extracting expression..." -#: expressions/dyn.py:77 expressions/pitd.py:114 expressions/tenc.py:79 +#: expressions/dyn.py:86 expressions/pitd.py:124 expressions/tenc.py:88 msgid "Expression extraction complete." msgstr "Expression extraction complete." -#: expressions/pitd.py:27 +#: expressions/pitd.py:28 msgid "Pitch Deviation (curve)" msgstr "Pitch Deviation (curve)" -#: expressions/pitd.py:29 +#: expressions/pitd.py:30 msgid "finest accuracy, fast, CPU only (ONNX Runtime)" msgstr "finest accuracy, fast, CPU only (ONNX Runtime)" -#: expressions/pitd.py:30 +#: expressions/pitd.py:31 msgid "fair accuracy, fastest, CPU only (ONNX Runtime)" msgstr "fair accuracy, fastest, CPU only (ONNX Runtime)" -#: expressions/pitd.py:31 +#: expressions/pitd.py:32 msgid "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)" msgstr "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)" -#: expressions/pitd.py:36 +#: expressions/pitd.py:33 +msgid "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)" +msgstr "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)" + +#: expressions/pitd.py:38 #, python-format msgid "" "**F0 detection backend** for extracting pitch from WAV files. Available " @@ -408,7 +424,7 @@ msgstr "" "%s\n" "\n" -#: expressions/pitd.py:37 +#: expressions/pitd.py:39 #, python-format msgid "" "Minimum **confidence level** for keeping detected pitch values in the " @@ -425,7 +441,7 @@ msgstr "" "%s\n" "\n" -#: expressions/pitd.py:38 +#: expressions/pitd.py:40 #, python-format msgid "" "Minimum **confidence level** for keeping detected pitch values in the " @@ -442,7 +458,7 @@ msgstr "" "%s\n" "\n" -#: expressions/pitd.py:40 +#: expressions/pitd.py:42 msgid "" "**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is" " an octave higher than the reference WAV, set to 12; if lower, set to " @@ -452,48 +468,48 @@ msgstr "" " an octave higher than the reference WAV, set to 12; if lower, set to " "-12. Omit to enable automatic shift estimation" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "confidence" msgstr "confidence" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "Pitch Extraction Confidence" msgstr "Pitch Extraction Confidence" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "Confidence" msgstr "Confidence" -#: expressions/pitd.py:47 +#: expressions/pitd.py:50 msgid "raw_pitch" msgstr "raw_pitch" -#: expressions/pitd.py:47 +#: expressions/pitd.py:50 msgid "Raw Pitch" msgstr "Raw Pitch" -#: expressions/pitd.py:47 expressions/pitd.py:48 +#: expressions/pitd.py:50 expressions/pitd.py:51 msgid "Pitch (Hz)" msgstr "Pitch (Hz)" -#: expressions/pitd.py:48 +#: expressions/pitd.py:51 msgid "aligned_pitch" msgstr "aligned_pitch" -#: expressions/pitd.py:48 +#: expressions/pitd.py:51 msgid "Aligned Pitch" msgstr "Aligned Pitch" -#: expressions/pitd.py:190 +#: expressions/pitd.py:199 #, python-brace-format msgid "Estimated Semitone-shift: {}" msgstr "Estimated Semitone-shift: {}" -#: expressions/tenc.py:25 +#: expressions/tenc.py:26 msgid "Tension (curve)" msgstr "Tension (curve)" -#: expressions/tenc.py:31 +#: expressions/tenc.py:32 msgid "" "**Bias** offset added to the expression curve. Positive values shift the " "curve upward; negative values shift it downward" @@ -513,22 +529,22 @@ msgstr "Loop region" msgid "Zoom" msgstr "Zoom" -#: utils/wavtool.py:87 +#: utils/wavtool.py:88 #, python-brace-format msgid "Loading F0 data from cache file: '{}'" msgstr "Loading F0 data from cache file: '{}'" -#: utils/wavtool.py:130 +#: utils/wavtool.py:133 #, python-brace-format msgid "F0 data saved to cache file: '{}'" msgstr "F0 data saved to cache file: '{}'" -#: utils/wavtool.py:304 +#: utils/wavtool.py:380 #, python-brace-format msgid "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" msgstr "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" -#: utils/wavtool.py:310 +#: utils/wavtool.py:386 #, python-brace-format msgid "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" msgstr "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" diff --git a/locales/zh_CN/LC_MESSAGES/app.po b/locales/zh_CN/LC_MESSAGES/app.po index c3db131..d9ba0be 100644 --- a/locales/zh_CN/LC_MESSAGES/app.po +++ b/locales/zh_CN/LC_MESSAGES/app.po @@ -7,7 +7,7 @@ msgid "" msgstr "" "Project-Id-Version: expressive\n" "Report-Msgid-Bugs-To: https://github.com/NewComer00/expressive/issues\n" -"POT-Creation-Date: 2026-04-11 19:51+0800\n" +"POT-Creation-Date: 2026-04-15 21:12+0800\n" "PO-Revision-Date: 2026-03-01 20:21+0800\n" "Last-Translator: NewComer00\n" "Language: zh_CN\n" @@ -149,59 +149,63 @@ msgstr "轨道编号" msgid "Expression Selection" msgstr "表情参数" -#: expressive_gui.py:549 expressive_gui.py:639 +#: expressive_gui.py:549 expressive_gui.py:647 msgid "Trim Silence" msgstr "剪除静音" -#: expressive_gui.py:554 expressive_gui.py:609 expressive_gui.py:644 +#: expressive_gui.py:552 expressive_gui.py:581 expressive_gui.py:650 +msgid "Spline Smoothing" +msgstr "样条曲线平滑" + +#: expressive_gui.py:557 expressive_gui.py:617 expressive_gui.py:655 msgid "Align Radius" msgstr "对齐半径" -#: expressive_gui.py:559 expressive_gui.py:620 expressive_gui.py:649 +#: expressive_gui.py:562 expressive_gui.py:628 expressive_gui.py:660 msgid "Smoothness" msgstr "平滑度" -#: expressive_gui.py:564 expressive_gui.py:625 expressive_gui.py:654 +#: expressive_gui.py:567 expressive_gui.py:633 expressive_gui.py:665 msgid "Scaler" msgstr "缩放因子" -#: expressive_gui.py:593 +#: expressive_gui.py:601 msgid "UTAU Confidence" msgstr "歌姬音频置信度" -#: expressive_gui.py:599 +#: expressive_gui.py:607 msgid "Reference Confidence" msgstr "参考音频置信度" -#: expressive_gui.py:605 +#: expressive_gui.py:613 msgid "Backend" msgstr "后端" -#: expressive_gui.py:614 +#: expressive_gui.py:622 msgid "Semitone Shift" msgstr "半音偏移" -#: expressive_gui.py:615 +#: expressive_gui.py:623 msgid "Auto Estimation" msgstr "自动估算" -#: expressive_gui.py:659 +#: expressive_gui.py:670 msgid "Bias" msgstr "偏置" -#: expressive_gui.py:667 +#: expressive_gui.py:678 msgid "Import Config" msgstr "导入配置" -#: expressive_gui.py:673 +#: expressive_gui.py:684 msgid "Export Config" msgstr "导出配置" -#: expressive_gui.py:684 +#: expressive_gui.py:695 msgid "Process" msgstr "开始处理" -#: expressive_gui.py:726 +#: expressive_gui.py:737 msgid "Migrate expressions from real singers to DiffSingers (GUI)" msgstr "将表情参数从真实歌手迁移到 DiffSinger 歌手,适用于图形用户界面(GUI)" @@ -280,102 +284,112 @@ msgstr "表情参数已写入 USTX 文件:'{}'" msgid "Expression result is empty. Skipping USTX update." msgstr "表情参数结果为空,USTX 文件将不会更新。" -#: expressions/dyn.py:25 +#: expressions/dyn.py:26 msgid "Dynamics (curve)" msgstr "动态曲线 Dynamics (curve)" -#: expressions/dyn.py:27 expressions/tenc.py:27 +#: expressions/dyn.py:28 expressions/tenc.py:28 msgid "" "**Trim silence** from the leading and trailing edges of the audio before " "extracting expression" msgstr "在提取表情特征前,**剪除**音频开头与结尾的**静音部分**" -#: expressions/dyn.py:28 expressions/pitd.py:39 expressions/tenc.py:28 +#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29 msgid "" "**Radius** for the FastDTW alignment algorithm; larger values allow more " "flexible alignment but increase computation time" msgstr "FastDTW 对齐算法的**半径**;值越大对齐越灵活,但计算时间也越长" -#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29 +#: expressions/dyn.py:30 expressions/pitd.py:43 expressions/tenc.py:30 msgid "" "Controls the **smoothness** of the expression curve using Gaussian " "filtering. Higher values produce smoother curves but may lose fine detail" msgstr "通过高斯滤波控制表情曲线的**平滑度**;值越大曲线越平滑,但可能丢失细节" -#: expressions/dyn.py:30 expressions/pitd.py:42 expressions/tenc.py:30 +#: expressions/dyn.py:31 expressions/pitd.py:44 expressions/tenc.py:31 msgid "" "**Scaling factor** applied to the expression curve. Values >1 amplify the" " expression, =1 keeps original intensity, <1 reduces it" msgstr "应用于表情曲线的**缩放因子**;大于 1 则放大,等于 1 则保持原强度,小于 1 则缩小" -#: expressions/dyn.py:33 expressions/dyn.py:35 expressions/pitd.py:45 -#: expressions/pitd.py:48 expressions/tenc.py:34 expressions/tenc.py:36 +#: expressions/dyn.py:32 expressions/pitd.py:45 expressions/tenc.py:33 +msgid "" +"Perform **spline smoothing** on the final expression curve for extra " +"smoothness" +msgstr "对提取出的表情曲线进行**样条曲线平滑**,以获得更自然的效果" + +#: expressions/dyn.py:35 expressions/dyn.py:37 expressions/pitd.py:48 +#: expressions/pitd.py:51 expressions/tenc.py:36 expressions/tenc.py:38 msgid "Tick" msgstr "时间刻度(Tick)" -#: expressions/dyn.py:34 expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/tenc.py:37 msgid "raw_rms" msgstr "原音频的 RMS" -#: expressions/dyn.py:34 expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/tenc.py:37 msgid "Raw RMS" msgstr "原音频的均方根能量(RMS)特征" -#: expressions/dyn.py:34 expressions/pitd.py:46 expressions/pitd.py:47 -#: expressions/tenc.py:35 +#: expressions/dyn.py:36 expressions/pitd.py:49 expressions/pitd.py:50 +#: expressions/tenc.py:37 msgid "Time (s)" msgstr "时间(秒)" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "RMS" msgstr "RMS" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46 -#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49 +#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "Reference" msgstr "参考音频" -#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46 -#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35 -#: expressions/tenc.py:36 +#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49 +#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37 +#: expressions/tenc.py:38 msgid "UTAU" msgstr "歌姬音频" -#: expressions/dyn.py:35 expressions/tenc.py:36 +#: expressions/dyn.py:37 expressions/tenc.py:38 msgid "aligned_rms" msgstr "对齐后的 RMS" -#: expressions/dyn.py:35 expressions/tenc.py:36 +#: expressions/dyn.py:37 expressions/tenc.py:38 msgid "Aligned RMS" msgstr "对齐后的均方根能量(RMS)特征" -#: expressions/dyn.py:45 expressions/pitd.py:61 expressions/tenc.py:47 +#: expressions/dyn.py:48 expressions/pitd.py:65 expressions/tenc.py:50 msgid "Extracting expression..." msgstr "正在提取表情参数..." -#: expressions/dyn.py:77 expressions/pitd.py:114 expressions/tenc.py:79 +#: expressions/dyn.py:86 expressions/pitd.py:124 expressions/tenc.py:88 msgid "Expression extraction complete." msgstr "表情参数提取完成。" -#: expressions/pitd.py:27 +#: expressions/pitd.py:28 msgid "Pitch Deviation (curve)" msgstr "音高偏差曲线 Pitch Deviation (curve)" -#: expressions/pitd.py:29 +#: expressions/pitd.py:30 msgid "finest accuracy, fast, CPU only (ONNX Runtime)" msgstr "精度最佳,较快,仅支持 CPU,基于 ONNX Runtime" -#: expressions/pitd.py:30 +#: expressions/pitd.py:31 msgid "fair accuracy, fastest, CPU only (ONNX Runtime)" msgstr "精度一般,最快,仅支持 CPU,基于 ONNX Runtime" -#: expressions/pitd.py:31 +#: expressions/pitd.py:32 msgid "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)" msgstr "精度较好,较慢,支持 CPU 和 NVIDIA GPU,基于 TensorFlow" -#: expressions/pitd.py:36 +#: expressions/pitd.py:33 +msgid "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)" +msgstr "以 rmvpe-onnx 为基准,综合了 swift-f0 的结果,仅支持 CPU,基于 ONNX Runtime" + +#: expressions/pitd.py:38 #, python-format msgid "" "**F0 detection backend** for extracting pitch from WAV files. Available " @@ -389,7 +403,7 @@ msgstr "" "%s\n" "\n" -#: expressions/pitd.py:37 +#: expressions/pitd.py:39 #, python-format msgid "" "Minimum **confidence level** for keeping detected pitch values in the " @@ -404,7 +418,7 @@ msgstr "" "%s\n" "\n" -#: expressions/pitd.py:38 +#: expressions/pitd.py:40 #, python-format msgid "" "Minimum **confidence level** for keeping detected pitch values in the " @@ -419,55 +433,55 @@ msgstr "" "%s\n" "\n" -#: expressions/pitd.py:40 +#: expressions/pitd.py:42 msgid "" "**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is" " an octave higher than the reference WAV, set to 12; if lower, set to " "-12. Omit to enable automatic shift estimation" msgstr "歌姬音频与参考音频之间的**半音偏移**;若歌姬比参考高一个八度则设为 12,低一个八度则设为 -12;忽略该参数则启用自动估算" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "confidence" msgstr "置信度" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "Pitch Extraction Confidence" msgstr "音高提取置信度" -#: expressions/pitd.py:46 +#: expressions/pitd.py:49 msgid "Confidence" msgstr "置信度" -#: expressions/pitd.py:47 +#: expressions/pitd.py:50 msgid "raw_pitch" msgstr "原音频的 Pitch" -#: expressions/pitd.py:47 +#: expressions/pitd.py:50 msgid "Raw Pitch" msgstr "原音频的音高(Pitch)特征" -#: expressions/pitd.py:47 expressions/pitd.py:48 +#: expressions/pitd.py:50 expressions/pitd.py:51 msgid "Pitch (Hz)" msgstr "音高(Hz)" -#: expressions/pitd.py:48 +#: expressions/pitd.py:51 msgid "aligned_pitch" msgstr "对齐后的 Pitch" -#: expressions/pitd.py:48 +#: expressions/pitd.py:51 msgid "Aligned Pitch" msgstr "对齐后的音高(Pitch)特征" -#: expressions/pitd.py:190 +#: expressions/pitd.py:199 #, python-brace-format msgid "Estimated Semitone-shift: {}" msgstr "估计的半音偏移:{}" -#: expressions/tenc.py:25 +#: expressions/tenc.py:26 msgid "Tension (curve)" msgstr "张力曲线 Tension (curve)" -#: expressions/tenc.py:31 +#: expressions/tenc.py:32 msgid "" "**Bias** offset added to the expression curve. Positive values shift the " "curve upward; negative values shift it downward" @@ -485,22 +499,22 @@ msgstr "选区循环播放" msgid "Zoom" msgstr "缩放" -#: utils/wavtool.py:87 +#: utils/wavtool.py:88 #, python-brace-format msgid "Loading F0 data from cache file: '{}'" msgstr "正在从缓存文件加载 F0 数据:'{}'" -#: utils/wavtool.py:130 +#: utils/wavtool.py:133 #, python-brace-format msgid "F0 data saved to cache file: '{}'" msgstr "F0 数据已保存到缓存文件:'{}'" -#: utils/wavtool.py:304 +#: utils/wavtool.py:380 #, python-brace-format msgid "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" msgstr "开始时间 {:.3f}秒 被截取到 {:.3f}秒(完整时长:{:.3f}秒) " -#: utils/wavtool.py:310 +#: utils/wavtool.py:386 #, python-brace-format msgid "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)" msgstr "结束时间 {:.3f}秒 被截取到 {:.3f}秒(完整时长:{:.3f}秒) " diff --git a/tests/test_seqtool.py b/tests/test_seqtool.py index 5df030a..5851c14 100644 --- a/tests/test_seqtool.py +++ b/tests/test_seqtool.py @@ -9,6 +9,7 @@ sequence_interval_union, unify_sequence_time, gaussian_filter1d_with_nan, + seq_spline_smoothing, align_sequence_tick, seq_dynamics_trends, seq_rcr, @@ -278,6 +279,98 @@ def test_smoothing_levels(self, sigma, should_smooth): assert_array_almost_equal(result, seq) +# --------------------------------------------------------------------------- +# seq_spline_smoothing +# --------------------------------------------------------------------------- + +class TestSeqSplineSmoothing: + """Test seq_spline_smoothing.""" + + def test_output_shape(self): + t = np.linspace(0, 1, 50) + v = np.sin(2 * np.pi * t) + result = seq_spline_smoothing(t, v) + assert result.shape == v.shape + + def test_smooths_noisy_signal(self): + np.random.seed(0) + t = np.linspace(0, 1, 100) + v = np.sin(2 * np.pi * t) + np.random.normal(0, 0.3, 100) + result = seq_spline_smoothing(t, v) + assert np.var(result) < np.var(v) + + def test_no_nan_policy(self): + t = np.linspace(0, 1, 20) + v = np.sin(2 * np.pi * t) + v[5] = np.nan + v[10] = np.nan + result = seq_spline_smoothing(t, v, nan_policy='no_nan') + assert not np.any(np.isnan(result)) + + def test_preserve_all_nan_policy(self): + t = np.linspace(0, 1, 20) + v = np.sin(2 * np.pi * t) + nan_indices = [3, 7, 15] + v[nan_indices] = np.nan + result = seq_spline_smoothing(t, v, nan_policy='preserve_all') + assert result.shape == v.shape + for i in nan_indices: + assert np.isnan(result[i]) + non_nan = [i for i in range(len(v)) if i not in nan_indices] + assert not np.any(np.isnan(result[non_nan])) + + def test_preserve_head_tail_nan_policy(self): + t = np.linspace(0, 1, 20) + v = np.sin(2 * np.pi * t) + # Leading, interior, and trailing NaNs + v[0] = np.nan + v[1] = np.nan + v[10] = np.nan + v[18] = np.nan + v[19] = np.nan + result = seq_spline_smoothing(t, v, nan_policy='preserve_head_tail') + assert result.shape == v.shape + # Leading and trailing NaNs preserved + assert np.isnan(result[0]) + assert np.isnan(result[1]) + assert np.isnan(result[18]) + assert np.isnan(result[19]) + # Interior NaN filled + assert not np.isnan(result[10]) + + def test_invalid_nan_policy_raises(self): + t = np.linspace(0, 1, 10) + v = np.ones(10) + with pytest.raises(ValueError, match="Unknown nan_policy"): + seq_spline_smoothing(t, v, nan_policy='invalid_policy') + + def test_manual_lam_smoother_than_auto(self): + """A large lam should produce a smoother (lower variance) result than auto.""" + np.random.seed(1) + t = np.linspace(0, 1, 100) + v = np.sin(2 * np.pi * t) + np.random.normal(0, 0.2, 100) + result_auto = seq_spline_smoothing(t, v, lam=None) + result_heavy = seq_spline_smoothing(t, v, lam=1e4) + assert np.var(result_heavy) <= np.var(result_auto) + + def test_no_nan_input_all_policies_agree(self): + """With no NaNs, all nan_policy values should return identical results.""" + t = np.linspace(0, 1, 30) + v = np.cos(2 * np.pi * t) + r_no_nan = seq_spline_smoothing(t, v, nan_policy='no_nan') + r_preserve = seq_spline_smoothing(t, v, nan_policy='preserve_all') + r_head_tail = seq_spline_smoothing(t, v, nan_policy='preserve_head_tail') + assert_array_almost_equal(r_no_nan, r_preserve) + assert_array_almost_equal(r_no_nan, r_head_tail) + + @pytest.mark.parametrize("nan_policy", ['no_nan', 'preserve_all', 'preserve_head_tail']) + def test_all_policies_accepted(self, nan_policy): + t = np.linspace(0, 1, 20) + v = np.sin(2 * np.pi * t) + result = seq_spline_smoothing(t, v, nan_policy=nan_policy) + assert result.shape == v.shape + + # --------------------------------------------------------------------------- # seq_dynamics_trends # --------------------------------------------------------------------------- diff --git a/tests/test_wavtool.py b/tests/test_wavtool.py index 5815435..e87c7db 100644 --- a/tests/test_wavtool.py +++ b/tests/test_wavtool.py @@ -4,6 +4,7 @@ import csv import os import tempfile +import types import unittest from unittest.mock import MagicMock, patch @@ -35,6 +36,16 @@ def _make_wav(duration: float = 5.0, sr: int = 22050) -> str: return tmp.name +def _make_tonal_wav(duration: float = 2.0, sr: int = 22050, freq: float = 440.0) -> str: + """Write a sine-wave WAV so RMS is non-trivially non-zero throughout.""" + t = np.linspace(0, duration, int(duration * sr), endpoint=False) + y = (0.5 * np.sin(2 * np.pi * freq * t)).astype(np.float32) + tmp = tempfile.NamedTemporaryFile(suffix=".wav", delete=False) + sf.write(tmp.name, y, sr) + tmp.close() + return tmp.name + + # --------------------------------------------------------------------------- # timestamp2sec # --------------------------------------------------------------------------- @@ -586,15 +597,6 @@ def tearDown(self): except FileNotFoundError: pass - def _make_tonal_wav(self, duration=2.0, sr=22050, freq=440.0) -> str: - """Write a sine-wave WAV so RMS is non-trivially non-zero throughout.""" - t = np.linspace(0, duration, int(duration * sr), endpoint=False) - y = (0.5 * np.sin(2 * np.pi * freq * t)).astype(np.float32) - tmp = tempfile.NamedTemporaryFile(suffix=".wav", delete=False) - sf.write(tmp.name, y, sr) - tmp.close() - return tmp.name - # --- return types and shapes --- def test_returns_tuple_of_two(self): @@ -630,6 +632,11 @@ def test_rms_time_is_monotonically_increasing(self): rms_time, _ = extract_wav_rms(self.wav) self.assertTrue(np.all(np.diff(rms_time) > 0)) + def test_rms_time_does_not_exceed_duration(self): + """All time points should fall within (or very close to) the file duration.""" + rms_time, _ = extract_wav_rms(self.wav) + self.assertLessEqual(rms_time[-1], 2.1) + # --- silence masking --- def test_silent_wav_masked_with_nan(self): @@ -653,7 +660,7 @@ def test_mask_silence_false_values_nonnegative(self): def test_tonal_wav_has_active_frames(self): """A sine wave should have some non-NaN (active) RMS frames.""" - wav = self._make_tonal_wav() + wav = _make_tonal_wav() try: _, rms = extract_wav_rms(wav, mask_silence=True) self.assertTrue(np.any(~np.isnan(rms))) @@ -661,7 +668,7 @@ def test_tonal_wav_has_active_frames(self): os.unlink(wav) def test_tonal_wav_active_frames_are_nonnegative(self): - wav = self._make_tonal_wav() + wav = _make_tonal_wav() try: _, rms = extract_wav_rms(wav, mask_silence=True) active = rms[~np.isnan(rms)] @@ -669,6 +676,15 @@ def test_tonal_wav_active_frames_are_nonnegative(self): finally: os.unlink(wav) + def test_tonal_wav_no_nan_when_mask_false(self): + """A fully tonal signal with mask_silence=False must have zero NaN values.""" + wav = _make_tonal_wav() + try: + _, rms = extract_wav_rms(wav, mask_silence=False) + self.assertFalse(np.any(np.isnan(rms))) + finally: + os.unlink(wav) + def test_leading_silence_masked(self): """Frames before the first active frame should be NaN when mask_silence=True.""" sr = 22050 @@ -745,15 +761,13 @@ def _patch_swift(self): mock_cls = MagicMock(return_value=mock_detector) return patch("swift_f0.SwiftF0", mock_cls, create=True) - def _patch_rmvpe(self): + def _patch_rmvpe(self, times=None, freqs=None, confs=None): """Patch RMVPE and soundfile so the rmvpe-onnx backend runs without - real model weights or audio I/O.""" - import types - - fake_timestamp = np.array(self._fake_times) - fake_frequency = np.array(self._fake_freqs) - fake_confidence = np.array(self._fake_confs) - fake_activation = np.zeros(len(self._fake_times)) + real model weights or audio I/O. Optionally override return values.""" + fake_timestamp = np.array(times if times is not None else self._fake_times) + fake_frequency = np.array(freqs if freqs is not None else self._fake_freqs) + fake_confidence = np.array(confs if confs is not None else self._fake_confs) + fake_activation = np.zeros(len(fake_timestamp)) fake_rmvpe_instance = MagicMock() fake_rmvpe_instance.predict.return_value = ( @@ -863,7 +877,6 @@ def test_crepe_backend_accepted(self): fake_conf = np.random.uniform(0.5, 1.0, self._n) fake_act = np.zeros((self._n, 360)) - import types fake_crepe_mod = types.ModuleType("crepe") fake_crepe_mod.predict = MagicMock( return_value=(fake_time, fake_freq, fake_conf, fake_act) @@ -878,6 +891,131 @@ def test_crepe_backend_accepted(self): self.assertEqual(len(time), len(freq)) self.assertEqual(len(freq), len(conf)) + # --- hybrid backend --- + + def test_hybrid_backend_accepted(self): + """hybrid backend must be accepted without raising ValueError.""" + with self._patch_rmvpe(), self._patch_swift(): + # soundfile is also used inside _merge_rmvpe_and_swift_f0 + extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + + def test_hybrid_returns_tuple_of_three(self): + with self._patch_rmvpe(), self._patch_swift(): + result = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertIsInstance(result, tuple) + self.assertEqual(len(result), 3) + + def test_hybrid_output_arrays_are_ndarrays(self): + with self._patch_rmvpe(), self._patch_swift(): + time, freq, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertIsInstance(time, np.ndarray) + self.assertIsInstance(freq, np.ndarray) + self.assertIsInstance(conf, np.ndarray) + + def test_hybrid_all_outputs_same_length(self): + with self._patch_rmvpe(), self._patch_swift(): + time, freq, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertEqual(len(time), len(freq)) + self.assertEqual(len(freq), len(conf)) + + def test_hybrid_output_length_matches_rmvpe_grid(self): + """Hybrid uses the rmvpe-onnx time grid, so output length == rmvpe output length.""" + with self._patch_rmvpe(), self._patch_swift(): + time, freq, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertEqual(len(time), self._n) + + def test_hybrid_time_values_are_float(self): + with self._patch_rmvpe(), self._patch_swift(): + time, _, _ = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertTrue(np.issubdtype(time.dtype, np.floating)) + + def test_hybrid_frequency_values_are_float(self): + with self._patch_rmvpe(), self._patch_swift(): + _, freq, _ = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertTrue(np.issubdtype(freq.dtype, np.floating)) + + def test_hybrid_confidence_values_are_float(self): + with self._patch_rmvpe(), self._patch_swift(): + _, _, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False) + self.assertTrue(np.issubdtype(conf.dtype, np.floating)) + + def test_hybrid_time_matches_rmvpe_time(self): + """The time axis returned by hybrid must equal the rmvpe-onnx time axis.""" + with self._patch_rmvpe(), self._patch_swift(): + time_hybrid, _, _ = extract_wav_frequency( + self.wav, backend="hybrid", use_cache=False + ) + with self._patch_rmvpe(): + time_rmvpe, _, _ = extract_wav_frequency( + self.wav, backend="rmvpe-onnx", use_cache=False + ) + np.testing.assert_array_almost_equal(time_hybrid, time_rmvpe) + + def test_hybrid_replaces_low_confidence_rmvpe_frames(self): + """When rmvpe confidence is low and swift-f0 confidence is high in a voiced + region, the hybrid output should include some swift-f0 frequency values.""" + n = 50 + times = list(np.linspace(0, 1.0, n)) + + # rmvpe: low confidence everywhere so swift-f0 replacements will be chosen + rmvpe_freqs = [200.0] * n + rmvpe_confs = [0.5] * n # below threshold 0.80 + + # swift-f0: high confidence + different frequency + swift_freqs = [400.0] * n + swift_confs = [0.99] * n # above threshold 0.95 + + # Build a tonal WAV so RMS > 0 (voiced region) + wav = _make_tonal_wav(duration=1.0, sr=22050, freq=440.0) + try: + with self._patch_rmvpe(times=times, freqs=rmvpe_freqs, confs=rmvpe_confs), \ + self._patch_swift(): + # Override swift-f0 mock with specific values + fake_result = MagicMock() + fake_result.timestamps.tolist.return_value = times + fake_result.pitch_hz.tolist.return_value = swift_freqs + fake_result.confidence.tolist.return_value = swift_confs + mock_detector = MagicMock() + mock_detector.detect_from_file.return_value = fake_result + mock_cls = MagicMock(return_value=mock_detector) + with patch("swift_f0.SwiftF0", mock_cls, create=True): + _, freq, _ = extract_wav_frequency(wav, backend="hybrid", use_cache=False) + + # At least some frames should have been replaced with 400 Hz + self.assertTrue(np.any(np.isclose(freq, 400.0)), + "Expected some frames to be replaced by swift-f0 (400 Hz)") + finally: + os.unlink(wav) + + def test_hybrid_keeps_rmvpe_frames_when_rmvpe_confident(self): + """When rmvpe confidence is high the hybrid output should retain rmvpe frequencies.""" + n = 50 + times = list(np.linspace(0, 1.0, n)) + rmvpe_freqs = [200.0] * n + rmvpe_confs = [0.95] * n # above threshold 0.80 — should not be replaced + + swift_freqs = [400.0] * n + swift_confs = [0.99] * n + + wav = _make_tonal_wav(duration=1.0, sr=22050, freq=440.0) + try: + with self._patch_rmvpe(times=times, freqs=rmvpe_freqs, confs=rmvpe_confs): + fake_result = MagicMock() + fake_result.timestamps.tolist.return_value = times + fake_result.pitch_hz.tolist.return_value = swift_freqs + fake_result.confidence.tolist.return_value = swift_confs + mock_detector = MagicMock() + mock_detector.detect_from_file.return_value = fake_result + mock_cls = MagicMock(return_value=mock_detector) + with patch("swift_f0.SwiftF0", mock_cls, create=True): + _, freq, _ = extract_wav_frequency(wav, backend="hybrid", use_cache=False) + + # All frames should stay at 200 Hz (rmvpe confident, no replacement) + self.assertTrue(np.all(np.isclose(freq, 200.0)), + "Expected rmvpe frequencies to be kept when confidence is high") + finally: + os.unlink(wav) + # --- caching --- def test_cache_file_written_when_use_cache_true(self): diff --git a/utils/seqtool.py b/utils/seqtool.py index 8352c56..3c975a5 100644 --- a/utils/seqtool.py +++ b/utils/seqtool.py @@ -3,7 +3,7 @@ import numpy as np from fastdtw import fastdtw # type: ignore -from scipy.interpolate import interp1d +from scipy.interpolate import interp1d, make_smoothing_spline from scipy.ndimage import gaussian_filter1d from scipy.stats import zscore @@ -127,7 +127,7 @@ def unify_sequence_time(seq_times, seq_vals, to_ticks=False): ] return unified_seq_time, tuple(unified_seqs_val) - unified_seq_ticks = _time_to_ticks_fn(unified_seq_time) + unified_seq_ticks = np.unique(_time_to_ticks_fn(unified_seq_time)) time_mapping = _ticks_to_time_fn(unified_seq_ticks) unified_seqs_val = [ interp1d(st, sv, fill_value="extrapolate")(time_mapping) # type: ignore @@ -161,6 +161,60 @@ def gaussian_filter1d_with_nan(seq, sigma, **kwargs): return seq +def seq_spline_smoothing(seq_time, seq_val, lam=None, nan_policy='preserve_all'): + """Smooth a sequence using an adaptive smoothing spline. + + Args: + seq_time (numpy.ndarray): Time values for the sequence. + seq_val (numpy.ndarray): Sequence values to smooth. + lam (float or None): Smoothing parameter. None selects automatically via GCV. + Higher values produce smoother results. + nan_policy (str): How to handle NaNs in the output. One of: + - 'no_nan': NaNs are excluded from fitting; output is fully predicted + with no NaNs. + - 'preserve_all': NaN positions are excluded from fitting and restored in + the output. + - 'preserve_head_tail': Only leading and trailing NaNs are restored; interior + NaNs are filled by the spline. + + Returns: + numpy.ndarray: Smoothed sequence values, same length as seq_time. + + Raises: + ValueError: If nan_policy is not recognized. + + Example: + >>> seq_smoothing(time, val) # auto smoothness, preserve all NaNs + >>> seq_smoothing(time, val, lam=0.1) # manual smoothness + >>> seq_smoothing(time, val, nan_policy='no_nan') # fully predicted, no NaNs + >>> seq_smoothing(time, val, nan_policy='preserve_head_tail') # only boundary NaNs restored + """ + nan_policies = {'no_nan', 'preserve_all', 'preserve_head_tail'} + if nan_policy not in nan_policies: + raise ValueError(f"Unknown nan_policy: {nan_policy!r}. Choose from: {nan_policies}.") + + seq_val = np.asarray(seq_val, dtype=float) + seq_time = np.asarray(seq_time, dtype=float) + nan_mask = np.isnan(seq_val) + + valid_time = seq_time[~nan_mask] + valid_val = seq_val[~nan_mask] + + result = make_smoothing_spline(valid_time, valid_val, lam=lam)(seq_time) + + if nan_policy == 'preserve_all': + result[nan_mask] = np.nan + + elif nan_policy == 'preserve_head_tail': + first_valid = np.argmax(~nan_mask) + last_valid = len(nan_mask) - np.argmax(~nan_mask[::-1]) - 1 + head_tail_mask = nan_mask.copy() + head_tail_mask[first_valid:last_valid + 1] = False + result[head_tail_mask] = np.nan + + return result + + def align_sequence_tick( query_time, queries, reference_time, references, align_radius=1 ): diff --git a/utils/wavtool.py b/utils/wavtool.py index a193b18..b778cf7 100644 --- a/utils/wavtool.py +++ b/utils/wavtool.py @@ -56,10 +56,11 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True): Args: file_path (str): Path to the WAV file. - backend (str, optional): Pitch detection backend. One of "crepe" or "swift-f0" or "rmvpe-onnx". + backend (str, optional): Pitch detection backend. "crepe" uses the CREPE model (requires TensorFlow, GPU-accelerated). "swift-f0" uses SwiftF0 (faster CPU inference, requires swift-f0 package). "rmvpe-onnx" uses RMVPE ONNX model (fast CPU inference, requires rmvpe-onnx package). + "hybrid" uses a hybrid strategy based on "rmvpe-onnx" and "swift-f0". Defaults to "rmvpe-onnx". use_cache (bool, optional): Whether to use cached data if available. Defaults to True. @@ -69,7 +70,7 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True): - frequency (np.ndarray of float): Detected pitch frequencies in Hz. Shape: (n_time_points). - confidence (np.ndarray of float): Confidence values for the detected pitches. Shape: (n_time_points). """ - _SUPPORTED_BACKENDS = ("crepe", "swift-f0", "rmvpe-onnx") + _SUPPORTED_BACKENDS = ("crepe", "swift-f0", "rmvpe-onnx", "hybrid") if backend not in _SUPPORTED_BACKENDS: raise ValueError(f"Unknown backend '{backend}'. Choose from: {_SUPPORTED_BACKENDS}") @@ -84,7 +85,7 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True): cache_path = cache_dir / f"{wav_hash}.{backend}.csv" if cache_path.is_file(): - print(_("Loading F0 data from cache file: '{}'").format(cache_path)) + print(f"[{backend}] " + _("Loading F0 data from cache file: '{}'").format(cache_path)) with open(cache_path, "r", newline="") as file: reader = csv.reader(file) next(reader) # Skip header @@ -119,6 +120,8 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True): time = timestamp.tolist() frequency = frequency.tolist() confidence = confidence.tolist() + elif backend == "hybrid": + time, frequency, confidence = _merge_rmvpe_and_swift_f0(file_path, use_cache) # Save data to cache if use_cache: @@ -127,11 +130,84 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True): writer.writerow(["Time (s)", "Frequency (Hz)", "Confidence"]) for t, f, c in zip(time, frequency, confidence, strict=False): writer.writerow([t, f, c]) - print(_("F0 data saved to cache file: '{}'").format(cache_path)) + print(f"[{backend}] " + _("F0 data saved to cache file: '{}'").format(cache_path)) return np.asarray(time), np.asarray(frequency), np.asarray(confidence) +def _merge_rmvpe_and_swift_f0(file_path, use_cache): + """Merge rmvpe-onnx and swift-f0 pitch predictions into a single output. + + Uses rmvpe-onnx as the base prediction and selectively replaces frames with + swift-f0 results where all three conditions are met: + 1. The frame falls within a voiced region (RMS energy >= Otsu threshold). + 2. rmvpe-onnx confidence is low, indicating uncertain prediction. + 3. swift-f0 confidence is high, indicating a reliable prediction. + + Voiced regions are detected by computing per-frame RMS energy with librosa, + then thresholding with Otsu's method to separate voiced from unvoiced frames. + swift-f0 frames are aligned to the rmvpe-onnx time grid via nearest-neighbour + lookup before comparison. + + Args: + file_path (str): Path to the WAV file. Passed directly to + extract_wav_frequency for both backends. + use_cache (bool): Whether to use cached predictions. Passed directly to + extract_wav_frequency for both backends. + + Returns: + tuple: (time, frequency, confidence), where: + - time (list of float): Time points in seconds from the rmvpe-onnx grid. + - frequency (list of float): Merged pitch frequencies in Hz. + - confidence (list of float): Confidence values corresponding to + whichever backend's frequency was selected per frame. + """ + _CONFIDENCE_THRESHOLDS = { + "rmvpe-onnx": 0.80, + "swift-f0": 0.95, + } + + from skimage.filters import threshold_otsu + from librosa.feature import rms as librosa_rms + + r_time, r_freq, r_conf = extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=use_cache) + s_time, s_freq, s_conf = extract_wav_frequency(file_path, backend="swift-f0", use_cache=use_cache) + + # --- Base prediction: start from rmvpe-onnx --- + out_freq = r_freq.copy() + out_conf = r_conf.copy() + + # --- Voiced region via Otsu threshold on RMS --- + import soundfile as sf + audio, sr = sf.read(file_path) + hop_length = 512 + frame_rms = librosa_rms(y=audio, hop_length=hop_length)[0] + rms_times = np.arange(len(frame_rms)) * hop_length / sr + otsu_thr = threshold_otsu(frame_rms) + # Snap RMS voiced mask to rmvpe-onnx time grid + rms_indices = np.searchsorted(rms_times, r_time).clip(0, len(frame_rms) - 1) + voiced_region = frame_rms[rms_indices] >= otsu_thr + + # --- Align swift-f0 to rmvpe-onnx time grid via nearest-neighbour lookup --- + snap_indices = np.searchsorted(s_time, r_time).clip(0, len(s_time) - 1) + prev_indices = np.maximum(snap_indices - 1, 0) + use_prev = np.abs(s_time[prev_indices] - r_time) < np.abs(s_time[snap_indices] - r_time) + snap_indices = np.where(use_prev, prev_indices, snap_indices) + + s_freq_aligned = s_freq[snap_indices] + s_conf_aligned = s_conf[snap_indices] + + # --- Replace: voiced region + rmvpe low confidence + swift-f0 high confidence --- + r_low = r_conf < _CONFIDENCE_THRESHOLDS["rmvpe-onnx"] + s_high = s_conf_aligned >= _CONFIDENCE_THRESHOLDS["swift-f0"] + replace_mask = voiced_region & r_low & s_high + + out_freq = np.where(replace_mask, s_freq_aligned, out_freq) + out_conf = np.where(replace_mask, s_conf_aligned, out_conf) + + return r_time.tolist(), out_freq.tolist(), out_conf.tolist() + + def extract_wav_rms(wav_path, mask_silence=True): """Extract RMS energy from a WAV file.