From 508dfe0d0bfdcfd130e48fa2ea06637b81e587ee Mon Sep 17 00:00:00 2001
From: NewCommer00
Date: Wed, 15 Apr 2026 23:49:46 +0800
Subject: [PATCH] feat(pitd, f0, docs): improve PITD and add hybrid F0 backend
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
BREAKING CHANGE: PITD scaler default changed from 2.0 to 1.0; PITD results prior to v0.9.0 are unreliable
feat(f0): add hybrid F0 backend with fallback; improve stability and reduce discontinuities
fix(pitd): fix overly flat PITD curves (issue #21); special thanks to @ma0shu for helping identify and diagnose this critical bug ❤
docs(readme): add v0.9.0+ warning; document scaler change; update troubleshooting
---
README.en.md | 71 +++++--
README.md | 42 ++++-
.../expressive_config.json" | 8 +-
.../expressive_config.json" | 4 +-
.../expressive_config.json" | 6 +-
expressions/dyn.py | 25 ++-
expressions/pitd.py | 75 +++++---
expressions/tenc.py | 29 ++-
expressive_gui.py | 23 ++-
locales/app.pot | 136 +++++++------
locales/en/LC_MESSAGES/app.po | 138 ++++++++------
locales/zh_CN/LC_MESSAGES/app.po | 136 +++++++------
tests/test_seqtool.py | 93 +++++++++
tests/test_wavtool.py | 178 ++++++++++++++++--
utils/seqtool.py | 58 +++++-
utils/wavtool.py | 84 ++++++++-
16 files changed, 811 insertions(+), 295 deletions(-)
diff --git a/README.en.md b/README.en.md
index 79a1693..1e8d5f2 100644
--- a/README.en.md
+++ b/README.en.md
@@ -7,6 +7,15 @@
+> [!WARNING]
+> 🚨 **Please Read Before Downloading** 🚨
+>
+> **It is strongly recommended to use v0.9.0 or later**. Earlier versions of the **PITD** expression parameter processing algorithm contain a [critical flaw](https://github.com/NewComer00/expressive/releases/tag/v0.9.0) that may result in **incorrect pitch curve generation**. To download the latest version, please visit the [Releases page](https://github.com/NewComer00/expressive/releases).
+>
+> For users migrating from an older version to `v0.9.0` or later, note that the default value of the **PITD Scaler** is now `1.0` (previously `2.0`). If you have an old configuration file, please set the **PITD Scaler** to `1.0`.
+>
+> **🎵 Thank you for using Expressive🎵**
+
# Expressive
**Expressive** is a [DiffSinger](https://github.com/openvpi/diffsinger) expression parameter importer developed for [OpenUtau](https://github.com/stakira/OpenUtau). It aims to extract expression parameters from real human vocals and import them into the appropriate tracks of your project.
@@ -21,18 +30,17 @@ The current version supports importing the following expression parameters:
| **Working with OpenUtau** | **Data Viewer** |
|:---:|:---:|
-|
|
|
+|
|
|
-> - *OpenUtau version from [keirokeer/OpenUtau-DiffSinger-Lunai](https://github.com/keirokeer/OpenUtau-DiffSinger-Lunai)*
-> - *Singer model from [yousa-ling-official-production/yousa-ling-diffsinger-v1](https://github.com/yousa-ling-official-production/yousa-ling-diffsinger-v1)*
+> - *Example from [`examples/明天会更好`](examples/明天会更好). Click to view details.*
> [!TIP]
>
> 👉 Click to expand the full voiced demo video 👈
>
->
+>
>
>
>
@@ -47,6 +55,8 @@ By default, this application uses [rmvpe-onnx](https://github.com/newcomer00/rmv
The [swift-f0](https://github.com/lars76/swift-f0) and [CREPE](https://github.com/marl/crepe) pitch extraction backends are also available. The former runs on CPU only and is the fastest option, though its accuracy is modest. The latter is a classic algorithm in the field and runs more slowly. In a CUDA environment, the CREPE backend will automatically enable GPU acceleration.
+There is also a newly added experimental **hybrid** backend available. The hybrid backend combines the prediction results of rmvpe-onnx and swift-f0, primarily using the pitch extraction results from rmvpe-onnx. In voiced segments of the audio, if the confidence of rmvpe-onnx is low and the confidence of swift-f0 is high, the result from swift-f0 is used for correction, improving the overall accuracy of pitch extraction.
+
> \* On Windows, TensorFlow 2.10 is the last version that supports GPU acceleration, and Python 3.10 is the highest Python version supported by its `.whl` files.
## 📌 Use Case
@@ -297,24 +307,53 @@ Relaunching the application should restore normal functionality, and this issue
#### Future Plan
The NiceGUI framework has begun improving its drag-and-drop support and should resolve this in a future release.
-### PITD expression curve is overall too flat
+---
+
+### PITD expression curve is overly flat
#### Symptom
-The extracted PITD expression curve is too flat, with almost no significant variation overall. Pitch changes in the reference vocal are not reflected in the expression curve.
-#### Possible Cause
-The two confidence thresholds in the PITD extractor are set **too high**, causing many pitch changes to be discarded.
+The extracted PITD expression curve is too flat, with almost no significant variation. Pitch changes in the reference vocal are not properly reflected in the curve.
-#### Solution
-First try using the best-performing rmvpe-onnx backend (with default confidence thresholds). If the issue persists, try lowering both confidence thresholds. In general, the **Utau vocal** is relatively clean, so it is advisable to first adjust the confidence threshold for the **Reference vocal**.
+#### Possible Causes
+
+1. In versions earlier than **v0.9.0**, there is an issue in the conversion between pitch and PITD values, which can cause the curve to appear overly flat.
+2. The two confidence thresholds in the PITD extractor are set **too high**, causing many pitch changes to be discarded.
+ You can observe missing segments in the original pitch curve in [`expressive-viewer`](#iewer).
+
+#### Solutions
-### PITD expression curve has sudden jumps or spikes at certain positions
+1. Please upgrade to **v0.9.0 or later**.
+2. First try using the best-performing **rmvpe-onnx** or **hybrid** backend (with default confidence thresholds).
+ If the issue persists, try lowering both confidence thresholds. You can use the pitch confidence curve in [`expressive-viewer`](#viewer) as a reference when tuning.
+ In general, the **Utau vocal** is relatively clean, so it is recommended to adjust the confidence threshold for the **reference vocal** first.
+
+#### Future Plans
+Incorporate semantic information into the PITD expression extraction algorithm.
+
+---
+
+### PITD expression curve has sudden jumps or spikes
#### Symptom
-The PITD expression curve changes too rapidly at certain positions, with very large jumps or spikes that clearly do not match natural vocal behavior.
-#### Possible Cause
-The two confidence thresholds in the PITD extractor are set **too low**, causing erroneous detection results to be accepted.
+The PITD expression curve changes too abruptly at certain positions, with large jumps or spikes that do not match natural vocal behavior.
-#### Solution
-First try using the best-performing rmvpe-onnx backend (with default confidence thresholds). If the issue persists, try increasing both confidence thresholds. In general, the **Utau vocal** is relatively clean, so it is advisable to first adjust the confidence threshold for the **Reference vocal**.
+#### Possible Causes
+
+1. In versions earlier than **v0.9.0**, there is an issue in the conversion between pitch and PITD values.
+2. There is noise in the reference audio around the corresponding timestamps.
+ You can observe abnormal spikes in the original pitch curve in [`expressive-viewer`](#viewer)
+3. The two confidence thresholds in the PITD extractor are set **too low**, causing incorrect detections to be accepted.
+ This may also appear as spikes in the pitch curve in [`expressive-viewer`](#viewer).
+
+#### Solutions
+
+1. Please upgrade to **v0.9.0 or later**.
+2. Try denoising the reference audio using tools such as [UVR](https://github.com/Anjok07/ultimatevocalremovergui) or [MSST](https://github.com/SUC-DriverOld/MSST-WebUI).
+3. First try using the best-performing **rmvpe-onnx** or **hybrid** backend (with default confidence thresholds).
+ If the issue persists, try increasing both confidence thresholds. You can use the pitch confidence curve in [`expressive-viewer`](#viewer) to guide your adjustments.
+ In general, the **Utau vocal** is relatively clean, so it is recommended to adjust the confidence threshold for the **reference vocal** first.
+
+#### Future Plans
+Incorporate semantic information into the PITD expression extraction algorithm.
diff --git a/README.md b/README.md
index 0b08bda..9103852 100644
--- a/README.md
+++ b/README.md
@@ -7,6 +7,15 @@
+> [!WARNING]
+> 🚨 **下载前请注意** 🚨
+>
+> **强烈建议您使用 `v0.9.0` 及以上版本**。早先版本的 **PITD** 表情参数处理算法存在[严重缺陷](https://github.com/NewComer00/expressive/releases/tag/v0.9.0),会导致**音高曲线绘制错误**。下载最新版本请前往 [Releases 页面](https://github.com/NewComer00/expressive/releases)。
+>
+> 对于从旧版本迁移到 `v0.9.0` 及以上版本的用户,新版本中 **PITD 缩放因子(Scaler)的默认值为 `1.0`**,不再是原来的 `2.0`。若您有旧版本的配置文件,请将 **PITD 缩放因子(Scaler)设置为 `1.0`**。
+>
+> **🎵 感谢您使用 Expressive🎵**
+
# Expressive
**Expressive** 是一个为 [OpenUtau](https://github.com/stakira/OpenUtau) 开发的 [DiffSinger](https://github.com/openvpi/diffsinger) 表情参数导入工具,旨在从真实人声中提取表情参数,并导入至工程的相应轨道。
@@ -21,18 +30,17 @@
| **工作流程** | **数据可视化** |
|:---:|:---:|
-|
|
|
+|
|
|
-> - *OpenUtau 版本来自 [keirokeer/OpenUtau-DiffSinger-Lunai](https://github.com/keirokeer/OpenUtau-DiffSinger-Lunai)*
-> - *歌手模型来自 [yousa-ling-official-production/yousa-ling-diffsinger-v1](https://github.com/yousa-ling-official-production/yousa-ling-diffsinger-v1)*
+> - *示例来自 [`examples/明天会更好`](examples/明天会更好),点击查看详情信息*
> [!TIP]
>
> 👉 点击展开完整有声演示视频 👈
>
->
+>
>
>
>
@@ -47,6 +55,8 @@
应用也提供了 [swift-f0](https://github.com/lars76/swift-f0) 与 [CREPE](https://github.com/marl/crepe) 音高提取后端。前者仅依赖 CPU,效果一般,但速度最快。后者是业内的经典算法,速度较慢。在 CUDA 环境下,CREPE 后端会自动启用 GPU 加速。
+应用还新增了一个实验性的 **hybrid** 后端。该后端融合了 rmvpe-onnx 与 swift-f0 的预测结果,以 rmvpe-onnx 的音高提取结果为主,在音频有声段中,如果 rmvpe-onnx 的置信度较低且 swift-f0 的置信度较高,则采用 swift-f0 的结果进行修正,从而提升整体音高提取的准确性。
+
> \* 在 Windows 平台下,TensorFlow 2.10 是最后一个支持 GPU 加速的版本,Python 3.10 是它的 `.whl` 文件支持的最高 Python 版本。
## 📌 使用场景
@@ -302,16 +312,25 @@ graph TB;
#### 未来计划
NiceGUI 框架已经开始着手改进文件拖拽支持,应该在未来的版本中能够解决此问题。
+---
+
### PITD 表情曲线整体变化过于平缓
#### 问题现象
提取出的 PITD 表情曲线过于平缓,整体上几乎没有大的起伏,参考人声中的音高变化并没有反映到表情曲线上。
#### 可能原因
-PITD 表情提取器中,两个置信度阈值设置**过高**,许多音高变化没有被采信。
+1. 在早于 `v0.9.0` 的版本中,PITD 表情曲线取值与音高之间的换算有问题,会导致 PITD 表情曲线整体非常平缓。
+2. PITD 表情提取器中,两个置信度阈值设置**过高**,许多音高变化没有被采信。您可以在 [`expressive-viewer`](#可视化工具viewer) 中观察到,原始的音高曲线中有很多不该出现的缺失部分。
#### 解决方案
-请先尝试使用效果最好的 rmvpe-onnx 后端(默认置信度阈值)。若问题仍在,尝试降低两个置信度阈值。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。
+1. 请下载安装 `v0.9.0` 及之后的版本。
+2. 请先尝试使用效果最好的 rmvpe-onnx 或 hybrid 后端(默认置信度阈值)。若问题仍在,尝试降低两个置信度阈值。您可以参考 [`expressive-viewer`](#可视化工具viewer) 的音高置信度曲线来辅助调整。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。
+
+#### 未来计划
+为 PITD 表情提取算法引入语义信息。
+
+---
### PITD 表情曲线在某些位置变化过快,出现跳跃或毛刺
@@ -319,7 +338,14 @@ PITD 表情提取器中,两个置信度阈值设置**过高**,许多音高
PITD 表情曲线在某些位置变化过快,出现非常大的跳跃或毛刺,明显不符合人声的变化规律。
#### 可能原因
-PITD 表情提取器中,两个置信度阈值设置**过低**,错误的识别结果被采信。
+1. 在早于 `v0.9.0` 的版本中,PITD 表情曲线取值与音高之间的换算有问题。
+2. 参考音频的对应时间戳附近有噪声。您可以在 [`expressive-viewer`](#可视化工具viewer) 中观察到,原始的音高曲线中有很多不该出现的尖刺。
+3. PITD 表情提取器中,两个置信度阈值设置**过低**,错误的识别结果被采信。您可以在 [`expressive-viewer`](#可视化工具viewer) 中观察到,原始的音高曲线中有很多不该出现的尖刺。
#### 解决方案
-请先尝试使用效果最好的 rmvpe-onnx 后端(默认置信度阈值)。若问题仍在,尝试增加两个置信度阈值。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。
+1. 请下载安装 `v0.9.0` 及之后的版本。
+2. 可使用 [UVR](https://github.com/Anjok07/ultimatevocalremovergui) 、[MSST](https://github.com/SUC-DriverOld/MSST-WebUI) 等工具对参考音频去噪声(denoise)。
+3. 请先尝试使用效果最好的 rmvpe-onnx 或 hybrid 后端(默认置信度阈值)。若问题仍在,尝试增加两个置信度阈值。您可以参考 [`expressive-viewer`](#可视化工具viewer) 的音高置信度曲线来辅助调整。一般来说,**歌姬音声**比较纯净,可以先调整**参考人声**的置信度阈值。
+
+#### 未来计划
+为 PITD 表情提取算法引入语义信息。
diff --git "a/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json" "b/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json"
index 2f6e244..1dd2b38 100644
--- "a/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json"
+++ "b/examples/\320\237\321\200\320\265\320\272\321\200\320\260\321\201\320\275\320\276\320\265 \320\224\320\260\320\273\320\265\320\272\320\276/expressive_config.json"
@@ -18,13 +18,13 @@
},
"pitd": {
"selected": true,
- "backend": "rmvpe-onnx",
+ "backend": "hybrid",
"confidence_utau": null,
"confidence_ref": null,
"align_radius": 1,
"semitone_shift": 0,
- "smoothness": 4,
- "scaler": 2.2
+ "smoothness": 2,
+ "scaler": 1.0
},
"tenc": {
"selected": true,
@@ -32,7 +32,7 @@
"align_radius": 1,
"smoothness": 6,
"scaler": 1.0,
- "bias": 10
+ "bias": 15
}
}
}
diff --git "a/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json" "b/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json"
index cd5bb2a..813ec20 100644
--- "a/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json"
+++ "b/examples/\343\203\206\343\203\210\343\203\252\343\202\271/expressive_config.json"
@@ -18,13 +18,13 @@
},
"pitd": {
"selected": true,
- "backend": "rmvpe-onnx",
+ "backend": "hybrid",
"confidence_utau": null,
"confidence_ref": null,
"align_radius": 1,
"semitone_shift": 0,
"smoothness": 2,
- "scaler": 2.0
+ "scaler": 1.0
},
"tenc": {
"selected": true,
diff --git "a/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json" "b/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json"
index 73a4124..f7418ae 100644
--- "a/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json"
+++ "b/examples/\346\230\216\345\244\251\344\274\232\346\233\264\345\245\275/expressive_config.json"
@@ -18,13 +18,13 @@
},
"pitd": {
"selected": true,
- "backend": "rmvpe-onnx",
+ "backend": "hybrid",
"confidence_utau": null,
"confidence_ref": null,
"align_radius": 1,
"semitone_shift": 0,
"smoothness": 2,
- "scaler": 2.0
+ "scaler": 1.0
},
"tenc": {
"selected": true,
@@ -32,7 +32,7 @@
"align_radius": 1,
"smoothness": 6,
"scaler": 1.2,
- "bias": 10
+ "bias": 15
}
}
}
diff --git a/expressions/dyn.py b/expressions/dyn.py
index cbb331a..3df9cd8 100644
--- a/expressions/dyn.py
+++ b/expressions/dyn.py
@@ -10,6 +10,7 @@
register_expression
)
from utils.seqtool import (
+ seq_spline_smoothing,
unify_sequence_time,
align_sequence_tick,
gaussian_filter1d_with_nan,
@@ -24,10 +25,11 @@ class DynLoader(ExpressionLoader):
expression_name = "dyn"
expression_info = _l("Dynamics (curve)")
args = SimpleNamespace(
- trim_silence = Args(name="trim_silence", type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501
- align_radius = Args(name="align_radius", type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501
- smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501
- scaler = Args(name="scaler" , type=float, default=1.5 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501
+ trim_silence = Args(name="trim_silence" , type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501
+ align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501
+ smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501
+ scaler = Args(name="scaler" , type=float, default=1.5 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501
+ spline_smoothing = Args(name="spline_smoothing", type=bool , default=True, help=_l("Perform **spline smoothing** on the final expression curve for extra smoothness")), # noqa: E501
)
plots = SimpleNamespace(
expression = Plot(tag=expression_info , title=expression_info , x_label=_l("Tick") , y_label=expression_name, legends=[expression_name] ), # noqa: E501
@@ -37,10 +39,11 @@ class DynLoader(ExpressionLoader):
def get_expression(
self,
- trim_silence = args.trim_silence.default,
- align_radius = args.align_radius.default,
- smoothness = args.smoothness .default,
- scaler = args.scaler .default,
+ trim_silence = args.trim_silence .default,
+ align_radius = args.align_radius .default,
+ smoothness = args.smoothness .default,
+ scaler = args.scaler .default,
+ spline_smoothing = args.spline_smoothing.default,
):
self.logger.info(_("Extracting expression..."))
@@ -68,6 +71,12 @@ def get_expression(
# Generate expression curve
dyn_val = get_experssion_dynamics(time_aligned_ref_rms, smoothness, scaler)
+ if spline_smoothing:
+ # Final spline smoothing of the expression curve
+ # NOTE: All NaN positions except the leading and trailing ones will be interpolated
+ # Only preserving NaN at the head/tail to avoid edge artifacts of spline smoothing
+ dyn_val = seq_spline_smoothing(dyn_tick, dyn_val, nan_policy='preserve_head_tail')
+
# Collect plots
self.collect_plot(self.plots.expression, (dyn_tick, dyn_val))
self.collect_plot(self.plots.raw_rms, (ref_time, ref_rms), (utau_time, utau_rms))
diff --git a/expressions/pitd.py b/expressions/pitd.py
index fc6bd82..5ae40dd 100644
--- a/expressions/pitd.py
+++ b/expressions/pitd.py
@@ -12,6 +12,7 @@
)
from utils.i18n import _, _l, _lf
from utils.seqtool import (
+ seq_spline_smoothing,
unify_sequence_time,
align_sequence_tick,
gaussian_filter1d_with_nan,
@@ -29,17 +30,19 @@ class PitdLoader(ExpressionLoader):
"rmvpe-onnx": _l("finest accuracy, fast, CPU only (ONNX Runtime)"),
"swift-f0": _l("fair accuracy, fastest, CPU only (ONNX Runtime)"),
"crepe": _l("good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)"),
+ "hybrid": _l("based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)"),
}
- confidence_utau_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.95, "crepe": 0.80}
- confidence_ref_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.93, "crepe": 0.60}
+ confidence_utau_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.95, "crepe": 0.80, "hybrid": 0.03}
+ confidence_ref_recommended = {"rmvpe-onnx": 0.03, "swift-f0": 0.93, "crepe": 0.60, "hybrid": 0.03}
args = SimpleNamespace(
- backend = Args(name="backend" , type=str , default="rmvpe-onnx", choices=list(backend_choices.keys()), help=_lf("**F0 detection backend** for extracting pitch from WAV files. Available options:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.backend_choices.items()]))), # noqa: E501
- confidence_utau = Args(name="confidence_utau", type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **UTAU** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_utau_recommended.items()]))), # noqa: E501
- confidence_ref = Args(name="confidence_ref" , type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **reference** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_ref_recommended.items()]))), # noqa: E501
- align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501
- semitone_shift = Args(name="semitone_shift" , type=int , default=None, help=_l("**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is an octave higher than the reference WAV, set to 12; if lower, set to -12. Omit to enable automatic shift estimation")), # noqa: E501
- smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501
- scaler = Args(name="scaler" , type=float, default=2.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501
+ backend = Args(name="backend" , type=str , default="rmvpe-onnx", choices=list(backend_choices.keys()), help=_lf("**F0 detection backend** for extracting pitch from WAV files. Available options:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.backend_choices.items()]))), # noqa: E501
+ confidence_utau = Args(name="confidence_utau" , type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **UTAU** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_utau_recommended.items()]))), # noqa: E501
+ confidence_ref = Args(name="confidence_ref" , type=float, default=None, help=_lf("Minimum **confidence level** for keeping detected pitch values in the **reference** WAV. Lower values retain more frames but may include errors. Omit to use the recommended value for the selected backend:\n\n%s\n\n", lambda: "\n".join([f"- `{k}`: {v}" for k, v in PitdLoader.confidence_ref_recommended.items()]))), # noqa: E501
+ align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501
+ semitone_shift = Args(name="semitone_shift" , type=int , default=None, help=_l("**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is an octave higher than the reference WAV, set to 12; if lower, set to -12. Omit to enable automatic shift estimation")), # noqa: E501
+ smoothness = Args(name="smoothness" , type=int , default=2 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501
+ scaler = Args(name="scaler" , type=float, default=1.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501
+ spline_smoothing = Args(name="spline_smoothing", type=bool , default=True, help=_l("Perform **spline smoothing** on the final expression curve for extra smoothness")), # noqa: E501
)
plots = SimpleNamespace(
expression = Plot(tag=expression_info , title=expression_info , x_label=_l("Tick") , y_label=expression_name , legends=[expression_name] ), # noqa: E501
@@ -50,13 +53,14 @@ class PitdLoader(ExpressionLoader):
def get_expression(
self,
- backend = args.backend .default,
- confidence_utau = args.confidence_utau.default,
- confidence_ref = args.confidence_ref .default,
- align_radius = args.align_radius .default,
- semitone_shift = args.semitone_shift .default,
- smoothness = args.smoothness .default,
- scaler = args.scaler .default,
+ backend = args.backend .default,
+ confidence_utau = args.confidence_utau .default,
+ confidence_ref = args.confidence_ref .default,
+ align_radius = args.align_radius .default,
+ semitone_shift = args.semitone_shift .default,
+ smoothness = args.smoothness .default,
+ scaler = args.scaler .default,
+ spline_smoothing = args.spline_smoothing.default,
):
self.logger.info(_("Extracting expression..."))
@@ -94,16 +98,22 @@ def get_expression(
time_aligned_ref_pitch,
unified_utau_pitch,
semitone_shift=semitone_shift,
- smoothness=smoothness,
)
# Calculate pitch delta for USTX pitch editing
pitd_val = get_pitch_delta(
time_pitch_aligned_ref_pitch,
unified_utau_pitch,
+ smoothness=smoothness,
scaler=scaler,
)
+ if spline_smoothing:
+ # Final spline smoothing of the expression curve
+ # NOTE: All NaN positions except the leading and trailing ones will be interpolated
+ # Only preserving NaN at the head/tail to avoid edge artifacts of spline smoothing
+ pitd_val = seq_spline_smoothing(pitd_tick, pitd_val, nan_policy='preserve_head_tail')
+
# Collect plots
self.collect_plot(self.plots.expression, (pitd_tick, pitd_val))
self.collect_plot(self.plots.confidence, (ref_time, ref_confidence), (utau_time, utau_confidence))
@@ -167,7 +177,7 @@ def get_wav_features(wav_path, backend="rmvpe-onnx", confidence_threshold=0.8, c
return wav_time, wav_pitch, wav_confidence, wav_features
-def align_sequence_pitch(query, reference, semitone_shift=None, smoothness=0):
+def align_sequence_pitch(query, reference, semitone_shift=None):
"""Align pitch sequences by shifting in semitones and applying smoothing.
Args:
@@ -175,7 +185,6 @@ def align_sequence_pitch(query, reference, semitone_shift=None, smoothness=0):
reference (numpy.ndarray): Target reference pitch values.
semitone_shift (int, optional): Semitones to shift the query pitch.
If None, estimated automatically.
- smoothness (int, optional): Smoothing sigma. Defaults to 0.
Returns:
tuple: (pitch_aligned_query, semitone_shift)
@@ -189,22 +198,30 @@ def align_sequence_pitch(query, reference, semitone_shift=None, smoothness=0):
)
print(_("Estimated Semitone-shift: {}").format(semitone_shift))
- pitch_aligned_query = gaussian_filter1d_with_nan(
- query * np.exp2(semitone_shift / 12),
- sigma=smoothness,
- )
+ pitch_aligned_query = query * np.exp2(semitone_shift / 12)
return pitch_aligned_query, semitone_shift
-def get_pitch_delta(query, reference, scaler=2.5):
+def get_pitch_delta(query, reference, smoothness=2, scaler=1.0):
"""Calculate the scaled pitch difference between two sequences.
+ PITD is expressed in cents (100 cents = 1 semitone).
+ The renderer applies PITD on top of the base pitch, so the delta
+ must be in the same unit OpenUtau expects for the PITD expression.
+
Args:
- query (numpy.ndarray): Pitch values from the query sequence.
- reference (numpy.ndarray): Pitch values from the reference sequence.
- scaler (float, optional): Scaling factor. Defaults to 2.5.
+ query (numpy.ndarray): Pitch values from the query sequence.
+ reference (numpy.ndarray): Pitch values from the reference sequence.
+ smoothness (int, optional): Smoothing sigma. Defaults to 2.
+ scaler (float, optional): Scaling factor. Defaults to 1.0.
Returns:
- numpy.ndarray: Scaled pitch difference values.
+ numpy.ndarray: Scaled pitch difference in cents, preserving NaN for unvoiced frames.
"""
- return scaler * (query - reference)
+ voiced = (query > 0) & (reference > 0)
+
+ delta = np.full_like(query, fill_value=np.nan)
+ delta[voiced] = 1200.0 * np.log2(query[voiced] / reference[voiced])
+
+ delta = gaussian_filter1d_with_nan(delta, sigma=smoothness)
+ return scaler * delta
diff --git a/expressions/tenc.py b/expressions/tenc.py
index 730d56e..9078a49 100644
--- a/expressions/tenc.py
+++ b/expressions/tenc.py
@@ -10,6 +10,7 @@
register_expression
)
from utils.seqtool import (
+ seq_spline_smoothing,
unify_sequence_time,
align_sequence_tick,
gaussian_filter1d_with_nan,
@@ -24,11 +25,12 @@ class TencLoader(ExpressionLoader):
expression_name = "tenc"
expression_info = _l("Tension (curve)")
args = SimpleNamespace(
- trim_silence = Args(name="trim_silence", type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501
- align_radius = Args(name="align_radius", type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501
- smoothness = Args(name="smoothness" , type=int , default=6 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501
- scaler = Args(name="scaler" , type=float, default=1.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501
- bias = Args(name="bias" , type=int , default=10 , help=_l("**Bias** offset added to the expression curve. Positive values shift the curve upward; negative values shift it downward")), # noqa: E501
+ trim_silence = Args(name="trim_silence" , type=bool , default=True, help=_l("**Trim silence** from the leading and trailing edges of the audio before extracting expression")), # noqa: E501
+ align_radius = Args(name="align_radius" , type=int , default=1 , help=_l("**Radius** for the FastDTW alignment algorithm; larger values allow more flexible alignment but increase computation time")), # noqa: E501
+ smoothness = Args(name="smoothness" , type=int , default=6 , help=_l("Controls the **smoothness** of the expression curve using Gaussian filtering. Higher values produce smoother curves but may lose fine detail")), # noqa: E501
+ scaler = Args(name="scaler" , type=float, default=1.0 , help=_l("**Scaling factor** applied to the expression curve. Values >1 amplify the expression, =1 keeps original intensity, <1 reduces it")), # noqa: E501
+ bias = Args(name="bias" , type=int , default=10 , help=_l("**Bias** offset added to the expression curve. Positive values shift the curve upward; negative values shift it downward")), # noqa: E501
+ spline_smoothing = Args(name="spline_smoothing", type=bool , default=True, help=_l("Perform **spline smoothing** on the final expression curve for extra smoothness")), # noqa: E501
)
plots = SimpleNamespace(
expression = Plot(tag=expression_info , title=expression_info , x_label=_l("Tick") , y_label=expression_name, legends=[expression_name] ), # noqa: E501
@@ -38,11 +40,12 @@ class TencLoader(ExpressionLoader):
def get_expression(
self,
- trim_silence = args.trim_silence.default,
- align_radius = args.align_radius.default,
- smoothness = args.smoothness .default,
- scaler = args.scaler .default,
- bias = args.bias .default,
+ trim_silence = args.trim_silence .default,
+ align_radius = args.align_radius .default,
+ smoothness = args.smoothness .default,
+ scaler = args.scaler .default,
+ bias = args.bias .default,
+ spline_smoothing = args.spline_smoothing.default,
):
self.logger.info(_("Extracting expression..."))
@@ -70,6 +73,12 @@ def get_expression(
# Generate expression curve
tenc_val = get_experssion_tension(time_aligned_ref_rms, smoothness, scaler, bias)
+ if spline_smoothing:
+ # Final spline smoothing of the expression curve
+ # NOTE: All NaN positions except the leading and trailing ones will be interpolated
+ # Only preserving NaN at the head/tail to avoid edge artifacts of spline smoothing
+ tenc_val = seq_spline_smoothing(tenc_tick, tenc_val, nan_policy='preserve_head_tail')
+
# Collect plots
self.collect_plot(self.plots.expression, (tenc_tick, tenc_val))
self.collect_plot(self.plots.raw_rms, (ref_time, ref_rms), (utau_time, utau_rms))
diff --git a/expressive_gui.py b/expressive_gui.py
index 58b4416..3d8a392 100644
--- a/expressive_gui.py
+++ b/expressive_gui.py
@@ -543,12 +543,15 @@ async def update_placeholders():
with ui.card().classes("w-full").bind_visibility_from(
state["expressions"]["dyn"], "selected"
):
- with ui.row().classes("w-full"):
- ui.label(dyn_info).classes("text-lg font-bold")
- ui.space()
+ ui.label(dyn_info).classes("text-lg font-bold")
+
+ with ui.grid(columns=2).classes("w-full"):
ui.switch(_("Trim Silence")).bind_value(
state["expressions"]["dyn"], "trim_silence",
).tooltip_md(dyn_args.trim_silence.help)
+ ui.switch(_("Spline Smoothing")).bind_value(
+ state["expressions"]["dyn"], "spline_smoothing",
+ ).tooltip_md(dyn_args.spline_smoothing.help)
with ui.grid(columns=3).classes("w-full"):
ui.number(label=_("Align Radius"), min=1, format="%d").bind_value(
@@ -574,6 +577,11 @@ async def update_placeholders():
):
ui.label(pitd_info).classes("text-lg font-bold")
+ with ui.grid(columns=2).classes("w-full"):
+ ui.switch(_("Spline Smoothing")).bind_value(
+ state["expressions"]["pitd"], "spline_smoothing",
+ ).tooltip_md(pitd_args.spline_smoothing.help)
+
with ui.grid(columns=3).classes("w-full"):
def on_backend_change(e):
nonlocal ui_confidence_utau, ui_confidence_ref
@@ -633,12 +641,15 @@ def on_backend_change(e):
with ui.card().classes("w-full").bind_visibility_from(
state["expressions"]["tenc"], "selected"
):
- with ui.row().classes("w-full"):
- ui.label(tenc_info).classes("text-lg font-bold")
- ui.space()
+ ui.label(tenc_info).classes("text-lg font-bold")
+
+ with ui.grid(columns=2).classes("w-full"):
ui.switch(_("Trim Silence")).bind_value(
state["expressions"]["tenc"], "trim_silence",
).tooltip_md(tenc_args.trim_silence.help)
+ ui.switch(_("Spline Smoothing")).bind_value(
+ state["expressions"]["tenc"], "spline_smoothing",
+ ).tooltip_md(tenc_args.spline_smoothing.help)
with ui.grid(columns=3).classes("w-full"):
ui.number(label=_("Align Radius"), min=1, format="%d").bind_value(
diff --git a/locales/app.pot b/locales/app.pot
index e9c0b29..12d076e 100644
--- a/locales/app.pot
+++ b/locales/app.pot
@@ -8,7 +8,7 @@ msgid ""
msgstr ""
"Project-Id-Version: expressive VERSION\n"
"Report-Msgid-Bugs-To: EMAIL@ADDRESS\n"
-"POT-Creation-Date: 2026-04-11 19:51+0800\n"
+"POT-Creation-Date: 2026-04-15 21:12+0800\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME \n"
"Language-Team: LANGUAGE \n"
@@ -148,59 +148,63 @@ msgstr ""
msgid "Expression Selection"
msgstr ""
-#: expressive_gui.py:549 expressive_gui.py:639
+#: expressive_gui.py:549 expressive_gui.py:647
msgid "Trim Silence"
msgstr ""
-#: expressive_gui.py:554 expressive_gui.py:609 expressive_gui.py:644
+#: expressive_gui.py:552 expressive_gui.py:581 expressive_gui.py:650
+msgid "Spline Smoothing"
+msgstr ""
+
+#: expressive_gui.py:557 expressive_gui.py:617 expressive_gui.py:655
msgid "Align Radius"
msgstr ""
-#: expressive_gui.py:559 expressive_gui.py:620 expressive_gui.py:649
+#: expressive_gui.py:562 expressive_gui.py:628 expressive_gui.py:660
msgid "Smoothness"
msgstr ""
-#: expressive_gui.py:564 expressive_gui.py:625 expressive_gui.py:654
+#: expressive_gui.py:567 expressive_gui.py:633 expressive_gui.py:665
msgid "Scaler"
msgstr ""
-#: expressive_gui.py:593
+#: expressive_gui.py:601
msgid "UTAU Confidence"
msgstr ""
-#: expressive_gui.py:599
+#: expressive_gui.py:607
msgid "Reference Confidence"
msgstr ""
-#: expressive_gui.py:605
+#: expressive_gui.py:613
msgid "Backend"
msgstr ""
-#: expressive_gui.py:614
+#: expressive_gui.py:622
msgid "Semitone Shift"
msgstr ""
-#: expressive_gui.py:615
+#: expressive_gui.py:623
msgid "Auto Estimation"
msgstr ""
-#: expressive_gui.py:659
+#: expressive_gui.py:670
msgid "Bias"
msgstr ""
-#: expressive_gui.py:667
+#: expressive_gui.py:678
msgid "Import Config"
msgstr ""
-#: expressive_gui.py:673
+#: expressive_gui.py:684
msgid "Export Config"
msgstr ""
-#: expressive_gui.py:684
+#: expressive_gui.py:695
msgid "Process"
msgstr ""
-#: expressive_gui.py:726
+#: expressive_gui.py:737
msgid "Migrate expressions from real singers to DiffSingers (GUI)"
msgstr ""
@@ -279,102 +283,112 @@ msgstr ""
msgid "Expression result is empty. Skipping USTX update."
msgstr ""
-#: expressions/dyn.py:25
+#: expressions/dyn.py:26
msgid "Dynamics (curve)"
msgstr ""
-#: expressions/dyn.py:27 expressions/tenc.py:27
+#: expressions/dyn.py:28 expressions/tenc.py:28
msgid ""
"**Trim silence** from the leading and trailing edges of the audio before "
"extracting expression"
msgstr ""
-#: expressions/dyn.py:28 expressions/pitd.py:39 expressions/tenc.py:28
+#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29
msgid ""
"**Radius** for the FastDTW alignment algorithm; larger values allow more "
"flexible alignment but increase computation time"
msgstr ""
-#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29
+#: expressions/dyn.py:30 expressions/pitd.py:43 expressions/tenc.py:30
msgid ""
"Controls the **smoothness** of the expression curve using Gaussian "
"filtering. Higher values produce smoother curves but may lose fine detail"
msgstr ""
-#: expressions/dyn.py:30 expressions/pitd.py:42 expressions/tenc.py:30
+#: expressions/dyn.py:31 expressions/pitd.py:44 expressions/tenc.py:31
msgid ""
"**Scaling factor** applied to the expression curve. Values >1 amplify the"
" expression, =1 keeps original intensity, <1 reduces it"
msgstr ""
-#: expressions/dyn.py:33 expressions/dyn.py:35 expressions/pitd.py:45
-#: expressions/pitd.py:48 expressions/tenc.py:34 expressions/tenc.py:36
+#: expressions/dyn.py:32 expressions/pitd.py:45 expressions/tenc.py:33
+msgid ""
+"Perform **spline smoothing** on the final expression curve for extra "
+"smoothness"
+msgstr ""
+
+#: expressions/dyn.py:35 expressions/dyn.py:37 expressions/pitd.py:48
+#: expressions/pitd.py:51 expressions/tenc.py:36 expressions/tenc.py:38
msgid "Tick"
msgstr ""
-#: expressions/dyn.py:34 expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/tenc.py:37
msgid "raw_rms"
msgstr ""
-#: expressions/dyn.py:34 expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/tenc.py:37
msgid "Raw RMS"
msgstr ""
-#: expressions/dyn.py:34 expressions/pitd.py:46 expressions/pitd.py:47
-#: expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/pitd.py:49 expressions/pitd.py:50
+#: expressions/tenc.py:37
msgid "Time (s)"
msgstr ""
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "RMS"
msgstr ""
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46
-#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49
+#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "Reference"
msgstr ""
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46
-#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49
+#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "UTAU"
msgstr ""
-#: expressions/dyn.py:35 expressions/tenc.py:36
+#: expressions/dyn.py:37 expressions/tenc.py:38
msgid "aligned_rms"
msgstr ""
-#: expressions/dyn.py:35 expressions/tenc.py:36
+#: expressions/dyn.py:37 expressions/tenc.py:38
msgid "Aligned RMS"
msgstr ""
-#: expressions/dyn.py:45 expressions/pitd.py:61 expressions/tenc.py:47
+#: expressions/dyn.py:48 expressions/pitd.py:65 expressions/tenc.py:50
msgid "Extracting expression..."
msgstr ""
-#: expressions/dyn.py:77 expressions/pitd.py:114 expressions/tenc.py:79
+#: expressions/dyn.py:86 expressions/pitd.py:124 expressions/tenc.py:88
msgid "Expression extraction complete."
msgstr ""
-#: expressions/pitd.py:27
+#: expressions/pitd.py:28
msgid "Pitch Deviation (curve)"
msgstr ""
-#: expressions/pitd.py:29
+#: expressions/pitd.py:30
msgid "finest accuracy, fast, CPU only (ONNX Runtime)"
msgstr ""
-#: expressions/pitd.py:30
+#: expressions/pitd.py:31
msgid "fair accuracy, fastest, CPU only (ONNX Runtime)"
msgstr ""
-#: expressions/pitd.py:31
+#: expressions/pitd.py:32
msgid "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)"
msgstr ""
-#: expressions/pitd.py:36
+#: expressions/pitd.py:33
+msgid "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)"
+msgstr ""
+
+#: expressions/pitd.py:38
#, python-format
msgid ""
"**F0 detection backend** for extracting pitch from WAV files. Available "
@@ -384,7 +398,7 @@ msgid ""
"\n"
msgstr ""
-#: expressions/pitd.py:37
+#: expressions/pitd.py:39
#, python-format
msgid ""
"Minimum **confidence level** for keeping detected pitch values in the "
@@ -395,7 +409,7 @@ msgid ""
"\n"
msgstr ""
-#: expressions/pitd.py:38
+#: expressions/pitd.py:40
#, python-format
msgid ""
"Minimum **confidence level** for keeping detected pitch values in the "
@@ -406,55 +420,55 @@ msgid ""
"\n"
msgstr ""
-#: expressions/pitd.py:40
+#: expressions/pitd.py:42
msgid ""
"**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is"
" an octave higher than the reference WAV, set to 12; if lower, set to "
"-12. Omit to enable automatic shift estimation"
msgstr ""
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "confidence"
msgstr ""
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "Pitch Extraction Confidence"
msgstr ""
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "Confidence"
msgstr ""
-#: expressions/pitd.py:47
+#: expressions/pitd.py:50
msgid "raw_pitch"
msgstr ""
-#: expressions/pitd.py:47
+#: expressions/pitd.py:50
msgid "Raw Pitch"
msgstr ""
-#: expressions/pitd.py:47 expressions/pitd.py:48
+#: expressions/pitd.py:50 expressions/pitd.py:51
msgid "Pitch (Hz)"
msgstr ""
-#: expressions/pitd.py:48
+#: expressions/pitd.py:51
msgid "aligned_pitch"
msgstr ""
-#: expressions/pitd.py:48
+#: expressions/pitd.py:51
msgid "Aligned Pitch"
msgstr ""
-#: expressions/pitd.py:190
+#: expressions/pitd.py:199
#, python-brace-format
msgid "Estimated Semitone-shift: {}"
msgstr ""
-#: expressions/tenc.py:25
+#: expressions/tenc.py:26
msgid "Tension (curve)"
msgstr ""
-#: expressions/tenc.py:31
+#: expressions/tenc.py:32
msgid ""
"**Bias** offset added to the expression curve. Positive values shift the "
"curve upward; negative values shift it downward"
@@ -472,22 +486,22 @@ msgstr ""
msgid "Zoom"
msgstr ""
-#: utils/wavtool.py:87
+#: utils/wavtool.py:88
#, python-brace-format
msgid "Loading F0 data from cache file: '{}'"
msgstr ""
-#: utils/wavtool.py:130
+#: utils/wavtool.py:133
#, python-brace-format
msgid "F0 data saved to cache file: '{}'"
msgstr ""
-#: utils/wavtool.py:304
+#: utils/wavtool.py:380
#, python-brace-format
msgid "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
msgstr ""
-#: utils/wavtool.py:310
+#: utils/wavtool.py:386
#, python-brace-format
msgid "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
msgstr ""
diff --git a/locales/en/LC_MESSAGES/app.po b/locales/en/LC_MESSAGES/app.po
index 7c47ad1..2621c73 100644
--- a/locales/en/LC_MESSAGES/app.po
+++ b/locales/en/LC_MESSAGES/app.po
@@ -7,7 +7,7 @@ msgid ""
msgstr ""
"Project-Id-Version: expressive\n"
"Report-Msgid-Bugs-To: https://github.com/NewComer00/expressive/issues\n"
-"POT-Creation-Date: 2026-04-11 19:51+0800\n"
+"POT-Creation-Date: 2026-04-15 21:12+0800\n"
"PO-Revision-Date: 2026-03-01 20:21+0800\n"
"Last-Translator: NewComer00\n"
"Language: en\n"
@@ -151,59 +151,63 @@ msgstr "Track Number"
msgid "Expression Selection"
msgstr "Expression Selection"
-#: expressive_gui.py:549 expressive_gui.py:639
+#: expressive_gui.py:549 expressive_gui.py:647
msgid "Trim Silence"
msgstr "Trim Silence"
-#: expressive_gui.py:554 expressive_gui.py:609 expressive_gui.py:644
+#: expressive_gui.py:552 expressive_gui.py:581 expressive_gui.py:650
+msgid "Spline Smoothing"
+msgstr "Spline Smoothing"
+
+#: expressive_gui.py:557 expressive_gui.py:617 expressive_gui.py:655
msgid "Align Radius"
msgstr "Align Radius"
-#: expressive_gui.py:559 expressive_gui.py:620 expressive_gui.py:649
+#: expressive_gui.py:562 expressive_gui.py:628 expressive_gui.py:660
msgid "Smoothness"
msgstr "Smoothness"
-#: expressive_gui.py:564 expressive_gui.py:625 expressive_gui.py:654
+#: expressive_gui.py:567 expressive_gui.py:633 expressive_gui.py:665
msgid "Scaler"
msgstr "Scaler"
-#: expressive_gui.py:593
+#: expressive_gui.py:601
msgid "UTAU Confidence"
msgstr "UTAU Confidence"
-#: expressive_gui.py:599
+#: expressive_gui.py:607
msgid "Reference Confidence"
msgstr "Reference Confidence"
-#: expressive_gui.py:605
+#: expressive_gui.py:613
msgid "Backend"
msgstr "Backend"
-#: expressive_gui.py:614
+#: expressive_gui.py:622
msgid "Semitone Shift"
msgstr "Semitone Shift"
-#: expressive_gui.py:615
+#: expressive_gui.py:623
msgid "Auto Estimation"
msgstr "Auto Estimation"
-#: expressive_gui.py:659
+#: expressive_gui.py:670
msgid "Bias"
msgstr "Bias"
-#: expressive_gui.py:667
+#: expressive_gui.py:678
msgid "Import Config"
msgstr "Import Config"
-#: expressive_gui.py:673
+#: expressive_gui.py:684
msgid "Export Config"
msgstr "Export Config"
-#: expressive_gui.py:684
+#: expressive_gui.py:695
msgid "Process"
msgstr "Process"
-#: expressive_gui.py:726
+#: expressive_gui.py:737
msgid "Migrate expressions from real singers to DiffSingers (GUI)"
msgstr "Migrate expressions from real singers to DiffSingers (GUI)"
@@ -290,11 +294,11 @@ msgstr "Expression written to USTX file: '{}'"
msgid "Expression result is empty. Skipping USTX update."
msgstr "Expression result is empty. Skipping USTX update."
-#: expressions/dyn.py:25
+#: expressions/dyn.py:26
msgid "Dynamics (curve)"
msgstr "Dynamics (curve)"
-#: expressions/dyn.py:27 expressions/tenc.py:27
+#: expressions/dyn.py:28 expressions/tenc.py:28
msgid ""
"**Trim silence** from the leading and trailing edges of the audio before "
"extracting expression"
@@ -302,7 +306,7 @@ msgstr ""
"**Trim silence** from the leading and trailing edges of the audio before "
"extracting expression"
-#: expressions/dyn.py:28 expressions/pitd.py:39 expressions/tenc.py:28
+#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29
msgid ""
"**Radius** for the FastDTW alignment algorithm; larger values allow more "
"flexible alignment but increase computation time"
@@ -310,7 +314,7 @@ msgstr ""
"**Radius** for the FastDTW alignment algorithm; larger values allow more "
"flexible alignment but increase computation time"
-#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29
+#: expressions/dyn.py:30 expressions/pitd.py:43 expressions/tenc.py:30
msgid ""
"Controls the **smoothness** of the expression curve using Gaussian "
"filtering. Higher values produce smoother curves but may lose fine detail"
@@ -318,7 +322,7 @@ msgstr ""
"Controls the **smoothness** of the expression curve using Gaussian "
"filtering. Higher values produce smoother curves but may lose fine detail"
-#: expressions/dyn.py:30 expressions/pitd.py:42 expressions/tenc.py:30
+#: expressions/dyn.py:31 expressions/pitd.py:44 expressions/tenc.py:31
msgid ""
"**Scaling factor** applied to the expression curve. Values >1 amplify the"
" expression, =1 keeps original intensity, <1 reduces it"
@@ -326,74 +330,86 @@ msgstr ""
"**Scaling factor** applied to the expression curve. Values >1 amplify the"
" expression, =1 keeps original intensity, <1 reduces it"
-#: expressions/dyn.py:33 expressions/dyn.py:35 expressions/pitd.py:45
-#: expressions/pitd.py:48 expressions/tenc.py:34 expressions/tenc.py:36
+#: expressions/dyn.py:32 expressions/pitd.py:45 expressions/tenc.py:33
+msgid ""
+"Perform **spline smoothing** on the final expression curve for extra "
+"smoothness"
+msgstr ""
+"Perform **spline smoothing** on the final expression curve for extra "
+"smoothness"
+
+#: expressions/dyn.py:35 expressions/dyn.py:37 expressions/pitd.py:48
+#: expressions/pitd.py:51 expressions/tenc.py:36 expressions/tenc.py:38
msgid "Tick"
msgstr "Tick"
-#: expressions/dyn.py:34 expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/tenc.py:37
msgid "raw_rms"
msgstr "raw_rms"
-#: expressions/dyn.py:34 expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/tenc.py:37
msgid "Raw RMS"
msgstr "Raw RMS"
-#: expressions/dyn.py:34 expressions/pitd.py:46 expressions/pitd.py:47
-#: expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/pitd.py:49 expressions/pitd.py:50
+#: expressions/tenc.py:37
msgid "Time (s)"
msgstr "Time (s)"
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "RMS"
msgstr "RMS"
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46
-#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49
+#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "Reference"
msgstr "Reference"
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46
-#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49
+#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "UTAU"
msgstr "UTAU"
-#: expressions/dyn.py:35 expressions/tenc.py:36
+#: expressions/dyn.py:37 expressions/tenc.py:38
msgid "aligned_rms"
msgstr "aligned_rms"
-#: expressions/dyn.py:35 expressions/tenc.py:36
+#: expressions/dyn.py:37 expressions/tenc.py:38
msgid "Aligned RMS"
msgstr "Aligned RMS"
-#: expressions/dyn.py:45 expressions/pitd.py:61 expressions/tenc.py:47
+#: expressions/dyn.py:48 expressions/pitd.py:65 expressions/tenc.py:50
msgid "Extracting expression..."
msgstr "Extracting expression..."
-#: expressions/dyn.py:77 expressions/pitd.py:114 expressions/tenc.py:79
+#: expressions/dyn.py:86 expressions/pitd.py:124 expressions/tenc.py:88
msgid "Expression extraction complete."
msgstr "Expression extraction complete."
-#: expressions/pitd.py:27
+#: expressions/pitd.py:28
msgid "Pitch Deviation (curve)"
msgstr "Pitch Deviation (curve)"
-#: expressions/pitd.py:29
+#: expressions/pitd.py:30
msgid "finest accuracy, fast, CPU only (ONNX Runtime)"
msgstr "finest accuracy, fast, CPU only (ONNX Runtime)"
-#: expressions/pitd.py:30
+#: expressions/pitd.py:31
msgid "fair accuracy, fastest, CPU only (ONNX Runtime)"
msgstr "fair accuracy, fastest, CPU only (ONNX Runtime)"
-#: expressions/pitd.py:31
+#: expressions/pitd.py:32
msgid "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)"
msgstr "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)"
-#: expressions/pitd.py:36
+#: expressions/pitd.py:33
+msgid "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)"
+msgstr "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)"
+
+#: expressions/pitd.py:38
#, python-format
msgid ""
"**F0 detection backend** for extracting pitch from WAV files. Available "
@@ -408,7 +424,7 @@ msgstr ""
"%s\n"
"\n"
-#: expressions/pitd.py:37
+#: expressions/pitd.py:39
#, python-format
msgid ""
"Minimum **confidence level** for keeping detected pitch values in the "
@@ -425,7 +441,7 @@ msgstr ""
"%s\n"
"\n"
-#: expressions/pitd.py:38
+#: expressions/pitd.py:40
#, python-format
msgid ""
"Minimum **confidence level** for keeping detected pitch values in the "
@@ -442,7 +458,7 @@ msgstr ""
"%s\n"
"\n"
-#: expressions/pitd.py:40
+#: expressions/pitd.py:42
msgid ""
"**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is"
" an octave higher than the reference WAV, set to 12; if lower, set to "
@@ -452,48 +468,48 @@ msgstr ""
" an octave higher than the reference WAV, set to 12; if lower, set to "
"-12. Omit to enable automatic shift estimation"
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "confidence"
msgstr "confidence"
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "Pitch Extraction Confidence"
msgstr "Pitch Extraction Confidence"
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "Confidence"
msgstr "Confidence"
-#: expressions/pitd.py:47
+#: expressions/pitd.py:50
msgid "raw_pitch"
msgstr "raw_pitch"
-#: expressions/pitd.py:47
+#: expressions/pitd.py:50
msgid "Raw Pitch"
msgstr "Raw Pitch"
-#: expressions/pitd.py:47 expressions/pitd.py:48
+#: expressions/pitd.py:50 expressions/pitd.py:51
msgid "Pitch (Hz)"
msgstr "Pitch (Hz)"
-#: expressions/pitd.py:48
+#: expressions/pitd.py:51
msgid "aligned_pitch"
msgstr "aligned_pitch"
-#: expressions/pitd.py:48
+#: expressions/pitd.py:51
msgid "Aligned Pitch"
msgstr "Aligned Pitch"
-#: expressions/pitd.py:190
+#: expressions/pitd.py:199
#, python-brace-format
msgid "Estimated Semitone-shift: {}"
msgstr "Estimated Semitone-shift: {}"
-#: expressions/tenc.py:25
+#: expressions/tenc.py:26
msgid "Tension (curve)"
msgstr "Tension (curve)"
-#: expressions/tenc.py:31
+#: expressions/tenc.py:32
msgid ""
"**Bias** offset added to the expression curve. Positive values shift the "
"curve upward; negative values shift it downward"
@@ -513,22 +529,22 @@ msgstr "Loop region"
msgid "Zoom"
msgstr "Zoom"
-#: utils/wavtool.py:87
+#: utils/wavtool.py:88
#, python-brace-format
msgid "Loading F0 data from cache file: '{}'"
msgstr "Loading F0 data from cache file: '{}'"
-#: utils/wavtool.py:130
+#: utils/wavtool.py:133
#, python-brace-format
msgid "F0 data saved to cache file: '{}'"
msgstr "F0 data saved to cache file: '{}'"
-#: utils/wavtool.py:304
+#: utils/wavtool.py:380
#, python-brace-format
msgid "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
msgstr "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
-#: utils/wavtool.py:310
+#: utils/wavtool.py:386
#, python-brace-format
msgid "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
msgstr "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
diff --git a/locales/zh_CN/LC_MESSAGES/app.po b/locales/zh_CN/LC_MESSAGES/app.po
index c3db131..d9ba0be 100644
--- a/locales/zh_CN/LC_MESSAGES/app.po
+++ b/locales/zh_CN/LC_MESSAGES/app.po
@@ -7,7 +7,7 @@ msgid ""
msgstr ""
"Project-Id-Version: expressive\n"
"Report-Msgid-Bugs-To: https://github.com/NewComer00/expressive/issues\n"
-"POT-Creation-Date: 2026-04-11 19:51+0800\n"
+"POT-Creation-Date: 2026-04-15 21:12+0800\n"
"PO-Revision-Date: 2026-03-01 20:21+0800\n"
"Last-Translator: NewComer00\n"
"Language: zh_CN\n"
@@ -149,59 +149,63 @@ msgstr "轨道编号"
msgid "Expression Selection"
msgstr "表情参数"
-#: expressive_gui.py:549 expressive_gui.py:639
+#: expressive_gui.py:549 expressive_gui.py:647
msgid "Trim Silence"
msgstr "剪除静音"
-#: expressive_gui.py:554 expressive_gui.py:609 expressive_gui.py:644
+#: expressive_gui.py:552 expressive_gui.py:581 expressive_gui.py:650
+msgid "Spline Smoothing"
+msgstr "样条曲线平滑"
+
+#: expressive_gui.py:557 expressive_gui.py:617 expressive_gui.py:655
msgid "Align Radius"
msgstr "对齐半径"
-#: expressive_gui.py:559 expressive_gui.py:620 expressive_gui.py:649
+#: expressive_gui.py:562 expressive_gui.py:628 expressive_gui.py:660
msgid "Smoothness"
msgstr "平滑度"
-#: expressive_gui.py:564 expressive_gui.py:625 expressive_gui.py:654
+#: expressive_gui.py:567 expressive_gui.py:633 expressive_gui.py:665
msgid "Scaler"
msgstr "缩放因子"
-#: expressive_gui.py:593
+#: expressive_gui.py:601
msgid "UTAU Confidence"
msgstr "歌姬音频置信度"
-#: expressive_gui.py:599
+#: expressive_gui.py:607
msgid "Reference Confidence"
msgstr "参考音频置信度"
-#: expressive_gui.py:605
+#: expressive_gui.py:613
msgid "Backend"
msgstr "后端"
-#: expressive_gui.py:614
+#: expressive_gui.py:622
msgid "Semitone Shift"
msgstr "半音偏移"
-#: expressive_gui.py:615
+#: expressive_gui.py:623
msgid "Auto Estimation"
msgstr "自动估算"
-#: expressive_gui.py:659
+#: expressive_gui.py:670
msgid "Bias"
msgstr "偏置"
-#: expressive_gui.py:667
+#: expressive_gui.py:678
msgid "Import Config"
msgstr "导入配置"
-#: expressive_gui.py:673
+#: expressive_gui.py:684
msgid "Export Config"
msgstr "导出配置"
-#: expressive_gui.py:684
+#: expressive_gui.py:695
msgid "Process"
msgstr "开始处理"
-#: expressive_gui.py:726
+#: expressive_gui.py:737
msgid "Migrate expressions from real singers to DiffSingers (GUI)"
msgstr "将表情参数从真实歌手迁移到 DiffSinger 歌手,适用于图形用户界面(GUI)"
@@ -280,102 +284,112 @@ msgstr "表情参数已写入 USTX 文件:'{}'"
msgid "Expression result is empty. Skipping USTX update."
msgstr "表情参数结果为空,USTX 文件将不会更新。"
-#: expressions/dyn.py:25
+#: expressions/dyn.py:26
msgid "Dynamics (curve)"
msgstr "动态曲线 Dynamics (curve)"
-#: expressions/dyn.py:27 expressions/tenc.py:27
+#: expressions/dyn.py:28 expressions/tenc.py:28
msgid ""
"**Trim silence** from the leading and trailing edges of the audio before "
"extracting expression"
msgstr "在提取表情特征前,**剪除**音频开头与结尾的**静音部分**"
-#: expressions/dyn.py:28 expressions/pitd.py:39 expressions/tenc.py:28
+#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29
msgid ""
"**Radius** for the FastDTW alignment algorithm; larger values allow more "
"flexible alignment but increase computation time"
msgstr "FastDTW 对齐算法的**半径**;值越大对齐越灵活,但计算时间也越长"
-#: expressions/dyn.py:29 expressions/pitd.py:41 expressions/tenc.py:29
+#: expressions/dyn.py:30 expressions/pitd.py:43 expressions/tenc.py:30
msgid ""
"Controls the **smoothness** of the expression curve using Gaussian "
"filtering. Higher values produce smoother curves but may lose fine detail"
msgstr "通过高斯滤波控制表情曲线的**平滑度**;值越大曲线越平滑,但可能丢失细节"
-#: expressions/dyn.py:30 expressions/pitd.py:42 expressions/tenc.py:30
+#: expressions/dyn.py:31 expressions/pitd.py:44 expressions/tenc.py:31
msgid ""
"**Scaling factor** applied to the expression curve. Values >1 amplify the"
" expression, =1 keeps original intensity, <1 reduces it"
msgstr "应用于表情曲线的**缩放因子**;大于 1 则放大,等于 1 则保持原强度,小于 1 则缩小"
-#: expressions/dyn.py:33 expressions/dyn.py:35 expressions/pitd.py:45
-#: expressions/pitd.py:48 expressions/tenc.py:34 expressions/tenc.py:36
+#: expressions/dyn.py:32 expressions/pitd.py:45 expressions/tenc.py:33
+msgid ""
+"Perform **spline smoothing** on the final expression curve for extra "
+"smoothness"
+msgstr "对提取出的表情曲线进行**样条曲线平滑**,以获得更自然的效果"
+
+#: expressions/dyn.py:35 expressions/dyn.py:37 expressions/pitd.py:48
+#: expressions/pitd.py:51 expressions/tenc.py:36 expressions/tenc.py:38
msgid "Tick"
msgstr "时间刻度(Tick)"
-#: expressions/dyn.py:34 expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/tenc.py:37
msgid "raw_rms"
msgstr "原音频的 RMS"
-#: expressions/dyn.py:34 expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/tenc.py:37
msgid "Raw RMS"
msgstr "原音频的均方根能量(RMS)特征"
-#: expressions/dyn.py:34 expressions/pitd.py:46 expressions/pitd.py:47
-#: expressions/tenc.py:35
+#: expressions/dyn.py:36 expressions/pitd.py:49 expressions/pitd.py:50
+#: expressions/tenc.py:37
msgid "Time (s)"
msgstr "时间(秒)"
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "RMS"
msgstr "RMS"
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46
-#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49
+#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "Reference"
msgstr "参考音频"
-#: expressions/dyn.py:34 expressions/dyn.py:35 expressions/pitd.py:46
-#: expressions/pitd.py:47 expressions/pitd.py:48 expressions/tenc.py:35
-#: expressions/tenc.py:36
+#: expressions/dyn.py:36 expressions/dyn.py:37 expressions/pitd.py:49
+#: expressions/pitd.py:50 expressions/pitd.py:51 expressions/tenc.py:37
+#: expressions/tenc.py:38
msgid "UTAU"
msgstr "歌姬音频"
-#: expressions/dyn.py:35 expressions/tenc.py:36
+#: expressions/dyn.py:37 expressions/tenc.py:38
msgid "aligned_rms"
msgstr "对齐后的 RMS"
-#: expressions/dyn.py:35 expressions/tenc.py:36
+#: expressions/dyn.py:37 expressions/tenc.py:38
msgid "Aligned RMS"
msgstr "对齐后的均方根能量(RMS)特征"
-#: expressions/dyn.py:45 expressions/pitd.py:61 expressions/tenc.py:47
+#: expressions/dyn.py:48 expressions/pitd.py:65 expressions/tenc.py:50
msgid "Extracting expression..."
msgstr "正在提取表情参数..."
-#: expressions/dyn.py:77 expressions/pitd.py:114 expressions/tenc.py:79
+#: expressions/dyn.py:86 expressions/pitd.py:124 expressions/tenc.py:88
msgid "Expression extraction complete."
msgstr "表情参数提取完成。"
-#: expressions/pitd.py:27
+#: expressions/pitd.py:28
msgid "Pitch Deviation (curve)"
msgstr "音高偏差曲线 Pitch Deviation (curve)"
-#: expressions/pitd.py:29
+#: expressions/pitd.py:30
msgid "finest accuracy, fast, CPU only (ONNX Runtime)"
msgstr "精度最佳,较快,仅支持 CPU,基于 ONNX Runtime"
-#: expressions/pitd.py:30
+#: expressions/pitd.py:31
msgid "fair accuracy, fastest, CPU only (ONNX Runtime)"
msgstr "精度一般,最快,仅支持 CPU,基于 ONNX Runtime"
-#: expressions/pitd.py:31
+#: expressions/pitd.py:32
msgid "good accuracy, slow, CPU & NVIDIA GPU (TensorFlow)"
msgstr "精度较好,较慢,支持 CPU 和 NVIDIA GPU,基于 TensorFlow"
-#: expressions/pitd.py:36
+#: expressions/pitd.py:33
+msgid "based on rmvpe-onnx, improved by swift-f0, CPU only (ONNX Runtime)"
+msgstr "以 rmvpe-onnx 为基准,综合了 swift-f0 的结果,仅支持 CPU,基于 ONNX Runtime"
+
+#: expressions/pitd.py:38
#, python-format
msgid ""
"**F0 detection backend** for extracting pitch from WAV files. Available "
@@ -389,7 +403,7 @@ msgstr ""
"%s\n"
"\n"
-#: expressions/pitd.py:37
+#: expressions/pitd.py:39
#, python-format
msgid ""
"Minimum **confidence level** for keeping detected pitch values in the "
@@ -404,7 +418,7 @@ msgstr ""
"%s\n"
"\n"
-#: expressions/pitd.py:38
+#: expressions/pitd.py:40
#, python-format
msgid ""
"Minimum **confidence level** for keeping detected pitch values in the "
@@ -419,55 +433,55 @@ msgstr ""
"%s\n"
"\n"
-#: expressions/pitd.py:40
+#: expressions/pitd.py:42
msgid ""
"**Semitone shift** between the UTAU and reference WAV. If the UTAU WAV is"
" an octave higher than the reference WAV, set to 12; if lower, set to "
"-12. Omit to enable automatic shift estimation"
msgstr "歌姬音频与参考音频之间的**半音偏移**;若歌姬比参考高一个八度则设为 12,低一个八度则设为 -12;忽略该参数则启用自动估算"
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "confidence"
msgstr "置信度"
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "Pitch Extraction Confidence"
msgstr "音高提取置信度"
-#: expressions/pitd.py:46
+#: expressions/pitd.py:49
msgid "Confidence"
msgstr "置信度"
-#: expressions/pitd.py:47
+#: expressions/pitd.py:50
msgid "raw_pitch"
msgstr "原音频的 Pitch"
-#: expressions/pitd.py:47
+#: expressions/pitd.py:50
msgid "Raw Pitch"
msgstr "原音频的音高(Pitch)特征"
-#: expressions/pitd.py:47 expressions/pitd.py:48
+#: expressions/pitd.py:50 expressions/pitd.py:51
msgid "Pitch (Hz)"
msgstr "音高(Hz)"
-#: expressions/pitd.py:48
+#: expressions/pitd.py:51
msgid "aligned_pitch"
msgstr "对齐后的 Pitch"
-#: expressions/pitd.py:48
+#: expressions/pitd.py:51
msgid "Aligned Pitch"
msgstr "对齐后的音高(Pitch)特征"
-#: expressions/pitd.py:190
+#: expressions/pitd.py:199
#, python-brace-format
msgid "Estimated Semitone-shift: {}"
msgstr "估计的半音偏移:{}"
-#: expressions/tenc.py:25
+#: expressions/tenc.py:26
msgid "Tension (curve)"
msgstr "张力曲线 Tension (curve)"
-#: expressions/tenc.py:31
+#: expressions/tenc.py:32
msgid ""
"**Bias** offset added to the expression curve. Positive values shift the "
"curve upward; negative values shift it downward"
@@ -485,22 +499,22 @@ msgstr "选区循环播放"
msgid "Zoom"
msgstr "缩放"
-#: utils/wavtool.py:87
+#: utils/wavtool.py:88
#, python-brace-format
msgid "Loading F0 data from cache file: '{}'"
msgstr "正在从缓存文件加载 F0 数据:'{}'"
-#: utils/wavtool.py:130
+#: utils/wavtool.py:133
#, python-brace-format
msgid "F0 data saved to cache file: '{}'"
msgstr "F0 数据已保存到缓存文件:'{}'"
-#: utils/wavtool.py:304
+#: utils/wavtool.py:380
#, python-brace-format
msgid "start {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
msgstr "开始时间 {:.3f}秒 被截取到 {:.3f}秒(完整时长:{:.3f}秒) "
-#: utils/wavtool.py:310
+#: utils/wavtool.py:386
#, python-brace-format
msgid "end {:.3f}s clamped to {:.3f}s (total duration: {:.3f}s)"
msgstr "结束时间 {:.3f}秒 被截取到 {:.3f}秒(完整时长:{:.3f}秒) "
diff --git a/tests/test_seqtool.py b/tests/test_seqtool.py
index 5df030a..5851c14 100644
--- a/tests/test_seqtool.py
+++ b/tests/test_seqtool.py
@@ -9,6 +9,7 @@
sequence_interval_union,
unify_sequence_time,
gaussian_filter1d_with_nan,
+ seq_spline_smoothing,
align_sequence_tick,
seq_dynamics_trends,
seq_rcr,
@@ -278,6 +279,98 @@ def test_smoothing_levels(self, sigma, should_smooth):
assert_array_almost_equal(result, seq)
+# ---------------------------------------------------------------------------
+# seq_spline_smoothing
+# ---------------------------------------------------------------------------
+
+class TestSeqSplineSmoothing:
+ """Test seq_spline_smoothing."""
+
+ def test_output_shape(self):
+ t = np.linspace(0, 1, 50)
+ v = np.sin(2 * np.pi * t)
+ result = seq_spline_smoothing(t, v)
+ assert result.shape == v.shape
+
+ def test_smooths_noisy_signal(self):
+ np.random.seed(0)
+ t = np.linspace(0, 1, 100)
+ v = np.sin(2 * np.pi * t) + np.random.normal(0, 0.3, 100)
+ result = seq_spline_smoothing(t, v)
+ assert np.var(result) < np.var(v)
+
+ def test_no_nan_policy(self):
+ t = np.linspace(0, 1, 20)
+ v = np.sin(2 * np.pi * t)
+ v[5] = np.nan
+ v[10] = np.nan
+ result = seq_spline_smoothing(t, v, nan_policy='no_nan')
+ assert not np.any(np.isnan(result))
+
+ def test_preserve_all_nan_policy(self):
+ t = np.linspace(0, 1, 20)
+ v = np.sin(2 * np.pi * t)
+ nan_indices = [3, 7, 15]
+ v[nan_indices] = np.nan
+ result = seq_spline_smoothing(t, v, nan_policy='preserve_all')
+ assert result.shape == v.shape
+ for i in nan_indices:
+ assert np.isnan(result[i])
+ non_nan = [i for i in range(len(v)) if i not in nan_indices]
+ assert not np.any(np.isnan(result[non_nan]))
+
+ def test_preserve_head_tail_nan_policy(self):
+ t = np.linspace(0, 1, 20)
+ v = np.sin(2 * np.pi * t)
+ # Leading, interior, and trailing NaNs
+ v[0] = np.nan
+ v[1] = np.nan
+ v[10] = np.nan
+ v[18] = np.nan
+ v[19] = np.nan
+ result = seq_spline_smoothing(t, v, nan_policy='preserve_head_tail')
+ assert result.shape == v.shape
+ # Leading and trailing NaNs preserved
+ assert np.isnan(result[0])
+ assert np.isnan(result[1])
+ assert np.isnan(result[18])
+ assert np.isnan(result[19])
+ # Interior NaN filled
+ assert not np.isnan(result[10])
+
+ def test_invalid_nan_policy_raises(self):
+ t = np.linspace(0, 1, 10)
+ v = np.ones(10)
+ with pytest.raises(ValueError, match="Unknown nan_policy"):
+ seq_spline_smoothing(t, v, nan_policy='invalid_policy')
+
+ def test_manual_lam_smoother_than_auto(self):
+ """A large lam should produce a smoother (lower variance) result than auto."""
+ np.random.seed(1)
+ t = np.linspace(0, 1, 100)
+ v = np.sin(2 * np.pi * t) + np.random.normal(0, 0.2, 100)
+ result_auto = seq_spline_smoothing(t, v, lam=None)
+ result_heavy = seq_spline_smoothing(t, v, lam=1e4)
+ assert np.var(result_heavy) <= np.var(result_auto)
+
+ def test_no_nan_input_all_policies_agree(self):
+ """With no NaNs, all nan_policy values should return identical results."""
+ t = np.linspace(0, 1, 30)
+ v = np.cos(2 * np.pi * t)
+ r_no_nan = seq_spline_smoothing(t, v, nan_policy='no_nan')
+ r_preserve = seq_spline_smoothing(t, v, nan_policy='preserve_all')
+ r_head_tail = seq_spline_smoothing(t, v, nan_policy='preserve_head_tail')
+ assert_array_almost_equal(r_no_nan, r_preserve)
+ assert_array_almost_equal(r_no_nan, r_head_tail)
+
+ @pytest.mark.parametrize("nan_policy", ['no_nan', 'preserve_all', 'preserve_head_tail'])
+ def test_all_policies_accepted(self, nan_policy):
+ t = np.linspace(0, 1, 20)
+ v = np.sin(2 * np.pi * t)
+ result = seq_spline_smoothing(t, v, nan_policy=nan_policy)
+ assert result.shape == v.shape
+
+
# ---------------------------------------------------------------------------
# seq_dynamics_trends
# ---------------------------------------------------------------------------
diff --git a/tests/test_wavtool.py b/tests/test_wavtool.py
index 5815435..e87c7db 100644
--- a/tests/test_wavtool.py
+++ b/tests/test_wavtool.py
@@ -4,6 +4,7 @@
import csv
import os
import tempfile
+import types
import unittest
from unittest.mock import MagicMock, patch
@@ -35,6 +36,16 @@ def _make_wav(duration: float = 5.0, sr: int = 22050) -> str:
return tmp.name
+def _make_tonal_wav(duration: float = 2.0, sr: int = 22050, freq: float = 440.0) -> str:
+ """Write a sine-wave WAV so RMS is non-trivially non-zero throughout."""
+ t = np.linspace(0, duration, int(duration * sr), endpoint=False)
+ y = (0.5 * np.sin(2 * np.pi * freq * t)).astype(np.float32)
+ tmp = tempfile.NamedTemporaryFile(suffix=".wav", delete=False)
+ sf.write(tmp.name, y, sr)
+ tmp.close()
+ return tmp.name
+
+
# ---------------------------------------------------------------------------
# timestamp2sec
# ---------------------------------------------------------------------------
@@ -586,15 +597,6 @@ def tearDown(self):
except FileNotFoundError:
pass
- def _make_tonal_wav(self, duration=2.0, sr=22050, freq=440.0) -> str:
- """Write a sine-wave WAV so RMS is non-trivially non-zero throughout."""
- t = np.linspace(0, duration, int(duration * sr), endpoint=False)
- y = (0.5 * np.sin(2 * np.pi * freq * t)).astype(np.float32)
- tmp = tempfile.NamedTemporaryFile(suffix=".wav", delete=False)
- sf.write(tmp.name, y, sr)
- tmp.close()
- return tmp.name
-
# --- return types and shapes ---
def test_returns_tuple_of_two(self):
@@ -630,6 +632,11 @@ def test_rms_time_is_monotonically_increasing(self):
rms_time, _ = extract_wav_rms(self.wav)
self.assertTrue(np.all(np.diff(rms_time) > 0))
+ def test_rms_time_does_not_exceed_duration(self):
+ """All time points should fall within (or very close to) the file duration."""
+ rms_time, _ = extract_wav_rms(self.wav)
+ self.assertLessEqual(rms_time[-1], 2.1)
+
# --- silence masking ---
def test_silent_wav_masked_with_nan(self):
@@ -653,7 +660,7 @@ def test_mask_silence_false_values_nonnegative(self):
def test_tonal_wav_has_active_frames(self):
"""A sine wave should have some non-NaN (active) RMS frames."""
- wav = self._make_tonal_wav()
+ wav = _make_tonal_wav()
try:
_, rms = extract_wav_rms(wav, mask_silence=True)
self.assertTrue(np.any(~np.isnan(rms)))
@@ -661,7 +668,7 @@ def test_tonal_wav_has_active_frames(self):
os.unlink(wav)
def test_tonal_wav_active_frames_are_nonnegative(self):
- wav = self._make_tonal_wav()
+ wav = _make_tonal_wav()
try:
_, rms = extract_wav_rms(wav, mask_silence=True)
active = rms[~np.isnan(rms)]
@@ -669,6 +676,15 @@ def test_tonal_wav_active_frames_are_nonnegative(self):
finally:
os.unlink(wav)
+ def test_tonal_wav_no_nan_when_mask_false(self):
+ """A fully tonal signal with mask_silence=False must have zero NaN values."""
+ wav = _make_tonal_wav()
+ try:
+ _, rms = extract_wav_rms(wav, mask_silence=False)
+ self.assertFalse(np.any(np.isnan(rms)))
+ finally:
+ os.unlink(wav)
+
def test_leading_silence_masked(self):
"""Frames before the first active frame should be NaN when mask_silence=True."""
sr = 22050
@@ -745,15 +761,13 @@ def _patch_swift(self):
mock_cls = MagicMock(return_value=mock_detector)
return patch("swift_f0.SwiftF0", mock_cls, create=True)
- def _patch_rmvpe(self):
+ def _patch_rmvpe(self, times=None, freqs=None, confs=None):
"""Patch RMVPE and soundfile so the rmvpe-onnx backend runs without
- real model weights or audio I/O."""
- import types
-
- fake_timestamp = np.array(self._fake_times)
- fake_frequency = np.array(self._fake_freqs)
- fake_confidence = np.array(self._fake_confs)
- fake_activation = np.zeros(len(self._fake_times))
+ real model weights or audio I/O. Optionally override return values."""
+ fake_timestamp = np.array(times if times is not None else self._fake_times)
+ fake_frequency = np.array(freqs if freqs is not None else self._fake_freqs)
+ fake_confidence = np.array(confs if confs is not None else self._fake_confs)
+ fake_activation = np.zeros(len(fake_timestamp))
fake_rmvpe_instance = MagicMock()
fake_rmvpe_instance.predict.return_value = (
@@ -863,7 +877,6 @@ def test_crepe_backend_accepted(self):
fake_conf = np.random.uniform(0.5, 1.0, self._n)
fake_act = np.zeros((self._n, 360))
- import types
fake_crepe_mod = types.ModuleType("crepe")
fake_crepe_mod.predict = MagicMock(
return_value=(fake_time, fake_freq, fake_conf, fake_act)
@@ -878,6 +891,131 @@ def test_crepe_backend_accepted(self):
self.assertEqual(len(time), len(freq))
self.assertEqual(len(freq), len(conf))
+ # --- hybrid backend ---
+
+ def test_hybrid_backend_accepted(self):
+ """hybrid backend must be accepted without raising ValueError."""
+ with self._patch_rmvpe(), self._patch_swift():
+ # soundfile is also used inside _merge_rmvpe_and_swift_f0
+ extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+
+ def test_hybrid_returns_tuple_of_three(self):
+ with self._patch_rmvpe(), self._patch_swift():
+ result = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertIsInstance(result, tuple)
+ self.assertEqual(len(result), 3)
+
+ def test_hybrid_output_arrays_are_ndarrays(self):
+ with self._patch_rmvpe(), self._patch_swift():
+ time, freq, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertIsInstance(time, np.ndarray)
+ self.assertIsInstance(freq, np.ndarray)
+ self.assertIsInstance(conf, np.ndarray)
+
+ def test_hybrid_all_outputs_same_length(self):
+ with self._patch_rmvpe(), self._patch_swift():
+ time, freq, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertEqual(len(time), len(freq))
+ self.assertEqual(len(freq), len(conf))
+
+ def test_hybrid_output_length_matches_rmvpe_grid(self):
+ """Hybrid uses the rmvpe-onnx time grid, so output length == rmvpe output length."""
+ with self._patch_rmvpe(), self._patch_swift():
+ time, freq, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertEqual(len(time), self._n)
+
+ def test_hybrid_time_values_are_float(self):
+ with self._patch_rmvpe(), self._patch_swift():
+ time, _, _ = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertTrue(np.issubdtype(time.dtype, np.floating))
+
+ def test_hybrid_frequency_values_are_float(self):
+ with self._patch_rmvpe(), self._patch_swift():
+ _, freq, _ = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertTrue(np.issubdtype(freq.dtype, np.floating))
+
+ def test_hybrid_confidence_values_are_float(self):
+ with self._patch_rmvpe(), self._patch_swift():
+ _, _, conf = extract_wav_frequency(self.wav, backend="hybrid", use_cache=False)
+ self.assertTrue(np.issubdtype(conf.dtype, np.floating))
+
+ def test_hybrid_time_matches_rmvpe_time(self):
+ """The time axis returned by hybrid must equal the rmvpe-onnx time axis."""
+ with self._patch_rmvpe(), self._patch_swift():
+ time_hybrid, _, _ = extract_wav_frequency(
+ self.wav, backend="hybrid", use_cache=False
+ )
+ with self._patch_rmvpe():
+ time_rmvpe, _, _ = extract_wav_frequency(
+ self.wav, backend="rmvpe-onnx", use_cache=False
+ )
+ np.testing.assert_array_almost_equal(time_hybrid, time_rmvpe)
+
+ def test_hybrid_replaces_low_confidence_rmvpe_frames(self):
+ """When rmvpe confidence is low and swift-f0 confidence is high in a voiced
+ region, the hybrid output should include some swift-f0 frequency values."""
+ n = 50
+ times = list(np.linspace(0, 1.0, n))
+
+ # rmvpe: low confidence everywhere so swift-f0 replacements will be chosen
+ rmvpe_freqs = [200.0] * n
+ rmvpe_confs = [0.5] * n # below threshold 0.80
+
+ # swift-f0: high confidence + different frequency
+ swift_freqs = [400.0] * n
+ swift_confs = [0.99] * n # above threshold 0.95
+
+ # Build a tonal WAV so RMS > 0 (voiced region)
+ wav = _make_tonal_wav(duration=1.0, sr=22050, freq=440.0)
+ try:
+ with self._patch_rmvpe(times=times, freqs=rmvpe_freqs, confs=rmvpe_confs), \
+ self._patch_swift():
+ # Override swift-f0 mock with specific values
+ fake_result = MagicMock()
+ fake_result.timestamps.tolist.return_value = times
+ fake_result.pitch_hz.tolist.return_value = swift_freqs
+ fake_result.confidence.tolist.return_value = swift_confs
+ mock_detector = MagicMock()
+ mock_detector.detect_from_file.return_value = fake_result
+ mock_cls = MagicMock(return_value=mock_detector)
+ with patch("swift_f0.SwiftF0", mock_cls, create=True):
+ _, freq, _ = extract_wav_frequency(wav, backend="hybrid", use_cache=False)
+
+ # At least some frames should have been replaced with 400 Hz
+ self.assertTrue(np.any(np.isclose(freq, 400.0)),
+ "Expected some frames to be replaced by swift-f0 (400 Hz)")
+ finally:
+ os.unlink(wav)
+
+ def test_hybrid_keeps_rmvpe_frames_when_rmvpe_confident(self):
+ """When rmvpe confidence is high the hybrid output should retain rmvpe frequencies."""
+ n = 50
+ times = list(np.linspace(0, 1.0, n))
+ rmvpe_freqs = [200.0] * n
+ rmvpe_confs = [0.95] * n # above threshold 0.80 — should not be replaced
+
+ swift_freqs = [400.0] * n
+ swift_confs = [0.99] * n
+
+ wav = _make_tonal_wav(duration=1.0, sr=22050, freq=440.0)
+ try:
+ with self._patch_rmvpe(times=times, freqs=rmvpe_freqs, confs=rmvpe_confs):
+ fake_result = MagicMock()
+ fake_result.timestamps.tolist.return_value = times
+ fake_result.pitch_hz.tolist.return_value = swift_freqs
+ fake_result.confidence.tolist.return_value = swift_confs
+ mock_detector = MagicMock()
+ mock_detector.detect_from_file.return_value = fake_result
+ mock_cls = MagicMock(return_value=mock_detector)
+ with patch("swift_f0.SwiftF0", mock_cls, create=True):
+ _, freq, _ = extract_wav_frequency(wav, backend="hybrid", use_cache=False)
+
+ # All frames should stay at 200 Hz (rmvpe confident, no replacement)
+ self.assertTrue(np.all(np.isclose(freq, 200.0)),
+ "Expected rmvpe frequencies to be kept when confidence is high")
+ finally:
+ os.unlink(wav)
+
# --- caching ---
def test_cache_file_written_when_use_cache_true(self):
diff --git a/utils/seqtool.py b/utils/seqtool.py
index 8352c56..3c975a5 100644
--- a/utils/seqtool.py
+++ b/utils/seqtool.py
@@ -3,7 +3,7 @@
import numpy as np
from fastdtw import fastdtw # type: ignore
-from scipy.interpolate import interp1d
+from scipy.interpolate import interp1d, make_smoothing_spline
from scipy.ndimage import gaussian_filter1d
from scipy.stats import zscore
@@ -127,7 +127,7 @@ def unify_sequence_time(seq_times, seq_vals, to_ticks=False):
]
return unified_seq_time, tuple(unified_seqs_val)
- unified_seq_ticks = _time_to_ticks_fn(unified_seq_time)
+ unified_seq_ticks = np.unique(_time_to_ticks_fn(unified_seq_time))
time_mapping = _ticks_to_time_fn(unified_seq_ticks)
unified_seqs_val = [
interp1d(st, sv, fill_value="extrapolate")(time_mapping) # type: ignore
@@ -161,6 +161,60 @@ def gaussian_filter1d_with_nan(seq, sigma, **kwargs):
return seq
+def seq_spline_smoothing(seq_time, seq_val, lam=None, nan_policy='preserve_all'):
+ """Smooth a sequence using an adaptive smoothing spline.
+
+ Args:
+ seq_time (numpy.ndarray): Time values for the sequence.
+ seq_val (numpy.ndarray): Sequence values to smooth.
+ lam (float or None): Smoothing parameter. None selects automatically via GCV.
+ Higher values produce smoother results.
+ nan_policy (str): How to handle NaNs in the output. One of:
+ - 'no_nan': NaNs are excluded from fitting; output is fully predicted
+ with no NaNs.
+ - 'preserve_all': NaN positions are excluded from fitting and restored in
+ the output.
+ - 'preserve_head_tail': Only leading and trailing NaNs are restored; interior
+ NaNs are filled by the spline.
+
+ Returns:
+ numpy.ndarray: Smoothed sequence values, same length as seq_time.
+
+ Raises:
+ ValueError: If nan_policy is not recognized.
+
+ Example:
+ >>> seq_smoothing(time, val) # auto smoothness, preserve all NaNs
+ >>> seq_smoothing(time, val, lam=0.1) # manual smoothness
+ >>> seq_smoothing(time, val, nan_policy='no_nan') # fully predicted, no NaNs
+ >>> seq_smoothing(time, val, nan_policy='preserve_head_tail') # only boundary NaNs restored
+ """
+ nan_policies = {'no_nan', 'preserve_all', 'preserve_head_tail'}
+ if nan_policy not in nan_policies:
+ raise ValueError(f"Unknown nan_policy: {nan_policy!r}. Choose from: {nan_policies}.")
+
+ seq_val = np.asarray(seq_val, dtype=float)
+ seq_time = np.asarray(seq_time, dtype=float)
+ nan_mask = np.isnan(seq_val)
+
+ valid_time = seq_time[~nan_mask]
+ valid_val = seq_val[~nan_mask]
+
+ result = make_smoothing_spline(valid_time, valid_val, lam=lam)(seq_time)
+
+ if nan_policy == 'preserve_all':
+ result[nan_mask] = np.nan
+
+ elif nan_policy == 'preserve_head_tail':
+ first_valid = np.argmax(~nan_mask)
+ last_valid = len(nan_mask) - np.argmax(~nan_mask[::-1]) - 1
+ head_tail_mask = nan_mask.copy()
+ head_tail_mask[first_valid:last_valid + 1] = False
+ result[head_tail_mask] = np.nan
+
+ return result
+
+
def align_sequence_tick(
query_time, queries, reference_time, references, align_radius=1
):
diff --git a/utils/wavtool.py b/utils/wavtool.py
index a193b18..b778cf7 100644
--- a/utils/wavtool.py
+++ b/utils/wavtool.py
@@ -56,10 +56,11 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True):
Args:
file_path (str): Path to the WAV file.
- backend (str, optional): Pitch detection backend. One of "crepe" or "swift-f0" or "rmvpe-onnx".
+ backend (str, optional): Pitch detection backend.
"crepe" uses the CREPE model (requires TensorFlow, GPU-accelerated).
"swift-f0" uses SwiftF0 (faster CPU inference, requires swift-f0 package).
"rmvpe-onnx" uses RMVPE ONNX model (fast CPU inference, requires rmvpe-onnx package).
+ "hybrid" uses a hybrid strategy based on "rmvpe-onnx" and "swift-f0".
Defaults to "rmvpe-onnx".
use_cache (bool, optional): Whether to use cached data if available. Defaults to True.
@@ -69,7 +70,7 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True):
- frequency (np.ndarray of float): Detected pitch frequencies in Hz. Shape: (n_time_points).
- confidence (np.ndarray of float): Confidence values for the detected pitches. Shape: (n_time_points).
"""
- _SUPPORTED_BACKENDS = ("crepe", "swift-f0", "rmvpe-onnx")
+ _SUPPORTED_BACKENDS = ("crepe", "swift-f0", "rmvpe-onnx", "hybrid")
if backend not in _SUPPORTED_BACKENDS:
raise ValueError(f"Unknown backend '{backend}'. Choose from: {_SUPPORTED_BACKENDS}")
@@ -84,7 +85,7 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True):
cache_path = cache_dir / f"{wav_hash}.{backend}.csv"
if cache_path.is_file():
- print(_("Loading F0 data from cache file: '{}'").format(cache_path))
+ print(f"[{backend}] " + _("Loading F0 data from cache file: '{}'").format(cache_path))
with open(cache_path, "r", newline="") as file:
reader = csv.reader(file)
next(reader) # Skip header
@@ -119,6 +120,8 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True):
time = timestamp.tolist()
frequency = frequency.tolist()
confidence = confidence.tolist()
+ elif backend == "hybrid":
+ time, frequency, confidence = _merge_rmvpe_and_swift_f0(file_path, use_cache)
# Save data to cache
if use_cache:
@@ -127,11 +130,84 @@ def extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=True):
writer.writerow(["Time (s)", "Frequency (Hz)", "Confidence"])
for t, f, c in zip(time, frequency, confidence, strict=False):
writer.writerow([t, f, c])
- print(_("F0 data saved to cache file: '{}'").format(cache_path))
+ print(f"[{backend}] " + _("F0 data saved to cache file: '{}'").format(cache_path))
return np.asarray(time), np.asarray(frequency), np.asarray(confidence)
+def _merge_rmvpe_and_swift_f0(file_path, use_cache):
+ """Merge rmvpe-onnx and swift-f0 pitch predictions into a single output.
+
+ Uses rmvpe-onnx as the base prediction and selectively replaces frames with
+ swift-f0 results where all three conditions are met:
+ 1. The frame falls within a voiced region (RMS energy >= Otsu threshold).
+ 2. rmvpe-onnx confidence is low, indicating uncertain prediction.
+ 3. swift-f0 confidence is high, indicating a reliable prediction.
+
+ Voiced regions are detected by computing per-frame RMS energy with librosa,
+ then thresholding with Otsu's method to separate voiced from unvoiced frames.
+ swift-f0 frames are aligned to the rmvpe-onnx time grid via nearest-neighbour
+ lookup before comparison.
+
+ Args:
+ file_path (str): Path to the WAV file. Passed directly to
+ extract_wav_frequency for both backends.
+ use_cache (bool): Whether to use cached predictions. Passed directly to
+ extract_wav_frequency for both backends.
+
+ Returns:
+ tuple: (time, frequency, confidence), where:
+ - time (list of float): Time points in seconds from the rmvpe-onnx grid.
+ - frequency (list of float): Merged pitch frequencies in Hz.
+ - confidence (list of float): Confidence values corresponding to
+ whichever backend's frequency was selected per frame.
+ """
+ _CONFIDENCE_THRESHOLDS = {
+ "rmvpe-onnx": 0.80,
+ "swift-f0": 0.95,
+ }
+
+ from skimage.filters import threshold_otsu
+ from librosa.feature import rms as librosa_rms
+
+ r_time, r_freq, r_conf = extract_wav_frequency(file_path, backend="rmvpe-onnx", use_cache=use_cache)
+ s_time, s_freq, s_conf = extract_wav_frequency(file_path, backend="swift-f0", use_cache=use_cache)
+
+ # --- Base prediction: start from rmvpe-onnx ---
+ out_freq = r_freq.copy()
+ out_conf = r_conf.copy()
+
+ # --- Voiced region via Otsu threshold on RMS ---
+ import soundfile as sf
+ audio, sr = sf.read(file_path)
+ hop_length = 512
+ frame_rms = librosa_rms(y=audio, hop_length=hop_length)[0]
+ rms_times = np.arange(len(frame_rms)) * hop_length / sr
+ otsu_thr = threshold_otsu(frame_rms)
+ # Snap RMS voiced mask to rmvpe-onnx time grid
+ rms_indices = np.searchsorted(rms_times, r_time).clip(0, len(frame_rms) - 1)
+ voiced_region = frame_rms[rms_indices] >= otsu_thr
+
+ # --- Align swift-f0 to rmvpe-onnx time grid via nearest-neighbour lookup ---
+ snap_indices = np.searchsorted(s_time, r_time).clip(0, len(s_time) - 1)
+ prev_indices = np.maximum(snap_indices - 1, 0)
+ use_prev = np.abs(s_time[prev_indices] - r_time) < np.abs(s_time[snap_indices] - r_time)
+ snap_indices = np.where(use_prev, prev_indices, snap_indices)
+
+ s_freq_aligned = s_freq[snap_indices]
+ s_conf_aligned = s_conf[snap_indices]
+
+ # --- Replace: voiced region + rmvpe low confidence + swift-f0 high confidence ---
+ r_low = r_conf < _CONFIDENCE_THRESHOLDS["rmvpe-onnx"]
+ s_high = s_conf_aligned >= _CONFIDENCE_THRESHOLDS["swift-f0"]
+ replace_mask = voiced_region & r_low & s_high
+
+ out_freq = np.where(replace_mask, s_freq_aligned, out_freq)
+ out_conf = np.where(replace_mask, s_conf_aligned, out_conf)
+
+ return r_time.tolist(), out_freq.tolist(), out_conf.tolist()
+
+
def extract_wav_rms(wav_path, mask_silence=True):
"""Extract RMS energy from a WAV file.