diff --git a/CHANGELOG.md b/CHANGELOG.md index 26f06cb3..c2d43be8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -26,6 +26,7 @@ Please view [SomeB1oody/RustyML](https://github.com/SomeB1oody/RustyML) for more - **Behavior change: a training step refuses a sublayer call that reaches another layer than the tree holds at its path.** A training pass records the layer of each `Ctx::sublayer` call. The model compares each record against the tree after the forward pass and after the backward pass. A call of sublayer `b` under the name of sublayer `a` therefore stops the step, also when `a` and `b` have the same type and shape. - **Behavior change: the model build checks the whole layer tree.** It refuses 2 arrays or 2 sublayers of 1 layer with 1 name. It also refuses an empty name, a name that holds a `.`, 2 sublayer rosters that disagree, 1 layer at 2 nodes of 1 tree, and a path more than 64 sublayers deep. A roster that lists its own layer therefore gives an error, not a stack overflow. The message names the layer path. - **Fix: `Sequential::weight` and `Graph::weight` accept 1 spelling per model position.** A path such as `00.kernel` or `+0.kernel` reaches no array. +- **Behavior change: `DBSCAN::predict` returns `Error::EmptyInput` for a matrix with 0 rows.** Before, it returned an empty array, and it skipped the feature-count check for that matrix. Every other model already returned `EmptyInput`. - **Fix: `DecisionTree::predict_one`, `DecisionTree::predict_proba_one`, and `IsolationForest::score_sample` return `Error::NonFinite` for a NaN or an infinite value.** Before, these single-sample methods checked only the length. A NaN then took the right branch at each split, and the method returned a normal-looking result. The matrix methods already refused such a value. - **Fix: `GroupNormalization` and `InstanceNormalization` refuse the same ranks in `compute_output_shape`, in the build, and in the forward pass.** Before, the build accepted rank 1 and rank 2, and the forward pass refused both, so a model built and then failed on its first batch. `GroupNormalization` now accepts rank 2 and higher. At rank 2 each group folds over its own channels alone. `InstanceNormalization` accepts rank 3 and higher, because at rank 2 its output would always equal `beta`. diff --git a/guide/en/src/Chapter-02/2.8._DBSCAN.md b/guide/en/src/Chapter-02/2.8._DBSCAN.md index d40d4786..6c0de2d4 100644 --- a/guide/en/src/Chapter-02/2.8._DBSCAN.md +++ b/guide/en/src/Chapter-02/2.8._DBSCAN.md @@ -181,7 +181,7 @@ fn main() { `predict` finds each query's nearest core point. It searches the core points saved during `fit`, not the full training set. It returns that core point's cluster label if the query is within `eps`. Otherwise, it returns `-1`. `predict` picks the single nearest core point. It does not check every core point within `eps`. The `eps` gate is inclusive. `predict` never creates a new cluster, never promotes a query to a core point, and never runs the density flood again. So `predict(x)` does *not* give the same result as adding `x` to the training data and calling `fit` again. Treat `predict` as a fast, approximate way to assign held-out points to clusters `fit` already found. Call `fit` again on the enlarged set to get true DBSCAN semantics. -Calling `predict` before `fit` gives `Error::NotFitted`. A feature-count mismatch gives `Error::DimensionMismatch`. A non-finite query value gives `Error::NonFinite`. Empty input returns an empty array. +Calling `predict` before `fit` gives `Error::NotFitted`. A feature-count mismatch gives `Error::DimensionMismatch`. A non-finite query value gives `Error::NonFinite`. A zero-row matrix gives `Error::EmptyInput`, the same as for `fit` and for every other model. ## 2.8.7. Two rings where KMeans fails diff --git a/guide/zh-Hans/src/Chapter-02/2.8._DBSCAN.md b/guide/zh-Hans/src/Chapter-02/2.8._DBSCAN.md index 82a2babe..d4479e05 100644 --- a/guide/zh-Hans/src/Chapter-02/2.8._DBSCAN.md +++ b/guide/zh-Hans/src/Chapter-02/2.8._DBSCAN.md @@ -181,7 +181,7 @@ fn main() { `predict` 会为每个查询点找到它最近的核心点。它只在 `fit` 期间存下的核心点里搜索,而不是整个训练集。如果查询点落在 `eps` 之内,就返回那个核心点的簇标签,否则返回 `-1`。`predict` 只挑最近的那一个核心点,不会去检查 `eps` 内的每一个核心点。`eps` 这道闸取闭区间。`predict` 从不创建新簇,从不把查询点提拔为核心点,也从不重新跑一遍密度洪泛。所以 `predict(x)` 的结果*并不等于*把 `x` 加进训练数据后再调用一次 `fit`。把 `predict` 当成一种快速、近似的手段,用来把留出的点分派到 `fit` 已经找到的那些簇里。如果想要真正的 DBSCAN 语义,就在扩大后的数据集上重新调用 `fit`。 -在 `fit` 之前调用 `predict` 会得到 `Error::NotFitted`。特征数不匹配会得到 `Error::DimensionMismatch`。查询里出现非有限值会得到 `Error::NonFinite`。空输入则返回空数组。 +在 `fit` 之前调用 `predict` 会得到 `Error::NotFitted`。特征数不匹配会得到 `Error::DimensionMismatch`。查询里出现非有限值会得到 `Error::NonFinite`。零行矩阵会得到 `Error::EmptyInput`,与 `fit` 和其他所有模型一致。 ## 2.8.7. KMeans 失手的双环 diff --git a/src/machine_learning/clustering/dbscan.rs b/src/machine_learning/clustering/dbscan.rs index c6f9a6b6..65d61082 100644 --- a/src/machine_learning/clustering/dbscan.rs +++ b/src/machine_learning/clustering/dbscan.rs @@ -379,6 +379,7 @@ impl DBSCAN { /// # Errors /// /// - `Error::NotFitted` - If the model has not been fitted yet + /// - `Error::EmptyInput` - If `new_data` holds no element /// - `Error::DimensionMismatch` - If feature dimensions do not match /// - `Error::NonFinite` - If the data contains non-finite values /// @@ -398,12 +399,7 @@ impl DBSCAN { .as_ref() .ok_or_else(|| Error::not_fitted("DBSCAN"))?; - // Empty input yields an empty result - if new_data.nrows() == 0 { - return Ok(Array1::from(vec![])); - } - - // Validate feature dimensions and finiteness against the fitted model + // Validate the row count, the feature dimensions, and finiteness against the fitted model validate_predict_input(new_data, core_points.ncols())?; // Assign each point to the cluster of its nearest core point within `eps` diff --git a/tests/machine_learning/dbscan.rs b/tests/machine_learning/dbscan.rs index 50d6b70e..87016ff7 100644 --- a/tests/machine_learning/dbscan.rs +++ b/tests/machine_learning/dbscan.rs @@ -90,16 +90,23 @@ fn constructor_default_values() { // predict() after fit: empty input and error paths -/// predict on empty new_data returns Ok with an empty array +/// predict refuses a 0-row matrix, also when its feature count differs from the model +/// +/// A 0-row matrix of 3 features against a model of 2 features must not return `Ok`. The +/// row check and the feature check run before any work, so neither one can be skipped #[test] -fn predict_empty_new_data_returns_empty_array() { - let train = two_blobs_noise(); +fn predict_on_a_0_row_matrix_returns_empty_input() { + let train = two_blobs_noise(); // 2 features let mut m = DBSCAN::new(0.5, 2).unwrap(); m.fit(&train).unwrap(); - let empty: Array2 = Array2::zeros((0, 2)); - let preds = m.predict(&empty).expect("expected Ok for empty new_data"); - assert_eq!(preds.len(), 0); + for features in [2, 3] { + let empty: Array2 = Array2::zeros((0, features)); + assert!( + matches!(m.predict(&empty), Err(Error::EmptyInput(_))), + "expected EmptyInput for a 0x{features} matrix" + ); + } } /// predict with wrong number of features returns DimensionMismatch diff --git a/tests/machine_learning/ml_infra.rs b/tests/machine_learning/ml_infra.rs index edb886e6..6ac03e4a 100644 --- a/tests/machine_learning/ml_infra.rs +++ b/tests/machine_learning/ml_infra.rs @@ -484,20 +484,14 @@ fn fit_on_non_finite_input_returns_non_finite() { } /// After fit, each method returns `EmptyInput` for a matrix with 0 rows. -/// -/// `DBSCAN::predict` is the exception: it returns an empty array. #[test] -fn fitted_methods_on_empty_input() { +fn fitted_methods_on_empty_input_return_empty_input() { let models = Models::fitted(); let x: Array2 = Array2::zeros((0, 2)); for (model, method, call) in matrix_method_calls() { - let expected = match (model, method) { - ("DBSCAN", "predict") => Outcome::Rows(0), - _ => Outcome::EmptyInput, - }; assert_eq!( Outcome::from(call(&models, &x)), - expected, + Outcome::EmptyInput, "{model}::{method} on a 0-row matrix" ); }