An experimental real-time audio upscaling engine that converts low-resolution PCM streams (e.g., 8-bit low sample rate) to high-fidelity output (e.g., 16-bit double sample rate) using a single Next-Generation Reservoir Computing (NGRC) neuron.
The core architecture is sample-rate agnostic, operating directly on normalized discrete-time delay tap vectors. Designed primarily for low-latency retro-computing emulation pipelines (e.g., Amiga PAULA audio upsampling) where traditional deep neural networks introduce prohibitive latency or compute overhead.
Traditional Reservoir Computing (RC) maps inputs into a high-dimensional random dynamical system (the reservoir) and trains a simple linear readout layer.
Next-Generation Reservoir Computing (NGRC) replaces the complex, recurrent reservoir with a deterministic polynomial feature expansion constructed from time-delay tap vectors of the input signal:
By evaluating unique non-linear monomials (up to degree
- Sample-Rate & Resolution Agnostic: Processes arbitrary input/output sample rates and quantization levels.
- Ultra-Low Latency: Zero recurrent state dependencies; runs frame-by-frame or sample-by-sample via direct matrix-vector inner products.
- OpenMP Parallelized C++ Inference: Optimized C++ engine capable of processing 3+ million samples/sec on a single thread.
- Deterministic & Explainable: Fully auditable weight vectors and polynomial tap maps (no black-box hidden states).
-
Compact Binary Serialization: Custom
.binweight layout for fast$O(1)$ engine loading.
- High Throughput: Replaces heavy matrix multiplications with a single non-linear feature map lookup and scalar dot product.
-
Exact Closed-Form Training: Solved directly using Ridge Regression (
$\mathbf{W} = (\mathbf{X}^T\mathbf{X} + \gamma \mathbf{I})^{-1}\mathbf{X}^T\mathbf{Y}$ ); no backpropagation or vanishing gradient issues. - Phase Alignment: Symmetric non-causal context windows eliminate phase shift relative to the original signal.
- Superior to SINC Interpolation: Learns non-linear harmonic reconstruction rather than relying purely on static frequency-domain band limiting.
-
Combinatorial Feature Explosion: Feature dimensionality scales as
$O(N^d)$ with respect to tap window size$N$ and polynomial degree$d$ . - Sensitivity to Lossy Input Noise: High-degree monomial cross-terms can amplify lossy codec noise floor artifacts (e.g., Opus/AAC) if inputs are not properly dithered and anti-aliased prior to downquantization.
.
├── infer_ngrc.cpp # High-performance OpenMP C++ inference engine
├── train_ngrc.py # Parallel grid-search training script
├── inference.py # Python-based inference implementation
├── export_weights.py # Serializes weights and tap tables to custom C++ .bin format
├── inspector.py # CLI utility to analyze weight energy and prune dead terms
├── compare_baselines.py # Quantitative benchmark tool (SNR, RMSE vs SINC/linear)
├── phase_alignment_checker.py # Cross-correlation phase delay alignment validator
├── visualize.py # Waveform and spectral comparison visualizer
├── generate_audio.py # Dataset decimation and quantization pipeline
└── youtube_audio.py # Dataset acquisition script for training audio