Preserving multi-way continuum structures for sample-efficient, noise-resilient, and physically interpretable AI in data-scarce scientific discovery.
Scientific measurements across neuroscience, medical engineering, and complex physical dynamics are governed by multi-dimensional laws, yielding observation arrays with coupled spatial, temporal, spectral, trial-wise, and subject-specific modes. Standard machine learning models excel in data-abundant Euclidean domains; however, scientific applications operate under sample-constrained cohorts, non-stationary domain shifts, and physical sensor perturbations.
Supported by the MEXT AI for Science SPReAD Program, this project investigates structure-preserving robust tensor representation learning. By integrating multi-mode tensor parameterization, transform-domain spectral regularization ($t$-product algebra and dual spectral sparsity), and dispersion-adaptive robust optimization, we formulate a mathematical framework delivering tight generalization bounds, resilience to sensor outages, and explicit physical factor disentanglement.
Comparing the geometric assumptions, sample complexity, and perturbation robustness of different representation paradigms.
Treats high-order physical tensors as 1D vectors or 2D matrices. While effective for uncorrelated tabular records, flattening breaks coordinate-free geometric symmetries and inflates unconstrained parameter counts to $O(CTFd)$.
Universal approximation is powerful when training samples are plentiful. In sample-constrained scientific regimes ($N \ll D$), unconstrained hypothesis spaces yield loose generalization bounds and high vulnerability to missing channels.
Directly formulates representations over multilinear manifolds. Low-rank factorizations and transform spectral priors constrain the hypothesis space, maximizing sample efficiency and preserving predictive stability under channel dropouts.
Four interconnected mathematical pillars spanning multi-way continuous tensorization, transform-domain spectral regularization, robust optimization, and interpretable factor recovery.
Raw physical observations (e.g. intracranial EEG arrays) are modeled directly as higher-order tensors $\mathcal{X} \in \mathbb{R}^{C \times T \times F \times N \times S}$, corresponding to Electrode Channels $\times$ Time Points $\times$ Frequency Bands $\times$ Trials $\times$ Subjects. Continuous multi-dimensional physical symmetries are preserved without ad-hoc vectorization.
We generalize matrix singular value decompositions to tensors via the algebraic $t$-product framework and ring-parametrized dual spectral sparsity. By applying unitary or learned orthogonal transforms $L$ along the frequency/mode axis, we induce compact singular value profiles that capture coherent physical dynamics.
Scientific environments suffer from unpredictable noise, sporadic channel outages, and high cross-subject variability. Our framework integrates robust objective formulations (e.g., tube-dispersion batch normalization, Schatten-$p$ norms, and sample reweighting) to preserve stability under up to 50% missing channels.
Unlike black-box representations, tensor decompositions yield explicit factor matrices corresponding to spatial topography, temporal waveforms, and spectral energy envelopes. Learned tubal ranks provide an objective quantitative metric of intrinsic physical complexity.
Single-variable geometry ablation and empirical validation on standard benchmarks (Sleep-EDF and BNCI-2a Motor Imagery) under strict Leave-One-Subject-Out (LOSO) evaluation.
Intracranial EEG provides high temporal resolution and localized spatial information directly from brain tissue. However, it represents an extreme archetype of scientific data challenges:
Comprehensive stress-testing protocols to validate sample efficiency, noise resilience, and domain adaptability:
Structured phased execution under the MEXT SPReAD funding period, from initial tensor pipelines to community know-how sharing.
Establishing standard intracranial EEG and electrophysiological tensor representation pipelines. Implementing $t$-product algebra kernels and automated time-frequency decomposition.
Integrating dual spectral sparsity regularization and learned unitary transform layers. Deriving Rademacher generalization bounds under small-sample constraints.
Stress-testing on missing-channel scenarios ($10\% \to 50\%$ channel dropouts). Evaluating transferability across heterogeneous patient cohorts.
Consolidating empirical results, open-sourcing modular PyTorch/NumPy tensor templates, and presenting findings at academic seminars and MEXT SPReAD reporting forums.
Key theoretical tools and baseline protocols from prior investigations that provide mathematical foundations and empirical context for this SPReAD study.
This work analyzes the singular value distribution of multi-dimensional tensors under orthogonal and unitary transformations. It establishes theoretical bounds demonstrating that appropriate transform domains induce compact spectral representations for structured signals. In the SPReAD project, these findings provide a theoretical reference for formulating the dual spectral sparsity regularization.
This study examines the geometric and algebraic properties of $t$-product tensor manifolds over commutative rings. It characterizes how ring-based multilinear operations influence representation capacity and manifold geometry. These geometric insights serve as an algebraic foundation for the multi-mode parameterization explored in this project.
This paper investigates low-rank transition operators designed to model shifts across tensor coefficient spaces. It provides an efficient framework for adapting structured models to varying data distributions across experimental environments. The method offers algorithmic references for addressing cross-subject variability in neural electrophysiology decoding.
This work investigates generalization error bounds for tensor-structured neural network architectures. It proves that applying mode-wise transformations prior to low-rank factorization tightens Rademacher complexity bounds. These theoretical bounds provide preliminary analytical tools for assessing sample efficiency in scientific data modeling.
This study explores sample-reweighted learning for identifying epileptic seizure onset zones from clinical intracranial EEG recordings. It examines practical experimental challenges arising from electrode noise and limited clinical sample cohorts. The findings provide empirical context and baseline protocols for the neural signal decoding component of this project.
Disseminating practical know-how, open pipelines, and empirical principles to empower researchers across broader scientific disciplines.
Clean PyTorch and NumPy script templates for direct multi-mode tensorization of continuous time-series, with automated time-frequency transformations.
Clear guidelines on choosing between standard vectorization, Tucker/CP models, and $t$-product ring algebras for different physical data distributions.
Translating neural engineering and iEEG validation protocols into reusable methodologies for other scientific domains such as climate dynamics and biophysics.
If you utilize concepts, formulations, or empirical benchmarks from the SPReAD project in your research, please cite:
@misc{wang2026spread,
title = {Robust Tensor Representation Learning for High-Dimensional Scientific Data},
author = {Wang, Andong},
howpublished = {MEXT AI for Science Program: SPReAD Initiative (Grant No. 26279343)},
institution = {RIKEN Center for Advanced Intelligence Project (RIKEN AIP)},
year = {2026},
url = {https://pingzaiwang.github.io/homepage/spread.html}
}
Correspondence: Andong Wang (w.a.d@outlook.com) · Personal Homepage ↗