MEXT AI for Science Program · SPReAD Initiative (Grant No. 26279343) Official Portal ↗

Robust Tensor Representation Learning for High-Dimensional Scientific Data

Preserving multi-way continuum structures for sample-efficient, noise-resilient, and physically interpretable AI in data-scarce scientific discovery.

Principal Investigator (PI): WANG Andong, Ph.D.
Research Scientist, Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project (RIKEN AIP)

Executive Summary & Scientific Scope

Scientific measurements across neuroscience, medical engineering, and complex physical dynamics are governed by multi-dimensional laws, yielding observation arrays with coupled spatial, temporal, spectral, trial-wise, and subject-specific modes. Standard machine learning models excel in data-abundant Euclidean domains; however, scientific applications operate under sample-constrained cohorts, non-stationary domain shifts, and physical sensor perturbations.

Supported by the MEXT AI for Science SPReAD Program, this project investigates structure-preserving robust tensor representation learning. By integrating multi-mode tensor parameterization, transform-domain spectral regularization ($t$-product algebra and dual spectral sparsity), and dispersion-adaptive robust optimization, we formulate a mathematical framework delivering tight generalization bounds, resilience to sensor outages, and explicit physical factor disentanglement.

Inductive Biases and Representation Structure in Scientific AI

Comparing the geometric assumptions, sample complexity, and perturbation robustness of different representation paradigms.

Vectorization & Unfolding

Euclidean Flattening

Treats high-order physical tensors as 1D vectors or 2D matrices. While effective for uncorrelated tabular records, flattening breaks coordinate-free geometric symmetries and inflates unconstrained parameter counts to $O(CTFd)$.

Generic Deep Models

Unconstrained Capacity

Universal approximation is powerful when training samples are plentiful. In sample-constrained scientific regimes ($N \ll D$), unconstrained hypothesis spaces yield loose generalization bounds and high vulnerability to missing channels.

Structure-Preserving Tensors

Algebraic & Spectral Inductive Biases

Directly formulates representations over multilinear manifolds. Low-rank factorizations and transform spectral priors constrain the hypothesis space, maximizing sample efficiency and preserving predictive stability under channel dropouts.

Training-Free Premise Probes on BNCI Motor Imagery and Sleep-EDF
Click to expand
Figure 1 (Training-Free Premise Probes): (a) Coupling Probe: Amplitude-phase coupling (AAC/PAC) is at chance on BNCI Motor Imagery (25.9% vs. 25% chance) but strongly discriminative on Sleep-EDF (65.5% / 53.9% vs. 20% chance). (b) Dispersion Probe: Per-band tubal dispersion accurately predicts a +18.6 pp advantage on Sleep ($p < 0.01$) versus +0.0 pp on Motor Imagery without heavy neural network training.

The SPReAD Tensor Representation Framework

Four interconnected mathematical pillars spanning multi-way continuous tensorization, transform-domain spectral regularization, robust optimization, and interpretable factor recovery.

SPReAD Project Overall Framework Architecture
Click to expand full resolution
Figure 2 (SPReAD Technical Architecture): Four sequential stages: (1) Multi-mode tensorization; (2) Transform-domain spectral regularization with $t$-SVD & dual sparsity; (3) Robust optimization across noise, missing channels, and subject domain shifts; (4) Physical factor decomposition and clinical/scientific insight extraction.
01

Multi-way Scientific Tensorization

Raw physical observations (e.g. intracranial EEG arrays) are modeled directly as higher-order tensors $\mathcal{X} \in \mathbb{R}^{C \times T \times F \times N \times S}$, corresponding to Electrode Channels $\times$ Time Points $\times$ Frequency Bands $\times$ Trials $\times$ Subjects. Continuous multi-dimensional physical symmetries are preserved without ad-hoc vectorization.

02

Transform-Domain Spectral Regularization

We generalize matrix singular value decompositions to tensors via the algebraic $t$-product framework and ring-parametrized dual spectral sparsity. By applying unitary or learned orthogonal transforms $L$ along the frequency/mode axis, we induce compact singular value profiles that capture coherent physical dynamics.

03

Robust Learning under Severe Perturbations

Scientific environments suffer from unpredictable noise, sporadic channel outages, and high cross-subject variability. Our framework integrates robust objective formulations (e.g., tube-dispersion batch normalization, Schatten-$p$ norms, and sample reweighting) to preserve stability under up to 50% missing channels.

04

Interpretable Factor & Rank Profiling

Unlike black-box representations, tensor decompositions yield explicit factor matrices corresponding to spatial topography, temporal waveforms, and spectral energy envelopes. Learned tubal ranks provide an objective quantitative metric of intrinsic physical complexity.

Proof of Concept: Empirical Neural Signal Validation

Single-variable geometry ablation and empirical validation on standard benchmarks (Sleep-EDF and BNCI-2a Motor Imagery) under strict Leave-One-Subject-Out (LOSO) evaluation.

Single-Variable Geometry & Mechanism Decomposition on BNCI and Sleep-EDF
Click to expand
Figure 3 (Geometry & Mechanism Single-Variable Decomposition): (a) BNCI2014_001 Motor Imagery ($n=9$ domains): Confirms an honest null result (Product $58.4\%$ vs $+L$ $57.0\%$ vs Tubal $56.3\%$, all $p > 0.65$ n.s.). (b) Sleep-EDF ($n=12$ subjects, LOSO): Our Tubal geometry achieves $78.2 \pm 5.5\%$ vs Product baseline $48.0 \pm 13.1\%$ ($+30.3\text{ pp}$, paired $p = 0.001$), driven primarily by Tube Dispersion alone ($+24.3\text{ pp}$, $p = 0.013$) rather than coupling alone ($+6.0\text{ pp}$, n.s.). Individual per-subject points shown with jitter.

Why iEEG is the Ideal Benchmark for AI4Science

Intracranial EEG provides high temporal resolution and localized spatial information directly from brain tissue. However, it represents an extreme archetype of scientific data challenges:

  • Extreme Sample Scarcity: Clinical annotations for Seizure Onset Zones (SOZ) require scarce expert epileptologist reviews.
  • Complex Multi-mode Geometry: Strong coupling across spatial electrode grids, rhythmic oscillation bands ($\alpha, \beta, \gamma$), and transient burst times.
  • Cross-Subject Distribution Shifts: Electrode implantation geometries vary substantially across patients.

Key Evaluation Metrics in the 180-Day PoC

Comprehensive stress-testing protocols to validate sample efficiency, noise resilience, and domain adaptability:

  • Discriminative Power: Classification balanced accuracy, AUC, and F1 scores for SOZ localization.
  • Channel Dropout Curve: Measuring performance stability under random missing channels ($0\% \to 50\%$).
  • Generalization: Leave-One-Subject-Out (LOSO) cross-patient validation.
  • Interpretability: Aligning extracted spatial-spectral tensor factors with post-surgical clinical findings.

180-Day Implementation & Milestones

Structured phased execution under the MEXT SPReAD funding period, from initial tensor pipelines to community know-how sharing.

Days 1 – 45

Phase 1: Multi-way Continuous Tensorization & $t$-SVD Pipelines

Establishing standard intracranial EEG and electrophysiological tensor representation pipelines. Implementing $t$-product algebra kernels and automated time-frequency decomposition.

Days 46 – 105

Phase 2: Transform-Domain Spectral Regularization & Dual Sparsity

Integrating dual spectral sparsity regularization and learned unitary transform layers. Deriving Rademacher generalization bounds under small-sample constraints.

Days 106 – 150

Phase 3: Robust Optimization & Cross-Subject Generalization

Stress-testing on missing-channel scenarios ($10\% \to 50\%$ channel dropouts). Evaluating transferability across heterogeneous patient cohorts.

Days 151 – 180

Phase 4: Synthesis, Clinical Validation & Community Deliverables

Consolidating empirical results, open-sourcing modular PyTorch/NumPy tensor templates, and presenting findings at academic seminars and MEXT SPReAD reporting forums.

Technical Groundwork & Prior References

Key theoretical tools and baseline protocols from prior investigations that provide mathematical foundations and empirical context for this SPReAD study.

Technical Reference · Spectral Regularization
Refining Dual Spectral Sparsity in Transformed Tensor Singular Values
Andong Wang, Yuning Qiu, Haonan Huang, Zhong Jin, Guoxu Zhou, Qibin Zhao
International Conference on Machine Learning (ICML), 2026.

This work analyzes the singular value distribution of multi-dimensional tensors under orthogonal and unitary transformations. It establishes theoretical bounds demonstrating that appropriate transform domains induce compact spectral representations for structured signals. In the SPReAD project, these findings provide a theoretical reference for formulating the dual spectral sparsity regularization.

Technical Reference · t-Product Algebra
Towards a Geometric Understanding of Tensor Learning via the t-Product
Andong Wang, Yuning Qiu, Haonan Huang, Zhong Jin, Guoxu Zhou, Qibin Zhao
Advances in Neural Information Processing Systems (NeurIPS), 2025.

This study examines the geometric and algebraic properties of $t$-product tensor manifolds over commutative rings. It characterizes how ring-based multilinear operations influence representation capacity and manifold geometry. These geometric insights serve as an algebraic foundation for the multi-mode parameterization explored in this project.

Technical Reference · Subspace Transitions
Low-Rank Tensor Transitions (LoRT) for Transferable Tensor Regression
Andong Wang, Yuning Qiu, Zhong Jin, Guoxu Zhou, Qibin Zhao
International Conference on Machine Learning (ICML), 2025.

This paper investigates low-rank transition operators designed to model shifts across tensor coefficient spaces. It provides an efficient framework for adapting structured models to varying data distributions across experimental environments. The method offers algorithmic references for addressing cross-subject variability in neural electrophysiology decoding.

Technical Reference · Generalization Bounds
Transformed Low-Rank Parameterization Can Help Robust Generalization for Tensor Neural Networks
Andong Wang, Chao Li, Meng Bai, Zhong Jin, Guoxu Zhou, Qibin Zhao
Advances in Neural Information Processing Systems (NeurIPS), 2023.

This work investigates generalization error bounds for tensor-structured neural network architectures. It proves that applying mode-wise transformations prior to low-rank factorization tightens Rademacher complexity bounds. These theoretical bounds provide preliminary analytical tools for assessing sample efficiency in scientific data modeling.

Technical Reference · Electrophysiology Baseline
Classification of Epileptic Seizure Onset Zone from iEEG by Reweighting Augmented Samples
Xuyang Zhao, Qibin Zhao, Andong Wang, Hidenori Sugano, Toshihisa Tanaka
Biomedical Signal Processing and Control, Vol. 120, 2026.

This study explores sample-reweighted learning for identifying epileptic seizure onset zones from clinical intracranial EEG recordings. It examines practical experimental challenges arising from electrode noise and limited clinical sample cohorts. The findings provide empirical context and baseline protocols for the neural signal decoding component of this project.

SPReAD AI for Science Community Deliverables

Disseminating practical know-how, open pipelines, and empirical principles to empower researchers across broader scientific disciplines.

Reusable Code Templates

Clean PyTorch and NumPy script templates for direct multi-mode tensorization of continuous time-series, with automated time-frequency transformations.

Practical Method Guidelines

Clear guidelines on choosing between standard vectorization, Tucker/CP models, and $t$-product ring algebras for different physical data distributions.

Cross-Disciplinary Workflows

Translating neural engineering and iEEG validation protocols into reusable methodologies for other scientific domains such as climate dynamics and biophysics.

Citation & Correspondence

If you utilize concepts, formulations, or empirical benchmarks from the SPReAD project in your research, please cite:

BibTeX Citation
@misc{wang2026spread,
  title        = {Robust Tensor Representation Learning for High-Dimensional Scientific Data},
  author       = {Wang, Andong},
  howpublished = {MEXT AI for Science Program: SPReAD Initiative (Grant No. 26279343)},
  institution  = {RIKEN Center for Advanced Intelligence Project (RIKEN AIP)},
  year         = {2026},
  url          = {https://pingzaiwang.github.io/homepage/spread.html}
}

Correspondence: Andong Wang (w.a.d@outlook.com) · Personal Homepage ↗

Copied to clipboard!