iS iStructTab ICPR 2026 · Image + Tabular
Published ICPR 2026 Multimodal Learning Image + Tabular Feature Sequencing GEDS OEMT

iStructTab

iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data formulates multimodal feature ordering as a Column Permutation Problem (CPP) . It introduces Graph-Enhanced Descriptor Sequencing (GEDS) to derive a structured global feature order and an Order-Aware Efficient Transformer with Memory Augmentation (OEMT) to preserve and supervise that ordering during downstream learning.

Al Zadid Sultan Bin Habib1, Md Younus Ahamed1, Prashnna Kumar Gyawali1, Gianfranco Doretto2, Donald A. Adjeroh1

1 West Virginia University    2 University of Utah

ICPR 2026 logo

International Conference on Pattern Recognition · ICPR 2026

Overview

Image-tabular multimodal learning combines visual representations with structured metadata, but the resulting fused feature vector is commonly treated as an arbitrary collection of coordinates. iStructTab asks whether the fused representation itself should first be explicitly structured.

We formulate multimodal feature sequencing as a Column Permutation Problem (CPP) and seek a feature ordering that reduces dispersion and encourages a more coherent multimodal representation. iStructTab first encodes image and tabular inputs, concatenates their representations, derives a global feature sequence using GEDS, and processes the reordered representation using OEMT.

CPP Formulation Multimodal feature ordering is treated as an explicit combinatorial structure-learning problem.
GEDS Mean-variance feature descriptors, cosine similarity, graph refinement, and scalar ranking produce a global feature permutation.
OEMT Order-aware pooling, memory augmentation, Linformer encoding, and sequencing supervision preserve the structured representation.
1.50 ± 0.76 Average Rank
2.21 ± 4.59 Average Regret
99.29% DVM Accuracy with Tuning

Architecture

Overview of iStructTab with image and tabular encoders, GEDS, and OEMT
Overview of iStructTab. Modality-specific image and tabular representations are concatenated into a unified feature matrix. GEDS computes statistical descriptors and graph-based relationships to derive the global feature permutation \(\pi_{\mathrm{GEDS}}\). The reordered representation is subsequently processed by OEMT.

Method at a Glance

1. Encode Modalities Image encoder + Transformer-based tabular representation.
2. Fuse Features Concatenate image and tabular embeddings into one unified feature matrix.
3. GEDS Descriptor graph → graph refinement → feature scores → argsort.
4. OEMT Order-aware pooling + memory tokens + Linformer + sequencing loss.
\[ \min_{\pi \in S_m} D(\pi) = \sum_{i \lt j} w_{ij} \left| \pi(i) - \pi(j) \right| \]

The objective defines a feature permutation \(\pi\) over the fused multimodal representation. In iStructTab, GEDS provides the practical sequencing procedure.

Graph-Enhanced Descriptor Sequencing (GEDS)

GEDS operates on the fused image-tabular representation . For each feature column \(f_j\), we compute a low-dimensional statistical descriptor based on its mean and variance.

\[ d_j = \begin{bmatrix} \mu_j \\ \sigma_j^2 \end{bmatrix} \in \mathbb{R}^{2} \]

The descriptors are normalized and compared using cosine similarity to construct a feature-descriptor graph. Graph convolution subsequently refines those descriptors.

\[ A = DD^\top \] \[ \hat{A} = D_{\mathrm{deg}}^{-1/2} (A+I) D_{\mathrm{deg}}^{-1/2} \] \[ D' = \sigma( \hat{A}DW ) \]

Each graph-refined descriptor is assigned a scalar score. Sorting these scores produces the final GEDS feature permutation.

\[ s_j = \left\| D'_{j:} \right\|_2 \] \[ \pi_{\mathrm{GEDS}} = \operatorname{argsort} ( s_1, s_2, \ldots, s_m ) \]
Statistical Structure Each fused feature is summarized using mean and variance.
Relational Structure Cosine similarity creates a graph over the feature descriptors.
Global Ordering Graph-refined descriptor norms produce one sequence over the fused representation.
iStructTab sequencing: Feature statistics → descriptor similarity graph → graph convolution → feature scores → argsort → global feature sequence.

Order-Aware Efficient Transformer with Memory Augmentation (OEMT)

OEMT processes the structured feature representation produced by GEDS. The ordered feature sequence is first transformed into token embeddings and compressed from \(m\) original ordered features to \(k \ll m\) summary tokens through order-aware pooling .

\[ \mathbf{P} = \operatorname{softmax} \left( \mathbf{W}_{\mathrm{pool}} \mathbf{T}_{\mathrm{in}}^\top \right) \] \[ \widetilde{\mathbf{T}}_{\mathrm{in}} = \mathbf{P} \mathbf{T}_{\mathrm{in}} \]

Learnable memory tokens are appended to the pooled ordered representation and processed using a Linformer-based encoder.

\[ \mathbf{T} = [ \widetilde{\mathbf{T}}_{\mathrm{in}}; \mathbf{T}_{\mathrm{mem}} ] \] \[ \mathbf{T}' = \operatorname{Encoder} ( \mathbf{T} ) \]

Classification uses the memory-token representation, while an additional sequencing head predicts the target feature order. Training therefore jointly optimizes classification and feature sequencing.

\[ \mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{CE}} ( y, \hat{y} ) + \lambda_{\mathrm{FS}} \left\| \mathbf{s}' - \boldsymbol{\beta} \right\|_2^2 \]

Results Across Six Image-Tabular Benchmarks

iStructTab is evaluated across six multimodal image-tabular datasets: DVM, HAM10000, Deep Lesion, Pokémon, CheXpert, and PetFinder.

Across the complete benchmark suite, iStructTab obtains an overall average rank of 1.50 ± 0.76 and average regret of 2.21 ± 4.59.

Model DVM HAM DLes Pok CheX Pet Avg. Rank ↓ Avg. Regret ↓
TIP 98.27* 70.39 69.17 63.40∗ 87.50 83.86 5.83 ± 4.21 6.86 ± 6.01
STiL 99.27 78.48 81.35 27.32 88.60 87.68 4.33 ± 2.92 8.51 ± 14.72
iStructTab 99.29* 85.23 83.75 68.29 76.15 87.75 1.50 ± 0.76 2.21 ± 4.59

Accuracy (%) on the image-tabular benchmarks. See the paper for the complete comparison including all image-only, tabular-only, and multimodal baselines. * denotes results taken from another paper.

Robustness and Inference Diagnostics

Beyond predictive performance, the paper examines robustness under noisy labels and inference-level perturbations. On HAM10000, iStructTab maintains stable performance as label corruption increases.

Method 10% 20% 40% 60% Mean
CUFIT 82.6 81.5 79.1 70.1 78.3
iStructTab 83.92 81.92 79.42 76.42 80.42
iStructTab HAM10000 reliability diagram
Calibration and reliability.
iStructTab HAM10000 nearest-neighbor agreement
Fused-space kNN agreement.
iStructTab HAM10000 confidence analysis
Prediction confidence.
iStructTab image perturbation robustness
Image perturbation robustness.
iStructTab tabular perturbation robustness
Tabular perturbation robustness.
iStructTab HAM10000 inference efficiency
Inference diagnostics.

Computational Efficiency

For a fused feature matrix \(F \in \mathbb{R}^{N \times m}\), GEDS computes feature statistics in \(\mathcal{O}(Nm)\) and the descriptor affinity graph in \(\mathcal{O}(m^2)\).

OEMT reduces the ordered feature sequence before applying a Linformer encoder, avoiding full quadratic self-attention over the original fused sequence.

In the DVM efficiency analysis, iStructTab reaches the low-complexity / high-accuracy Pareto region and achieves the strongest reported normalized efficiency score.

iStructTab DVM weighted efficiency comparison
Weighted efficiency comparison.
iStructTab accuracy efficiency Pareto frontier
Accuracy-efficiency Pareto frontier.
iStructTab performance versus complexity
Accuracy versus composite complexity.
iStructTab normalized efficiency profile
Normalized efficiency profile.

Component Ablation

The DVM ablation isolates the roles of GEDS, memory augmentation, and sequencing supervision. The strongest degradation occurs when explicit feature sequencing is removed, indicating that the gains cannot simply be attributed to the image backbone.

Variant Without Tuning With Tuning
Concat Fuse, no sequence / memory 90.42 89.60
iStructTab, ResNet-50 + GEDS 96.85 99.29
iStructTab, ResNeXt-50 + GEDS 91.88 93.33
iStructTab without memory tokens 92.84 87.60
iStructTab without sequencing loss 96.85 96.78
iStructTab without feature sequencing 80.80 82.98
No sequence with memory loss 13.91 16.86

From iStructTab to iSyncTab

iStructTab was our first approach to explicit feature sequencing for image-tabular multimodal learning. It demonstrated that the organization of the fused multimodal feature representation itself can be an important inductive bias.

Our later iSyncTab work, published at ECCV 2026 , advances this research direction by making explicit cross-modal alignment part of the sequencing algorithm itself.

How are iStructTab and iSyncTab different?

iStructTab · ICPR 2026

  • Core sequencing: Graph-Enhanced Descriptor Sequencing (GEDS).
  • Modality handling: image and tabular representations are concatenated before sequencing.
  • Feature descriptors: principally mean and variance.
  • Structural mechanism: descriptor cosine similarity graph followed by graph convolution.
  • Ordering: graph-refined descriptor norms are converted into scalar feature scores and globally sorted.
  • Output: one global feature sequence over the fused representation.
  • Downstream architecture: OEMT uses order-aware pooling, memory tokens, Linformer encoding, and memory-derived sequencing supervision.
Pipeline
GEDS → statistical descriptors → similarity graph → graph convolution → scalar ranking → global permutation → OEMT

iSyncTab · ECCV 2026

  • Core sequencing: Neural Synchrony-guided Paired Feature Sequencing (NS-PFS).
  • Modality handling: visual and tabular feature streams remain explicitly separated during sequencing.
  • Structure discovery: visual and tabular features are clustered separately.
  • Cross-modal alignment: energy coherence and centroid / phase-analogue similarity measure neural synchrony.
  • Matching: the Hungarian algorithm performs optimal cross-modal bipartite cluster assignment.
  • Output: matched visual-tabular clusters produce synchronized local sequences that are combined into a global permutation.
  • Downstream architecture: OMT preserves the NS-PFS ordered token sequence and supervises the ordered data token states.
Pipeline
NS-PFS → modality-specific clustering → neural synchrony → Hungarian matching → synchronized local CPPs → global permutation → OMT

The key difference is therefore methodological. GEDS performs descriptor/graph-based global ranking over an already fused representation, whereas NS-PFS explicitly discovers and aligns modality-specific structures before constructing the synchronized feature sequence.

The downstream sequence-processing architectures also differ. OEMT compresses the GEDS-ordered representation using order-aware pooling before memory-augmented Linformer processing, whereas iSyncTab's OMT retains the NS-PFS-ordered multimodal token sequence and applies sequencing supervision directly to the ordered data-token representations.

The two works investigate the same broader research direction using distinct sequencing algorithms, distinct cross-modal alignment mechanisms, and different downstream sequence-processing designs.

Code and Resources

pip install istructtab

Takeaway

Structure First GEDS Explicitly sequences the fused multimodal feature space rather than treating column arrangement as arbitrary.
Order-Aware Processing OEMT Preserves the feature structure through pooling, memory augmentation, efficient attention, and sequencing supervision.
Research Progression → iSyncTab Extends the broader feature-sequencing direction toward explicit cross-modal synchronization and paired feature alignment.

Acknowledgements

We gratefully acknowledge the organizations and individuals who supported this research and its presentation at ICPR 2026.

NSF Support This work was supported in part by the U.S. National Science Foundation (NSF) under Award #1920920 , Award #2125872 , and Award #2223793 .
ICPR 2026 Registration Support We gratefully acknowledge the International Association for Pattern Recognition (IAPR) for providing registration support for ICPR 2026.
Reviewers We sincerely thank the anonymous ICPR 2026 reviewers for their careful evaluation, constructive comments, and valuable feedback that helped improve this work.

Citation

Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Kumar Gyawali, Gianfranco Doretto, and Donald A. Adjeroh. iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data. International Conference on Pattern Recognition (ICPR), 2026.

@inproceedings{habib2026istructtab,
  title     = {iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data},
  author    = {Habib, Al Zadid Sultan Bin and Ahamed, Md Younus and Gyawali, Prashnna Kumar and Doretto, Gianfranco and Adjeroh, Donald A.},
  booktitle = {Proceedings of the International Conference on Pattern Recognition (ICPR)},
  year      = {2026},
  doi       = {10.1007/978-3-032-31404-8_43}
}
↑