iStructTab: Structured Feature Sequencing for
Multimodal Learning of Image and Tabular Data
formulates multimodal feature ordering as a
Column Permutation Problem (CPP)
.
It introduces
Graph-Enhanced Descriptor Sequencing (GEDS)
to derive a structured global feature order and an
Order-Aware Efficient Transformer with Memory
Augmentation (OEMT)
to preserve and supervise that ordering during
downstream learning.
Al Zadid Sultan Bin Habib1,
Md Younus Ahamed1,
Prashnna Kumar Gyawali1,
Gianfranco Doretto2,
Donald A. Adjeroh1
International Conference on Pattern Recognition
· ICPR 2026
Overview
Image-tabular multimodal learning combines visual
representations with structured metadata, but the
resulting fused feature vector is commonly treated
as an arbitrary collection of coordinates.
iStructTab asks whether the fused representation
itself should first be explicitly structured.
We formulate multimodal feature sequencing as a
Column Permutation Problem (CPP)
and seek a feature ordering that reduces dispersion
and encourages a more coherent multimodal
representation.
iStructTab first encodes image and tabular inputs,
concatenates their representations, derives a global
feature sequence using GEDS, and processes the
reordered representation using OEMT.
CPP Formulation
Multimodal feature ordering is treated as an
explicit combinatorial structure-learning
problem.
GEDS
Mean-variance feature descriptors, cosine
similarity, graph refinement, and scalar
ranking produce a global feature permutation.
OEMT
Order-aware pooling, memory augmentation,
Linformer encoding, and sequencing supervision
preserve the structured representation.
1.50 ± 0.76
Average Rank
2.21 ± 4.59
Average Regret
99.29%
DVM Accuracy with Tuning
Architecture
Overview of iStructTab. Modality-specific image
and tabular representations are concatenated into
a unified feature matrix. GEDS computes statistical
descriptors and graph-based relationships to derive
the global feature permutation
\(\pi_{\mathrm{GEDS}}\). The reordered representation
is subsequently processed by OEMT.
The objective defines a feature permutation
\(\pi\) over the fused multimodal representation.
In iStructTab, GEDS provides the practical sequencing
procedure.
Graph-Enhanced Descriptor Sequencing (GEDS)
GEDS operates on the
fused image-tabular representation
.
For each feature column \(f_j\), we compute a
low-dimensional statistical descriptor based on
its mean and variance.
The descriptors are normalized and compared using
cosine similarity to construct a feature-descriptor
graph. Graph convolution subsequently refines those
descriptors.
\[
A = DD^\top
\]
\[
\hat{A}
=
D_{\mathrm{deg}}^{-1/2}
(A+I)
D_{\mathrm{deg}}^{-1/2}
\]
\[
D'
=
\sigma(
\hat{A}DW
)
\]
Each graph-refined descriptor is assigned a scalar
score. Sorting these scores produces the final
GEDS feature permutation.
Order-Aware Efficient Transformer with Memory
Augmentation (OEMT)
OEMT processes the structured feature representation
produced by GEDS.
The ordered feature sequence is first transformed
into token embeddings and compressed from \(m\)
original ordered features to \(k \ll m\) summary
tokens through
order-aware pooling
.
Classification uses the memory-token representation,
while an additional sequencing head predicts the
target feature order.
Training therefore jointly optimizes classification
and feature sequencing.
iStructTab is evaluated across six multimodal
image-tabular datasets:
DVM,
HAM10000,
Deep Lesion,
Pokémon,
CheXpert,
and PetFinder.
Across the complete benchmark suite, iStructTab
obtains an overall
average rank of 1.50 ± 0.76
and
average regret of 2.21 ± 4.59.
Model
DVM
HAM
DLes
Pok
CheX
Pet
Avg. Rank ↓
Avg. Regret ↓
TIP
98.27*
70.39
69.17
63.40∗
87.50
83.86
5.83 ± 4.21
6.86 ± 6.01
STiL
99.27
78.48
81.35
27.32
88.60
87.68
4.33 ± 2.92
8.51 ± 14.72
iStructTab
99.29*
85.23
83.75
68.29
76.15
87.75
1.50 ± 0.76
2.21 ± 4.59
Accuracy (%) on the image-tabular benchmarks.
See the paper for the complete comparison including
all image-only, tabular-only, and multimodal
baselines. * denotes results taken from another paper.
Robustness and Inference Diagnostics
Beyond predictive performance, the paper examines
robustness under noisy labels and inference-level
perturbations.
On HAM10000, iStructTab maintains stable performance
as label corruption increases.
For a fused feature matrix
\(F \in \mathbb{R}^{N \times m}\),
GEDS computes feature statistics in
\(\mathcal{O}(Nm)\) and the descriptor affinity graph
in \(\mathcal{O}(m^2)\).
OEMT reduces the ordered feature sequence before
applying a Linformer encoder, avoiding full
quadratic self-attention over the original fused
sequence.
In the DVM efficiency analysis, iStructTab reaches
the low-complexity / high-accuracy Pareto region and
achieves the strongest reported normalized
efficiency score.
The DVM ablation isolates the roles of GEDS,
memory augmentation, and sequencing supervision.
The strongest degradation occurs when explicit
feature sequencing is removed, indicating that the
gains cannot simply be attributed to the image
backbone.
Variant
Without Tuning
With Tuning
Concat Fuse, no sequence / memory
90.42
89.60
iStructTab, ResNet-50 + GEDS
96.85
99.29
iStructTab, ResNeXt-50 + GEDS
91.88
93.33
iStructTab without memory tokens
92.84
87.60
iStructTab without sequencing loss
96.85
96.78
iStructTab without feature sequencing
80.80
82.98
No sequence with memory loss
13.91
16.86
From iStructTab to iSyncTab
iStructTab was our first approach to explicit
feature sequencing for image-tabular multimodal
learning.
It demonstrated that the organization of the fused
multimodal feature representation itself can be an
important inductive bias.
Our later
iSyncTab
work, published at
ECCV 2026
,
advances this research direction by making
explicit cross-modal alignment part of the
sequencing algorithm itself.
Modality handling:
visual and tabular feature streams remain
explicitly separated during sequencing.
Structure discovery:
visual and tabular features are clustered
separately.
Cross-modal alignment:
energy coherence and centroid /
phase-analogue similarity measure neural
synchrony.
Matching:
the Hungarian algorithm performs optimal
cross-modal bipartite cluster assignment.
Output:
matched visual-tabular clusters produce
synchronized local sequences that are
combined into a global permutation.
Downstream architecture:
OMT preserves the NS-PFS ordered token
sequence and supervises the ordered data
token states.
Pipeline
NS-PFS → modality-specific clustering →
neural synchrony → Hungarian matching →
synchronized local CPPs →
global permutation → OMT
The key difference is therefore methodological.
GEDS performs descriptor/graph-based global
ranking over an already fused representation,
whereas
NS-PFS explicitly discovers and aligns
modality-specific structures before constructing
the synchronized feature sequence.
The downstream sequence-processing architectures
also differ. OEMT compresses the GEDS-ordered
representation using order-aware pooling before
memory-augmented Linformer processing, whereas
iSyncTab's OMT retains the NS-PFS-ordered
multimodal token sequence and applies sequencing
supervision directly to the ordered data-token
representations.
The two works investigate
the same broader research direction using
distinct sequencing algorithms,
distinct cross-modal alignment mechanisms,
and different downstream sequence-processing
designs.
Structure First
GEDS
Explicitly sequences the fused multimodal feature
space rather than treating column arrangement as
arbitrary.
Order-Aware Processing
OEMT
Preserves the feature structure through pooling,
memory augmentation, efficient attention, and
sequencing supervision.
Research Progression
→ iSyncTab
Extends the broader feature-sequencing direction
toward explicit cross-modal synchronization and
paired feature alignment.
Acknowledgements
We gratefully acknowledge the organizations and individuals
who supported this research and its presentation at ICPR 2026.
ICPR 2026 Registration Support
We gratefully acknowledge the
International Association for Pattern Recognition (IAPR)
for providing registration support for
ICPR 2026.
Reviewers
We sincerely thank the
anonymous ICPR 2026 reviewers
for their careful evaluation, constructive comments,
and valuable feedback that helped improve this work.
Citation
Al Zadid Sultan Bin Habib,
Md Younus Ahamed,
Prashnna Kumar Gyawali,
Gianfranco Doretto,
and Donald A. Adjeroh.
iStructTab: Structured Feature Sequencing for
Multimodal Learning of Image and Tabular Data.
International Conference on Pattern Recognition
(ICPR), 2026.
@inproceedings{habib2026istructtab,
title = {iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data},
author = {Habib, Al Zadid Sultan Bin and Ahamed, Md Younus and Gyawali, Prashnna Kumar and Doretto, Gianfranco and Adjeroh, Donald A.},
booktitle = {Proceedings of the International Conference on Pattern Recognition (ICPR)},
year = {2026},
doi = {10.1007/978-3-032-31404-8_43}
}