iSyncTab
iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony treats multimodal feature integration as a structured sequencing problem. It aligns image and tabular feature clusters with neural synchrony-inspired pairing, constructs a synchronized cross-modal order, and processes the ordered tokens with an Order-aware Memory-augmented Transformer.
1 West Virginia University 2 University of Utah * Corresponding authors
European Conference on Computer Vision · Malmö · September 8-12, 2026
Overview
Multimodal image-tabular learning typically combines visual and structured features without an explicit cross-modal order. iSyncTab instead frames feature sequencing as a Column Permutation Problem (CPP) and learns an ordered multimodal representation before sequence-sensitive fusion.
Neural Synchrony-guided Paired Feature Sequencing clusters both modalities, computes an energy-similarity synchrony matrix, and pairs clusters with Hungarian matching.
The Order-aware Memory-augmented Transformer uses Linformer attention, learnable memory tokens, and order-consistency regularization over the synchronized feature sequence.
The learned permutation reduces feature dispersion and provides an explicit structural inductive bias for image-tabular fusion rather than treating the fused representation as arbitrarily ordered.
Architecture
Method
iSyncTab starts from a fused image-tabular feature matrix \(X_{\mathrm{mm}}\in\mathbb{R}^{N\times m}\) and seeks a permutation \(\pi\in S_m\) that reduces pairwise dispersion.
The full pipeline couples this structured ordering with neural synchrony-inspired cross-modal alignment and an order-aware transformer backbone.
NS-PFS: Neural Synchrony-guided Paired Feature Sequencing
NS-PFS output
Produces the final synchronized cross-modal permutation used to order visual and tabular tokens before OMT.
For visual cluster \(I_a\) and tabular cluster \(T_b\), iSyncTab combines energy coherence \(E_{ab}\) with centroid similarity \(S_{ab}\). Their product defines the synchrony score \(\Psi_{ab}\).
OMT: Order-aware Memory-augmented Transformer
OMT output
Produces memory-based prediction states and sequence-aware ordered token representations.
OMT uses the NS-PFS permutation as an explicit structural constraint. Memory-token states provide the global representation for classification, while ordered data-token states are used by the sequencing head.
Main Results
Across 15 benchmarks, iSyncTab achieves the strongest overall ranking and zero average regret, with particularly strong gains when image and tabular modalities provide complementary information.
| Model | DVM | HAM | DLes | Pok | CheX | Pet | Avg. Rank ↓ | Avg. Regret ↓ |
|---|---|---|---|---|---|---|---|---|
| DAFT | 74.22 | 74.88 | 77.07 | 18.05 | 67.75 | 45.45 | 7.42 ± 2.71 | 29.27 ± 19.86 |
| Interact Fuse | 78.58 | 84.87 | 70.68 | 11.71 | 50.90 | 52.44 | 8.33 ± 4.71 | 30.65 ± 22.37 |
| MMCL | 85.79 | 70.09 | 67.34 | 38.05 | 49.00 | 29.72 | 10.67 ± 5.19 | 27.75 ± 16.03 |
| TIP | 98.27 | 70.39 | 69.17 | 37.56 | 87.50 | 83.86 | 6.00 ± 4.07 | 10.08 ± 8.09 |
| STiL | 99.27 | 78.48 | 81.35 | 27.32 | 88.60 | 87.68 | 4.42 ± 2.83 | 11.73 ± 20.49 |
| iSyncTab | 99.60 | 86.82 | 85.63 | 42.44 | 88.60 | 88.02 | 1.08 ± 0.19 | 0.00 ± 0.00 |
Test accuracy (%) on the six image-tabular benchmarks. This webpage shows the multimodal comparison subset from the full paper table; the paper also reports tabular-only and image-only baselines.
Efficiency and Accuracy-Cost Trade-offs
The paper evaluates multimodal models on DVM using accuracy, FLOPs, parameter count, peak inference GPU memory, and a composite complexity proxy. iSyncTab achieves the strongest accuracy while remaining competitive in compute and memory rather than relying on a larger model.
Component Ablation
Removing memory tokens, the sequencing loss, or feature sequencing reduces performance, while random column shuffling causes the strongest degradation. These experiments support the combined contribution of NS-PFS and OMT.
| Variant | DVM w/o tuning | DVM w/ tuning | HAM w/o tuning | HAM w/ tuning |
|---|---|---|---|---|
| Concat Fuse | 96.50 | 96.20 | 76.58 | 78.42 |
| iSyncTab: ResNet-50 + NS-PFS + OMT | 99.80 | 99.60 | 78.72 | 86.82 |
| iSyncTab w/o memory tokens | 92.68 | 87.60 | 77.90 | 80.12 |
| iSyncTab w/o sequencing-loss term | 97.80 | 98.70 | 78.06 | 84.36 |
| iSyncTab w/o feature sequencing | 96.80 | 96.50 | 78.16 | 78.60 |
| iSyncTab w/ random column shuffle | 86.72 | 82.31 | 62.06 | 62.16 |
HAM10000 Calibration, Robustness, and Inference Diagnostics
Inference-level analysis on HAM10000 evaluates calibration, neighborhood consistency, confidence-based coverage, robustness to image blur and tabular perturbations, and efficiency as sequence length changes.
Cross-Modal Generality and NS-PFS Sensitivity
Beyond image-tabular learning, the sequencing mechanism was also evaluated on audio-video data using RAVDESS. The Linformer-ResNet50 variant reaches 97.56% best test accuracy and \( 95.25\pm1.34\% \) five-fold mean accuracy.
Sensitivity analysis varies the sequencing metric and cluster count while keeping the architecture and optimization settings fixed.
iSyncTab Presentation
Watch our ECCV 2026 presentation of iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony . The presentation introduces the motivation behind cross-modal feature sequencing, NS-PFS, the Order-aware Memory-augmented Transformer, and the main experimental findings.
Code, Package, and Model Release
The public iSyncTab release includes the image-tabular implementation, a general audio-video extension, experiment notebooks, PyPI installation, and HAM10000 model artifacts hosted on Hugging Face.
Install from PyPI
pip install isynctab from isynctab import iSyncTab, iSyncTab_AV from isynctab import iSyncTabAV
Summary
iSyncTab treats cross-modal image-tabular integration as a structured feature-sequencing problem rather than an arbitrary concatenation problem. NS-PFS aligns visual and tabular feature groups through energy and phase-analogue synchrony, Hungarian matching, and local CPP sequencing. OMT then uses the synchronized order with learnable memory tokens, Linformer attention, and an auxiliary order-consistency objective.
Across six image-tabular benchmarks, the framework achieves the strongest overall ranking and zero average regret. Component ablations, efficiency analysis, HAM10000 inference diagnostics, and audio-video experiments further support the role and generality of synchrony-guided feature sequencing.
Acknowledgements
Citation
Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Kumar Gyawali, Gianfranco Doretto, and Donald A. Adjeroh. iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony. Proceedings of the European Conference on Computer Vision (ECCV), 2026.
@inproceedings{habib2026isynctab,
title = {iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony},
author = {Habib, Al Zadid Sultan Bin and Ahamed, Md Younus and Gyawali, Prashnna Kumar and Doretto, Gianfranco and Adjeroh, Donald A.},
booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
year = {2026},
doi = {10.1007/978-3-032-37035-8}
}