iS iSyncTab Neural synchrony-guided cross-modal feature sequencing
Published ECCV 2026 Multimodal Learning Image + Tabular Audio + Video NeuroAI Feature Sequencing

iSyncTab

iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony treats multimodal feature integration as a structured sequencing problem. It aligns image and tabular feature clusters with neural synchrony-inspired pairing, constructs a synchronized cross-modal order, and processes the ordered tokens with an Order-aware Memory-augmented Transformer.

Al Zadid Sultan Bin Habib1,*, Md Younus Ahamed1, Prashnna Kumar Gyawali1, Gianfranco Doretto2, Donald A. Adjeroh1,*

1 West Virginia University    2 University of Utah    * Corresponding authors

ECCV 2026 Malmö

European Conference on Computer Vision · Malmö · September 8-12, 2026

Overview

Multimodal image-tabular learning typically combines visual and structured features without an explicit cross-modal order. iSyncTab instead frames feature sequencing as a Column Permutation Problem (CPP) and learns an ordered multimodal representation before sequence-sensitive fusion.

NS-PFS

Neural Synchrony-guided Paired Feature Sequencing clusters both modalities, computes an energy-similarity synchrony matrix, and pairs clusters with Hungarian matching.

OMT

The Order-aware Memory-augmented Transformer uses Linformer attention, learnable memory tokens, and order-consistency regularization over the synchronized feature sequence.

Cross-modal structure

The learned permutation reduces feature dispersion and provides an explicit structural inductive bias for image-tabular fusion rather than treating the fused representation as arbitrarily ordered.

Architecture

iSyncTab architecture with image and tabular encoders, NS-PFS feature sequencing, and OMT prediction
End-to-end iSyncTab architecture. Image and tabular encoders produce modality-specific feature matrices. NS-PFS clusters and synchronizes cross-modal feature groups, derives a global permutation, and OMT processes the ordered multimodal token sequence with classification and sequencing-consistency objectives.

Method

iSyncTab starts from a fused image-tabular feature matrix \(X_{\mathrm{mm}}\in\mathbb{R}^{N\times m}\) and seeks a permutation \(\pi\in S_m\) that reduces pairwise dispersion.

Column Permutation Problem \[ \min_{\pi\in S_m} D(\pi) = \sum_{1\le u\lt v\le m} w_{uv}\, \lvert \pi(u)-\pi(v) \rvert \]

The full pipeline couples this structured ordering with neural synchrony-inspired cross-modal alignment and an order-aware transformer backbone.

NS-PFS: Neural Synchrony-guided Paired Feature Sequencing

1. Modality-specific features Image and tabular encoders produce feature matrices.
↓
2. Cross-modal clustering Feature columns are transposed and clustered separately for image and tabular modalities.
↓
3. Synchrony matrix Cluster energy coherence and centroid similarity define cross-modal synchrony scores.
↓
4. Hungarian pairing Visual and tabular clusters are matched one-to-one by maximum total synchrony.
↓
5. Local CPP sequencing Each matched joint cluster is ordered under a synchrony-modulated dispersion objective.
↓
6. Global synchronized permutation Ranked local sequences are concatenated to form the final NS-PFS order.

NS-PFS output

Produces the final synchronized cross-modal permutation used to order visual and tabular tokens before OMT.

For visual cluster \(I_a\) and tabular cluster \(T_b\), iSyncTab combines energy coherence \(E_{ab}\) with centroid similarity \(S_{ab}\). Their product defines the synchrony score \(\Psi_{ab}\).

Energy and phase-analogue synchrony \[ E_{ab} = 1- \frac{ |\xi_{I_a}-\xi_{T_b}| }{ \xi_{I_a}+\xi_{T_b}+\epsilon } \] \[ S_{ab} = \frac{ \langle \mu_{I_a}, \mu_{T_b} \rangle }{ \|\mu_{I_a}\| \|\mu_{T_b}\| } \] \[ \Psi_{ab} = E_{ab} S_{ab} \]
Neural synchrony analogy Cross-modal groups interact more strongly when they are both energy-compatible and directionally aligned, mirroring coherence-based communication in neural systems.
Metric-generalized sequencing NS-PFS supports feature-wise statistical, distance, and information-theoretic measures, including variance, energy, Euclidean, Manhattan, cosine, correlation, KL, and JS.
Efficient cross-modal pairing Hungarian assignment operates over modality-specific feature clusters rather than the entire feature set.
Global order Matched local sequences are concatenated as \( \pi_{\mathrm{NS\!-\!PFS}} = [ \pi^\star_{J_{(1)}} \Vert \cdots \Vert \pi^\star_{J_{(R)}} ] \).

OMT: Order-aware Memory-augmented Transformer

NS-PFS ordered tokens Tabular and visual tokens are permuted into the synchronized cross-modal order.
↓
Learnable memory tokens Global memory tokens are prepended as compact readout states.
↓
Linformer encoder Linear-complexity attention models the ordered multimodal token sequence.
↓
Joint objective Classification and sequence-consistency heads are optimized together.

OMT output

Produces memory-based prediction states and sequence-aware ordered token representations.

OMT uses the NS-PFS permutation as an explicit structural constraint. Memory-token states provide the global representation for classification, while ordered data-token states are used by the sequencing head.

Order-aware training objective \[ \mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{CE}} + \lambda_{\mathrm{FS}} \frac{1}{BL} \sum_{\iota=1}^{B} \sum_{\ell=1}^{L} \left( q_{\iota\ell} - \beta_{\pi,\ell} \right)^2 \]
Memory-based readout Learnable memory tokens aggregate information from the synchronized multimodal sequence.
Sequence consistency The auxiliary loss encourages OMT to retain the structural order derived by NS-PFS.
Scalable attention Linformer reduces the attention cost from quadratic self-attention to linear scaling in sequence length for fixed projection rank.

Main Results

Across 15 benchmarks, iSyncTab achieves the strongest overall ranking and zero average regret, with particularly strong gains when image and tabular modalities provide complementary information.

1.08 Average Rank
0.00 Average Regret
15 Benchmarks
Consistent multimodal performance iSyncTab reaches 99.60% on DVM, 86.82% on HAM10000, 85.63% on DeepLesion, 42.44% on Pokémon, 88.60% on CheXpert, and 88.02% on Pet Finder.
Best overall ranking Avg. Rank: \(1.08\pm0.19\). Avg. Regret: \(0.00\pm0.00\).
Robust under label noise On HAM10000, iSyncTab achieves the highest mean accuracy across the tested 10%, 20%, 40%, and 60% label-noise settings.
Model DVM HAM DLes Pok CheX Pet Avg. Rank ↓ Avg. Regret ↓
DAFT 74.22 74.88 77.07 18.05 67.75 45.45 7.42 ± 2.71 29.27 ± 19.86
Interact Fuse 78.58 84.87 70.68 11.71 50.90 52.44 8.33 ± 4.71 30.65 ± 22.37
MMCL 85.79 70.09 67.34 38.05 49.00 29.72 10.67 ± 5.19 27.75 ± 16.03
TIP 98.27 70.39 69.17 37.56 87.50 83.86 6.00 ± 4.07 10.08 ± 8.09
STiL 99.27 78.48 81.35 27.32 88.60 87.68 4.42 ± 2.83 11.73 ± 20.49
iSyncTab 99.60 86.82 85.63 42.44 88.60 88.02 1.08 ± 0.19 0.00 ± 0.00

Test accuracy (%) on the six image-tabular benchmarks. This webpage shows the multimodal comparison subset from the full paper table; the paper also reports tabular-only and image-only baselines.

Efficiency and Accuracy-Cost Trade-offs

The paper evaluates multimodal models on DVM using accuracy, FLOPs, parameter count, peak inference GPU memory, and a composite complexity proxy. iSyncTab achieves the strongest accuracy while remaining competitive in compute and memory rather than relying on a larger model.

iSyncTab efficiency ranking
Weighted efficiency ranking.
iSyncTab accuracy complexity Pareto frontier
Pareto frontier of accuracy versus composite complexity.
Accuracy versus composite complexity
Accuracy versus composite complexity.
Normalized efficiency radar profile
Normalized accuracy and resource radar profile.
Accuracy versus FLOPs
Test accuracy versus FLOPs.
Accuracy versus model parameters
Test accuracy versus trainable parameters.
Parameters versus FLOPs
Joint scaling of parameters and FLOPs.
Peak inference GPU memory
Peak inference-time GPU memory footprint.

Component Ablation

Removing memory tokens, the sequencing loss, or feature sequencing reduces performance, while random column shuffling causes the strongest degradation. These experiments support the combined contribution of NS-PFS and OMT.

Variant DVM w/o tuning DVM w/ tuning HAM w/o tuning HAM w/ tuning
Concat Fuse 96.50 96.20 76.58 78.42
iSyncTab: ResNet-50 + NS-PFS + OMT 99.80 99.60 78.72 86.82
iSyncTab w/o memory tokens 92.68 87.60 77.90 80.12
iSyncTab w/o sequencing-loss term 97.80 98.70 78.06 84.36
iSyncTab w/o feature sequencing 96.80 96.50 78.16 78.60
iSyncTab w/ random column shuffle 86.72 82.31 62.06 62.16

HAM10000 Calibration, Robustness, and Inference Diagnostics

Inference-level analysis on HAM10000 evaluates calibration, neighborhood consistency, confidence-based coverage, robustness to image blur and tabular perturbations, and efficiency as sequence length changes.

HAM10000 reliability diagram
Reliability diagram and calibration behavior.
HAM10000 kNN label agreement
Fused-embedding kNN label agreement.
HAM10000 confidence accuracy coverage curve
Confidence versus accuracy and coverage.
HAM10000 robustness to image blur
Robustness to Gaussian image blur.
HAM10000 robustness to tabular perturbations
Robustness to tabular shuffling and dropping.
HAM10000 efficiency versus sequence length
Efficiency and memory behavior versus sequence length.
HAM10000 trained weights and the sanitized public reproducibility checkpoint are hosted separately on Hugging Face. Original HAM10000 images and metadata are not redistributed with the model release.

Cross-Modal Generality and NS-PFS Sensitivity

Beyond image-tabular learning, the sequencing mechanism was also evaluated on audio-video data using RAVDESS. The Linformer-ResNet50 variant reaches 97.56% best test accuracy and \( 95.25\pm1.34\% \) five-fold mean accuracy.

Sensitivity analysis varies the sequencing metric and cluster count while keeping the architecture and optimization settings fixed.

NS-PFS sensitivity on HAM10000 image-tabular learning
NS-PFS metric and cluster-count sensitivity on HAM10000 image-tabular learning.
NS-PFS sensitivity on RAVDESS audio-video learning
NS-PFS sensitivity on RAVDESS audio-video learning.
ECCV 2026

iSyncTab Presentation

▶ Watch on YouTube

Watch our ECCV 2026 presentation of iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony . The presentation introduces the motivation behind cross-modal feature sequencing, NS-PFS, the Order-aware Memory-augmented Transformer, and the main experimental findings.

Code, Package, and Model Release

The public iSyncTab release includes the image-tabular implementation, a general audio-video extension, experiment notebooks, PyPI installation, and HAM10000 model artifacts hosted on Hugging Face.

GitHub Source code, experiment notebooks, architecture, reproducibility resources, and package configuration. Open repository →
PyPI Install the public iSyncTab Python package directly with pip. Open PyPI →
Hugging Face HAM10000 trained model weights, public checkpoint, configuration, and release metadata. Open model release →

Install from PyPI

pip install isynctab

from isynctab import iSyncTab, iSyncTab_AV
from isynctab import iSyncTabAV

Summary

iSyncTab treats cross-modal image-tabular integration as a structured feature-sequencing problem rather than an arbitrary concatenation problem. NS-PFS aligns visual and tabular feature groups through energy and phase-analogue synchrony, Hungarian matching, and local CPP sequencing. OMT then uses the synchronized order with learnable memory tokens, Linformer attention, and an auxiliary order-consistency objective.

Across six image-tabular benchmarks, the framework achieves the strongest overall ranking and zero average regret. Component ablations, efficiency analysis, HAM10000 inference diagnostics, and audio-video experiments further support the role and generality of synchrony-guided feature sequencing.

Acknowledgements

Computational Support We also gratefully acknowledge Zaigham Abbas Randhawa , Pouya Iranmanesh , and Ghazaleh Mirzaee for their assistance with the GPU clusters.

Citation

Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Kumar Gyawali, Gianfranco Doretto, and Donald A. Adjeroh. iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony. Proceedings of the European Conference on Computer Vision (ECCV), 2026.

@inproceedings{habib2026isynctab,
  title     = {iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony},
  author    = {Habib, Al Zadid Sultan Bin and Ahamed, Md Younus and Gyawali, Prashnna Kumar and Doretto, Gianfranco and Adjeroh, Donald A.},
  booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
  year      = {2026},
  doi       = {10.1007/978-3-032-37035-8}
}
↑