VLDB 2026 Research / reviewers in the wild / expert
Xi Xiao 0003
dblp:83/6642-3
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0000-0931-6982ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair DisentanglementabstractWhile deep generative models have significantly advanced representation learning, they may inherit or amplify biases and fairness issues by encoding sensitive attributes alongside predictive features. Enforcing strict independence in disentanglement is often unrealistic when target and sensitive factors are naturally correlated. To address this challenge, we propose CAD-VAE(Correlation-Aware Disentangled VAE), which introduces a correlated latent code to capture the information shared between the target and sensitive attributes. Given this correlated latent, our method effectively separates overlapping factors without extra domain knowledge by directly minimizing the conditional mutual information between target and sensitive codes. A relevance-driven optimization strategy refines the correlated code by efficiently capturing essential correlated features and eliminating redundancy. Extensive experiments on benchmark datasets demonstrate that CAD-VAE produces fairer representations, realistic counterfactuals, and improved fairness-aware image editing. Chenrui Ma, Xi Xiao 0003, Tianyang Wang 0004, Xiao Wang 0004, Yanning Shen |
AAAI | 2 |
| 2026 | Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model AdaptationabstractXi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, Hao Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xi Xiao 0003, Chenrui Ma, Yunbei Zhang, Chen Liu 0020, Zhuxuanzi Wang, Yanshu Li, Guosheng Hu, Tianyang Wang 0004 |
ACL (1) | 1 |
| 2026 | 4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's DiagnosisabstractMultimodal neuroimaging provides complementary structural and functional insights into both human brain organization and disease-related dynamics. Recent studies demonstrate enhanced diagnostic sensitivity for Alzheimer’s disease (AD) through synergistic integration of neuroimaging data (e.g., sMRI, fMRI) with tabular data (e.g., behavioral and cognitive tests). However, the intrinsic heterogeneity across modalities (e.g., 4D spatiotemporal fMRI dynamics vs. 3D anatomical sMRI structure) presents critical challenges for discriminative feature fusion, often leading to information loss or biased fusion. To bridge this gap, we propose M2M-AlignNet: a multimodal co-attention network with latent alignment for early AD diagnosis using sMRI and fMRI. At the core of our approach is a multi-patch-to-multi-patch (M2M) contrastive loss function that quantifies and reduces representational discrepancies via weighted patch correspondence, explicitly aligning fMRI components across brain regions with their sMRI structural substrates without one-to-one constraints. Additionally, we propose a latent-as-query co-attention module to autonomously discover fusion patterns, circumventing modality prioritization biases while minimizing feature redundancy. We conduct extensive experiments to confirm the effectiveness of our method and highlight the correspondence between fMRI and sMRI as AD biomarkers. Yuxiang Wei 0004, Yanteng Zhang, Xi Xiao 0003, Tianyang Wang 0004, Xiao Wang 0004, Vince D. Calhoun |
WACV | 3 |
| 2026 | Self-Supervised Visual Prompting for Cross-Domain Road Damage DetectionabstractThe deployment of automated pavement defect detection is often hindered by poor cross-domain generalization. Supervised detectors achieve strong in-domain accuracy but require costly re-annotation for new environments, while standard self-supervised methods capture generic features and remain vulnerable to domain shift. We propose PROBE, a self-supervised framework that visually probes target domains without labels. PROBE introduces a Self-supervised Prompt Enhancement Module (SPEM), which derives defect-aware prompts from unlabeled target data to guide a frozen ViT backbone, and a Domain-Aware Prompt Alignment (DAPA) objective, which aligns prompt-conditioned source and target representations. Experiments on four challenging benchmarks show that PROBE consistently outperforms strong supervised, self-supervised, and adaptation baselines, achieving robust zero-shot transfer, improved resilience to domain variations, and high data efficiency in few-shot adaptation. These results highlight self-supervised prompting as a practical direction for building scalable and adaptive visual inspection systems. Source code is publicly available: https://github.com/xixiaouab/PROBE/tree/main Xi Xiao 0003, Zhuxuanzi Wang, Mingqiao Mo, Chen Liu 0020, Chenrui Ma, Yanshu Li, Smita Krishnaswamy, Xiao Wang 0004, Tianyang Wang 0004 |
WACV | 1 |
| 2026 | RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage UnderstandingabstractAccurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understanding that textual information can provide. To address this limitation, we introduce RoadBench, the first multimodal benchmark for comprehensive road damage understanding. This dataset pairs high-resolution images of road damages with detailed textual descriptions, providing a richer context for model training. We also present RoadCLIP, a novel vision-language model that builds upon CLIP by integrating domain-specific enhancements. It includes a disease-aware positional encoding that captures spatial patterns of road defects and a mechanism for injecting road-condition priors to refine the model’s understanding of road damages. We further employ a GPT-driven data generation pipeline to expand the image–text pairs in Road-Bench, greatly increasing data diversity without exhaustive manual annotation. Experiments demonstrate that Road-CLIP achieves state-of-the-art performance on road damage recognition tasks, significantly outperforming existing vision-only models by 19.2%. These results highlight the advantages of integrating visual and textual information for enhanced road condition analysis, setting new benchmarks for the field and paving the way for more effective infrastructure monitoring through multimodal learning. Xi Xiao 0003, Yunbei Zhang, Janet Wang, Yuxiang Wei 0004, Hengjia Li, Yanshu Li, Xiao Wang 0004, Swalpa Kumar Roy, Tianyang Wang 0004 |
WACV | 1 |
| 2025 | TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage DetectionabstractObject detection has witnessed remarkable advancements over the past decade, largely driven by breakthroughs in deep learning and the proliferation of large-scale datasets. However, the domain of road damage detection remains relatively underexplored, despite its critical significance for applications such as infrastructure maintenance and road safety. This paper addresses this gap by introducing a novel top-down benchmark that offers a complementary perspective to existing datasets, specifically tailored for road damage detection. Our proposed Top-Down Road Damage Detection Dataset (TD-RD) includes three primary categories of road damage—cracks, potholes, and patches—captured from an top-down viewpoint. The dataset consists of 7,088 high-resolution images, encompassing 12,882 annotated instances of road damage. Additionally, we present a novel real-time object detection framework, TD-YOLOV10, designed to handle the unique challenges posed by the TD-RD dataset. Comparative studies with state-of-the-art models demonstrate competitive baseline results. By releasing TD-RD, we aim to accelerate research in this crucial area. A sample of the dataset will be made publicly available upon the paper’s acceptance. Xi Xiao 0003, Zhengji Li, Houjie Lin, Swalpa Kumar Roy, Tianyang Wang 0004, Min Xu 0009 |
ICASSP | 1 |
| 2025 | Visual Instance-aware Prompt TuningabstractVisual Prompt Tuning (VPT) has emerged as a parameter-efficient fine-tuning paradigm for vision transformers, with conventional approaches utilizing dataset-level prompts that remain the same across all input instances. We observe that this strategy results in sub-optimal performance due to high variance in downstream datasets. To address this challenge, we propose Visual Instance-aware Prompt Tuning (ViaPT), which generates instance-aware prompts based on each individual input and fuses them with dataset-level prompts, leveraging Principal Component Analysis (PCA) to retain important prompting information. Moreover, we reveal that VPT-Deep and VPT-Shallow represent two corner cases based on a conceptual understanding, in which they fail to effectively capture instance-specific information, while random dimension reduction on prompts only yields performance between the two extremes. Instead, ViaPT overcomes these limitations by balancing dataset-level and instance-level knowledge, while reducing the amount of learnable parameters compared to VPT-Deep. Extensive experiments across 34 diverse datasets demonstrate that our method consistently outperforms state-of-the-art baselines, establishing a new paradigm for analyzing and optimizing visual prompts for vision transformers. Xi Xiao 0003, Yunbei Zhang, Xingjian Li 0002, Tianyang Wang 0004, Xiao Wang 0004, Yuxiang Wei 0004, Jihun Hamm, Min Xu 0009 |
ACM Multimedia | 1 |
| 2025 | MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual DecodingabstractDecoding visual experiences from fMRI offers a powerful avenue to understand human perception and develop advanced brain-computer interfaces. However, current progress often prioritizes maximizing reconstruction fidelity while overlooking interpretability, an essential aspect for deriving neuroscientific insight. To address this gap, we propose MoRE-Brain, a neuro-inspired framework designed for high-fidelity, adaptable, and interpretable visual reconstruction. MoRE-Brain uniquely employs a hierarchical Mixture-of-Experts architecture where distinct experts process fMRI signals from functionally related voxel groups, mimicking specialized brain networks. The experts are first trained to encode fMRI into the frozen CLIP space. A finetuned diffusion model then synthesizes images, guided by expert outputs through a novel dual-stage routing mechanism that dynamically weighs expert contributions across the diffusion process. MoRE-Brain offers three main advancements: First, it introduces a novel Mixture-of-Experts architecture grounded in brain network principles for neuro-decoding. Second, it achieves efficient cross-subject generalization by sharing core expert networks while adapting only subject-specific routers. Third, it provides enhanced mechanistic insight, as the explicit routing reveals precisely how different modeled brain regions shape the semantic and spatial attributes of the reconstructed image. Extensive experiments validate MoRE-Brain’s high reconstruction fidelity, with bottleneck analyses further demonstrating its effective utilization of fMRI signals, distinguishing genuine neural decoding from over-reliance on generative priors. Consequently, MoRE-Brain marks a substantial advance towards more generalizable and interpretable fMRI-based visual decoding. Yuxiang Wei 0004, Yanteng Zhang, Xi Xiao 0003, Tianyang Wang 0004, Xiao Wang 0004, Vince D. Calhoun |
NeurIPS | 3 |
| 2025 | ETT-CKGE: Efficient Task-Driven Tokens for Continual Knowledge Graph Embedding
Lijing Zhu, Qizhen Lan, Qing Tian 0003, Xi Xiao 0003, Tiehang Duan, Cui Tao, Shuteng Niu |
ECML/PKDD (6) | 8 |
| 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate DownscalingabstractSparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data. Xiao Wang 0004, Jong-Youl Choi, Takuya Kurihana, Isaac Lyngaas, Hong-Jun Yoon, Xi Xiao 0003, David Pugmire, Nasik Muhammad Nafi, Aristeidis Tsaris, Ashwin M. Aji, Maliha Hossain, Mohamed Wahib, Dali Wang, Peter E. Thornton, Prasanna Balaprakash, Moetasim Ashfaq, Dan Lu 0001 |
SC | 6 |
| 2024 | HGTDP-DTA: Hybrid Graph-Transformer with Dynamic Prompt for Drug-Target Binding Affinity Prediction
Xi Xiao 0003, Lijing Zhu, Gaofei Chen, Zhengji Li, Tianyang Wang 0004, Min Xu 0009 |
ICONIP (4) | 1 |
| 2024 | Multi-dimension Transformer with Attention-based Filtering for Medical Image SegmentationabstractThe accurate segmentation of medical images is crucial for diagnosing and treating diseases. Recent studies demonstrate that vision transformer-based methods have significantly improved performance in medical image segmentation, primarily due to their superior ability to establish global relationships among features and adaptability to various inputs. However, these methods struggle with the low signal-to-noise ratio inherent to medical images. Additionally, the effective utilization of channel and spatial information, which are essential for medical image segmentation, is limited by the representation capacity of self-attention. To address these challenges, we propose a Multi-dimension Transformer with Attention-based Filtering (MDT-AF), which redesigns the patch embedding and self-attention mechanism for medical image segmentation. MDT-AF incorporates an attention-based feature filtering mechanism into the patch embedding blocks and employs a coarse-to-fine process to mitigate the impact of a low signal-to-noise ratio. To better capture complex structures in medical images, MDT-AF extends self-attention and introduces an interaction mechanism to build and enhance feature relationships between dimensions, which can achieve richer feature representations across the spatial and channel dimensions. Experimental results on three public medical image segmentation benchmarks show that MDT-AF achieves state-of-the-art (SOTA) performance. Xi Xiao 0003, Qizhen Lan, Xuanyao Huang, Qing Tian 0003, Swalpa Kumar Roy, Tianyang Wang 0004 |
ICTAI | 2 |