EDBT 2026 Demo / reviewers in the wild / expert
Qi'ao Xu
dblp:360/1815
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0002-5567-438XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 39% Video understanding and tracking · 30% Vision and language · 30% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
domain generalization |
1.0 | 1 | 2026 | EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering · AAAI 2026 |
Computer vision › Video understanding and tracking › video question answering
egocentric video question answering |
1.0 | 1 | 2026 | EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.3 | 1 | 2026 | EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.0fine-tuning · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question AnsweringabstractRecent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deployment inevitably encounters domain shifts, where target domains differ substantially in both visual style and semantic content. To bridge this gap, we introduce EgoCross, a comprehensive benchmark designed to evaluate the cross-domain generalization of MLLMs in EgocentricQA. EgoCross covers four diverse and challenging domains, including surgery, industry, extreme sports, and animal perspective, representing realistic and high-impact application scenarios. It comprises approximately 1,000 QA pairs across 798 video clips, spanning four key QA tasks: prediction, recognition, localization, and counting. Each QA pair provides both OpenQA and CloseQA formats to support fine-grained evaluation. Extensive experiments show that most existing MLLMs, whether general-purpose or egocentric-specialized, struggle to generalize to domains beyond daily life, highlighting the limitations of current models. Furthermore, we conduct several pilot studies, e.g., fine-tuning and reinforcement learning, to explore potential improvements. We hope EgoCross and our accompanying analysis will serve as a foundation for advancing domain-adaptive, robust egocentric video understanding. Yuqian Fu, Tianwen Qian, Qi'ao Xu, Silong Dai, Danda Pani Paudel, Luc Van Gool, Xiaoling Wang 0004 |
AAAI | 4 |
| 2025 | GTPC-SSCD: Gate-guided Two-level Perturbation Consistency-based Semi-Supervised Change DetectionabstractSemi-supervised change detection (SSCD) utilizes partially labeled data and abundant unlabeled data to detect differences between multi-temporal remote sensing images. The mainstream SSCD methods based on consistency regularization have limitations. They perform perturbations mainly at a single level, restricting the utilization of unlabeled data and failing to fully tap its potential. In this paper, we introduce a novel Gate-guided Two-level Perturbation Consistency regularization-based SSCD method (GTPC-SSCD). It simultaneously maintains strong-to-weak consistency at the image level and perturbation consistency at the feature level, enhancing the utilization efficiency of unlabeled data. Moreover, we develop a hardness analysis-based gating mechanism to assess the training complexity of different samples and determine the necessity of performing feature perturbations for each sample. Through this differential treatment, the network can explore the potential of unlabeled data more efficiently. Extensive experiments conducted on six benchmark CD datasets demonstrate the superiority of our GTPC-SSCD over seven state-of-the-art methods. Qi'ao Xu, Zongyu Guo, Rui Huang 0006, Yuxiang Zhang 0003 |
ICME | 2 |
| 2025 | HSACNet: Hierarchical Scale-Aware Consistency Regularized Semi-Supervised Change DetectionabstractSemi-Supervised change detection (SSCD) aims to detect changes between bi-temporal remote sensing images by utilizing limited labeled data and abundant unlabeled data. Existing methods struggle in complex scenarios, exhibiting poor performance when confronted with noisy data. They typically neglect intra-layer multi-scale features while emphasizing inter-layer fusion, harming the integrity of change objects with different scales. In this paper, we propose HSACNet, a Hierarchical Scale-Aware Consistency regularized Network for SSCD. Specifically, we integrate Segment Anything Model 2 (SAM2), using its Hiera backbone as the encoder to extract inter-layer multi-scale features and applying adapters for parameter-efficient fine-tuning. Moreover, we design a Scale-Aware Differential Attention Module (SADAM) that can precisely capture intra-layer multi-scale change features and suppress noise. Additionally, a dual-augmentation consistency regularization strategy is adopted to effectively utilize the unlabeled data. Extensive experiments across four CD benchmarks demonstrate that our HSACNet achieves state-of-the-art performance, with reduced parameters and computational cost. Qi'ao Xu, Pengfei Wang 0009, Tianwen Qian, Xiaoling Wang 0004 |
ICME | 1 |
| 2024 | Cross Branch Feature Fusion Decoder for Consistency Regularization-Based Semi-Supervised Change DetectionabstractSemi-supervised change detection (SSCD) utilizes partially labeled data and a large amount of unlabeled data to detect changes. However, the transformer-based SSCD network does not perform as well as the convolution-based SSCD network due to the lack of labeled data. To overcome this limitation, we introduce a new decoder called Cross Branch Feature Fusion CBFF, which combines the strengths of both local convolutional branch and global transformer branch. The convolutional branch is easy to learn and can produce high-quality features with a small amount of labeled data. The transformer branch, on the other hand, can extract global context features but is hard to learn without a lot of labeled data. Using CBFF, we build our SSCD model based on a strong-to-weak consistency strategy. Through comprehensive experiments on WHU-CD and LEVIR-CD datasets, we have demonstrated the superiority of our method over seven state-of-the-art SSCD methods. Qi'ao Xu, Jingcheng Zeng, Sihua Gao |
ICASSP | 2 |
| 2023 | HQFS: High-Quality Feature Selection for Accurate Change Detection
Qi'ao Xu, Qingyi Zhao, Rui Huang 0006, Yuxiang Zhang 0003 |
ICIG (1) | 2 |