Zehao Xiao

dblp:225/5426 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Adaptive Latent Decomposition for Domain Generalization in Time Series Forecasting
abstract
Time series forecasting is essential in many real-world applications, yet developing models that generalize well to unseen and related domains—such as forecasting web traffic on new web sites/platforms or predicting e-commerce demand in new regions—remains a big challenge. Prior work addressing this problem, known as domain generalization, focuses on identifying common patterns but often overlooks the complex characteristics of time series data and fails to use the available information from test samples of unseen domains. We propose a novel approach, adaptive latent decomposition (ALD) for domain generalization in time series forecasting, which consists of a decomposed variational autoencoder (VAE) and an adaptive inference mechanism to improve predictive performance in unseen domains. The decomposed VAE involves a learnable kernel-selection mechanism that learns latent variables of decomposed components of time series, i.e., trend-cyclical and seasonal components. These latent variables capture hidden temporal dependencies in time series data, allowing forecasting models to learn general patterns from various seen domains. The adaptive inference mechanism bridges the gap between seen and unseen domains with a sample-wise optimization strategy specifically designed for time series forecasting. ALD builds a latent variable-aware pre-trained model and tailors it for each test sample, improving the generalization on unseen test domains. We validate ALD across six real-world datasets, from online behavior to complex temporal systems, demonstrating its superior generalization performance compared to state-of-the-art methods.
Songgaojun Deng, Zehao Xiao, Maarten de Rijke
ACM Trans. Knowl. Discov. Data2
2025 DynaPrompt: Dynamic Test-Time Prompt Tuning
abstract
Test-time prompt tuning enhances zero-shot generalization of vision-language models but tends to ignore the relatedness among test samples during inference. Online test-time prompt tuning provides a simple way to leverage the information in previous test samples, albeit with the risk of prompt collapse due to error accumulation. To enhance test-time prompt tuning, we propose DynaPrompt, short for dynamic test-time prompt tuning, exploiting relevant data distribution information while reducing error accumulation. Built on an online prompt buffer, DynaPrompt adaptively selects and optimizes the relevant prompts for each test sample during tuning. Specifically, we introduce a dynamic prompt selection strategy based on two metrics: prediction entropy and probability difference. For unseen test data information, we develop dynamic prompt appending, which allows the buffer to append new prompts and delete the inactive ones. By doing so, the prompts are optimized to exploit beneficial information on specific test data, while alleviating error accumulation. Experiments on fourteen datasets demonstrate the effectiveness of dynamic test-time prompt tuning.
Zehao Xiao, Shilin Yan, Jack Hong, Jiayin Cai, Yao Hu 0002, Cheems Wang, Cees Snoek
ICLR1
2025 Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes
abstract
Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentations and (2) quantifying predictive uncertainty to help users identify unreliable regions. In this work, we propose \emph{NPISeg3D}, a novel probabilistic framework that builds upon Neural Processes (NPs) to address these challenges. Specifically, NPISeg3D introduces a hierarchical latent variable structure with scene-specific and object-specific latent variables to enhance few-shot generalization by capturing both global context and object-specific characteristics. Additionally, we design a probabilistic prototype modulator that adaptively modulates click prototypes with object-specific latent variables, improving the model’s ability to capture object-aware context and quantify predictive uncertainty. Experiments on four 3D point cloud datasets demonstrate that NPISeg3D achieves superior segmentation performance with fewer clicks while providing reliable uncertainty estimations.
Jie Liu 0043, Pan Zhou 0002, Zehao Xiao, Wenzhe Yin, Jan-Jakob Sonke, Efstratios Gavves
ICML3
2025 GeneralizeFormer: Layer-Adaptive Model Generation Across Test-Time Distribution Shifts
abstract
We consider the problem of test-time domain generalization, where a model is trained on several source domains and adjusted on target domains never seen during training. Different from the common methods that fine-tune the model or adjust the classifier parameters online, we propose to generate multiple layer parameters on the fly during inference by a lightweight meta-learned transformer, which we call GeneralizeFormer: The layer-wise parameters are generated per target batch without fine-tuning or online adjustment. By doing so, our method is more effective in dynamic scenarios with multiple target distributions and also avoids forgetting valuable source distribution characteristics. Moreover, by considering layer-wise gradients, the proposed method adapts itself to various distribution shifts. To reduce the computational and time cost, we fix the convolutional parameters while only generating parameters of the Batch Normalization layers and the linear classifier. Experiments on six widely used domain generalization datasets demonstrate the benefits and abilities of the proposed method to efficiently handle various distribution shifts, generalize in dynamic scenarios, and avoid forgetting. Our code is available: https://github.com/ambekarsameer96/generalizeformer
Sameer Ambekar, Zehao Xiao, Xiantong Zhen, Cees Snoek
WACV2
2024 Any-Shift Prompting for Generalization Over Distributions
abstract
Image-language models with prompt learning have shown remarkable advances in numerous downstream vision tasks. Nevertheless, conventional prompt learning methods overfit their training distribution and lose the generalization ability on test distributions. To improve generalization across various distribution shifts, we propose any-shift prompting: a general probabilistic inference framework that considers the relationship between training and test distributions during prompt learning. We explicitly connect training and test distributions in the latent space by constructing training and test prompts in a hierarchical architecture. Within this framework, the test prompt exploits the distribution relationships to guide the generalization of the CLIP image-language model from training to any test distribution. To effectively encode the distribution information and their relationships, we further introduce a transformer inference network with a pseudo-shift training mechanism. The network generates the tailored test prompt with both training and test information in a feed forward pass, avoiding extra training costs at test time. Extensive experiments on twenty-three datasets demonstrate the effectiveness of any-shift prompting on the generalization over various distribution shifts.
Zehao Xiao, Mohammad Mahdi Derakhshani, Shengcai Liao, Cees Snoek
CVPR1
2024 GO4Align: Group Optimization for Multi-Task Alignment
abstract
This paper proposes **GO4Align**, a multi-task optimization approach that tackles task imbalance by explicitly aligning the optimization across tasks. To achieve this, we design an adaptive group risk minimization strategy, comprising two techniques in implementation: (i) dynamical group assignment, which clusters similar tasks based on task interactions; (ii) risk-guided group indicators, which exploit consistent task correlations with risk information from previous iterations. Comprehensive experimental results on diverse benchmarks demonstrate our method's performance superiority with even lower computational costs.
Qi Wang 0009, Zehao Xiao, Nanne van Noord, Marcel Worring
NeurIPS3
2024 RI-PCGrad: Optimizing multi-task learning with rescaling and impartial projecting conflict gradients
Fanyun Meng, Zehao Xiao
Appl. Intell.2
2024 Regret analysis of an online majorized semi-proximal ADMM for online composite optimization
Zehao Xiao
J. Glob. Optim.1
2023 Energy-Based Test Sample Adaptation for Domain Generalization
Zehao Xiao, Xiantong Zhen, Shengcai Liao, Cees Snoek
ICLR1
2023 ProtoDiff: Learning to Learn Prototypical Networks by Task-Guided Diffusion
abstract
Prototype-based meta-learning has emerged as a powerful technique for addressing few-shot learning challenges. However, estimating a deterministic prototype using a simple average function from a limited number of examples remains a fragile process. To overcome this limitation, we introduce ProtoDiff, a novel framework that leverages a task-guided diffusion model during the meta-training phase to gradually generate prototypes, thereby providing efficient class representations. Specifically, a set of prototypes is optimized to achieve per-task prototype overfitting, enabling accurately obtaining the overfitted prototypes for individual tasks. Furthermore, we introduce a task-guided diffusion process within the prototype space, enabling the meta-learning of a generative process that transitions from a vanilla prototype to an overfitted prototype. ProtoDiff gradually generates task-specific prototypes from random noise during the meta-test stage, conditioned on the limited samples available for the new task. Furthermore, to expedite training and enhance ProtoDiff's performance, we propose the utilization of residual prototype learning, which leverages the sparsity of the residual prototype. We conduct thorough ablation studies to demonstrate its ability to accurately capture the underlying prototype distribution and enhance generalization. The new state-of-the-art performance on within-domain, cross-domain, and few-task few-shot classification further substantiates the benefit of ProtoDiff.
Yingjun Du, Zehao Xiao, Shengcai Liao, Cees Snoek
NeurIPS2
2022 Learning to Generalize across Domains on Single Test Samples
Zehao Xiao, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICLR1
2022 Association Graph Learning for Multi-Task Classification with Category Shifts
abstract
In this paper, we focus on multi-task classification, where related classification tasks share the same label space and are learned simultaneously. In particular, we tackle a new setting, which is more realistic than currently addressed in the literature, where categories shift from training to test data. Hence, individual tasks do not contain complete training data for the categories in the test set. To generalize to such test data, it is crucial for individual tasks to leverage knowledge from related tasks. To this end, we propose learning an association graph to transfer knowledge among tasks for missing classes. We construct the association graph with nodes representing tasks, classes and instances, and encode the relationships among the nodes in the edges to guide their mutual knowledge transfer. By message passing on the association graph, our model enhances the categorical information of each instance, making it more discriminative. To avoid spurious correlations between task and class nodes in the graph, we introduce an assignment entropy maximization that encourages each class node to balance its edge weights. This enables all tasks to fully utilize the categorical information from related tasks. An extensive evaluation on three general benchmarks and a medical dataset for skin lesion classification reveals that our method consistently performs better than representative baselines.
Zehao Xiao, Xiantong Zhen, Cees Snoek, Marcel Worring
NeurIPS2
2022 Light weight object detector based on composite attention residual network and boundary location loss
Zehao Xiao, Enzeng Dong, Jigang Tong, Zenghui Wang 0001
Neurocomputing1
2022 Spherical Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) is a highly non-trivial task to generalize from seen to unseen classes. In this paper, we propose spherical zero-shot learning (SZSL) to address the major challenges in ZSL. By decoupling the similarity metric in the spherical embedding space into radius and angle, our SZSL can map classes to hyperspherical surfaces of different radiuses, which greatly increases its flexibility. Specifically, we introduce the spherical alignment on angles to spread classes as uniformly as possible to alleviate the hubness problem and simultaneously preserve the inter-class semantic structure to make the alignment more reasonable. We also introduce the spherical calibration with a minimum entropy based regularizer by adopting a larger radius for unseen classes than seen classes to reduce the prediction bias. Extensive experiments on five middle-scale benchmarks and large-scale ImageNet dataset demonstrate that the proposed approach consistently achieves superior performance for the traditional and generalized settings of ZSL.
Zehao Xiao, Xiantong Zhen, Lei Zhang 0093
IEEE Trans. Circuits Syst. Video Technol.2
2021 A Bit More Bayesian: Domain-Invariant Learning with Uncertainty
abstract
Domain generalization is challenging due to the domain shift and the uncertainty caused by the inaccessibility of target domain data. In this paper, we address both challenges with a probabilistic framework based on variational Bayesian inference, by incorporating uncertainty into neural network weights. We couple domain invariance in a probabilistic formula with the variational Bayesian inference. This enables us to explore domain-invariant learning in a principled way. Specifically, we derive domain-invariant representations and classifiers, which are jointly established in a two-layer Bayesian neural network. We empirically demonstrate the effectiveness of our proposal on four widely used cross-domain visual recognition benchmarks. Ablation studies validate the synergistic benefits of our Bayesian treatment when jointly learning domain-invariant representations and classifiers for domain generalization. Further, our method consistently delivers state-of-the-art mean accuracy on all benchmarks.
Zehao Xiao, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICML1
2020 Heterogenous output regression network for direct face alignment
Xiantong Zhen, Mengyang Yu, Zehao Xiao, Lei Zhang 0093, Ling Shao 0001
Pattern Recognit.3
2019 Crowd Counting and Density Estimation by Trellis Encoder-Decoder Networks
abstract
Crowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-fold. First, we develop a new trellis architecture that incorporates multiple decoding paths to hierarchically aggregate features at different encoding stages, which improves the representative capability of convolutional features for large variations in objects. Second, we employ dense skip connections interleaved across paths to facilitate sufficient multi-scale feature fusions, which also helps TEDnet to absorb the supervision information. Third, we propose a new combinatorial loss to enforce similarities in local coherence and spatial correlation between maps. By distributedly imposing this combinatorial loss on intermediate outputs, TEDnet can improve the back-propagation process and alleviate the gradient vanishing problem. Finally, on four widely-used benchmarks, our TEDnet achieves the best overall performance in terms of both density map quality and counting accuracy, with an improvement up to 14% in MAE metric. These results validate the effectiveness of TEDnet for crowd counting.
Zehao Xiao, Baochang Zhang 0001, Xiantong Zhen, Xianbin Cao 0001, David S. Doermann, Ling Shao 0001
CVPR2
2019 Relational Attention Network for Crowd Counting
abstract
Crowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision. Density estimation is a popular strategy for crowd counting, where conventional density estimation methods perform pixel-wise regression without explicitly accounting the interdependence of pixels. As a result, independent pixel-wise predictions can be noisy and inconsistent. In order to address such an issue, we propose a Relational Attention Network (RANet) with a self-attention mechanism for capturing interdependence of pixels. The RANet enhances the self-attention mechanism by accounting both short-range and long-range interdependence of pixels, where we respectively denote these implementations as local self-attention (LSA) and global self-attention (GSA). We further introduce a relation module to fuse LSA and GSA to achieve more informative aggregated feature representations. We conduct extensive experiments on four public datasets, including ShanghaiTech A, ShanghaiTech B, UCF-CC-50 and UCF-QNRF. Experimental results on all datasets suggest RANet consistently reduces estimation errors and surpasses the state-of-the-art approaches by large margins.
Zehao Xiao, Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001
ICCV3
2019 ASiam: adaptive Siamese regression tracking with adversarial template generation and motion-based failure recovery
abstract
Object tracking is challenged by the varying appearances of targets and the real‐time requirement. Siamese regression trackers, being one of the most popular tracking paradigms, excel in efficiency but suffer at adaptability to cope with appearance variations. To improve their adaptability, the authors propose a new adaptive Siamese (ASiam) tracker, which integrates a novel adversarial template generation module and a motion‐based failure recovery module. The template generation module exploits the temporal coherence and evolution of target appearance variations encoded in preceding tracklets and then generates an adaptive target template online which approximates the varying target in the coming frame. This generation module is optimised via adversarial learning to achieve accurate appearance prediction and sharp template quality. The generated template, together with a search region, are fed into a Siamese tracking backbone to compute an appearance response map via dense similarity computation in a sliding‐window way. At frames where the Siamese tracking fails, the failure recovery module is invoked to perform deep frame differencing motion detection to provide a motion response map. By fusing different response maps, the drifted tracker can be re‐calibrated. Extensive experiments on the OTB2013, OTB2015, and VOT2016 datasets prove the accuracy and efficiency of the proposed tracker.
Zehao Xiao, Baochang Zhang 0001, Xianbin Cao 0001
IET Image Process.2
2019 Attentional Information Fusion Networks for Cross-Scene Power Line Detection
abstract
The power line is one of the most hazardous obstacles for low-altitude aircrafts. As aircrafts usually encounter scenes like never before during the flight, cross-scene power line detection is the key for their flight safety. However, compared to regular object detection tasks, cross-scene power line detection is extremely challenging due to its weak visual appearance and widespread existence. In this letter, we propose a cross-scene power line detection method based on attentional information fusion networks. Specifically, we construct a fully convolutional network with attention and information fusion mechanism for cross-scene detection. The two main modules make full use of the semantic and location information, which enables the model to focus more on power lines rather than the unexpected scenes. To the best of author knowledge, our method establishes the first end-to-end convolutional architecture for pixelwise power line detection. Experimental results have shown that our method outperforms previous methods by large margins for cross-scene power line detection.
Yan Li 0054, Zehao Xiao, Xiantong Zhen, Xianbin Cao 0001
IEEE Geosci. Remote. Sens. Lett.2
2018 In Defense of Single-column Networks for Crowd Counting
Ze Wang 0008, Zehao Xiao, Qiang Qiu 0001, Xiantong Zhen, Xianbin Cao 0001
BMVC2