EDBT 2026 Demo / reviewers in the wild / expert
Xuehao Wang
dblp:272/4397
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRISM: Prior-enhanced Inference for Spatial Transcriptomic Cell Type MappingabstractMOTIVATION: Cell type annotation in spatial transcriptomics (ST) is fundamental for deciphering complex tissue organization and spatially resolved biological processes. Most existing methods perform ST cell type annotation by transferring labels from single-cell RNA-seq (scRNA) data to ST data, but typically rely on weakly constrained representations that neglect structured spatial dependencies and treat marker gene selection as an isolated preprocessing step. This renders them vulnerable to substantial domain gaps as well as platform-specific noise, resulting in unstable predictions and limited biological interpretability. RESULTS: To address these issues, we propose Prior-enhanced Inference for Spatial Transcriptomic Cell Type Mapping (PRISM), a novel three-stage framework integrating biological prior construction, pseudo-label generation, and multi-level ST refinement. First, PRISM constructs a cross-domain biological prior to explicitly extract marker genes to enforce positive biological discriminability. Next, it adopts a prior-enhanced self-training strategy, where scRNA-trained ensembles generate reliable pseudo-label candidates for ST data, serving as a robust anchor for cross-domain adaptation. Finally, the framework consolidates high-quality ensemble predictions selected via metric-guided evaluation, encodes spatial information, and optimizes the model under dual-directional biological constraints. Extensive experiments on eleven ST datasets across six platforms, two species, and multiple tissue contexts validate PRISM. Specifically, on the five labeled benchmarks, PRISM shows strong overall performance under both Accuracy and Macro-F1 evaluation across brain and non-brain tissues. Moreover, under fully label-free settings, PRISM achieves the best overall composite rank across all datasets, demonstrating strong robustness to domain shift and platform heterogeneity. AVAILABILITY AND IMPLEMENTATION: PRISM is available at https://github.com/lilab-ai4s/PRISM and https://doi.org/10.5281/zenodo.20529683. Yiheng Xu, Xuehao Wang, Congcong Ge |
Bioinform. | 2 |
| 2026 | Mamba-guided lightweight capsule network for chemical hazards recognition
Xingrong Li, Fulai Zhang, Xuehao Wang, Donghao Cheng |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Mitigating scale imbalance and conflicting gradients in deep multi-task learning
Yuepeng Jiang, Yunhao Gou, Xuehao Wang |
Frontiers Comput. Sci. | 4 |
| 2025 | Sharpness-Aware Black-Box OptimizationabstractBlack-box optimization algorithms have been widely used in various machine learning problems, including reinforcement learning and prompt fine-tuning. However, directly optimizing the training loss value, as commonly done in existing black-box optimization methods, could lead to suboptimal model quality and generalization performance. To address those problems in black-box optimization, we propose a novel Sharpness-Aware Black-box Optimization (SABO) algorithm, which applies a sharpness-aware minimization strategy to improve the model generalization. Specifically, the proposed SABO method first reparameterizes the objective function by its expectation over a Gaussian distribution. Then it iteratively updates the parameterized distribution by approximated stochastic gradients of the maximum objective value within a small neighborhood around the current solution in the Gaussian distribution space. Theoretically, we prove the convergence rate and generalization bound of the proposed SABO algorithm. Empirically, extensive experiments on the black-box prompt fine-tuning tasks demonstrate the effectiveness of the proposed SABO method in improving model generalization performance. Feiyang Ye 0001, Yueming Lyu, Xuehao Wang, Masashi Sugiyama, Yu Zhang 0006, Ivor W. Tsang |
ICLR | 3 |
| 2025 | HeadMap: Locating and Enhancing Knowledge Circuits in LLMsabstractLarge language models (LLMs), through pretraining on extensive corpora, encompass rich semantic knowledge and exhibit the potential for efficient adaptation to diverse downstream tasks. However, the intrinsic mechanisms underlying LLMs remain unexplored, limiting the efficacy of applying these models to downstream tasks. In this paper, we explore the intrinsic mechanisms of LLMs from the perspective of knowledge circuits. Specifically, considering layer dependencies, we propose a layer-conditioned locating algorithm to identify a series of attention heads, which is a knowledge circuit of some tasks. Experiments demonstrate that simply masking a small portion of attention heads in the knowledge circuit can significantly reduce the model's ability to make correct predictions. This suggests that the knowledge flow within the knowledge circuit plays a critical role when the model makes a correct prediction. Inspired by this observation, we propose a novel parameter-efficient fine-tuning method called HeadMap, which maps the activations of these critical heads in the located knowledge circuit to the residual stream by two linear layers, thus enhancing knowledge flow from the knowledge circuit in the residual stream. Extensive experiments conducted on diverse datasets demonstrate the efficiency and efficacy of the proposed method. Our code is available at https://github.com/XuehaoWangFi/HeadMap. Xuehao Wang, Binghuai Lin |
ICLR | 1 |
| 2025 | MTSAM: Multi-Task Fine-Tuning for Segment Anything ModelabstractThe Segment Anything Model (SAM), with its remarkable zero-shot capability, has the potential to be a foundation model for multi-task learning. However, adopting SAM to multi-task learning faces two challenges: (a) SAM has difficulty generating task-specific outputs with different channel numbers, and (b) how to fine-tune SAM to adapt multiple downstream tasks simultaneously remains unexplored. To address these two challenges, in this paper, we propose the Multi-Task SAM (MTSAM) framework, which enables SAM to work as a foundation model for multi-task learning. MTSAM modifies SAM's architecture by removing the prompt encoder and implementing task-specific no-mask embeddings and mask decoders, enabling the generation of task-specific outputs. Furthermore, we introduce Tensorized low-Rank Adaptation (ToRA) to perform multi-task fine-tuning on SAM. Specifically, ToRA injects an update parameter tensor into each layer of the encoder in SAM and leverages a low-rank tensor decomposition method to incorporate both task-shared and task-specific information.
Extensive experiments conducted on benchmark datasets substantiate the efficacy of MTSAM in enhancing the performance of multi-task learning. Our code is available at https://github.com/XuehaoWangFi/MTSAM. Xuehao Wang, Zhan Zhuang, Feiyang Ye 0001, Yu Zhang 0006 |
ICLR | 1 |
| 2025 | Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link PredictionabstractMessage-passing graph neural networks (MPNNs) and structural features (SFs) are cornerstones for the link prediction task. However, as a common and intuitive mode of understanding, the potential of visual perception has been overlooked in the MPNN community. For the first time, we equip MPNNs with vision structural awareness by proposing an effective framework called Graph Vision Network (GVN), along with a more efficient variant (E-GVN). Extensive empirical results demonstrate that with the proposed frameworks, GVN consistently benefits from the vision enhancement across seven link prediction datasets, including challenging large-scale graphs. Such improvements are compatible with existing state-of-the-art (SOTA) methods and GVNs achieve new SOTA results, thereby underscoring a promising novel direction for link prediction. Yanbin Wei, Xuehao Wang, Zhan Zhuang, Yang Chen 0031, Shuhao Chen, Yulong Zhang 0005, James T. Kwok, Yu Zhang 0006 |
ICML | 2 |
| 2025 | Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank AdaptationabstractLow-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters’ activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter’s marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto. Zhan Zhuang, Xiequn Wang, Yulong Zhang 0005, Qiushi Huang, Shuhao Chen, Xuehao Wang, Yanbin Wei, Yuhe Nie, Kede Ma, Yu Zhang 0006, Ying Wei 0001 |
ICML | 7 |
| 2025 | MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionabstractHuman Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a novel self-supervised framework that enhances interpretability by tokenizing inertial measurement unit signals into semantically meaningful motion primitives and leverages a Transformer architecture to learn rich temporal representations. MoPFormer comprises two stages. The first stage is to partition multi-channel sensor streams into short segments and quantize them into discrete ``motion primitive'' codewords, while the second stage enriches those tokenized sequences through a context-aware embedding module and then processes them with a Transformer encoder. The proposed MoPFormer can be pre-trained using a masked motion-modeling objective that reconstructs missing primitives, enabling it to develop robust representations across diverse sensor configurations. Experiments on six HAR benchmarks demonstrate that MoPFormer not only outperforms state-of-the-art methods but also successfully generalizes across multiple datasets. More importantly, the learned motion primitives significantly enhance both interpretability and cross-dataset performance by capturing fundamental movement patterns that remain consistent across similar activities, regardless of dataset origin. Zhan Zhuang, Xuehao Wang |
NeurIPS | 3 |
| 2025 | Advancing MRI segmentation with CLIP-driven semi-supervised learning and semantic alignment
Kexuan Li, Jingjuan Liu, Xuehao Wang, Yuanbo He, Huadan Xue, Aimin Hao, Shuai Li 0001 |
Neurocomputing | 5 |
| 2025 | SkyEye: Multi-Modal Perception Based Video Stitching for Multi-UAV Surveillance SystemabstractUsing multiple Unmanned Aerial Vehicles (UAVs) in video surveillance greatly enhances real-time monitoring of large areas. However, images captured by UAVs are separate and limited in view, making stitching crucial for a comprehensive perspective. Current methods combine sensor and visual modalities for video stitching but face challenges in robustness and real-time performance. Environmental disturbances increase errors in sensor and visual data, reducing accuracy and stability, while single-frame stitching incurs high computational costs and slow speeds. To address these issues, we propose SkyEye , a real-time video stitching method for UAV surveillance systems based on multi-modal perception. Specifically, we design a mutual verification method to assess the quality of visual and inertial data. Frames with higher confidence are prioritized for precise stitching. To reduce computational costs, we employ a reference frame scheme that reuses the perspective transformation and feature points of the reference frame for subsequent frames. Meanwhile, we design a video stitching framework based on a per-frame parallel unit to speed up the algorithm execution. We have implemented a prototype of SkyEye and carried out extensive evaluations. Experiment results show that SkyEye outperforms state-of-the-art methods, improving speed by 66.1% and accuracy by 28.2%. Kequan Lin, Yanling Bu, Xuehao Wang, Chenyu Ling, Lei Xie 0004, Yafeng Yin 0002, Sanglu Lu |
ACM Trans. Sens. Networks | 3 |
| 2024 | Semi-Supervised Medical Image Segmentation with Cross-View Consistency and Contrastive LearningabstractMedical image segmentation plays a crucial role in many clinical applications. To alleviate the dependency on massive annotations, semi-supervised learning has attracted increasing attention. However, these methods face significant intra-class and inter-class variation and do not fully utilize the critical multi-view information inherent in medical images. This study proposes a novel network, CV-Net, which integrates multi-view information for semi-supervised medical image segmentation. Concretely, the network is based on Mean-Teacher architecture which largely narrows the empirical distribution gap between labeled and unlabeled data. The proposed cross-view consistency regularization module incorporates a dual-branch attention architecture to integrate consistent semantics while focusing on details, enhancing feature extraction capabilities. The proposed bi-semantic contrastive learning module leverages limited labels and explore pseudo-labels to define semantically similar regions, enhancing the representation capacity. Experiments conducted on two datasets demonstrated the effectiveness of the proposed network. CV-Net showed significant improvements across four metrics, evident with both 5% and 10% labeled data. Specifically, with 5% labeled data, the mean Dice increased by 1.37%. Compared with previous state-of-the-art methods, CV-Net achieved the best results, notably reducing both intra-class and inter-class errors. Kexuan Li, Jingjuan Liu, Xuehao Wang, Huadan Xue, Aimin Hao, Shuai Li 0001 |
BIBM | 5 |
| 2024 | Adaptive Stochastic Gradient Algorithm for Black-box Multi-Objective LearningabstractMulti-objective optimization (MOO) has become an influential framework for various machine learning problems, including reinforcement learning and multi-task learning. In this paper, we study the black-box multi-objective optimization problem, where we aim to optimize multiple potentially conflicting objectives with function queries only. To address this challenging problem and find a Pareto optimal solution or the Pareto stationary solution,
we propose a novel adaptive stochastic gradient algorithm for black-box MOO, called ASMG.
Specifically, we use the stochastic gradient approximation method to obtain the gradient for the distribution parameters of the Gaussian smoothed MOO with function queries only. Subsequently, an adaptive weight is employed to aggregate all stochastic gradients to optimize all objective functions effectively.
Theoretically, we explicitly provide the connection between the original MOO problem and the corresponding Gaussian smoothed MOO problem and prove the convergence rate for the proposed ASMG algorithm in both convex and non-convex scenarios.
Empirically, the proposed ASMG method achieves competitive performance on multiple numerical benchmark problems. Additionally, the state-of-the-art performance on the black-box multi-task learning problem demonstrates the effectiveness of the proposed ASMG method. Feiyang Ye 0001, Yueming Lyu, Xuehao Wang, Yu Zhang 0006, Ivor W. Tsang |
ICLR | 3 |
| 2024 | Time-Varying LoRA: Towards Effective Cross-Domain Fine-Tuning of Diffusion ModelsabstractLarge-scale diffusion models are adept at generating high-fidelity images and facilitating image editing and interpolation. However, they have limitations when tasked with generating images in dynamic, evolving domains. In this paper, we introduce Terra, a novel Time-varying low-rank adapter that offers a fine-tuning framework specifically tailored for domain flow generation. The key innovation of Terra lies in its construction of a continuous parameter manifold through a time variable, with its expressive power analyzed theoretically. This framework not only enables interpolation of image content and style but also offers a generation-based approach to address the domain shift problems in unsupervised domain adaptation and domain generalization. Specifically, Terra transforms images from the source domain to the target domain and generates interpolated domains with various styles to bridge the gap between domains and enhance the model generalization, respectively. We conduct extensive experiments on various benchmark datasets, empirically demonstrate the effectiveness of Terra. Our source code is publicly available on https://github.com/zwebzone/terra. Zhan Zhuang, Yulong Zhang 0005, Xuehao Wang, Jiangang Lu, Ying Wei 0001, Yu Zhang 0006 |
NeurIPS | 3 |
| 2024 | Enhancing Sharpness-Aware Minimization by Learning Perturbation Radius
Xuehao Wang, Weisen Jiang, Yu Zhang 0006 |
ECML/PKDD (2) | 1 |
| 2023 | Multi-Task Learning via Time-Aware Neural ODEabstractMulti-Task Learning (MTL) is a well-established paradigm for learning shared models for a diverse set of tasks. Moreover, MTL improves data efficiency by jointly training all tasks simultaneously. However, directly optimizing the losses of all the tasks may lead to imbalanced performance on all the tasks due to the competition among tasks for the shared parameters in MTL models. Many MTL methods try to mitigate this problem by dynamically weighting task losses or manipulating task gradients. Different from existing studies, in this paper, we propose a Neural Ordinal diffeRential equation based Multi-tAsk Learning (NORMAL) method to alleviate this issue by modeling task-specific feature transformations from the perspective of dynamic flows built on the Neural Ordinary Differential Equation (NODE). Specifically, the proposed NORMAL model designs a time-aware neural ODE block to learn task-specific time information, which determines task positions of feature transformations in the dynamic flow, in NODE automatically via gradient descent methods. In this way, the proposed NORMAL model handles the problem of competing shared parameters by learning task positions. Moreover, the learned task positions can be used to measure the relevance among different tasks. Extensive experiments show that the proposed NORMAL model outperforms state-of-the-art MTL models. Feiyang Ye 0001, Xuehao Wang, Yu Zhang 0006, Ivor W. Tsang |
IJCAI | 2 |
| 2023 | Modality Profile - A New Critical Aspect to be Considered When Generating RGB-D Salient Object Detection Training SetabstractIt is widely acknowledged that selecting appropriate training data is crucial for obtaining good results in real-world testing, more so than utilizing complex network architectures. However, in the field of RGB-D SOD research, researchers have primarily focused on enhancing network architectures and have given less consideration to the choice of training and testing datasets, which may not translate well in practical applications. This paper aims to address an existing issue - how can we automatically generate a data-driven RGB-D SOD training dataset? We propose that in addition to scene similarity, the concept of "modality profile'' should be taken into account. The term "modality profile'' refers to the complementary status of modalities within a given dataset. A training dataset with a modality profile similar to the test dataset can significantly improve performance. To address this, we present a viable solution for automatically generating a training dataset with any desired modality profile in a weakly supervised manner. Our method also provides high-quality pseudo-GTs for all RGB-D images obtained from the web, making it suitable for training RGB-D SOD models. Extensive quantitative evaluations demonstrate the significance of the proposed "modality profile'' and confirm the superiority of the newly constructed training set guided by our "modality profile''. All codes, datasets, and results are available at this link. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
ACM Multimedia | 1 |
| 2021 | Depth quality-aware selective saliency fusion for RGB-D image salient object detection
Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
Neurocomputing | 1 |
| 2021 | Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object DetectionabstractExisting RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Yuming Fang 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |