VLDB 2026 Research / reviewers in the wild / expert
Haiying Xia
dblp:31/283
· DBLP profile ↗
45ranked-venue papers
18as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion RecognitionabstractMusic emotion recognition is a key task in symbolic music understanding (SMER). Recent approaches have shown promising results by fine-tuning large-scale pre-trained models (e.g., MIDIBERT, a benchmark in symbolic music understanding) to map musical semantics to emotional labels. While these models effectively capture distributional musical semantics, they often overlook tonal structures, particularly musical modes, which play a critical role in emotional perception according to music psychology. In this paper, we investigate the representational capacity of MIDIBERT and identify its limitations in capturing mode-emotion associations. To address this issue, we propose a Mode-Guided Enhancement (MoGE) strategy that incorporates psychological insights on mode into the model. Specifically, we first conduct a mode augmentation analysis, which reveals that MIDIBERT fails to effectively encode emotion-mode correlations. Motivated by this observation, we further identify the MIDIBERT layer that shows the weakest emotion relevance and introduce a Mode-guided Feature-wise linear modulation injection (MoFi) framework to inject explicit mode features, thereby enhancing the model's capability in emotional representation and inference. Extensive experiments on the EMOPIA and VGMIDI datasets demonstrate that our mode injection strategy significantly improves SMER performance, achieving accuracies of 75.2% and 59.1%, respectively. These results validate the effectiveness of mode-guided modeling in symbolic music emotion recognition. Haiying Xia, Yumei Tan, Shuxiang Song 0001 |
AAAI | 1 |
| 2026 | Motion-Aware Object Tracking via Motion and Geometry-Aware CuesabstractUnderstanding motion is essential for visual object tracking, especially in complex and dynamic scenarios. Yet, many existing methods rely on simplistic strategies such as template updates or temporal feature propagation, often overlooking the deeper modeling of motion information. To mitigate this limitation, we introduce a motion-aware spatio-temporal framework that enhances motion perception by explicitly matching motion patterns and modeling inter-frame motion relationships. Central to our design is a motion pattern dictionary, which encodes a diverse set of representative motion cues as learnable features. During tracking, features from the search region interact with the dictionary to retrieve the most relevant motion patterns, allowing the model to adapt to the current motion state. A dedicated decoder further incorporates temporal correlations to refine motion awareness. To complement motion modeling, we embed geometric cues into the search region features, which strengthens spatial perception, reduces ambiguity under occlusion, and improves foreground-background separation. Extensive evaluations on seven challenging benchmarks demonstrate the effectiveness of our design. In particular, MoDTrack_384 surpasses recent SOTA trackers on LaSOT by 1.2% in AUC, highlighting the benefits of motion pattern modeling and geometry-guided enhancement in mitigating tracking drift. Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Yufei Tan, Haiying Xia, Shuxiang Song 0001 |
AAAI | 6 |
| 2026 | AVSCNet: A dual-branch network for synchronization detection and content consistency learning in audio-video forgery detection
Guangwei Zhu, Haiying Xia, Shuxiang Song 0001 |
Neurocomputing | 3 |
| 2026 | Dynamic cross-instance context mining for multimodal sentiment analysis
Haiying Xia, Youyong Cheng, Yumei Tan, Shuxiang Song 0001 |
Inf. Process. Manag. | 1 |
| 2026 | HEL-Net: Heterogeneous Ensemble Learning for comprehensive diabetic retinopathy multi-lesion segmentation via Mamba-UNet
Lingyu Wu, Haiying Xia, Shuxiang Song 0001 |
Image Vis. Comput. | 2 |
| 2026 | Global-local co-regularization network for facial action unit detection
Yumei Tan, Haiying Xia, Shuxiang Song 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | Mamba-Driven Diffusion Model for Salient Object Detection in Optical Remote Sensing ImagesabstractExisting Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) methods mainly rely on a semantic segmentation paradigm, which relies on pixel-wise probabilities, leading to overconfident mispredictions. In contrast, the random sampling process of the diffusion model allows multiple possible predictions to be drawn from the mask distribution, effectively alleviating this problem. However, existing diffusion models mainly use Transformers as conditional feature extraction networks. Although they are good at global modeling, they have limited ability to handle long-range dependencies due to computational complexity. To overcome these challenges, we introduce MambaDif, an innovative diffusion model architecture based on Mamba. Specifically, we regard ORSI-SOD as a conditional mask generation task leveraging the diffusion model and achieving target distribution matching by adding noise to the mask and iteratively denoising it to match the target distribution. Then, we adopt Mamba to extract global features, efficiently process long sequences, and capture global contextual information with linear complexity. In addition, we introduce the global-local feature collaborative completion module (GLM), which combines the ability of convolutional layers to extract local features with the advantage of Mamba in capturing long-range dependencies, thereby achieving excellent denoising performance. Extensive experiments show that MambaDif outperforms SOTA methods in eight evaluation metrics on two standard datasets (EORSSD and ORSSD). We also report the generalization performance of the model on the challenging ORSI-4199 to evaluate its robustness. Bineng Zhong 0001, Qihua Liang, Yufei Tan, Haiying Xia, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Low-Rank Adaptive Structural Priors for Generalizable Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR), a serious complication of diabetes, is one of the primary causes of vision loss in retinal vascular diseases. While deep learning has been widely used for DR grading, their performance declines significantly when applied to data outside the training distribution due to domain shifts. Domain generalization (DG) addresses this challenge, yet most existing DG methods neglect lesion-specific features, limiting diagnostic accuracy.In this paper, we propose a novel approach that enhances existing DG methods by incorporating structural priors, inspired by the observation that DR grading is heavily dependent on vessel and lesion structures. We introduce Low-rank Adaptive Structural Priors (LoASP), a plug-and-play framework designed for seamless integration with existing DG models. LoASP improves generalization by learning adaptive structural representations that are finely tuned to the complexities of DR diagnosis.Extensive experiments on eight diverse datasets validate its effectiveness in both single-source and multisource domain scenarios. Visualizations show the learned priors align with vessel/lesion structures, enhancing interpretability and diagnostic relevance. Yunxuan Wang, Yumei Tan, Haiying Xia |
IJCNN | 5 |
| 2025 | Explicit Context Reasoning with Supervision for Visual TrackingabstractContextual reasoning with constraints is crucial for enhancing temporal consistency in cross-frame modeling for visual tracking. However, mainstream tracking algorithms typically associate context by merely stacking historical information without explicitly supervising the association process, making it difficult to effectively model the target's evolving dynamics. To alleviate this problem, we propose RSTrack, which explicitly models and supervises context reasoning via three core mechanisms. 1) Context Reasoning Mechanism : Constructs a target state reasoning pipeline, converting unconstrained contextual associations into a temporal reasoning process that predicts the current representation based on historical target states, thereby enhancing temporal consistency. 2) Forward Supervision Strategy : Utilizes true target features as anchors to constrain the reasoning pipeline, guiding the predicted output toward the true target distribution and suppressing drift in the context reasoning process. 3) Efficient State Modeling : Employs a compression-reconstruction mechanism to extract the core features of the target, removing redundant information across frames and preventing ineffective contextual associations. These three mechanisms collaborate to effectively alleviate the issue of contextual association divergence in traditional temporal modeling. Experimental results show that RSTrack achieves state-of-the-art performance on multiple benchmark datasets while maintaining real-time running speeds. Our code is available at https://github.com/GXNU-ZhongLab/RSTrack. Fansheng Zeng, Bineng Zhong 0001, Haiying Xia, Yufei Tan, Xiantao Hu, Liangtao Shi, Shuxiang Song 0001 |
ACM Multimedia | 3 |
| 2025 | TEMSA:Text enhanced modal representation learning for multimodal sentiment analysis
Shuxiang Song 0001, Yumei Tan, Haiying Xia |
Comput. Vis. Image Underst. | 4 |
| 2025 | Unifying Motion and Appearance Cues for Visual Tracking via Shared QueriesabstractThe rich motion and appearance cues between consecutive frames are crucial for robust visual tracking. However, most existing tracking methods are still limited in designing different components to separately employ corresponding cues and even ignore one of them. This makes them difficult to maintain effective interaction between different cues, thus hindering the models from fostering a comprehensive understanding of the target objects. To address these issues, we propose a unified spatio-temporal cues learning framework (named USCLTrack) that comprehensively mines the variation patterns of targets between consecutive frames in complex video streams. Specifically, USCLTrack firstly aggregates motion and appearance cues into shared queries to provide the bridge of interaction between both cues. Then, it directly generates object locations on the condition of these shared queries in an autoregressive manner, unifying different cues to guide future inferences. To effectively learn multiple spatio-temporal cues aggregated in the shared queries, we develop a spatio-temporal attention mechanism. This mechanism integrates motion cues with appearance cues according to the time steps for ensuring temporal consistency. Moreover, it concurrently captures motion trends and appearance changes to facilitate the understanding of the target objects. Extensive experiments on eight popular tracking benchmarks validate the effectiveness of the proposed USCLTrack. Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Haiying Xia, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Robust consistency learning for facial expression recognition under label noise
Yumei Tan, Haiying Xia, Shuxiang Song 0001 |
Vis. Comput. | 2 |
| 2024 | DTIL-Net: Dual-Task Interactive Learning Network for Automated Grading of Diabetic Retinopathy and Macular Edema
Yumei Tan, Shuxiang Song 0001, Haiying Xia |
PRCV (14) | 4 |
| 2024 | CFMISA: Cross-Modal Fusion of Modal Invariant and Specific Representations for Multimodal Sentiment Analysis
Haiying Xia, Yumei Tan |
PRCV (3) | 1 |
| 2024 | Joint pyramidal perceptual attention and hierarchical consistency constraint for gaze estimation
Haiying Xia, Zhuolin Gong, Yumei Tan, Shuxiang Song 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | Dual-consistency constraints network for noisy facial expression recognition
Haiying Xia, Chunhai Su, Shuxiang Song 0001, Yumei Tan |
Image Vis. Comput. | 1 |
| 2024 | Learning informative and discriminative semantic features for robust facial expression recognition
Yumei Tan, Haiying Xia, Shuxiang Song 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Hard semantic mask strategy for automatic facial action unit recognition with teacher-student model
Zichen Liang, Haiying Xia, Yumei Tan, Shuxiang Song 0001 |
Multim. Syst. | 2 |
| 2024 | Harmonious Mutual Learning for Facial Emotion RecognitionabstractAbstract Facial emotion recognition in the wild is an important task in computer vision, but it still remains challenging since the influence of backgrounds, occlusions and illumination variations in facial images, as well as the ambiguity of expressions. This paper proposes a harmonious mutual learning framework for emotion recognition, mainly through utilizing attention mechanisms and probability distributions without utilizing additional information. Specifically, this paper builds an architecture with two emotion recognition networks and makes progressive cooperation and interaction between them. We first integrate self-mutual attention module into the backbone to learn discriminative features against the influence from emotion-irrelevant facial information. In this process, we deploy spatial attention module and convolutional block attention module for the two networks respectively, guiding to enhanced and supplementary learning of attention. Further, in the classification head, we propose to learn the latent ground-truth emotion probability distributions using softmax function with temperature to characterize the expression ambiguity. On this basis, a probability distribution distillation learning module is constructed to perform class semantic interaction using bi-directional KL loss, allowing mutual calibration for the two networks. Experimental results on three public datasets show the superiority of the proposed method compared to state-of-the-art ones. Yanling Gan, Luhui Xu, Haiying Xia |
Neural Process. Lett. | 3 |
| 2024 | Learning from feature and label spaces' bias for uncertainty-adaptive facial emotion recognition
Luhui Xu, Yanling Gan, Haiying Xia |
Pattern Recognit. Lett. | 3 |
| 2024 | Feature fusion of multi-granularity and multi-scale for facial expression recognition
Haiying Xia, Lidan Lu, Shuxiang Song 0001 |
Vis. Comput. | 1 |
| 2023 | GroupSeg: An Efficient Grouping Transformer Network for Polyp SegmentationabstractPrecise polyp segmentation is extremely challenging due to the diverse appearances and shaps’s of polyps along with severe light imbalance. To this end, we propose a novel network architecture, GroupSeg, which aims to learn robust representations by fully exploring global contextual representations and texture features of polyps through grouping segmentation to improve polyp segmentation performance. Specifically, GroupSeg groups features from the transformer encoder into progressively larger segments of different shapes, making full use of global contextual information to generate segmentation maps from fine to coarse, and thus can automatically adapt to polyps of different shapes and sizes, and is highly adaptive to the distribution of different polyp datasets. To further mine texture features to improve the performance of polyp segmentation, we introduce a Grouping Feature Aggregation Module (GFA), which adaptively mines the clues of local pixels and grouping feature segments, making the grouping features more accurate.GroupSeg demonstrates superior performance compared to state-of-the-art methods on ETIS, Endoscene, and ColonDB datasets, while high competitiveness on ClinicDB. It offers propsective outcomes in polyp segmentation. Mingwen Zhang, Haiying Xia, Yumei Tan |
BIBM | 2 |
| 2023 | ST-VQA: shrinkage transformer with accurate alignment for visual question answering
Haiying Xia, Richeng Lan, Hai-Sheng Li 0001, Shuxiang Song 0001 |
Appl. Intell. | 1 |
| 2023 | A multi-scale gated network for retinal hemorrhage detection
Haiying Xia, Zengyan Rao, Zuoshan Zhou |
Appl. Intell. | 1 |
| 2023 | Three-dimensional quantum wavelet transforms
Hai-Sheng Li 0001, Guiqiong Li, Haiying Xia |
Frontiers Comput. Sci. | 3 |
| 2023 | RT-Net: Region-Enhanced Attention Transformer Network for Polyp Segmentation
Yilin Qin, Haiying Xia, Shuxiang Song 0001 |
Neural Process. Lett. | 2 |
| 2022 | HT-Net: hierarchical context-attention transformer network for medical ct image segmentation
Haiying Xia, Yumei Tan, Hai-Sheng Li 0001, Shuxiang Song 0001 |
Appl. Intell. | 2 |
| 2022 | MC-Net: multi-scale context-attention network for medical CT image segmentation
Haiying Xia, Hai-Sheng Li 0001, Shuxiang Song 0001 |
Appl. Intell. | 1 |
| 2022 | MFC-Net: Multi-scale fusion coding network for Image Deblurring
Haiying Xia, Yumei Tan, Shuxiang Song 0001 |
Appl. Intell. | 1 |
| 2022 | Collaborative learning network for head pose estimation
Haiying Xia, Luhui Xu, Yanling Gan |
Image Vis. Comput. | 1 |
| 2022 | HRNet: A hierarchical recurrent convolution neural network for retinal vessel segmentation
Haiying Xia, Lingyu Wu, Hai-Sheng Li 0001, Shuxiang Song 0001 |
Multim. Tools Appl. | 1 |
| 2021 | A multi-scale segmentation-to-classification network for tiny microaneurysm detection in fundus images
Haiying Xia, Shuxiang Song 0001, Hai-Sheng Li 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Combination of multi-scale and residual learning in deep CNN for image denoisingabstractTo better restore a clean image from a noise observation under high noise levels, the authors propose an image denoising network based on the combination of multi‐scale and residual learning. Instead of using filters with different large sizes in traditional multi‐scale schemes, they arrange multi‐layer convolutions with the filters of the same size to speed up the model. Some dilated convolutions of different rates are combined with the common convolutions to enrich the extracted features in multi‐layer convolutions. Furthermore, they cascade the multi‐layer convolutions with residual blocks to improve the performance of image denoising. Their extensive evaluations on several challenging datasets demonstrate that the proposed model outperforms the state‐of‐art methods under all different noise levels in terms of peak signal‐to‐noise ratio, and the visual effects achieved by the proposed model are also better than the competing methods. Haiying Xia, Fuyu Zhu, Hai-Sheng Li 0001, Shuxiang Song 0001, Xiangwei Mou |
IET Image Process. | 1 |
| 2020 | Style transfer for QR code
Hai-Sheng Li 0001, Fan Xue, Haiying Xia |
Multim. Tools Appl. | 3 |
| 2020 | Md-Net: Multi-scale Dilated Convolution Network for CT Images Segmentation
Haiying Xia, Weifan Sun, Shuxiang Song 0001, Xiangwei Mou |
Neural Process. Lett. | 1 |
| 2019 | Quantum multi-level wavelet transforms
Hai-Sheng Li 0001, Haiying Xia, Shuxiang Song 0001 |
Inf. Sci. | 3 |
| 2019 | Quantum vision representations and multi-dimensional quantum transforms
Hai-Sheng Li 0001, Shuxiang Song 0001, Huiling Peng, Haiying Xia |
Inf. Sci. | 5 |
| 2019 | Fast template matching based on deformable best-buddies similarity measure
Haiying Xia, Wenxian Zhao, Frank Jiang 0001, Hai-Sheng Li 0001, Jing Xin |
Multim. Tools Appl. | 1 |
| 2018 | Retinal Vessel Segmentation via A Coarse-to-fine Convolutional Neural Network
Haiying Xia, Ruibin Zhuge, Hai-Sheng Li 0001 |
BIBM | 1 |
| 2018 | Fast Single Image De-raining via a Weighted Residual Network
Ruibin Zhuge, Haiying Xia, Hai-Sheng Li 0001, Shuxiang Song 0001 |
ICONIP (6) | 2 |
| 2018 | Ensemble One-Dimensional Convolution Neural Networks for Skeleton-Based Action RecognitionabstractThis letter proposes an ensemble neural network (Ensem-NN) for skeleton-based action recognition. The Ensem-NN is introduced based on the idea of ensemble learning, “two heads are better than one.” According to the property of skeleton sequences, we design one-dimensional convolution neural network with residual structure asBase-Net. From entirety to local, from focus to motion, we designed four different subnets based on theBase-Netto extract diverse features. The first subnet is aTwo-stream Entirety Net, which performs on the entirety skeleton and explores both temporal and spatial features. The second is aBody-part Net, which can extract fine-grained spatial and temporal features. The third is anAttention Net, in which a channel-wised attention mechanism can learn important frames and feature channels.Frame-difference Net, as the fourth subnet, aims at exploring motion features. Finally, the four subnets are fused as one ensemble network. Experimental results show that the proposed Ensem-NN performs better than state-of-the-art methods on three widely used datasets. Yangyang Xu 0004, Jun Cheng 0002, Lei Wang 0018, Haiying Xia, Feng Liu 0013, Dapeng Tao |
IEEE Signal Process. Lett. | 4 |
| 2017 | A new binary hybrid particle swarm optimization with wavelet mutation
Frank Jiang 0001, Haiying Xia, Quang-Anh Tran, Quang Minh Ha, Nhat-Quang Tran, Jiankun Hu |
Knowl. Based Syst. | 2 |
| 2016 | Robust retinal vessel segmentation via clustering-based patch mapping functionsabstractRobust vessel segmentation of fundus images is of great interest for better diagnosis of many diseases like diabetic retinopathy, retinopathy of prematurity, vein occlusions and so on. In this paper, we propose a novel example-based vessel segmentation method, based on learning the mapping relationship between fundus images and their corresponding ground truths. Firstly, the training images and their corresponding ground truths are divided into patches and clustered. Secondly, the mapping functions for each cluster are computed in a simple and efficient way from the training patches to their manual segmentation patches. Finally, Vessel segmentation are reconstructed by the simple mapping functions. Experimental results show that our method is efficient and can achieve competitive performance for vessel segmentation problems. Haiying Xia, Shuaifei Deng, Minqi Li, Frank Jiang 0001 |
BIBM | 1 |
| 2016 | An adaptive vehicular epidemic routing method based on attractor selection model
Daxin Tian, Jianshan Zhou, Haiying Xia |
Ad Hoc Networks | 5 |
| 2015 | A Dynamic and Self-Adaptive Network Selection Method for Multimode Communications in Heterogeneous Vehicular TelematicsabstractWith the increasing demands for vehicle-to-vehicle and vehicle-to-infrastructure communications in intelligent transportation systems, new generation of vehicular telematics inevitably depends on the cooperation of heterogeneous wireless networks. In heterogeneous vehicular telematics, the network selection is an important step to the realization of multimode communications that use multiple access technologies and multiple radios in a collaborative manner. This paper presents an innovative network selection solution for the fundamental technological requirement of multimode communications in heterogeneous vehicular telematics. To guarantee the QoS satisfaction of multiple mobile users and the efficient utilization and fair allocation of heterogeneous network resources in a global sense, a dynamic and self-adaptive method for network selection is proposed. It is biologically inspired by the cellular gene network, which enables terminals to dynamically select an appropriate access network according to the variety of QoS requirements and to the dynamic conditions of various available networks. The experimental results prove the effectiveness of the bioinspired scheme and confirm that the proposed network selection method provides better global performance when compared with the utility function method with greedy optimization. Daxin Tian, Jianshan Zhou, Yingrong Lu, Haiying Xia, Zhenguo Yi |
IEEE Trans. Intell. Transp. Syst. | 5 |