EDBT 2026 Demo / reviewers in the wild / expert
Xuemei Jia
dblp:125/5093
· DBLP profile ↗
18ranked-venue papers
3as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order TransferabstractAction recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition robustness. To address this, we introduce a multi-view formulation tailored for UAV altitudes and empirically uncover a distinctive partial order among views, where recognition accuracy consistently declines as altitude increases. This key observation motivates the proposed Aero Partial Order Guided Network (Aerorder), which explicitly models and exploits the hierarchical structure of UAV views to enhance cross-altitude action recognition. Aerorder comprises three main components: (1) a View Partition (VP) module that groups views by altitude using the head-to-body ratio; (2) an Order-aware Feature Decoupling (OFD) module that disentangles action-relevant and view-specific representations under partial order guidance; and (3) an Action Partial Order Guide (APOG) that progressively transfers knowledge from easier (low-altitude) to harder (high-altitude) views. Extensive experiments on Drone-Action, MOD20, and UAV validate the superiority of Aerorder, achieving consistent improvements over state-of-the-art methods, up to 4.7% and 1.3% gains on Drone-Action and MOD20, respectively. Wenxuan Liu 0008, Zhuo Zhou, Xuemei Jia, Siyuan Yang 0001, Wenxin Huang, Xian Zhong, Chia-Wen Lin |
AAAI | 3 |
| 2026 | A Multiagent Reasoning Framework for Classical Chinese Question Answering With Large Language ModelsabstractBackground Understanding Classical Chinese remains a major challenge in Chinese education, especially in the National College Entrance Examination (NCEE). Although large language models (LLMs) exhibit strong reasoning capabilities, their performance on exam‐style Classical Chinese questions still suffers from instability and limited accuracy. Methods We propose a multiagent reasoning framework based on LLMs for Classical Chinese question answering. For each question type, standardized reasoning procedures are defined, and specialized agents are trained for subtasks including word interpretation, grammatical analysis, translation, and semantic summarization. A two‐round reasoning mechanism, consisting of an initial response followed by refinement using standard answers, is introduced to enhance consistency and robustness. Results Experiments on Gaokao‐style Classical Chinese questions demonstrate that the proposed framework achieves higher accuracy and greater reasoning stability than single‐agent systems and general‐purpose LLMs. In objective tasks, it outperforms strong Chinese‐oriented models such as Qwen‐Max and Baichuan‐4 by up to 6.8%. Conclusions The proposed multiagent framework improves both the interpretability and reliability of LLM‐based Classical Chinese understanding. It shows strong potential for applications in intelligent tutoring systems, curriculum support, and cognitive modeling of human‐like reasoning in educational contexts. Bin Nong, Xuemei Jia, Hongmeng Chen, Yulin Deng |
Int. J. Intell. Syst. | 3 |
| 2026 | Robust mixed-degradation person Re-identification via structural consistency distillation
Wenxin Huang, Wenxuan Liu 0008, Xuemei Jia, Xian Zhong |
Pattern Recognit. | 4 |
| 2025 | Balancing Privacy and Performance: A Many-in-One Approach for Image AnonymizationabstractThe effective utilization of data through Deep Neural Networks (DNNs) has profoundly influenced various aspects of society. The growing demand for high-quality, particularly personalized, data has spurred research efforts to prevent data leakage and protect privacy in recent years. Early privacy-preserving methods primarily relied on instance-wise modifications, such as erasing or obfuscating essential features for de-identification. However, this approach highlights an inherent trade-off: minimal modification offers insufficient privacy protection, while excessive modification significantly degrades task performance. In this paper, we propose a novel Recombining for Obfuscation (FRO) approach to address this trade-off. Unlike existing methods that generate one anonymized instance by perturbing the original data on a one-to-one basis, our FRO approach generates an anonymized instance by reassembling mixed ID-related features from multiple original data sources on a many-in-one basis. Instead of introducing additional noise for de-identification, our approach leverages the existing non-polluted features from other instances to anonymize data. Extensive experiments on identity identification tasks demonstrate that FRO outperforms previous state-of-the-art methods, not only in utility performance but also in visual anonymization. Xuemei Jia, Jiawei Du 0002, Hui Wei 0004, Ruinian Xue, Zheng Wang 0007, Hongyuan Zhu 0002, Jun Chen 0001 |
AAAI | 1 |
| 2025 | Dynamic and static mutual fitting for action recognition
Wenxuan Liu 0008, Xuemei Jia, Xian Zhong, Kui Jiang, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 2 |
| 2025 | Motion-Consistent Representation Learning for UAV-Based Action RecognitionabstractAction recognition aims to identify action categories in trimmed videos captured by multimedia devices, which often suffer from jitter, especially in uncrewed aerial vehicle (UAV) applications. Existing methods typically ignore the effect of jitter on actor motion or rely on external stabilization tools trained on large-scale unstable video datasets that may not be tailored to specific tasks. To address this, we propose the Stabilization-enhanced Recognition Network (StaRNet), an end-to-end framework that integrates video stabilization and contrastive learning. Inspired by traditional stabilizers, StaRNet’s Motion-aware Stabilization Module (MSM) constructs positive and negative video pairs to model instability: positive pairs use optical flow to estimate frame motion and refine rigid motion via keyframe estimation for motion-aware stabilization, while negative pairs assess temporal consistency using motion cues to boost classification. Moreover, we introduce a Motion-aware Constraint (MC) that regulates dynamic stabilization to adapt to varying motion patterns and enrich action representations. Experiments on UAV benchmarks show that StaRNet outperforms state-of-the-art methods and substantially enhances video stabilization. The code is available athttps://github.com/lwxfight/-StaRNet Wenxuan Liu 0008, Xian Zhong, Yihan Dai, Xuemei Jia, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Fragrant: frequency-auxiliary guided relational attention network for low-light action recognition
Wenxuan Liu 0008, Xuemei Jia, Yihao Ju, Yakun Ju, Kui Jiang, Shifeng Wu, Luo Zhong, Xian Zhong |
Vis. Comput. | 2 |
| 2024 | Physical Adversarial Attack Meets Computer Vision: A Decade SurveyabstractDespite the impressive achievements of Deep Neural Networks (DNNs) in computer vision, their vulnerability to adversarial attacks remains a critical concern. Extensive research has demonstrated that incorporating sophisticated perturbations into input images can lead to a catastrophic degradation in DNNs' performance. This perplexing phenomenon not only exists in the digital space but also in the physical world. Consequently, it becomes imperative to evaluate the security of DNNs-based systems to ensure their safe deployment in real-world scenarios, particularly in security-sensitive applications. To facilitate a profound understanding of this topic, this paper presents a comprehensive overview of physical adversarial attacks. First, we distill four general steps for launching physical adversarial attacks. Building upon this foundation, we uncover the pervasive role of artifacts carrying adversarial perturbations in the physical world. These artifacts influence each step. To denote them, we introduce a new term: adversarial medium. Then, we take the first step to systematically evaluate the performance of physical adversarial attacks, taking the adversarial medium as a first attempt. Our proposed evaluation metric, hiPAA, comprises six perspectives: Effectiveness, Stealthiness, Robustness, Practicability, Aesthetics, and Economics. We also provide comparative results across task categories, together with insightful observations and suggestions for future research directions. Hui Wei 0004, Hao Tang 0005, Xuemei Jia, Zhixiang Wang 0001, Hanxun Yu, Zhubo Li, Shin'ichi Satoh 0001, Luc Van Gool, Zheng Wang 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | ICLR: Instance Credibility-Based Label Refinement for label noisy person re-identification
Xian Zhong, Xuemei Jia, Wenxin Huang, Wenxuan Liu 0008, Shuaipeng Su, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 3 |
| 2023 | HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable DesignabstractAdversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solutions are often impractical and attention-grabbing. To address the need for a physically practical and stealthy adversarial attack, we introduce HotCold Block, a novel physical attack for infrared detectors that hide persons utilizing the wearable Warming Paste and Cooling Paste. By attaching these readily available temperature-controlled materials to the body, HotCold Block evades human eyes efficiently. Moreover, unlike existing methods that build adversarial patches with complex texture and structure features, HotCold Block utilizes an SSP-oriented adversarial optimization algorithm that enables attacks with pure color blocks and explores the influence of size, shape, and position on attack performance. Extensive experimental results in both digital and physical environments demonstrate the performance of our proposed HotCold Block. Code is available: https://github.com/weihui1308/HOTCOLDBlock. Hui Wei 0004, Zhixiang Wang 0001, Xuemei Jia, Yinqiang Zheng, Hao Tang 0005, Shin'ichi Satoh 0001, Zheng Wang 0007 |
AAAI | 3 |
| 2023 | Neighborhood Information-Based Label Refinement for Person Re-Identification with Label NoiseabstractThe existing excellent person re-identification (Re-ID) model is still affected by the samples with the incorrect labels. It is difficult to accurately annotate person images in the real scene, resulting in label noise. To avoid fitting to the noisy labels, a common solution in Re-ID is to replace the original label with the label predicted by the deep model. Unfortunately, similar samples of different identities with the same label are due to label noise, which is challenging for the model to distinguish them. Neighborhood information can optimize noisy labels through neighborhood labels and similarity between samples. This paper proposes a label refinement module based on neighborhood information (LRNI) for person Re-ID with label noise. Specifically, we first use the pre-trained model to extract features and calculate the similarity between samples. Rather than treating samples as isolated, the similarity used as label propagation weight and neighborhood labels are combined to optimize noisy labels. To further reduce the influence of label noise, we design a hard sample re-weighting (HSR) strategy to balance the learning of noisy and boundary samples. Experimental results under different noise settings demonstrate our method's effectiveness in the person Re-ID task. Xian Zhong, Shuaipeng Su, Wenxuan Liu 0008, Xuemei Jia, Wenxin Huang, Mengdie Wang |
ICASSP | 4 |
| 2023 | Beyond the Parts: Learning Coarse-to-Fine Adaptive Alignment Representation for Person SearchabstractPerson search is a time-consuming computer vision task that entails locating and recognizing query people in scenic pictures. Body components are commonly mismatched during matching due to position variation, occlusions, and partially absent body parts, resulting in unsatisfactory person search results. Existing approaches for extracting local characteristics of the human body using keypoint information are unable to handle the search job when distinct body parts are misaligned, ignoring to exploit multiple granularities, which is crucial in the person search process. Moreover, the alignment learning methods learn body part features with fixed and equal weights, ignoring the beneficial contextual information, e.g., the umbrella carried by the pedestrian, which supplements compelling clues for identifying the person. In this paper, we propose a Coarse-to-Fine Adaptive Alignment Representation (CFA 2 R) network for learning multiple granular features in misaligned person search in the coarse-to-fine perspective. To exploit more beneficial body parts and related context of the cropped pedestrians, we design a Part-Attentional Progressive Module (PAPM) to guide the network to focus on informative body parts and positive accessorial regions. Besides, we propose a Re-weighting Alignment Module (RAM) shedding light on more contributive parts instead of treating them equally. Specifically, adaptive re-weighted but not fixed part features are reconstructed by Re-weighting Reconstruction module, considering that different parts serve unequally during image matching. Extensive experiments conducted on CUHK-SYSU and PRW datasets demonstrate competitive performance of our proposed method. Wenxin Huang, Xuemei Jia, Xian Zhong, Xiao Wang 0029, Kui Jiang, Zheng Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | VCD: View-Constraint Disentanglement for Action RecognitionabstractAction recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is tantamount to track the view-invariant representation as it evolves with the performed action. However, the views’ and actions’ performance always complement each other, once simply looking for the view-invariant representation may cause some behavior information to be lost. In this paper, we propose the View-Constraint Disentanglement (VCD) framework for cross-view action recognition. Specifically, Constraint Disentanglement Module (CDM) is utilized to learn an action-invariant representation by discretizing view-specific representation and its normal distribution, which resolves the entangled relationship between view and action. Moreover, a novel Adaptive Distribution Module (ADM) is intended to befit enhance the high-correlation viewpoint variation information and refine the suitable weight. Extensive experiments are conducted on public benchmarks, indicating that our approach achieves better performance than other state-of-the-art approaches. Xian Zhong, Zhuo Zhou, Wenxuan Liu 0008, Kui Jiang, Xuemei Jia, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 5 |
| 2022 | Patching Your Clothes: Semantic-Aware Learning for Cloth-Changed Person Re-Identification
Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang |
MMM (2) | 1 |
| 2022 | Actor-Aware Alignment Network for Action RecognitionabstractAction recognition has attracted growing interest recently. It suffers from the problem that complex and diverse environments may disturb the extraction of action features. Existing methods propose to explore the temporal associations to alleviate the issue. However, they cannot handle long-range frames, and the rigid techniques are powerless against the differences caused by the deformation of the actors. To this end, we propose the Actor-Aware Alignment Network (A$^{3}$Net), which helps locate the action region. Specifically, through the intra-snippet correction, we afford the local segment alignment frames. The inter-snippet is designed to rectify the results, avoiding the occlusion situation that may appear in the local snippet. In addition, we consider intra-alignment short-range adjustive frames and long-range context frames between different snippets, which allows our A$^{3}$Net network to achieve the effect of focusing on long-range frame information. Multiple Reasoning Attention (MRA) modules are introduced to integrate features along the temporal dimension to keep the video spatio-temporal consistent. Extensive experiments conducted on three widely-used public benchmarks,UCF101,HMDB51, andInfAR, indicate that the excellence of our approach over other state-of-the-art models in wild scenarios. Wenxuan Liu 0008, Xian Zhong, Xuemei Jia, Kui Jiang, Chia-Wen Lin |
IEEE Signal Process. Lett. | 3 |
| 2022 | Grayscale Enhancement Colorization Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is an emerging and challenging cross-modality image matching problem because of the explosive surveillance data in night-time surveillance applications. To handle the large modality gap, various generative adversarial network models have been developed to eliminate the cross-modality variations based on a cross-modal image generation framework. However, the lack of point-wise cross-modality ground-truths makes it extremely challenging to learn such a cross-modal image generator. To address these problems, we learn the correspondence between single-channel infrared images and three-channel visible images by generating intermediate grayscale images as auxiliary information to colorize the single-modality infrared images. We propose a grayscale enhancement colorization network (GECNet) to bridge the modality gap by retaining the structure of the colored image which contains rich information. To simulate the infrared-to-visible transformation, the point-wise transformed grayscale images greatly enhance the colorization process. Our experiments conducted on two visible-infrared cross-modality person re-identification datasets demonstrate the superiority of the proposed method over the state-of-the-arts. Xian Zhong, Tianyou Lu, Wenxin Huang, Mang Ye, Xuemei Jia, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Complementary Data Augmentation for Cloth-Changing Person Re-IdentificationabstractThis paper studies the challenging person re-identification (Re-ID) task under the cloth-changing scenario, where the same identity (ID) suffers from uncertain cloth changes. To learn cloth- and ID-invariant features, it is crucial to collect abundant training data with varying clothes, which is difficult in practice. To alleviate the reliance on rich data collection, we reinforce the feature learning process by designing powerful complementary data augmentation strategies, including positive and negative data augmentation. Specifically, the positive augmentation fulfills the ID space by randomly patching the person images with different clothes, simulating rich appearance to enhance the robustness against clothes variations. For negative augmentation, its basic idea is to randomly generate out-of-distribution synthetic samples by combining various appearance and posture factors from real samples. The designed strategies seamlessly reinforce the feature learning without additional information introduction. Extensive experiments conducted on both cloth-changing and -unchanging tasks demonstrate the superiority of our proposed method, consistently improving the accuracy over various baselines. Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang |
IEEE Trans. Image Process. | 1 |
| 2021 | Random Walk Erasing with Attention Calibration for Action Recognition
Yuze Tian, Xian Zhong, Wenxuan Liu 0008, Xuemei Jia, Mang Ye |
PRICAI (3) | 4 |