EDBT 2026 Demo / reviewers in the wild / expert
Wenxuan Liu 0008
dblp:197/5243-8
· DBLP profile ↗
41ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0002-4417-6628ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 29 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order TransferabstractAction recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition robustness. To address this, we introduce a multi-view formulation tailored for UAV altitudes and empirically uncover a distinctive partial order among views, where recognition accuracy consistently declines as altitude increases. This key observation motivates the proposed Aero Partial Order Guided Network (Aerorder), which explicitly models and exploits the hierarchical structure of UAV views to enhance cross-altitude action recognition. Aerorder comprises three main components: (1) a View Partition (VP) module that groups views by altitude using the head-to-body ratio; (2) an Order-aware Feature Decoupling (OFD) module that disentangles action-relevant and view-specific representations under partial order guidance; and (3) an Action Partial Order Guide (APOG) that progressively transfers knowledge from easier (low-altitude) to harder (high-altitude) views. Extensive experiments on Drone-Action, MOD20, and UAV validate the superiority of Aerorder, achieving consistent improvements over state-of-the-art methods, up to 4.7% and 1.3% gains on Drone-Action and MOD20, respectively. Wenxuan Liu 0008, Zhuo Zhou, Xuemei Jia, Siyuan Yang 0001, Wenxin Huang, Xian Zhong, Chia-Wen Lin |
AAAI | 1 |
| 2026 | From temporal thumbnail to semantics: Debiasing multi-view action recognition
Zixian Zhu, Wenxuan Liu 0008, Xu Wang 0015, Bingyi Liu, Xiaohan Yu 0001, Xian Zhong |
Pattern Recognit. | 3 |
| 2026 | AM40: Enhancing action recognition through matting-driven interaction analysis
Wenxuan Liu 0008, Kui Jiang, Siyuan Yang 0001, Chia-Wen Lin, Xian Zhong |
Pattern Recognit. | 2 |
| 2026 | Robust mixed-degradation person Re-identification via structural consistency distillation
Wenxin Huang, Wenxuan Liu 0008, Xuemei Jia, Xian Zhong |
Pattern Recognit. | 3 |
| 2026 | PhyTrace: Tracing Physical Inconsistency in AI-Generated Images via ISP EmulationabstractThe high realism of AI-generated images has emerged as a significant cybersecurity threat. While existing detection methods have achieved some success, most rely on fixed models that are incapable of adapting to new data or generating model updates. This paper overcomes these limitations by shifting the focus to the fundamental imaging process of real images: Image Signal Processing (ISP). Unlike real images, AI-generated images do not undergo this process, making them more susceptible to physical variations within ISP modules. By analyzing how ISP sub-modules influence the physical characteristics of imaging, we simulate the ISP mapping process to amplify the differences in physical responses between real and AI-generated images during ISP transformations. Tracing these physical differences, we propose PhyTrace, a novel training-free method for detecting AI-generated images. PhyTrace enforces physical consistency constraints within ISP, operates independently of specific datasets and generative models, and effectively detects a wide range of AI-generated images. PhyTrace reveals distinct distribution patterns of real and AI-generated images. Extensive experiments on 18 test sets demonstrate that our method outperforms prior approaches in average precision and generalization, offering a robust solution for AI-generated image detection in open-world scenarios. Wenxuan Liu 0008, Danni Xu, Joey Tianyi Zhou, Zheng Wang 0007 |
IEEE Trans. Image Process. | 2 |
| 2026 | TCP: Text-Guided Cascade Network for Pedestrian Crossing Intention PredictionabstractPedestrian crossing intention prediction is crucial for ensuring safety in intelligent transportation systems, especially in autonomous driving scenarios. Most existing methods rely primarily on visual information; however, the quality of visual data deteriorates significantly at long distances due to limited resolution. Although multi-modal approaches can mitigate this issue by incorporating additional sensory data, they inevitably introduce extra computational overhead. To address these challenges, we propose a lightweight cascaded model for pedestrian crossing intention prediction based on text-trajectory alignment. The model employs a cascaded architecture that jointly performs coordinate and intention prediction, while leveraging a pre-trained large language model (LLM) to generate textual descriptions of videos, thereby enriching trajectory features. Furthermore, a center-aware classification module is integrated to enhance inter-class separability and intra-class compactness. Extensive experiments onJAADandPIEdemonstrate state-of-the-art performance: our method achieves 91% accuracy onPIEand 89% onJAAD, matching or surpassing recent multi-modal approaches with substantially fewer inputs. The source code will be released athttps://github.com/xyhhappy/TCP-prediction Wenxuan Liu 0008, Wenxin Huang, Ryan Wen Liu, Xian Zhong |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Pioneering Explainable Video Fact-Checking with a New Dataset and Multi-role Multimodal Model ApproachabstractExisting video fact-checking datasets often lack detailed evidence and explanations, compromising the reliability and interpretability of fact-checking methods. To address these gaps, we developed a novel dataset featuring comprehensive annotations for each news item, including veracity labels, the rationales behind these labels, and supporting evidence. This dataset significantly enhances models' ability to accurately identify and explain video content. We also present an explainable automatic framework 3MFact, utilizing Multi-role Multimodal Models for video Fact-checking. Our framework iteratively gathers and synthesizes online evidence to progressively determine the veracity label, generating three key outputs: veracity label, rationale, and supported evidence. We aim for this work to be a pioneering effort, providing robust support for the field of video fact-checking. Kaipeng Niu, Danni Xu, Bingjian Yang, Wenxuan Liu 0008, Zheng Wang 0007 |
AAAI | 4 |
| 2025 | Anomize: Better Open Vocabulary Video Anomaly DetectionabstractOpen Vocabulary Video Anomaly Detection (OVVAD) seeks to detect and classify both base and novel anomalies. However, existing methods face two specific challenges related to novel anomalies. The first challenge is detection ambiguity, where the model struggles to assign accurate anomaly scores to unfamiliar anomalies. The second challenge is categorization confusion, where novel anomalies are often misclassified as visually similar base instances. To address these challenges, we explore supplementary information from multiple sources to mitigate detection ambiguity by leveraging multiple levels of visual data alongside matching textual information. Furthermore, we propose incorporating label relations to guide the encoding of new labels, thereby improving alignment between novel videos and their corresponding labels, which helps reduce categorization confusion. The resulting Anomize framework effectively tackles these issues, achieving superior performance on UCF-Crime and XD-Violence datasets, demonstrating its effectiveness in OVVAD. Wenxuan Liu 0008, Ruixu Zhang, Yuran Wang 0003, Xian Zhong, Zheng Wang 0007 |
CVPR | 2 |
| 2025 | Differential Coding for Training-Free ANN-to-SNN ConversionabstractSpiking Neural Networks (SNNs) exhibit significant potential due to their low energy consumption. Converting Artificial Neural Networks (ANNs) to SNNs is an efficient way to achieve high-performance SNNs. However, many conversion methods are based on rate coding, which requires numerous spikes and longer time-steps compared to directly trained SNNs, leading to increased energy consumption and latency. This article introduces differential coding for ANN-to-SNN conversion, a novel coding scheme that reduces spike counts and energy consumption by transmitting changes in rate information rather than rates directly, and explores its application across various layers. Additionally, the threshold iteration method is proposed to optimize thresholds based on activation distribution when converting Rectified Linear Units (ReLUs) to spiking neurons. Experimental results on various Convolutional Neural Networks (CNNs) and Transformers demonstrate that the proposed differential coding significantly improves accuracy while reducing energy consumption, particularly when combined with the threshold iteration method, achieving state-of-the-art performance. The source codes of the proposed method are available at https://github.com/h-z-h-cell/ANN-to-SNN-DCGS. Zihan Huang, Wei Fang 0006, Tong Bu, Zecheng Hao, Wenxuan Liu 0008, Yuanhong Tang, Zhaofei Yu, Tiejun Huang 0001 |
ICML | 6 |
| 2025 | SOTA: Spike-Navigated Optimal TrAnsport Saliency Region Detection in Composite-bias VideosabstractExisting saliency detection methods struggle in real-world scenarios due to motion blur and occlusions. In contrast, spike cameras, with their high temporal resolution, significantly enhance visual saliency maps. However, the composite noise inherent to spike camera imaging introduces discontinuities in saliency detection. Low-quality samples further distort model predictions, leading to saliency bias. To address these challenges, we propose Spike-navigated Optimal TrAnsport Saliency Region Detection (SOTA), a framework that leverages the strengths of spike cameras while mitigating biases in both spatial and temporal dimensions. Our method introduces Spike-based Micro-debias (SM) to capture subtle frame-to-frame variations and preserve critical details, even under minimal scene or lighting changes. Additionally, Spike-based Global-debias (SG) refines predictions by reducing inconsistencies across diverse conditions. Extensive experiments on real and synthetic datasets demonstrate that SOTA outperforms existing methods by eliminating composite noise bias. Our code and dataset will be released at https://github.com/lwxfight/sota. Wenxuan Liu 0008, Xian Zhong, Zhaofei Yu, Tiejun Huang 0001 |
IJCAI | 1 |
| 2025 | SegTraj: A Segmented-Trajectory-Aware Spatio-Temporal Graph Convolutional Network for Social Group DetectionabstractSocial group detection aims to identify groups of individuals exhibiting social behavior from multi-individual trajectory data. Recent approaches often determine group correlations based on global trajectory similarity, while temporal dynamics can cause diverging member trajectories and undermine similarity-based measures. Other methods model pairwise interaction strengths to capture group relations, focusing only on explicit direct interactions while ignoring implicit indirect interactions. To address temporal variability of group structures, we decompose long trajectories into multiple semantic sub-trajectories, enabling the capture of dynamic characteristics. Furthermore, to explore implicit indirect interactions, we introduce a unified spatio-temporal graph structure that models both direct and indirect interactions among individuals. In addition, considering the contextual influence of the neighborhood of an individual, we incorporate neighborhood information into the trajectory representation process. Based on these insights, we propose a Segmented-Trajectory-Aware Spatio-Temporal Graph Convolutional Network (SegTraj). This framework uniformly models explicit and implicit interactions through a spatio-temporal graph, and fuses individual trajectories with contextual neighborhood information for fine-grained representation of group relationships. Extensive experiments on three datasets covering both synthetic and real-world scenarios demonstrate that SegTraj significantly outperforms baseline methods. The code is available at https://github.com/DC0827/SegTraj. Xiongwei Dang, Wenxuan Liu 0008, Xian Zhong, Zheng Wang 0007 |
ACM Multimedia | 2 |
| 2025 | SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention LocalizationabstractTemporal Intention Localization (TIL) is crucial for video surveillance, focusing on identifying varying levels of suspicious intention to enhance security monitoring. However, existing discrete classification methods fail to capture the continuous progression of suspicious intentions, limiting early intervention and explainability. In this paper, we reconceptualize hidden intention modeling by shifting from discrete classification to continuous regression and propose Suspicion Progression Analysis Network (SPAN), which capture the fluctuations and progression of hidden intentions over time. Specifically, when analyzing the temporal progression of suspicion, we discover that suspicion exhibits long-term dependency and cumulative effects across extended sequences, characteristics significantly similar to the settings in Temporal Point Process (TPP) theory. Based on these insights, we formalize a suspicion score formula that models continuous changes while accounting for temporal characteristics. We also propose Suspicion Coefficient Modulation to adjust suspicion coefficients using multimodal information, reflecting different effects of suspicious actions. Notably, we introduce a Concept-Anchored Mapping method to quantify associations between suspicious actions and predefined intention concepts, enabling understanding of not just actions occurring but also their potential underlying intentions. Extensive experiments on the HAI dataset show that SPAN significantly outperforms existing methods, reducing MSE by 19.8% and improving average mAP by 1.78%,. Notably, SPAN achieves a 2.74% mAP gain in low-frequency cases, indicating superior capability in capturing subtle behavioral changes.Compared to discrete classification systems, out continuous suspicion modeling method enables earlier detection and more proactive interventions, substantially enhancing both system explainability and practical utility in security applications. Yuran Wang 0003, Ruixu Zhang, Yue Li 0038, Wenxuan Liu 0008, Zheng Wang 0007 |
ACM Multimedia | 5 |
| 2025 | A New Dataset and Benchmark for Grounding Multimodal MisinformationabstractThe proliferation of online misinformation videos poses serious societal risks. Current datasets and detection methods primarily target binary classification or single-modality localization based on post-processed data, lacking the interpretability needed to counter persuasive misinformation. In this paper, we introduce the task of Grounding Multimodal Misinformation (GroundMM), which verifies multimodal content and localizes misleading segments across modalities. We present the first real-world dataset for this task, GroundLie360, featuring a taxonomy of misinformation types, fine-grained annotations across text, speech, and visuals, and validation with Snopes evidence and annotator reasoning. We also propose a VLM-based, QA-driven baseline, FakeMark, using single and cross-modal cues for effective detection and grounding. Our experiments highlight the challenges of this task and lay a foundation for explainable multimodal misinformation detection. Dataset will be released at https://github.com/yangbingjian/GroundLie360. Bingjian Yang, Danni Xu, Kaipeng Niu, Wenxuan Liu 0008, Zheng Wang 0007, Mohan Kankanhalli |
ACM Multimedia | 4 |
| 2025 | Beyond the Individual: Introducing Group Intention Forecasting with SHOT DatasetabstractIntention recognition has traditionally focused on individual intentions, overlooking the complexities of collective intentions in group settings. To address this limitation, we introduce the concept of group intention, which represents shared goals emerging through the actions of multiple individuals, and Group Intention Forecasting (GIF), a novel task that forecasts when group intentions will occur by analyzing individual actions and interactions before the collective goal becomes apparent. To investigate GIF in a specific scenario, we propose SHOT, the first large-scale dataset for GIF, consisting of 1,979 basketball video clips captured from 5 camera views and annotated with 6 types of individual attributes. SHOT is designed with 3 key characteristics: multi-individual information, multi-view adaptability, and multi-level intention, making it well-suited for studying emerging group intentions. Furthermore, we introduce GIFT (Group Intention ForecasTer), a framework that extracts fine-grained individual features and models evolving group dynamics to forecast intention emergence. Experimental results confirm the effectiveness of SHOT and GIFT, establishing a strong foundation for future research in group intention forecasting. The dataset is available at https://xinyi-hu.github.io/SHOT\_DATASET. Ruixu Zhang, Yuran Wang 0003, Chaoyu Mai, Wenxuan Liu 0008, Danni Xu, Xian Zhong, Zheng Wang 0007 |
ACM Multimedia | 5 |
| 2025 | Dynamic and static mutual fitting for action recognition
Wenxuan Liu 0008, Xuemei Jia, Xian Zhong, Kui Jiang, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 1 |
| 2025 | Motion-Consistent Representation Learning for UAV-Based Action RecognitionabstractAction recognition aims to identify action categories in trimmed videos captured by multimedia devices, which often suffer from jitter, especially in uncrewed aerial vehicle (UAV) applications. Existing methods typically ignore the effect of jitter on actor motion or rely on external stabilization tools trained on large-scale unstable video datasets that may not be tailored to specific tasks. To address this, we propose the Stabilization-enhanced Recognition Network (StaRNet), an end-to-end framework that integrates video stabilization and contrastive learning. Inspired by traditional stabilizers, StaRNet’s Motion-aware Stabilization Module (MSM) constructs positive and negative video pairs to model instability: positive pairs use optical flow to estimate frame motion and refine rigid motion via keyframe estimation for motion-aware stabilization, while negative pairs assess temporal consistency using motion cues to boost classification. Moreover, we introduce a Motion-aware Constraint (MC) that regulates dynamic stabilization to adapt to varying motion patterns and enrich action representations. Experiments on UAV benchmarks show that StaRNet outperforms state-of-the-art methods and substantially enhances video stabilization. The code is available athttps://github.com/lwxfight/-StaRNet Wenxuan Liu 0008, Xian Zhong, Yihan Dai, Xuemei Jia, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Fragrant: frequency-auxiliary guided relational attention network for low-light action recognition
Wenxuan Liu 0008, Xuemei Jia, Yihao Ju, Yakun Ju, Kui Jiang, Shifeng Wu, Luo Zhong, Xian Zhong |
Vis. Comput. | 1 |
| 2024 | Bi-Causal: Group Activity Recognition via Bidirectional CausalityabstractCurrent approaches in Group Activity Recognition (GAR) predominantly emphasize Human Relations (HRs) while often neglecting the impact of Human-Object Inter-actions (HOIs). This study prioritizes the consideration of both HRs and HOIs, emphasizing their interdependence. Notably, employing Granger Causality Tests reveals the presence of bidirectional causality between HRs and HOIs. Leveraging this insight, we propose a Bidirectional-Causal GAR network. This network establishes a causality commu-nication channel while modeling relations and interactions, enabling reciprocal enhancement between human-object interactions and human relations, ensuring their mutual consistency. Additionally, an Interaction Module is devised to effectively capture the dynamic nature of human-object interactions. Comprehensive experiments conducted on two publicly available datasets showcase the superiority of our proposed method over state-of-the-art approaches. Our project page: https://angzong.github.io/bi-causal.github.io/ Youliang Zhang, Wenxuan Liu 0008, Danni Xu, Zhuo Zhou, Zheng Wang 0007 |
CVPR | 2 |
| 2024 | Localization of Image Splicing Under Segment Anything Model With Integrated Compression and Edge ArtifactsabstractThe localization of image splicing involves identifying pixels in an image that have been spliced from other images, necessitating the discernment of splicing features. Despite significant advancements driven by the rise of social media and deep learning, existing methods exhibit limitations, often neglecting the integration of coarse and precise features and lacking the ability to understand objects. This leads to erroneous predictions in identifying spliced regions. This paper proposes Segment Anything Model with Integrated Compression and Edge artifacts (SAM-ICE) for the localization of image splicing, addressing these limitations by fusing forged edge features and compression artifact features. Leveraging SAM’s object understanding ability, our method identifies spliced regions using the fused features as guidance. Specifically, we employ Edge Artifact Extractor (EAE) to extract fine high-frequency edge splicing features and Compression Artifact Extractor (CAE) to extract coarse compression artifact features. By combining these features, our method utilizes coarse-fine features to accurately pinpoint the spliced portions of the image. Experimental results demonstrate the superior accuracy, robustness, and generalizability of our method compared to the state-of-the-arts. Ruhao Zhao, Xian Zhong, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ICIP | 4 |
| 2024 | Predicting the Unseen: A Novel Dataset for Hidden Intention Localization in Pre-abnormal Analysis
Zehao Qi, Ruixu Zhang, Wenxuan Liu 0008, Zheng Wang 0007 |
ACM Multimedia | 4 |
| 2024 | Towards Low-latency Event-based Visual Recognition with Hybrid Step-wise Distillation Spiking Neural NetworksabstractSpiking neural networks (SNNs) have garnered significant attention for their low power consumption and high biological interpretability. Their rich spatio-temporal information processing capability and event-driven nature make them ideally well-suited for neuromorphic datasets. However, current SNNs struggle to balance accuracy and latency in classifying these datasets. In this paper, we propose Hybrid Step-wise Distillation (HSD) method, tailored for neuromorphic datasets, to mitigate the notable decline in performance at lower time steps. Our work disentangles the dependency between the number of event frames and the time steps of SNNs, utilizing more event frames during the training stage to improve performance, while using fewer event frames during the inference stage to reduce latency. Nevertheless, the average output of SNNs across all time steps is susceptible to individual time step with abnormal outputs, particularly at extremely low time steps. To tackle this issue, we implement Step-wise Knowledge Distillation (SKD) module that considers variations in the output distribution of SNNs at each time step. Empirical evidence demonstrates that our method yields competitive performance in classification tasks on neuromorphic datasets, especially at lower time steps. Our code will be available at: https://github.com/hsw0929/HSD. Xian Zhong, Shengwang Hu, Wenxuan Liu 0008, Wenxin Huang, Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001 |
ACM Multimedia | 3 |
| 2024 | ICLR: Instance Credibility-Based Label Refinement for label noisy person re-identification
Xian Zhong, Xuemei Jia, Wenxin Huang, Wenxuan Liu 0008, Shuaipeng Su, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 5 |
| 2023 | Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic SegmentationabstractWhile enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to handle the more realistic multi-target domain adaptive semantic segmentation (MT-DASS) tasks. To solve this problem, we propose a Bi-Alignment framework based on Transformation (BAT). Specifically, we employ the Fourier style transform to convert the style of the source domain to that of the target domain without training any style transfer networks. In this way, we transform the single-peak distributed source domain into a multi-peak distribution that resembles the multi-target domain. Then, we perform fine-grained global and local dual distribution alignment between the same style of source-target domain pairs to achieve a multi-to-multi distribution alignment. Finally, self-training is utilized to further improve the network’s discriminability. Experimental results show that our approach achieves competitive results over state-of-the-art methods. Xian Zhong, Jing Xiao 0004, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 5 |
| 2023 | Neighborhood Information-Based Label Refinement for Person Re-Identification with Label NoiseabstractThe existing excellent person re-identification (Re-ID) model is still affected by the samples with the incorrect labels. It is difficult to accurately annotate person images in the real scene, resulting in label noise. To avoid fitting to the noisy labels, a common solution in Re-ID is to replace the original label with the label predicted by the deep model. Unfortunately, similar samples of different identities with the same label are due to label noise, which is challenging for the model to distinguish them. Neighborhood information can optimize noisy labels through neighborhood labels and similarity between samples. This paper proposes a label refinement module based on neighborhood information (LRNI) for person Re-ID with label noise. Specifically, we first use the pre-trained model to extract features and calculate the similarity between samples. Rather than treating samples as isolated, the similarity used as label propagation weight and neighborhood labels are combined to optimize noisy labels. To further reduce the influence of label noise, we design a hard sample re-weighting (HSR) strategy to balance the learning of noisy and boundary samples. Experimental results under different noise settings demonstrate our method's effectiveness in the person Re-ID task. Xian Zhong, Shuaipeng Su, Wenxuan Liu 0008, Xuemei Jia, Wenxin Huang, Mengdie Wang |
ICASSP | 3 |
| 2023 | Background-Weakening Consistency Regularization for Semi-Supervised Video Action DetectionabstractConsistency-based techniques have produced state-of-the-art results in semi-supervised action detection. When the model false detects the dynamic information in the background as an action, spatio-temporal consistency calculations can hardly reflect this false detection result. We consider weakening the dynamic information in the augmented video background to reduce its spatio-temporal consistency with the dynamic information in the original video background. Thus we propose a Background-Weakening with Calibration Constraint (BWCC) framework, which highlights the negative impact of information in the background of false detection by calculating the consistency of the predictions of the background weakened video and the original video. Specifically, Background Weaken (BW) module judges the foreground and background of the video based on the initial predictions of the model and makes adjustments to the video background. To mitigate the effects of the misjudgments result in weakened action pixels, we additionally introduce a model that does not undergo background weakening to aid training through Calibration Constraint (CC) module. Our approach achieves competitive performance over existing leading approaches on two action detection datasets, UCF101-24 and JHMDB-21. Xian Zhong, Aoyu Yi, Wenxuan Liu 0008, Wenxin Huang, Chengming Zou, Zheng Wang 0007 |
ICASSP | 3 |
| 2023 | Implicit Attention-Based Cross-Modal Collaborative Learning for Action RecognitionabstractHuman action recognition is an active research topic in recent years. Multiple modalities often convey heterogeneous but potentially complementary action information that single modality does not hold. Some efforts have been resoted to explore cross-modal representation to promote the modeling capability, but with limited improvement due to the simple fusion of different modalities. To this end, we propose an impliCit attention-based Cross-modal Collaborative Learning (C3L) for action recognition. Specifically, we apply a Modality Generalization network with Grayscale enhancement (MGG) to learn specific modality representation and interaction (infrared and RGB). Then, we construct a unified representation space through the Uniform Modality Representation module (UMR), which preserves the modality information while enhancing the overall representation ability. Finally, feature extractors adaptively leverage modality-specific knowledge to realize cross-modal collaborative learning. Extensive experiments conducted on three widely-used public benchmarks InfAR, HMDB51, and UCF101, demonstrate the effectiveness and strength of our proposed method. Jianghao Zhang, Xian Zhong, Wenxuan Liu 0008, Kui Jiang, Zhengwei Yang 0001, Zheng Wang 0007 |
ICIP | 3 |
| 2023 | DAWN: Direction-aware Attention Wavelet Network for Image DerainingabstractSingle image deraining aims to remove rain perturbation while restoring the clean background scene from a rain image. However, existing methods tend to produce blurry and over-smooth outputs, lacking some textural details. Wavelet transform can depict the contextual and textural information of an image at different levels, showing impressive capability of learning structural information in the images to avoid artifacts, and thus has been recently explored to consider the inherent overlap of background and rain perturbation in both the pixel domain and the frequency embedding space. However, the existing wavelet-based methods ignore the heterogeneous degradation for different coefficients due to the inherent directional characteristics of rain streaks, leading to inter-frequency conflicts and compromised deraining results. To address this issue, we propose a novel Direction-aware Attention Wavelet Network (DAWN) for rain streaks removal. DAWN has several key distinctions from existing wavelet transform-based methods: 1) introducing the vector decomposition to parameterize the learning procedure, where the rain streaks are derived into the vertical (V) and horizontal (H) components to learn the specific representation; 2) a novel direction-aware attention module (DAM) to fit the projection and transformation parameters to characterize the direction-specific rain components, which helps accurate texture restoration; 3) exploring practical composite constraints on the structure, details, and chrominance aspects for high-quality background restoration. Our proposed DAWN delivers significant performance gains on nine datasets across image deraining and object detection tasks, exceeding the state-of-the-art method MPRNet by 0.88 dB in PSNR on the Test1200 dataset with only 35.5% computation cost. Kui Jiang, Wenxuan Liu 0008, Zheng Wang 0007, Xian Zhong, Junjun Jiang, Chia-Wen Lin |
ACM Multimedia | 2 |
| 2023 | Uncovering the Unseen: Discover Hidden Intentions by Micro-Behavior Graph ReasoningabstractThis paper introduces a new and challenging Hidden Intention Discovery (HID) task. Unlike existing intention recognition tasks, which are based on obvious visual representations to identify common intentions for normal behavior, HID focuses on discovering hidden intentions when humans try to hide their intentions for abnormal behavior. HID presents a unique challenge in that hidden intentions lack the obvious visual representations to distinguish them from normal intentions. Fortunately, from a sociological and psychological perspective, we find that the difference between hidden and normal intentions can be reasoned from multiple micro-behaviors, such as gaze, attention, and facial expressions. Therefore, we first discover the relationship between micro-behavior and hidden intentions and use graph structure to reason about hidden intentions. To facilitate research in the field of HID, we also constructed a seminal dataset containing a hidden intention annotation of a typical theft scenario for HID. Extensive experiments show that the proposed network improves performance on the HID task by 9.9% over the state-of-the-art method SBP. Zhuo Zhou, Wenxuan Liu 0008, Danni Xu, Zheng Wang 0007, Jian Zhao 0006 |
ACM Multimedia | 2 |
| 2023 | SCPNet: Self-constrained parallelism network for keypoint-based lightweight object detection
Xian Zhong, Mengdie Wang, Wenxuan Liu 0008, Jingling Yuan, Wenxin Huang |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Dual-Recommendation Disentanglement Network for View Fuzz in Action RecognitionabstractMulti-view action recognition aims to identify action categories from given clues. Existing studies ignore the negative influences of fuzzy views between view and action in disentangling, commonly arising the mistaken recognition results. To this end, we regard the observed image as the composition of the view and action components, and give full play to the advantages of multiple views via the adaptive cooperative representation among these two components, forming a Dual-Recommendation Disentanglement Network (DRDN) for multi-view action recognition. Specifically, 1) For the action, we leverage a multi-level Specific Information Recommendation (SIR) to enhance the interaction among intricate activities and views. SIR offers a more comprehensive representation of activities, measuring the trade-off between global and local information. 2) For the view, we utilize a Pyramid Dynamic Recommendation (PDR) to learn a complete and detailed global representation by transferring features from different views. It is explicitly restricted to resist the fuzzy noise influence, focusing on positive knowledge from other views. Our DRDN aims for complete action and view representation, where PDR directly guides action to disentangle with view features and SIR considers mutual exclusivity of view and action clues. Extensive experiments have indicated that the multi-view action recognition method DRDN we proposed achieves state-of-the-art performance over powerful competitors on several standard benchmarks. The code will be available at https://github.com/51cloud/DRDN. Wenxuan Liu 0008, Xian Zhong, Zhuo Zhou, Kui Jiang, Zheng Wang 0007, Chia-Wen Lin |
IEEE Trans. Image Process. | 1 |
| 2022 | VCD: View-Constraint Disentanglement for Action RecognitionabstractAction recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is tantamount to track the view-invariant representation as it evolves with the performed action. However, the views’ and actions’ performance always complement each other, once simply looking for the view-invariant representation may cause some behavior information to be lost. In this paper, we propose the View-Constraint Disentanglement (VCD) framework for cross-view action recognition. Specifically, Constraint Disentanglement Module (CDM) is utilized to learn an action-invariant representation by discretizing view-specific representation and its normal distribution, which resolves the entangled relationship between view and action. Moreover, a novel Adaptive Distribution Module (ADM) is intended to befit enhance the high-correlation viewpoint variation information and refine the suitable weight. Extensive experiments are conducted on public benchmarks, indicating that our approach achieves better performance than other state-of-the-art approaches. Xian Zhong, Zhuo Zhou, Wenxuan Liu 0008, Kui Jiang, Xuemei Jia, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 3 |
| 2022 | Graph-Based Structural Attributes for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID), which aims to identify the same vehicle across different surveillance cameras, is a significant application in urban operation and security. Although the existing methods have noticed the importance of local features, near-duplicated cases are still hard to be handled. The reason lies that the attribute features and personalized structure of vehicles are often ignored. In this paper, we propose a graph-based structural attribute network (GSAN), which contains an attribute feature extraction module (AFEM) and a dual-grained structural relation module (DSRM). The AFEM aims to obtain attribute features of vehicles with structural information between attributes, and the DSRM aims to make the attribute features able to represent structural relation information between parts and attributes. The result on representative datasets shows that GSAN achieves competitive improvements over the state-of-the-art methods. We also collect a dataset of vehicle images with attribute annotations. Our dataset and code are released at https://github.com/HappyBoBo0331/GSAN. Rongbo Zhang, Xian Zhong, Xiao Wang 0029, Wenxin Huang, Wenxuan Liu 0008 |
ICME | 5 |
| 2022 | Patching Your Clothes: Semantic-Aware Learning for Cloth-Changed Person Re-Identification
Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang |
MMM (2) | 4 |
| 2022 | Actor-Aware Alignment Network for Action RecognitionabstractAction recognition has attracted growing interest recently. It suffers from the problem that complex and diverse environments may disturb the extraction of action features. Existing methods propose to explore the temporal associations to alleviate the issue. However, they cannot handle long-range frames, and the rigid techniques are powerless against the differences caused by the deformation of the actors. To this end, we propose the Actor-Aware Alignment Network (A$^{3}$Net), which helps locate the action region. Specifically, through the intra-snippet correction, we afford the local segment alignment frames. The inter-snippet is designed to rectify the results, avoiding the occlusion situation that may appear in the local snippet. In addition, we consider intra-alignment short-range adjustive frames and long-range context frames between different snippets, which allows our A$^{3}$Net network to achieve the effect of focusing on long-range frame information. Multiple Reasoning Attention (MRA) modules are introduced to integrate features along the temporal dimension to keep the video spatio-temporal consistent. Extensive experiments conducted on three widely-used public benchmarks,UCF101,HMDB51, andInfAR, indicate that the excellence of our approach over other state-of-the-art models in wild scenarios. Wenxuan Liu 0008, Xian Zhong, Xuemei Jia, Kui Jiang, Chia-Wen Lin |
IEEE Signal Process. Lett. | 1 |
| 2022 | Complementary Data Augmentation for Cloth-Changing Person Re-IdentificationabstractThis paper studies the challenging person re-identification (Re-ID) task under the cloth-changing scenario, where the same identity (ID) suffers from uncertain cloth changes. To learn cloth- and ID-invariant features, it is crucial to collect abundant training data with varying clothes, which is difficult in practice. To alleviate the reliance on rich data collection, we reinforce the feature learning process by designing powerful complementary data augmentation strategies, including positive and negative data augmentation. Specifically, the positive augmentation fulfills the ID space by randomly patching the person images with different clothes, simulating rich appearance to enhance the robustness against clothes variations. For negative augmentation, its basic idea is to randomly generate out-of-distribution synthetic samples by combining various appearance and posture factors from real samples. The designed strategies seamlessly reinforce the feature learning without additional information introduction. Extensive experiments conducted on both cloth-changing and -unchanging tasks demonstrate the superiority of our proposed method, consistently improving the accuracy over various baselines. Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang |
IEEE Trans. Image Process. | 4 |
| 2021 | Unsupervised Vehicle Search in the Wild: A New BenchmarkabstractIn urban surveillance systems, finding a specific vehicle in video frames efficiently and accurately has always been an essential part of traffic supervision and criminal investigation. Existing studies focus on vehicle re-identification (re-ID), but vehicle search is still underexploited. These methods depend on the locations of many vehicles (bounding boxes) that are not available in most real-world applications. Therefore, the unsupervised joint study of vehicle location and identification for the observed scene is a pressing need. Inspired by person search, we conduct a study on the vehicle search while considering four main discrepancies among them, summarized as: 1) It is challenging to select the candidate regions for the observed vehicle due to the perspective differences (front or side); 2) The sides of the same type of vehicles are almost the same, resulting in smaller inter-class; 3) Lacking satisfied dataset for vehicle search to meet the practical scenarios; 4) Supervised search publishing methods rely on datasets with expensive annotations. To address these issues, we have established a new vehicle search dataset. We design an unsupervised framework on this benchmark dataset to generate pseudo labels for further training existing vehicle re-ID or person search models. Experimental results reveal that these methods turn less effective on vehicle search tasks. Therefore, the vehicle search task needs to be further developed, and this dataset can advance the research of vehicle search. Https://github.com/zsl1997/VSW. Xian Zhong, Xiao Wang 0029, Kui Jiang, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ACM Multimedia | 5 |
| 2021 | Random Walk Erasing with Attention Calibration for Action Recognition
Yuze Tian, Xian Zhong, Wenxuan Liu 0008, Xuemei Jia, Mang Ye |
PRICAI (3) | 3 |
| 2021 | Subspace Enhancement and Colorization Network for Infrared Video Action Recognition
Xian Zhong, Wenxuan Liu 0008, Zhengwei Yang 0001, Luo Zhong |
PRICAI (3) | 3 |
| 2021 | Attention-guided image captioning with adaptive global and local feature fusion
Xian Zhong, Guozhang Nie, Wenxin Huang, Wenxuan Liu 0008, Chia-Wen Lin |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Local Facial Makeup Transfer via Disentangled Representation
Zhaoyang Sun, Shengwu Xiong 0001, Wenxuan Liu 0008 |
ACCV (4) | 5 |
| 2020 | Visible-infrared Person Re-identification via Colorization-based Siamese Generative Adversarial NetworkabstractWith explosive surveillance data during day and night, visible-infrared person re-identification (VI-ReID) is an emerging challenge due to the apparent cross-modality discrepancy between visible and infrared images. Existing VI-ReID work mainly focuses on learning a robust feature to represent a person in both modalities despite the modality gap cannot be effectively eliminated. Recent research works have proposed various generative adversarial network (GAN) models to transfer the visible modality to another unified modality, aiming to bridge the cross-modality gap. However, they neglect the information loss caused by transferring the domain of visible images which is significant for identification. To effectively address the problems, we observe that key information such as textures and semantics in an infrared image can help to color the image itself and the colored infrared image maintains rich information from infrared image while reducing the discrepancy with the visible image. We therefore propose a colorization-based Siamese generative adversarial network (CoSiGAN) for VI-ReID to bridge the cross-modality gap, by retaining the identity of the colored infrared image. Furthermore, we also propose a feature-level fusion model to supplement the transfer loss of colorization. The experiments conducted on two cross-modality person re-identification datasets demonstrate the superiority of the proposed method compared with the state-of-the-arts. Xian Zhong, Tianyou Lu, Wenxin Huang, Jingling Yuan, Wenxuan Liu 0008, Chia-Wen Lin |
ICMR | 5 |