Wenxin Huang

dblp:126/5563 · DBLP profile ↗
← Back
55ranked-venue papers
6as first author
42since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing
abstract
Knowledge Tracing (KT) aims to dynamically model a student’s mastery of knowledge concepts based on their historical learning interactions. Most current methods rely on single-point estimates, which cannot distinguish true ability from outburst or carelessness, creating ambiguity in judging mastery. To address this issue, we propose a Knowledge Mastery-State Disambiguation for Knowledge Tracing model (KeenKT), which represents a student’s knowledge state at each interaction using a Normal-Inverse-Gaussian (NIG) distribution, thereby capturing the fluctuations in student learning behaviors. Furthermore, we design an NIG-distance-based attention mechanism to model the dynamic evolution of the knowledge state. In addition, we introduce a diffusion-based denoising reconstruction loss and a distributional contrastive learning loss to enhance the model’s robustness. Extensive experiments on six public datasets demonstrate that KeenKT outperforms state-of-the-art KT models in terms of prediction accuracy and sensitivity to behavioral fluctuations. The proposed method yields the maximum AUC improvement of 5.85% and the maximum ACC improvement of 6.89%.
Zhifei Li 0009, Lifan Chen, Jiali Yi, Xiaoju Hou, Wenxin Huang, Miao Zhang 0036, Kui Xiao
AAAI6
2026 Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer
abstract
Action recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition robustness. To address this, we introduce a multi-view formulation tailored for UAV altitudes and empirically uncover a distinctive partial order among views, where recognition accuracy consistently declines as altitude increases. This key observation motivates the proposed Aero Partial Order Guided Network (Aerorder), which explicitly models and exploits the hierarchical structure of UAV views to enhance cross-altitude action recognition. Aerorder comprises three main components: (1) a View Partition (VP) module that groups views by altitude using the head-to-body ratio; (2) an Order-aware Feature Decoupling (OFD) module that disentangles action-relevant and view-specific representations under partial order guidance; and (3) an Action Partial Order Guide (APOG) that progressively transfers knowledge from easier (low-altitude) to harder (high-altitude) views. Extensive experiments on Drone-Action, MOD20, and UAV validate the superiority of Aerorder, achieving consistent improvements over state-of-the-art methods, up to 4.7% and 1.3% gains on Drone-Action and MOD20, respectively.
Wenxuan Liu 0008, Zhuo Zhou, Xuemei Jia, Siyuan Yang 0001, Wenxin Huang, Xian Zhong, Chia-Wen Lin
AAAI5
2026 LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models
abstract
Jian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Xi Yang, Changjinli, Yonghua Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Changjin Li, Yonghua Lin
ACL (1)5
2026 See what you seek: Semantic contextual integration for cloth-changing person re-identification
Wenxin Huang, Xian Zhong, Jingling Yuan, Alex Chichung Kot
Pattern Recognit.1
2026 Robust mixed-degradation person Re-identification via structural consistency distillation
Wenxin Huang, Wenxuan Liu 0008, Xuemei Jia, Xian Zhong
Pattern Recognit.2
2026 TCP: Text-Guided Cascade Network for Pedestrian Crossing Intention Prediction
abstract
Pedestrian crossing intention prediction is crucial for ensuring safety in intelligent transportation systems, especially in autonomous driving scenarios. Most existing methods rely primarily on visual information; however, the quality of visual data deteriorates significantly at long distances due to limited resolution. Although multi-modal approaches can mitigate this issue by incorporating additional sensory data, they inevitably introduce extra computational overhead. To address these challenges, we propose a lightweight cascaded model for pedestrian crossing intention prediction based on text-trajectory alignment. The model employs a cascaded architecture that jointly performs coordinate and intention prediction, while leveraging a pre-trained large language model (LLM) to generate textual descriptions of videos, thereby enriching trajectory features. Furthermore, a center-aware classification module is integrated to enhance inter-class separability and intra-class compactness. Extensive experiments onJAADandPIEdemonstrate state-of-the-art performance: our method achieves 91% accuracy onPIEand 89% onJAAD, matching or surpassing recent multi-modal approaches with substantially fewer inputs. The source code will be released athttps://github.com/xyhhappy/TCP-prediction
Wenxuan Liu 0008, Wenxin Huang, Ryan Wen Liu, Xian Zhong
IEEE Trans. Intell. Transp. Syst.3
2025 StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models
abstract
The rapid development of multimodal large language models (MLLMs) has positioned visual storytelling as a crucial area in content creation. However, existing models often struggle to maintain temporal, spatial, and narrative coherence across image sequences, and they frequently lack the depth and engagement of human-authored stories. To address these challenges, we propose Story with Large Language-and-Vision Alignment (StoryLLaVA), a novel framework for enhancing visual storytelling. Our approach introduces a topic-driven narrative optimizer that improves both the training data and MLLM models by integrating image descriptions, topic generation, and GPT-4-based refinements. Furthermore, we employ a preference-based ranked story sampling method that aligns model outputs with human storytelling preferences through positive-negative pairing. These two phases of the framework differ in their training methods: the former uses supervised fine-tuning, while the latter incorporates reinforcement learning with positive and negative sample pairs. Experimental results demonstrate that StoryLLaVA outperforms current models in visual relevance, coherence, and fluency, with LLM-based evaluations confirming the generation of richer and more engaging narratives. The enhanced dataset and model will be made publicly available soon.
Zhiding Xiao, Wenxin Huang, Xian Zhong
COLING3
2025 A Cooperative Safety-Enhanced Control Framework for Driving Assistance in the Internet of Vehicles
abstract
For the Internet of Vehicles (IoV), driving safety applications require reliable and up-to-date knowledge of the state of vehicles and traffic. A single vehicle cannot meet all the reliability requirements because of the limited capability of information acquisition. Thus, cooperation among vehicles for information sharing is essential. However, due to the high dynamic network topology and harsh channel conditions, maintaining long-term cooperation is not feasible. Only the messages that most affect the driving state can obtain the transmission opportunity for avoiding network congestion. In this paper, we propose a cooperative safety-enhanced control framework (SCF). This framework concentrates on the construction of dynamic and adaptive cooperation among vehicles and evaluates the key feature parameters to achieve an optimal safety utility for feedback control over the driving state. We construct a general multi-layer solution framework for driving assistance in SCF. First, we construct multiple temporary cooperative platoons to coordinate adjacent vehicles and realize a relatively uniform driving state. The cooperative platoon maintains short-term stability for vehicle sensing and tracing. Second, we propose a utility evaluation model for extracting the key feature parameters related to the driving state, which is the basis of the optimization for message transmission and driving control. Third, we design a two-level joint optimization mechanism for the deep fusion of the multi-source heterogeneous data to maximize the total utility of driving safety. Finally, we propose an adaptive feedback control model for the cooperative platoon, which actively adjusts the driving control strategy and the message transmission strategy in a real-time manner. Then the optimal driving assistant decision can be made. Extensive simulation results show that SCF outperforms related communication mechanisms for safe driving in the IoV, demonstrating that SCF can effectively enhance driving assistance control.
Yan Zhang 0077, Chao Yang 0043, Zhifei Li 0009, Kui Xiao, Miao Zhang 0036, Wenxin Huang, Hao Chen 0134, Jianhua Song, Xian Zhong, Haobo Ma
ICMR7
2025 YES: You should Examine Suspect cues for low-light object detection
Wenxin Huang, Xian Zhong
Comput. Vis. Image Underst.2
2025 Semi-supervised lithography hotspot detection based on feature fusion and residual attention
Xinzhong Xiao, Wenxin Huang, Ruijun Ma 0002, Fuxin Tang, Pan Qi, Huaguo Liang
Integr.3
2025 Uncertainty-Aware With Adaptive Geometric Correction for Multimodal Land-Cover Classification
abstract
Land cover classification (LCC) is a fundamental task in remote sensing and geographic information science. Multi-modal fusion has shown great potential for enhancing LCC performance, for example, by combining optical and synthetic aperture radar (SAR) imagery to leverage their complementary strengths. However, two key challenges hinder effective fusion:1) local geometric mismatches caused by distinct imaging geometries, and2) inconsistent reliability (the ability of a modality to deliver accurate and stable information) in LCC arising from different modalities and their acquisition conditions. To address these issues, we propose Uncertainty-Aware Fusion with Adaptive Geometric Correction (UAG), which comprises three main components. First, the Adaptive Geometric Correction Module (AGCM) applies learnable pixel shifts to establish bidirectional local correlations between multiscale optical and SAR features, thereby mitigating spatial inconsistencies. Second, the Adaptive Uncertainty-Aware Dynamic Fusion Module (ADFM) employs evidential deep learning to model uncertainty, defined as the extent of reliability deficiency, for each modality using the Dirichlet distribution and subjective logic, enabling confidence-aware feature weighting. Third, a lightweight multiscale decoder integrates hierarchical features through a hybrid MLP-convolutional architecture, improving both segmentation efficiency and accuracy. We evaluate UAG on WHU-OPT-SAR and DFC23 datasets, where experimental results demonstrate substantial improvements over state-of-the-art methods. The code will be released at https://github.com/cccwbin/UAGNet.
Xu Wang 0015, Yi Xiao 0003, Wenxin Huang, Bihan Wen, Xian Zhong
IEEE Trans. Geosci. Remote. Sens.4
2025 On-Orbit Characterization of FY-4B AGRI Calibration Components for Solar Diffuser Degradation Estimation
abstract
The Advanced Geostationary Radiation Imager (AGRI) of the FY-4B satellite uses a set of reflective solar band (RSB) onboard calibration components, with a solar diffuser (SD) and an SD reflectance degradation monitor (SDRDM). The SDRDM is primarily used to monitor the degradation of the SD bidirectional reflectance distribution function (BRDF) during the satellite’s on-orbit period and to provide degradation correction coefficients. In this work, the degradation coefficients of the SD BRDF in the visible near-infrared (NIR) band are calculated on the basis of the observation dataset obtained with the SDRDM from August 2021 to March 2023. The time series of calculated SD degradation coefficients exhibits notable anomalous fluctuations due to the cosine error in the incidence angle associated with the placement of the calibration device and the variation in stray light with the illumination angle. A systematic degradation coefficient evaluation method based on the optical mechanical structure characteristics and working principle of the SDRDM is proposed. The SD BRDF degradation coefficient is obtained by combining time-series data with the same geometric lighting conditions and interband degradation coefficient ratios and removing the relative angular changes and associated effects. With the stable CH3 band as the reference band, geometric factors are reconstructed for other bands for subsequent SD degradation monitoring. Finally, the monitoring results are applied to observed SD data for verification, and from August 2021 to March 2023, the SD BRDF degradation was 2.3% at 0.45–$0.55~\mu $m and almost no degradation at 0.55–$0.90~\mu $m.
Xiaolong Si, Xiuju Li, Shiwei Bao, Changpei Han, Wenxin Huang
IEEE Trans. Geosci. Remote. Sens.7
2025 Local-Global Sparse Transformer for Road Extraction From Remote Sensing Imagery
abstract
Accurate road extraction from high-resolution remote sensing imagery is crucial for applications such as road network generation, urban planning, autonomous driving, and military operations. Although deep learning has significantly advanced segmentation performance, road extraction remains challenging because roads exhibit geometric and structural characteristics that differ markedly from general objects. We propose a local–global sparse transformer network (LGST-Net) that leverages sparse attention (SA) to capture both local details and global context. First, a coarse-to-fine feature-enhanced preprocessing (C2F-FEP) module extracts low-level features at a fine-grained scale during the model’s initial stage. Next, we design the LGST backbone, which comprises three novel components: local shift-mask SA (LSM-SA), global compressed sparse attention (GC-SA), and a spatial-channel gate. Extensive experiments onDeepGlobeandRoadTracerdatasets demonstrate that LGST-Net outperforms state-of-the-art methods. Our code will be publicly available athttps://github.com/JaymeWX/Road_LGST-Net
Xu Wang 0015, Wenxin Huang, Xian Zhong
IEEE Trans. Geosci. Remote. Sens.3
2024 Localization of Image Splicing Under Segment Anything Model With Integrated Compression and Edge Artifacts
abstract
The localization of image splicing involves identifying pixels in an image that have been spliced from other images, necessitating the discernment of splicing features. Despite significant advancements driven by the rise of social media and deep learning, existing methods exhibit limitations, often neglecting the integration of coarse and precise features and lacking the ability to understand objects. This leads to erroneous predictions in identifying spliced regions. This paper proposes Segment Anything Model with Integrated Compression and Edge artifacts (SAM-ICE) for the localization of image splicing, addressing these limitations by fusing forged edge features and compression artifact features. Leveraging SAM’s object understanding ability, our method identifies spliced regions using the fused features as guidance. Specifically, we employ Edge Artifact Extractor (EAE) to extract fine high-frequency edge splicing features and Compression Artifact Extractor (CAE) to extract coarse compression artifact features. By combining these features, our method utilizes coarse-fine features to accurately pinpoint the spliced portions of the image. Experimental results demonstrate the superior accuracy, robustness, and generalizability of our method compared to the state-of-the-arts.
Ruhao Zhao, Xian Zhong, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007
ICIP5
2024 Towards Low-latency Event-based Visual Recognition with Hybrid Step-wise Distillation Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have garnered significant attention for their low power consumption and high biological interpretability. Their rich spatio-temporal information processing capability and event-driven nature make them ideally well-suited for neuromorphic datasets. However, current SNNs struggle to balance accuracy and latency in classifying these datasets. In this paper, we propose Hybrid Step-wise Distillation (HSD) method, tailored for neuromorphic datasets, to mitigate the notable decline in performance at lower time steps. Our work disentangles the dependency between the number of event frames and the time steps of SNNs, utilizing more event frames during the training stage to improve performance, while using fewer event frames during the inference stage to reduce latency. Nevertheless, the average output of SNNs across all time steps is susceptible to individual time step with abnormal outputs, particularly at extremely low time steps. To tackle this issue, we implement Step-wise Knowledge Distillation (SKD) module that considers variations in the output distribution of SNNs at each time step. Empirical evidence demonstrates that our method yields competitive performance in classification tasks on neuromorphic datasets, especially at lower time steps. Our code will be available at: https://github.com/hsw0929/HSD.
Xian Zhong, Shengwang Hu, Wenxuan Liu 0008, Wenxin Huang, Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001
ACM Multimedia4
2024 Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight Adjustment
abstract
The sharp increase in data-related expenses has motivated research into condensing datasets while retaining the most informative features. Dataset distillation has thus recently come to the fore. This paradigm generates synthetic datasets that are representative enough to replace the original dataset in training a neural network. To avoid redundancy in these synthetic datasets, it is crucial that each element contains unique features and remains diverse from others during the synthesis stage. In this paper, we provide a thorough theoretical and empirical analysis of diversity within synthesized datasets. We argue that enhancing diversity can improve the parallelizable yet isolated synthesizing approach. Specifically, we introduce a novel method that employs dynamic and directed weight adjustment techniques to modulate the synthesis process, thereby maximizing the representativeness and diversity of each synthetic instance. Our method ensures that each batch of synthetic data mirrors the characteristics of a large, varying subset of the original dataset. Extensive experiments across multiple datasets, including CIFAR, Tiny-ImageNet, and ImageNet-1K, demonstrate the superior performance of our method, highlighting its effectiveness in producing diverse and representative synthetic datasets with minimal computational expense. Our code is available at https://github.com/AngusDujw/Diversity-Driven-Synthesis.
Jiawei Du 0002, Xin Zhang 0092, Wenxin Huang, Joey Tianyi Zhou
NeurIPS4
2024 Multi-view hyperspectral image classification via weighted sparse representation
Zhifei Li 0009, Wenxin Huang
Multim. Tools Appl.4
2024 ICLR: Instance Credibility-Based Label Refinement for label noisy person re-identification
Xian Zhong, Xuemei Jia, Wenxin Huang, Wenxuan Liu 0008, Shuaipeng Su, Xiaohan Yu 0001, Mang Ye
Pattern Recognit.4
2024 TCSA: Efficient Localization of Busy-Wait Synchronization Bugs for Latency-Critical Applications
abstract
Busy-wait synchronization is often used for latency-critical applications to ensure low latency. Unfortunately, its performance bugs due to thread contention may lead to request failures or even system crashes. Localizing the performance bugs of busy-wait synchronization is not trivial because we have to pinpoint the exact moment of occurrence from a relatively long measurement period and simultaneously identify candidate busy-wait threads from numerous concurrent threads. Existing methods often rely on hotspot-driven analysis of lock-related functions, but they still need extensive manual work to localize busy-wait threads. This paper proposes timing call stack analysis (TCSA), an efficient approach to localizing busy-wait synchronization bugs. The key idea is to time-serialize the function call stacks of applications and identify consecutive identical call stacks to catch busy-wait threads. TCSA can handle any application regardless of its programming language and identify various busy-wait patterns, including spinlocks, chaining spinlocks, futexes, and safepoint checks within the Java Virtual Machine. Compared to the state-of-the-art, TCSA can effectively diminish the quantity of examined records (e.g., threads and functions) by 1 to 3 orders of magnitude. TCSA has been deployed to a large cloud service provider, demonstrating its effectiveness, efficiency, and practicality in four real latency-critical applications.
Ning Li 0054, Jianmei Guo, Bo Huang 0002, Chengdong Li, Wenxin Huang
IEEE Trans. Parallel Distributed Syst.7
2023 Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation
abstract
Video captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success. However, they focus on exploiting foreground semantics, ignoring the potential negative impact of video background disturbance to caption generation, i.e., the entities and the actions are misjudged by a similar video background. To ameliorate this issue, we propose Entity-Action Relocation (EAR) to enhance the adaptability of entities and actions to various backgrounds by giving them the background. Specifically, for an extracted original video feature, we construct a mixed background for all entities and actions to form a distracting video feature sample. After that, contrastive learning is applied to pull the generated caption of the original representations and of the distracting representations closer, and to push the former away from the generated caption of other videos, explicitly concentrating on the entities and actions of the current video scene. Extensive experiments on two public datasets (MSR-VTT and MSVD) demonstrate that dealing with background disturbance for video can obtain a competitive caption generation effect.
Xian Zhong, Shuqin Chen, Wenxin Huang, Lin Li 0001
ICASSP5
2023 Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation
abstract
While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to handle the more realistic multi-target domain adaptive semantic segmentation (MT-DASS) tasks. To solve this problem, we propose a Bi-Alignment framework based on Transformation (BAT). Specifically, we employ the Fourier style transform to convert the style of the source domain to that of the target domain without training any style transfer networks. In this way, we transform the single-peak distributed source domain into a multi-peak distribution that resembles the multi-target domain. Then, we perform fine-grained global and local dual distribution alignment between the same style of source-target domain pairs to achieve a multi-to-multi distribution alignment. Finally, self-training is utilized to further improve the network’s discriminability. Experimental results show that our approach achieves competitive results over state-of-the-art methods.
Xian Zhong, Jing Xiao 0004, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007
ICASSP6
2023 Neighborhood Information-Based Label Refinement for Person Re-Identification with Label Noise
abstract
The existing excellent person re-identification (Re-ID) model is still affected by the samples with the incorrect labels. It is difficult to accurately annotate person images in the real scene, resulting in label noise. To avoid fitting to the noisy labels, a common solution in Re-ID is to replace the original label with the label predicted by the deep model. Unfortunately, similar samples of different identities with the same label are due to label noise, which is challenging for the model to distinguish them. Neighborhood information can optimize noisy labels through neighborhood labels and similarity between samples. This paper proposes a label refinement module based on neighborhood information (LRNI) for person Re-ID with label noise. Specifically, we first use the pre-trained model to extract features and calculate the similarity between samples. Rather than treating samples as isolated, the similarity used as label propagation weight and neighborhood labels are combined to optimize noisy labels. To further reduce the influence of label noise, we design a hard sample re-weighting (HSR) strategy to balance the learning of noisy and boundary samples. Experimental results under different noise settings demonstrate our method's effectiveness in the person Re-ID task.
Xian Zhong, Shuaipeng Su, Wenxuan Liu 0008, Xuemei Jia, Wenxin Huang, Mengdie Wang
ICASSP5
2023 Background-Weakening Consistency Regularization for Semi-Supervised Video Action Detection
abstract
Consistency-based techniques have produced state-of-the-art results in semi-supervised action detection. When the model false detects the dynamic information in the background as an action, spatio-temporal consistency calculations can hardly reflect this false detection result. We consider weakening the dynamic information in the augmented video background to reduce its spatio-temporal consistency with the dynamic information in the original video background. Thus we propose a Background-Weakening with Calibration Constraint (BWCC) framework, which highlights the negative impact of information in the background of false detection by calculating the consistency of the predictions of the background weakened video and the original video. Specifically, Background Weaken (BW) module judges the foreground and background of the video based on the initial predictions of the model and makes adjustments to the video background. To mitigate the effects of the misjudgments result in weakened action pixels, we additionally introduce a model that does not undergo background weakening to aid training through Calibration Constraint (CC) module. Our approach achieves competitive performance over existing leading approaches on two action detection datasets, UCF101-24 and JHMDB-21.
Xian Zhong, Aoyu Yi, Wenxuan Liu 0008, Wenxin Huang, Chengming Zou, Zheng Wang 0007
ICASSP4
2023 CSEC: A Chinese Semantic Error Correction Dataset for Written Correction
Wenxin Huang, Mengxiang Wang, Guangya Liu, Jianxing Yu, Huaijie Zhu, Jian Yin 0001
ICONIP (5)1
2023 SCPNet: Self-constrained parallelism network for keypoint-based lightweight object detection
Xian Zhong, Mengdie Wang, Wenxuan Liu 0008, Jingling Yuan, Wenxin Huang
J. Vis. Commun. Image Represent.5
2023 Visual Exposes You: Pedestrian Trajectory Prediction Meets Visual Intention
abstract
Pedestrian trajectory prediction in multiple scenarios is of immense importance in autonomous driving and disentanglement of human behavior but is limited in catching human intention and initiative. Most previous works tend to predict the trajectory using only 2D coordinates, which generally cause two common problems: a) Overlooking the subjective initiative, including sudden swerve and erratic movement; b) A potential challenge called abnormal collision caused by unlabeled pedestrians on dataset is not being identified and resolved, which would ruin the model prediction. To break those limitations, we introduce visual localization and orientation as Visual Intention Knowledge to help the trajectory prediction, which is learned directly from visual scenarios. It benefits to comprehend human intention and formulates decision-making processes. Moreover, by learning from the visual information and decision-making policy, we construct the Visual Intention Knowledge associated spatio-temporal Transformer (VIKT) to predict human trajectory by combining the intention knowledge with the novel Transformer. Extensive experimental results demonstrate that our VIKT model could achieve competitive performance by the Visual Intention Knowledge through optimizing the model prediction compared with state-of-the-art methods in terms of prediction accuracy on ETH/UCY and SDD benchmarks.
Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Kui Jiang, Ryan Wen Liu, Zheng Wang 0007
IEEE Trans. Intell. Transp. Syst.4
2023 Graph Complemented Latent Representation for Few-Shot Image Classification
abstract
Few-shot learning is a tough topic to solve since obtaining a large number of training samples in real applications is challenging. It has attracted increasing attention recently. Meta-learning is a prominent way to address this issue, intending to adapt predictors as base-learners to new tasks swiftly. However, a key challenge of meta-learning is its lack of expressive capacity, which stems from the difficulty of extracting general information from a small number of training samples. As a result, the generalizability of meta-learners trained from high-dimensional parameter spaces is frequently limited. To learn a better representation, we propose a graph complemented latent representation (GCLR) network for few-shot image classification. In particular, we embed the representation into a latent space, in which the latent codes are reconstructed using variational information to enrich the representation. In this way, the latent representation can achieve better generalizability. Another benefit is that, because the latent space is formed using variational inference, it cooperates well with various base-learners, boosting robustness. To make full use of the relation between samples in each category, a graph neural network (GNN) is also incorporated to improve relation mining. Consequently, our end-to-end framework delivers competitive performance on three few-shot learning benchmarks for image classification.
Xian Zhong, Mang Ye, Wenxin Huang, Chia-Wen Lin
IEEE Trans. Multim.4
2023 Beyond the Parts: Learning Coarse-to-Fine Adaptive Alignment Representation for Person Search
abstract
Person search is a time-consuming computer vision task that entails locating and recognizing query people in scenic pictures. Body components are commonly mismatched during matching due to position variation, occlusions, and partially absent body parts, resulting in unsatisfactory person search results. Existing approaches for extracting local characteristics of the human body using keypoint information are unable to handle the search job when distinct body parts are misaligned, ignoring to exploit multiple granularities, which is crucial in the person search process. Moreover, the alignment learning methods learn body part features with fixed and equal weights, ignoring the beneficial contextual information, e.g., the umbrella carried by the pedestrian, which supplements compelling clues for identifying the person. In this paper, we propose a Coarse-to-Fine Adaptive Alignment Representation (CFA 2 R) network for learning multiple granular features in misaligned person search in the coarse-to-fine perspective. To exploit more beneficial body parts and related context of the cropped pedestrians, we design a Part-Attentional Progressive Module (PAPM) to guide the network to focus on informative body parts and positive accessorial regions. Besides, we propose a Re-weighting Alignment Module (RAM) shedding light on more contributive parts instead of treating them equally. Specifically, adaptive re-weighted but not fixed part features are reconstructed by Re-weighting Reconstruction module, considering that different parts serve unequally during image matching. Extensive experiments conducted on CUHK-SYSU and PRW datasets demonstrate competitive performance of our proposed method.
Wenxin Huang, Xuemei Jia, Xian Zhong, Xiao Wang 0029, Kui Jiang, Zheng Wang 0007
ACM Trans. Multim. Comput. Commun. Appl.1
2022 VCD: View-Constraint Disentanglement for Action Recognition
abstract
Action recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is tantamount to track the view-invariant representation as it evolves with the performed action. However, the views’ and actions’ performance always complement each other, once simply looking for the view-invariant representation may cause some behavior information to be lost. In this paper, we propose the View-Constraint Disentanglement (VCD) framework for cross-view action recognition. Specifically, Constraint Disentanglement Module (CDM) is utilized to learn an action-invariant representation by discretizing view-specific representation and its normal distribution, which resolves the entangled relationship between view and action. Moreover, a novel Adaptive Distribution Module (ADM) is intended to befit enhance the high-correlation viewpoint variation information and refine the suitable weight. Extensive experiments are conducted on public benchmarks, indicating that our approach achieves better performance than other state-of-the-art approaches.
Xian Zhong, Zhuo Zhou, Wenxuan Liu 0008, Kui Jiang, Xuemei Jia, Wenxin Huang, Zheng Wang 0007
ICASSP6
2022 Graph-Based Structural Attributes for Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID), which aims to identify the same vehicle across different surveillance cameras, is a significant application in urban operation and security. Although the existing methods have noticed the importance of local features, near-duplicated cases are still hard to be handled. The reason lies that the attribute features and personalized structure of vehicles are often ignored. In this paper, we propose a graph-based structural attribute network (GSAN), which contains an attribute feature extraction module (AFEM) and a dual-grained structural relation module (DSRM). The AFEM aims to obtain attribute features of vehicles with structural information between attributes, and the DSRM aims to make the attribute features able to represent structural relation information between parts and attributes. The result on representative datasets shows that GSAN achieves competitive improvements over the state-of-the-art methods. We also collect a dataset of vehicle images with attribute annotations. Our dataset and code are released at https://github.com/HappyBoBo0331/GSAN.
Rongbo Zhang, Xian Zhong, Xiao Wang 0029, Wenxin Huang, Wenxuan Liu 0008
ICME4
2022 Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene Understanding
abstract
Scene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and methods: 1) Manually synthetic rainy samples with empirically settings and human subjective assumptions; 2) Limited rainy conditions, including the rain patterns, intensity, and degradation factors; 3) Separated training manners for image deraining and semantic segmentation. To break these limitations, we pioneer a real, comprehensive, and well-annotated scene understanding dataset under rainy weather, named Rainy WCity. It covers various rain patterns and their bring-in negative visual effects, covering wiper, droplet, reflection, refraction, shadow, windshield-blurring, etc. In addition, to alleviate dependence on paired training samples, we design an unsupervised contrastive learning network for real image deraining and the final rainy scene semantic segmentation via multi-task joint optimization. A comprehensive comparison analysis is also provided, which shows that scene understanding in rainy weather is a largely open problem. Finally, we summarize our general observations, identify open research challenges, and point out future directions.
Xian Zhong, Shidong Tu, Xianzheng Ma, Kui Jiang, Wenxin Huang, Zheng Wang 0007
IJCAI5
2022 Patching Your Clothes: Semantic-Aware Learning for Cloth-Changed Person Re-Identification
Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang
MMM (2)5
2022 Global Temporal Attention Optimization for Human Trajectory Prediction
abstract
Predicting human trajectory is one of the key knowledge required for autonomous driving and social robots in real scenarios. Recent studies based on Transformer networks have shown a great ability to model social behaviors. As far as we know, global trajectory information has an essential influence on prediction at a certain step. However, these methods only rely on the previous trajectory states/attention but ignore the important following states/attention of the trajectory for each pedestrian, which will generally collapse on some irregular movements (e.g. acceleration, deceleration, and motionless). To solve this issue, we propose a Global Temporal Attention optimization model (GTAO), which activates the utilization of the following states/attention of the trajectory, and jointly and iteratively optimizes the preliminary trajectory prediction through a global temporal attention (GTA) module. To effectively address the decline in the generalizability and abnormal processing of the model, we further introduce global temporal guidance (GTG) module to instruct the GTA to learn the features closer to realistic trajectories. Experimental results on commonly used real-world human trajectory prediction datasets (ETH and UCY) indicate that our GTAO can achieve better performance in terms of prediction accuracy.
Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Zheng Wang 0007
SMC5
2022 Grayscale Enhancement Colorization Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is an emerging and challenging cross-modality image matching problem because of the explosive surveillance data in night-time surveillance applications. To handle the large modality gap, various generative adversarial network models have been developed to eliminate the cross-modality variations based on a cross-modal image generation framework. However, the lack of point-wise cross-modality ground-truths makes it extremely challenging to learn such a cross-modal image generator. To address these problems, we learn the correspondence between single-channel infrared images and three-channel visible images by generating intermediate grayscale images as auxiliary information to colorize the single-modality infrared images. We propose a grayscale enhancement colorization network (GECNet) to bridge the modality gap by retaining the structure of the colored image which contains rich information. To simulate the infrared-to-visible transformation, the point-wise transformed grayscale images greatly enhance the colorization process. Our experiments conducted on two visible-infrared cross-modality person re-identification datasets demonstrate the superiority of the proposed method over the state-of-the-arts.
Xian Zhong, Tianyou Lu, Wenxin Huang, Mang Ye, Xuemei Jia, Chia-Wen Lin
IEEE Trans. Circuits Syst. Video Technol.3
2022 Complementary Data Augmentation for Cloth-Changing Person Re-Identification
abstract
This paper studies the challenging person re-identification (Re-ID) task under the cloth-changing scenario, where the same identity (ID) suffers from uncertain cloth changes. To learn cloth- and ID-invariant features, it is crucial to collect abundant training data with varying clothes, which is difficult in practice. To alleviate the reliance on rich data collection, we reinforce the feature learning process by designing powerful complementary data augmentation strategies, including positive and negative data augmentation. Specifically, the positive augmentation fulfills the ID space by randomly patching the person images with different clothes, simulating rich appearance to enhance the robustness against clothes variations. For negative augmentation, its basic idea is to randomly generate out-of-distribution synthetic samples by combining various appearance and posture factors from real samples. The designed strategies seamlessly reinforce the feature learning without additional information introduction. Extensive experiments conducted on both cloth-changing and -unchanging tasks demonstrate the superiority of our proposed method, consistently improving the accuracy over various baselines.
Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang
IEEE Trans. Image Process.5
2021 Part-Aligned Network with Background for Misaligned Person Search
abstract
Person search is a significant computer vision task that requires addressing person detection and re-identification simultaneously. Body parts are frequently misaligned due to variation poses, occlusions, and partial missing, leading to the unsatisfied results of person search. Existing methods usually extract local features from the human body by the key point information, that cannot tackle the recognition task between a pair of persons with different body parts due to misalignment. Moreover, these methods overlook background information (e.g. the carries and the background reference object) which can also supplement effective features for representing the person. In this paper, we propose a part-aligned network with background (PANB) to address this misalignment issue. To learn local fine-grained features of different body parts, we fine-tune a parsing network to divide the body region into seven parts. In particular, our proposed method considers extracting the background features as the eighth part features to extract more robust representations, which is more rational and efficient. Furthermore, we design a reconstruction method to align the parts existing in both the query image and the cropped gallery image. Extensive experiments show that our proposed method achieves competitive performance on CUHK-SYSU and PRW datasets.
Xian Zhong, Wenxin Huang, Xiao Wang 0029, Jingling Yuan
ICASSP3
2021 Person Retrieval in Physical World
abstract
Person re-identification (re-ID) gains plenty of achievements as a retrieval problem in constrained camera networks. However, most of the researches are concentrated on visual appearance, they still suffer from the complicated environments in unconstrained urban/campus surveillance scenario due to unreliable visual representations with extremely challenging problems as lack of training samples, amounts of irrelevant crowds, etc. Besides, most of the existing person re-ID datasets neglect the physical truth of realistic investigation application: 1) investigators search only few suspects among amounts of crowds. Moreover, he may not go through by every camera in the surveillance area and may appear in the same camera several times; and 2) the corresponding characteristic in multi-space of the same ID can be verified with each other. Therefore, we propose a person retrieval in physical world (PRPW) dataset with large-scale unconstrained surveillance scenario. It contains over 1.4 million bounding boxes, including 20 labeled IDs and numerous irrelevant crowds captured by 86 cameras. Furthermore, over 30,000 records of 20 mobile trajectories are collected in this dataset, and the 20 mobile trajectories are partially overlapped while passing by 86 cameras. Finally, based on two common senses and a verification experiment, we provide a proposal to tackle with PRPW task on the basis of trajectory association which utilizes global optimization to compensate for the errors caused by visual expression on local observation points. The comparison experiments with two typical unsupervised person re-ID methods are implemented on the constructed dataset.
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Xian Zhong
ICME1
2021 Auxiliary Bi-Level Graph Representation for Cross-Modal Image-Text Retrieval
abstract
Image-text retrieval is one of the most common tasks in multimodal retrieval. It suffers from the problem of information imbalance between modalities, which is so-called modality gap. It remains challenging because prior methods cannot bridge the gap reasonably. With the help of scene graph, we start by designing an auxiliary bi-level graph representation (ABGR) pipeline that can fully mine the potential information and reduce the information redundancy. By doing so, each modality will be represented by lexical word graph that carries the main content of the information. Specifically, we design a graph feature enhancement (GFE) module to embed the graph-structured information in a common subspace while exploring the relationship between lexical words. As a result, a better representation for both image and text can be obtained, which helps us to evaluate the similarity between images and texts more reasonably. Experimental results conducted on two benchmark datasets Flickr30K and MS-COCO demonstrate the effectiveness of our proposed model for cross-modal retrieval task.
Xian Zhong, Zhengwei Yang 0001, Mang Ye, Wenxin Huang, Jingling Yuan, Chia-Wen Lin
ICME4
2021 Unsupervised Vehicle Search in the Wild: A New Benchmark
abstract
In urban surveillance systems, finding a specific vehicle in video frames efficiently and accurately has always been an essential part of traffic supervision and criminal investigation. Existing studies focus on vehicle re-identification (re-ID), but vehicle search is still underexploited. These methods depend on the locations of many vehicles (bounding boxes) that are not available in most real-world applications. Therefore, the unsupervised joint study of vehicle location and identification for the observed scene is a pressing need. Inspired by person search, we conduct a study on the vehicle search while considering four main discrepancies among them, summarized as: 1) It is challenging to select the candidate regions for the observed vehicle due to the perspective differences (front or side); 2) The sides of the same type of vehicles are almost the same, resulting in smaller inter-class; 3) Lacking satisfied dataset for vehicle search to meet the practical scenarios; 4) Supervised search publishing methods rely on datasets with expensive annotations. To address these issues, we have established a new vehicle search dataset. We design an unsupervised framework on this benchmark dataset to generate pseudo labels for further training existing vehicle re-ID or person search models. Experimental results reveal that these methods turn less effective on vehicle search tasks. Therefore, the vehicle search task needs to be further developed, and this dataset can advance the research of vehicle search. Https://github.com/zsl1997/VSW.
Xian Zhong, Xiao Wang 0029, Kui Jiang, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007
ACM Multimedia6
2021 Attention-guided image captioning with adaptive global and local feature fusion
Xian Zhong, Guozhang Nie, Wenxin Huang, Wenxuan Liu 0008, Chia-Wen Lin
J. Vis. Commun. Image Represent.3
2021 Occluded suspect search via channel-guided mechanism
Wenxin Huang, Ruimin Hu, Xiao Wang 0029, Chao Liang 0001, Jun Chen 0001
Neural Comput. Appl.1
2021 Trajectory Association for Person Re-identification
Ruimin Hu, Wenxin Huang, Dengshi Li, Xiaochen Wang 0001, Chenhao Hu
Neural Process. Lett.3
2020 Multi-Scale Residual Network for Image Classification
abstract
Multi-scale approach representing image objects at various levels-of-details has been applied to various computer vision tasks. Existing image classification approaches place more emphasis on multi-scale convolution kernels, and overlook multi-scale feature maps. As such, some shallower information of the network will not be fully utilized. In this paper, we propose the Multi-Scale Residual (MSR) module that integrates multi-scale feature maps of the underlying information to the last layer of Convolutional Neural Network. Our proposed method significantly enhances the characteristics of the information in the final classification. Extensive experiments conducted on CIFAR100, Tiny-ImageNet and large-scale CalTech-256 datasets demonstrate the effectiveness of our method compared with Res-Family.
Xian Zhong, Oubo Gong, Wenxin Huang, Jingling Yuan, Ryan Wen Liu
ICASSP3
2020 Dual-Direction Perception and Collaboration Network for Near-Online Multi-Object Tracking
abstract
Backward tracks have been exploited to improve performance of multi-object tracking (MOT). The existing method brings a stable similarity measurement but neglects unreliable detection. Exploiting predictions of forward tracks has emerged as a popular approach to tackle the task of tracking-by-detection. However, it's observed that missing detection has not been solved well enough which would significantly influence the tracking accuracy. Thus, obtaining more proposals from dual-direction tracking and predictions of tracks is concerned to address the problem of missing detection. In this paper, we propose a dual-direction perception and collaboration network (DPCNet) for MOT that exploits forward and backward tracking to collaboratively track objects. It collects candidates from the dual directions so that they can complement each other in different scenarios. Moreover, we propose a near-online tracking model based on DPCNet to improve the efficiency, which batches the tracking and makes forward and backward tracking in parallel. Experiments conducted on MOT challenge benchmarks demonstrate that the proposed method outperforms the state-of-the-arts.
Xian Zhong, Weijian Ruan, Wenxin Huang, Jingling Yuan
ICIP4
2020 Complementing Representation Deficiency in Few-shot Image Classification: A Meta-Learning Approach
abstract
Few-shot learning is a challenging problem that has attracted more and more attention recently since abundant training samples are difficult to obtain in practical applications. Meta-learning has been proposed to address this issue, which focuses on quickly adapting a predictor as a base-learner to new tasks, given limited labeled samples. However, a critical challenge for meta-learning is the representation deficiency since it is hard to discover common information from a small number of training samples or even one, as is the representation of key features from such little information. As a result, a meta-learner cannot be trained well in a high-dimensional parameter space to generalize to new tasks. Existing methods mostly resort to extracting less expressive features so as to avoid the representation deficiency. Aiming at learning better representations, we propose a meta-learning approach with complemented representations network (MCRNet) for few-shot image classification. In particular, we embed a latent space, where latent codes are reconstructed with extra representation information to complement the representation deficiency. Furthermore, the latent space is established with variational inference, collaborating well with different base-learners, and can be extended to other models. Finally, our end-to-end framework achieves the state-of-the-art performance in image classification on three standard few-shot learning datasets.
Xian Zhong, Wenxin Huang, Lin Li 0001, Shuqin Chen, Chia-Wen Lin
ICPR3
2020 Visible-infrared Person Re-identification via Colorization-based Siamese Generative Adversarial Network
abstract
With explosive surveillance data during day and night, visible-infrared person re-identification (VI-ReID) is an emerging challenge due to the apparent cross-modality discrepancy between visible and infrared images. Existing VI-ReID work mainly focuses on learning a robust feature to represent a person in both modalities despite the modality gap cannot be effectively eliminated. Recent research works have proposed various generative adversarial network (GAN) models to transfer the visible modality to another unified modality, aiming to bridge the cross-modality gap. However, they neglect the information loss caused by transferring the domain of visible images which is significant for identification. To effectively address the problems, we observe that key information such as textures and semantics in an infrared image can help to color the image itself and the colored infrared image maintains rich information from infrared image while reducing the discrepancy with the visible image. We therefore propose a colorization-based Siamese generative adversarial network (CoSiGAN) for VI-ReID to bridge the cross-modality gap, by retaining the identity of the colored infrared image. Furthermore, we also propose a feature-level fusion model to supplement the transfer loss of colorization. The experiments conducted on two cross-modality person re-identification datasets demonstrate the superiority of the proposed method compared with the state-of-the-arts.
Xian Zhong, Tianyou Lu, Wenxin Huang, Jingling Yuan, Wenxuan Liu 0008, Chia-Wen Lin
ICMR3
2020 HMM-Based Person Re-identification in Large-Scale Open Scenario
Ruimin Hu, Wenxin Huang, Xiaochen Wang 0001, Dengshi Li
MMM (1)3
2020 Video Human Behavior Recognition Based on ISA Deep Network Model
abstract
Vision-based behavior recognition is the analysis and recognition of human behavior in video. It has been widely used in many aspects such as multimedia information retrieval, behavior monitoring, and robot perception. This paper uses the Independent Subspace Analysis (ISA) deep network model feature extraction method, which is based on the ISA model and neural network theory, and combines data preprocessing methods, [Formula: see text]-means clustering methods, and Support Vector Machine (SVM) classifiers to achieve video classification and identification of human behavior. The ISA-based deep network model feature extraction method is an unsupervised learning method that can obtain behavior characteristics with good invariance and characterization capabilities in video human behavior. The experiment was conducted on the basis of the Hollywood2 human behavior data set. This experiment was compared with other commonly used human behavior feature extraction and recognition methods. The experimental results validated the effectiveness and advantages of this method in the classification and recognition of human behavior.
Xian Zhong, Wenxin Huang, Ruiqi Luo
Int. J. Pattern Recognit. Artif. Intell.2
2019 Multi-Similarity Re-Ranking for Person Re-Identification
abstract
Re-ranking has been proved an effective method to boost the performance of person re-identification. Existing works focus on contextual or graph-based similarity to improve the initial ranking result. The former mainly concentrates on more accurate similarity description but neglects the manifold constraint. While, the later centers on solving similarities with manifold constraint, which acquires several accurate top ranks. In this paper, we propose a novel method which not only takes contextual similarity to generate top ranks accurately but also refines the ranks based on graph-based similarity. Specifically, given initial Euclidean distances between a probe and galleries, we mine contextual and graph-based similarities respectively and then re-rank all galleries with a diffusion procedure under constraints of both similarities. Experiments on two person re-ID datasets demonstrate that our method outperforms state-of-the-art re-ranking approaches in person re-identification.
Longxiang Jiang, Chao Liang 0001, Dongshu Xu, Wenxin Huang
ICIP4
2019 Squeeze-and-Excitation Wide Residual Networks in Image Classification
abstract
The depth and width of the network have been investigated to influence the performance of image classification during the resent research. Wide residual networks (WRNs) have proved that the performance of classification can be improved by the width of the networks. With consideration of the significance, expanding the width is to increase the number of channels. However, not all the channels are needed. Meanwhile, much channel information will be lost while exploiting the global average pooling at the end of WRNs for image representations because the mean value is only related to the first order information. With the two considerations stated above, we propose squeeze-and-excitation WRNs which are based on the global covariance pooling (SE-WRNs-GVP). A residual Squeeze-and-Excitation block (rSE-block) can make up for the lost information due to global average pooling in SE-block. Then, informative channels of WRNs will be utilized. Finally, the global covariance pooling at the end of WRNs characterizes the correlations of feature channels for more discriminative representations. A SE-block with dropout is proposed to avoid over-fitting. We conduct experiments on CIFAR10 and CIFAR100 datasets and achieve a better performance without increasing the model complexity.
Xian Zhong, Oubo Gong, Wenxin Huang, Lin Li 0001, Hongxia Xia
ICIP3
2019 Deep Multi-label Hashing for Image Retrieval
abstract
Due to its low storage cost and fast query speed, hashing has been widely applied to approximate nearest neighbor search for large-scale image retrieval, while deep hashing further improves the retrieval quality by learning a good image representation. However, existing deep hash methods simplify multi-label images into single-label processing, so the rich semantic information from multi-label is ignored. Meanwhile, the imbalance of similarity information leads to the wrong sample weight in the loss function, which makes unsatisfactory training performance and lower recall rate. In this paper, we propose Deep Multi-Label Hashing (DMLH) model that generates binary hash codes which retain the semantic relationship of multi-label of the image. The contributions of this new model mainly include the following two aspects: (1) A novel sample weight calculation model adaptively adjusts the weight of the sample pair by calculating the semantic similarity of the multi-label image pairs. (2) The sample weight cross-entropy loss function, which is designed according to the similarity of the image, adjusts the balance of similar image pairs and dissimilar image pairs. Extensive experiments demonstrate that the proposed method can generate hash codes which achieve better retrieval performance on two benchmark datasets, NUS-WIDE and MS-COCO.
Xian Zhong, Jiachen Li 0002, Wenxin Huang
ICTAI3
2019 Poses Guide Spatiotemporal Model for Vehicle Re-identification
Xian Zhong, Meng Feng, Wenxin Huang, Zheng Wang 0007, Shin'ichi Satoh 0001
MMM (2)3
2016 Camera Network Based Person Re-identification by Leveraging Spatial-Temporal Constraint and Multiple Cameras Relations
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Xian Zhong, Chunjie Zhang 0001
MMM (1)1
2016 Spatial Constrained Fine-Grained Color Name for Person Re-identification
Yang Yang 0062, Yuhong Yang 0001, Mang Ye, Wenxin Huang, Zheng Wang 0007, Chao Liang 0001, Chunjie Zhang 0001
MMM (1)4
2015 Multi-Level Fusion for Person Re-identification with Incomplete Marks
abstract
Most video surveillance suspect investigation systems rely on the videos taken in different camera views. Actually, besides the videos, in the investigation process, investigators also manually label some marks, which, albeit incomplete, can be quite accurate and helpful in identifying persons. This paper studies the problem of Person Re-identification with Incomplete Marks (PRIM), aiming at ranking the persons in the gallery according to both the videos and incomplete marks. This problem is solved by a multi-step fusion algorithm, which consists of three key steps: (i) The early fusing step exploits both visual features and marked attributes to predict a complete and precise attribute vector. (ii) Based on the statistical attribute d ominance and saliency phenomena, a dominance-saliency matching model is suggested for measuring the distance between attribute vectors. (iii) The gallery is ranked separately by using visual features and attribute vectors, and the overall ranking list is the result of a late fusion. Experiments conducted on VIPeR dataset have validated the effectiveness of the proposed method in all the three key steps. The results also show that through introducing marks, the retrieval accuracy is significantly improved.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Wenxin Huang
ACM Multimedia5