EDBT 2026 Demo / reviewers in the wild / expert
Xian Zhong
dblp:87/6023
· DBLP profile ↗
108ranked-venue papers
28as first author
85since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 68 · 22 first-author · 57 since 2021Artificial intelligence and machine learning · 34 · 7 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Computer networks · 6 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order TransferabstractAction recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition robustness. To address this, we introduce a multi-view formulation tailored for UAV altitudes and empirically uncover a distinctive partial order among views, where recognition accuracy consistently declines as altitude increases. This key observation motivates the proposed Aero Partial Order Guided Network (Aerorder), which explicitly models and exploits the hierarchical structure of UAV views to enhance cross-altitude action recognition. Aerorder comprises three main components: (1) a View Partition (VP) module that groups views by altitude using the head-to-body ratio; (2) an Order-aware Feature Decoupling (OFD) module that disentangles action-relevant and view-specific representations under partial order guidance; and (3) an Action Partial Order Guide (APOG) that progressively transfers knowledge from easier (low-altitude) to harder (high-altitude) views. Extensive experiments on Drone-Action, MOD20, and UAV validate the superiority of Aerorder, achieving consistent improvements over state-of-the-art methods, up to 4.7% and 1.3% gains on Drone-Action and MOD20, respectively. Wenxuan Liu 0008, Zhuo Zhou, Xuemei Jia, Siyuan Yang 0001, Wenxin Huang, Xian Zhong, Chia-Wen Lin |
AAAI | 6 |
| 2026 | TLC-Plan: A Two-Level Codebook Based Network for End-to-End Vector Floorplan GenerationabstractAbstract Automated floorplan generation aims to improve design quality, architectural efficiency, and sustainability by jointly modeling global spatial organization and precise geometric detail. However, existing approaches operate in raster space and rely on post hoc vectorization, which introduces structural inconsistencies and hinders end‐to‐end learning. Motivated by compositional spatial reasoning, we propose TLC‐Plan, a hierarchical generative model that directly synthesizes vector floorplans from input boundaries, aligning with human architectural workflows based on modular and reusable patterns. TLC‐Plan employs a two‐level VQ‐VAE to encode global layouts as semantically labeled room bounding boxes and to refine local geometries using polygon‐level codes. This hierarchy is unified in a CodeTree representation, while an autoregressive transformer samples codes conditioned on the boundary to generate diverse and topologically valid designs, without requiring explicit room topology or dimensional priors. Extensive experiments show state‐of‐the‐art performance on RPLAN dataset (FID = 1.84, MSE = 2.06) and leading results on LIFULL dataset. The proposed framework advances constraint‐aware and scalable vector floorplan generation for real‐world architectural applications. Source code and trained models are released at https://github.com/rosolose/TLC‐PLAN . Biao Xiong, Qiegen Liu, Xian Zhong |
Comput. Graph. Forum | 5 |
| 2026 | Ask and focus more: Question-prompt uncertainty allocation for dual-controllable video captioning
Shuqin Chen, Xingrui Yang 0003, Xiaohan Yu 0001, Xian Zhong |
Pattern Recognit. | 6 |
| 2026 | From temporal thumbnail to semantics: Debiasing multi-view action recognition
Zixian Zhu, Wenxuan Liu 0008, Xu Wang 0015, Bingyi Liu, Xiaohan Yu 0001, Xian Zhong |
Pattern Recognit. | 7 |
| 2026 | See what you seek: Semantic contextual integration for cloth-changing person re-identification
Wenxin Huang, Xian Zhong, Jingling Yuan, Alex Chichung Kot |
Pattern Recognit. | 2 |
| 2026 | AM40: Enhancing action recognition through matting-driven interaction analysis
Wenxuan Liu 0008, Kui Jiang, Siyuan Yang 0001, Chia-Wen Lin, Xian Zhong |
Pattern Recognit. | 7 |
| 2026 | ETV-Attack: Efficient text-driven visual-variable adversarial attacks on visual question answering with pre-trained language models
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang |
Pattern Recognit. | 3 |
| 2026 | Refined generation-based framework for consistent and reliable visual question answering
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang |
Pattern Recognit. | 3 |
| 2026 | Robust mixed-degradation person Re-identification via structural consistency distillation
Wenxin Huang, Wenxuan Liu 0008, Xuemei Jia, Xian Zhong |
Pattern Recognit. | 7 |
| 2026 | RegenTrack: Distance-Adaptive Regeneration Pool Matching for Drone-Based Crowd TrackingabstractDrone-based crowd tracking remains challenging due to low object distinctiveness, high crowd density, frequent occlusions, and the difficulty of maintaining continuous and precise localization from aerial views. To address these challenges, we proposeRegenTrack, a tracking framework that integrates distance-adaptive fusion with Regeneration Pool (RegenPool) matching. RegenTrack dynamically balances appearance and motion cues for trajectory-object association. Appearance cues are extracted from multi-frame fused trajectory and object features, while motion cues are derived by comparing predicted and observed positions via a motion network. Unmatched trajectories are temporarily stored in RegenPool and re-matched in subsequent frames to mitigate object loss caused by occlusion. Meanwhile, RegenPool evaluates unmatched objects through multi-frame observations to determine whether they should be promoted to new trajectories. In addition, a two-stage cascade matching strategy is employed to further enhance trajectory continuity and stability. Experiments on DRONECROWD, DRONEBIRD, and CROHD demonstrate that RegenTrack achieves notable improvements in both tracking accuracy and reliability for drone-based crowd scenarios. The code will be released at https://github.com/Zebrabeast/RegenTrack. Jingling Yuan, Huilin Zhu, Jinqiao Wang, Xian Zhong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | TCP: Text-Guided Cascade Network for Pedestrian Crossing Intention PredictionabstractPedestrian crossing intention prediction is crucial for ensuring safety in intelligent transportation systems, especially in autonomous driving scenarios. Most existing methods rely primarily on visual information; however, the quality of visual data deteriorates significantly at long distances due to limited resolution. Although multi-modal approaches can mitigate this issue by incorporating additional sensory data, they inevitably introduce extra computational overhead. To address these challenges, we propose a lightweight cascaded model for pedestrian crossing intention prediction based on text-trajectory alignment. The model employs a cascaded architecture that jointly performs coordinate and intention prediction, while leveraging a pre-trained large language model (LLM) to generate textual descriptions of videos, thereby enriching trajectory features. Furthermore, a center-aware classification module is integrated to enhance inter-class separability and intra-class compactness. Extensive experiments onJAADandPIEdemonstrate state-of-the-art performance: our method achieves 91% accuracy onPIEand 89% onJAAD, matching or surpassing recent multi-modal approaches with substantially fewer inputs. The source code will be released athttps://github.com/xyhhappy/TCP-prediction Wenxuan Liu 0008, Wenxin Huang, Ryan Wen Liu, Xian Zhong |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph CaptioningabstractVideo paragraph captioning (VPC) aims to generate coherent, detailed narratives that accurately reflect a video's content. However, existing methods typically depend on coarse-grained event correlations and neglect the nuanced spatio-temporal interactions critical for comprehensive understanding. Refined verbs and prepositions, encoding actions and spatial relations, are essential for clear, consistent descriptions. To address these issues, we propose the Fine-Grained Lexical-Centric Semantic Network (FLS-Net), which emphasizes verbs and prepositions linked to salient objects to improve spatio-temporal coherence across events. FLS-Net integrates a multi-lexical synergy mechanism, leveraging nouns obtained via multi-modal matching, and employs a Verb-Guided Event Consistency Module (VECM) alongside a Preposition-Driven Relation Representation Module (PRRM). A cyclic encoder-decoder architecture further enforces event consistency, significantly boosting VPC performance. Extensive experiments onActivityNet CaptionsandYouCook2demonstrate FLS-Net's superiority over state-of-the-art approaches. The source code is available athttps://github.com/yangxingrui/FLS. Shuqin Chen, Xian Zhong, Xingrui Yang 0003, Bin Sheng 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 2 |
| 2026 | Sparse Mixture of Mambas for Domain Generalized Atomic Electron Tomography AugmentationabstractAtomic electron tomography (AET) is essential for characterizing the atomic structure of functional materials. However, raw 3-D tomograms often exhibit severe artifacts caused by geometric constraints and low radiation doses. Although point-attention-based ensemble augmentation methods effectively remove artifacts in simulated datasets with varying structure factors, they struggle with the complexity of real tomograms that demand multidomain feature learning. Moreover, existing models degrade in multidomain scenarios and incur high parameter counts introduced by point-attention mechanisms, which increase hardware demands. To address these challenges, we propose a sparse mixture of Mambas (MoMambas), a novel 3-D augmentation method that enhances domain generalization. MoMambas decouple domain-specific parameters by integrating a sparse mixture-of-experts (MoE) framework with Mamba-based experts, resolve positional ambiguity in sparse input sequences through positional information enhancement, and boost MoE accuracy via a multihead routing algorithm. Our approach achieves a 22% accuracy improvement over state-of-the-art AET augmentation methods in multidomain learning, reduces the parameter count to just 2.9% of the original, and lowers computational cost by 6%. Codes and data are publicly available at https://github.com/yuy38457/MoMambas. Jingling Yuan, Xian Zhong, Qixuan Zhao, Liqiang Mai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2026 | Concise Object-word Visuals as Effective Cues for Visual Question AnsweringabstractIn Visual Question Answering (VQA) , both the image and its accompanying question serve as the primary sources of information for the model. Conventional approaches typically rely heavily on dense visual representations for reasoning and answer prediction. However, when the visual and textual modalities are imbalanced or semantically misaligned, such disparities hinder effective multimodal learning and inference. To address this issue, we propose a multimodal information adjustment method, the Visual Text Information Adjuster (ViTA) . ViTA investigates the impact of embedding textual cues within images on the VQA process and promotes cross-modal balance to improve accuracy. Specifically, since image content often dominates over question content, ViTA adjusts the balance by either masking visual information or augmenting it with object-word visual cues directly embedded in the image. Experimental results validate our hypothesis and further demonstrate that ViTA can serve as an effective data augmentation strategy, yielding measurable improvements across multiple VQA models. The code will be released at https://github.com/xqx23/ViTA . Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Rubing Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | QG-STR: Training-Time Optimized Question-Guided Scene Text Recognition via Visual Question AnsweringabstractScene Text Spotting (STS) aims to transcribe text embedded in natural images, typically encompassing Scene Text Detection (STD) and Scene Text Recognition (STR) . Advances in image understanding have made end-to-end text spotting increasingly viable. Concurrently, multimodal research has highlighted the potential of vision-language reasoning tasks, such as Visual Question Answering (VQA) . To leverage multimodal reasoning for STR, we propose a training-time question-guided STR framework that integrates VQA, termed Question-Guided STR (QG-STR) . The framework unifies STR, Visual Question Generation (VQG) , and VQA within a single architecture, enabling multimodal reasoning to enhance text-spotting performance. Specifically, visual understanding and logical reasoning are used as supervisory signals during training to improve text recognition accuracy and boost end-to-end text spotting. QG-STR is model-agnostic and compatible with diverse STR and VQA architectures, employing question guidance solely as a training-time supervision mechanism. During inference, the STR module functions independently without requiring external questions. Extensive experiments on Total-Text , ICDAR2015 , ICDAR2013 , and CTW1500 validate the effectiveness of QG-STR. Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Rubing Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | Occlusion-aware vehicle re-identification via sparse token modeling under multi-perspective cues
Yuanbang Li, Xianyao Ping, Rong Peng, Xian Zhong |
Vis. Comput. | 5 |
| 2026 | Vessel re-identification via background structure learning
Minxuan Yang, Zhenzhu Fu, Xian Zhong |
Vis. Comput. | 5 |
| 2025 | StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language ModelsabstractThe rapid development of multimodal large language models (MLLMs) has positioned visual storytelling as a crucial area in content creation. However, existing models often struggle to maintain temporal, spatial, and narrative coherence across image sequences, and they frequently lack the depth and engagement of human-authored stories. To address these challenges, we propose Story with Large Language-and-Vision Alignment (StoryLLaVA), a novel framework for enhancing visual storytelling. Our approach introduces a topic-driven narrative optimizer that improves both the training data and MLLM models by integrating image descriptions, topic generation, and GPT-4-based refinements. Furthermore, we employ a preference-based ranked story sampling method that aligns model outputs with human storytelling preferences through positive-negative pairing. These two phases of the framework differ in their training methods: the former uses supervised fine-tuning, while the latter incorporates reinforcement learning with positive and negative sample pairs. Experimental results demonstrate that StoryLLaVA outperforms current models in visual relevance, coherence, and fluency, with LLM-based evaluations confirming the generation of richer and more engaging narratives. The enhanced dataset and model will be made publicly available soon. Zhiding Xiao, Wenxin Huang, Xian Zhong |
COLING | 4 |
| 2025 | Anomize: Better Open Vocabulary Video Anomaly DetectionabstractOpen Vocabulary Video Anomaly Detection (OVVAD) seeks to detect and classify both base and novel anomalies. However, existing methods face two specific challenges related to novel anomalies. The first challenge is detection ambiguity, where the model struggles to assign accurate anomaly scores to unfamiliar anomalies. The second challenge is categorization confusion, where novel anomalies are often misclassified as visually similar base instances. To address these challenges, we explore supplementary information from multiple sources to mitigate detection ambiguity by leveraging multiple levels of visual data alongside matching textual information. Furthermore, we propose incorporating label relations to guide the encoding of new labels, thereby improving alignment between novel videos and their corresponding labels, which helps reduce categorization confusion. The resulting Anomize framework effectively tackles these issues, achieving superior performance on UCF-Crime and XD-Violence datasets, demonstrating its effectiveness in OVVAD. Wenxuan Liu 0008, Ruixu Zhang, Yuran Wang 0003, Xian Zhong, Zheng Wang 0007 |
CVPR | 6 |
| 2025 | STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have gained significant attention due to their biological plausibility and energy efficiency, making them promising alternatives to Artificial Neural Networks (ANNs). However, the performance gap between SNNs and ANNs remains a substantial challenge hindering the widespread adoption of SNNs. In this paper, we propose a Spatial-Temporal Attention Aggregator SNN (STAA-SNN) framework, which dynamically focuses on and captures both spatial and temporal dependencies. First, we introduce a spike-driven self-attention mechanism specifically designed for SNNs. Additionally, we pioneeringly incorporate position encoding to integrate latent temporal relationships into the incoming features. For spatial-temporal information aggregation, we employ step attention to selectively amplify relevant features to variant steps. Finally, we implement a time-step random dropout strategy to avoid local optima. The framework demonstrates exceptional performance across diverse datasets and exhibits strong generalization capabilities. Notably, STAA-SNN achieves state-of-the-art results on neuromorphic datasets CIFAR10-DVS of 82.10% and with performances of 97.14%, 82.05% and 70.40% on the static datasets CIFAR-10, CIFAR-100 and ImageNet, respectively. Furthermore, this model exhibits improved performance ranging from 0.33% to 2.80% with fewer time steps. Tianqing Zhang, Kairong Yu, Xian Zhong, Hongwei Wang 0001, Qi Xu 0008, Qiang Zhang 0008 |
CVPR | 3 |
| 2025 | LGNet: Linear Graph Representation for Efficient Cold-Start RecommendationsabstractGraph Convolutional Networks (GCNs) demonstrate significant potential in recommendation systems but face difficulties with the cold-start problem, especially in integrating new nodes during inference. The typical solution leverages meta-learning for few-shot learning, though it often fails to fully capture collaborative filtering between nodes. In this paper, we revisit the node embedding propagation algorithm in GCNs, emphasizing the importance of collaborative filtering and elucidating the relation between high-order and low-order embeddings. Given the substantial interaction data required for training recommendation models, maintaining a simple model structure remains crucial. To address these challenges, we propose the Linear Graph Network (LGNet), which theoretically compresses multi-layer GCNs into a single layer, enabling the embedding of new nodes during inference. Experimental results on benchmark datasets for link prediction and user cold-start tasks demonstrate that LGNet outperforms existing methods. The code will be available at https://github.com/kunbeibei/LGNet. Ruiqi Luo, Bangchao Wang, Xian Zhong |
ICASSP | 5 |
| 2025 | Synergistic Integration of Cross-Spatial Learning for Lightweight Crack DetectionabstractEfficient crack segmentation is crucial for engineering surface inspection, especially on edge devices where both accuracy and computational efficiency are essential. To address the challenges posed by crack directionality and blurred edges while enhancing performance, we propose a lightweight segmentation model, SCCU-Net, based on cross-spatial synergistic learning and coordinate awareness. The model integrates coordinate information and perceptual sets into a hybrid attention mechanism, significantly boosting segmentation accuracy. We introduce a Mapping Attention Gate (MAG), which utilizes gating signals from fine-grained features to guide cross-spatial learning, alongside an Adaptive Skip Fusion (ASF) to ensure smooth feature integration while optimizing computational resources. Extensive experiments on six benchmark datasets demonstrate that SCCU-Net consistently outperforms existing lightweight models, setting a new benchmark for crack detection on edge devices. Code is available at https://github.com/lsyyy20000830/SCCU-Net. Senyao Li, Jingling Yuan, Huilin Zhu, Xian Zhong |
ICASSP | 4 |
| 2025 | Bridging the One-to-Many Gap: Multi-label Semantic Learning and Relay for Video CaptioningabstractMany commonly used video captioning datasets contain multiple caption annotations per video. When training with cross-entropy loss, the model encounters ambiguity because the same input is mapped to different targets, leading to confusion. To address this issue, we propose the Multi-label Semantic Learning and Relay (MSLR) framework, which transforms the one-to-many relation between videos and their descriptions into a one-to-one mapping. Specifically, MSLR introduces two modules after the decoder. The Multi-label Multi-granularity Learning (MML) module integrates sentence-level granularity from multiple descriptions and employs attention mechanisms to extract word-level granularity weights, thereby capturing complementary semantic information from diverse perspectives. The Multi-label Semantic Relay (MSR) module subsequently leverages a parameter-sharing mechanism to feed complementary semantics back into the decoder, thus preventing the generation of overly generic descriptions during inference. Compared to state-of-the-art lightweight methods, MSLR achieves highly competitive results. The code is available at https://github.com/hyk0320/MSLR. Shuqin Chen, Yikang Hu, Zhixin Sun, Liangjun Yu, Xian Zhong |
ICME | 6 |
| 2025 | SOTA: Spike-Navigated Optimal TrAnsport Saliency Region Detection in Composite-bias VideosabstractExisting saliency detection methods struggle in real-world scenarios due to motion blur and occlusions. In contrast, spike cameras, with their high temporal resolution, significantly enhance visual saliency maps. However, the composite noise inherent to spike camera imaging introduces discontinuities in saliency detection. Low-quality samples further distort model predictions, leading to saliency bias. To address these challenges, we propose Spike-navigated Optimal TrAnsport Saliency Region Detection (SOTA), a framework that leverages the strengths of spike cameras while mitigating biases in both spatial and temporal dimensions. Our method introduces Spike-based Micro-debias (SM) to capture subtle frame-to-frame variations and preserve critical details, even under minimal scene or lighting changes. Additionally, Spike-based Global-debias (SG) refines predictions by reducing inconsistencies across diverse conditions. Extensive experiments on real and synthetic datasets demonstrate that SOTA outperforms existing methods by eliminating composite noise bias. Our code and dataset will be released at https://github.com/lwxfight/sota. Wenxuan Liu 0008, Xian Zhong, Zhaofei Yu, Tiejun Huang 0001 |
IJCAI | 4 |
| 2025 | A Cooperative Safety-Enhanced Control Framework for Driving Assistance in the Internet of VehiclesabstractFor the Internet of Vehicles (IoV), driving safety applications require reliable and up-to-date knowledge of the state of vehicles and traffic. A single vehicle cannot meet all the reliability requirements because of the limited capability of information acquisition. Thus, cooperation among vehicles for information sharing is essential. However, due to the high dynamic network topology and harsh channel conditions, maintaining long-term cooperation is not feasible. Only the messages that most affect the driving state can obtain the transmission opportunity for avoiding network congestion. In this paper, we propose a cooperative safety-enhanced control framework (SCF). This framework concentrates on the construction of dynamic and adaptive cooperation among vehicles and evaluates the key feature parameters to achieve an optimal safety utility for feedback control over the driving state. We construct a general multi-layer solution framework for driving assistance in SCF. First, we construct multiple temporary cooperative platoons to coordinate adjacent vehicles and realize a relatively uniform driving state. The cooperative platoon maintains short-term stability for vehicle sensing and tracing. Second, we propose a utility evaluation model for extracting the key feature parameters related to the driving state, which is the basis of the optimization for message transmission and driving control. Third, we design a two-level joint optimization mechanism for the deep fusion of the multi-source heterogeneous data to maximize the total utility of driving safety. Finally, we propose an adaptive feedback control model for the cooperative platoon, which actively adjusts the driving control strategy and the message transmission strategy in a real-time manner. Then the optimal driving assistant decision can be made. Extensive simulation results show that SCF outperforms related communication mechanisms for safe driving in the IoV, demonstrating that SCF can effectively enhance driving assistance control. Yan Zhang 0077, Chao Yang 0043, Zhifei Li 0009, Kui Xiao, Miao Zhang 0036, Wenxin Huang, Hao Chen 0134, Jianhua Song, Xian Zhong, Haobo Ma |
ICMR | 12 |
| 2025 | Step-wise Soft Alignment Enhanced Procedural Text Generation from Long Instructional VideosabstractWith the rise of generative models, video-language cross-modal applications have seen significant growth. Generating procedural text from instructional videos has become a crucial task, playing a key role in both understanding visual scenes and supporting practical applications. The sequential nature of video clips is particularly important, as entities may appear across multiple clips, reflecting fine-grained intra-modal self-similarity. However, most existing training methods treat other clips in a sequence as negative samples when a target is specified, neglecting their step-wise correlations. To address this limitation, we introduce Step-wise Soft Alignment via OpTimal TrAnsport (SATA), which constructs soft positive pairs to mitigate the issue. SATA first generates a step-wise similarity matrix by leveraging visual representations and generated procedural text. It then aligns the step-wise distributions between procedural text and video clips using optimal transport. The resulting transport distance serves as a weight, treating these pairs as soft positives for contrastive learning, ultimately improving the accuracy of procedural text generation. Our experiments on the publicly available YouCookII and ActivityNet Captions datasets demonstrate the effectiveness of SATA, achieving absolute improvements of 0.5% to 1.7% and 0.6% to 14.9% in paragraph-level evaluation, respectively. Lin Li 0001, Xian Zhong, Xiaohui Tao 0001, Jianquan Liu |
ICMR | 3 |
| 2025 | SegTraj: A Segmented-Trajectory-Aware Spatio-Temporal Graph Convolutional Network for Social Group DetectionabstractSocial group detection aims to identify groups of individuals exhibiting social behavior from multi-individual trajectory data. Recent approaches often determine group correlations based on global trajectory similarity, while temporal dynamics can cause diverging member trajectories and undermine similarity-based measures. Other methods model pairwise interaction strengths to capture group relations, focusing only on explicit direct interactions while ignoring implicit indirect interactions. To address temporal variability of group structures, we decompose long trajectories into multiple semantic sub-trajectories, enabling the capture of dynamic characteristics. Furthermore, to explore implicit indirect interactions, we introduce a unified spatio-temporal graph structure that models both direct and indirect interactions among individuals. In addition, considering the contextual influence of the neighborhood of an individual, we incorporate neighborhood information into the trajectory representation process. Based on these insights, we propose a Segmented-Trajectory-Aware Spatio-Temporal Graph Convolutional Network (SegTraj). This framework uniformly models explicit and implicit interactions through a spatio-temporal graph, and fuses individual trajectories with contextual neighborhood information for fine-grained representation of group relationships. Extensive experiments on three datasets covering both synthetic and real-world scenarios demonstrate that SegTraj significantly outperforms baseline methods. The code is available at https://github.com/DC0827/SegTraj. Xiongwei Dang, Wenxuan Liu 0008, Xian Zhong, Zheng Wang 0007 |
ACM Multimedia | 3 |
| 2025 | Beyond the Individual: Introducing Group Intention Forecasting with SHOT DatasetabstractIntention recognition has traditionally focused on individual intentions, overlooking the complexities of collective intentions in group settings. To address this limitation, we introduce the concept of group intention, which represents shared goals emerging through the actions of multiple individuals, and Group Intention Forecasting (GIF), a novel task that forecasts when group intentions will occur by analyzing individual actions and interactions before the collective goal becomes apparent. To investigate GIF in a specific scenario, we propose SHOT, the first large-scale dataset for GIF, consisting of 1,979 basketball video clips captured from 5 camera views and annotated with 6 types of individual attributes. SHOT is designed with 3 key characteristics: multi-individual information, multi-view adaptability, and multi-level intention, making it well-suited for studying emerging group intentions. Furthermore, we introduce GIFT (Group Intention ForecasTer), a framework that extracts fine-grained individual features and models evolving group dynamics to forecast intention emergence. Experimental results confirm the effectiveness of SHOT and GIFT, establishing a strong foundation for future research in group intention forecasting. The dataset is available at https://xinyi-hu.github.io/SHOT\_DATASET. Ruixu Zhang, Yuran Wang 0003, Chaoyu Mai, Wenxuan Liu 0008, Danni Xu, Xian Zhong, Zheng Wang 0007 |
ACM Multimedia | 7 |
| 2025 | YES: You should Examine Suspect cues for low-light object detection
Wenxin Huang, Xian Zhong |
Comput. Vis. Image Underst. | 6 |
| 2025 | SU-YOLO: Spiking neural network for efficient underwater object detection
Guoqiang Gong, Xiaobo Ding, Xian Zhong |
Neurocomputing | 5 |
| 2025 | Dynamic and static mutual fitting for action recognition
Wenxuan Liu 0008, Xuemei Jia, Xian Zhong, Kui Jiang, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 3 |
| 2025 | Uncertainty-Aware With Adaptive Geometric Correction for Multimodal Land-Cover ClassificationabstractLand cover classification (LCC) is a fundamental task in remote sensing and geographic information science. Multi-modal fusion has shown great potential for enhancing LCC performance, for example, by combining optical and synthetic aperture radar (SAR) imagery to leverage their complementary strengths. However, two key challenges hinder effective fusion:1) local geometric mismatches caused by distinct imaging geometries, and2) inconsistent reliability (the ability of a modality to deliver accurate and stable information) in LCC arising from different modalities and their acquisition conditions. To address these issues, we propose Uncertainty-Aware Fusion with Adaptive Geometric Correction (UAG), which comprises three main components. First, the Adaptive Geometric Correction Module (AGCM) applies learnable pixel shifts to establish bidirectional local correlations between multiscale optical and SAR features, thereby mitigating spatial inconsistencies. Second, the Adaptive Uncertainty-Aware Dynamic Fusion Module (ADFM) employs evidential deep learning to model uncertainty, defined as the extent of reliability deficiency, for each modality using the Dirichlet distribution and subjective logic, enabling confidence-aware feature weighting. Third, a lightweight multiscale decoder integrates hierarchical features through a hybrid MLP-convolutional architecture, improving both segmentation efficiency and accuracy. We evaluate UAG on WHU-OPT-SAR and DFC23 datasets, where experimental results demonstrate substantial improvements over state-of-the-art methods. The code will be released at https://github.com/cccwbin/UAGNet. Xu Wang 0015, Yi Xiao 0003, Wenxin Huang, Bihan Wen, Xian Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | MFAE-YOLO: Multifeature Attention-Enhanced Network for Remote Sensing Images Object DetectionabstractObject detection in aerial remote sensing images is essential for applications such as traffic management, public security, and ecological monitoring. However, existing methods struggle to handle multi-scale objects, complex backgrounds, and frequent occlusions, resulting in inadequate global feature extraction and unstable multi-scale loss calculations. To address these challenges, we propose the Multi-Feature Attention-Enhanced YOLO (MFAE-YOLO) network. Our approach introduces a Global Feature Fusion Processing (GFFP) module to enhance global feature extraction. The backbone network incorporates Fusion of Channel, Pixel, and Spatial (FCPS) attention modules to strengthen feature representation, while C2F-Feature Pool Extraction Units (C2F-FPEU) optimize pooling-based feature extraction. Additionally, we propose an Accurate IoU (AIoU) loss function to refine bounding box regression. Experiments on NWPU VHR-10, RSOD, and DIOR datasets demonstrate that MFAE-YOLO surpasses state-of-the-art methods, achieving mAP50values of 94.7%, 94.8%, and 67.0%, respectively. Ablation studies further validate the contributions of each module. The code is available at https://github.com/yiboCode/MFAE-YOLO. Xin Cheng 0020, Ning Xu 0006, Xu Wang 0015, Xian Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Local-Global Sparse Transformer for Road Extraction From Remote Sensing ImageryabstractAccurate road extraction from high-resolution remote sensing imagery is crucial for applications such as road network generation, urban planning, autonomous driving, and military operations. Although deep learning has significantly advanced segmentation performance, road extraction remains challenging because roads exhibit geometric and structural characteristics that differ markedly from general objects. We propose a local–global sparse transformer network (LGST-Net) that leverages sparse attention (SA) to capture both local details and global context. First, a coarse-to-fine feature-enhanced preprocessing (C2F-FEP) module extracts low-level features at a fine-grained scale during the model’s initial stage. Next, we design the LGST backbone, which comprises three novel components: local shift-mask SA (LSM-SA), global compressed sparse attention (GC-SA), and a spatial-channel gate. Extensive experiments onDeepGlobeandRoadTracerdatasets demonstrate that LGST-Net outperforms state-of-the-art methods. Our code will be publicly available athttps://github.com/JaymeWX/Road_LGST-Net Xu Wang 0015, Wenxin Huang, Xian Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Multi-Granularity Distribution Alignment for Cross-Domain Crowd CountingabstractUnsupervised domain adaptation enables the transfer of knowledge from a labeled source domain to an unlabeled target domain, and its application in crowd counting is gaining momentum. Current methods typically align distributions across domains to address inter-domain disparities at a global level. However, these methods often struggle with significant intra-domain gaps caused by domain-agnostic factors such as density, surveillance angles, and scale, leading to inaccurate alignment and unnecessary computational burdens, especially in large-scale training scenarios. To address these challenges, we propose the Multi-Granularity Optimal Transport (MGOT) distribution alignment framework, which aligns domain-agnostic factors across domains at different granularities. The motivation behind multi-granularity is to capture fine-grained domain-agnostic variations within domains. Our method proceeds in three phases: first, clustering coarse-grained features based on intra-domain similarity; second, aligning the granular clusters using an optimal transport framework and constructing a mapping from cluster centers to finer patch levels between domains; and third, re-weighting the aligned distribution for model refinement in domain adaptation. Extensive experiments across twelve cross-domain benchmarks show that our method outperforms existing state-of-the-art methods in adaptive crowd counting. The code will be available at https://github.com/HopooLinZ/MGOT. Xian Zhong, Lingyue Qiu, Huilin Zhu, Jingling Yuan, Shengfeng He, Zheng Wang 0007 |
IEEE Trans. Image Process. | 1 |
| 2025 | Motion-Consistent Representation Learning for UAV-Based Action RecognitionabstractAction recognition aims to identify action categories in trimmed videos captured by multimedia devices, which often suffer from jitter, especially in uncrewed aerial vehicle (UAV) applications. Existing methods typically ignore the effect of jitter on actor motion or rely on external stabilization tools trained on large-scale unstable video datasets that may not be tailored to specific tasks. To address this, we propose the Stabilization-enhanced Recognition Network (StaRNet), an end-to-end framework that integrates video stabilization and contrastive learning. Inspired by traditional stabilizers, StaRNet’s Motion-aware Stabilization Module (MSM) constructs positive and negative video pairs to model instability: positive pairs use optical flow to estimate frame motion and refine rigid motion via keyframe estimation for motion-aware stabilization, while negative pairs assess temporal consistency using motion cues to boost classification. Moreover, we introduce a Motion-aware Constraint (MC) that regulates dynamic stabilization to adapt to varying motion patterns and enrich action representations. Experiments on UAV benchmarks show that StaRNet outperforms state-of-the-art methods and substantially enhances video stabilization. The code is available athttps://github.com/lwxfight/-StaRNet Wenxuan Liu 0008, Xian Zhong, Yihan Dai, Xuemei Jia, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Fragrant: frequency-auxiliary guided relational attention network for low-light action recognition
Wenxuan Liu 0008, Xuemei Jia, Yihao Ju, Yakun Ju, Kui Jiang, Shifeng Wu, Luo Zhong, Xian Zhong |
Vis. Comput. | 8 |
| 2024 | Zero-Shot Object Counting with Good Exemplars
Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Yu Guo 0008, Zheng Wang 0007, Xian Zhong, Shengfeng He |
ECCV (5) | 6 |
| 2024 | Localization of Image Splicing Under Segment Anything Model With Integrated Compression and Edge ArtifactsabstractThe localization of image splicing involves identifying pixels in an image that have been spliced from other images, necessitating the discernment of splicing features. Despite significant advancements driven by the rise of social media and deep learning, existing methods exhibit limitations, often neglecting the integration of coarse and precise features and lacking the ability to understand objects. This leads to erroneous predictions in identifying spliced regions. This paper proposes Segment Anything Model with Integrated Compression and Edge artifacts (SAM-ICE) for the localization of image splicing, addressing these limitations by fusing forged edge features and compression artifact features. Leveraging SAM’s object understanding ability, our method identifies spliced regions using the fused features as guidance. Specifically, we employ Edge Artifact Extractor (EAE) to extract fine high-frequency edge splicing features and Compression Artifact Extractor (CAE) to extract coarse compression artifact features. By combining these features, our method utilizes coarse-fine features to accurately pinpoint the spliced portions of the image. Experimental results demonstrate the superior accuracy, robustness, and generalizability of our method compared to the state-of-the-arts. Ruhao Zhao, Xian Zhong, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ICIP | 2 |
| 2024 | DenseTrack: Drone-Based Crowd Tracking via Density-Aware Motion-Appearance SynergyabstractDrone-based crowd tracking faces difficulties in accurately identifying and monitoring objects from an aerial perspective, largely due to their small size and close proximity to each other, which complicates both localization and tracking. To address these challenges, we present the Density-aware Tracking (DenseTrack) framework. DenseTrack capitalizes on crowd counting to precisely determine object locations, blending visual and motion cues to improve the tracking of small-scale objects. It specifically addresses the problem of cross-frame motion to enhance tracking accuracy and dependability. DenseTrack employs crowd density estimates as anchors for exact object localization within video frames. These estimates are merged with motion and position information from the tracking network, with motion offsets serving as key tracking cues. Moreover, DenseTrack enhances the ability to distinguish small-scale objects using insights from the visual-language model, integrating appearance with motion cues. The framework utilizes the Hungarian algorithm to ensure the accurate matching of individuals across frames. Demonstrated on DroneCrowd dataset, our approach exhibits superior performance, confirming its effectiveness in scenarios captured by drones. Our code will be available at: https://github.com/Zebrabeast/DenseTrack. Huilin Zhu, Jingling Yuan, Guangli Xiang, Xian Zhong, Shengfeng He |
ACM Multimedia | 5 |
| 2024 | Towards Low-latency Event-based Visual Recognition with Hybrid Step-wise Distillation Spiking Neural NetworksabstractSpiking neural networks (SNNs) have garnered significant attention for their low power consumption and high biological interpretability. Their rich spatio-temporal information processing capability and event-driven nature make them ideally well-suited for neuromorphic datasets. However, current SNNs struggle to balance accuracy and latency in classifying these datasets. In this paper, we propose Hybrid Step-wise Distillation (HSD) method, tailored for neuromorphic datasets, to mitigate the notable decline in performance at lower time steps. Our work disentangles the dependency between the number of event frames and the time steps of SNNs, utilizing more event frames during the training stage to improve performance, while using fewer event frames during the inference stage to reduce latency. Nevertheless, the average output of SNNs across all time steps is susceptible to individual time step with abnormal outputs, particularly at extremely low time steps. To tackle this issue, we implement Step-wise Knowledge Distillation (SKD) module that considers variations in the output distribution of SNNs at each time step. Empirical evidence demonstrates that our method yields competitive performance in classification tasks on neuromorphic datasets, especially at lower time steps. Our code will be available at: https://github.com/hsw0929/HSD. Xian Zhong, Shengwang Hu, Wenxuan Liu 0008, Wenxin Huang, Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001 |
ACM Multimedia | 1 |
| 2024 | ICLR: Instance Credibility-Based Label Refinement for label noisy person re-identification
Xian Zhong, Xuemei Jia, Wenxin Huang, Wenxuan Liu 0008, Shuaipeng Su, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 1 |
| 2024 | Study of stability and object tracking of traffic video image for smart cities
Yongfeng Xing, Zhong Luo, Xian Zhong |
Pers. Ubiquitous Comput. | 3 |
| 2024 | Find Gold in Sand: Fine-Grained Similarity Mining for Domain-Adaptive Crowd CountingabstractThe domain shift of crowd scenes significantly hinders the application of crowd counting models in open scenarios. Although domain adaptation methods for crowd counting have bridged this gap to some extent, they ignore one of the significant causes of domain shift, which is the inter-domain data distribution bias. We discover that there exists a connection between the known and unknown distribution, which can be utilized by similarity mining to address the domain shift. However, there are still challenges related to insufficient and inaccurate similarity mining. In this article, a novel Fine-grained Inter-domain Similarity Mining (FSIM) framework is proposed. To comprehensively explore the similar distributions between source and target domains, we propose a Multi-scale Distribution Alignment (MDA) module based on diffusion retrieval. To enhance the reliability of inter- domain similarity mining, we propose a Multi-retrieval Refinement (MR) module based on evidence theory, which serves as an uncertainty measurement method. Eventually, to eliminate the data distribution bias, we perform model retraining using a similar distribution. Extensive experiments conducted on five standard crowd counting benchmarks, SHA, SHB, QNRF, NWPU, and JHU-CROWD++, show that the proposed FSIM has strong generalizability. Huilin Zhu, Jingling Yuan, Xian Zhong, Zheng Wang 0007 |
IEEE Trans. Multim. | 3 |
| 2024 | Action-aware Linguistic Skeleton Optimization Network for Non-autoregressive Video CaptioningabstractNon-autoregressive video captioning methods generate visual words in parallel but often overlook semantic correlations among them, especially regarding verbs, leading to lower caption quality. To address this, we integrate action information of highlighted objects to enhance semantic connections among visual words. Our proposed Action-aware Language Skeleton Optimization Network (ALSO-Net) tackles the challenge of extracting action information across frames, improving understanding of complex context-dependent video actions and reducing sentence inconsistencies. ALSO-Net incorporates a linguistic skeleton tag generator to refine semantic correlations and a video action predictor to enhance verb prediction accuracy in video captions. We also address issues of unsatisfactory caption length and quality by jointly optimizing different levels of motion prediction loss. Experimental evaluation on prominent video captioning datasets demonstrates that ALSO-Net outperforms baseline methods by a significant margin and achieves competitive performance compared to state-of-the-art autoregressive methods with smaller model complexity and faster inference time. Shuqin Chen, Xian Zhong, Lei Zhu 0003, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Refined Semantic Enhancement towards Frequency Diffusion for Video CaptioningabstractVideo captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at low-frequency tokens, which rarely occur but carry critical semantics, playing a vital role in the detailed generation. In this paper, we introduce a novel Refined Semantic enhancement method towards Frequency Diffusion (RSFD), a captioning model that constantly perceives the linguistic representation of the infrequent tokens. Concretely, a Frequency-Aware Diffusion (FAD) module is proposed to comprehend the semantics of low-frequency tokens to break through generation limitations. In this way, the caption is refined by promoting the absorption of tokens with insufficient occurrence. Based on FAD, we design a Divergent Semantic Supervisor (DSS) module to compensate for the information loss of high-frequency tokens brought by the diffusion process, where the semantics of low-frequency tokens is further emphasized to alleviate the long-tailed problem. Extensive experiments indicate that RSFD outperforms the state-of-the-art methods on two benchmark datasets, i.e., MSR-VTT and MSVD, demonstrate that the enhancement of low-frequency tokens semantics can obtain a competitive generation effect. Code is available at https://github.com/lzp870/RSFD. Xian Zhong, Shuqin Chen, Kui Jiang, Chen Chen 0001, Mang Ye |
AAAI | 1 |
| 2023 | Good is Bad: Causality Inspired Cloth-debiasing for Cloth-changing Person Re-identificationabstractEntangled representation of clothing and identity (ID)-intrinsic clues are potentially concomitant in conventional person Re- IDentification (ReID). Nevertheless, eliminating the negative impact of clothing on ID remains challenging due to the lack of theory and the difficulty of isolating the exact implications. In this paper, a causality-based Auto-Intervention Model, referred to as AIM11Codes will publicly available at https://github.com/BoomShakaY/AIM-CCReID., is first proposed to mitigate clothing bias for robust cloth-changing person ReID (CC-ReID). Specifically, we analyze the effect of clothing on the model inference and adopt a dual-branch model to simulate causal intervention. Progressively, clothing bias is eliminated automatically with model training. AIM is encouraged to learn more discriminative ID clues that are free from clothing bias. Extensive experiments on two standard CC-ReID datasets demonstrate the superiority of the proposed AIM over other state-of-the-art methods. Zhengwei Yang 0001, Xian Zhong, Zheng Wang 0007 |
CVPR | 3 |
| 2023 | Background Disturbance Mitigation for Video Captioning Via Entity-Action RelocationabstractVideo captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success. However, they focus on exploiting foreground semantics, ignoring the potential negative impact of video background disturbance to caption generation, i.e., the entities and the actions are misjudged by a similar video background. To ameliorate this issue, we propose Entity-Action Relocation (EAR) to enhance the adaptability of entities and actions to various backgrounds by giving them the background. Specifically, for an extracted original video feature, we construct a mixed background for all entities and actions to form a distracting video feature sample. After that, contrastive learning is applied to pull the generated caption of the original representations and of the distracting representations closer, and to push the former away from the generated caption of other videos, explicitly concentrating on the entities and actions of the current video scene. Extensive experiments on two public datasets (MSR-VTT and MSVD) demonstrate that dealing with background disturbance for video can obtain a competitive caption generation effect. Xian Zhong, Shuqin Chen, Wenxin Huang, Lin Li 0001 |
ICASSP | 2 |
| 2023 | Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic SegmentationabstractWhile enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to handle the more realistic multi-target domain adaptive semantic segmentation (MT-DASS) tasks. To solve this problem, we propose a Bi-Alignment framework based on Transformation (BAT). Specifically, we employ the Fourier style transform to convert the style of the source domain to that of the target domain without training any style transfer networks. In this way, we transform the single-peak distributed source domain into a multi-peak distribution that resembles the multi-target domain. Then, we perform fine-grained global and local dual distribution alignment between the same style of source-target domain pairs to achieve a multi-to-multi distribution alignment. Finally, self-training is utilized to further improve the network’s discriminability. Experimental results show that our approach achieves competitive results over state-of-the-art methods. Xian Zhong, Jing Xiao 0004, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 1 |
| 2023 | Neighborhood Information-Based Label Refinement for Person Re-Identification with Label NoiseabstractThe existing excellent person re-identification (Re-ID) model is still affected by the samples with the incorrect labels. It is difficult to accurately annotate person images in the real scene, resulting in label noise. To avoid fitting to the noisy labels, a common solution in Re-ID is to replace the original label with the label predicted by the deep model. Unfortunately, similar samples of different identities with the same label are due to label noise, which is challenging for the model to distinguish them. Neighborhood information can optimize noisy labels through neighborhood labels and similarity between samples. This paper proposes a label refinement module based on neighborhood information (LRNI) for person Re-ID with label noise. Specifically, we first use the pre-trained model to extract features and calculate the similarity between samples. Rather than treating samples as isolated, the similarity used as label propagation weight and neighborhood labels are combined to optimize noisy labels. To further reduce the influence of label noise, we design a hard sample re-weighting (HSR) strategy to balance the learning of noisy and boundary samples. Experimental results under different noise settings demonstrate our method's effectiveness in the person Re-ID task. Xian Zhong, Shuaipeng Su, Wenxuan Liu 0008, Xuemei Jia, Wenxin Huang, Mengdie Wang |
ICASSP | 1 |
| 2023 | Background-Weakening Consistency Regularization for Semi-Supervised Video Action DetectionabstractConsistency-based techniques have produced state-of-the-art results in semi-supervised action detection. When the model false detects the dynamic information in the background as an action, spatio-temporal consistency calculations can hardly reflect this false detection result. We consider weakening the dynamic information in the augmented video background to reduce its spatio-temporal consistency with the dynamic information in the original video background. Thus we propose a Background-Weakening with Calibration Constraint (BWCC) framework, which highlights the negative impact of information in the background of false detection by calculating the consistency of the predictions of the background weakened video and the original video. Specifically, Background Weaken (BW) module judges the foreground and background of the video based on the initial predictions of the model and makes adjustments to the video background. To mitigate the effects of the misjudgments result in weakened action pixels, we additionally introduce a model that does not undergo background weakening to aid training through Calibration Constraint (CC) module. Our approach achieves competitive performance over existing leading approaches on two action detection datasets, UCF101-24 and JHMDB-21. Xian Zhong, Aoyu Yi, Wenxuan Liu 0008, Wenxin Huang, Chengming Zou, Zheng Wang 0007 |
ICASSP | 1 |
| 2023 | Implicit Attention-Based Cross-Modal Collaborative Learning for Action RecognitionabstractHuman action recognition is an active research topic in recent years. Multiple modalities often convey heterogeneous but potentially complementary action information that single modality does not hold. Some efforts have been resoted to explore cross-modal representation to promote the modeling capability, but with limited improvement due to the simple fusion of different modalities. To this end, we propose an impliCit attention-based Cross-modal Collaborative Learning (C3L) for action recognition. Specifically, we apply a Modality Generalization network with Grayscale enhancement (MGG) to learn specific modality representation and interaction (infrared and RGB). Then, we construct a unified representation space through the Uniform Modality Representation module (UMR), which preserves the modality information while enhancing the overall representation ability. Finally, feature extractors adaptively leverage modality-specific knowledge to realize cross-modal collaborative learning. Extensive experiments conducted on three widely-used public benchmarks InfAR, HMDB51, and UCF101, demonstrate the effectiveness and strength of our proposed method. Jianghao Zhang, Xian Zhong, Wenxuan Liu 0008, Kui Jiang, Zhengwei Yang 0001, Zheng Wang 0007 |
ICIP | 2 |
| 2023 | DAWN: Direction-aware Attention Wavelet Network for Image DerainingabstractSingle image deraining aims to remove rain perturbation while restoring the clean background scene from a rain image. However, existing methods tend to produce blurry and over-smooth outputs, lacking some textural details. Wavelet transform can depict the contextual and textural information of an image at different levels, showing impressive capability of learning structural information in the images to avoid artifacts, and thus has been recently explored to consider the inherent overlap of background and rain perturbation in both the pixel domain and the frequency embedding space. However, the existing wavelet-based methods ignore the heterogeneous degradation for different coefficients due to the inherent directional characteristics of rain streaks, leading to inter-frequency conflicts and compromised deraining results. To address this issue, we propose a novel Direction-aware Attention Wavelet Network (DAWN) for rain streaks removal. DAWN has several key distinctions from existing wavelet transform-based methods: 1) introducing the vector decomposition to parameterize the learning procedure, where the rain streaks are derived into the vertical (V) and horizontal (H) components to learn the specific representation; 2) a novel direction-aware attention module (DAM) to fit the projection and transformation parameters to characterize the direction-specific rain components, which helps accurate texture restoration; 3) exploring practical composite constraints on the structure, details, and chrominance aspects for high-quality background restoration. Our proposed DAWN delivers significant performance gains on nine datasets across image deraining and object detection tasks, exceeding the state-of-the-art method MPRNet by 0.88 dB in PSNR on the Test1200 dataset with only 35.5% computation cost. Kui Jiang, Wenxuan Liu 0008, Zheng Wang 0007, Xian Zhong, Junjun Jiang, Chia-Wen Lin |
ACM Multimedia | 4 |
| 2023 | DAOT: Domain-Agnostically Aligned Optimal Transport for Domain-Adaptive Crowd CountingabstractDomain adaptation is commonly employed in crowd counting to bridge the domain gaps between different datasets. However, existing domain adaptation methods tend to focus on inter-dataset differences while overlooking the intra-differences within the same dataset, leading to additional learning ambiguities. These domain-agnostic factors,e.g., density, surveillance perspective, and scale, can cause significant in-domain variations, and the misalignment of these factors across domains can lead to a drop in performance in cross-domain crowd counting. To address this issue, we propose a Domain-agnostically Aligned Optimal Transport (DAOT) strategy that aligns domain-agnostic factors between domains. The DAOT consists of three steps. First, individual-level differences in domain-agnostic factors are measured using structural similarity (SSIM). Second, the optimal transfer (OT) strategy is employed to smooth out these differences and find the optimal domain-to-domain misalignment, with outlier individuals removed via a virtual "dustbin'' column. Third, knowledge is transferred based on the aligned domain-agnostic factors, and the model is retrained for domain adaptation to bridge the gap across domains. We conduct extensive experiments on five standard crowd-counting benchmarks and demonstrate that the proposed method has strong generalizability across diverse datasets. Our code will be available at: https://github.com/HopooLinZ/DAOT/. Huilin Zhu, Jingling Yuan, Xian Zhong, Zhengwei Yang 0001, Zheng Wang 0007, Shengfeng He |
ACM Multimedia | 3 |
| 2023 | Generating live commentary for marine traffic scenarios based on multi-model learning
Rui Zhang 0066, Yifan Zhuo, Kezhong Liu, Xian Zhong, Shaohua Wan 0001 |
Comput. Commun. | 5 |
| 2023 | SCPNet: Self-constrained parallelism network for keypoint-based lightweight object detection
Xian Zhong, Mengdie Wang, Wenxuan Liu 0008, Jingling Yuan, Wenxin Huang |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Dual-Recommendation Disentanglement Network for View Fuzz in Action RecognitionabstractMulti-view action recognition aims to identify action categories from given clues. Existing studies ignore the negative influences of fuzzy views between view and action in disentangling, commonly arising the mistaken recognition results. To this end, we regard the observed image as the composition of the view and action components, and give full play to the advantages of multiple views via the adaptive cooperative representation among these two components, forming a Dual-Recommendation Disentanglement Network (DRDN) for multi-view action recognition. Specifically, 1) For the action, we leverage a multi-level Specific Information Recommendation (SIR) to enhance the interaction among intricate activities and views. SIR offers a more comprehensive representation of activities, measuring the trade-off between global and local information. 2) For the view, we utilize a Pyramid Dynamic Recommendation (PDR) to learn a complete and detailed global representation by transferring features from different views. It is explicitly restricted to resist the fuzzy noise influence, focusing on positive knowledge from other views. Our DRDN aims for complete action and view representation, where PDR directly guides action to disentangle with view features and SIR considers mutual exclusivity of view and action clues. Extensive experiments have indicated that the multi-view action recognition method DRDN we proposed achieves state-of-the-art performance over powerful competitors on several standard benchmarks. The code will be available at https://github.com/51cloud/DRDN. Wenxuan Liu 0008, Xian Zhong, Zhuo Zhou, Kui Jiang, Zheng Wang 0007, Chia-Wen Lin |
IEEE Trans. Image Process. | 2 |
| 2023 | Win-Win by Competition: Auxiliary-Free Cloth-Changing Person Re-IdentificationabstractRecent person Re-IDentification (ReID) systems have been challenged by changes in personnel clothing, leading to the study of Cloth-Changing person ReID (CC-ReID). Commonly used techniques involve incorporating auxiliary information (e.g., body masks, gait, skeleton, and keypoints) to accurately identify the target pedestrian. However, the effectiveness of these methods heavily relies on the quality of auxiliary information and comes at the cost of additional computational resources, ultimately increasing system complexity. This paper focuses on achieving CC-ReID by effectively leveraging the information concealed within the image. To this end, we introduce an Auxiliary-free Competitive IDentification (ACID) model. It achieves a win-win situation by enriching the identity (ID)-preserving information conveyed by the appearance and structure features while maintaining holistic efficiency. In detail, we build a hierarchical competitive strategy that progressively accumulates meticulous ID cues with discriminating feature extraction at the global, channel, and pixel levels during model inference. After mining the hierarchical discriminative clues for appearance and structure features, these enhanced ID-relevant features are crosswise integrated to reconstruct images for reducing intra-class variations. Finally, by combing with self- and cross-ID penalties, the ACID is trained under a generative adversarial learning framework to effectively minimize the distribution discrepancy between the generated data and real-world data. Experimental results on four public cloth-changing datasets (i.e., PRCC-ReID, VC-Cloth, LTCC-ReID, and Celeb-ReID) demonstrate the proposed ACID can achieve superior performance over state-of-the-art methods. The code is available soon at: https://github.com/BoomShakaY/Win-CCReID. Zhengwei Yang 0001, Xian Zhong, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Visual Exposes You: Pedestrian Trajectory Prediction Meets Visual IntentionabstractPedestrian trajectory prediction in multiple scenarios is of immense importance in autonomous driving and disentanglement of human behavior but is limited in catching human intention and initiative. Most previous works tend to predict the trajectory using only 2D coordinates, which generally cause two common problems: a) Overlooking the subjective initiative, including sudden swerve and erratic movement; b) A potential challenge called abnormal collision caused by unlabeled pedestrians on dataset is not being identified and resolved, which would ruin the model prediction. To break those limitations, we introduce visual localization and orientation as Visual Intention Knowledge to help the trajectory prediction, which is learned directly from visual scenarios. It benefits to comprehend human intention and formulates decision-making processes. Moreover, by learning from the visual information and decision-making policy, we construct the Visual Intention Knowledge associated spatio-temporal Transformer (VIKT) to predict human trajectory by combining the intention knowledge with the novel Transformer. Extensive experimental results demonstrate that our VIKT model could achieve competitive performance by the Visual Intention Knowledge through optimizing the model prediction compared with state-of-the-art methods in terms of prediction accuracy on ETH/UCY and SDD benchmarks. Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Kui Jiang, Ryan Wen Liu, Zheng Wang 0007 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Graph Complemented Latent Representation for Few-Shot Image ClassificationabstractFew-shot learning is a tough topic to solve since obtaining a large number of training samples in real applications is challenging. It has attracted increasing attention recently. Meta-learning is a prominent way to address this issue, intending to adapt predictors as base-learners to new tasks swiftly. However, a key challenge of meta-learning is its lack of expressive capacity, which stems from the difficulty of extracting general information from a small number of training samples. As a result, the generalizability of meta-learners trained from high-dimensional parameter spaces is frequently limited. To learn a better representation, we propose a graph complemented latent representation (GCLR) network for few-shot image classification. In particular, we embed the representation into a latent space, in which the latent codes are reconstructed using variational information to enrich the representation. In this way, the latent representation can achieve better generalizability. Another benefit is that, because the latent space is formed using variational inference, it cooperates well with various base-learners, boosting robustness. To make full use of the relation between samples in each category, a graph neural network (GNN) is also incorporated to improve relation mining. Consequently, our end-to-end framework delivers competitive performance on three few-shot learning benchmarks for image classification. Xian Zhong, Mang Ye, Wenxin Huang, Chia-Wen Lin |
IEEE Trans. Multim. | 1 |
| 2023 | Beyond the Parts: Learning Coarse-to-Fine Adaptive Alignment Representation for Person SearchabstractPerson search is a time-consuming computer vision task that entails locating and recognizing query people in scenic pictures. Body components are commonly mismatched during matching due to position variation, occlusions, and partially absent body parts, resulting in unsatisfactory person search results. Existing approaches for extracting local characteristics of the human body using keypoint information are unable to handle the search job when distinct body parts are misaligned, ignoring to exploit multiple granularities, which is crucial in the person search process. Moreover, the alignment learning methods learn body part features with fixed and equal weights, ignoring the beneficial contextual information, e.g., the umbrella carried by the pedestrian, which supplements compelling clues for identifying the person. In this paper, we propose a Coarse-to-Fine Adaptive Alignment Representation (CFA 2 R) network for learning multiple granular features in misaligned person search in the coarse-to-fine perspective. To exploit more beneficial body parts and related context of the cropped pedestrians, we design a Part-Attentional Progressive Module (PAPM) to guide the network to focus on informative body parts and positive accessorial regions. Besides, we propose a Re-weighting Alignment Module (RAM) shedding light on more contributive parts instead of treating them equally. Specifically, adaptive re-weighted but not fixed part features are reconstructed by Re-weighting Reconstruction module, considering that different parts serve unequally during image matching. Extensive experiments conducted on CUHK-SYSU and PRW datasets demonstrate competitive performance of our proposed method. Wenxin Huang, Xuemei Jia, Xian Zhong, Xiao Wang 0029, Kui Jiang, Zheng Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | VCD: View-Constraint Disentanglement for Action RecognitionabstractAction recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is tantamount to track the view-invariant representation as it evolves with the performed action. However, the views’ and actions’ performance always complement each other, once simply looking for the view-invariant representation may cause some behavior information to be lost. In this paper, we propose the View-Constraint Disentanglement (VCD) framework for cross-view action recognition. Specifically, Constraint Disentanglement Module (CDM) is utilized to learn an action-invariant representation by discretizing view-specific representation and its normal distribution, which resolves the entangled relationship between view and action. Moreover, a novel Adaptive Distribution Module (ADM) is intended to befit enhance the high-correlation viewpoint variation information and refine the suitable weight. Extensive experiments are conducted on public benchmarks, indicating that our approach achieves better performance than other state-of-the-art approaches. Xian Zhong, Zhuo Zhou, Wenxuan Liu 0008, Kui Jiang, Xuemei Jia, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 1 |
| 2022 | Attentive Decoupling Network for Cloth-Changing Re-IdentificationabstractRecently, Cloth-Changing person Re-IDentification (CC-ReID) plays a vital role in the public security system and social livelihood, and suffers the problem of considerable intra-class variation. This paper demonstrates that coarse-grained appearance and body shape features are helpful for CC-ReID. We propose an Attentive DeCoupling (ADC) Network for CC-ReID without auxiliary information. The proposed network is built on two core designs. First, a joint identification structure is proposed to retain ID-relevant information at appearance and shape levels. Second, Competitive Attention (CA) is adopted, where the model progressively updates attention to accumulate sound cues for discriminating identity (ID). The proposed decoupling process is continuously improved through constant self-defeating competition of the network. Experimental results on the public cloth-changing dataset show the proposed method's effectiveness and generalizability. Zhengwei Yang 0001, Xian Zhong, Hong Liu 0009, Zhun Zhong, Zheng Wang 0007 |
ICME | 2 |
| 2022 | Graph-Based Structural Attributes for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID), which aims to identify the same vehicle across different surveillance cameras, is a significant application in urban operation and security. Although the existing methods have noticed the importance of local features, near-duplicated cases are still hard to be handled. The reason lies that the attribute features and personalized structure of vehicles are often ignored. In this paper, we propose a graph-based structural attribute network (GSAN), which contains an attribute feature extraction module (AFEM) and a dual-grained structural relation module (DSRM). The AFEM aims to obtain attribute features of vehicles with structural information between attributes, and the DSRM aims to make the attribute features able to represent structural relation information between parts and attributes. The result on representative datasets shows that GSAN achieves competitive improvements over the state-of-the-art methods. We also collect a dataset of vehicle images with attribute annotations. Our dataset and code are released at https://github.com/HappyBoBo0331/GSAN. Rongbo Zhang, Xian Zhong, Xiao Wang 0029, Wenxin Huang, Wenxuan Liu 0008 |
ICME | 2 |
| 2022 | Dual-Scale Alignment-Based Transformer on Linguistic Skeleton Tags for Non-Autoregressive Video CaptioningabstractDue to the characteristic of one-time parallel generation of a caption, non-autoregressive video captioning lacks strong dependencies between words. Although using guideline of scene-related visual words can promote caption generation, the semantic relations among visual words are barely explored, limiting the accurate representation. To this end, we propose a Dual-Scale Alignment-based transformer on Linguistic Skeleton Tags (DSA-LST), which alleviates the defect above in the form of visual words group (several words representing a video frame). Different groups represent different semantic dependencies by attention. We utilize linguistic skeleton tags (i.e., several groups) as sentence-level supervision for visual words sequence. For visual words group to accurately express a specific frame, we further design dual scales of visual-language bi-direction alignment to achieve internal relevance of the tags. Extensive experiments conducted on widely used datasets: MSVD and MSR-VTT demonstrate the effectiveness of our method when compared with existing approaches. Xian Zhong, Shuqin Chen, Zhixin Sun, Huantao Zheng, Kui Jiang |
ICME | 1 |
| 2022 | A Novel Spatial-Spectral Random Forest Algorithm for Pine WILT MonitoringabstractPine wilt disease is one of the most dangerous forest diseases. Because of its strong infectivity and harm, it is very important to find out and stop it in time. In this paper, a novel spatial-spectral random forest (SRF) algorithm for pine wilt monitoring is proposed, for solving the problem of small manual detection range, long investigation time, and untimely discovery of the diseased tree. The proposed method organically combines spatial features with spectral information to quickly and efficiently mark the location of diseased trees. In this way, the online monitoring of the target area using the data of the Beijing-2 satellite is realized. This paper analyses the location of diseased trees and provides early warnings for disease-prone trees. The accuracy of the proposed algorithm is 86.66%, by the confusion matrix analysis. Yali Zhang 0001, Wei Feng 0004, Yinghui Quan, Xian Zhong, Yijia Song, Qiang Li 0029, Gabriel Dauphin, Yong Wang 0011, Mengdao Xing |
IGARSS | 4 |
| 2022 | Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene UnderstandingabstractScene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and methods: 1) Manually synthetic rainy samples with empirically settings and human subjective assumptions; 2) Limited rainy conditions, including the rain patterns, intensity, and degradation factors; 3) Separated training manners for image deraining and semantic segmentation. To break these limitations, we pioneer a real, comprehensive, and well-annotated scene understanding dataset under rainy weather, named Rainy WCity. It covers various rain patterns and their bring-in negative visual effects, covering wiper, droplet, reflection, refraction, shadow, windshield-blurring, etc. In addition, to alleviate dependence on paired training samples, we design an unsupervised contrastive learning network for real image deraining and the final rainy scene semantic segmentation via multi-task joint optimization. A comprehensive comparison analysis is also provided, which shows that scene understanding in rainy weather is a largely open problem. Finally, we summarize our general observations, identify open research challenges, and point out future directions. Xian Zhong, Shidong Tu, Xianzheng Ma, Kui Jiang, Wenxin Huang, Zheng Wang 0007 |
IJCAI | 1 |
| 2022 | Fine-Grained Fragment Diffusion for Cross Domain Crowd CountingabstractDeep learning improves the performance of crowd counting, but model migration remains a tricky challenge. Due to the reliance on training data and inherent domain shift, model application to unseen scenarios is tough. To facilitate the problem, this paper proposes a cross-domain Fine-Grained Fragment Diffusion model (FGFD) that explores feature-level fine-grained similarities of crowd distributions between different fragments to bridge the cross-domain gap (content-level coarse-grained dissimilarities). Specifically, we obtain features of fragments in both source and target domains, and then perform the alignment of the crowd distribution across different domains. With the assistance of the diffusion of crowd distribution, it is able to label unseen domain fragments and make source domain close to target domain, which is fed back to the model to reduce the domain discrepancy. By monitoring the distribution alignment, the distribution perception model is updated, then the performance of distribution alignment is improved. During the model inference, the gap between different domains is gradually alleviated. Multiple sets of migration experiments show that the proposed method achieves competitive results with other state-of-the-art domain-transfer methods. Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Xian Zhong, Zheng Wang 0007 |
ACM Multimedia | 4 |
| 2022 | Joint Re-Detection and Re-Identification for Multi-Object Tracking
Xian Zhong, Jingling Yuan, Luo Zhong |
MMM (1) | 2 |
| 2022 | Patching Your Clothes: Semantic-Aware Learning for Cloth-Changed Person Re-Identification
Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang |
MMM (2) | 2 |
| 2022 | Global Temporal Attention Optimization for Human Trajectory PredictionabstractPredicting human trajectory is one of the key knowledge required for autonomous driving and social robots in real scenarios. Recent studies based on Transformer networks have shown a great ability to model social behaviors. As far as we know, global trajectory information has an essential influence on prediction at a certain step. However, these methods only rely on the previous trajectory states/attention but ignore the important following states/attention of the trajectory for each pedestrian, which will generally collapse on some irregular movements (e.g. acceleration, deceleration, and motionless). To solve this issue, we propose a Global Temporal Attention optimization model (GTAO), which activates the utilization of the following states/attention of the trajectory, and jointly and iteratively optimizes the preliminary trajectory prediction through a global temporal attention (GTA) module. To effectively address the decline in the generalizability and abnormal processing of the model, we further introduce global temporal guidance (GTG) module to instruct the GTA to learn the features closer to realistic trajectories. Experimental results on commonly used real-world human trajectory prediction datasets (ETH and UCY) indicate that our GTAO can achieve better performance in terms of prediction accuracy. Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Zheng Wang 0007 |
SMC | 2 |
| 2022 | Conceptual semantic enhanced representation learning for event recognition in still imagesabstractImage event recognition is different from object recognition, behaviour recognition and scene recognition. Event is a more advanced concept than object, behaviour and scene. Regarding semantics loss in image event recognition, this paper first proposes a WordNet-based optimization algorithm for concept semantics similarity and describes the semantics relations between different concepts by taking account of such following four impact factors in the WordNet tree as concept semantics distance, concept node depth, concept node density and concept semantics overlap ratio. On that basis, an image event recognition algorithm (CS-IER) based on concept score is proposed, while multi-view learning is applied to fuse concept score and inter-conceptual semantics relations. However, if a higher erroneous concept score is given using CNN, multi-view learning will also augment the concept score approximate to its erroneous concept semantics, thereby leading to the distortion of image representation information. To address this problem, CNN is used to extract channel information to obtain the local features of the image, and it is further fused with the optimized concept score features, so as to form the final image representation information and complete the image event recognition. In experiments, the effectiveness of the proposed algorithm on three datasets is verified. Ruiqi Luo, Bangchao Wang, Zaihui Deng, Xian Zhong |
Connect. Sci. | 5 |
| 2022 | Actor-Aware Alignment Network for Action RecognitionabstractAction recognition has attracted growing interest recently. It suffers from the problem that complex and diverse environments may disturb the extraction of action features. Existing methods propose to explore the temporal associations to alleviate the issue. However, they cannot handle long-range frames, and the rigid techniques are powerless against the differences caused by the deformation of the actors. To this end, we propose the Actor-Aware Alignment Network (A$^{3}$Net), which helps locate the action region. Specifically, through the intra-snippet correction, we afford the local segment alignment frames. The inter-snippet is designed to rectify the results, avoiding the occlusion situation that may appear in the local snippet. In addition, we consider intra-alignment short-range adjustive frames and long-range context frames between different snippets, which allows our A$^{3}$Net network to achieve the effect of focusing on long-range frame information. Multiple Reasoning Attention (MRA) modules are introduced to integrate features along the temporal dimension to keep the video spatio-temporal consistent. Extensive experiments conducted on three widely-used public benchmarks,UCF101,HMDB51, andInfAR, indicate that the excellence of our approach over other state-of-the-art models in wild scenarios. Wenxuan Liu 0008, Xian Zhong, Xuemei Jia, Kui Jiang, Chia-Wen Lin |
IEEE Signal Process. Lett. | 2 |
| 2022 | Grayscale Enhancement Colorization Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is an emerging and challenging cross-modality image matching problem because of the explosive surveillance data in night-time surveillance applications. To handle the large modality gap, various generative adversarial network models have been developed to eliminate the cross-modality variations based on a cross-modal image generation framework. However, the lack of point-wise cross-modality ground-truths makes it extremely challenging to learn such a cross-modal image generator. To address these problems, we learn the correspondence between single-channel infrared images and three-channel visible images by generating intermediate grayscale images as auxiliary information to colorize the single-modality infrared images. We propose a grayscale enhancement colorization network (GECNet) to bridge the modality gap by retaining the structure of the colored image which contains rich information. To simulate the infrared-to-visible transformation, the point-wise transformed grayscale images greatly enhance the colorization process. Our experiments conducted on two visible-infrared cross-modality person re-identification datasets demonstrate the superiority of the proposed method over the state-of-the-arts. Xian Zhong, Tianyou Lu, Wenxin Huang, Mang Ye, Xuemei Jia, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Complementary Data Augmentation for Cloth-Changing Person Re-IdentificationabstractThis paper studies the challenging person re-identification (Re-ID) task under the cloth-changing scenario, where the same identity (ID) suffers from uncertain cloth changes. To learn cloth- and ID-invariant features, it is crucial to collect abundant training data with varying clothes, which is difficult in practice. To alleviate the reliance on rich data collection, we reinforce the feature learning process by designing powerful complementary data augmentation strategies, including positive and negative data augmentation. Specifically, the positive augmentation fulfills the ID space by randomly patching the person images with different clothes, simulating rich appearance to enhance the robustness against clothes variations. For negative augmentation, its basic idea is to randomly generate out-of-distribution synthetic samples by combining various appearance and posture factors from real samples. The designed strategies seamlessly reinforce the feature learning without additional information introduction. Extensive experiments conducted on both cloth-changing and -unchanging tasks demonstrate the superiority of our proposed method, consistently improving the accuracy over various baselines. Xuemei Jia, Xian Zhong, Mang Ye, Wenxuan Liu 0008, Wenxin Huang |
IEEE Trans. Image Process. | 2 |
| 2021 | Modeling Context-Guided Visual and Linguistic Semantic Feature for Video Captioning
Zhixin Sun, Xian Zhong, Shuqin Chen, Duxiu Feng |
ICANN (5) | 2 |
| 2021 | Part-Aligned Network with Background for Misaligned Person SearchabstractPerson search is a significant computer vision task that requires addressing person detection and re-identification simultaneously. Body parts are frequently misaligned due to variation poses, occlusions, and partial missing, leading to the unsatisfied results of person search. Existing methods usually extract local features from the human body by the key point information, that cannot tackle the recognition task between a pair of persons with different body parts due to misalignment. Moreover, these methods overlook background information (e.g. the carries and the background reference object) which can also supplement effective features for representing the person. In this paper, we propose a part-aligned network with background (PANB) to address this misalignment issue. To learn local fine-grained features of different body parts, we fine-tune a parsing network to divide the body region into seven parts. In particular, our proposed method considers extracting the background features as the eighth part features to extract more robust representations, which is more rational and efficient. Furthermore, we design a reconstruction method to align the parts existing in both the query image and the cropped gallery image. Extensive experiments show that our proposed method achieves competitive performance on CUHK-SYSU and PRW datasets. Xian Zhong, Wenxin Huang, Xiao Wang 0029, Jingling Yuan |
ICASSP | 1 |
| 2021 | Person Retrieval in Physical WorldabstractPerson re-identification (re-ID) gains plenty of achievements as a retrieval problem in constrained camera networks. However, most of the researches are concentrated on visual appearance, they still suffer from the complicated environments in unconstrained urban/campus surveillance scenario due to unreliable visual representations with extremely challenging problems as lack of training samples, amounts of irrelevant crowds, etc. Besides, most of the existing person re-ID datasets neglect the physical truth of realistic investigation application: 1) investigators search only few suspects among amounts of crowds. Moreover, he may not go through by every camera in the surveillance area and may appear in the same camera several times; and 2) the corresponding characteristic in multi-space of the same ID can be verified with each other. Therefore, we propose a person retrieval in physical world (PRPW) dataset with large-scale unconstrained surveillance scenario. It contains over 1.4 million bounding boxes, including 20 labeled IDs and numerous irrelevant crowds captured by 86 cameras. Furthermore, over 30,000 records of 20 mobile trajectories are collected in this dataset, and the 20 mobile trajectories are partially overlapped while passing by 86 cameras. Finally, based on two common senses and a verification experiment, we provide a proposal to tackle with PRPW task on the basis of trajectory association which utilizes global optimization to compensate for the errors caused by visual expression on local observation points. The comparison experiments with two typical unsupervised person re-ID methods are implemented on the constructed dataset. Wenxin Huang, Ruimin Hu, Chao Liang 0001, Xian Zhong |
ICME | 5 |
| 2021 | Auxiliary Bi-Level Graph Representation for Cross-Modal Image-Text RetrievalabstractImage-text retrieval is one of the most common tasks in multimodal retrieval. It suffers from the problem of information imbalance between modalities, which is so-called modality gap. It remains challenging because prior methods cannot bridge the gap reasonably. With the help of scene graph, we start by designing an auxiliary bi-level graph representation (ABGR) pipeline that can fully mine the potential information and reduce the information redundancy. By doing so, each modality will be represented by lexical word graph that carries the main content of the information. Specifically, we design a graph feature enhancement (GFE) module to embed the graph-structured information in a common subspace while exploring the relationship between lexical words. As a result, a better representation for both image and text can be obtained, which helps us to evaluate the similarity between images and texts more reasonably. Experimental results conducted on two benchmark datasets Flickr30K and MS-COCO demonstrate the effectiveness of our proposed model for cross-modal retrieval task. Xian Zhong, Zhengwei Yang 0001, Mang Ye, Wenxin Huang, Jingling Yuan, Chia-Wen Lin |
ICME | 1 |
| 2021 | Imbalanced Multi-Class Classification of Hyperspectral Image Based on Smote and Deep Rotation ForestabstractIn this paper, a novel Synthetic Minority Oversampling Technique based Deep Rotation Forest(SMOTE-DRoF) algorithm is proposed for the classification of imbalanced hyperspectral image data. It builds a multi -level forests cascade model by training a balanced dataset generated by SMOTE. In this model, each level of the random forest produces misclassification information of the data which are used as guidance information to adjust the sample weight adaptively for the next level. Experiment results on the hyperspectral image Indian Pines AVRIS and University of Pavia ROSIS demonstrate that the proposed method can get better performance than support vector machine, random forest, rotation forest, SMOTE combined random forest, and SMOTE combined rotation forest in imbalance learning. Xian Zhong, Yinghui Quan, Wei Feng 0004, Qiang Li 0029, Gabriel Dauphin, Mengdao Xing |
IGARSS | 1 |
| 2021 | Unsupervised Vehicle Search in the Wild: A New BenchmarkabstractIn urban surveillance systems, finding a specific vehicle in video frames efficiently and accurately has always been an essential part of traffic supervision and criminal investigation. Existing studies focus on vehicle re-identification (re-ID), but vehicle search is still underexploited. These methods depend on the locations of many vehicles (bounding boxes) that are not available in most real-world applications. Therefore, the unsupervised joint study of vehicle location and identification for the observed scene is a pressing need. Inspired by person search, we conduct a study on the vehicle search while considering four main discrepancies among them, summarized as: 1) It is challenging to select the candidate regions for the observed vehicle due to the perspective differences (front or side); 2) The sides of the same type of vehicles are almost the same, resulting in smaller inter-class; 3) Lacking satisfied dataset for vehicle search to meet the practical scenarios; 4) Supervised search publishing methods rely on datasets with expensive annotations. To address these issues, we have established a new vehicle search dataset. We design an unsupervised framework on this benchmark dataset to generate pseudo labels for further training existing vehicle re-ID or person search models. Experimental results reveal that these methods turn less effective on vehicle search tasks. Therefore, the vehicle search task needs to be further developed, and this dataset can advance the research of vehicle search. Https://github.com/zsl1997/VSW. Xian Zhong, Xiao Wang 0029, Kui Jiang, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ACM Multimedia | 1 |
| 2021 | Local-enhanced Multi-resolution Representation Learning for Vehicle Re-identificationabstractIn real traffic scenarios, the changes of vehicle resolution that the camera captures tend to be relatively obvious considering the distances to the vehicle, different directions, and height of the camera. When the resolution difference exists between the probe and the gallery vehicle, the resolution mismatch will occur, which will seriously influence the performance of the vehicle re-identification (Re-ID). This problem is also known as multi-resolution vehicle Re-ID. An effective strategy is equivalent to utilize image super-resolution to handle the resolution gap. However, existing methods conduct super-resolution on global images instead of local representation of each image, leading to much more noisy information generated from the background and illumination variations. In our work, a local-enhanced multi-resolution representation learning (LMRL) is therefore proposed to address these problems by combining the training of local-enhanced super-resolution (LSR) module and local-guided contrastive learning (LCL) module. Specifically, we use a parsing network to parse a vehicle into four different parts to extract local-enhanced vehicle representation. And then, the LSR module, which consists of two auto-encoders that share parameters, transforms low-resolution images into high-resolution in both global and local branches. LCL module can learn discriminative vehicle representation by contrasting local representation between the high-resolution reconstructed image and the ground truth. We evaluate our approach on two public datasets that contain vehicle images at a wide range of resolutions, in which our approach shows significant superiority to the existing solution. Xian Zhong, Jingling Yuan, Rongbo Zhang, Duxiu Feng, Luo Zhong |
MMAsia | 2 |
| 2021 | Random Walk Erasing with Attention Calibration for Action Recognition
Yuze Tian, Xian Zhong, Wenxuan Liu 0008, Xuemei Jia, Mang Ye |
PRICAI (3) | 2 |
| 2021 | Subspace Enhancement and Colorization Network for Infrared Video Action Recognition
Xian Zhong, Wenxuan Liu 0008, Zhengwei Yang 0001, Luo Zhong |
PRICAI (3) | 2 |
| 2021 | Attention-guided image captioning with adaptive global and local feature fusion
Xian Zhong, Guozhang Nie, Wenxin Huang, Wenxuan Liu 0008, Chia-Wen Lin |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | Multi-Scale Residual Network for Image ClassificationabstractMulti-scale approach representing image objects at various levels-of-details has been applied to various computer vision tasks. Existing image classification approaches place more emphasis on multi-scale convolution kernels, and overlook multi-scale feature maps. As such, some shallower information of the network will not be fully utilized. In this paper, we propose the Multi-Scale Residual (MSR) module that integrates multi-scale feature maps of the underlying information to the last layer of Convolutional Neural Network. Our proposed method significantly enhances the characteristics of the information in the final classification. Extensive experiments conducted on CIFAR100, Tiny-ImageNet and large-scale CalTech-256 datasets demonstrate the effectiveness of our method compared with Res-Family. Xian Zhong, Oubo Gong, Wenxin Huang, Jingling Yuan, Ryan Wen Liu |
ICASSP | 1 |
| 2020 | Dual-Direction Perception and Collaboration Network for Near-Online Multi-Object TrackingabstractBackward tracks have been exploited to improve performance of multi-object tracking (MOT). The existing method brings a stable similarity measurement but neglects unreliable detection. Exploiting predictions of forward tracks has emerged as a popular approach to tackle the task of tracking-by-detection. However, it's observed that missing detection has not been solved well enough which would significantly influence the tracking accuracy. Thus, obtaining more proposals from dual-direction tracking and predictions of tracks is concerned to address the problem of missing detection. In this paper, we propose a dual-direction perception and collaboration network (DPCNet) for MOT that exploits forward and backward tracking to collaboratively track objects. It collects candidates from the dual directions so that they can complement each other in different scenarios. Moreover, we propose a near-online tracking model based on DPCNet to improve the efficiency, which batches the tracking and makes forward and backward tracking in parallel. Experiments conducted on MOT challenge benchmarks demonstrate that the proposed method outperforms the state-of-the-arts. Xian Zhong, Weijian Ruan, Wenxin Huang, Jingling Yuan |
ICIP | 1 |
| 2020 | A Lightweight High-Resolution Representation Backbone For Real-Time Keypoint-Based Object DetectionabstractThe keypoint based detectors are a relatively new object detection mechanism, avoiding the complicated computation related to anchor box and achieving state-of-the-art accuracy. However, inference speed is a major drawback of these detectors because of the heavy backbone network. In this paper, we design a novel lightweight backbone named DNet for keypoint-based detection and propose a real-time object detection network. In the backbone part, DNet is able to maintain high-resolution feature maps throughout the process and gradually extract and integrate features across scales. In the detection part, we detect a center keypoint and a pair of corners to predict the bounding boxes, and completely avoid the complicated computation related to anchor boxes. Compared with state-of-the-art real-time detectors, our network achieves superior performance with 30.0% AP on COCO benchmark at 21. 5ms. In addition, the experimental results show that our network is capable of running real-time on embedded devices. Jiansheng Dong, Jingling Yuan, Lin Li 0001, Xian Zhong |
ICME | 4 |
| 2020 | Complementing Representation Deficiency in Few-shot Image Classification: A Meta-Learning ApproachabstractFew-shot learning is a challenging problem that has attracted more and more attention recently since abundant training samples are difficult to obtain in practical applications. Meta-learning has been proposed to address this issue, which focuses on quickly adapting a predictor as a base-learner to new tasks, given limited labeled samples. However, a critical challenge for meta-learning is the representation deficiency since it is hard to discover common information from a small number of training samples or even one, as is the representation of key features from such little information. As a result, a meta-learner cannot be trained well in a high-dimensional parameter space to generalize to new tasks. Existing methods mostly resort to extracting less expressive features so as to avoid the representation deficiency. Aiming at learning better representations, we propose a meta-learning approach with complemented representations network (MCRNet) for few-shot image classification. In particular, we embed a latent space, where latent codes are reconstructed with extra representation information to complement the representation deficiency. Furthermore, the latent space is established with variational inference, collaborating well with different base-learners, and can be extended to other models. Finally, our end-to-end framework achieves the state-of-the-art performance in image classification on three standard few-shot learning datasets. Xian Zhong, Wenxin Huang, Lin Li 0001, Shuqin Chen, Chia-Wen Lin |
ICPR | 1 |
| 2020 | Two-Step Ensemble Based Class Noise Cleaning Method for Hyperspectral Image ClassificationabstractThe presence of noise is often unavoidable and has been a serious nuisance factor that needs to be taken into account in the hyperspectral image classification. Effective noise handling is one of the most difficult problems in data classification. Ensemble-based filtering has been demonstrated successful in dealing with the class noise problem. In this paper, a novel two-step ensemble-based data filtering method is proposed to improve the hyperspectral image classification accuracy in the presence of class noise. The proposed method is a combination of noise redundancy classifiers and sensitive algorithms. The experimental results on two public hyperspectral datasets demonstrate the effectiveness of the proposed approach. Wei Feng 0004, Yinghui Quan, Gabriel Dauphin, Xian Zhong, Qiang Li 0029, Mengdao Xing, Wenjiang Huang |
IGARSS | 4 |
| 2020 | Optimizing Queries over Video via Lightweight Keypoint-based Object DetectionabstractRecent advancements in convolutional neural networks based object detection have enabled analyzing the mounting video data with high accuracy. However, inference speed is a major drawback of these video analysis system because of the heavy object detectors. To address the computational and practicability challenges of video analysis, we propose FastQ, a system for efficient querying over video at scale. Given a target video, FastQ can automatically label the category and number of objects for each frame. We introduce a novel lightweight object detector named FDet to improve the efficiency of query system. First, a difference detector filters the frames whose difference is less than the threshold. Second, FDet is employed to efficiently label the remaining frames. To reduce inference time, FDet detects a center keypoint and a pair of corners from the feature map generated by a lightweight backbone to predict the bounding boxes. FDet completely avoid the complicated computation related to anchor boxes. Compared with state-of-the-art real-time detectors, FDet achieves superior performance with 29.1% AP on COCO benchmark at 25.3ms. Experiments show that FastQ achieves 150 times to 300 times speed-ups while maintaining more than 90% accuracy in video queries. Jiansheng Dong, Jingling Yuan, Lin Li 0001, Xian Zhong, Weiru Liu |
ICMR | 4 |
| 2020 | Visible-infrared Person Re-identification via Colorization-based Siamese Generative Adversarial NetworkabstractWith explosive surveillance data during day and night, visible-infrared person re-identification (VI-ReID) is an emerging challenge due to the apparent cross-modality discrepancy between visible and infrared images. Existing VI-ReID work mainly focuses on learning a robust feature to represent a person in both modalities despite the modality gap cannot be effectively eliminated. Recent research works have proposed various generative adversarial network (GAN) models to transfer the visible modality to another unified modality, aiming to bridge the cross-modality gap. However, they neglect the information loss caused by transferring the domain of visible images which is significant for identification. To effectively address the problems, we observe that key information such as textures and semantics in an infrared image can help to color the image itself and the colored infrared image maintains rich information from infrared image while reducing the discrepancy with the visible image. We therefore propose a colorization-based Siamese generative adversarial network (CoSiGAN) for VI-ReID to bridge the cross-modality gap, by retaining the identity of the colored infrared image. Furthermore, we also propose a feature-level fusion model to supplement the transfer loss of colorization. The experiments conducted on two cross-modality person re-identification datasets demonstrate the superiority of the proposed method compared with the state-of-the-arts. Xian Zhong, Tianyou Lu, Wenxin Huang, Jingling Yuan, Wenxuan Liu 0008, Chia-Wen Lin |
ICMR | 1 |
| 2020 | Video Human Behavior Recognition Based on ISA Deep Network ModelabstractVision-based behavior recognition is the analysis and recognition of human behavior in video. It has been widely used in many aspects such as multimedia information retrieval, behavior monitoring, and robot perception. This paper uses the Independent Subspace Analysis (ISA) deep network model feature extraction method, which is based on the ISA model and neural network theory, and combines data preprocessing methods, [Formula: see text]-means clustering methods, and Support Vector Machine (SVM) classifiers to achieve video classification and identification of human behavior. The ISA-based deep network model feature extraction method is an unsupervised learning method that can obtain behavior characteristics with good invariance and characterization capabilities in video human behavior. The experiment was conducted on the basis of the Hollywood2 human behavior data set. This experiment was compared with other commonly used human behavior feature extraction and recognition methods. The experimental results validated the effectiveness and advantages of this method in the classification and recognition of human behavior. Xian Zhong, Wenxin Huang, Ruiqi Luo |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2020 | Image classification with a MSF dropout
Ruiqi Luo, Xian Zhong, Enxiao Chen |
Multim. Tools Appl. | 2 |
| 2020 | An emotion classification algorithm based on SPT-CapsNet
Xian Zhong, Jinhang Liu, Lin Li 0001, Shuqin Chen, Yuyu Dong, Bingqing Wu, Luo Zhong |
Neural Comput. Appl. | 1 |
| 2020 | Adaptively Converting Auxiliary Attributes and Textual Embedding for Video Captioning Based on BiLSTM
Shuqin Chen, Xian Zhong, Lin Li 0001, Luo Zhong |
Neural Process. Lett. | 2 |
| 2020 | Image-to-video person re-identification with cross-modal embeddings
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Jianwen Xiang |
Pattern Recognit. Lett. | 3 |
| 2020 | An Encoder-Decoder Network Based FCN Architecture for Semantic SegmentationabstractIn recent years, the convolutional neural network (CNN) has made remarkable achievements in semantic segmentation. The method of semantic segmentation has a desirable application prospect. Nowadays, the methods mostly use an encoder-decoder architecture as a way of generating pixel by pixel segmentation prediction. The encoder is for extracting feature maps and decoder for recovering feature map resolution. An improved semantic segmentation method on the basis of the encoder-decoder architecture is proposed. We can get better segmentation accuracy on several hard classes and reduce the computational complexity significantly. This is possible by modifying the backbone and some refining techniques. Finally, after some processing, the framework has achieved good performance in many datasets. In comparison with the traditional architecture, our architecture does not need additional decoding layer and further reuses the encoder weight, thus reducing the complete quantity of parameters needed for processing. In this paper, a modified focal loss function is also put forward, as a replacement for the cross-entropy function to achieve a better treatment of the imbalance problem of the training data. In addition, more context information is added to the decode module as a way of improving the segmentation results. Experiments prove that the presented method can get better segmentation results. As an integral part of a smart city, multimedia information plays an important role. Semantic segmentation is an important basic technology for building a smart city. Yongfeng Xing, Luo Zhong, Xian Zhong |
Wirel. Commun. Mob. Comput. | 3 |
| 2019 | Squeeze-and-Excitation Wide Residual Networks in Image ClassificationabstractThe depth and width of the network have been investigated to influence the performance of image classification during the resent research. Wide residual networks (WRNs) have proved that the performance of classification can be improved by the width of the networks. With consideration of the significance, expanding the width is to increase the number of channels. However, not all the channels are needed. Meanwhile, much channel information will be lost while exploiting the global average pooling at the end of WRNs for image representations because the mean value is only related to the first order information. With the two considerations stated above, we propose squeeze-and-excitation WRNs which are based on the global covariance pooling (SE-WRNs-GVP). A residual Squeeze-and-Excitation block (rSE-block) can make up for the lost information due to global average pooling in SE-block. Then, informative channels of WRNs will be utilized. Finally, the global covariance pooling at the end of WRNs characterizes the correlations of feature channels for more discriminative representations. A SE-block with dropout is proposed to avoid over-fitting. We conduct experiments on CIFAR10 and CIFAR100 datasets and achieve a better performance without increasing the model complexity. Xian Zhong, Oubo Gong, Wenxin Huang, Lin Li 0001, Hongxia Xia |
ICIP | 1 |
| 2019 | An Efficient Semantic Segmentation Method using Pyramid ShuffleNet V2 with Vortex PoolingabstractEfficient and accurate semantic segmentation is particularly important especially for applications like autonomous driving which requires real-time inference speed and high performance. Many works try to compromise spatial resolution to achieve real-time inference speed, which leads to poor performance. As a result, real-time segmentation task for embedded devices is still an open problem. In this paper, we focus on building a network with better performance possible while still achieve real-time inference speed. We first use a pyramid kernel size to capture more spatial information instead of using just a 3×3 kernel size for DWConvolution in ShuffleNet v2. Meanwhile, an efficient Vortex Pooling module is employed to aggregate the contextual information and generate high-resolution features. Compared with other state-of-the-art real-time semantic segmentation networks, the proposed network achieves similar inference speed and better performance on embedded device. Specifically, we achieve state-of-the-art 73.46% mean IoU on Cityscapes test dataset, for a 768×1024 input, a speed of 46.1 frames per second on NVIDIA Jetson AGX Xavier embedded development board is achieved. Jiansheng Dong, Jingling Yuan, Lin Li 0001, Xian Zhong, Weiru Liu |
ICTAI | 4 |
| 2019 | Deep Multi-label Hashing for Image RetrievalabstractDue to its low storage cost and fast query speed, hashing has been widely applied to approximate nearest neighbor search for large-scale image retrieval, while deep hashing further improves the retrieval quality by learning a good image representation. However, existing deep hash methods simplify multi-label images into single-label processing, so the rich semantic information from multi-label is ignored. Meanwhile, the imbalance of similarity information leads to the wrong sample weight in the loss function, which makes unsatisfactory training performance and lower recall rate. In this paper, we propose Deep Multi-Label Hashing (DMLH) model that generates binary hash codes which retain the semantic relationship of multi-label of the image. The contributions of this new model mainly include the following two aspects: (1) A novel sample weight calculation model adaptively adjusts the weight of the sample pair by calculating the semantic similarity of the multi-label image pairs. (2) The sample weight cross-entropy loss function, which is designed according to the similarity of the image, adjusts the balance of similar image pairs and dissimilar image pairs. Extensive experiments demonstrate that the proposed method can generate hash codes which achieve better retrieval performance on two benchmark datasets, NUS-WIDE and MS-COCO. Xian Zhong, Jiachen Li 0002, Wenxin Huang |
ICTAI | 1 |
| 2019 | Poses Guide Spatiotemporal Model for Vehicle Re-identification
Xian Zhong, Meng Feng, Wenxin Huang, Zheng Wang 0007, Shin'ichi Satoh 0001 |
MMM (2) | 1 |
| 2019 | An Energy Dynamic Control Algorithm Based on Reinforcement Learning for Data CentersabstractIn recent years, how to use renewable energy to reduce the energy cost of internet data center (IDC) has been an urgent problem to be solved. More and more solutions are beginning to consider machine learning, but many of the existing methods need to take advantage of some future information, which is difficult to obtain in the actual operation process. In this paper, we focus on reducing the energy cost of IDC by controlling the energy flow of renewable energy without any future information. we propose an efficient energy dynamic control algorithm based on the theory of reinforcement learning, which approximates the optimal solution by learning the feedback of historical control decisions. For the purpose of avoiding overestimation, improving the convergence ability of the algorithm, we use the double [Formula: see text]-method to further optimize. The extensive experimental results show that our algorithm can on average save the energy cost by 18.3% and reduce the rate of grid intervention by 26.2% compared with other algorithms, and thus has good application prospects. Yao Xiang, Jingling Yuan, Ruiqi Luo, Xian Zhong, Tao Li 0006 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2019 | Enhancing multimodal deep representation learning by fixed model reuse
Zhongwei Xie, Lin Li 0001, Xian Zhong, Yang He 0003, Luo Zhong |
Multim. Tools Appl. | 3 |
| 2019 | Civil engineering supervision video retrieval method optimization based on spectral clustering and R-tree
Shifeng Wu, Huazhu Song, Gui Cheng, Xian Zhong |
Neural Comput. Appl. | 4 |
| 2018 | A Hybrid Model Reuse Training Approach for Multilingual OCR
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Qing Xie 0002, Jianwen Xiang |
WISE (1) | 3 |
| 2017 | Complete Tolerance Relation Based Filling Algorithm Using SparkabstractWith the advent of cloud computing, renewable energy is integrated into data center power supply systems increasingly. The power statistics collection may not be available due to the instability of renewable energy, which results in incomplete data. The incomplete energy data will significantly disturb the management of data centers. We further propose a filling algorithm based on complete tolerance class. The algorithm expands the traditional tolerance relation, and fills the missing values of the energy data, which ensures the data integrity. By taking good advantage of in-Memory Computing, We further parallelize and optimize our algorithm using Spark. The experiment results demonstrate that our algorithm outperforms other general filling algorithms in terms of filling accuracy. The proposed algorithm also shows good performance as the missing rate rises up. Jingling Yuan, Yao Xiang, Xian Zhong, Mincheng Chen, Tao Li 0006 |
ICDCS | 3 |
| 2016 | Camera Network Based Person Re-identification by Leveraging Spatial-Temporal Constraint and Multiple Cameras Relations
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Xian Zhong, Chunjie Zhang 0001 |
MMM (1) | 6 |