VLDB 2026 Research / reviewers in the wild / expert
Xiaopeng Hong
dblp:06/592
· DBLP profile ↗
144ranked-venue papers
8as first author
83since 2021 · last 2026
0000-0002-0611-0636ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 103 · 6 first-author · 56 since 2021Artificial intelligence and machine learning · 86 · 5 first-author · 48 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time Series Forecasting via Direct Per-Step Probability Distribution ModelingabstractDeep neural network-based time series prediction models have recently demonstrated superior capabilities in capturing complex temporal dependencies. However, it is challenging for these models to account for uncertainty associated with their predictions, because they directly output scalar values at each time step. To address such a challenge, we propose a novel model named interleaved dual-branch Probability Distribution Network (interPDN), which directly constructs discrete probability distributions per step instead of a scalar. The regression output at each time step is derived by computing the expectation of the predictive distribution on a predefined support set. To mitigate prediction anomalies, a dual-branch architecture is introduced with interleaved support sets, augmented by coarse temporal-scale branches for long-term trend forecasting. Outputs from another branch are treated as auxiliary signals to impose self-supervised consistency constraints on the current branch's prediction. Extensive experiments on multiple real-world datasets demonstrate the superior performance of interPDN. Linghao Kong, Xiaopeng Hong |
AAAI | 2 |
| 2026 | 2D Gaussians Spatial Transport for Point-supervised Density RegressionabstractThis paper introduces Gaussian Spatial Transport (GST), a novel framework that leverages Gaussian splatting to facilitate transport from the probability measure in the image coordinate space to the annotation map. We propose a Gaussian splatting-based method to estimate pixel-annotation correspondence, which is then used to compute a transport plan derived from Bayesian probability. To integrate the resulting transport plan into standard network optimization in typical computer vision tasks, we derive a loss function that measures discrepancy after transport. Extensive experiments on representative computer vision tasks, including crowd counting and landmark detection, validate the effectiveness of our approach. Compared to conventional optimal transport schemes, GST eliminates iterative transport plan computation during training, significantly improving efficiency. Miao Shang, Xiaopeng Hong |
AAAI | 2 |
| 2026 | Sample-Aware Knowledge Association and Enhancement for Open-Vocabulary Continual Learning
Zhilin Zhu 0001, Zhiheng Ma, Yabin Wang 0001, Yaguang Song, Yaowei Wang 0001, Xiaopeng Hong |
Int. J. Comput. Vis. | 6 |
| 2026 | Penny-Wise and Pound-Foolish in AI-Generated Image DetectionabstractThe rise of AI-generated images has sparked serious concerns about their potential misuse across various domains, prompting the urgent need for robust detection methods. Despite advancements, many current approaches prioritize short-term gains at the expense of long-term effectiveness. This paper critiques the overly specialized approach of fine-tuning pre-trained models for short-term gains on a single AI image dataset, while disregarding the long-term imperative of achieving generalization and knowledge retention. To address this trade-off issue, we propose a novel learning framework (PoundNet) for the generalization of AI-generated image detection on a pre-trained vision-language model. PoundNet incorporates a learnable prompt design and a balanced objective to preserve broad knowledge from upstream tasks (object classification) while enhancing generalization for downstream tasks (AI-generated image detection). We train PoundNet on a single standard AI image dataset, following common practice in the literature. We then evaluate its performance across 10 large-scale public AI-generated image detection datasets with 5 main evaluation metrics, forming the largest benchmark test set for assessing the generalization ability of AI-generated image detection models, to our knowledge. The comprehensive benchmark evaluation demonstrates that PoundNet successfully balances generalization with knowledge retention, achieving a remarkable relative improvement of 19% in AI-generated image detection performance compared to state-of-the-art methods, while maintaining a strong performance of 63% on object classification tasks. Yabin Wang 0001, Zhiwu Huang, Zhou Su 0001, Adam Prügel-Bennett, Xiaopeng Hong |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Dual-Attention based prompt generation and catalyzing for instance-wise continual learning
Xiaopeng Hong, Yabin Wang 0001, Zhiheng Ma, Jinfeng Yang, Dongmei Jiang, Yaowei Wang 0001 |
Pattern Recognit. | 2 |
| 2026 | Linguistic profiling of deepfakes: An open database for next-Generation deepfake detection
Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhiwu Huang |
Pattern Recognit. | 2 |
| 2026 | Asymmetric modal fusion for multi-modal crowd counting
Xiaopeng Hong, Zhiheng Ma, Yabin Wang 0001 |
Pattern Recognit. | 2 |
| 2026 | Task Memory Sinkhorn Neuralization under varying measure distributions
Xiaopeng Hong, Shuangxiu Li, Wangmeng Zuo, Xiaopeng Fan 0001 |
Pattern Recognit. | 2 |
| 2026 | A Survey on Deep Learning for Group-Level Emotion RecognitionabstractWith the rapid advancement of artificial intelligence, group-level emotion recognition (GER) has emerged as an important domain in human behavior analysis. Early GER methods primarily relied on handcrafted features. However, the recent success of deep learning has shifted the focus toward neural network-based solution, enabling more effective exploitation of the rich visual and contextual cues in group images and videos. Unlike individual-level emotion recognition, GER must account for the diversity and dynamics of multiple individuals within varied social contexts. Over the past decade, numerous deep learning-based methods have been proposed, achieving substantial performance gains. This survey provides a comprehensive review of deep learning-centric review of GER, introducing a new taxonomy that spans representation learning, graph-based modeling, attention and transformer architectures, and multimodal fusion strategies. We summarize benchmark datasets, outline prevailing GER pipelines, and consolidate performance trends from recent state-of-the-art approaches. In addition, we discuss the integration of foundation models and large language model-guided multimodal reasoning into GER. Key challenges are identified, and potential research directions are proposed to support the development of robust, real-world GER systems. This work aims to serve as a pivotal reference for future research in this evolving field. Xiaohua Huang 0003, Xiaopeng Hong, Qirong Mao, Wenming Zheng, Abhinav Dhall |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | Language Supervised Multi-Camera Multi-Object TrackingabstractRecent multi-camera multi-object tracking (MCMOT) algorithms are primarily trained using per-detection identity annotations, which are complicated to obtain. In contrast, labeling a language description per-object is a more natural and human-friendly way. In this paper, we explore MCMOT in a language-supervised manner (LS-MCMOT) and propose a novel approach LaVST, which performs language-to-vision weakly-supervised learning based on reliable pseudo-labels generated via tracklet-level cross-modality matching. In addition, we design an ID-aware projection self-correction mechanism to correct inaccurate image-to-ground projection in a self-supervised manner. The models trained with our approach exhibit promising performance in LS-MCMOT. Surprisingly, they perform favorably against state-of-the-art identity-supervised methods, especially in cross-dataset evaluation (with an average gain by 20.0% in IDF1), underscoring the potential of language annotations in MCMOT. Codes and language annotations will be available here. Kaige Mao, Xiaopeng Hong, Xiaopeng Fan 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2026 | Continual Conceptual Entity Learning for Text-to-Image Generative ModelsabstractCurrent Text-to-Image generative models struggle to continuously learn multiple distinct entities or concepts, limiting their scalability and hindering practical deployment in dynamic environments. We formulate this task as Continual Conceptual Entity Learning (CEL) and propose a novel framework called Continual Entity Adapter Learning (CEAL). CEAL leverages a compact set of tunable parameters, termed SuperLoRA, to efficient and scalable learning of new entities. We propose a dynamic rank-increasing strategy to train the SuperLoRA, balancing computational efficiency with performance. To evaluate our method, we create three benchmarks encompassing generic objects, human faces, and artistic styles. Experimental results demonstrate that CEAL effectively learns new entities while preserving prior knowledge, outperforming existing methods in both entity fidelity and parameter efficiency. Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhou Su 0001, Zhiwu Huang |
IEEE Trans. Multim. | 2 |
| 2026 | Hierarchical Concept Bottleneck With Compensation Concept LearningabstractConcept Bottleneck Models (CBMs) enhance the interpretability of deep neural networks by mapping images to human-understandable concepts and then using the concepts to make predictions. While they improve transparency, existing CBMs primarily explain only the final layer's features, limiting the interpretability of intermediate layers. Additionally, constructing a comprehensive concept set remains a challenging task, further constraining model performance. In this paper, we investigate the assignment of concept granularity across model layers and propose theHierarchicalConceptBottleneckModel (HCBM) to enhance interpretability. HCBM introduces a Hybrid Concept Bottleneck Layer (HCBL) at each layer, consisting of a Predefined Concept Bottleneck (PCB) that maps visual features to concepts of corresponding granularity and a Compensation Concept Bottleneck (CCB) which incorporates the concept frequency loss and the concept semantic loss to capture compensation concepts for improving performance. Extensive experiments demonstrate that HCBM outperforms state-of-the-art methods. It is worth noting that the HCBM with CLIP RN50 as the backbone outperforms the black-box model. Miao Shang, Kaige Mao, Xiaopeng Hong, Xuhui Huang |
IEEE Trans. Multim. | 4 |
| 2025 | Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and RegularizationabstractAudio-Visual Learning (AVL) aims at the audio-visual perception with both audio and vision modalities. AVL also suffers from data insufficiency in many applications as with other unimodal tasks. Concurrently, AVL often needs to continuously learn over time rather than all knowledge simultaneously. Considering the above two perspectives, our work mainly focuses on benchmarking the unexplored Few-Shot Audio-Visual Class-Incremental Learning (FS-AVCIL), i.e., continually perceiving novel categories described by a limited number of labeled examples with audio and visual modalities. Firstly, we provide the detailed task configuration together with a thorough analysis of the challenges in FS-AVCIL: (1) how to efficiently learn and fuse multimodal information with limited labeled examples; and (2) how to alleviate catastrophic forgetting cross-modal semantic correlations with limited data. Then, we propose an efficient framework based on Vision Transformer to solve FS-AVCIL. This framework contains two parts: temporal-residual prompting for audio-visual synergy adapter and temporal prompt regularization. Specifically, temporal-residual prompting is incorporated into the audio-visual adapter to efficiently finetune the pre-trained foundation model with limited data and capture audio-visual correlation by learning temporal-relevant prompts. Besides, we regularize temporal-relevant prompts to memorize previous knowledge by fully using the temporal knowledge from various perspectives. This framework is validated in audio-visual classification tasks under the FS-AVCIL scenario, and extensive experiments demonstrate its superior performance. Yawen Cui, Zitong Yu, Guanjie Huang, Xiaopeng Hong |
AAAI | 5 |
| 2025 | ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge EditingabstractLarge multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current multimodal knowledge editing evaluations are limited in scope and potentially biased, focusing on narrow tasks and failing to assess the impact on in-domain samples. To address these issues, we introduce ComprehendEdit, a comprehensive benchmark comprising eight diverse tasks from multiple datasets. We propose two novel metrics: Knowledge Generalization Index (KGI) and Knowledge Preservation Index (KPI), which evaluate editing effects on in-domain samples without relying on AI-synthetic samples. Based on insights from our framework, we establish Hierarchical In-Context Editing (HICE), a baseline method employing a two-stage approach that balances performance across all metrics. This study provides a more comprehensive evaluation framework for multimodal knowledge editing, reveals unique challenges in this field, and offers a baseline method demonstrating improved performance. Our work opens new perspectives for future research and provides a foundation for developing more robust and effective editing techniques for MLLMs. Yaohui Ma, Xiaopeng Hong, Shizhou Zhang, Huiyun Li, Zhilin Zhu 0001, Wei Luo 0014, Zhiheng Ma |
AAAI | 2 |
| 2025 | Specifying What You Know or Not for Multi-Label Class-Incremental LearningabstractExisting class incremental learning is mainly designed for single-label classification task, which is ill-equipped for multi-label scenarios due to the inherent contradiction of learning objectives for samples with incomplete labels. We argue that the main challenge to overcome this contradiction in multi-label class-incremental learning (MLCIL) lies in the model's inability to clearly distinguish between known and unknown knowledge. This ambiguity hinders the model's ability to retain historical knowledge, master current classes, and prepare for future learning simultaneously. In this paper, we target at specifying what is known or not to accommodate Historical, Current, and Prospective knowledge for MLCIL and propose a novel framework termed as HCP. Specifically, (i) we clarify the known classes by dynamic feature purification and recall enhancement with distribution prior, enhancing the precision and retention of known information. (ii) We design prospective knowledge mining to probe the unknown, preparing the model for future learning. Extensive experiments validate that our method effectively alleviates catastrophic forgetting in MLCIL, surpassing the previous state-of-the-art by 3.3% on average accuracy for MS-COCO B0-C10 setting without replay buffers. Aoting Zhang, Dongbao Yang, Xiaopeng Hong, Yu Zhou 0015 |
AAAI | 4 |
| 2025 | DCA: Dividing and Conquering Amnesia in Incremental Object DetectionabstractIncremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous classes. Existing methods achieve certain success by improving knowledge distillation and exemplar replay for transformer-based detection frameworks, but the intrinsic forgetting mechanisms remain underexplored. In this paper, we dive into the cause of forgetting and discover forgetting imbalance between localization and recognition in transformer-based IOD, which means that localization is less-forgetting and can generalize to future classes, whereas catastrophic forgetting occurs primarily on recognition. Based on these insights, we propose a Divide-and-Conquer Amnesia (DCA) strategy, which redesigns the transformer-based IOD into a localization-then-recognition process. DCA can well maintain and transfer the localization ability, leaving decoupled fragile recognition to be specially conquered. To reduce feature drift in recognition, we leverage semantic knowledge encoded in pre-trained language models to anchor class representations within a unified feature space across incremental tasks. This involves designing a duplex classifier fusion and embedding class semantic features into the recognition decoding process in the form of queries. Extensive experiments validate that our approach achieves state-of-the-art performance, especially for long-term incremental scenarios. For example, under the four-step setting on MS-COCO, our DCA strategy significantly improves the final AP by 6.9%. Aoting Zhang, Dongbao Yang, Xiaopeng Hong, Miao Shang, Yu Zhou 0015 |
AAAI | 4 |
| 2025 | Free Lunch Enhancements for Multi-modal Crowd CountingabstractThis paper addresses multi-modal crowd counting with a novel ‘free lunch’ training enhancement strategy that requires no additional data, parameters, or increased inference complexity. First, we introduce a cross-modal alignment technique as a plug-in post-processing step for the pre-trained backbone network, enhancing the model’s ability to capture shared information across modalities. Second, we incorporate a regional density supervision mechanism during the fine-tuning stage, which differentiates features in regions with varying crowd densities. Extensive experiments on three multi-modal crowd counting datasets validate our approach, making it the first to achieve an MAE below 10 on RGBT-CC. The code is available at https://github.com/HenryCilence/Free-Lunch-Multimodal-Counting. Haoliang Meng, Xiaopeng Hong, Zhengqin Lai, Miao Shang |
CVPR | 2 |
| 2025 | T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingabstractZero-Shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-language models like CLIP, but often exhibit limited sensitivity to text prompts. We present T21 Count, a diffusion-based framework that lever-ages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step demising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the demising U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/chal5yq/T2lCount. Yifei Qian, Zhongliang Guo 0001, Bowen Deng 0006, Chun Tong Lei, Shuai Zhao 0007, Chun Pong Lau 0001, Xiaopeng Hong, Michael P. Pound |
CVPR | 7 |
| 2025 | OpenSDI: Spotting Diffusion-Generated Images in the Open WorldabstractThis paper identifies OpenSDI, a challenge for spotting diffusion-generated images in open-world settings. In response to this challenge, we define a new benchmark, the OpenSDI dataset (OpenSDID), which stands out from existing datasets due to its diverse use of large vision-language models that simulate open-world diffusion-based manipulations. Another outstanding feature of OpenSDID is its inclusion of both detection and localization tasks for images manipulated globally and locally by diffusion models. To address the OpenSDI challenge, we propose a Synergizing Pretrained Models (SPM) scheme to build up a mixture of foundation models. This approach exploits a collaboration mechanism with multiple pretrained foundation models to enhance generalization in the OpenSDI context, moving beyond traditional training by synergizing multiple pretrained models through prompting and attending strategies. Building on this scheme, we introduce MaskCLIP, an SPM-based model that aligns Contrastive Language-Image Pre-Training (CLIP) with Masked Autoencoder (MAE). Extensive evaluations on OpenSDID show that MaskCLIP significantly outperforms current state-of-the-art methods for the OpenSDI challenge, achieving remarkable relative improvements of 14.23% in IoU (14.11% in F1) and 2.05% in accuracy (2.38% in F1) compared to the second-best model in localization and detection tasks, respectively. Our dataset and code are available at https://github.com/iamwangyabin/OpenSDI. Yabin Wang 0001, Zhiwu Huang, Xiaopeng Hong |
CVPR | 3 |
| 2025 | MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question AnsweringabstractFacial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In recent years, substantial advancements have been made in the areas of ME recognition, spotting, and generation. However, conventional approaches that treat spotting and recognition as separate tasks are suboptimal, particularly for analyzing long-duration videos in realistic settings. Concurrently, the emergence of multimodal large language models (MLLMs) and large vision-language models (LVLMs) offers promising new avenues for enhancing ME analysis through their powerful multimodal reasoning capabilities. The ME grand challenge (MEGC) 2025 introduces two tasks that reflect these evolving research directions: (1) ME spot-then-recognize (ME-STR), which integrates ME spotting and subsequent recognition in a unified sequential pipeline; and (2) ME visual question answering (ME-VQA), which explores ME understanding through visual question answering, leveraging MLLMs or LVLMs to address diverse question types related to MEs. All participating algorithms are required to run on this test set and submit their results on a leaderboard. More details are available at https://megc2025.github.io. Xinqi Fan, Jingting Li 0001, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaopeng Hong, Adrian K. Davison |
ACM Multimedia | 7 |
| 2025 | Agent-MER: A Cognitive Agent with Hierarchical Deliberation for Open-Vocabulary Multimodal Emotion RecognitionabstractThis paper focuses on Open-Vocabulary Multimodal Emotion Recognition (OV-MER) and is dedicated to solving the two challenges it faces: concept semantic misalignment and incomplete coverage of fine-grained emotion categories. To address this, we propose a novel cognitive agent framework (Agent-MER), which reframes the OV-MER task as a problem to be solved by an agent that mimics the human cognitive process through knowledge-guided deliberation. We first construct a hierarchical Emotion Tree to serve as the agent's knowledge base. Building on this, we design a Knowledge-Guided Hierarchical Deliberation reasoning process. This process systematically explores the entire emotional landscape through a three-level, coarse-to-fine iterative reasoning process, enabling the identification of a richer and deeper range of emotions. Finally, a Self-Consistent Voting mechanism is employed to aggregate the results from multiple reasoning runs, ensuring the robustness of the final output. Experiments conducted in the MER2025 Challenge demonstrate that our proposed method achieved a top-ranking score of 61.04%, securing first place and significantly outperforming existing baselines. This work not only provides an effective solution for OV-MER but also opens up new avenues for developing more human-like affective intelligence systems. Zhengqin Lai, Zhilin Zhu 0001, Xiaopeng Hong, Yaowei Wang 0001 |
ACM Multimedia | 3 |
| 2025 | Semi-Supervised Counting via Pixel-by-Pixel Density Distribution ModelingabstractThis paper focuses on semi-supervised crowd counting, where only a small portion of the training data are labeled. We formulate the pixel-wise density value to regress as a probability distribution, instead of a single deterministic value. On this basis, we propose a semi-supervised crowd counting model. First, we design a pixel-wise distribution matching loss to measure the differences in the pixel-wise density distributions between the prediction and the ground-truth; Second, we enhance the transformer decoder by using density tokens to specialize the forwards of decoders w.r.t. different density intervals; Third, we design the interleaving consistency self-supervised learning mechanism to learn from unlabeled data efficiently. Extensive experiments on four datasets are performed to show that our method clearly outperforms the competitors by a large margin under various labeled ratio settings. Zhiheng Ma, Rongrong Ji, Yaowei Wang 0001, Zhou Su 0001, Xiaopeng Hong, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Perspective-assisted prototype-based learning for semi-supervised crowd counting
Yifei Qian, Liangfei Zhang, Zhongliang Guo 0001, Xiaopeng Hong, Ognjen Arandjelovic, Carl Donovan |
Pattern Recognit. | 4 |
| 2025 | Energy Efficient Multi-Robot Task Allocation Constrained by Time Window and PrecedenceabstractTo meet the demands in terms of energy-efficient and fast production and delivery of goods, robotic fleets began to populate warehouses and industrial environments. To maximize the profitability of the operations, multi-robot systems are required to coordinate agents and avoid downtime efficiently. In this paper, agent coordination is formulated as a multi-robot task allocation (MRTA) problem with time and precedence constraints. The method capitalizes on a graph method to build a measure graph reflecting the sparsity of tasks and a precedence graph, which includes the task constraints, to group the tasks into batches. A batch solver is provided to obtain the final solutions to the MRTA. In this way, the sustainability and environmental impact of logistics operations can be improved by reducing the number of robots needed to complete tasks and also by assigning tasks closest to the robot location, reducing the amount of time and the total energy required for the robots to complete the job. Extensive experiments on both uniformly distributed and sparse data sets prove the effectiveness of the proposed algorithm compared to state-of-the-art algorithms such as MIP and TePSSI.Note to Practitioners—This paper was motivated by the problem of minimizing the energy consumption of multi-robot systems in the execution of complex tasks, which requires, in the most general case, the motion of the robot to a target location and further on-site operations. This scenario is particularly relevant in smart, automated warehouses, where mobile robots are repeatedly demanded to store or dispatch goods in a structured environment, where operation duration and future requests are known a priori. The paper formulates this problem by means of a batched multi-robot task allocation (BMRTA) optimization, which can include time windows and precedence constraints jointly. First, the task constraints are encoded into two graphs and then combined to group subtasks together in batches. Then, each batch is solved separately, minimizing the overall energy required to achieve the tasks in the batch. Although the optimality of the solution is ensured only locally, i.e., within the same batch, the task clustering improves the computational efficiency with respect to global approaches, especially in large-sized problems. Experimental results demonstrated that when comparing BMRTA with literature approaches such as MIP and TePSSI, not only the energy consumption but also the total travel distance can be minimized while the total duration of the tasks remains comparable. Lixuan Zhang, Jianzhuang Zhao, Edoardo Lamon, Yabin Wang 0001, Xiaopeng Hong |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Joint Memory Optimization for Continual LearningabstractContinual learning, focusing on sequential knowledge acquisition and retention, necessitates efficient memory management. This paper introduces a holistic approach, diverging from traditional methods that separately optimize neural network and replay buffer memory. We aim to enhance overall memory efficiency, addressing neural network parameters and replay buffer concurrently within strict memory constraints. This is achieved by harnessing neural network parameter redundancies and employing compression techniques like pruning and quantization, allowing data replay storage without extra memory overhead. Balancing memory use across components is challenging due to the complex search space of combined tasks. We tackle this by conceptualizing it as a bi-level optimization problem, integrating all tasks under a single objective, thus optimizing memory use and managing the interplay between different components. We employ a synergy of optimization techniques to solve this challenging bi-level optimization problem. Our experimental findings affirm the superior performance of our proposed method, outperforming existing techniques such as prompt-based, feature-replay, exemplar-replay, and regularization-based methods under stringent memory constraints, consistently across various datasets and neural network architectures. Zhiheng Ma, Yaohui Ma, Xiaopeng Hong, Huiyun Li, Shizhou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Swiss Army Knife for Tracking by Natural Language SpecificationabstractTracking by natural language specification requires trackers to jointly perform grounding and tracking tasks. Existing methods either use separate models or a single shared network, failing to account for the link and diversity between tasks jointly. In this paper, we propose a novel framework that performs dynamic task switching to customize its network path routing for each task within a unified model. For this purpose, we design a task-switchable attention module, which enables the acquisition of modal relation patterns with different dominant modalities for each task via dynamic task switching. In addition, to alleviate the inconsistency between the static language description and the dynamic target appearance during tracking, we propose a language renovation mechanism that renovates the initial language online via visual-context-aware linguistic prompting. Extensive experimental results on five datasets demonstrate that the proposed method performs favorably against state-of-the-art approaches for both grounding and tracking. Our project will be available at: https://github.com/mkg1204/SAKTrack. Kaige Mao, Xiaopeng Hong, Xiaopeng Fan 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2025 | Multi-Scale Spiking Pyramid Wireless Communication Framework for Food RecognitionabstractFood recognition applications in human health have recently garnered significant attention in the field of computer vision. With the advancement of mobile devices, robust food recognition in wireless communication has become a practical and challenging application scenario. We propose a novel Multi-scale Spiking Pyramid Transmission Network (MSPTN) to tackle this challenge. The MSPTN learns diverse and complementary local and global feature maps simultaneously, generating a comprehensive description of food images that capture the correlations of feed-specific features. The feature sender uses a three-layer Spiking Neural Network (SNN). The proposed sender compresses features into sparse and discrete spike trains, significantly reducing the required transmission bandwidth and improving channel utilization and energy efficiency. Our model introduces the Compressed Factorized Bilinear block (CFB), which employs a low-rank feature approximation to reduce computational complexity and feature transmission volume while preserving the discriminate features. The enhancement reasoning module is proposed to enhance the received features by projecting them into a higher-dimensional space and utilizing the self-attention mechanism and sum pooling to compress them back to the original dimension. We conduct extensive experiments on the ETH Food-101 and Food2k datasets. Our results reveal that the MSPTN demonstrates state-of-the-art recognition performance, even with binary spike trains. Meanwhile, the MSPTN also exhibits remarkable robustness in wireless communication scenarios. With the combination of CFB, SNN, and EFB, our model achieves significant efficiency gains, including a nearly nine-fold decrease in feature transmission volume and a three-fold improvement in runtime & computational memory speed. Wenrui Li 0001, Jiahui Li 0001, Mengyao Ma, Xiaopeng Hong, Xiaopeng Fan 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Spatial-Temporal Saliency Guided Unbiased Contrastive Learning for Video Scene Graph GenerationabstractAccurately detecting objects and their interrelationships for Video Scene Graph Generation (VidSGG) confronts two primary challenges. The first involves the identification of active objects interacting with humans from the numerous background objects, while the second challenge is long-tailed distribution among predicate classes. To tackle these challenges, we propose STABILE, a novel framework with a spatial-temporal saliency-guided contrastive learning scheme. For the first challenge, STABILE features an active object retriever that includes an object saliency fusion block for enhancing object embeddings with motion cues alongside an object temporal encoder to capture temporal dependencies. For the second challenge, STABILE introduces an unbiased relationship representation learning module with an Unbiased Multi-Label (UML) contrastive loss to mitigate the effect of long-tailed distribution. With the enhancements in both aspects, STABILE substantially boosts the accuracy of scene graph generation. Extensive experiments demonstrate the superiority of STABILE, setting new benchmarks in the field by offering enhanced accuracy and unbiased scene graph generation. Weijun Zhuang, Bowen Dong 0001, Zhilin Zhu 0001, Zhijun Li 0002, Jie Liu 0001, Yaowei Wang 0001, Xiaopeng Hong, Xin Li 0034, Wangmeng Zuo |
IEEE Trans. Multim. | 7 |
| 2025 | Robust and Rotation-Equivariant Contrastive LearningabstractContrastive learning (CL) methods achieve great success by learning the invariant representation from various transformations. However, rotation transformations are considered harmful to CL and are rarely used, which results in failure when the objects show unseen orientations. This article proposes a representation focus shift network (RefosNet), which adds the rotation transformations to CL methods to improve the robustness of representation. First, the RefosNet constructs the rotation-equivariant mapping between the features of the original image and the rotated ones. Then, the RefosNet learns semantic-invariant representations (SIRs) based on explicitly decoupling the rotation-invariant features and the rotation-equivariant features. Moreover, an adaptive gradient passivation strategy is introduced to gradually shift the representation focus to invariant representations. This strategy can prevent catastrophic forgetting of the rotation equivariance, which is beneficial to the generalization of representations in both seen and unseen orientations. We adapt the baseline methods (i.e., "SimCLR" and "momentum contrast (MoCo) v2") to work with RefosNet to verify the performance. Extensive experimental results show that our method achieves significant improvements on the task of recognition. On ObjectNet-13 with unseen orientations, RefosNet gains 7.12% in terms of classification accuracy compared with SimCLR. On datasets in seen orientation, the performance improves by 5.5% on ImageNet-100, 7.29% on STL10, and 1.93% on CIFAR10. In addition, RefosNet has strong generalization on Place205, PASCAL VOC, and Caltech 101. Our method has also achieved satisfactory results in image retrieval tasks. Gairui Bai, Wei Xi 0003, Xiaopeng Hong, Songwen Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Multidimensional Measure Matching for Crowd CountingabstractThis article addresses the challenge of scale variations in crowd-counting problems from a multidimensional measure-theoretic perspective. We start by formulating crowd counting as a measure-matching problem, based on the assumption that discrete measures can express the scattered ground truth and the predicted density map. In this context, we introduce the Sinkhorn counting loss and extend it to the semi-balanced form, which alleviates the problems including entropic bias, distance destruction, and amount constraints. We then model the measure matching under the multidimensional space, in order to learn the counting from both location and scale. To achieve this, we extend the traditional 2-D coordinate support to 3-D, incorporating an additional axis to represent scale information, where a pyramid-based structure will be leveraged to learn the scale value for the predicted density. Extensive experiments on four challenging crowd-counting datasets, namely, ShanghaiTech A, UCF-QNRF, JHU++, and NWPU have validated the proposed method. Code is released at https://github.com/LoraLinH/Multidimensional-Measure-Matching-for-Crowd-Counting. Xiaopeng Hong, Zhiheng Ma, Yaowei Wang 0001, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Image Captioning via Dynamic Path CustomizationabstractThis article explores a novel dynamic network for vision and language (V&L) tasks, where the inferring structure is customized on the fly for different inputs. Most previous state-of-the-art (SOTA) approaches are static and handcrafted networks, which not only heavily rely on expert knowledge but also ignore the semantic diversity of input samples, therefore resulting in suboptimal performance. To address these issues, we propose a novel Dynamic Transformer Network (DTNet) for image captioning, which dynamically assigns customized paths to different samples, leading to discriminative yet accurate captions. Specifically, to build a rich routing space and improve routing efficiency, we introduce five types of basic cells and group them into two separate routing spaces according to their operating domains, i.e., spatial and channel. Then, we design a Spatial-Channel Joint Router (SCJR), which endows the model with the capability of path customization based on both spatial and channel information of the input sample. To validate the effectiveness of our proposed DTNet, we conduct extensive experiments on the MS-COCO dataset and achieve new SOTA performance on both the Karpathy split and the online test server. The source code is publicly available at https://github.com/xmu-xiaoma666/DTNet. Jiayi Ji, Xiaoshuai Sun, Yiyi Zhou, Xiaopeng Hong, Yongjian Wu 0001, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Gramformer: Learning Crowd Counting via Graph-Modulated TransformerabstractTransformer has been popular in recent crowd counting work since it breaks the limited receptive field of traditional CNNs. However, since crowd images always contain a large number of similar patches, the self-attention mechanism in Transformer tends to find a homogenized solution where the attention maps of almost all patches are identical. In this paper, we address this problem by proposing Gramformer: a graph-modulated transformer to enhance the network by adjusting the attention and input node features respectively on the basis of two different types of graphs. Firstly, an attention graph is proposed to diverse attention maps to attend to complementary information. The graph is building upon the dissimilarities between patches, modulating the attention in an anti-similarity fashion. Secondly, a feature-based centrality encoding is proposed to discover the centrality positions or importance of nodes. We encode them with a proposed centrality indices scheme to modulate the node features and similarity relationships. Extensive experiments on four challenging crowd counting datasets have validated the competitiveness of the proposed method. Code is available at https://github.com/LoraLinH/Gramformer. Zhiheng Ma, Xiaopeng Hong, Qinnan Shangguan, Deyu Meng |
AAAI | 3 |
| 2024 | Multi-modal Crowd Counting via Modal Emulation
Xiaopeng Hong, Zhiheng Ma, Yabin Wang 0001, Xiaopeng Fan 0001 |
BMVC | 2 |
| 2024 | Multi-modal Crowd Counting via a Broker Modality
Haoliang Meng, Xiaopeng Hong, Miao Shang, Wangmeng Zuo |
ECCV (74) | 2 |
| 2024 | GLAD: Towards Better Reconstruction with Global and Local Adaptive Diffusion Models for Unsupervised Anomaly Detection
Hang Yao 0001, Ming Liu 0018, Zhicun Yin, Zifei Yan, Xiaopeng Hong, Wangmeng Zuo |
ECCV (71) | 5 |
| 2024 | Reshaping the Online Data Buffering and Organizing Mechanism for Continual Test-Time Adaptation
Zhilin Zhu 0001, Xiaopeng Hong, Zhiheng Ma, Weijun Zhuang, Yaohui Ma, Yaowei Wang 0001 |
ECCV (82) | 2 |
| 2024 | MEGC2024: ACM Multimedia 2024 Facial Micro-Expression Grand ChallengeabstractFacial micro-expressions (MEs) are involuntary spontaneous movements of the face that typically appear in high-stakes situations where a person attempts to conceal a certain emotion from being known. A decade after the inception of the widely used CASME II and SMIC datasets, research in computational analysis of MEs has now advanced toward new pathways, exploring problems crucial to model generalization and real-world practicality. It is often challenging to design robust algorithms or models for spotting micro-expressions due to the high variability across diverse cultural backgrounds. Also, treating spotting and recognition as separate tasks is undesirable when handling long-spanning videos under realistic settings. This Grand Challenge comprises two distinct tracks: the Cross-Cultural Spotting (CCS) track, and the Spot-Then-Recognize (STR) track. All participating solutions submitted their results to a leaderboard, and several submissions performed well surpassing their respective baseline results. More details are available at: https://megc2024.github.io. John See, Jingting Li 0001, Adrian K. Davison, Gen-Bing Liong, Moi Hoon Yap, Wen-Huang Cheng, Xiaopeng Hong |
ACM Multimedia | 8 |
| 2024 | Boosting Semi-supervised Crowd Counting with Scale-based Active LearningabstractThe core of active semi-supervised crowd counting is the sample selection criteria. However, the scale factor has been neglected in active learning approaches despite the fact that the scale of heads varies drastically in the crowd images. In this paper, we propose a simple yet effective active labeling strategy to explicitly select informative unlabeled images, guided by the intra-scale uncertainty and inter-scale inconsistency metrics. The intra-scale uncertainty is quantified through the sum of the query-level entropy of images at different scales. Images are initially ranked based on this uncertainty for preselection. Inter-scale inconsistency is measured by the divergence between the query-level predictions of upscaled and downscaled images, allowing for the identification of the most informative images exhibiting the highest inconsistency. Additionally, we implement a progressive updating scheme for the semi-supervised crowd counting framework, in which the pseudo-labels for unlabeled images are refined iteratively. It further improves the counting accuracy. Through extensive experiments on widely used benchmarks, the proposed approach has demonstrated superior performance compared to previous state-of-the-art semi-supervised and active semi-supervised crowd counting methods. Shiwei Zhang 0004, Wei Ke 0003, Shuai Liu 0016, Xiaopeng Hong, Tong Zhang 0023 |
ACM Multimedia | 4 |
| 2024 | Label-Efficient Emotion and Sentiment AnalysisabstractEmotion and sentiment analysis (ESA) assists machines to serve humans more intelligently. However, collecting large-scale high-quality datasets for training ESA models in a supervised manner is expensive, time-consuming, and difficult in practice. This tutorial focuses on the label-efficient ESA (LeESA) learning methods. Specifically, we first introduce the stimuli and characteristics of emotion and then illustrate seven typical training paradigms, followed by applications and future directions of LeESA. Sicheng Zhao, Guoli Jia, Xiaopeng Hong, Jianhua Tao 0001 |
ACM Multimedia | 3 |
| 2024 | Token-based deep reinforcement learning for Heterogeneous VRP with Service Time Constraints
Xiaopeng Hong, Yabin Wang 0001, Junzhou Zhao, Guanghui Sun, Baoxing Qin |
Knowl. Based Syst. | 2 |
| 2024 | Semi-Supervised Crowd Counting With Contextual Modeling: Facilitating Holistic Understanding of Crowd ScenesabstractTo alleviate the heavy annotation burden for training a reliable crowd counting model and thus make the model more practicable and accurate by being able to benefit from more data, this paper presents a new semi-supervised method based on the mean teacher framework. When there is a scarcity of labeled data available, the model is prone to overfit local patches. Within such contexts, the conventional approach of solely improving the accuracy of local patch predictions through unlabeled data proves inadequate. Consequently, we propose a more nuanced approach: fostering the model’s intrinsic ‘subitizing’ capability. This ability allows the model to accurately estimate the count in regions by leveraging its understanding of the crowd scenes, mirroring the human cognitive process. To achieve this goal, we apply masking on unlabeled data, guiding the model to make predictions for these masked patches based on the holistic cues. Furthermore, to help with feature learning, herein we incorporate a fine-grained density classification task. Our method is general and applicable to most existing crowd counting methods as it doesn’t have strict structural or loss constraints. In addition, we observe that the model trained with our framework shows strong contextual modeling capabilities, which allows it to make robust predictions even when some local details of patches are lost. Our method achieves the state-of-the-art performance, surpassing previous approaches by a large margin on challenging benchmarks such as ShanghaiTech A and UCF-QNRF. The code is available at: https://github.com/cha15yq/MRC-Crowd. Yifei Qian, Xiaopeng Hong, Zhongliang Guo 0001, Ognjen Arandjelovic, Carl Donovan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Exploiting Multi-Scale Parallel Self-Attention and Local Variation via Dual-Branch Transformer-CNN Structure for Face Super-ResolutionabstractRecently, deep learning technique has been widely employed to deal with face super-resolution (FSR) problem. It aims to predict the nonlinear relationship between the low-resolution (LR) face images and corresponding high-resolution (HR) ones, which could recover the high-frequency details from the LR degraded textures. However, either CNN-based or Transformer-based approaches mostly enhance the details by exploiting the relationship of local pixels or patches on LR features, the nonlocal features are not fully taken into account for producing high-frequency textures. To improve the above problem, we design a novel dual-branch module which consists of Transformer and CNN respectively. The Transformer branch extracts multiple scale feature embeddings and explores local and nonlocal self-attention simultaneously. Thus, the parallel self-attention mechanism has superior capabilities to capture the local and nonlocal dependencies on face image in the face reconstruction. Furthermore, the traditional CNNs usually extract features by combining pixels in a local convolutional kernel, it may be not effective to recover lost high-frequency details since the variations of local pixels are not well measured, which is important in recovering vivid edges and contours. To this end, we propose the local variation based attention block on the CNN branch, which could enhance the capabilities by directly extracting features from the variation of neighboring pixels. Finally, the Transformer-branch and CNN-branch are combined together by the modulation block to fuse both nonlocal and local advantages from two branches. Experimental results demonstrate the effectiveness of the proposed method when compared with state-of-the-art approaches. Jingang Shi, Yusi Wang, Zitong Yu, Guanxin Li, Xiaopeng Hong, Fei Wang 0037, Yihong Gong |
IEEE Trans. Multim. | 5 |
| 2024 | Brain Cognition-Inspired Dual-Pathway CNN Architecture for Image ClassificationabstractInspired by the global-local information processing mechanism in the human visual system, we propose a novel convolutional neural network (CNN) architecture named cognition-inspired network (CogNet) that consists of a global pathway, a local pathway, and a top-down modulator. We first use a common CNN block to form the local pathway that aims to extract fine local features of the input image. Then, we use a transformer encoder to form the global pathway to capture global structural and contextual information among local parts in the input image. Finally, we construct the learnable top-down modulator where fine local features of the local pathway are modulated by global representations of the global pathway. For ease of use, we encapsulate the dual-pathway computation and modulation process into a building block, called the global-local block (GL block), and a CogNet of any depth can be constructed by stacking a necessary number of GL blocks one after another. Extensive experimental evaluations have revealed that the proposed CogNets have achieved the state-of-the-art performance accuracies on all the six benchmark datasets and are very effective for overcoming the "texture bias" and the "semantic confusion" problems faced by many CNN models. Songlin Dong, Yihong Gong, Jingang Shi, Miao Shang, Xing Wei 0001, Xiaopeng Hong, Tiangang Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Deep Class-Incremental Learning From Decentralized DataabstractIn this article, we focus on a new and challenging decentralized machine learning paradigm in which there are continuous inflows of data to be addressed and the data are stored in multiple repositories. We initiate the study of data-decentralized class-incremental learning (DCIL) by making the following contributions. First, we formulate the DCIL problem and develop the experimental protocol. Second, we introduce a paradigm to create a basic decentralized counterpart of typical (centralized) CIL approaches, and as a result, establish a benchmark for the DCIL study. Third, we further propose a decentralized composite knowledge incremental distillation (DCID) framework to transfer knowledge from historical models and multiple local sites to the general model continually. DCID consists of three main components, namely, local CIL, collaborated knowledge distillation (KD) among local models, and aggregated KD from local models to the general one. We comprehensively investigate our DCID framework by using a different implementation of the three components. Extensive experimental results demonstrate the effectiveness of our DCID framework. The source code of the baseline methods and the proposed DCIL is available at https://github.com/Vision-Intelligence-and-Robots-Group/DCIL. Songlin Dong, Jinjie Chen, Qi Tian 0001, Yihong Gong, Xiaopeng Hong |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without InterferenceabstractThis paper focuses on the prevalent stage interference and stage performance imbalance of incremental learning. To avoid obvious stage learning bottlenecks, we propose a new incremental learning framework, which leverages a series of stage-isolated classifiers to perform the learning task at each stage, without interference from others. To be concrete, to aggregate multiple stage classifiers as a uniform one impartially, we first introduce a temperature-controlled energy metric for indicating the confidence score levels of the stage classifiers. We then propose an anchor-based energy self-normalization strategy to ensure the stage classifiers work at the same energy level. Finally, we design a voting-based inference augmentation strategy for robust inference. The proposed method is rehearsal-free and can work for almost all incremental learning scenarios. We evaluate the proposed method on four large datasets. Extensive results demonstrate the superiority of the proposed method in setting up new state-of-the-art overall performance. Code is available at https://github.com/iamwangyabin/ESN. Yabin Wang 0001, Zhiheng Ma, Zhiwu Huang, Yaowei Wang 0001, Zhou Su 0001, Xiaopeng Hong |
AAAI | 6 |
| 2023 | One-Shot Replay: Boosting Incremental Object Detection via Retrospecting One ObjectabstractModern object detectors are ill-equipped to incrementally learn new emerging object classes over time due to the well-known phenomenon of catastrophic forgetting. Due to data privacy or limited storage, few or no images of the old data can be stored for replay. In this paper, we design a novel One-Shot Replay (OSR) method for incremental object detection, which is an augmentation-based method. Rather than storing original images, only one object-level sample for each old class is stored to reduce memory usage significantly, and we find that copy-paste is a harmonious way to replay for incremental object detection. In the incremental learning procedure, diverse augmented samples with co-occurrence of old and new objects to existing training data are generated. To introduce more variants for objects of old classes, we propose two augmentation modules. The object augmentation module aims to enhance the ability of the detector to perceive potential unknown objects. The feature augmentation module explores the relations between old and new classes and augments the feature space via analogy. Extensive experimental results on VOC2007 and COCO demonstrate that OSR can outperform the state-of-the-art incremental object detection methods without using extra wild data. Dongbao Yang, Yu Zhou 0015, Xiaopeng Hong, Aoting Zhang, Weiping Wang 0005 |
AAAI | 3 |
| 2023 | MEGC2023: ACM Multimedia 2023 ME Grand ChallengeabstractFacial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. Unfortunately, the small sample problem severely limits the automation of ME analysis. Furthermore, due to the weak and transient nature of MEs, it is difficult for models to distinguish it from other types of facial actions. Therefore, ME in long videos is a challenging task, and the current performance cannot meet the practical application requirements. Addressing these issues, this challenge focuses on ME and the macro-expression (MaE) spotting task. This year, in order to evaluate algorithms' performance more fairly, based on CAS(ME)2, SAMM Long Videos, SMIC-E-long, CAS(ME)3 and 4DME, we build an unseen cross-cultural long-video test set. All participating algorithms are required to run on this test set and submit their results on a leaderboard with a baseline result. Adrian K. Davison, Jingting Li 0001, Moi Hoon Yap, John See, Wen-Huang Cheng, Xiaopeng Hong |
ACM Multimedia | 7 |
| 2023 | FME '23: 3rd Facial Micro-Expression WorkshopabstractMicro-expressions are facial movements that are extremely short and not easily detected, which often reflect the genuine emotions of individuals. Micro-expressions are important cues for understanding real human emotions and can be used for non-contact, non-perceptual deception detection, or abnormal emotion recognition. It has broad application prospects in national security, judicial practice, health prevention, and clinical practice. However, micro-expression feature extraction and learning are highly challenging because they are typically short in duration, low intensity, and have local facial asymmetry. In addition, the intelligent micro-expression analysis combined with deep learning technology is also plagued by the problem of relatively small data samples. Not only is micro-expression elicitation very difficult, micro-expression annotation is also very time-consuming and laborious. More importantly, the micro-expression generation mechanism is not yet clear, which shackles the application of micro-expressions in real scenarios. FME'23 is the inaugural workshop in this area of research, with the aim of promoting interactions between researchers and scholars from within this niche area of research. This year we hope to discuss the growing ethical conversations when using face data, and how we can come to a consensus on micro-expression standards within affective computing. Adrian K. Davison, Jingting Li 0001, Moi Hoon Yap, John See, Wen-Huang Cheng, Xiaopeng Hong |
ACM Multimedia | 7 |
| 2023 | Pseudo Object Replay and Mining for Incremental Object DetectionabstractIncremental object detection (IOD) aims to mitigate catastrophic forgetting for object detectors when incrementally learning to detect new emerging object classes without using original training data. Most existing IOD methods benefit from the assumption that unlabeled old-class objects may co-occur with labeled new-class objects in the new training data. However, in practical scenarios, old-class objects may be absent, which is called non co-occurrence IOD. In this paper, we propose a pseudo object replay and mining method (PseudoRM) to handle the co-occurrence dependent problem, reducing the performance degradation caused by the absence of old-class objects. The new training data can be augmented by co-occurring fake (old-class) and real (new-class) objects with a patch-level data-free generation method in the pseudo object replay stage. To fully use existing training data, we propose pseudo object mining to explore false positives for transferring useful instance-level knowledge. In the incremental learning procedure, a generative distillation is introduced to distill image-level knowledge for balancing stability and plasticity. Experimental results on PASCAL VOC and COCO demonstrate that PseudoRM can effectively boost the performance on both co-occurrence and non co-occurrence scenarios without using old samples or extra wild data. Dongbao Yang, Yu Zhou 0015, Xiaopeng Hong, Aoting Zhang, Linchengxi Zeng, Weiping Wang 0005 |
ACM Multimedia | 3 |
| 2023 | A Continual Deepfake Detection Benchmark: Dataset, Methods, and EssentialsabstractThere have been emerging a number of benchmarks and techniques for the detection of deepfakes. However, very few works study the detection of incrementally appearing deepfakes in the real-world scenarios. To simulate the wild scenes, this paper suggests a continual deepfake detection benchmark (CDDB) over a new collection of deepfakes from both known and unknown generative models. The suggested CDDB designs multiple evaluations on the detection over easy, hard, and long sequence of deepfake tasks, with a set of appropriate measures. In addition, we exploit multiple approaches to adapt multiclass incremental learning methods, commonly used in the continual visual recognition, to the continual deepfake detection problem. We evaluate existing methods, including their adapted ones, on the proposed CDDB. Within the proposed benchmark, we explore some commonly known essentials of standard continual learning. Our study provides new insights on these essentials in the context of continual deepfake detection. The suggested CDDB is clearly more challenging than the existing benchmarks, which thus offers a suitable evaluation avenue to the future research. Both data and code are available at https://github.com/Coral79/CDDB. Chuqiao Li, Zhiwu Huang, Danda Pani Paudel, Yabin Wang 0001, Mohamad Shahbazi, Xiaopeng Hong, Luc Van Gool |
WACV | 6 |
| 2023 | Toward Label-Efficient Emotion and Sentiment AnalysisabstractEmotion and sentiment play a central role in various human activities, such as perception, decision-making, social interaction, and logical reasoning. Developing artificial emotional intelligence (AEI) for machines is becoming a bottleneck in human–computer interaction. The first step of AEI is to recognize the emotion and sentiment that are conveyed in different affective signals. Traditional supervised emotion and sentiment analysis (ESA) methods, especially deep learning-based ones, usually require large-scale labeled training data. However, due to the essential subjectivity, complexity, uncertainty and ambiguity, and subtlety, collecting such annotations is expensive, time-consuming, and difficult in practice. In this article, we introduce label-efficient ESA from the computational perspective. First, we present a hierarchical taxonomy for label-efficient learning based on the availability of sample labels, emotion categories, and data domains during training. Second, for each of the seven paradigms, i.e., unsupervised, semisupervised, weakly supervised, low-shot, incremental, domain-adaptive, and domain-generalizable ESA, we give the definition, summarize existing methods, and present our views on the quantitative and qualitative comparison. Finally, we provide several promising real-world applications, followed by unsolved challenges and potential future directions. Sicheng Zhao, Xiaopeng Hong, Jufeng Yang, Guiguang Ding |
Proc. IEEE | 2 |
| 2023 | Editorial for pattern recognition letters special issue on face-based emotion understanding
Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong |
Pattern Recognit. Lett. | 5 |
| 2023 | Semi-Supervised Crowd Counting via Multiple Representation LearningabstractThere has been a growing interest in counting crowds through computer vision and machine learning techniques in recent years. Despite that significant progress has been made, most existing methods heavily rely on fully-supervised learning and require a lot of labeled data. To alleviate the reliance, we focus on the semi-supervised learning paradigm. Usually, crowd counting is converted to a density estimation problem. The model is trained to predict a density map and obtains the total count by accumulating densities over all the locations. In particular, we find that there could be multiple density map representations for a given image in a way that they differ in probability distribution forms but reach a consensus on their total counts. Therefore, we propose multiple representation learning to train several models. Each model focuses on a specific density representation and utilizes the count consistency between models to supervise unlabeled data. To bypass the explicit density regression problem, which makes a strong parametric assumption on the underlying density distribution, we propose an implicit density representation method based on the kernel mean embedding. Extensive experiments demonstrate that our approach outperforms state-of-the-art semi-supervised methods significantly. Xing Wei 0001, Yunfeng Qiu, Zhiheng Ma, Xiaopeng Hong, Yihong Gong |
IEEE Trans. Image Process. | 4 |
| 2023 | Model Behavior Preserving for Class-Incremental LearningabstractDeep models have shown to be vulnerable to catastrophic forgetting, a phenomenon that the recognition performance on old data degrades when a pre-trained model is fine-tuned on new data. Knowledge distillation (KD) is a popular incremental approach to alleviate catastrophic forgetting. However, it usually fixes the absolute values of neural responses for isolated historical instances, without considering the intrinsic structure of the responses by a convolutional neural network (CNN) model. To overcome this limitation, we recognize the importance of the global property of the whole instance set and treat it as a behavior characteristic of a CNN model relevant to model incremental learning. On this basis: 1) we design an instance neighborhood-preserving (INP) loss to maintain the order of pair-wise instance similarities of the old model in the feature space; 2) we devise a label priority-preserving (LPP) loss to preserve the label ranking lists within instance-wise label probability vectors in the output space; and 3) we introduce an efficient derivable ranking algorithm for calculating the two loss functions. Extensive experiments conducted on CIFAR100 and ImageNet show that our approach achieves the state-of-the-art performance. Xiaopeng Hong, Songlin Dong, Jingang Shi, Yihong Gong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Segmentation Assisted U-shaped Multi-scale Transformer for Crowd Counting
Yifei Qian, Liangfei Zhang, Xiaopeng Hong, Carl Donovan, Ognjen Arandjelovic |
BMVC | 3 |
| 2022 | Boosting Crowd Counting via Multifaceted AttentionabstractThis paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size attention of recent vision transformers can well handle this kind of variations. To address this problem, we propose a Multifaceted Attention Network (MAN) to improve transformer models in local spatial relation encoding. MAN incorporates global attention from vanilla transformer, learnable local attention, and instance attention into a counting model. Firstly, the local Learnable Region Attention (LRA) is proposed to assign attention exclusive for each feature location dynamically. Secondly, we design the Local Attention Regularization to supervise the training of LRA by minimizing the deviation among the attention for different feature locations. Finally, we provide an Instance Attention mechanism to focus on the most important instances dynamically during training. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++, and NWPU have validated the proposed method. Code: https://github.com/LoraLinH/Boosting-Crowd-Counting-via-Multifaceted-Attention. Zhiheng Ma, Rongrong Ji, Yaowei Wang 0001, Xiaopeng Hong |
CVPR | 5 |
| 2022 | IDPT: Interconnected Dual Pyramid Transformer for Face Super-ResolutionabstractFace Super-resolution (FSR) task works for generating high-resolution (HR) face images from the corresponding low-resolution (LR) inputs, which has received a lot of attentions because of the wide application prospects. However, due to the diversity of facial texture and the difficulty of reconstructing detailed content from degraded images, FSR technology is still far away from being solved. In this paper, we propose a novel and effective face super-resolution framework based on Transformer, namely Interconnected Dual Pyramid Transformer (IDPT). Instead of straightly stacking cascaded feature reconstruction blocks, the proposed IDPT designs the pyramid encoder/decoder Transformer architecture to extract coarse and detailed facial textures respectively, while the relationship between the dual pyramid Transformers is further explored by a bottom pyramid feature extractor. The pyramid encoder/decoder structure is devised to adapt various characteristics of textures in different spatial spaces hierarchically. A novel fusing modulation module is inserted in each spatial layer to guide the refinement of detailed texture by the corresponding coarse texture, while fusing the shallow-layer coarse feature and corresponding deep-layer detailed feature simultaneously. Extensive experiments and visualizations on various datasets demonstrate the superiority of the proposed method for face super-resolution tasks. Jingang Shi, Yusi Wang, Songlin Dong, Xiaopeng Hong, Zitong Yu, Fei Wang 0037, Changxin Wang, Yihong Gong |
IJCAI | 4 |
| 2022 | FME '22: 2nd Workshop on Facial Micro-Expression: Advanced Techniques for Multi-Modal Facial Expression AnalysisabstractMicro-expressions are facial movements that are extremely short and not easily detected, which often reflect the genuine emotions of individuals. Micro-expressions are important cues for understanding real human emotions and can be used for non-contact non-perceptual deception detection, or abnormal emotion recognition. It has broad application prospects in national security, judicial practice, health prevention, clinical practice, etc. However, micro-expression feature extraction and learning are highly challenging because micro-expressions have the characteristics of short duration, low intensity, and local asymmetry. In addition, the intelligent micro-expression analysis combined with deep learning technology is also plagued by the problem of small samples. Not only is micro-expression elicitation very difficult, micro-expression annotation is also very time-consuming and laborious. More importantly, the micro-expression generation mechanism is not yet clear, which shackles the application of micro-expressions in real scenarios. FME'22 is the inaugural workshop in this area of research, with the aim of promoting interactions between researchers and scholars from within this niche area of research and also including those from broader, general areas of expression and psychology research. The complete FME'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552465. Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong |
ACM Multimedia | 5 |
| 2022 | MEGC2022: ACM Multimedia 2022 Micro-Expression Grand ChallengeabstractFacial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. Unfortunately, the small sample problem severely limits the automation of ME analysis. Furthermore, due to the brief and subtle nature of ME, ME spotting is a challenging task, and the performance is still not satisfactory yet. This challenge focuses on two tasks, i.e., the micro- and macro-expression spotting task, and the ME Generation task. Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong, Adrian K. Davison, Yante Li, Zizhao Dong |
ACM Multimedia | 5 |
| 2022 | Semi-supervised Crowd Counting via Density AgencyabstractIn this paper, we propose a new agency-guided semi-supervised counting approach. First, we build a learnable auxiliary structure, namely the density agency to bring the recognized foreground regional features close to corresponding density sub-classes (agents) and push away background ones. Second, we propose a density-guided contrastive learning loss to consolidate the backbone feature extractor. Third, we build a regression head by using a transformer structure to refine the foreground features further. Finally, an efficient noise depression loss is provided to minimize the negative influence of annotation noises. Extensive experiments on four challenging crowd counting datasets demonstrate that our method achieves superior performance to the state-of-the-art semi-supervised counting methods by a large margin. The code is available at https://github.com/LoraLinH/Semi-supervised-Crowd-Counting-via-Density-Agency. Zhiheng Ma, Xiaopeng Hong, Yaowei Wang 0001, Zhou Su 0001 |
ACM Multimedia | 3 |
| 2022 | S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningabstractState-of-the-art deep neural networks are still struggling to address the catastrophic forgetting problem in continual learning. In this paper, we propose one simple paradigm (named as S-Prompting) and two concrete approaches to highly reduce the forgetting degree in one of the most typical continual learning scenarios, i.e., domain increment learning (DIL). The key idea of the paradigm is to learn prompts independently across domains with pre-trained transformers, avoiding the use of exemplars that commonly appear in conventional methods. This results in a win-win game where the prompting can achieve the best for each domain. The independent prompting across domains only requests one single cross-entropy loss for training and one simple K-NN operation as a domain identifier for inference. The learning paradigm derives an image prompt learning approach and a novel language-image prompt learning approach. Owning an excellent scalability (0.03% parameter increase per domain), the best of our approaches achieves a remarkable relative improvement (an average of about 30%) over the best of the state-of-the-art exemplar-free methods for three standard DIL tasks, and even surpasses the best of them relatively by about 6% in average when they use exemplars. Source code is available at https://github.com/iamwangyabin/S-Prompts. Yabin Wang 0001, Zhiwu Huang, Xiaopeng Hong |
NeurIPS | 3 |
| 2022 | Short and Long Range Relation Based Spatio-Temporal Transformer for Micro-Expression RecognitionabstractBeing spontaneous, micro-expressions are useful in the inference of a person's true emotions even if an attempt is made to conceal them. Due to their short duration and low intensity, the recognition of micro-expressions is a difficult task in affective computing. The early work based on handcrafted spatio-temporal features which showed some promise, has recently been superseded by different deep learning approaches which now compete for the state of the art performance. Nevertheless, the problem of capturing both local and global spatio-temporal patterns remains challenging. To this end, herein we propose a novel spatio-temporal transformer architecture – to the best of our knowledge, the first purely transformer based approach (i.e., void of any convolutional network use) for micro-expression recognition. The architecture comprises a spatial encoder which learns spatial patterns, a temporal aggregator for temporal dimension analysis, and a classification head. A comprehensive evaluation on three widely used spontaneous micro-expression data sets, namely SMIC-HS, CASME II and SAMM, shows that the proposed approach consistently outperforms the state of the art, and is the first framework in the published literature on micro-expression recognition to achieve the unweighted F1-score greater than 0.9 on any of the aforementioned data sets. The source code is available athttps://github.com/Vision-Intelligence-and-Robots-Group/SLSTT. Liangfei Zhang, Xiaopeng Hong, Ognjen Arandjelovic, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Disentangling Task-Oriented Representations for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to address the domain-shift problem between a labeled source domain and an unlabeled target domain. Many efforts have been made to eliminate the mismatch between the distributions of training and testing data by learning domain-invariant representations. However, the learned representations are usually not task-oriented, i.e., being class-discriminative and domain-transferable simultaneously. This drawback limits the flexibility of UDA in complicated open-set tasks where no labels are shared between domains. In this paper, we break the concept of task-orientation into task-relevance and task-irrelevance, and propose a dynamic task-oriented disentangling network (DTDN) to learn disentangled representations in an end-to-end fashion for UDA. The dynamic disentangling network effectively disentangles data representations into two components: the task-relevant ones embedding critical information associated with the task across domains, and the task-irrelevant ones with the remaining non-transferable or disturbing information. These two components are regularized by a group of task-specific objective functions across domains. Such regularization explicitly encourages disentangling and avoids the use of generative models or decoders. Experiments in complicated, open-set scenarios (retrieval tasks) and empirical benchmarks (classification tasks) demonstrate that the proposed method captures rich disentangled information and achieves superior performance. Pingyang Dai, Peixian Chen, Qiong Wu 0012, Xiaopeng Hong, Qixiang Ye, Qi Tian 0001, Chia-Wen Lin, Rongrong Ji |
IEEE Trans. Image Process. | 4 |
| 2022 | Identity-Quantity Harmonic Multi-Object TrackingabstractThe data association problem of multi-object tracking (MOT) aims to assign IDentity (ID) labels to detections and infer a complete trajectory for each target. Most existing methods assume that each detection corresponds to a unique target and thus cannot handle situations when multiple targets occur in a single detection due to detection failure in crowded scenes. To relax this strong assumption for practical applications, we formulate the MOT as a Maximizing An Identity-Quantity Posterior (MAIQP) problem on the basis of associating each detection with an identity and a quantity characteristic and then provide solutions to tackle two key problems arising. Firstly, a local target quantification module is introduced to count the number of targets within one detection. Secondly, we propose an identity-quantity harmony mechanism to reconcile the two characteristics. On this basis, we develop a novel Identity-Quantity HArmonic Tracking (IQHAT) framework that allows assigning multiple ID labels to detections containing several targets. Through extensive experimental evaluations on five benchmark datasets, we demonstrate the superiority of the proposed method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
IEEE Trans. Image Process. | 3 |
| 2022 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problem in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from three aspects. First, we establish a CDMER experimental evaluation protocol aiming to allow the researchers to conveniently work on this topic and evaluate their proposed methods under the same standard. Second, we conduct benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating CDMER problem from two different perspectives. Third, we propose a novel DA method called region selective transfer regression (RSTR) to deal with the CDMER task. The overall superior performance of RSTR over the state-of-the-art DA methods demonstrates that taking into consideration the facial local region information used in RSTR contributes to developing effective DA methods for dealing with CDMER problem. Tong Zhang 0015, Yuan Zong, Wenming Zheng, C. L. Philip Chen, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | ECCNAS: Efficient Crowd Counting Neural Architecture SearchabstractRecent solutions to crowd counting problems have already achieved promising performance across various benchmarks. However, applying these approaches to real-world applications is still challenging, because they are computation intensive and lack the flexibility to meet various resource budgets. In this article, we propose an efficient crowd counting neural architecture search (ECCNAS) framework to search efficient crowd counting network structures, which can fill this research gap. A novel search from pre-trained strategy enables our cross-task NAS to explore the significantly large and flexible search space with less search time and get more proper network structures. Moreover, our well-designed search space can intrinsically provide candidate neural network structures with high performance and efficiency. In order to search network structures according to hardwares with different computational performance, we develop a novel latency cost estimation algorithm in our ECCNAS. Experiments show our searched models get an excellent trade-off between computational complexity and accuracy and have the potential to deploy in practical scenarios with various resource budgets. We reduce the computational cost, in terms of multiply-and-accumulate (MACs), by up to 96% with comparable accuracy. And we further designed experiments to validate the efficiency and the stability improvement of our proposed search from pre-trained strategy. Yabin Wang 0001, Zhiheng Ma, Xing Wei 0001, Yaowei Wang 0001, Xiaopeng Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Few-Shot Class-Incremental Learning via Relation Knowledge DistillationabstractIn this paper, we focus on the challenging few-shot class incremental learning (FSCIL) problem, which requires to transfer knowledge from old tasks to new ones and solves catastrophic forgetting. We propose the exemplar relation distillation incremental learning framework to balance the tasks of old-knowledge preserving and new-knowledge adaptation. First, we construct an exemplar relation graph to represent the knowledge learned by the original network and update gradually for new tasks learning. Then an exemplar relation loss function for discovering the relation knowledge between different classes is introduced to learn and transfer the structural information in relation graph. A large number of experiments demonstrate that relation knowledge does exist in the exemplars and our approach outperforms other state-of-the-art class-incremental learning methods on the CIFAR100, miniImageNet, and CUB200 datasets. Songlin Dong, Xiaopeng Hong, Xinyuan Chang, Xing Wei 0001, Yihong Gong |
AAAI | 2 |
| 2021 | Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingabstractThis paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples are available in the target domain. The key issue is how to utilize unlabelled videos in the target domain for knowledge learning and transferring from the source domain. To tackle this problem, we propose a novel Error-aware Density Isomorphism REConstruction Network (EDIREC-Net) for cross-domain crowd counting. EDIREC-Net jointly transfers a pre-trained counting model to target domains using a density isomorphism reconstruction objective and models the reconstruction erroneousness by error reasoning. Specifically, as crowd flows in videos are consecutive, the density maps in adjacent frames turn out to be isomorphic. On this basis, we regard the density isomorphism reconstruction error as a self-supervised signal to transfer the pre-trained counting models to different target domains. Moreover, we leverage an estimation-reconstruction consistency to monitor the density reconstruction erroneousness and suppress unreliable density reconstructions during training. Experimental results on four benchmark datasets demonstrate the superiority of the proposed method and ablation studies investigate the efficiency and robustness. The source code is available at https://github.com/GehenHe/EDIREC-Net. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
AAAI | 4 |
| 2021 | Learning to Count via Unbalanced Optimal TransportabstractCounting dense crowds through computer vision technology has attracted widespread attention. Most crowd counting datasets use point annotations. In this paper, we formulate crowd counting as a measure regression problem to minimize the distance between two measures with different supports and unequal total mass. Specifically, we adopt the unbalanced optimal transport distance, which remains stable under spatial perturbations, to quantify the discrepancy between predicted density maps and point annotations. An efficient optimization algorithm based on the regularized semi-dual formulation of UOT is introduced, which alternatively learns the optimal transportation and optimizes the density regressor. The quantitative and qualitative results illustrate that our method achieves state-of-the-art counting and localization performance. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yunfeng Qiu, Yihong Gong |
AAAI | 3 |
| 2021 | Image-to-Image Translation via Hierarchical Style DisentanglementabstractRecently, image-to-image translation has made significant progress in achieving both multi-label (i.e., translation conditioned on different labels) and multi-style (i.e., generation with diverse styles) tasks. However, due to the unexplored independence and exclusiveness in the labels, existing endeavors are defeated by involving uncontrolled manipulations to the translation results. In this paper, we propose Hierarchical Style Disentanglement (HiSD) to address this issue. Specifically, we organize the labels into a hierarchical tree structure, in which independent tags, exclusive attributes, and disentangled styles are allocated from top to bottom. Correspondingly, a new translation process is designed to adapt the above structure, in which the styles are identified for controllable translations. Both qualitative and quantitative results on the CelebA-HQ dataset verify the ability of the proposed HiSD. The code has been released at https://github.com/imlixinyang/HiSD. Shengchuan Zhang, Jie Hu 0018, Liujuan Cao, Xiaopeng Hong, Xudong Mao, Feiyue Huang, Yongjian Wu 0001, Rongrong Ji |
CVPR | 5 |
| 2021 | Aha! Adaptive History-driven Attack for Decision-based Black-box ModelsabstractThe decision-based black-box attack means to craft adversarial examples with only the top-1 label of the victim model available. A common practice is to start from a large perturbation and then iteratively reduce it with a deterministic direction and a random one while keeping it adversarial. The limited information obtained from each query and inefficient direction sampling impede attack efficiency, making it hard to obtain a small enough perturbation within a limited number of queries. To tackle this problem, we propose a novel attack method termed Adaptive History-driven Attack (AHA) which gathers information from all historical queries as the prior for current sampling. Moreover, to balance between the deterministic direction and the random one, we dynamically adjust the coefficient according to the ratio of the actual magnitude reduction to the expected one. Such a strategy improves the success rate of queries during optimization, letting adversarial examples move swiftly along the decision boundary. Our method can also integrate with subspace optimization like dimension reduction to further improve efficiency. Extensive experiments on both ImageNet and CelebA datasets demonstrate that our method achieves at least 24.3% lower magnitude of perturbation on average with the same number of queries. Finally, we prove the practical potential of our method by evaluating it on popular defense methods and a real-world system provided by MEGVII Face++. Jie Li 0052, Rongrong Ji, Peixian Chen, Baochang Zhang 0001, Xiaopeng Hong, Shaoxin Li 0001, Feiyue Huang, Yongjian Wu 0001 |
ICCV | 5 |
| 2021 | Towards A Universal Model for Cross-Dataset Crowd CountingabstractThis paper proposes to handle the practical problem of learning a universal model for crowd counting across scenes and datasets. We dissect that the crux of this problem is the catastrophic sensitivity of crowd counters to scale shift, which is very common in the real world and caused by factors such as different scene layouts and image resolutions. Therefore it is difficult to train a universal model that can be applied to various scenes. To address this problem, we propose scale alignment as a prime module for establishing a novel crowd counting framework. We derive a closed-form solution to get the optimal image rescaling factors for alignment by minimizing the distances between their scale distributions. A novel neural network together with a loss function based on an efficient sliced Wasserstein distance is also proposed for scale distribution estimation. Benefiting from the proposed method, we have learned a universal model that generally works well on several datasets where can even outperform state-of-the-art models that are particularly fine-tuned for each dataset significantly. Experiments also demonstrate the much better generalizability of our model to unseen scenes. Zhiheng Ma, Xiaopeng Hong, Xing Wei 0001, Yunfeng Qiu, Yihong Gong |
ICCV | 2 |
| 2021 | Anomaly Detection Via Self-Organizing MapabstractAnomaly detection plays a key role in industrial manufacturing for product quality control. Traditional methods for anomaly detection are rule-based with limited generalization ability. Recent methods based on supervised deep learning are more powerful but require large-scale annotated datasets for training. In practice, abnormal products are rare thus it is very difficult to train a deep model in a fully supervised way. In this paper, we propose a novel unsupervised anomaly detection approach based on Self-organizing Map (SOM). Our method, Self-organizing Map for Anomaly Detection (SOMAD) maintains normal characteristics by using topological memory based on multi-scale features. SOMAD achieves state-of-the-art performance on unsupervised anomaly detection and localization on the MVTec dataset. Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 5 |
| 2021 | Class Incremental Learning for Video Action ClassificationabstractClass Incremental Learning (CIL) is a hot topic in machine learning for CNN models to learn new classes incrementally. However, most of the CIL studies are for image classification and object recognition tasks and few CIL studies are available for video action classification. To mitigate this problem, in this paper, we present a new Grow When Required network (GWR) based video CIL framework for action classification. GWR learns knowledge incrementally by modeling the manifold of video frames for each encountered action class in feature space. We also introduce a Knowledge Consolidation (KC) method to separate the feature manifolds of old class and new class and introduce an associative matrix for label prediction. Experimental results on KTH and Weizmann demonstrate the effectiveness of the framework. Jiawei Ma, Jianxing Ma, Xiaopeng Hong, Yihong Gong |
ICIP | 4 |
| 2021 | Direct Measure Matching for Crowd CountingabstractTraditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussian kernel sizes. In this paper, we propose a new measure-based counting approach to regress the predicted density maps to the scattered point-annotated ground truth directly. First, crowd counting is formulated as a measure matching problem. Second, we derive a semi-balanced form of Sinkhorn divergence, based on which a Sinkhorn counting loss is designed for measure matching. Third, we propose a self-supervised mechanism by devising a Sinkhorn scale consistency loss to resist scale changes. Finally, an efficient optimization method is provided to minimize the overall loss function. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++ and NWPU have validated the proposed method. Xiaopeng Hong, Zhiheng Ma, Xing Wei 0001, Yunfeng Qiu, Yaowei Wang 0001, Yihong Gong |
IJCAI | 2 |
| 2021 | Kohonen Self-Organizing Map based Route Planning: A RevisitabstractIn this paper, we revisit the long-standing Traveling Salesman Problem (TSP) and focus on the challenging, yet practical route planning problem with limited computational resources. We make contributions to TSP, one of the most famous NP-hard problems by providing a new improved approximate solution, which we term TOpology Preserving Self-Organizing Map (TOPSOM). TOPSOM well preserves the topology of the node map to be traversed by maintaining the continuity of nodes and the distances between them. In addition, to satisfy the requirements of convex hull, we design an elastic competitive Hebbian learning rule. TOPSOM can solve large-scale TSPs with high precision and high efficiency with limited computational costs. Extensive experimental results on mainstream route planning benchmarks including TSPLIB and National TSP’s show that our method consistently outperforms baseline methods, by up to 7.7% in terms of the Percent Deviation of Mean solution to best known solution. Qingshu Guan, Xiaopeng Hong, Wei Ke 0003, Liangfei Zhang, Guanghui Sun, Yihong Gong |
IROS | 2 |
| 2021 | Few-shot Learning for Multi-Modality TasksabstractRecent deep learning methods rely on a large amount of labeled data to achieve high performance. These methods may be impractical in some scenarios, where manual data annotation is costly or the samples of certain categories are scarce (e.g., tumor lesions, endangered animals and rare individual activities). When only limited annotated samples are available, these methods usually suffer from the overfitting problem severely, which degrades the performance significantly. In contrast, humans can recognize the objects in the images rapidly and correctly with their prior knowledge after exposed to only a few annotated samples. To simulate the learning schema of humans and relieve the reliance on the large-scale annotation benchmarks, researchers start shifting towards the few-shot learning problem: they try to learn a model to correctly recognize novel categories with only a few annotated samples. Jie Chen 0001, Qixiang Ye, Xiaoshan Yang, Shaohua Kevin Zhou, Xiaopeng Hong, Li Zhang 0040 |
ACM Multimedia | 5 |
| 2021 | FME'21: 1st Workshop on Facial Micro-Expression: Advanced Techniques for Facial Expressions Generation and SpottingabstractFacial micro-expressions (FMEs) are involuntary facial movements that occur spontaneously when a person experiences an emotion but tries to suppress or repress the facial expression and usually occur in high-risk situations. Thus, FMEs are very short in duration, an important feature that distinguishes them from ordinary facial expressions. And MEs are considered to be one of the most valuable cues for complex human emotion understanding and lie detection. Since 2014, the computational analysis and automation of MEs have been an emerging area of face research. The workshop will explore various dimensions of the human mind through emotion understanding and FME analysis, as well as extended research based on multi modal approaches. Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong |
ACM Multimedia | 5 |
| 2021 | Structural Knowledge Organization and Transfer for Class-Incremental LearningabstractDeep models are vulnerable to catastrophic forgetting when fine-tuned on new data. Popular distillation-based methods usually neglect the relations between data samples and may eventually forget essential structural knowledge. To solve these shortcomings, we propose a structural graph knowledge distillation based incremental learning framework to preserve both the positions of samples and their relations. Firstly, a memory knowledge graph (MKG) is generated to fully characterize the structural knowledge of historical tasks. Secondly, we develop a graph interpolation mechanism to enrich the domain of knowledge and alleviate the inter-class sample imbalance issue. Thirdly, we introduce structural graph knowledge distillation to transfer the knowledge of historical tasks. Comprehensive experiments on three datasets validate the proposed method. Xiaopeng Hong, Songlin Dong, Jingang Shi, Yihong Gong |
MMAsia | 2 |
| 2021 | Micro-expression spotting: A new benchmarkabstractMicro-expressions (MEs) are brief and involuntary facial expressions that occur when people are trying to hide their true feelings or conceal their emotions. Based on psychology research, MEs play an important role in understanding genuine emotions, which leads to many potential applications. Therefore, ME analysis has become an attractive topic for various research areas, such as psychology, law enforcement, and psychotherapy. In the computer vision field, the study of MEs can be divided into two main tasks, spotting and recognition, which are used to identify positions of MEs in videos and determine the emotion category of the detected MEs, respectively. Recently, although much research has been done, no fully automatic system for analyzing MEs has yet been constructed on a practical level for two main reasons: most of the research on MEs only focuses on the recognition part, while abandoning the spotting task; current public datasets for ME spotting are not challenging enough to support developing a robust spotting algorithm. The contributions of this paper are threefold: (1) we introduce an extension of the SMIC-E database, namely the SMIC-E-Long database, which is a new challenging benchmark for ME spotting; (2) we suggest a new evaluation protocol that standardizes the comparison of various ME spotting techniques; (3) extensive experiments with handcrafted and deep learning-based approaches on the SMIC-E-Long database are performed for baseline evaluation. Thuong-Khanh Tran, Quang Nhat Vo, Xiaopeng Hong, Guoying Zhao 0001 |
Neurocomputing | 3 |
| 2021 | Tripool: Graph triplet pooling for 3D skeleton-based action recognitionabstractGraph Convolutional Network (GCN) has already been successfully applied to skeleton-based action recognition. However, current GCNs in this task are lack of pooling operations such that the architectures are inherently flat, which not only increases the computational complexity but also requires larger memory space to keep the entire graph embedding. More seriously, a flat architecture forces the high-level semantic feature representations to have the same physical structure of the low-level input skeletons, which we argue is unreasonable and harmful for the final performance. To address these issues, we propose Tripool, a novel graph pooling method for 3D action recognition from skeleton data. Tripool provides to optimize a triplet pooling loss, in which both graph topology and global graph context are taken into consideration, to learn a hierarchical graph representation. The training process of graph pooling is efficient since it optimizes the graph topology by minimizing an upper bound of the pooling loss. Besides, Tripool also automatically generates an embedding matrix since the graph is changed after pooling. On one hand, Tripool reduces the computational cost by removing the redundant nodes. On the other hand it overcomes the limitation of the topology constrain for the high-level semantic representations, thus improves the final performance. Tripool can be combined with various graph neural networks in an end-to-end fashion. Comprehensive experiments on two current largest scale 3D datasets are conducted to evaluate our method. With our Tripool, we consistently get the best results in terms of various performance measures. Wei Peng 0009, Xiaopeng Hong, Guoying Zhao 0001 |
Pattern Recognit. | 2 |
| 2021 | Beyond Universal Person Re-Identification AttackabstractDeep learning-based person re-identification (Re-ID) has made great progress and achieved high performance recently. In this paper, we make the first attempt to examine the vulnerability of current person Re-ID models against a dangerous attack method, i.e., the universal adversarial perturbation (UAP) attack, which has been shown to fool classification models with a little overhead. We propose a more universal adversarial perturbation (MUAP) method for both image-agnostic and model-insensitive person Re-ID attack. Firstly, we adopt a list-wise attack objective function to disrupt the similarity ranking list directly. Secondly, we propose a model-insensitive mechanism for cross-model attack. Extensive experiments show that the proposed attack approach achieves high attack performance and outperforms other state of the arts by large margin in cross-model scenario. The results also demonstrate the vulnerability of current Re-ID models to MUAP and further suggest the need of designing more robust Re-ID models. Xing Wei 0001, Rongrong Ji, Xiaopeng Hong, Qi Tian 0001, Yihong Gong |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Analogy-Detail Networks for Object RecognitionabstractThe human visual system can recognize object categories accurately and efficiently and is robust to complex textures and noises. To mimic the analogy-detail dual-pathway human visual cognitive mechanism revealed in recent cognitive science studies, in this article, we propose a novel convolutional neural network (CNN) architecture named analogy-detail networks (ADNets) for accurate object recognition. ADNets disentangle the visual information and process them separately using two pathways: the analogy pathway extracts coarse and global features representing the gist (i.e., shape and topology) of the object, while the detail pathway extracts fine and local features representing the details (i.e., texture and edges) for determining object categories. We modularize the architecture and encapsulate the two pathways into the analogy-detail block as the CNN building block to construct ADNets. For implementation, we propose a general principle that transmutes typical CNN structures into the ADNet architecture and applies the transmutation on representative baseline CNNs. Extensive experiments on CIFAR10, CIFAR100, street view house numbers, and ImageNet data sets demonstrate that ADNets significantly reduce the test error rates of the baseline CNNs by up to 5.76% and outperform other state-of-the-art architectures. Comprehensive analysis and visualizations further demonstrate that ADNets are interpretable and have a better shape-texture tradeoff for recognizing the objects with complex textures. Xiaopeng Hong, Weiwei Shi 0003, Xinyuan Chang, Yihong Gong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Infrared-Visible Cross-Modal Person Re-Identification with an X ModalityabstractThis paper focuses on the emerging Infrared-Visible cross-modal person re-identification task (IV-ReID), which takes infrared images as input and matches with visible color images. IV-ReID is important yet challenging, as there is a significant gap between the visible and infrared images. To reduce this ‘gap’, we introduce an auxiliary X modality as an assistant and reformulate infrared-visible dual-mode cross-modal learning as an X-Infrared-Visible three-mode learning problem. The X modality restates from RGB channels to a format with which cross-modal learning can be easily performed. With this idea, we propose an X-Infrared-Visible (XIV) ReID cross-modal learning framework. Firstly, the X modality is generated by a lightweight network, which is learnt in a self-supervised manner with the labels inherited from visible images. Secondly, under the XIV framework, cross-modal learning is guided by a carefully designed modality gap constraint, with information exchanged cross the visible, X, and infrared modalities. Extensive experiments are performed on two challenging datasets SYSU-MM01 and RegDB to evaluate the proposed XIV-ReID approach. Experimental results show that our method considerably achieves an absolute gain of over 7% in terms of rank 1 and mAP even compared with the latest state-of-the-art methods. Diangang Li, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
AAAI | 3 |
| 2020 | Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingabstractHuman action recognition from skeleton data, fuelled by the Graph Convolutional Network (GCN) with its powerful capability of modeling non-Euclidean data, has attracted lots of attention. However, many existing GCNs provide a pre-defined graph structure and share it through the entire network, which can loss implicit joint correlations especially for the higher-level features. Besides, the mainstream spectral GCN is approximated by one-order hop such that higher-order connections are not well involved. All of these require huge efforts to design a better GCN architecture. To address these problems, we turn to Neural Architecture Search (NAS) and propose the first automatically designed GCN for this task. Specifically, we explore the spatial-temporal correlations between nodes and build a search space with multiple dynamic graph modules. Besides, we introduce multiple-hop modules and expect to break the limitation of representational capacity caused by one-order approximation. Moreover, a corresponding sampling- and memory-efficient evolution strategy is proposed to search in this space. The resulted architecture proves the effectiveness of the higher-order approximation and the layer-wise dynamic graph modules. To evaluate the performance of the searched model, we conduct extensive experiments on two very large scale skeleton-based action recognition datasets. The results show that our model gets the state-of-the-art results in term of given metrics. Wei Peng 0009, Xiaopeng Hong, Haoyu Chen 0001, Guoying Zhao 0001 |
AAAI | 2 |
| 2020 | Bi-Objective Continual Learning: Learning 'New' While Consolidating 'Known'abstractIn this paper, we propose a novel single-task continual learning framework named Bi-Objective Continual Learning (BOCL). BOCL aims at both consolidating historical knowledge and learning from new data. On one hand, we propose to preserve the old knowledge using a small set of pillars, and develop the pillar consolidation (PLC) loss to preserve the old knowledge and to alleviate the catastrophic forgetting problem. On the other hand, we develop the contrastive pillar (CPL) loss term to improve the classification performance, and examine several data sampling strategies for efficient onsite learning from ‘new’ with a reasonable amount of computational resources. Comprehensive experiments on CIFAR10/100, CORe50 and a subset of ImageNet validate the BOCL framework. We also reveal the performance accuracy of different sampling strategies when used to finetune a given CNN model. The code will be released. Xiaopeng Hong, Xinyuan Chang, Yihong Gong |
AAAI | 2 |
| 2020 | Superpixel Masking and Inpainting for Self-Supervised Anomaly Detection
Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
BMVC | 6 |
| 2020 | Noise-Aware Fully Webly Supervised Object DetectionabstractWe investigate the emerging task of learning object detectors with sole image-level labels on the web without requiring any other supervision like precise annotations or additional images from well-annotated benchmark datasets. Such a task, termed as fully webly supervised object detection, is extremely challenging, since image-level labels on the web are always noisy, leading to poor performance of the learned detectors. In this work, we propose an end-to-end framework to jointly learn webly supervised detectors and reduce the negative impact of noisy labels. Such noise is heterogeneous, which is further categorized into two types, namely background noise and foreground noise. Regarding the background noise, we propose a residual learning structure incorporated with weakly supervised detection, which decomposes background noise and models clean data. To explicitly learn the residual feature between clean data and noisy labels, we further propose a spatially-sensitive entropy criterion, which exploits the conditional distribution of detection results to estimate the confidence of background categories being noise. Regarding the foreground noise, a bagging-mixup learning is introduced, which suppresses foreground noisy signals from incorrectly labelled images, whilst maintaining the diversity of training data. We evaluate the proposed approach on popular benchmark datasets by training detectors on web images, which are retrieved by the corresponding category tags from photo-sharing sites. Extensive experiments show that our method achieves significant improvements over the state-of-the-art methods. Yunhang Shen, Rongrong Ji, Xiaopeng Hong, Feng Zheng 0001, Jianzhuang Liu, Mingliang Xu 0001, Qi Tian 0001 |
CVPR | 4 |
| 2020 | Few-Shot Class-Incremental LearningabstractThe ability to incrementally learn new classes is crucial to the development of real-world artificial intelligence systems. In this paper, we focus on a challenging but practical few-shot class-incremental learning (FSCIL) problem. FSCIL requires CNN models to incrementally learn new classes from very few labelled samples, without forgetting the previously learned ones. To address this problem, we represent the knowledge using a neural gas (NG) network, which can learn and preserve the topology of the feature manifold formed by different classes. On this basis, we propose the TOpology-Preserving knowledge InCrementer (TOPIC) framework. TOPIC mitigates the forgetting of the old classes by stabilizing NG's topology and improves the representation learning for few-shot new classes by growing and adapting NG to new training samples. Comprehensive experimental results demonstrate that our proposed method significantly outperforms other state-of-the-art class-incremental learning methods on CIFAR100, miniImageNet, and CUB200 datasets. Xiaopeng Hong, Xinyuan Chang, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 2 |
| 2020 | Topology-Preserving Class-Incremental Learning
Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Yihong Gong |
ECCV (19) | 3 |
| 2020 | MEGC2020 - The Third Facial Micro-Expression Grand ChallengeabstractThe recent emergence of automatic facial micro-expression analysis has attracted a lot of attention in the last five years. Compared to the advances made in micro-expression recognition, the task of micro-expression spotting from long videos is tremendously in need of more effective methods. This paper summarises the 3rd Facial Micro-Expression Grand Challenge (MEGC 2020) held in conjunction with the 15th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2020. In this workshop, we propose a new challenge of spotting both macro- and micro-expressions from long videos, to spur the community to develop new techniques for micro-expression spotting and also to extend facial micro-expression analysis to more complex real-world scenarios where micro-expressions are likely to be intertwined among normal expressions. In this paper, we outline the evaluation protocols for the challenge task, and describe the datasets involved. Then, we summarize the methods from the accepted challenge papers, present the comparison and analysis of results, as well as future directions. Jingting Li 0001, Moi Hoon Yap, John See, Xiaopeng Hong |
FG | 5 |
| 2020 | Complex Spatial-Temporal Attention Aggregation For Video Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to match pedestrian tracklets of the same identity captured by different cameras. Existing works usually compute the video-level feature representation via simple frame-level feature aggregation, such as average pooling and max pooling. However, the performance of such methods degenerates severely under low signal-noise ratio and partial occlusions. In this paper, we propose a novel Complex Spatial-Temporal Attention Aggregation (CAA), which fully exploits the discriminative information in spatial-temporal dimension via the combination of two aggregation method, namely region-aware aggregation and region-regardless aggregation. We evaluate the proposed method in three widely used video Re-ID datasets, including MARS, iLIDS-VID, and PRID-2011. The experimental results demonstrate that the proposed method outperforms the state of the arts. Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 3 |
| 2020 | Class-Incremental Learning with Topological Schemas of Memory SpacesabstractClass-incremental learning (CIL) aims to incrementally learn a unified classifier for new classes emerging, which suffers from the catastrophic forgetting problem. To alleviate forgetting and improve the recognition performance, we propose a novel CIL framework, named the topological schemas model (TSM). TSM consists of a Gaussian mixture model arranged on 2D grids (2D-GMM) as the memory of the learned knowledge. To train the 2D-GMM model, we develop a novel competitive expectation-maximization (CEM) method, which contains a global topology embedding step and a local expectation-maximization fine-tuning step. Meanwhile, we choose the image samples of old classes that have the maximum posterior probability with respect to each Gaussian distribution as the episodic points. When finetuning for new classes, we propose the memory preservation loss (MPL) term to ensure episodic points still have maximum probabilities with respect to the corresponding Gaussian distribution. MPL preserves the distribution of 2D-GMM for old knowledge during incremental learning and alleviates catastrophic forgetting. Comprehensive experimental evaluations on two popular CIL benchmarks CIFAR100 and subImageNet demonstrate the superiority of our TSM. Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Wei Ke 0003, Yihong Gong |
ICPR | 3 |
| 2020 | Polynomial Universal Adversarial Perturbations for Person Re-IdentificationabstractIn this paper, we focus on Universal Adversarial Perturbations (UAP) attack on state-of-the-art person re-identification (Re-ID) methods. Existing UAP methods usually compute a perturbation image and add it to the images of interest. Such a simple constant form greatly limits the attack power. To address this problem, we extend the formulation of UAP to a polynomial form and propose the Polynomial Universal Adversarial Perturbation (PUAP). Unlike traditional UAP methods which only rely on the additive perturbation signal, the proposed PUAP consists of both an additive perturbation and a multiplicative modulation factor. The additive perturbation produces the fundamental component of the signal, while the multiplicative factor modulates the perturbation signal in line with the unit impulse pattern of the input image. Moreover, we introduce a Pearson correlation coefficient loss to generate universal perturbations, for disrupting the outputs of person Re-ID models. Extensive experiments on DukeMTMC-reID, Market-1501, and MARS show that the proposed method can efficiently improve the attack performance, especially when the magnitude of UAP is constrained to a relatively small value. Xing Wei 0001, Rongrong Ji, Xiaopeng Hong, Yihong Gong |
ICPR | 4 |
| 2020 | Learning Scales from Points: A Scale-aware Probabilistic Model for Crowd CountingabstractCounting people automatically through computer vision technology is a challenging task. Recently, convolution neural network (CNN) based methods have made significant progress. Nonetheless, large scale variations of instances caused by, for example, perspective effects remain unsolved. Moreover, it is problematic to estimate scales with only point annotations. In this paper, we propose a scale-aware probabilistic model to handle this problem. Unlike previous methods that generate a single density map where instances of various scales are processed indiscriminately, we propose a density pyramid network (DPN), where each pyramid level handles instances within a particular scale range. Furthermore, we propose a scale distribution estimator (SDE) to learn scales of people from input data, under the weak supervision of point annotations. Finally, we adopt an instance-level probabilistic scale-aware model (IPSM) to guide the multi-scale training of DPN explicitly. Qualitative and quantitative experimental results demonstrate the effectiveness of the proposed method, which achieves competitive results on four widely used benchmarks. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ACM Multimedia | 3 |
| 2020 | Co-Attentive Lifting for Infrared-Visible Person Re-IdentificationabstractInfrared-visible cross-modality person re-identification (IV-ReID) has attracted much attention with the popularity of dual-mode video surveillance systems, where the RGB mode works in the daytime and automatically switches to the infrared mode at night. Despite its significant application value, IV-ReID remains a difficult problem mainly due to two great challenges. First, it is difficult to identify persons in the infrared image, which lacks color and texture clues. Second, there is a significant gap between the infrared and visible modalities where appearances of the same person vary considerably. This paper proposes a novel attention-based approach to handle the two difficulties in a unified framework. 1) We propose an attention lifting mechanism to learn discriminative features in each modality. 2) We propose a co-attentive learning mechanism to bridge the gap between the two modalities. Our method only makes slight modifications of a given backbone network and requires small computation overhead while improving the performance significantly. We conduct extensive experiments to demonstrate the superiority of our proposed method. Xing Wei 0001, Diangang Li, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
ACM Multimedia | 3 |
| 2020 | K-armed Bandit based Multi-Modal Network Architecture Search for Visual Question AnsweringabstractIn this paper, we propose a cross-modal network architecture search (NAS) algorithm for VQA, termed as k-Armed Bandit based NAS (KAB-NAS). KAB-NAS regards the design of each layer as a k-armed bandit problem and updates the preference of each candidate via numerous samplings in a single-shot search framework. To establish an effective search space, we further propose a new architecture termed Automatic Graph Attention Network (AGAN), and extend the popular self-attention layer with three graph structures, denoted as dense-graph, co-graph and separate-graph.These graph layers are used to form the direction of information propagation in the graph network, and their optimal combinations are searched by KAB-NAS. To evaluate KAB-NAS and AGAN, we conduct extensive experiments on two VQA benchmark datasets, i.e., VQA2.0 and GQA, and also test AGAN with the popular BERT-style pre-training. The experimental results show that with the help of KAB-NAS, AGAN can achieve the state-of-the-art performance on both benchmark datasets with much fewer parameters and computations. Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Gen Luo, Xiaopeng Hong, Jinsong Su, Xinghao Ding, Ling Shao 0001 |
ACM Multimedia | 5 |
| 2020 | Transductive semi-supervised metric learning for person re-identification
Xinyuan Chang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
Pattern Recognit. | 4 |
| 2020 | Multi-Target Multi-Camera Tracking by Tracklet-to-Target AssignmentabstractThis paper focuses on the Multi-Target Multi-Camera Tracking task (MTMCT), which aims at tracking multiple targets within a multi-camera network. As the trajectory of each target is inherently split into multiple sub-trajectories (namely local tracklets) in a multi-camera network, a major challenge of MTMCT is how to accurately match the local tracklets generated within each camera across different cameras and generate a complete global trajectory for each target, i.e., the cross-camera tracklet matching problem. We solve the cross-camera tracklet matching problem by TRACklet-to-Target Assignment (TRACTA), and propose the Restricted Non-negative Matrix Factorization (RNMF) algorithm to compute the optimal assignment solution that meets a set of constraints, which should be in force in practice. TRACTA can correct the tracking errors caused by occlusions and missed detections in local tracklets, and produce a complete global trajectory for each target across all the cameras. Moreover, we also develop an analytical way of estimating the total number of targets in the camera network, which plays an important role to compute the tracklet-to-target assignment. Experimental evaluations and ablation studies on four MTMCT benchmark datasets show the superiority of the proposed TRACTA method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Weiwei Shi 0003, Yihong Gong |
IEEE Trans. Image Process. | 3 |
| 2020 | 3D Skeletal Gesture Recognition via Hidden States ExplorationabstractTemporal dynamics is an open issue for modeling human body gestures. A solution is resorting to the generative models, such as the hidden Markov model (HMM). Nevertheless, most of the work assumes fixed anchors for each hidden state, which make it hard to describe the explicit temporal structure of gestures. Based on the observation that a gesture is a time series with distinctly defined phases, we propose a new formulation to build temporal compositions of gestures by the low-rank matrix decomposition. The only assumption is that the gesture's "hold" phases with static poses are linearly correlated among each other. As such, a gesture sequence could be segmented into temporal states with semantically meaningful and discriminative concepts. Furthermore, different to traditional HMMs which tend to use specific distance metric for clustering and ignore the temporal contextual information when estimating the emission probability, we utilize the long short-term memory to learn probability distributions over states of HMM. The proposed method is validated on multiple challenging datasets. Experiments demonstrate that our approach can effectively work on a wide range of gestures, and achieve state-of-the-art performance. Xin Liu 0012, Henglin Shi, Xiaopeng Hong, Haoyu Chen 0001, Dacheng Tao, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-ExpressionsabstractRecently, the recognition task of spontaneous facial micro-expressions has attracted much attention with its various real-world applications. Plenty of handcrafted or learned features have been employed for a variety of classifiers and achieved promising performances for recognizing micro-expressions. However, the micro-expression recognition is still challenging due to the subtle spatiotemporal changes of micro-expressions. To exploit the merits of deep learning, we propose a novel deep recurrent convolutional networks based micro-expression recognition approach, capturing the spatiotemporal deformations of micro-expression sequence. Specifically, the proposed deep model is constituted of several recurrent convolutional layers for extracting visual features and a classificatory layer for recognition. It is optimized by an end-to-end manner and obviates manual feature design. To handle sequential data, we exploit two ways to extend the connectivity of convolutional networks across temporal domain, in which the spatiotemporal deformations are modeled in views of facial appearance and geometry separately. Besides, to overcome the shortcomings of limited and imbalanced training samples, two temporal data augmentation strategies as well as a balanced loss are jointly used for our deep network. By performing the experiments on three spontaneous micro-expression datasets, we verify the effectiveness of our proposed micro-expression recognition approach compared to the state-of-the-art methods. Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Corrections to "Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-Expressions"abstractPresents corrections to the author's information in the above named paper. Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | A Boost in Revealing Subtle Facial Expressions: A Consolidated Eulerian FrameworkabstractFacial Micro-expression Recognition (MER) distinguishes the underlying emotional states of spontaneous subtle facial expressions. Automatic MER is challenging because that the intensity of subtle facial muscle movement is extremely low and the duration of ME is transient.Recent works adopt motion magnification or temporal interpolation to resolve these issues. Nevertheless, existing works divide them into two separate modules due to their non-linearity. Though such operation eases the difficulty in implementation, it ignores their underlying connections and thus results in inevitable losses in both accuracy and speed. Instead, in this paper, we propose a consolidated Eulerian framework to reveal the subtle facial movements. It expands the temporal duration and amplifies the muscle movements in micro-expressions simultaneously. Compared to existing approaches, the proposed method can not only process ME clips more efficiently but also make subtle ME movements more distinguishable. Experiments on two public MER databases indicate that our model outperforms the state-of-the-art in both speed and accuracy. Wei Peng 0009, Xiaopeng Hong, Yingyue Xu, Guoying Zhao 0001 |
FG | 2 |
| 2019 | MEGC 2019 - The Second Facial Micro-Expressions Grand ChallengeabstractAutomatic facial micro-expression (ME) analysis is a growing field of research that has gained much attention in the last five years. With many recent works testing on limited data, there is a need to spur better approaches that are both robust and effective. This paper summarises the 2nd Facial Micro-Expression Grand Challenge (MEGC 2019) held in conjunction with the 14th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2019. In this workshop, we proposed challenges for two micro-expression (ME) tasks- spotting and recognition, with the aim of encouraging rigorous evaluation and development of new robust techniques that can accommodate data captured across a variety of settings. In this paper, we outline the evaluation protocols for the two challenge tasks, the datasets involved, and an analysis of the best performing works from the participating teams, together with a summary of results. Finally, we highlight some possible future directions. John See, Moi Hoon Yap, Jingting Li 0001, Xiaopeng Hong |
FG | 4 |
| 2019 | Universal Perturbation Attack Against Image RetrievalabstractUniversal adversarial perturbations (UAPs), a.k.a. input-agnostic perturbations, has been proved to exist and be able to fool cutting-edge deep learning models on most of the data samples. Existing UAP methods mainly focus on attacking image classification models. Nevertheless, little attention has been paid to attacking image retrieval systems. In this paper, we make the first attempt in attacking image retrieval systems. Concretely, image retrieval attack is to make the retrieval system return irrelevant images to the query at the top ranking list. It plays an important role to corrupt the neighbourhood relationships among features in image retrieval attack. To this end, we propose a novel method to generate retrieval-against UAP to break the neighbourhood relationships of image features via degrading the corresponding ranking metric. To expand the attack method to scenarios with varying input sizes or untouchable network parameters, a multi-scale random resizing scheme and a ranking distillation strategy are proposed. We evaluate the proposed method on four widely-used image retrieval datasets, and report a significant performance drop in terms of different metrics, such as mAP and mP@10. Finally, we test our attack methods on the real-world visual search engine, i.e., Google Images, which demonstrates the practical potentials of our methods. Jie Li 0052, Rongrong Ji, Hong Liu 0009, Xiaopeng Hong, Yue Gao 0002, Qi Tian 0001 |
ICCV | 4 |
| 2019 | Bayesian Loss for Crowd Count Estimation With Point SupervisionabstractIn crowd counting datasets, each person is annotated by a point, which is usually the center of the head. And the task is to estimate the total count in a crowd scene. Most of the state-of-the-art methods are based on density map estimation, which convert the sparse point annotations into a “ground truth” density map through a Gaussian kernel, and then use it as the learning target to train a density map estimator. However, such a "ground-truth" density map is imperfect due to occlusions, perspective effects, variations in object shapes, etc. On the contrary, we propose Bayesian loss, a novel loss function which constructs a density contribution probability model from the point annotations. Instead of constraining the value at every pixel in the density map, the proposed training loss adopts a more reliable supervision on the count expectation at each annotated point. Without bells and whistles, the loss function makes substantial improvements over the baseline loss on all tested datasets. Moreover, our proposed loss function equipped with a standard backbone network, without using any external detectors or multi-scale architectures, plays favourably against the state of the arts. Our method outperforms previous best approaches by a large margin on the latest and largest UCF-QNRF dataset. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICCV | 3 |
| 2019 | Structured Modeling of Joint Deep Feature and Prediction Refinement for Salient Object DetectionabstractRecent saliency models extensively explore to incorporate multi-scale contextual information from Convolutional Neural Networks (CNNs). Besides direct fusion strategies, many approaches introduce message-passing to enhance CNN features or predictions. However, the messages are mainly transmitted in two ways, by feature-to-feature passing, and by prediction-to-prediction passing. In this paper, we add message-passing between features and predictions and propose a deep unified CRF saliency model . We design a novel cascade CRFs architecture with CNN to jointly refine deep features and predictions at each scale and progressively compute a final refined saliency map. We formulate the CRF graphical model that involves message-passing of feature-feature, feature-prediction, and prediction-prediction, from the coarse scale to the finer scale, to update the features and the corresponding predictions. Also, we formulate the mean-field updates for joint end-to-end model training with CNN through back propagation. The proposed deep unified CRF saliency model is evaluated over six datasets and shows highly competitive performance among the state of the arts. Yingyue Xu, Dan Xu 0002, Xiaopeng Hong, Wanli Ouyang, Rongrong Ji, Min Xu 0001, Guoying Zhao 0001 |
ICCV | 3 |
| 2019 | Remote Heart Rate Measurement From Highly Compressed Facial Videos: An End-to-End Deep Learning Solution With Video EnhancementabstractRemote photoplethysmography (rPPG), which aims at measuring heart activities without any contact, has great potential in many applications (e.g., remote healthcare). Existing rPPG approaches rely on analyzing very fine details of facial videos, which are prone to be affected by video compression. Here we propose a two-stage, end-to-end method using hidden rPPG information enhancement and attention networks, which is the first attempt to counter video compression loss and recover rPPG signals from highly compressed videos. The method includes two parts: 1) a Spatio-Temporal Video Enhancement Network (STVEN) for video enhancement, and 2) an rPPG network (rPPGNet) for rPPG signal recovery. The rPPGNet can work on its own for robust rPPG measurement, and the STVEN network can be added and jointly trained to further boost the performance especially on highly compressed videos. Comprehensive experiments are performed on two benchmark datasets to show that, 1) the proposed method not only achieves superior performance on compressed videos with high-quality videos pair, 2) it also generalizes well on novel data with only compressed videos available, which implies the promising potential for real-world applications. Zitong Yu, Wei Peng 0009, Xiaopeng Hong, Guoying Zhao 0001 |
ICCV | 4 |
| 2019 | Video Action Recognition Via Neural Architecture SearchingabstractDeep neural networks have achieved great success for video analysis and understanding. However, designing a high-performance neural architecture requires substantial efforts and expertise. In this paper, we make the first attempt to let algorithm automatically design neural networks for video action recognition tasks. Specifically, a spatio-temporal network is developed in a differentiable space modeled by a directed acyclic graph, thus a gradient-based strategy can be performed to search an optimal architecture. Nonetheless, it is computationally expensive, since the computational burden to evaluate each architecture candidate is still heavy. To alleviate this issue, we, for the video input, introduce a temporal segment approach to reduce the computational cost without losing global video information. For the architecture, we explore in an efficient search space by introducing pseudo 3D operators. Experiments show that, our architecture outperforms popular neural architectures, under the training from scratch protocol, on the challenging UCF101 dataset, surprisingly, with only around one percentage of parameters of its manual-design counterparts. Wei Peng 0009, Xiaopeng Hong, Guoying Zhao 0001 |
ICIP | 2 |
| 2019 | A Part Power Set Model for Scale-Free Person RetrievalabstractRecently, person re-identification (re-ID) has attracted increasing research attention, which has broad application prospects in video surveillance and beyond. To this end, most existing methods highly relied on well-aligned pedestrian images and hand-engineered part-based model on the coarsest feature map. In this paper, to lighten the restriction of such fixed and coarse input alignment, an end-to-end part power set model with multi-scale features is proposed, which captures the discriminative parts of pedestrians from global to local, and from coarse to fine, enabling part-based scale-free person re-ID. In particular, we first factorize the visual appearance by enumerating $k$-combinations for all $k$ of $n$ body parts to exploit rich global and partial information to learn discriminative feature maps. Then, a combination ranking module is introduced to guide the model training with all combinations of body parts, which alternates between ranking combinations and estimating an appearance model. To enable scale-free input, we further exploit the pyramid architecture of deep networks to construct multi-scale feature maps with a feasible amount of extra cost in term of memory and time. Extensive experiments on the mainstream evaluation datasets, including Market-1501, DukeMTMC-reID and CUHK03, validate that our method achieves the state-of-the-art performance. Yunhang Shen, Rongrong Ji, Xiaopeng Hong, Feng Zheng 0001, Yongjian Wu 0001, Feiyue Huang |
IJCAI | 3 |
| 2019 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problems in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from two aspects. First, we establish a CDMER experimental evaluation protocol and provide a standard platform for evaluating their proposed methods. Second, we conduct extensive benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating the CDMER problem from two different perspectives and deeply analyze and discuss the experimental results. In addition, all the data and codes involving CDMER in this paper are released on our project website: http://aip.seu.edu.cn/cdmer. Yuan Zong, Wenming Zheng, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
ICMR | 3 |
| 2019 | Hidden States Exploration for 3D Skeleton-Based Gesture Recognitionabstract3D skeletal data has recently attracted wide attention in human behavior analysis for its robustness to variant scenes, while accurate gesture recognition is still challenging. The main reason lies in the high intra-class variance caused by temporal dynamics. A solution is resorting to the generative models, such as the hidden Markov model (HMM). However, existing methods commonly assume fixed anchors for each hidden state, which is hard to depict the explicit temporal structure of gestures. Based on the observation that a gesture is a time series with distinctly defined phases, we propose a new formulation to build temporal compositions of gestures by the low-rank matrix decomposition. The only assumption is that the gesture's "hold" phases with static poses are linearly correlated among each other. As such, a gesture sequence could be segmented into temporal states with semantically meaningful and discriminative concepts. Furthermore, different to traditional HMMs which tend to use specific distance metric for clustering and ignore the temporal contextual information when estimating the emission probability, the Long Short-Term Memory (LSTM) is utilized to learn probability distributions over states of HMM. The proposed method is validated on two challenging datasets. Experiments demonstrate that our approach can effectively work on a wide range of gestures and actions, and achieve state-of-the-art performance. Xin Liu 0012, Henglin Shi, Xiaopeng Hong, Haoyu Chen 0001, Dacheng Tao, Guoying Zhao 0001 |
WACV | 3 |
| 2019 | Zero-Shot Learning Via Recurrent Knowledge TransferabstractZero-shot learning (ZSL) which aims to learn new concepts without any labeled training data is a promising solution to large-scale concept learning. Recently, many works implement zero-shot learning by transferring structural knowledge from the semantic embedding space to the image feature space. However, we observe that such direct knowledge transfer may suffer from the space shift problem in the form of the inconsistency of geometric structures in the training and testing spaces. To alleviate this problem, we propose a novel method which actualizes recurrent knowledge transfer (RecKT) between the two spaces. Specifically, we unite the two spaces into the joint embedding space in which unseen image data are missing. The proposed method provides a synthesis-refinement mechanism to learn the shared subspace structure (SSS) and synthesize missing data simultaneously in the joint embedding space. The synthesized unseen image data are utilized to construct the classifier for unseen classes. Experimental results show that our method outperforms the state-of-the-art on three popular datasets. The ablation experiment and visualization of the learning process illustrate how our method can alleviate the space shift problem. By product, our method provides a perspective to interpret the ZSL performance by implementing subspace clustering on the learned SSS. Bo Zhao 0015, Xinwei Sun 0001, Xiaopeng Hong, Yuan Yao 0011, Yizhou Wang 0001 |
WACV | 3 |
| 2019 | A spatial-aware joint optic disc and cup segmentation method
Qing Liu 0003, Xiaopeng Hong, Shuo Li 0001, Zailiang Chen 0001, Guoying Zhao 0001, Beiji Zou 0001 |
Neurocomputing | 2 |
| 2019 | Saliency Integration: An Arbitrator ModelabstractSaliency integration has attracted much attention on unifying saliency maps from multiple saliency models. Previous offline integration methods usually face two challenges: 1) if most of the candidate saliency models misjudge the saliency on an image, the integration result will lean heavily on those inferior candidate models; and 2) an unawareness of the ground truth saliency labels brings difficulty in estimating the expertise of each candidate model. To address these problems, in this paper, we propose an arbitrator model (AM) for saliency integration. First, we incorporate the consensus of multiple saliency models and the external knowledge into a reference map to effectively rectify the misleading by candidate models. Second, our quest for ways of estimating the expertise of the saliency models without ground truth labels gives rise to two distinct online model-expertise estimation methods. Finally, we derive a Bayesian integration framework to reconcile the saliency models of varying expertise and the reference map. To extensively evaluate the proposed AM model, we test 27 state-of-the-art saliency models, covering both traditional and deep learning ones, on various combinations over four datasets. The evaluation results show that the AM model improves the performance substantially compared to the existing state-of-the-art integration methods, regardless of the chosen candidate saliency models. Yingyue Xu, Xiaopeng Hong, Fatih Porikli, Xin Liu 0012, Jie Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Bidirectional Long Short-Term Memory Variational Autoencoder
Henglin Shi, Xin Liu 0012, Xiaopeng Hong, Guoying Zhao 0001 |
BMVC | 3 |
| 2018 | Facial Micro-Expressions Grand Challenge 2018 SummaryabstractThis paper summarises the Facial Micro-Expression Grand Challenge (MEGC 2018) held in conjunction with the 13th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2018. In this workshop, we aim to stimulate new ideas and techniques for facial micro-expression analysis by proposing a new cross-database challenge. Two state-of-the-art datasets, CASME II and SAMM, are used to validate the performance of existing and new algorithms. Also, the challenge advocates the recognition of micro-expressions based on AU-centric objective classes rather than emotional classes. We present a summary and analysis of the baseline results using LBP-TOP, HOOF and 3DHOG, together with results from the challenge submissions. Moi Hoon Yap, John See, Xiaopeng Hong |
FG | 3 |
| 2018 | Sparse Tikhonov-Regularized Hashing for Multi-Modal LearningabstractThis paper mainly focuses on the role of regularization in Multi-Modal Learning (MML). Existing MML studies devote most of the efforts in maximizing the consensus of models from cues of different modalities. However, regularization methods are still far from fully explored. To fill in this gap, we propose a compact and efficient coding solution, termed by sparse Tikhonov-Regularized Hashing (STRH). The STRH enforces both the ℓ0-norm induced sparsity constraints and the Tikhonov regularization on the binary solution vectors which maximize cross-modal correlation. In addition, we raise the concerns on the challenging testing scenario of `Multi-modal Learning and Single-modal Prediction' (MLSP). Finally, we demonstrate that the STRH is an efficient hashing solutions by showing its superiority under the MLSP scenario. Lei Tian 0002, Xiaopeng Hong, Chunxiao Fan 0001, Yue Ming 0001, Matti Pietikäinen, Guoying Zhao 0001 |
ICIP | 2 |
| 2018 | Incorporating high-level and low-level cues for pain intensity estimationabstractPain is a transient physical reaction that exhibits on human faces. Automatic pain intensity estimation is of great importance in clinical and health-care applications. Pain expression is identified by a set of deformations of facial features. Hence, features are essential for pain estimation. In this paper, we propose a novel method that encodes low-level descriptors and powerful high-level deep features by a weighting process, to form an efficient representation of facial images. To obtain a powerful and compact low-level representation, we explore the way of using second-order pooling over the local descriptors. Instead of direct concatenation, we develop an efficient fusion approach that unites the low-level local descriptors and the high-level deep features. To the best of our knowledge, this is the first approach that incorporates the low-level local statistics together with the high-level deep features in pain intensity estimation. Experiments are evaluated on the benchmark databases of pain. The results demonstrate that the proposed low-to-high-level representation outperforms other methods and achieves promising results. Ruijing Yang, Xiaopeng Hong, Jinye Peng 0001, Xiaoyi Feng, Guoying Zhao 0001 |
ICPR | 2 |
| 2018 | Sparse projections matrix binary descriptors for face recognition
Chunxiao Fan 0001, Lei Tian 0002, Yue Ming 0001, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 4 |
| 2018 | Saliency detection via bi-directional propagation
Yingyue Xu, Xiaopeng Hong, Xin Liu 0012, Guoying Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Towards Reading Hidden Emotions: A Comparative Study of Spontaneous Micro-Expression Spotting and Recognition MethodsabstractMicro-expressions (MEs) are rapid, involuntary facial expressions which reveal emotions that people do not intend to show. Studying MEs is valuable as recognizing them has many important applications, particularly in forensic science and psychotherapy. However, analyzing spontaneous MEs is very challenging due to their short duration and low intensity. Automatic ME analysis includes two tasks: ME spotting and ME recognition. For ME spotting, previous studies have focused on posed rather than spontaneous videos. For ME recognition, the performance of previous studies is low. To address these challenges, we make the following contributions: (i) We propose the first method for spotting spontaneous MEs in long videos (by exploiting feature difference contrast). This method is training free and works on arbitrary unseen videos. (ii) We present an advanced ME recognition framework, which outperforms previous work by a large margin on two challenging spontaneous ME databases (SMIC and CASMEII). (iii) We propose the first automatic ME analysis system (MESR), which can spot and recognize MEs from spontaneous video data. Finally, we show our method outperforms humans in the ME recognition task by a large margin, and achieves comparable performance to humans at the very challenging task of spotting and then recognizing spontaneous MEs. Xiaopeng Hong, Antti Moilanen, Xiaohua Huang 0003, Tomas Pfister, Guoying Zhao 0001, Matti Pietikäinen |
IEEE Trans. Affect. Comput. | 2 |
| 2018 | Background Subtraction Using Spatio-Temporal Group Sparsity RecoveryabstractBackground subtraction is a key step in a wide spectrum of video applications, such as object tracking and human behavior analysis. Compressive sensing-based methods, which make little specific assumptions about the background, have recently attracted wide attention in background subtraction. Within the framework of compressive sensing, background subtraction is solved as a decomposition and optimization problem, where the foreground is typically modeled as pixel-wised sparse outliers. However, in real videos, foreground pixels are often not randomly distributed, but instead, group clustered. Moreover, due to costly computational expenses, most compressive sensing-based methods are unable to process frames online. In this paper, we take into account the group properties of foreground signals in both spatial and temporal domains, and propose a greedy pursuit-based method called spatio-temporal group sparsity recovery, which prunes data residues in an iterative process, according to both sparsity and group clustering priors, rather than merely sparsity. Furthermore, a random strategy for background dictionary learning is used to handle complex background variations, while foreground-free training is not required. Finally, we propose a two-pass framework to achieve online processing. The proposed method is validated on multiple challenging video sequences. Experiments demonstrate that our approach effectively works on a wide range of complex scenarios and achieves a state-of-the-art performance with far fewer computations. Xin Liu 0012, Jiawen Yao, Xiaopeng Hong, Xiaohua Huang 0003, Ziheng Zhou 0003, Chun Qi, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Sliding Window Based Micro-expression Spotting: A Benchmark
Thuong-Khanh Tran, Xiaopeng Hong, Guoying Zhao 0001 |
ACIVS | 2 |
| 2017 | Hierarchical Contour Closure-Based Holistic Salient Object DetectionabstractMost existing salient object detection methods compute the saliency for pixels, patches, or superpixels by contrast. Such fine-grained contrast-based salient object detection methods are stuck with saliency attenuation of the salient object and saliency overestimation of the background when the image is complicated. To better compute the saliency for complicated images, we propose a hierarchical contour closure-based holistic salient object detection method, in which two saliency cues, i.e., closure completeness and closure reliability, are thoroughly exploited. The former pops out the holistic homogeneous regions bounded by completely closed outer contours, and the latter highlights the holistic homogeneous regions bounded by averagely highly reliable outer contours. Accordingly, we propose two computational schemes to compute the corresponding saliency maps in a hierarchical segmentation space. Finally, we propose a framework to combine the two saliency maps, obtaining the final saliency map. Experimental results on three publicly available datasets show that even each single saliency map is able to reach the state-of-the-art performance. Furthermore, our framework, which combines two saliency maps, outperforms the state of the arts. Additionally, we show that the proposed framework can be easily used to extend existing methods and further improve their performances substantially. Qing Liu 0003, Xiaopeng Hong, Beiji Zou 0001, Jie Chen 0001, Zailiang Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Selective deep features for micro-expression recognitionabstractMicro-expression recognition is a challenging task in computer vision field due to the repressed facial appearance and short duration. Previous work for micro-expression recognition have used hand-crafted features like LBP-TOP, Gabor filter and optical flow. This paper is the first work to explore the possible use of deep learning for micro-expression recognition task. Due to the lack of data for micro-expression, training a CNN model from micro-expression data is not feasible. Instead, transfer learning from objects and facial expressions based CNN models are used. The aim is to use feature selection to remove the irrelevant deep features for our task. This work extends evolutionary algorithms to search an optimal set of deep features so that it does not overfit the training data and generalizes well for the test data. Promising results are presented for various micro-expression datasets. Devangini Patel, Xiaopeng Hong, Guoying Zhao 0001 |
ICPR | 2 |
| 2016 | Capturing correlations of local features for image representation
Xiaopeng Hong, Guoying Zhao 0001, Stefanos Zafeiriou, Maja Pantic, Matti Pietikäinen |
Neurocomputing | 1 |
| 2016 | Spontaneous facial micro-expression analysis using Spatiotemporal Completed Local Quantized Patterns
Xiaohua Huang 0003, Guoying Zhao 0001, Xiaopeng Hong, Wenming Zheng, Matti Pietikäinen |
Neurocomputing | 3 |
| 2016 | A unified 3D face authentication framework based on robust local mesh SIFT feature
Yue Ming 0001, Xiaopeng Hong |
Neurocomputing | 2 |
| 2016 | Dynamic texture and scene classification by transferring deep image features
Xianbiao Qi, Chun-Guang Li, Guoying Zhao 0001, Xiaopeng Hong, Matti Pietikäinen |
Neurocomputing | 4 |
| 2016 | Higher order partial least squares for object tracking: A 4D-tracking method
Bineng Zhong 0001, Xiangnan Yang, Yingju Shen, Cheng Wang 0020, Tian Wang 0001, Zhen Cui 0001, Hongbo Zhang 0002, Xiaopeng Hong, Duansheng Chen |
Neurocomputing | 8 |
| 2015 | A Task-Driven Eye Tracking Dataset for Visual Attention Analysis
Yingyue Xu, Xiaopeng Hong, Qiuhai He, Guoying Zhao 0001, Matti Pietikäinen |
ACIVS | 2 |
| 2014 | Improved Spatiotemporal Local Monogenic Binary Pattern for Emotion Recognition in The WildabstractLocal binary pattern from three orthogonal planes (LBP-TOP) has been widely used in emotion recognition in the wild. However, it suffers from illumination and pose changes. This paper mainly focuses on the robustness of LBP-TOP to unconstrained environment. Recent proposed method, spatiotemporal local monogenic binary pattern (STLMBP), was verified to work promisingly in different illumination conditions. Thus this paper proposes an improved spatiotemporal feature descriptor based on STLMBP. The improved descriptor uses not only magnitude and orientation, but also the phase information, which provide complementary information. In detail, the magnitude, orientation and phase images are obtained by using an effective monogenic filter, and multiple feature vectors are finally fused by multiple kernel learning. STLMBP and the proposed method are evaluated in the Acted Facial Expression in the Wild as part of the 2014 Emotion Recognition in the Wild Challenge. They achieve competitive results, with an accuracy gain of 6.35% and 7.65% above the challenge baseline (LBP-TOP) over video. Xiaohua Huang 0003, Qiuhai He, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
ICMI | 3 |
| 2014 | Pose Estimation via Complex-Frequency Domain Analysis of Image Gradient OrientationsabstractHead Pose Estimation (HPE) has recently attracted a lot of interests in various computer vision applications. One challenging problem for accurate HPE is to model the intrinsic variations among poses, and suppress the extraneous variations derived from other factors, such as the illumination changes, outliers, and noise. To this end, this paper proposes a simple and efficient facial description for head pose estimation from images. To handle the illumination changes, we characterize each image pixel by its image gradient orientation (IGO), rather than the intensity, which is sensitive to illumination changes. We then carry out complex-frequency domain analysis of the IGO image via the two-dimensional image transform, such as the 2D Discrete Cosine Transform (DCT2), to encode the spatial configuration of image gradient orientations. The proposed facial description is called IGO-DCT2. It is robust to illumination changes, outliers, and noise. In addition, it is learning free and computationally efficient. Finally, the fine-grain head pose estimation is regarded as a regression problem and off-the-shelf non-linear regression models are used to learn the mapping from the feature space to the continuous pose labels. Experimental results show the proposed facial description achieves highly competitive results on the publicly available FacePix dataset. Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 1 |
| 2014 | Structured partial least squares for simultaneous object tracking and segmentation
Bineng Zhong 0001, Xiao-Tong Yuan, Rongrong Ji, Yan Yan 0001, Zhen Cui 0001, Xiaopeng Hong, Yan Chen 0017, Tian Wang 0001, Duansheng Chen |
Neurocomputing | 6 |
| 2014 | A review of recent advances in visual speech decoding
Ziheng Zhou 0003, Guoying Zhao 0001, Xiaopeng Hong, Matti Pietikäinen |
Image Vis. Comput. | 3 |
| 2014 | A Compact Representation of Visual Speech Data Using Latent VariablesabstractThe problem of visual speech recognition involves the decoding of the video dynamics of a talking mouth in a high-dimensional visual space. In this paper, we propose a generative latent variable model to provide a compact representation of visual speech data. The model uses latent variables to separately represent the interspeaker variations of visual appearances and those caused by uttering within images, and incorporates the structural information of the visual data through placing priors of the latent variables along a curve embedded within a path graph. Ziheng Zhou 0003, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Combining LBP Difference and Feature Correlation for Texture DescriptionabstractEffective characterization of texture images requires exploiting multiple visual cues from the image appearance. The local binary pattern (LBP) and its variants achieve great success in texture description. However, because the LBP(-like) feature is an index of discrete patterns rather than a numerical feature, it is difficult to combine the LBP(-like) feature with other discriminative ones by a compact descriptor. To overcome the problem derived from the nonnumerical constraint of the LBP, this paper proposes a numerical variant accordingly, named the LBP difference (LBPD). The LBPD characterizes the extent to which one LBP varies from the average local structure of an image region of interest. It is simple, rotation invariant, and computationally efficient. To achieve enhanced performance, we combine the LBPD with other discriminative cues by a covariance matrix. The proposed descriptor, termed the covariance and LBPD descriptor (COV-LBPD), is able to capture the intrinsic correlation between the LBPD and other features in a compact manner. Experimental results show that the COV-LBPD achieves promising results on publicly available data sets. Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen, Xilin Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Combining local and global correlation for texture description
Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen, Xilin Chen 0001 |
ICPR | 1 |
| 2010 | Boosted Sigma Set for Pedestrian DetectionabstractThis paper presents a new method to detect pedestrian in still image using Sigma sets as image region descriptors in the boosting framework. Sigma set encodes second order statistics of an image region implicitly in the form of a point set. Compared with the covariance matrix, the traditional second order statistics based region descriptor, which requires computationally demanding operations based on Riemannian manifold, Sigma set preserves similar robustness and discriminative power more efficiently because the classification on Sigma sets can be directly performed in vector space. Experimental results on the INRIA and the Daimler Chrysler pedestrian datasets show the effectiveness and efficiency of the proposed method. Xiaopeng Hong, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001 |
ICPR | 1 |
| 2010 | A Sample Pre-mapping Method Enhancing Boosting for Object DetectionabstractWe propose a novel method to improve the training efficiency and accuracy of boosted classifiers for object detection. The key step of the proposed method is a sample pre-mapping on original space by referring to the selected `reference sample' before feeding into weak classifiers. The reference sample corresponds to an approximation of the optimal separating hyper-plane in an implicit high dimensional space, so that the resulting classifier could achieve the performance similar to kernel method, while spending the computation cost of linear classifier in both training and detection. We employ two different non-linear mappings to verify the proposed method under boosting framework. Experimental results show that the proposed approach achieves performance comparable with the common used methods on public datasets in both pedestrian detection and car detection. Xiaopeng Hong, Cherkeng Heng, Luhong Liang, Xilin Chen 0001 |
ICPR | 2 |
| 2010 | Sigma Set Based Implicit Online Learning for Object TrackingabstractThis letter presents a novel object tracking approach within the Bayesian inference framework through implicit online learning. In our approach, the target is represented by multiple patches, each of which is encoded by a powerful and efficient region descriptor called Sigma set. To model each target patch, we propose to utilize the online one-class support vector machine algorithm, named Implicit online Learning with Kernels Model (ILKM). ILKM is simple, efficient, and capable of learning a robust online target predictor in the presence of appearance changes. Responses of ILKMs related to multiple target patches are fused by an arbitrator with an inference of possible partial occlusions, to make the decision and trigger the model update. Experimental results demonstrate that the proposed tracking approach is effective and efficient in ever-changing and cluttered scenes. Xiaopeng Hong, Hong Chang 0001, Shiguang Shan, Bineng Zhong 0001, Xilin Chen 0001, Wen Gao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2009 | Sigma Set: A small second order statistical region descriptorabstractGiven an image region of pixels, second order statistics can be used to construct a descriptor for object representation. One example is the covariance matrix descriptor, which shows high discriminative power and good robustness in many computer vision applications. However, operations for the covariance matrix on Riemannian manifolds are usually computationally demanding. This paper proposes a novel second order statistics based region descriptor, named “Sigma Set”, in the form of a small set of vectors, which can be uniquely constructed through Cholesky decomposition on the covariance matrix. Sigma Set is of low dimension, powerful and robust. Moreover, compared with the covariance matrix, Sigma Set is not only more efficient in distance evaluation and average calculation, but also easier to be enriched with first order statistics. Experimental results in texture classification and object tracking verify the effectiveness and efficiency of this novel object descriptor. Xiaopeng Hong, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
CVPR | 1 |
| 2005 | An Information Acquiring Channel - Lip Movement
Xiaopeng Hong, Hongxun Yao, Qinghui Liu |
ACII | 1 |