Weihong Ren

dblp:177/7000 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0003-3839-0078ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 15 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different head poses, and so on. To address the above problems, current approaches rely on extensive learnable parameters and complex model architectures, which inevitably lead to overfitting and cause the FER model to focus on non-discriminative facial regions. In this work, we propose an HKAFER model that can adaptively enhance visual expression representations through efficiently fine-tuning the image encoder in large Visual Foundation Models (VFMs) and Vision-Language Models (VLMs). Specifically, we establish Heterogeneous Kronecker Adaptation (HeKA), which consists of multi-scale adapters based on Kronecker product in a parallel manner, offering significantly diverse subspaces to learn the incremental matrices. Besides, we also propose Dual-Branch Interactive Router (DBIR) to dynamically assign the weights of adapters, which promotes collaboration and information flow among them. In this way, our HKAFER can effectively capture robust spatial features and the regional associations. Experimental results demonstrate that our proposed model not only outperforms state-of-the-art methods on several FER benchmarks but also uses significantly fewer trainable parameters.
Yu Gao 0010, Haoyu Ji 0001, Zhiyong Wang 0009, Wenze Huang, Xueting Liu 0009, Weihong Ren, Honghai Liu 0001
AAAI8
2026 Domain Consistency Representation Learning for Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) exhibits a contradictory relationship between intra-domain discrimination and inter-domain gaps when learning from continuous data. Intra-domain discrimination focuses on individual nuances (i.e., clothing type, accessories,etc.), while inter-domain gaps emphasize domain consistency. Achieving a trade-off between maximizing intra-domain discrimination and minimizing inter-domain gaps is a crucial challenge for improving LReID performance. Most existing methods strive to reduce inter-domain gaps through knowledge distillation to maintain domain consistency. However, they often ignore intra-domain discrimination. To address this challenge, we propose a novel domain consistency representation learning (DCR) model that explores global and attribute-wise representations as a bridge to balance intra-domain discrimination and inter-domain gaps. At the intra-domain level, we explore the complementary relationship between global and attribute-wise representations to improve discrimination among similar identities. Excessive learning intra-domain discrimination can lead to catastrophic forgetting. We further develop an attribute-oriented anti-forgetting (AF) strategy that explores attribute-wise representations to enhance inter-domain consistency, and propose a knowledge consolidation (KC) strategy to facilitate knowledge transfer. Extensive experiments show that our DCR achieves superior performance compared to state-of-the-art LReID methods. Our code is available at https://github.com/LiuShiBen/DCR.
Shiben Liu, Huijie Fan, Qiang Wang 0015, Weihong Ren, Yandong Tang, Yang Cong
IEEE Trans. Circuits Syst. Video Technol.4
2026 Topology-Motion Decoupling Framework With Textual Regularization for Skeleton-Based Temporal Action Segmentation
abstract
Skeleton-based temporal action segmentation aims to capture key information in long skeleton motion sequences to temporally segment and identify actions at a fine-grained level. Existing approaches have achieved promising results by improving the modeling of topological spatial relationships and long-term temporal dependencies. However, current methods often overlook the distinct nature of motion and topological information, applying a monolithic modeling paradigm to both. This approach fails to fully exploit their differential contributions to precise boundary localization and effective class discrimination. To address these limitations, we propose a novel Topology-Motion Decoupling Framework (TMD). Our framework incorporates three key designs. First, an auxiliary Differential Motion Perception Branch explicitly models the temporal gradients of skeletal trajectory to decouple boundary-sensitive motion features. Second, we introduce two effective fusion modules that integrate the complementary features from both branches for mutual enhancement. Finally, a Boundary-Aware Textual Regularization scheme leverages a dual set of semantic prompts for boundary/non-boundary to differentially guide the feature learning process. By design, TMD explicitly mitigates semantic and temporal confusion between actions, thereby enhancing inter-class discriminability and boundary awareness. Extensive experiments on five challenging public datasets demonstrate that our TMD achieves state-of-the-art performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Context Modeling With Multimodal Prompts for Emotion Recognition in Conversation
abstract
Emotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. After the extensive exploration of the text modality, visual and audio information has attracted considerable attention. Most of the existing approaches adopt either the attention mechanism or graph neural networks to conduct multi-modal fusion by directly utilizing pre-extracted single modal features, but they ignore the inherent priors and characteristics (highly related to emotions) contained in different modalities. For example, for visual modality, facial expression can well reveal a person's emotional state, and the intonation, speech rate and volume of audio modality also reflect emotional fluctuation. Thus, in this work, we aim to promote context fusion among multi-modal features by exploring inherent modal priors. Firstly, we create textual descriptions for different emotion categories belonging to different modalities, denoted as multimodal prompts which can be generated either from common sense or using large language models (e.g., ChatGPT). Then, we propose an adaptive gating fusion module which dynamically learns the weights between unimodal features and prompts, and allows the multimodal prompts to participate in encoding and enriching modal information. Finally, to facilitate multimodal fusion, we design a multimodal progressive encoder to learn inter-modal interactions among conversational utterances. Experimental results show that our model outperforms state-of-the-art models in ERC on two popular benchmark datasets.
Weihong Ren, Yu Gao 0010, Jianzhuang Liu, Honghai Liu 0001
IEEE Trans. Multim.2
2025 A New Federated Learning Framework Against Gradient Inversion Attacks
abstract
Federated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety of privacy-preserving methods have been integrated into FL to thwart such attacks, such as Secure Multi-party Computing (SMC), Homomorphic Encryption (HE), and Differential Privacy (DP). Despite their ability to protect data privacy, these approaches inherently involve substantial privacy-utility trade-offs. By revisiting the key to privacy exposure in FL under GIA, which lies in the frequent sharing of model gradients that contain private data, we take a new perspective by designing a novel privacy preserve FL framework that effectively ``breaks the direct connection'' between the shared parameters and the local private data to defend against GIA. Specifically, we propose a Hypernetwork Federated Learning (HyperFL) framework that utilizes hypernetworks to generate the parameters of the local model and only the hypernetwork parameters are uploaded to the server for aggregation. Theoretical analyses demonstrate the convergence rate of the proposed HyperFL, while extensive experimental results show the privacy-preserving capability and comparable performance of HyperFL.
Pengxin Guo 0001, Shuang Zeng, Xiaodan Zhang 0003, Weihong Ren, Yuyin Zhou, Liangqiong Qu
AAAI5
2025 InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection
abstract
Recently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representations for HOIs from pretrained LFMs. However, such prompt-based approaches focus on transferring HOI-specific knowledge, but unexplore the potential reasoning capabilities of LFMs, which can provide informative context for ambiguous and open-world interaction recognition. In this paper, we propose InstructHOI, a novel method that leverages context-aware instructions to guide multi-modal reasoning for HOI detection. Specifically, to bridge knowledge gap and enhance reasoning abilities, we first perform HOI-domain fine-tuning on a pretrained multi-modal LFM, using a generated dataset with 140K interaction-reasoning image-text pairs. Then, we develop a Context-aware Instruction Generator (CIG) to guide interaction reasoning. Unlike traditional language-only instructions, CIG first mines visual interactive context at the human-object level, which is then fused with linguistic instructions, forming multi-modal reasoning guidance. Furthermore, an Interest Token Selector (ITS) is adopted to adaptively filter image tokens based on context-aware instructions, thereby aligning reasoning process with interaction regions. Extensive experiments on two public benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.
Jinguo Luo, Weihong Ren, Quanlong Zheng, Zhenlong Yuan, Zhiyong Wang 0009, Haonan Lu, Honghai Liu 0001
NeurIPS2
2025 JADFER: Exploring Spatial-Contextual Interaction With Joint Attention Dropping for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) aims to categorize emotional expressions depicted on a human face, and is a challenging task under unconstrained conditions, such as face occlusions and pose variations. Recent methods usually adopt self attention or cross attention to explore global or local relationships among different level features. However, these methods are inclined to focus on the redundant facial regions, causing model overfitting. To address this problem, we propose a new FER model named JADFER, which drops the joint attention in the weight matrix to adaptively enhance facial expression representations. Specifically, our JADFER model consists of three components: Spatial Branch (SB), Contextual Branch (CB), and Spatial-Contextual Interaction (SCI). First, SB runs$N$paths in parallel, where a Variety loss is designed to guide the paths of SB to focus on different discriminative regions. Meanwhile, CB abstracts the contextual facial representations using self attention with Joint Attention Dropping (JAD). Then, the SCI adopts the spatial features from SB to query the contextual representations from CB through cross attention with JAD, which regulates the attention weights by dropping the similar activations to further enhance the facial embeddings. Experimental results demonstrate that the proposed model outperforms the state-of-the-art methods on several FER benchmarks.
Yu Gao 0010, Weihong Ren, Weibo Jiang, Honghai Liu 0001
IEEE Trans. Affect. Comput.2
2025 Facial Expression Monitoring via Fine-Grained Vision-Language Alignment
abstract
In the fields of health care and clinical monitoring, vision-based Facial Expression Recognition (FER) has achieved significantly progress, but it still faces the challenge of poor generalization ability under unconstrained conditions of occlusions and pose variation. Recently, Vision-Language Model (VLM) has greatly advanced the FER task. However, the existing VLM-based FER methods typically leverage a hard-crafted prompt (e.g., “a photo of [class]”) and only focus on the holistic semantic alignment, which may suffer from modal heterogeneity. In this work, we propose a fine-grained vision-language model via Prompt Masking for FER (PMFER). Specifically, for each expression, we first create fine-grained prompts using facial action units to guide the image encoder to learn discriminative representations. Further, to finely align text prompts and visual action units, we randomly drop a phrase description in the prompts and then predict the dropped phrase by conducting modal cross attention, implicitly promoting fine-grained vision-language alignment. In addition, we also design a modal-adversarial strategy to holistically eliminate the modal difference between visual and textual embeddings in a common latent space. Experimental results demonstrate that our PMFER model outperforms the state-of-the-art methods on several FER benchmarks, especially under the conditions of occlusions and pose variations. Note to Practitioners—Facial expression recognition is very important in health care and clinical monitoring, which provides an useful tool to assess the psychological and physiological conditions of patients. Although FER has made significant progress with the development of deep learning technologies, it still faces problems in the complex environments (e.g., occlusions and pose variations). To address the above issues, we propose a novel FER method in this work based on the recent vision-language model. It takes RGB image and text prompts as input and finally predicts the expression classification. Different from the existing methods, the proposed PMFER can enable fine-grained modal alignment for facial key units. Compared with the state-of-the-art methods on the public datasets, it can achieve better results, especially under the conditions of occlusions and pose variations. Also, we evaluate the proposed method on a real-world pain dataset, and the results demonstrate that PMFER has a good generalization and can be applied to health care.
Weihong Ren, Yu Gao 0010, Xi'ai Chen, Zhi Han, Zhiyong Wang 0009, Jiaole Wang, Honghai Liu 0001
IEEE Trans Autom. Sci. Eng.1
2025 Multiscale Skeleton-Based Temporal Action Segmentation Using Hierarchical Temporal Modeling and Prediction Ensemble
abstract
Skeleton-based temporal action segmentation (TAS) decomposes untrimmed skeleton sequence into meaningful segments. The variance in temporal scale challenges the skeleton modeling network to seek a balance between over-segmentation and under-segmentation. Current methods often rely on parallel multiscale feature extractors and additional refinement modules to mitigate the multiscale issue, which brings significant computations and complexity. To address these issues, this article proposes multiscale skeleton-based TAS (MSTAS), consisting of temporal probability pyramid (TPP) and smoothed multiscale ensemble (SME). TPP represents each action as a collection of multiscale probability distributions using a U-shape hierarchical temporal pyramid. Subsequently, SME takes the average of distributions instead of deploying additional refinement stages to achieve action segmentation. Considering the over-confident issue that exists in each scale, SME incorporates a novel label smoothing phase to improve the probability distributions by dynamically calibrating the confidence of each scale. Experimental results on four public datasets show that the MSTAS achieves state-of-the-art performance with less computation overheads, such as +1.1% accuracy and +2.8% [email protected] on the challenging LARa dataset with 70% fewer parameters and 80% fewer GFLOPS. Benefiting from confidence calibration, the MSTAS efficiently utilizes more temporal scales while keeping better calibration for ambiguous action instances. Additionally, the U-shape pyramid demonstrates a strong compatibility with classical refinement module, enabling the efficient extraction of multiscale motion representations.
Bowen Chen 0004, Haoyu Ji 0001, Weihong Ren, Qiyi Tong, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Cybern.4
2025 Interaction-Aware Transformer Network for Human-Object Interaction Detection
abstract
human-object interaction (HOI) detection tackles the problem of joint localization and classification of HOIs. Recent HOI detection methods are mainly based on transformer networks, where the explicit priors at the object level (e.g., scene layout, object appearance, or category) are usually fed into the transformer to improve the object query ability. Though these methods have achieved remarkable results, they did not pay enough attention to the implicit action-level information, which is the fundamental element of HOI. In this work, we propose an interaction-aware transformer network (IATN) to obtain the interaction-aware query, by jointly utilizing implicit action-level priors and explicit object-level priors. Specifically, we design an action-aware module (AAM) to aggregate implicit action priors from the scene level and instance level, respectively. Then, we design an action-oriented graph (AOG), where human feature and object feature are graph nodes and action semantics represent graph edges, to aggregate priors jointly from action level and object level. Afterwards, the interaction-aware query is acquired and finally adopted to obtain the HOI predictions. Besides, we leverage knowledge distillation to enhance the action-level priors by transferring the final HOI predictions to the intermediate features. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed interaction-aware model.
Weibo Jiang, Weihong Ren, Jiandong Tian, Hanwei Ma, Bowen Chen 0004, Honghai Liu 0001
IEEE Trans. Cybern.2
2025 LFUID: Light Field-Based Underwater Image Formation, Restoration, and Real-World Dataset
abstract
Underwater imaging in turbid environments presents significant challenges for industrial applications due to optical distortions that severely degrade image quality. This article addresses the fundamental information deficit in traditional single-image restoration methods by introducing the first comprehensive Light Field Underwater Image Dataset, comprising 1356 images captured across diverse environments with 1820× 720× 14× 14× 3 resolution. We develop a novel physics-based image formation model that extends scattering principles to the 4-D light field domain, incorporating water attenuation coefficients and scattering properties while accounting for light field camera characteristics. Our model features three key components: an ambient underwater optical constant, a phase function for angular scattering, and a pixel position modulation function for spatial variations in backscatter. Based on this model, we propose a restoration algorithm that leverages complementary information across subaperture views. Experimental results demonstrate that our approach significantly outperforms state-of-the-art underwater image enhancement techniques across multiple metrics and diverse underwater scenes, establishing a promising new direction for underwater imaging applications.
Shijun Zhou, Yiming Su, Doneyue Wang, Weihong Ren, Jiandong Tian
IEEE Trans. Ind. Informatics4
2025 Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
abstract
Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish dependencies among joints as well as frames, and utilize one-hot encoding with cross-entropy loss for frame-wise classification supervision. However, these methods overlook the intrinsic correlations among joints and actions within skeletal features, leading to a limited understanding of human movements. To address this, we propose a Text-Derived Relational Graph-Enhanced Network (TRG-Net) that leverages prior graphs generated by Large Language Models (LLM) to enhance both modeling and supervision. For modeling, the Dynamic Spatio-Temporal Fusion Modeling (DSFM) method incorporates Text-Derived Joint Graphs (TJG) with channel- and frame-level dynamic adaptation to effectively model spatial relations, while integrating spatio-temporal core features during temporal modeling. For supervision, the Absolute-Relative Inter-Class Supervision (ARIS) method employs contrastive learning between action features and text embeddings to regularize the absolute class distributions, and utilizes Text-Derived Action Graphs (TAG) to capture the relative inter-class relationships among action features. Additionally, we propose a Spatial-Aware Enhancement Processing (SAEP) method, which incorporates random joint occlusion and axial rotation to enhance spatial generalization. Performance evaluations on four public datasets demonstrate that TRG-Net achieves state-of-the-art results.
Haoyu Ji 0001, Bowen Chen 0004, Weihong Ren, Wenze Huang, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Image Process.3
2025 Synergistic Prompting Learning for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection, as a foundational task in human-centric understanding, aims to detect interactive triplets in real-world scenarios. To better distinguish diverse HOIs within an open-world context, current HOI detectors utilize pre-trained Visual-Language Models (VLMs) to extract prior knowledge through textual prompts (i.e., descriptive texts for each HOI instance). However, relying on predetermined descriptive texts, such approaches only acquire a fixed set of textual knowledge for HOI prediction, consequently resulting in inferior performance and limited generalization. To remedy this, we propose a novel VLM-based method, which jointly performs prompting learning from both visual and textual perspectives and synergizes visual-textual prompting for HOI detection. Initially, we design a hierarchical adaptation architecture to perform progressive prompting: visual prompting is facilitated through gradual token migration from VLM's image encoder, while textual prompting is initialized with progressively leveled interaction descriptions. In addition, to synergize the visual-textual prompting learning, a text-supervising and image-tuning loop is introduced, in which the text-supervising stage guides visual prompting learning through contrastive learning and the image-tuning stage refines textual prompting by modal matching. Finally, we employ an interaction-aware knowledge merging mechanism to effectively transfer visual-textual knowledge encapsulated within synergistic prompting for HOI detection. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.
Jinguo Luo, Weihong Ren, Zhiyong Wang 0009, Xi'ai Chen, Huijie Fan, Zhi Han, Honghai Liu 0001
IEEE Trans. Image Process.2
2025 Iris Geometric Transformation Guided Deep Appearance-Based Gaze Estimation
abstract
The geometric alterations in the iris's appearance are intricately linked to the gaze direction. However, current deep appearance-based gaze estimation methods mainly rely on latent feature sharing to leverage iris features for improving deep representation learning, often neglecting the explicit modeling of their geometric relationships. To address this issue, this paper revisits the physiological structure of the eyeball and introduces a set of geometric assumptions, such as "the normal vector of the iris center approximates the gaze direction". Building on these assumptions, we propose an Iris Geometric Transformation Guided Gaze estimation (IGTG-Gaze) module, which establishes an explicit geometric parameter sharing mechanism to link gaze direction and sparse iris landmark coordinates directly. Extensive experimental results demonstrate that IGTG-Gaze seamlessly integrates into various deep neural networks, flexibly extends from sparse iris landmarks to dense eye mesh, and consistently achieves leading performance in both within- and cross-dataset evaluations, all while maintaining end-to-end optimization. These advantages highlight IGTG-Gaze as a practical and effective approach for enhancing deep gaze representation from appearance.
Zhiyong Wang 0009, Weihong Ren, Honghai Liu 0001
IEEE Trans. Image Process.3
2025 Snippet-Aware Transformer With Multiple Action Elements for Skeleton-Based Action Segmentation
abstract
The skeleton-based temporal action segmentation (STAS) aims to densely segment and classify human actions within lengthy untrimmed skeletal motion sequences. Current methods primarily rely on graph convolutional networks (GCNs) for intraframe spatial modeling and temporal convolutional networks (TCNs) for interframe temporal modeling to discern motion patterns. However, these approaches often overlook the distinctive nature of essential action elements across various actions, including engaged core body parts and key subactions. This oversight limits the ability to distinguish different actions within a given sequence. To address these limitations, the snippet-aware Transformer with multiple action element (ME-ST) is proposed to enhance the discrimination and segmentation among actions, which leverages intrasnippet attention along joints and sequences to identify core joints and key subactions at different scales. Specifically, in terms of the spatial domain, the intrasnippet cross-joint attention (CJA) module divides the sequence into distinct snippets and computes attention to establish intricate joint semantic relationships, emphasizing the identification of core motion joints. In terms of the temporal domain, in the encoder, the intrasnippet cross-frame attention (CFA) module segments the sequence in a blockwise expansion manner and establishes interframe relationships to highlight the most discriminative frames. In the decoder, clip-level representations at various temporal scales are initially generated through an hourglass-like sampling process, followed by the intrasnippet cross-scale attention (CSA) module to integrate the key clip information across different time scales. The performance evaluation on five public datasets demonstrates that ME-ST achieves state-of-the-art (SOTA) performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplets. However, most of these methods focus on seeking for self-triplet aggregation, but ignore the potential cross-triplet dependencies, resulting in ambiguity of action prediction. In this work, we propose to explore Self- and Cross-Triplet Correlations (SCTC) for HOI detection. Specifically, we regard each triplet proposal as a graph where Human, Object represent nodes and Action indicates edge, to aggregate self-triplet correlation. Also, we try to explore cross-triplet dependencies by jointly considering instance-level, semantic-level, and layout-level relations. Besides, we leverage the CLIP model to assist our SCTC obtain interaction-aware feature by knowledge distillation, which provides useful action clues for HOI detection. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed SCTC.
Weibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu, Zhiyong Wang 0009, Honghai Liu 0001
AAAI2
2024 Discovering Syntactic Interaction Clues for Human-Object Interaction Detection
abstract
Recently, Vision-Language Model (VLM) has greatly ad-vanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g., a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However, such approaches, only encoding the action-specific text prompts in vocabulary level, may suffer from learning ambiguity without exploring the fine-grained clues from the perspective of interaction context. In this paper, we propose a novel method to discover Syntactic Interaction Clues for HOI detection (SICHOI) by using VLM. Specifically, we first investigate what are the essen-tial elements for an interaction context, and then establish a syntactic interaction bank from three levels: spatial relationship, action-oriented posture and situational condition. Further, to align visual features with the syntactic interaction bank, we adopt a multi-view extractor to jointly aggre-gate visual features from instance, interaction, and image levels accordingly. In addition, we also introduce a dual cross-attention decoder to perform context propagation be-tween text knowledge and visual features, thereby enhancing the HOI detection. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on HICO-DET and V-COCO.
Jinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen, Qiang Wang 0015, Zhi Han, Honghai Liu 0001
CVPR2
2024 Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation
Haoyu Ji 0001, Bowen Chen 0004, Xinglong Xu, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
ECCV (54)4
2024 MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition
abstract
In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Language Model for Emotion Recognition to learn discriminative representations adaptively. Specifically, different from typically leveraging a hand-crafted prompt (e.g., "a photo of a [class] person"), we first establish Multi-Level Prompts from three aspects: facial expression, human posture and situational condition using large language models, like ChatGPT. Correspondingly, we extract the visual tokens from three levels: the face, body, and context. Further, to achieve fine-grained alignment at each level, we adopt textual tokens from the positive and the hard negative to query visual tokens, predicting whether a pair of image and text is matched. Experimental results demonstrate that our MLPER model outperforms the state-of-the-art methods on several ER benchmarks, especially under the conditions of occlusions and pose variations.
Yu Gao 0010, Weihong Ren, Xinglong Xu, Zhiyong Wang 0009, Honghai Liu 0001
IROS2
2024 GroupTrack: Multi-Object Tracking by Using Group Motion Patterns
abstract
The main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effectively solve long-term tracking due to individually motion or appearance modeling for each single target. In this paper, we propose a novel 2D MOT tracker GroupTrack, to learn reliable motion state for each target using group motion patterns. Specifically, for each tracklet, we first choose its neighboring ones to form a group of motion patterns, which can provide informative clues for the motion estimation of the current tracklet. Then, we apply the group motion patterns to perform tracklet prediction and data association. By integrating prior from neighboring motion patterns into the data association process, GroupTrack provides a new paradigm for target motion modeling in extremely crowded and occluded scenarios. Through extensive experiments on the public MOT17 and MOT20 datasets, we demonstrate the effectiveness of our approach in challenging scenarios and show state-of-the-art performance at various MOT metrics.
Xinglong Xu, Weihong Ren, Gan Sun, Haoyu Ji 0001, Yu Gao 0010, Honghai Liu 0001
IROS2
2024 Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open World
abstract
Universal object detectors aim to detect any object in any scene without human annotation, exhibiting superior generalization. However, the current universal object detectors show degraded performance in harsh weather, and their insufficient real-time capabilities limit their application. In this paper, we present Uni-YOLO, a universal detector designed for complex scenes with real-time performance. Uni-YOLO is a one-stage object detector that uses general object confidence to distinguish between objects and backgrounds, and employs a grid cell regression method for real-time detection. To improve its robustness in harsh weather conditions, the input of Uni-YOLO is adaptively enhanced with a physical model-based enhancement module. During training and inference, Uni-YOLO is guided by the extensive knowledge of the vision-language model CLIP. An object augmentation method is proposed to improve generalization in training by utilizing multiple source datasets with heterogeneous annotations. Furthermore, an online self-enhancement method is proposed to allow Uni-YOLO to further focus on specific objects through self-supervised fine-tuning in a given scene. Extensive experiments on public benchmarks and a UAV deployment are conducted to validate its superiority and practical value.
Weihong Ren, Xi'ai Chen, Huijie Fan, Yandong Tang, Zhi Han
ACM Multimedia2
2024 Learning Self- and Cross-Triplet Context Clues for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection aims to infer interactions between humans and objects, and it is very important for scene analysis and understanding. The existing methods usually focus on exploring instance-level (e.g., object appearance) or interaction-level (e.g., action semantic) features to conduct interaction prediction. However, most of these methods only consider the self-triplet feature aggregation, which may lead to learning ambiguity without exploring the cross-triplet context exchange. In this paper, from both visual and textual perspectives, we propose a novel method to jointly explore self-and cross-triplet interaction context clues for HOI detection. First, we employ a graph neural network to perform self-triplet aggregation, where human and object features represent graph nodes and visual interaction feature and textual prior knowledge are acted as two different edges. Furthermore, we also attempt to explore cross-triplet context exchange by incorporating symbiotic and layout relationships among different HOI triplets. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones and achieves the impressive performance of 40.32 mAP on HICO-DET and 69.1 mAP on V-COCO datasets, respectively.
Weihong Ren, Jinguo Luo, Weibo Jiang, Liangqiong Qu, Zhi Han, Jiandong Tian, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Dual Regression-Enhanced Gaze Target Detection in the Wild
abstract
Gaze is a vital feature in analyzing natural human behavior and social interaction. Existing gaze target detection studies learn gaze from gaze orientations and scene cues via a neural network to model gaze in unconstrained scenes. Though achieve decent accuracy, these studies either employ complex model architectures or leverage additional depth information, which limits the model application. This article proposes a simple and effective gaze target detection model that employs dual regression to improve detection accuracy while maintaining low model complexity. Specifically, in the training phase, the model parameters are optimized under the supervision of coordinate labels and corresponding Gaussian-smoothed heatmap labels. In the inference phase, the model outputs the gaze target in the form of coordinates as prediction rather than heatmaps. Extensive experimental results on within-dataset and cross-dataset evaluations on public datasets and clinical data of autism screening demonstrate that our model has high accuracy and inference speed with solid generalization capabilities.
Zhiyong Wang 0009, Weihong Ren, Xiu Xu, Honghai Liu 0001
IEEE Trans. Cybern.6
2023 WSCFER: Improving Facial Expression Representations by Weak Supervised Contrastive Learning
abstract
The major challenge of Facial Expression Recog-nition (FER) is to learn class discriminative representations, and the existing works mainly address it by designing various classification networks from class level. However, learning representations at class level is limited due to the inconspicuous class discrimination among different facial expressions. Thus, in this paper, we propose a Weak Supervised Contrastive learning FER (WSCFER) method to improve facial expression representations by simultaneously learning instance-level representations which are highly complementary to the general class-level representations. Specifically, our proposed WSCFER consists of three components: a major task for FER classification, an auxiliary task for Weak Supervised Contrastive (WSC) learning which pulls augmented samples of the same image together while pushing apart instance samples from different classes, and a Partial Consistency Loss (PCL) for optimizing the two embedding spaces from both the class level and the instance level. We compare WSC with some state-of-the-art contrastive methods and find that it can efficiently learn instance-level representations but avoid overemphasizing irrelevant parts, which is crucial for FER. WSCFER achieves superior performance on several in-the-wild databases, and it also shows the promising potential for learning representations under noisy annotations.
Bowen Chen 0004, Xiu Xu, Weihong Ren, Honghai Liu 0001
IROS5
2023 Multi-Scale Attention Learning Network for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) aims to identify emotional expressions in human faces, and it is a fundamental task in computer vision. Recently, some methods apply Vision Transformer (ViT) to FER and have achieved promising results. However, FER still suffers from two key issues: inter-class similarity and intra-class discrepancy. To address the issues, in this letter, we propose a Multi-Scale Attention Learning Network (MALN) based on ViT, which can learn facial expression embeddings in a multi-scale manner. Specifically, we adopt a multi-branch ViT architecture to adaptively explore multi-scale correlations without self-attention. Furthermore, we also design a Scale Distinction Loss (SDL) to dynamically regulate facial embeddings from multiple branches, which can guide ViT to capture discriminative facial regions. Experimental results on three public datasets (inluding RAF-DB, AffectNet and FERPlus) demonstrate the effectiveness of our proposed MALN for FER
Weihong Ren, Yu Gao 0010, Weibo Jiang, Honghai Liu 0001
IEEE Signal Process. Lett.2
2022 AutoENP: An Auto Rating Pipeline for Expressing Needs via Pointing Protocol
abstract
Early screening for ASD (Autism Spectrum Disorder) is crucial and also challenging due to the limited medical resource. Expressing Needs with Pointing (ENP) is a low-cost yet effective protocol for early screening. However, the current methods need to manually trim video for analyzing ENP protocol, which is labour-intensive. Also, they detect discriminative signs with separately high-level clues (e.g., pose, object detection), but ignore the temporal action relationships between child and clinician, which usually leads to invalid detection. In contrast to previous approaches, we propose an Auto Rating Pipeline for Expressing Needs via Pointing Protocol, named AutoENP. Specifically, we introduce action segmentation into early screening, to capture temporal interaction relationships without manually intervention. To detect fine-grained hand motions, we fuse global, local and fine-grained features to fully understand the screening scene. Besides, we integrate focal loss and center loss to improve the detection accuracy for rare actions. To evaluate the proposed pipeline, we collected 22 ENP videos containing 7 actions with above 40,000 frames. Experimental results demonstrate that our model achieves 82.1% and 84.7% action accuracy for child and clinician, respectively. Moreover, 18 in 22 children’s ENP levels are reported correctly against the clinician’s diagnoses.
Bowen Chen 0004, Weihong Ren, Honghai Liu 0001, Huiping Li 0004, Xiu Xu, Bingrui Zhou
ICPR2
2021 Tracking-by-Counting: Using Network Flows on Crowd Density Maps for Tracking Multiple Targets
abstract
State-of-the-art multi-object tracking (MOT) methods follow the tracking-by-detection paradigm, where object trajectories are obtained by associating per-frame outputs of object detectors. In crowded scenes, however, detectors often fail to obtain accurate detections due to heavy occlusions and high crowd density. In this paper, we propose a new MOT paradigm, tracking-by-counting, tailored for crowded scenes. Using crowd density maps, we jointly model detection, counting, and tracking of multiple targets as a network flow program, which simultaneously finds the global optimal detections and trajectories of multiple targets over the whole video. This is in contrast to prior MOT methods that either ignore the crowd density and thus are prone to errors in crowded scenes, or rely on a suboptimal two-step process using heuristic density-aware point-tracks for matching targets. Our approach yields promising results on public benchmarks of various domains including people tracking, cell tracking, and fish tracking.
Weihong Ren, Xinchao Wang, Jiandong Tian, Yandong Tang, Antoni B. Chan
IEEE Trans. Image Process.1
2021 Recurrent Generative Adversarial Network for Face Completion
abstract
Most recently-proposed face completion algorithms use high-level features extracted from convolutional neural networks (CNNs) to recover semantic texture content. Although the completed face is natural-looking, the synthesized content still lacks lots of high-frequency details, since the high-level features cannot supply sufficient spatial information for details recovery. To tackle this limitation, in this paper, we propose aRecurrentGenerativeAdversarialNetwork (RGAN) for face completion. Unlike previous algorithms, RGAN can take full advantage of multi-level features, and further provide advanced representations from multiple perspectives, which can well restore spatial information and details in face completion. Specifically, our RGAN model is composed of a CompletionNet and a DisctiminationNet, where the CompletionNet consists of two deep CNNs and a recurrent neural network (RNN). The first deep CNN is presented to learn the internal regulations of a masked image and represent it with multi-level features. The RNN model then exploits the relationships among the multi-level features and transfers these features in another domain, which can be used to complete the face image. Benefiting from bidirectional short links, another CNN is used to fuse multi-level features transferred from RNN and reconstruct the face image in different scales. Meanwhile, two context discrimination networks in the DisctiminationNet are adopted to ensure the completed image consistency globally and locally. Experimental results on benchmark datasets demonstrate qualitatively and quantitatively that our model performs better than the state-of-the-art face completion models, and simultaneously generates realistic image content and high-frequency details. The code will be released available soon.
Qiang Wang 0015, Huijie Fan, Gan Sun, Weihong Ren, Yandong Tang
IEEE Trans. Multim.4
2020 Dually Connected Deraining Net Using Pixel-Wise Attention
abstract
Recent single image deraining methods either use a recurrent mechanism to gradually learn the mapping between clear images and rainy images, or focus on designing various loss functions to supervise the learning process. In this letter, we propose a dually connected deraining net using pixel-wise attention, for single image rain removal. Specifically, the deraining net adopts an encoder-decoder net as a backbone, which can effectively learn a residual rain-streaks map by jointly using skip sum connection and skip concatenation connection. The dual connections enable the deraining net to promote information flow between layers, and thus can allow it to discriminate and localize the rain streaks. To preserve image details, the decoded features are weighted by the learnable pixel-wise attention for adaptively recalibrating their responses. Experimental results on synthetic datasets demonstrate that the proposed model outperforms the recent state-of-the-art deraining methods.
Weihong Ren, Jiandong Tian, Qiang Wang 0015, Yandong Tang
IEEE Signal Process. Lett.1
2018 Fusing Crowd Density Maps and Visual Object Trackers for People Tracking in Crowd Scenes
abstract
While visual tracking has been greatly improved over the recent years, crowd scenes remain particularly challenging for people tracking due to heavy occlusions, high crowd density, and significant appearance variation. To address these challenges, we first design a Sparse Kernelized Correlation Filter (S-KCF) to suppress target response variations caused by occlusions and illumination changes, and spurious responses due to similar distractor objects. We then propose a people tracking framework that fuses the S-KCF response map with an estimated crowd density map using a convolutional neural network (CNN), yielding a refined response map. To train the fusion CNN, we propose a two-stage strategy to gradually optimize the parameters. The first stage is to train a preliminary model in batch mode with image patches selected around the targets, and the second stage is to fine-tune the preliminary model using the real frame-by-frame tracking process. Our density fusion framework can significantly improves people tracking in crowd scenes, and can also be combined with other trackers to improve the tracking performance. We validate our framework on two crowd video datasets.
Weihong Ren, Yandong Tang, Antoni B. Chan
CVPR1
2018 Snowflake Removal for Videos via Global and Local Low-Rank Decomposition
abstract
Falling snow not only blocks human vision, but also significantly degrades the effectiveness of computer vision systems in outdoor environment. In this paper, we aim to remove snowflakes in videos by using the global and local low-rank property of snowflake-removed scenes. The stationary background and the mixture of moving foreground as well as falling snowflake are extracted via the global low-rank matrix decomposition. Some snowflake features, such as its color and size, are used to separate out the snowflakes from other moving objects. Then, the mean absolute difference based patch matching is applied to align every same moving object over frames to grab its low-rank structure. As such, the falling snowflake in front of moving objects can be removed via the local low-rank decomposition. Finally, the snowflake removed videos are generated by pasting moving foreground to stationary backgrounds. Experiments show that our method can remove snowflakes effectively and outperforms the comparison methods.
Jiandong Tian, Zhi Han, Weihong Ren, Xiai Chen, Yandong Tang
IEEE Trans. Multim.3
2017 Video Desnowing and Deraining Based on Matrix Decomposition
abstract
The existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snowflakes/rain streaks. In this paper, we propose a model based on matrix decomposition for video desnowing and deraining to solve the problems mentioned above. We divide snowflakes/rain streaks into two categories: sparse ones and dense ones. With background fluctuations and optical flow information, the detection of moving objects and sparse snowflakes/rain streaks is formulated as a multi-label Markov Random Fields (MRFs). As for dense snowflakes/rain streaks, they are considered to obey Gaussian distribution. The snowflakes/rain streaks, including sparse ones and dense ones, in scene backgrounds are removed by low-rank representation of the backgrounds. Meanwhile, a group sparsity term in our model is designed to filter snow/rain pixels within the moving objects. Experimental results show that our proposed model performs better than the state-of-the-art methods for snow and rain removal.
Weihong Ren, Jiandong Tian, Zhi Han, Antoni B. Chan, Yandong Tang
CVPR1
2017 Specular Reflection Separation With Color-Lines Constraint
abstract
According to dichromatic reflection model, the previous methods of specular reflection separation in image processing often separate specular reflection from a single image using patch-based priors. Due to lack of global information, these methods often cannot completely separate the specular component of an image and are incline to degrade image textures. In this paper, we derive a global color-lines constraint from dichromatic reflection model to effectively recover specular and diffuse reflection. Our key observation is from that each image pixel lies along a color line in normalized RGB space and the different color lines representing distinct diffuse chromaticities intersect at one point, namely, the illumination chromaticity. For pixels along the same color line, they spread over the entire image and their distances to the illumination chromaticity reflect the amount of specular reflection components. With global (non-local) information from these color lines, our method can effectively separate specular and diffuse reflection components in a pixelwise way for a single image, and it is suitable for real-time applications. Our experimental results on synthetic and real images show that our method performs better than the state-of-the-art methods to separate specular reflection.
Weihong Ren, Jiandong Tian, Yandong Tang
IEEE Trans. Image Process.1
2016 Blind Deconvolution With Nonlocal Similarity and l0 Sparsity for Noisy Image
abstract
The blind image deconvolution techniques with sparsity prior in gradient domain are sensitive to noise, even a small amount of noise. To address this problem, in this letter, we propose a novel blind deconvolution model that combines low-rank property, nonlocal similarity, and l0sparsity prior. Low-rank property makes the proposed deblurring model robust to image noise. The joint utilization of nonlocal similarity and l0sparsity prior has improved the accuracy of blur kernel estimation and restores the fine image details. A numerical method is also given to solve the proposed problem. Experimental results on synthetic and real data show that our algorithm performs better against with the state-of-the-art methods for both noise and noise-free images.
Weihong Ren, Jiandong Tian, Yandong Tang
IEEE Signal Process. Lett.1