Lei Chen 0069

dblp:09/3666-69 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-4279-3892ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Bi-Handover: A Unified Vision-Based Paradigm for Reliable Bidirectional Human-Robot Object Handover
abstract
Reliable object handover between humans and robots represents a fundamental capability for collaborative robotic systems. However, the diversity of human hand poses and object properties often leads to unstable grasping and unsafe interactions, posing significant challenges for robust human-robot collaboration. To address these, we propose Bi-Handover, a novel paradigm that enables bidirectional, reliable object transfer by constructing stable, safe intermediate handover states. Our framework maps human hand postures to parallel gripper grasping configurations with equivalent grasping capabilities. The diverse grasp patterns employed by both the object giver and receiver generate multiple intermediate handover states that critically determine task success. We evaluate the robustness of these intermediate states using an integrated methodology that combines grasp stability prediction with safety quantification, ultimately selecting the optimal state to ensure reliable handover performance. Bi-Handover is the first approach to achieve reliable bidirectional handover of arbitrary objects, demonstrating substantial performance gains over existing baselines through extensive experimental validation.
Ziwei Wang 0010, Lei Chen 0069, Jie Zhou 0001, Jiwen Lu
IEEE Trans Autom. Sci. Eng.3
2026 Efficient Arbitrary-Scale Super-Resolution With Compact Gaussian Splatting
abstract
Arbitrary-scale super-resolution is an essential image upsampling task, typically tackled with implicit neural representations. Unlike INR-based methods that rely on slow per-pixel decoding, Gaussian Splatting is promising for arbitrary-scale super-resolution when the explicit region-based nature of GS allows for highly efficient rendering via a lightweight decoder. Recently, Gaussian splatting has outperformed implicit neural methods in 3D scenes. However, existing attempts to address the super-resolution problem with Gaussian splatting face efficiency and accuracy challenges, such as redundant Gaussian primitives and discrete pixel sampling. Efficient Gaussian applications may result in continuous texture constraints due to limited feature richness in explicit fields, particularly the mismatch between the learned Gaussian fields and out-of-distribution sampling rates. Moreover, insufficient discrete sampling based on the given upscale factor may fail to accurately represent the splatted Gaussian field in screen space, causing high-frequency signal redundancy and aliasing. To address these challenges, we propose a compact Gaussian splatting method for efficient arbitrary-scale super-resolution, CGSSR. It constructs an efficient Gaussian space by distilling the image into a reduced number of primitives, each represented by a compact, low-dimensional feature embedding. To balance detail preservation and anti-aliasing, we introduce a scale-aware smoothing filter to regulate splatting frequency. Extensive experiments show that CGSSR achieves superior performance over existing Gaussian-based methods, especially at large scales, with higher efficiency.
Jingyi Zhang 0008, Jiajun Dong, Shuai Shen, Yansong Tang, Lei Chen 0069, Jiwen Lu
IEEE Trans. Circuits Syst. Video Technol.6
2026 MSP-Grasp: Multiscale Perceptual Framework for 6-DoF Grasping in Cluttered Environments
abstract
Autonomous grasping in cluttered environments represents one of the most challenging problems in robotic manipulation. The presence of occlusions and densely packed objects significantly complicates perception and substantially increases the risk of collisions. To address these challenges, we present a novel multiscale perception-based framework for robust robotic grasping in dense, cluttered environments. Our methodology adopts a multiscale progressive perception architecture: First, a global context awareness module analyzes the distribution of viable grasping opportunities, assesses collision risks, and evaluates spatial optimization potential to systematically identify optimal grasping regions. Second, a regional collision prediction module provides intermediate-scale analysis, effectively reducing collision incidents through enhanced spatial awareness. Finally, a local grasp evaluation module refines grasp selection by optimizing stability metrics and predicting the probability of grasp success. Comprehensive experiments across simulated environments and real-world scenarios demonstrate that our approach achieves substantially superior performance compared to existing baselines, confirming its effectiveness and applicability.
Ziwei Wang 0010, Lei Chen 0069, Jie Zhou 0001, Jiwen Lu
IEEE Trans. Ind. Informatics3
2026 Toward Generalizable Forgery Detection and Reasoning
abstract
Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a generalizable forgery detection model. Moreover, since every pixel in an AI-generated image is synthesized, traditional saliency-based forgery explanation methods are not well suited for this task. To address these challenges, we formulate detection and explanation as a unified Forgery Detection and Reasoning task (FDR-Task), leveraging Multi-Modal Large Language Models (MLLMs) to provide accurate detection through reliable reasoning over forgery attributes. To facilitate this task, we introduce the Multi-Modal Forgery Reasoning dataset (MMFR-Dataset), a large-scale dataset containing 120K images across 10 generative models, with 378K reasoning annotations on forgery attributes, enabling comprehensive evaluation of the FDR-Task. Furthermore, we propose FakeReasoning, a forgery detection and reasoning framework with three key components: 1) a dual-branch visual encoder that integrates CLIP and DINO to capture both high-level semantics and low-level artifacts; 2) a Forgery-Aware Feature Fusion Module that leverages DINO's attention maps and cross-attention mechanisms to guide MLLMs toward forgery-related clues; 3) a Classification Probability Mapper that couples language modeling and forgery detection, enhancing overall performance. Experiments across multiple generative models demonstrate that FakeReasoning not only achieves robust generalization but also outperforms state-of-the-art methods on both detection and reasoning tasks. The code is available at: https://github.com/PRIS-CV/FakeReasoning.
Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin, Muxi Diao, Lei Chen 0069, Kongming Liang, Zhanyu Ma
IEEE Trans. Image Process.6
2025 FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
abstract
Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional approaches using image diffusion models fall short in handling video dynamics, particularly for challenging temporal edits like motion adjustments. While current video diffusion models produce high-quality results, adapting them for efficient editing remains difficult due to the heavy computational demands that prevent the direct application of previous image editing techniques. To overcome these limitations, we introduce FADE—a training-free yet highly effective video editing approach that fully leverages the inherent priors from pre-trained video diffusion models via frequency-aware factorization. Rather than simply using these models, we first analyze the attention patterns within the video model to reveal how video priors are distributed across different components. Building on these insights, we propose a factorization strategy to optimize each component’s specialized role. Furthermore, we devise spectrum-guided modulation to refine the sampling trajectory with frequency domain cues, preventing information leakage and supporting efficient, versatile edits while preserving the basic spatial and temporal structure. Extensive experiments on real-world videos demonstrate that our method consistently delivers high-quality, realistic and temporally coherent editing results both qualitatively and quantitatively. Code is available at https://github.com/EternalEvan/FADE.
Yixuan Zhu, Haolin Wang 0006, Shilin Ma, Wenliang Zhao, Yansong Tang, Lei Chen 0069, Jie Zhou 0001
CVPR6
2025 D3QE: Learning Discrete Distribution Discrepancy-Aware Quantization Error for Autoregressive-Generated Image Detection
Yanran Zhang, Bingyao Yu, Yu Zheng 0015, Wenzhao Zheng, Yueqi Duan, Lei Chen 0069, Jie Zhou 0001, Jiwen Lu
ICCV6
2025 Learning Counterfactually Decoupled Attention for Open-World Model Attribution
abstract
In this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region partitioning or feature space, which could be confounded by the spurious statistical correlations and struggle with novel attacks in open-world scenarios. To address this, CDAL explicitly models the causal relationships between the attentional visual traces and source model attribution, and counterfactually decouples the discriminative model-specific artifacts from confounding source biases for comparison. In this way, the resulting causal effect provides a quantification on the quality of learned attention maps, thus encouraging the network to capture essential generation patterns that generalize to unseen source models by maximizing the effect. Extensive experiments on existing open-world model attribution benchmarks show that with minimal computational overhead, our method consistently improves state-of-the-art models by large margins, particularly for unseen novel attacks. Source code: https://github.com/yzheng97/CDAL.
Yu Zheng 0015, Boyang Gong, Fanye Kong, Yueqi Duan, Bingyao Yu, Wenzhao Zheng, Lei Chen 0069, Jiwen Lu, Jie Zhou 0001
ICCV7
2025 InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
abstract
Image enhancement finds wide-ranging applications in real-world scenarios due to complex environments and the inherent limitations of imaging devices. Recent diffusion-based methods yield promising outcomes but necessitate prolonged and computationally intensive iterative sampling. In response, we propose InstaRevive, a straightforward yet powerful image enhancement framework that employs score-based diffusion distillation to harness potent generative capability and minimize the sampling steps. To fully exploit the potential of the pre-trained diffusion model, we devise a practical and effective diffusion distillation pipeline using dynamic noise control to address inaccuracies in updating direction during score matching. Our noise control strategy enables a dynamic diffusing scope, facilitating precise learning of denoising trajectories within the diffusion model and ensuring accurate distribution matching gradients during training. Additionally, to enrich guidance for the generative power, we incorporate textual prompts via image captioning as auxiliary conditions, fostering further exploration of the diffusion model. Extensive experiments substantiate the efficacy of our framework across a diverse array of challenging tasks and datasets, unveiling the compelling efficacy and efficiency of InstaRevive in delivering high-quality and visually appealing results.
Yixuan Zhu, Haolin Wang 0006, Wenliang Zhao, Yansong Tang, Jingxuan Niu, Lei Chen 0069, Jie Zhou 0001, Jiwen Lu
ICLR7
2025 SPAN: A Salient Patch-Clue Aware Network for Cross-Domain Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) plays a critical role in ensuring the security of face recognition system from different kinds of presentation attacks. Most existing FAS research faces several limitations: 1) insufficient consideration of the role of local fine-grained information, 2) the assumption that spoofing patterns are uniformly distributed across the entire image, neglecting the uneven distribution of spoofing clues, and 3) an overemphasis on intra-domain scenarios, leading to limited generalization capabilities for unseen domains. In this paper, we propose a Salient Patch-Clue Aware Network (SPAN) for cross-domain face anti-spoofing to tackle the aforementioned issues. Specifically, we use all patches cropped from the complete image as input to the FAS network, enabling the network to focus on local information while avoiding information loss. Additionally, we propose a patch perception mechanism to extract key regions containing salient spoofing clues, such as reflections and edges, thereby reducing interference from irrelevant information. Furthermore, we introduce a pixel perception mechanism to capture finer-grained details. Based on these two mechanisms, we design a Salient Clue Perception Module (SCPM). We conduct cross-domain experiments on CASIA-FASD, Idiap Replay-Attack, MSU-MFSD, and OULU-NPU datasets. Our method achieves state-of-the-art HTER on seven protocols, especially excelling on M&I to C and M&I to O, surpassing the second place by 9.79% and 6.19%, showcasing strong generalization capability. The codes are available at https://github.com/SPAN2025/SPAN.
Liangfeng Zhang, Lei Chen 0069, Jinhui Lin, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
IJCNN3
2024 Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article Gronding
abstract
Weakly Supervised temporal Article Grounding (WSAG) is a challenging and practical task in video understanding. Specifically, given a video and a relevant article, whose sentences are at different semantic scales, WSAG aims to localize corresponding video segments for all “groundable” sentences. Compared to other grounding tasks, e.g., localizing one target segment with respect to a given sentence query, WSAG confronts an essential obstacle rooted in the intricate multi-scale information inherent within both textual and visual modalities. Existing methods overlook the modeling and alignment of such structured information present in multi-scale video segments and hierarchical textual content. To this end, we propose a Multi-Scale Video-Text Correspondence Learning (MVTCL) framework, which enhances the grounding performance in complex scenes by modeling multi-scale semantic correspondence both within and between modalities. Specifically, MVTCL initially aggregates video content spanning distinct temporal scales and leverages hierarchical textual relationships in both temporal and semantic dimensions via a semantic calibration module. Then multi-scale contrastive learning module is introduced to generate more discriminative representations by selecting typical contexts and performing inter-video contrastive learning. Through the multi-scale semantic calibration architecture and supervision design, our method achieves new state-of-the-art performance on existing WSAG benchmarks.
Wenjia Geng, Yong Liu 0033, Lei Chen 0069, Sujia Wang, Jie Zhou 0001, Yansong Tang
AAAI3
2024 Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
abstract
In this paper, we investigate a new problem called narrative action evaluation (NAE). NAE aims to generate professional commentary that evaluates the execution of an action. Unlike traditional tasks such as score-based action qual-ity assessment and video captioning involving superficial sentences, NAE focuses on creating detailed narratives in natural language. These narratives provide intricate descriptions of actions along with objective evaluations. NAE is a more challenging task because it requires both narrative flex-ibility and evaluation rigor. One existing possible solution is to use multi-task learning, where narrative language and evaluative information are predicted separately. However, this approach results in reduced performance for individual tasks because of variations between tasks and differences in modality between language information and evaluation information. To address this, we propose a prompt-guided multimodal interaction framework. This framework utilizes a pair of transformers to facilitate the interaction between different modalities of information. It also uses prompts to transform the score regression task into a video-text matching task, thus enabling task interactivity. To support further research in this field, we re-annotate the MTL-AQA and FineGym datasets with high-quality and comprehensive action narration. Additionally, we establish benchmarks for NAE. Extensive experiment results prove that our method outperforms separate learning methods and naive multi-task learning methods. Data and code are released at here.
Sule Bai, Guangyi Chen 0002, Lei Chen 0069, Jiwen Lu, Junle Wang, Yansong Tang
CVPR4
2024 Learning Dual-Level Deformable Implicit Representation for Real-World Scale Arbitrary Super-Resolution
Muheng Li, Jixuan Fan, Lei Chen 0069, Yansong Tang, Jiwen Lu, Jie Zhou 0001
ECCV (69)4
2024 Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting
abstract
Repetitive action counting (RAC) aims to estimate the number of class-agnostic action occurrences in a video without exemplars. Most current RAC methods rely on a raw frame-to-frame similarity representation for period prediction. However, this approach can be significantly disrupted by common noise such as action interruptions and inconsistencies, leading to sub-optimal counting performance in realistic scenarios. In this paper, we introduce a foreground localization optimization objective into similarity representation learning to obtain more robust and efficient video features. We propose a Localization-Aware Multi-Scale Representation Learning (LMRL) framework. Specifically, we apply a Multi-Scale Period-Aware Representation (MPR) with a scale-specific design to accommodate various action frequencies and learn more flexible temporal correlations. Furthermore, we introduce the Repetition Foreground Localization (RFL) method, which enhances the representation by coarsely identifying periodic actions and incorporating global semantic information. These two modules can be jointly optimized, resulting in a more discerning periodic action representation. Our approach significantly reduces the impact of noise, thereby improving counting accuracy. Additionally, the framework is designed to be scalable and adaptable to different types of video content. Experimental results on the RepCountA and UCFRep datasets demonstrate that our proposed method effectively handles repetitive action counting.
Sujia Wang, Xiangwei Shen, Yansong Tang, Wenjia Geng, Lei Chen 0069
VCIP6
2024 Frame-part-activated deep reinforcement learning for Action Prediction
Lei Chen 0069, Zhanjie Song
Pattern Recognit. Lett.1
2024 Sample Weighting with Hierarchical Equalization Loss for Dense Object Detection
abstract
Label assignment (LA) is one of the essential phases in the object detection paradigm and aims to classify samples as foreground or background. Current LA strategies generally discriminate samples by explicit thresholds and then calculate weighted losses based on their significances. However, existing methods mostly neglect to consider the importance of samples comprehensively due to the uneven distribution of objects and the limitations of detector structures. In this paper, we propose a hierarchical equalization loss (HEL) by reconsidering the underlying factors affecting sample weights. First, we mitigate sample imbalance at three progressive levels. (1) Task level. We propose task-reconciled weights (TRW) to overcome the effects caused by inter-task inconsistencies (i.e., the inherent differences of classification and localization). (2) Instance level. We propose instance-aware normalization (IAN) for reconstructing the distribution of sample weights within an instance to suppress environmental noise. (3) Pyramid level. We propose hierarchical modulation (HM) to alleviate the unbalanced distribution of multi-scale objects on feature pyramids. Then, we stack the above three mechanisms and formulate the effective weighted loss. Moreover, we propose a staggered candidate bag construction (SCBC) mechanism to further improve the robustness of our method. Without adding any extra overhead, HEL can improve the performance of representative detectors by an impressive margin. Equipped with HEL, a single “ResNet-50+FPN+Head” detector can achieve a performance of 41.9 AP on COCO under 1× schedule, outperforming other existing LA methods. Extensive experiments conducted on multiple backbones and datasets demonstrate the effectiveness of our method.
Jia-Wei Ma, Lei Chen 0069, Shu Tian, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Multim.3
2023 Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space Learning
abstract
In this paper, we propose Skip-Plan, a condensed action space learning method for procedure planning in instructional videos. Current procedure planning methods all stick to the state-action pair prediction at every timestep and generate actions adjacently. Although it coincides with human intuition, such a methodology consistently struggles with high-dimensional state supervision and error accumulation on action sequences. In this work, we abstract the procedure planning problem as a mathematical chain model. By skipping uncertain nodes and edges in action chains, we transfer long and complex sequence functions into short but reliable ones in two ways. First, we skip all the intermediate state supervision and only focus on action predictions. Second, we decompose relatively long chains into multiple short sub-chains by skipping unreliable intermediate actions. By this means, our model explores all sorts of reliable sub-relations within an action sequence in the condensed action space. Extensive experiments show Skip-Plan achieves state-of-the-art performance on the CrossTask and COIN benchmarks for procedure planning.
Wenjia Geng, Muheng Li, Lei Chen 0069, Yansong Tang, Jiwen Lu, Jie Zhou 0001
ICCV4
2023 Arbitrary Shape Text Detection via Segmentation With Probability Maps
abstract
Arbitrary shape text detection is a challenging task due to the significantly varied sizes and aspect ratios, arbitrary orientations or shapes, inaccurate annotations, etc. Due to the scalability of pixel-level prediction, segmentation-based methods can adapt to various shape texts and hence attracted considerable attention recently. However, accurate pixel-level annotations of texts are formidable, and the existing datasets for scene text detection only provide coarse-grained boundary annotations. Consequently, numerous misclassified text pixels or background pixels inside annotations always exist, degrading the performance of segmentation-based text detection methods. Generally speaking, whether a pixel belongs to text or not is highly related to the distance with the adjacent annotation boundary. With this observation, in this paper, we propose an innovative and robust segmentation-based detection method via probability maps for accurately detecting text instances. To be concrete, we adopt a Sigmoid Alpha Function (SAF) to transfer the distances between boundaries and their inside pixels to a probability map. However, one probability map can not cover complex probability distributions well because of the uncertainty of coarse-grained text boundary annotations. Therefore, we adopt a group of probability maps computed by a series of Sigmoid Alpha Functions to describe the possible probability distributions. In addition, we propose an iterative model to learn to predict and assimilate probability maps for providing enough information to reconstruct text instances. Finally, simple region growth algorithms are adopted to aggregate probability maps to complete text instances. Experimental results demonstrate that our method achieves state-of-the-art performance in terms of detection accuracy on several benchmarks. Notably, our method with Watershed Algorithm as post-processing achieves the best F-measure on Total-Text (88.79%), CTW1500 (85.75%), and MSRA-TD500 (88.93%). Besides, our method achieves promising performance on multi-oriented datasets (ICDAR2015) and multilingual datasets (ICDAR2017-MLT). Code is available at: https://github.com/GXYM/TextPMs.
Shi-Xue Zhang, Xiaobin Zhu 0001, Lei Chen 0069, Jie-Bo Hou, Xu-Cheng Yin
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Hypersphere guided embedding for masked face recognition
Xiaobin Zhu 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin, Lei Chen 0069
Pattern Recognit. Lett.6
2022 Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos
abstract
Action recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically meaningful human activities. Conventional action recognition approaches focus on analyzing single actions. However, they fail to fully reason about the contextual relations between adjacent actions, which provide potential temporal logic for understanding long videos. In this paper, we propose a prompt-based framework, Bridge-Prompt (Br-Prompt), to model the semantics across adjacent actions, so that it simultaneously exploits both out-of-context and contextual information from a series of ordinal actions in instructional videos. More specifically, we reformulate the individual action labels as integrated text prompts for super-vision, which bridge the gap between individual action semantics. The generated text prompts are paired with corresponding video clips, and together co-train the text encoder and the video encoder via a contrastive approach. The learned vision encoder has a stronger capability for ordinal-action-related downstream tasks, e.g. action segmentation and human activity recognition. We evaluate the performances of our approach on several video datasets: Georgia Tech Egocentric Activities (GTEA), 50Salads, and the Breakfast dataset. Br-Prompt achieves state-of-the-art on multiple benchmarks. Code is available at: https://github.com/ttlmh/Bridge-Prompt.
Muheng Li, Lei Chen 0069, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou 0001, Jiwen Lu
CVPR2
2022 Uncertainty-Aware Representation Learning for Action Segmentation
abstract
In this paper, we propose an uncertainty-aware representation Learning (UARL) method for action segmentation. Most existing action segmentation methods exploit continuity information of the action period to predict frame-level labels, which ignores the temporal ambiguity of the transition region between two actions. Moreover, similar periods of different actions, e.g., the beginning of some actions, will confuse the network if they are annotated with different labels, which causes spatial ambiguity. To address this, we design the UARL to exploit the transitional expression between two action periods by uncertainty learning. Specially, we model every frame of actions with an active distribution that represents the probabilities of different actions, which captures the uncertainty of the action and exploits the tendency during the action. We evaluate our method on three popular action prediction datasets: Breakfast, Georgia Tech Egocentric Activities (GTEA), and 50Salads. The experimental results demonstrate that our method achieves the performance with state-of-the-art.
Lei Chen 0069, Muheng Li, Yueqi Duan, Jie Zhou 0001, Jiwen Lu
IJCAI1
2022 Ambiguousness-Aware State Evolution for Action Prediction
abstract
In this paper, we propose an ambiguousness-aware state evolution (AASE) method which represents the uncertainty of the input sequence and evolves the subsequent skeletons to generate a reasonable full-length sequence for action prediction. Unlike most existing methods that enforce partial sequences with the labels of full-length videos and ignore the semantic information of the subsequent action, we develop an evolution method by predicting the instructional actions and generating the reasonable candidate subsequent actions, so that the ambiguity of the full sequence’s label supervising for the partial actions can be effectively alleviated. Our method generates the rational subsequent actions under the instructional action class to complement the partially observed action sequence. We design two criteria for a rational generation: 1) the instruction of subsequent action keeps the semantic consistency with the observed sequence; 2) the generation sequence is satisfied with the distribution of the sequence of real data. Moreover, we design an uncertainty module to decide the instructional action class for the generation network. AASE predicts instructional actions with uncertainty learning and evolves different instructional actions by generating the subsequent skeletons, which find the most probable action to represent the partially observed action by learning the way of perceiving the tendency of the ongoing action. We conduct experiments on seven widely used action datasets: NTU-60, NTU-120, UCF101, UT-Interaction, BIT, PKU-MMD and HMDB51, and our experimental results clearly demonstrate that our method achieves very competitive performance with state-of-the-art.
Lei Chen 0069, Jiwen Lu, Zhanjie Song, Jie Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Order-Constrained Representation Learning for Instructional Video Prediction
abstract
In this paper, we propose a weakly-supervised approach called Order-Constrained Representation Learning (OCRL) to predict future actions from instructional videos by observing incomplete steps of actions. Most conventional methods focus on predicting actions based on partially observed video frames, which mainly study low-level semantics such as motion consistency. Unlike performing a single action, completing a task in an instructional video usually requires several steps of action and longer periods. Motivated by the fact that the order of action steps is key to learning task semantics, we develop a new frame of contrastive loss, called StepNCE, to integrate the shared semantic information between step order and task semantics under the framework of the memory bank-based momentum-updating algorithm. Specifically, we learn the video representations from step order-rearranged trimmed video clips based on the proposed task-consistency rule and order-consistency rule. Our StepNCE loss can be used to pre-train a video feature encoder, which is then fine-tuned to carry out the instructional video prediction task. Our approach digs deeper into the sequential logic between different action steps with respect to a certain task, which is able to promote the video understanding methods to a new semantic level. We evaluate our method on five popular instructional video and action prediction datasets: COIN, CrossTask, UT-Interaction, BIT-Interaction, and ActivityNet v1.2, and the results show that our approach gains improvements from conventional prediction methods.
Muheng Li, Lei Chen 0069, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Recurrent Semantic Preserving Generation for Action Prediction
abstract
In this paper, we propose a recurrent semantic preserving generation (RSPG) method for action prediction. Unlike most existing methods which don't make full use of information from partially observed sequences, we develop a generation architecture to complement the sequence of skeletons for predicting the action, which can exploit more potential information of the movement tendency. Our method learns to capture the tendency of observed sequences and complement the subsequent action with adversarial learning under some constrains, which preserves the consistency between the generation sequence and the observed sequence. By generating the subsequent action, our method can predict the action with the most probability. Moreover, the redundant generation introduces the noise and disturbs the prediction. The insufficient generation cannot exploit the potential information for improving the effect of predicting the action. Our RSPG controls the generation step in a recurrent manner for maximizing the discriminative information of actions, which can adapt to the variable length of different actions. We evaluate our method on four popular action datasets: NTU, UCF101, BIT, and UT-Interaction, and experimental results show that our method achieves very competitive performance with the state-of-the-art.
Lei Chen 0069, Jiwen Lu, Zhanjie Song, Jie Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Deep Embedding Learning With Discriminative Sampling Policy
abstract
Deep embedding learning aims to learn a distance metric for effective similarity measurement, which has achieved promising performance in various tasks. As the vast majority of training samples produce gradients with magnitudes close to zero, hard example mining is usually employed to improve the effectiveness and efficiency of the training procedure. However, most existing sampling methods are designed by hand, which ignores the dependence between examples and suffer from exhaustive searching. In this paper, we propose a deep embedding with discriminative sampling policy (DE-DSP) learning framework by simultaneously training two models: a deep sampler network that learns effective sampling strategies, and a feature embedding that maps samples to the feature space. Rather than exhaustively calculating the hardness of all the examples for mining through forward-propagation, the deep sampler network exploits the strong prior of relations among samples to learn discriminative sampling policy in an more efficient manner. Experimental results demonstrate faster convergence and stronger discriminative power of our DE-DSP framework under different embedding objectives.
Yueqi Duan, Lei Chen 0069, Jiwen Lu, Jie Zhou 0001
CVPR2
2019 Learning principal orientations and residual descriptor for action recognition
Lei Chen 0069, Zhanjie Song, Jiwen Lu, Jie Zhou 0001
Pattern Recognit.1
2018 Part-Activated Deep Reinforcement Learning for Action Prediction
Lei Chen 0069, Jiwen Lu, Zhanjie Song, Jie Zhou 0001
ECCV (3)1