Ling-An Zeng

dblp:272/5365 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-8125-8024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching
abstract
To guide a learner in mastering action skills, it is crucial for a coach to 1) reason through the learner's action execution and technical points (TechPoints), and 2) provide detailed, comprehensible feedback on what is done well and what can be improved. However, existing score-based action assessment methods are still far from reaching this practical scenario. To bridge this gap, we investigate a new task termed Descriptive Action Coaching (DescCoach) which requires the model to provide detailed commentary on what is done well and what can be improved beyond a simple quality score for action execution. To this end, we first build a new dataset named EE4D-DescCoach. Through an automatic annotation pipeline, our dataset goes beyond the existing action assessment datasets by providing detailed TechPoint-level commentary. Furthermore, we propose TechCoach, a new framework that explicitly incorporates TechPoint-level reasoning into the DescCoach process. The central to our method lies in the Context-aware TechPoint Reasoner, which enables TechCoach to learn TechPoint-related quality representation by querying visual context under the supervision of TechPoint-level coaching commentary. By leveraging the visual context and the TechPoint-related quality representation, a unified TechPoint-aware Action Assessor is then employed to provide the overall coaching commentary together with the quality score. Combining all of these, we establish a new benchmark for DescCoach and evaluate the effectiveness of our method through extensive experiments.
Yuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin, Yu-Ming Tang, Wei-Shi Zheng 0001
AAAI3
2025 Light-T2M: A Lightweight and Fast Model for Text-to-motion Generation
abstract
Despite the significant role text-to-motion (T2M) generation plays across various applications, current methods involve a large number of parameters and suffer from slow inference speeds, leading to high usage costs. To address this, we aim to design a lightweight model to reduce usage costs. First, unlike existing works that focus solely on global information modeling, we recognize the importance of local information modeling in the T2M task by reconsidering the intrinsic properties of human motion, leading us to propose a lightweight Local Information Modeling Module. Second, we are the first to introduce Mamba to the T2M task, reducing the number of parameters and GPU memory demands, and we have designed a novel Pseudo-bidirectional Scan to replicate the effects of a bidirectional scan without increasing parameter count. Moreover, we propose a novel Adaptive Textual Information Injector that more effectively integrates textual information into the motion during generation. By integrating the aforementioned designs, we propose a lightweight and fast model named Light-T2M. Compared to the state-of-the-art method, MoMask, our Light-T2M model features just 10% of the parameters (4.48M vs 44.85M) and achieves a 16% faster inference time (0.152s vs 0.180s), while surpassing MoMask with an FID of 0.040 (vs. 0.045) on HumanML3D dataset and 0.161 (vs. 0.228) on KIT-ML dataset.
Ling-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi Zheng 0001
AAAI1
2025 ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation
abstract
We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using full-body poses as tokens, we argue that explicitly modeling joint-level interactions is more natural and effective for generating realistic HOIs, as it directly captures the geometric and semantic relationships between joints, rather than modeling interactions in the latent pose space. To this end, ChainHOI introduces a novel joint graph to capture potential interactions with objects, and a Generative Spatiotemporal Graph Convolution Network to explicitly model interactions at the joint level. Furthermore, we propose a Kinematics-based Interaction Module that explicitly models interactions at the kinetic chain level, ensuring more realistic and biomechanically coherent motions. Evaluations on two public datasets demonstrate that ChainHOI significantly outperforms previous methods, generating more realistic, and semantically consistent HOIs. Code is available here.
Ling-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu, Yu-Ming Tang, Jingke Meng, Wei-Shi Zheng 0001
CVPR1
2025 Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction Framework
abstract
Bimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their actions. However, we think bimanual manipulation involves not only coordinated tasks but also various uncoordinated tasks that do not require explicit cooperation during execution, such as grasping objects with the closest hand, which integrated control frameworks ignore to consider due to their enforced cooperation in the early inputs. In this paper, we propose a novel decoupled interaction framework that considers the characteristics of different tasks in bimanual manipulation. The key insight of our framework is to assign an independent model to each arm to enhance the learning of uncoordinated tasks, while introducing a selective interaction module that adaptively learns weights from its own arm to improve the learning of coordinated tasks. Extensive experiments on seven tasks in the RoboTwin dataset demonstrate that: (1) Our framework achieves outstanding performance, with a 23.5% boost over the SOTA method. (2) Our framework is flexible and can be seamlessly integrated into existing methods. (3) Our framework can be effectively extended to multi-agent manipulation tasks, achieving a 28% boost over the integrated control SOTA. (4) The performance boost stems from the decoupled design itself, surpassing the SOTA by 16.5% in success rate with only 1/6 of the model size.
Jian-Jian Jiang, Xiao-Ming Wu 0002, Yi-Xiang He, Ling-An Zeng, Yi-Lin Wei, Wei-Shi Zheng 0001
ICCV4
2025 AffordDexGrasp: Open-Set Language-Guided Dexterous Grasp With Generalizable-Instructive Affordance
Yi-Lin Wei, Mu Lin, Jian-Jian Jiang, Xiao-Ming Wu 0002, Ling-An Zeng, Wei-Shi Zheng 0001
ICCV6
2025 Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation
abstract
We propose a novel approach for generating text-guided human-object interactions (HOIs) that achieves explicit joint-level interaction modeling in a computationally efficient manner. Previous methods represent the entire human body as a single token, making it difficult to capture fine-grained joint-level interactions and resulting in unrealistic HOIs. However, treating each individual joint as a token would yield over twenty times more tokens, increasing computational overhead. To address these challenges, we introduce an Efficient Explicit Joint-level Interaction Model (EJIM). EJIM features a Dual-branch HOI Mamba that separately and efficiently models spatiotemporal HOI information, as well as a Dual-branch Condition Injector for integrating text semantics and object geometry into human and object motions. Furthermore, we design a Dynamic Interaction Block and a progressive masking mechanism to iteratively filter out irrelevant joints, ensuring accurate and nuanced interaction modeling. Extensive quantitative and qualitative evaluations on public datasets demonstrate that EJIM surpasses previous works by a large margin while using only 5% of the inference time. Code is available here.
Guohong Huang, Ling-An Zeng, Zexin Zheng, Shengbo Gu, Wei-Shi Zheng 0001
ICME2
2025 Supplementary Material for "NoiseActor: A Noise-Action Collaborative Framework for Privacy-Preserving Action Recognition without Privacy Labels"
abstract
Our supplementary material is organized into the following sections: • Section II provides the details of the LOCATION module. • Section III provides the evaluation protocol for the SBU dataset. • Section IV provides the evaluation protocol for cross-dataset experiment on the UCF101 dataset and the VISPR dataset. • Section V provides implementation details for each datasets. • Section VI provides more anonymized frames on the SBU dataset. • Section VII provides more anonymized frames on the UCF101 dataset. • Section VIII provides details about anonymized video file. • Section IX provides ablation studies of the adapter.
Xiao Li 0074, Xiao-Ming Wu 0002, Delong Zhang, Kun-Yu Lin, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001
ICME6
2025 Progressive Human Motion Generation Based on Text and Few Motion Frames
abstract
Although existing text-to-motion (T2M) methods can produce realistic human motion from text description, it is still difficult to align the generated motion with the desired postures since using text alone is insufficient for precisely describing diverse postures. To achieve more controllable generation, an intuitive way is to allow the user to input a few motion frames describing precise desired postures. Thus, we explore a new Text-Frame-to-Motion (TF2M) generation task that aims to generate motions from text and very few given frames. Intuitively, the closer a frame is to a given frame, the lower the uncertainty of this frame is when conditioned on this given frame. Hence, we propose a novel Progressive Motion Generation (PMG) method to progressively generate a motion from the frames with low uncertainty to those with high uncertainty in multiple stages. During each stage, new frames are generated by a Text-Frame Guided Generator conditioned on frame-aware semantics of the text, given frames, and frames generated in previous stages. Additionally, to alleviate the train-test gap caused by multi-stage accumulation of incorrectly generated frames during testing, we propose a Pseudo-frame Replacement Strategy for training. Experimental results show that our PMG outperforms existing T2M generation methods by a large margin with even one given frame, validating the effectiveness of our PMG. Code is available here.
Ling-An Zeng, Gaojie Wu, Ancong Wu, Jianfang Hu, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng, Jingke Meng, Wei-Shi Zheng 0001
ECCV (20)4
2024 Privacy-Preserving Action Recognition: A Survey
Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001
PRCV (7)4
2024 Continual Action Assessment via Task-Consistent Score-Discriminative Feature Distribution Modeling
abstract
Action Quality Assessment (AQA) is a task that tries to answer how well an action is carried out. While remarkable progress has been achieved, existing works on AQA assume that all the training data are visible for training at one time, but do not enable continual learning on assessing new technical actions. In this work, we address such a Continual Learning problem in AQA (Continual-AQA), which urges a unified model to learn AQA tasks sequentially without forgetting. Our idea for modeling Continual-AQA is to sequentially learn a task-consistent score-discriminative feature distribution, in which the latent features express a strong correlation with the score labels regardless of the task or action types. From this perspective, we aim to mitigate the forgetting in Continual-AQA from two aspects. Firstly, to fuse the features of new and previous data into a score-discriminative distribution, a novel Feature-Score Correlation-Aware Rehearsal is proposed to store and reuse data from previous tasks with limited memory size. Secondly, an Action General-Specific Graph is developed to learn and decouple the action-general and action-specific knowledge so that the task-consistent score-discriminative features can be better extracted across various tasks. Extensive experiments are conducted to evaluate the contributions of proposed components. The comparisons with the existing continual learning methods additionally verify the effectiveness and versatility of our approach. Data and code are available at https://github.com/iSEE-Laboratory/Continual-AQA.
Yuan-Ming Li, Ling-An Zeng, Jingke Meng, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multimodal Action Quality Assessment
abstract
Action quality assessment (AQA) is to assess how well an action is performed. Previous works perform modelling by only the use of visual information, ignoring audio information. We argue that although AQA is highly dependent on visual information, the audio is useful complementary information for improving the score regression accuracy, especially for sports with background music, such as figure skating and rhythmic gymnastics. To leverage multimodal information for AQA, i.e., RGB, optical flow and audio information, we propose a Progressive Adaptive Multimodal Fusion Network (PAMFN) that separately models modality-specific information and mixed-modality information. Our model consists of with three modality-specific branches that independently explore modality-specific information and a mixed-modality branch that progressively aggregates the modality-specific information from the modality-specific branches. To build the bridge between modality-specific branches and the mixed-modality branch, three novel modules are proposed. First, a Modality-specific Feature Decoder module is designed to selectively transfer modality-specific information to the mixed-modality branch. Second, when exploring the interaction between modality-specific information, we argue that using an invariant multimodal fusion policy may lead to suboptimal results, so as to take the potential diversity in different parts of an action into consideration. Therefore, an Adaptive Fusion Module is proposed to learn adaptive multimodal fusion policies in different parts of an action. This module consists of several FusionNets for exploring different multimodal fusion strategies and a PolicyNet for deciding which FusionNets are enabled. Third, a module called Cross-modal Feature Decoder is designed to transfer cross-modal features generated by Adaptive Fusion Module to the mixed-modality branch. Our extensive experiments validate the efficacy of the proposed method, and our method achieves state-of-the-art performance on two public datasets. Code is available at https://github.com/qinghuannn/PAMFN.
Ling-An Zeng, Wei-Shi Zheng 0001
IEEE Trans. Image Process.1
2024 Adaptive Weight Generator for Multi-Task Image Recognition by Task Grouping Prompt
abstract
Comparing to adapting the pre-trained backbone to a single image recognition task, multi-task image recognition enables the backbone to perform better when the tasks are related. An interesting research field in multi-task learning (MTL) is to learn the parameter sharing pattern among the involved tasks. Most existing works obtain the sharing pattern, ignoring the task grouping information among the involved tasks. In this work, we aim to build the task parameter sharing pattern based on automatically acquiring the task grouping information. The task grouping information together with the task specific information is then utilized to yield the task adaptive weights. Our method, called Task Grouping prompt-based Adaptive Weight generator (TGAW), consists of Prompt-based Task Representation (PTR) and Prompt-based Weight Generator (PWG). The PTR is modeled as task prompts consisting of task grouping prompt and task specific prompt. The task grouping prompt is automatically chosen from a candidate pool for each task, and tasks selecting the same grouping prompts are divided into the same group. Then, PWG generates task adaptive weights based on the task prompts. The experimental results show that TGAW achieves comparable performance with less than 30% amount of trainable parameters of the pre-trained backbone.
Gaojie Wu, Ling-An Zeng, Jingke Meng, Wei-Shi Zheng 0001
IEEE Trans. Multim.2
2022 Likert Scoring with Grade Decoupling for Long-term Action Assessment
abstract
Long-term action quality assessment is a task of evaluating how well an action is performed, namely, estimating a quality score from a long video. Intuitively, long-term actions generally involve parts exhibiting different levels of skill, and we call the levels of skill as performance grades. For example, technical highlights and faults may appear in the same long-term action. Hence, the final score should be determined by the comprehensive effect of different grades exhibited in the video. To explore this latent relationship, we design a novel Likert scoring paradigm in-spired by the Likert scale in psychometrics, in which we quantify the grades explicitly and generate the final quality score by combining the quantitative values and the corresponding responses estimated from the video, instead of performing direct regression. Moreover, we extract grade-specific features, which will be used to estimate the responses of each grade, through a Transformer decoder architecture with diverse learnable queries. The whole model is named as Grade-decoupling Likert Transformer (GDLT), and we achieve state-of-the-art results on two long-term action assessment datasets.11Project page https://isee-ai.cn/-angchi/CVPR22_GDLT.html
Angchi Xu, Ling-An Zeng, Wei-Shi Zheng 0001
CVPR2
2020 Hybrid Dynamic-static Context-aware Attention Network for Action Assessment in Long Videos
abstract
The objective of action quality assessment is to score sports videos. However, most existing works focus only on video dynamic information (i.e., motion information) but ignore the specific postures that an athlete is performing in a video, which is important for action assessment in long videos. In this work, we present a novel hybrid dynAmic-static Context-aware attenTION NETwork (ACTION-NET) for action assessment in long videos. To learn more discriminative representations for videos, we not only learn the video dynamic information but also focus on the static postures of the detected athletes in specific frames, which represent the action quality at certain moments, along with the help of the proposed hybrid dynamic-static architecture. Moreover, we leverage a context-aware attention module consisting of a temporal instance-wise graph convolutional network unit and an attention unit for both streams to extract more robust stream features, where the former is for exploring the relations between instances and the latter for assigning a proper weight to each instance. Finally, we combine the features of the two streams to regress the final video score, supervised by ground-truth scores given by experts. Additionally, we have collected and annotated the new Rhythmic Gymnastics dataset, which contains videos of four different types of gymnastics routines, for evaluation of action quality assessment in long videos. Extensive experimental results validate the efficacy of our proposed method, which outperforms related approaches.
Ling-An Zeng, Fa-Ting Hong, Wei-Shi Zheng 0001, Qi-Zhi Yu, Wei Zeng 0006, Yaowei Wang 0001, Jian-Huang Lai
ACM Multimedia1