Yin-Dong Zheng

dblp:249/8371 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-2437-9807ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Feature matters: Revisiting channel attention for Temporal Action Detection
Guo Chen 0006, Yin-Dong Zheng, Jiahao Wang 0005, Tong Lu 0002
Pattern Recognit.2
2026 HandMLP: A GAT-Enhanced MLP Network for Robust 3-D Hand Pose Estimation
abstract
3-D hand pose estimation plays a crucial role in various applications, including virtual interaction, augmented reality, and human–robot collaboration. However, the complex articulation of hand joints and frequent self-occlusions make this task highly challenging. Existing methods based on convolutional or graph architectures are effective for local modeling, but often struggle to jointly capture global context and structured joint dependencies. To address this limitation, we proposeHandMLP, a novel framework that integrates multilayer perceptron (MLP)-based global modeling with graph attention for structured joint reasoning. Specifically, the MLP module enhances global semantic representation and nonlinear fitting capacity, while the graph attention module explicitly models local structural relationships and dynamic dependencies among hand joints. In addition, we introduce a voting-based initialization strategy to generate coarse pose hypotheses, which are progressively refined through collaborative reasoning, improving robustness under sparse or noisy point clouds. Extensive experiments on three benchmark datasets demonstrate thatHandMLPconsistently outperforms state-of-the-art methods, achieving superior accuracy and robustness in 3-D hand pose estimation.
Changbo Gao, Yin-Dong Zheng, Mingxi Zhuang, Yanzi Miao, Hesheng Wang 0001
IEEE Trans. Ind. Informatics3
2025 Egocentric Object-Interaction Anticipation with Retentive and Predictive Learning
abstract
Egocentric object-interaction anticipation is critical for applications like augmented reality and robotics, but existing methods struggle with misaligned egocentric encoding, insufficient supervision, and underutilized historical context. These limitations stem from a lack of focus on retention, i.e., retaining long-term object-centric interactions, and prediction, i.e., future-centric encoding and future uncertainty modeling. We introduce EgoAnticipator, a novel Retentive and Predictive Learning framework that addresses these challenges. Our approach combines retentive pre-training for domain-specific encoding, predictive pre-training for future uncertainty modeling, and mirror distillation to transfer future-informed knowledge. Additionally, we propose long-term memory prompting to integrate historical interaction cues. We evaluate the effectiveness of our framework using the Ego4D short-term object interaction anticipation benchmark, covering both STAv1 and STAv2. Extensive experiments demonstrate that our framework outperforms existing methods, while ablation studies highlight the effectiveness of each design inside our retentive and predictive learning framework.
Guo Chen 0006, Yifei Huang 0002, Yin-Dong Zheng, Jiahao Wang 0005, Tong Lu 0002
IJCAI3
2025 Zero-Shot Temporal Interaction Localization for Egocentric Videos
abstract
Locating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer. Current temporal action localization methods typically rely on annotated action and object categories of interactions for optimization, which leads to domain bias and low deployment efficiency. Although some recent works have achieved zero-shot temporal action localization (ZS-TAL) with large vision-language models (VLMs), their coarse-grained estimations and open-loop pipelines hinder further performance improvements for temporal interaction localization (TIL). To address these issues, we propose a novel zero-shot TIL approach dubbed EgoLoc to locate the timings of grasp actions for human-object interaction in egocentric videos. EgoLoc introduces a self-adaptive sampling strategy to generate reasonable visual prompts for VLM reasoning. By absorbing both 2D and 3D observations, it directly samples high-quality initial guesses around the possible contact/separation timestamps of HOI according to 3D hand velocities, leading to high inference accuracy and efficiency. In addition, EgoLoc generates closed-loop feedback from visual and dynamic cues to further refine the localization results. Comprehensive experiments on the publicly available dataset and our newly proposed benchmark demonstrate that EgoLoc achieves better temporal interaction localization for egocentric videos compared to state-of-the-art baselines. We will release our code and relevant data as open-source at https://github.com/IRMVLab/EgoLoc.
Erhang Zhang, Junyi Ma, Yin-Dong Zheng, Yixuan Zhou 0003
IROS3
2023 ELAN: Enhancing Temporal Action Detection with Location Awareness
abstract
Current query-based temporal action detection methods lack multiple levels of location awareness, leading to performance degradation. In this paper, we present a novel query-based method called Enhanced Location-Aware Network (ELAN) for temporal action detection. ELAN adopts a lightweight convolution-based encoder, termed Temporal Location-Aware (TLA) encoder, to model temporal continuous location-aware context. Moreover, ELAN can re-aware the location-related context inside and between queries through our proposed Instance Location-Aware (ILA) decoder. As a result, ELAN can learn strong position discrimination of actions and effectively eliminates the ambiguity caused by sparse action decoding, yielding significant improvement in detection performance. ELAN achieves state-of-the-art performance on two temporal action detection benchmarks, including THUMOS-14 and ActivityNet-1.3.
Guo Chen 0006, Yin-Dong Zheng, Zhe Chen 0017, Jiahao Wang 0005, Tong Lu 0002
ICME2
2023 MRSN: Multi-Relation Support Network for Video Action Detection
abstract
Action detection is a challenging video understanding task, requiring modeling spatio-temporal and interaction relations. Current methods usually model actor-actor and actor-context relations separately, ignoring their complementarity and mutual support. To solve this problem, we propose a novel network called Multi-Relation Support Network (MRSN). In MRSN, Actor-Context Relation Encoder (ACRE) and Actor-Actor Relation Encoder (AARE) model the actor-context and actor-actor relation separately. Then Relation Support Encoder (RSE) computes the supports between the two relations and performs relation-level interactions. Finally, Relation Consensus Module (RCM) enhances two relations with the long-term relations from the Long-term Relation Bank (LRB) and yields a consensus. Our experiments demonstrate that modeling relations separately and performing relation-level interactions can achieve and outperformer state-of-the-art results on two challenging video datasets: AVA and UCF101-24.
Yin-Dong Zheng, Guo Chen 0006, Minglei Yuan, Tong Lu 0002
ICME1
2023 BasicTAD: An astounding RGB-Only baseline for temporal action detection
Min Yang 0011, Guo Chen 0006, Yin-Dong Zheng, Tong Lu 0002, Limin Wang 0002
Comput. Vis. Image Underst.3
2022 DCAN: Improving Temporal Action Detection via Dual Context Aggregation
abstract
Temporal action detection aims to locate the boundaries of action in the video. The current method based on boundary matching enumerates and calculates all possible boundary matchings to generate proposals. However, these methods neglect the long-range context aggregation in boundary prediction. At the same time, due to the similar semantics of adjacent matchings, local semantic aggregation of densely-generated matchings cannot improve semantic richness and discrimination. In this paper, we propose the end-to-end proposal generation method named Dual Context Aggregation Network (DCAN) to aggregate context on two levels, namely, boundary level and proposal level, for generating high-quality action proposals, thereby improving the performance of temporal action detection. Specifically, we design the Multi-Path Temporal Context Aggregation (MTCA) to achieve smooth context aggregation on boundary level and precise evaluation of boundaries. For matching evaluation, Coarse-to-fine Matching (CFM) is designed to aggregate context on the proposal level and refine the matching map from coarse to fine. We conduct extensive experiments on ActivityNet v1.3 and THUMOS-14. DCAN obtains an average mAP of 35.39% on ActivityNet v1.3 and reaches mAP 54.14% at [email protected] on THUMOS-14, which demonstrates DCAN can generate high-quality proposals and achieve state-of-the-art performance. We release the code at https://github.com/cg1177/DCAN.
Guo Chen 0006, Yin-Dong Zheng, Limin Wang 0002, Tong Lu 0002
AAAI2
2022 Uncertainty-Based Network for Few-Shot Image Classification
abstract
The transductive inference is an effective technique in the few-shot learning task, where query sets update prototypes to improve themselves. However, these methods optimize the model by considering only the classification scores of the query instances as confidence while ignoring the uncertainty of these classification scores. In this paper, we propose a novel method called Uncertainty-Based Network, which models the uncertainty of classification results with the help of mutual information. Specifically, we first data augment and classify the query instance and calculate the mutual information of these classification scores. Then, mutual information is used as uncertainty to assign weights to classification scores, and the iterative update strategy based on classification scores and uncertainties assigns the optimal weights to query instances in prototype optimization. Extensive results on four benchmarks show that Uncertainty-Based Network achieves comparable performance in classification accuracy compared to state-of-the-art methods.
Minglei Yuan, Chunhao Cai, Yin-Dong Zheng, Tao Wang 0052, Tong Lu 0002, Wenbin Li 0006
ICME4
2020 Dynamic Sampling Networks for Efficient Action Recognition in Videos
abstract
The existing action recognition methods are mainly based on clip-level classifiers such as two-stream CNNs or 3D CNNs, which are trained from the randomly selected clips and applied to densely sampled clips during testing. However, this standard setting might be suboptimal for training classifiers and also requires huge computational overhead when deployed in practice. To address these issues, we propose a new framework for action recognition in videos, called Dynamic Sampling Networks (DSN), by designing a dynamic sampling module to improve the discriminative power of learned clip-level classifiers and as well increase the inference efficiency during testing. Specifically, DSN is composed of a sampling module and a classification module, whose objective is to learn a sampling policy to on-the-fly select which clips to keep and train a clip-level classifier to perform action recognition based on these selected clips, respectively. In particular, given an input video, we train an observation network in an associative reinforcement learning setting to maximize the rewards of the selected clips with a correct prediction. We perform extensive experiments to study different aspects of the DSN framework on four action recognition datasets: UCF101, HMDB51, THUMOS14, and ActivityNet v1.3. The experimental results demonstrate that DSN is able to greatly improve the inference efficiency by only using less than half of the clips, which can still obtain a slightly better or comparable recognition accuracy to the state-of-the-art approaches.
Yin-Dong Zheng, Zhaoyang Liu 0001, Tong Lu 0002, Limin Wang 0002
IEEE Trans. Image Process.1
2019 A Novel Group-Aware Pruning Method for Few-shot Learning
abstract
Few-shot learning, which focuses on solving machine learning problems in the setting of scarce data, has become a hot spot in the field of neural network and machine learning recently. Inspired by BlockDrop that performs pruning compression on the network using reinforcement learning, in this paper, we present a Group-Aware Pruning (GAP) method with new unifying pruning strategies and different training/testing ways to fit few-shot learning. The proposed GAP consists of three modules, that is, a pruning module, a strategy consensus module (SCM), and a classification module. The whole support set is first fed into the pruning module to get a pruning strategy for each image. Next, the strategies are fed into SCM to fuse into a group strategy for further pruning. When the classification module makes a prediction for a query image, it no longer needs re-output the pruning strategy for the query image since pruning strategies have been unified. Note that in SCM, strategy unification is proposed to achieve the group-aware strategy, which assures an existing deep network that has been pruned by the group-aware strategy work well for few-shot learning problems. Additionally, the testing will be speeded up due to the fact that only one classification model is required in the proposed GAP. Experiments on two benchmark datasets, namely, Omingnet and miniImageNet, and comparisons with the existing state-of-the-art methods show that the accuracy of the proposed GAP is 4.94% higher than the state-of-the-art methods on the 5-way 5-shot task, which shows the effectiveness of the proposed method.
Yin-Dong Zheng, Yun-Tao Ma, Ruo-Ze Liu
IJCNN1