Cheng Zhang 0020

dblp:82/6384-20 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-6831-5103ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 23% Image recognition and object detection · 17% Efficient and distributed learning · 14%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
transformer
1.522025
TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image Classification · IEEE Trans. Multim. 2025
TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via Transformers · CVPR 2023
Computer vision › Face, body and person analysis
head pose estimation
1.322023
Orientation Cues-Aware Facial Relationship Representation for Head Pose Estimation via Transformer · IEEE Trans. Image Process. 2023
TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via Transformers · CVPR 2023
Computer vision › 3D vision › 3d scene understanding
3d visual grounding
1.012026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
1.012026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › scheduling
task scheduling
1.012026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Computer vision › Image recognition and object detection › image classification › fine-grained image classification
fine-grained bird image classification
0.912025
TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image Classification · IEEE Trans. Multim. 2025
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.912025
TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image Classification · IEEE Trans. Multim. 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image Classification · IEEE Trans. Multim. 2025
Machine learning › Efficient and distributed learning
model compression
0.812024
Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression · ECCV (47) 2024
Machine learning › Efficient and distributed learning › model compression
token compression
0.812024
Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression · ECCV (47) 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312026
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution · AAAI 2026
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
multi-view 3d object detection
0.212024
Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression · ECCV (47) 2024

Methods — techniques the papers use, named apart from their topics

orientation tokens · 1.3scheduling token mechanism · 1.0hierarchy stage feature aggregation · 0.9feature in feature abstraction · 0.9vision transformer · 0.8transformer · 0.7token guide multiloss · 0.7multi-loss function · 0.7
YearPublicationVenuePosition
2026 Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
abstract
Task scheduling has become increasingly critical for embodied AI, where agents need to follow natural language instructions and execute actions efficiently in 3D physical worlds. Existing datasets for task planning in 3D environments often simplify the problem, lacking operations research knowledge for task scheduling and 3D grounding for real-world applications. In this work, we propose Operations Research Knowledge-based 3D Grounded Task Scheduling (OKS3D), a new task that requires synerization of language understanding, 3D grounding, and efficiency optimization for embodied agents. OKS3D reflects real-world demands by requiring agents to generate efficient, step-by-step schedules that are grounded in 3D space. To facilitate research on OKS3D, we construct a large-scale dataset called OKS3D-60K, comprising 60K tasks across 4K real-world scenes. Furthermore, we propose GRANT, an embodied multi-modal large language model equipped with a simple yet effective scheduling token mechanism to generate efficient task schedules and grounded actions. Extensive experiments on the OKS3D-60K dataset validate the effectiveness of GRANT across language understanding, 3D grounding, and scheduling efficiency.
Dingkang Liang, Cheng Zhang 0020, Xiaopeng Xu, Jianzhong Ju, Zhenbo Luo, Xiang Bai
AAAI2
2025 TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image Classification
abstract
Fine-grained bird image classification (FBIC) is not only meaningful for endangered bird observation and protection but also a prevalent task for image classification in multimedia processing and computer vision. However, FBIC suffers from several challenges, such as bird molting, complex background, and arbitrary bird posture. To effectively tackle these challenges, we present a novel invariant cues-aware feature concentration Transformer (TransIFC), which learns invariant and core information in bird images. To this end, two novel modules are proposed to leverage the characteristics of bird images, namely, the hierarchy stage feature aggregation (HSFA) module and the feature in feature abstraction (FFA) module. The HSFA module aggregates the multiscale information of bird images by concatenating multilayer features. The FFA module extracts the invariant cues of birds through feature selection based on discrimination scores. Transformer is employed as the backbone to reveal the long-dependent semantic relationships in bird images. Moreover, abundant visualizations are provided to prove the interpretability of the HSFA and FFA modules in TransIFC. Comprehensive experiments demonstrate that TransIFC can achieve state-of-the-art performance on the CUB-200-2011 dataset (91.0%) and the NABirds dataset (90.9%). Finally, extended experiments have been conducted on the Stanford Cars dataset to suggest the potential of generalizing our method on other fine-grained visual classification tasks.
Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Bochen Xie, Tingting Liu 0006, Youfu Li 0001
IEEE Trans. Multim.2
2024 Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression
Dingyuan Zhang, Dingkang Liang, Zichang Tan, Xiaoqing Ye, Cheng Zhang 0020, Jingdong Wang 0001, Xiang Bai
ECCV (47)5
2024 MMATrans: Muscle Movement Aware Representation Learning for Facial Expression Recognition via Transformers
abstract
How to automatically recognize facial expression has caused concerns in industrial human–robot interaction. However, facial expression recognition (FER) is susceptible to problems, such as occlusion, arbitrary orientations, and illumination. To effectively address these challenges in FER, we present a novel facial muscle movement aware representation learning that can learn the semantic relationships of facial muscle movements in facial expression images. Two key findings are revealed: 1) muscle movements from different facial regions often show semantic relationships; and 2) not all facial muscle regions have equal contributions for different facial expressions. On this basis, this model presents two novel modules, namely, discriminative feature generation (DFG) and muscle relationship mining (MRM). Specifically, in DFG, the memory of our model for mislabeling decreases. In MRM, muscle–motion interaction among diverse facial regions is learned through visual transformers (MMATrans). Experiments on three in-the-wild FER datasets (RAF-DB, FERPlus, and AffectNet) show that our MMATrans yields better performance compared with state-of-the-art methods.
Hai Liu 0004, Qiyun Zhou, Cheng Zhang 0020, Junyan Zhu, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001
IEEE Trans. Ind. Informatics3
2023 TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via Transformers
abstract
Head pose estimation (HPE) has been widely used in the fields of human machine interaction, self-driving, and attention estimation. However, existing methods cannot deal with extreme head pose randomness and serious occlusions. To address these challenges, we identify three cues from head images, namely, neighborhood similarities, significant facial changes, and critical minority relationships. To leverage the observed findings, we propose a novel critical minority relationship-aware method based on the Transformer architecture in which the facial part relationships can be learned. Specifically, we design several orientation tokens to explicitly encode the basic orientation regions. Meanwhile, a novel token guide multiloss function is designed to guide the orientation tokens as they learn the desired regional similarities and relationships. We evaluate the proposed method on three challenging benchmark HPE datasets. Experiments show that our method achieves better performance compared with state-of-the-art methods. Our code is publicly available at https://github.com/zc2023/TokenHPE.
Cheng Zhang 0020, Hai Liu 0004, Yongjian Deng, Bochen Xie, Youfu Li 0001
CVPR1
2023 Orientation Cues-Aware Facial Relationship Representation for Head Pose Estimation via Transformer
abstract
Head pose estimation (HPE) is an indispensable upstream task in the fields of human-machine interaction, self-driving, and attention detection. However, practical head pose applications suffer from several challenges, such as severe occlusion, low illumination, and extreme orientations. To address these challenges, we identify three cues from head images, namely, critical minority relationships, neighborhood orientation relationships, and significant facial changes. On the basis of the three cues, two key insights on head poses are revealed: 1) intra-orientation relationship and 2) cross-orientation relationship. To leverage two key insights above, a novel relationship-driven method is proposed based on the Transformer architecture, in which facial and orientation relationships can be learned. Specifically, we design several orientation tokens to explicitly encode basic orientation regions. Besides, a novel token guide multi-loss function is accordingly designed to guide the orientation tokens as they learn the desired regional similarities and relationships. Experimental results on three challenging benchmark HPE datasets show that our proposed TokenHPE achieves state-of-the-art performance. Moreover, qualitative visualizations are provided to verify the effectiveness of the token-learning methodology.
Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001
IEEE Trans. Image Process.2