Huayi Zhou 0001

dblp:182/4224-1 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-2220-7286ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HACMatch: Semi-supervised rotation regression with hardness-aware curriculum pseudo labeling
Huayi Zhou 0001, Suizhi Huang, Yue Ding 0001, Hongtao Lu 0001
Comput. Vis. Image Underst.2
2026 Semi-Supervised Unconstrained Head Pose Estimation in the Wild
abstract
Existing research on unconstrained in-the-wild head pose estimation suffers from the flaws of its datasets, which consist of either numerous samples by non-realistic synthesis or constrained collection, or small-scale natural images yet with plausible manual annotations. This makes fully-supervised solutions compromised due to the reliance on generous labels. To alleviate it, we propose the first semi-supervised unconstrained head pose estimation method SemiUHPE, which can leverage abundant easily available unlabeled head images. Technically, we choose semi-supervised rotation regression and adapt it to the error-sensitive and label-scarce problem of unconstrained head pose. Our method is based on the observation that the aspect-ratio invariant cropping of wild heads is superior to previous landmark-based affine alignment given that landmarks of unconstrained human heads are usually unavailable, especially for underexplored non-frontal heads. Instead of using a pre-fixed threshold to filter out pseudo labeled heads, we propose dynamic entropy based filtering to adaptively remove unlabeled outliers as training progresses by updating the threshold in multiple stages. We then revisit the design of weak-strong augmentations and improve it by devising two novel head-oriented strong augmentations, termed pose-irrelevant cut-occlusion and pose-altering rotation consistency respectively. Extensive experiments and ablation studies show that SemiUHPE outperforms its counterparts greatly on public benchmarks under both the front-range and full-range settings. Furthermore, our proposed method is also beneficial for solving other closely related problems, including generic object rotation regression and 3D head reconstruction, demonstrating good versatility and extensibility.
Huayi Zhou 0001, Fei Jiang 0006, Yong Rui, Hongtao Lu 0001, Kui Jia
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 YOTO++: Learning Long-Horizon Closed-Loop Bimanual Manipulation From One-Shot Human Video Demonstrations
abstract
Bimanual robotic manipulation remains a fundamental challenge due to the inherent complexity of dual-arm coordination and high-dimensional action spaces. This paper presents the extended YOTO++ (You Only Teach Once), which is a unified one-shot learning framework for teaching bimanual skills directly from third-person human video demonstrations. Our method extracts structured 3D hand motions using binocular vision and distills them into compact, keyframe-based trajectories for dual-arm execution. We develop a scalable demonstration proliferation strategy that synthetically augments one-shot demonstrations into diverse training samples, enabling effective learning of a customized bimanual diffusion policy. Extensive evaluations across a broad spectrum of long-horizon bimanual tasks, including asynchronous, synchronous, contact-rich, and non-prehensile scenarios, demonstrate strong generalization to novel skills and objects. We further introduce a visual alignment mechanism at the initial manipulation stage for closed-loop control, enabling the system to timely adapt to perturbations during execution. We validate the framework on an unseen dual-arm robotic platform to show seamless cross-embodiment transfer without additional retraining. YOTO++ achieves impressive performance in accuracy, robustness, and scalability, advancing the practical deployment of general-purpose bimanual manipulation systems.
Huayi Zhou 0001, Yunxin Tai, Yueci Deng, Guiliang Liu, Kui Jia
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Unsupervised Domain Adaptive Hand Mesh Reconstruction of 2D Images in the Wild
Xinyi Hou, Huayi Zhou 0001, Yue Ding 0001, Hongtao Lu 0001
ICANN (2)2
2025 DexScale: Automating Data Scaling for Sim2Real Generalizable Robot Control
abstract
A critical prerequisite for achieving generalizable robot control is the availability of a large-scale robot training dataset. Due to the expense of collecting realistic robotic data, recent studies explored simulating and recording robot skills in virtual environments. While simulated data can be generated at higher speeds, lower costs, and larger scales, the applicability of such simulated data remains questionable due to the gap between simulated and realistic environments. To advance the Sim2Real generalization, in this study, we present DexScale, a data engine designed to perform automatic skills simulation and scaling for learning deployable robot manipulation policies. Specifically, DexScale ensures the usability of simulated skills by integrating diverse forms of realistic data into the simulated environment, preserving semantic alignment with the target tasks. For each simulated skill in the environment, DexScale facilitates effective Sim2Real data scaling by automating the process of domain randomization and adaptation. Tuned by the scaled dataset, the control policy achieves zero-shot Sim2Real generalization across diverse tasks, multiple robot embodiments, and widely studied policy model architectures, highlighting its importance in advancing Sim2Real embodied intelligence.
Guiliang Liu, Yueci Deng, Runyi Zhao, Huayi Zhou 0001, Jietao Chen, Ruiyan Xu, Yunxin Tai, Kui Jia
ICML4
2025 GAT-Grasp: Gesture-Driven Affordance Transfer for Task-Aware Robotic Grasping
abstract
Achieving precise and generalizable grasping across diverse objects and environments is essential for intelligent and collaborative robotic systems. However, existing approaches often struggle with ambiguous affordance reasoning and limited adaptability to unseen objects, leading to suboptimal grasp execution. In this work, we propose GAT-Grasp, a gesture-driven grasping framework that directly utilizes human hand gestures to guide the generation of task-specific grasp poses with appropriate positioning and orientation. Specifically, we introduce a retrieval-based affordance transfer paradigm, leveraging the implicit correlation between hand gestures and object affordances to extract grasping knowledge from large-scale human-object interaction videos. By eliminating the reliance on pre-given object priors, GAT-Grasp enables zero-shot generalization to novel objects and cluttered environments. Real-World evaluations confirm its robustness across diverse and unseen scenarios, demonstrating reliable grasp execution in complex task settings.
Huayi Zhou 0001, Xinyue Yao, Guiliang Liu, Kui Jia
IROS2
2024 PBADet: A One-Stage Anchor-Free Approach for Part-Body Association
abstract
The detection of human parts (e.g., hands, face) and their correct association with individuals is an essential task, e.g., for ubiquitous human-machine interfaces and action recognition. Traditional methods often employ multi-stage processes, rely on cumbersome anchor-based systems, or do not scale well to larger part sets. This paper presents PBADet, a novel one-stage, anchor-free approach for part-body association detection. Building upon the anchor-free object representation across multi-scale feature maps, we introduce a singular part-to-body center offset that effectively encapsulates the relationship between parts and their parent bodies. Our design is inherently versatile and capable of managing multiple parts-to-body associations without compromising on detection accuracy or robustness. Comprehensive experiments on various datasets underscore the efficacy of our approach, which not only outperforms existing state-of-the-art techniques but also offers a more streamlined and efficient solution to the part-body association challenge.
Zhongpai Gao, Huayi Zhou 0001, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001
ICLR2
2024 Joint Multi-person Body Detection and Orientation Estimation Via One Unified Embedding
Yiyang Han, Huayi Zhou 0001
PRCV (9)3
2024 BPJDet: Extended Object Representation for Generic Body-Part Joint Detection
abstract
Detection of human body and its parts has been intensively studied. However, most of CNNs-based detectors are trained independently, making it difficult to associate detected parts with body. In this paper, we focus on the joint detection of human body and its parts. Specifically, we propose a novel extended object representation integrating center-offsets of body parts, and construct an end-to-end generic Body-Part Joint Detector (BPJDet). In this way, body-part associations are neatly embedded in a unified representation containing both semantic and geometric contents. Therefore, we can optimize multi-loss to tackle multi-tasks synergistically. Moreover, this representation is suitable for anchor-based and anchor-free detectors. BPJDet does not suffer from error-prone post matching, and keeps a better trade-off between speed and accuracy. Furthermore, BPJDet can be generalized to detect body-part or body-parts of either human or quadruped animals. To verify the superiority of BPJDet, we conduct experiments on datasets of body-part (CityPersons, CrowdHuman and BodyHands) and body-parts (COCOHumanParts and Animals5C). While keeping high detection accuracy, BPJDet achieves state-of-the-art association performance on all datasets. Besides, we show benefits of advanced body-part association capability by improving performance of two representative downstream applications: accurate crowd head detection and hand contact estimation.
Huayi Zhou 0001, Fei Jiang 0006, Jiaxin Si, Yue Ding 0001, Hongtao Lu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Stuart: Individualized Classroom Observation of Students with Automatic Behavior Recognition And Tracking
abstract
Each student matters, but it is hardly for instructors to observe all the students during the courses and provide helps to the needed ones immediately. In this paper, we present StuArt, a novel automatic system designed for the individualized classroom observation, which empowers instructors to concern the learning status of each student. StuArt can recognize five representative student behaviors (hand-raising, standing, sleeping, yawning, and smiling) that are highly related to the engagement and track their variation trends during the course. To protect the privacy of students, all the variation trends are indexed by the seat numbers without any personal identification information. Furthermore, StuArt adopts various user-friendly visualization designs to help instructors quickly understand the individual and whole learning status. Experimental results on real classroom videos have demonstrated the superiority and robustness of the embedded algorithms. We expect our system promoting the development of large-scale individualized guidance of students. More information is in https://github.com/hnuzhy/StuArt.
Huayi Zhou 0001, Fei Jiang 0006, Jiaxin Si, Lili Xiong, Hongtao Lu 0001
ICASSP1
2023 Body-Part Joint Detection and Association via Extended Object Representation
abstract
The detection of human body and its related parts (e.g., face, head or hands) have been intensively studied and greatly improved since the breakthrough of deep CNNs. However, most of these detectors are trained independently, making it a challenging task to associate detected body parts with people. This paper focuses on the problem of joint detection of human body and its corresponding parts. Specifically, we propose a novel extended object representation that integrates the center location offsets of body or its parts, and construct a dense single-stage anchor-based Body-Part Joint Detector (BPJDet). Body-part associations in BPJDet are embedded into the unified representation which contains both the semantic and geometric information. Therefore, BPJDet does not suffer from error-prone association post-matching, and has a better accuracy-speed trade-off. Furthermore, BPJDet can be seamlessly generalized to jointly detect any body part. To verify the effectiveness and superiority of our method, we conduct extensive experiments on the CityPersons, CrowdHuman and BodyHands datasets. The proposed BPJDet detector achieves state-of-the-art association performance on these three benchmarks while maintains high accuracy of detection. Code is released in https://github.com/hnuzhy/BPJDet.
Huayi Zhou 0001, Fei Jiang 0006, Hongtao Lu 0001
ICME1
2023 SSDA-YOLO: Semi-supervised domain adaptive YOLO for cross-domain object detection
Huayi Zhou 0001, Fei Jiang 0006, Hongtao Lu 0001
Comput. Vis. Image Underst.1
2018 Who Are Raising Their Hands? Hand-Raiser Seeking Based on Object Detection and Pose Estimation
abstract
In this paper, we propose an automatic hand-raiser recognition algorithm to show who raise their hands in real classroom scenarios, which is of great importance for further analyzing the learning states of individuals. To recognize the hand-raisers, we divide the hand-raiser recognition into three subproblems, including hand-raising detection, pose estimation, and matching the raised hands to students. Several challenges exist while dealing with the above-mentioned subproblems, such as low resolution of the back row for keypoints detection, the motion distortion caused by hand raising in pose estimation, and various complex situations for matching. To solve these challenges, we first adopt an improved R-FCN algorithm for hand-raising detection, whose effectiveness has been demonstrated. Secondly, we present a novel PAF-based pose estimation algorithm for detecting keypoints of human bodies. The proposed PAF adds scale search and modified weight metric to adapt to the real and complex scenarios. Specifically, scale search improves the detection effect at low resolution by pooling human characteristics in different sizes of pictures, and modified weight metric reasonably utilizes the directional vectors of possible limb connections to optimize the case of motion distortion. Thirdly, a heuristic matching strategy based on the location of hand-raising and keypoints information is proposed to recognize the hand-raisers. Experimental results on six teaching videos in real classrooms have demonstrated the efficiency of the proposed algorithm, and 83% recognition accuracy indicates the potential applications in real classrooms.
Huayi Zhou 0001, Fei Jiang 0006, Ruimin Shen
ACML1