Binyuan Huang

dblp:297/6767 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HierarchicalGeoCount: Hierarchical scale perception for zero-shot object counting in remote sensing
Binyuan Huang
J. Vis. Commun. Image Represent.1
2026 Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
abstract
The field of autonomous driving increasingly demands high-quality annotated video training data. In this paper, we propose Panacea+, a powerful and universally applicable framework for generating video data in driving scenes. Built upon the foundation of our previous work, Panacea, Panacea+ adopts a multi-view appearance noise prior mechanism and a super-resolution module for enhanced consistency and increased resolution. Extensive experiments show that the generated video samples from Panacea+ greatly benefit a wide range of tasks on different datasets, including 3D object tracking, 3D object detection, and lane detection tasks on the nuScenes and Argoverse 2 dataset. These results strongly prove Panacea+ to be a valuable data generation framework for autonomous driving.
Yuqing Wen, Yingfei Liu, Binyuan Huang, Fan Jia 0006, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.4
2026 Watch Where You Move: Region-Aware Dynamic Aggregation and Excitation for Gait Recognition
abstract
Deep learning-based gait recognition has achieved great success in various applications. The key to accurate gait recognition lies in considering the unique and diverse behavior patterns indifferent motion regions, especially when covariates affect visual appearance. However, existing methods typically use predefined regions for temporal modeling, with fixed or equivalent temporal scales assigned to different types of regions, which makes it difficult to model motion regions that change dynamically over time and adapt to their specific patterns. To tackle this problem, we introduce a Region-aware Dynamic Aggregation and Excitation framework (GaitRDAE) that automatically searches for motion regions, assigns adaptive temporal scales and applies corresponding attention. Specifically, the framework includes two core modules: the Region-aware Dynamic Aggregation (RDA) module, which dynamically searches the optimal temporal receptive field for each region, and the Region-aware Dynamic Excitation (RDE) module, which emphasizes the learning of motion regions containing more stable behavior patterns while suppressing attention to static regions that are more susceptible to covariates. Experimental results show that GaitRDAE achieves state-of-the-art performance on several benchmark datasets. The source code will be published athttps://github.com/HUAFOR/GaitRDAE.
Binyuan Huang, Yongdong Luo, Xianda Guo, Xiawu Zheng, Jiahui Pan 0003, Chengju Zhou
IEEE Trans. Multim.1
2025 SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
abstract
Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production in a way that could continuously improve autonomous driving applications. We investigate the impact of scaling up the quantity of generative data on the performance of downstream perception models and find that enhancing data diversity plays a crucial role in effectively scaling generative data production. Therefore, we have developed a novel model equipped with a subject control mechanism, which allows the generative model to leverage diverse external data sources for producing varied and useful data. Extensive evaluations confirm SubjectDrive's efficacy in generating scalable autonomous driving training data, marking a significant step toward revolutionizing data production methods in this field.
Binyuan Huang, Yuqing Wen, Yaosi Hu, Yingfei Liu, Fan Jia 0006, Weixin Mao, Tiancai Wang, Chi Zhang 0026, Chang Wen Chen, Zhenzhong Chen 0001, Xiangyu Zhang 0005
AAAI1
2025 Attention-aware spatio-temporal learning for multi-view gait-based age estimation and gender classification
abstract
Abstract Recently, gait‐based age and gender recognition have attracted considerable attention in the fields of advertisement marketing and surveillance retrieval due to the unique advantage that gaits can be perceived at a long distance. Intuitively, age and gender can be recognised by observing people's static shape (e.g. different hairstyles between males and females) and dynamic motion (e.g. different walking velocities between the elderly and youth). However, most of the existing gait‐based age and gender recognition methods are based on Gait Energy Image (GEI), which loses the capability of explicitly modelling temporal dynamic information and is not robust to the multi‐view recognition that inevitably happens in a real application. Therefore, in this study, an Attention‐aware Spatio‐Temporal Learning (ASTL) framework is proposed, which employs a silhouette sequence as input to learn essential and invariable spatial‐temporal gait representations. More specifically, a Multi‐Scale Temporal Aggregation (MSTA) module provides an effective scheme for dynamic gait description by exploring and aggregating multi‐scale temporal interval information, which is a core supplement to spatial representation. Then, a Multiple Attention Aggregation (MAA) module is designed to help the network focus on the most discriminatory information along temporal, spatial and channel dimensions. Finally, a Multimodal Collaborative Learning (MCL) block gives full play to the advantages of different modal features through a multimodal cooperative learning strategy. The mean absolute error (MAE) for the age estimation and the correct classification rate (CCR) for the gender classification on OU‐MVLP achieve 6.68 years and 97%, respectively, demonstrating the superiority of the method. Ablation experiments and visualisation results also prove the effectiveness of the three individual modules in their framework.
Binyuan Huang, Yongdong Luo, Jiahui Xie, Jiahui Pan 0003, Chengju Zhou
IET Comput. Vis.1
2024 GaitCTCG: cross-view gait recognition via cascaded residual temporal shift and comprehensive multi-granularity learning
Binyuan Huang, Chengju Zhou, Lewei He, Jiahui Pan 0003
Appl. Intell.1
2023 ICE-GCN: An interactional channel excitation-enhanced graph convolutional network for skeleton-based action recognition
abstract
Abstract Thanks to the development of depth sensors and pose estimation algorithms, skeleton-based action recognition has become prevalent in the computer vision community. Most of the existing works are based on spatio-temporal graph convolutional network frameworks, which learn and treat all spatial or temporal features equally, ignoring the interaction with channel dimension to explore different contributions of different spatio-temporal patterns along the channel direction and thus losing the ability to distinguish confusing actions with subtle differences. In this paper, an interactional channel excitation (ICE) module is proposed to explore discriminative spatio-temporal features of actions by adaptively recalibrating channel-wise pattern maps. More specifically, a channel-wise spatial excitation (CSE) is incorporated to capture the crucial body global structure patterns to excite the spatial-sensitive channels. A channel-wise temporal excitation (CTE) is designed to learn temporal inter-frame dynamics information to excite the temporal-sensitive channels. ICE enhances different backbones as a plug-and-play module. Furthermore, we systematically investigate the strategies of graph topology and argue that complementary information is necessary for sophisticated action description. Finally, together equipped with ICE, an interactional channel excited graph convolutional network with complementary topology (ICE-GCN) is proposed and evaluated on three large-scale datasets, NTU RGB+D 60, NTU RGB+D 120, and Kinetics-Skeleton. Extensive experimental results and ablation studies demonstrate that our method outperforms other SOTAs and proves the effectiveness of individual sub-modules. The code will be published at https://github.com/shuxiwang/ICE-GCN .
Shuxi Wang, Jiahui Pan 0003, Binyuan Huang, Pingzhi Liu, Zina Li, Chengju Zhou
Mach. Vis. Appl.3
2022 GaitMSTP: Multi-Granularity Spatio-Temporal Pyramid for Gait Recognition Under Complex Covariation Conditions
abstract
Gait, with its unique advantage of remote perception without any cooperation from the perceived subject, has become a popular biometric modality for human identity authentication. Diverse spatial representation and temporal modeling are crucial information for gait recognition, especially under covariation conditions. However, most existing algorithms do not fully and explicitly exploit the rich spatial-temporal clues in the gait sequences, leading to a decline in the discriminative ability of gait feature representations. In this paper, we propose a GaitMSTP network for gait recognition under complex covariation conditions, which explicitly models spatio-temporal representations at multi-granularity and multi-sematic levels. More specifically, a Multi-Granularity Temporal Pyramid (MGTP) module is incorporated to extract features at different temporal granularity, which simulates coarse- and fine-grained motion patterns at diverse temporal scales. A Multi-Granularity Spatial Pyramid (MGSP) is designed to capture global and local features at multiple spatial locations and scales. In addition, the multi-granularity spatial-temporal features extracted from shallow to deep semantic levels are further used for supervised learning, aiming to exploit both high-level and low-level of multi-sematic gait characteristics. Extensive experiments on the CASIA-B dataset show that our method outperforms the state-of-the-art algorithms for gait recognition in all scenarios.
Binyuan Huang, Chengju Zhou, Jiahui Pan 0003
IJCB1
2021 HID 2021: Competition on Human Identification at a Distance 2021
abstract
The Competition on Human Identification at a Distance 2021 (HID 2021) is to promote the research in human identification at a distance and to provide a benchmark to evaluate different methods. HID 2021 is the second follow-up from the first one, HID 2020. The dataset size and the evaluation protocal are the same with the previous competition, but the data in the test set has been changed. The paper firstly introduces the dataset and the evaluation protocol, then describes the methods from the top teams and their results. The methods show how to achieve state-of-the-art performance on gait recognition. The results in HID 2021 are better than those in HID 2020. From the comparisons and analysis, some useful conclusions can be drawn. We hope more improvements can be achieved by better followup competitions.
Shiqi Yu 0001, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Edel B. García Reyes, Feng Zheng 0001, Md. Atiqur Rahman Ahad, Beibei Lin, Haijun Xiong, Binyuan Huang
IJCB11