Hongzhi Gao

dblp:154/3784 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-5821-2407ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 20% Trustworthy machine learning · 18% 3D vision · 11%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter modules
0.912025
VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping · AAAI 2025
Computer vision › Segmentation and scene understanding
dense prediction
0.912025
VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping · AAAI 2025
Computer vision › Image recognition and object detection
object detection
0.912025
VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping · AAAI 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping · AAAI 2025
Computer vision › 3D vision
3d object detection
0.812024
Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection · AAAI 2024
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial defense
0.812024
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety · ACL (1) 2024
Natural language and speech › Language models and text generation
large language model
0.812024
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety · ACL (1) 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent safety
0.812024
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety · ACL (1) 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety · ACL (1) 2024
Machine learning › Deep learning architectures and training
downsampling
0.712023
FouriDown: Factoring Down-Sampling into Shuffling and Superposing · NeurIPS 2023
Image and video processing › image restoration
image deblurring
0.712023
FouriDown: Factoring Down-Sampling into Shuffling and Superposing · NeurIPS 2023
Image and video processing
image restoration
0.712023
FouriDown: Factoring Down-Sampling into Shuffling and Superposing · NeurIPS 2023
Image and video processing › image enhancement
low-light image enhancement
0.712023
FouriDown: Factoring Down-Sampling into Shuffling and Superposing · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.312025
VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping · AAAI 2025
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection
0.212024
Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection · AAAI 2024
Robotics › Autonomous driving
perception
0.212024
Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection · AAAI 2024

Methods — techniques the papers use, named apart from their topics

2d discrete fourier transform · 1.3parameter-efficient fine-tuning · 0.9hybrid operation mapping · 0.9teacher-student framework · 0.8self-supervised learning · 0.8psychological profiling · 0.8pseudo-labeling · 0.8defense evaluation · 0.8cross-modal fusion · 0.8adversarial attack · 0.8channel shuffling · 0.7
YearPublicationVenuePosition
2025 VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping
abstract
Although pre-trained large vision foundation models (VFM) yield superior results on various downstream tasks, full fine-tuning is often impractical due to its high computational cost and storage requirements. Recent advancements in parameter-efficient fine-tuning (PEFT) of VFM for image classification show significant promise. However, the application of PEFT techniques to dense prediction tasks remains largely unexplored. Our analysis of existing methods reveals that the underlying premise of utilizing low-rank parameter matrices, despite their efficacy in specific applications, may not be adequately suitable for dense prediction tasks. To this end, we propose a novel PEFT learning approach tailored for dense prediction tasks, namely VFM-Adapter. Specifically, the VFM-Adapter introduces a hybrid operation mapping technique that seamlessly integrates local information with global modeling to the adapter module. It capitalizes on the distinct inductive biases inherent in different operations. Additionally, we dynamically generate parameters for the VFM-Adapter, enabling flexibility of feature extraction given specific inputs. To validate the efficacy of VFM-Adapter, we conduct extensive experiments across object detection, semantic segmentation, and instance segmentation tasks. Results on multiple benchmarks consistently demonstrate the superiority of our method over previous approaches. Notably, with only three percent of the trainable parameters of the SAM-Base backbone, our approach achieves competitive or even superior performance compared to full fine-tuning. The code will be available.
Hongzhi Gao, Lin Chen 0019, Jiaming Liu 0003, Feng Zhao 0004
AAAI4
2024 Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection
abstract
Training high-accuracy 3D detectors necessitates massive labeled 3D annotations with 7 degree-of-freedom, which is laborious and time-consuming. Therefore, the form of point annotations is proposed to offer significant prospects for practical applications in 3D detection, which is not only more accessible and less expensive but also provides strong spatial information for object localization. In this paper, we empirically discover that it is non-trivial to merely adapt Point-DETR to its 3D form, encountering two main bottlenecks: 1) it fails to encode strong 3D prior into the model, and 2) it generates low-quality pseudo labels in distant regions due to the extreme sparsity of LiDAR points. To overcome these challenges, we introduce Point-DETR3D, a teacher-student framework for weakly semi-supervised 3D detection, designed to fully capitalize on point-wise supervision within a constrained instance-wise annotation budget. Different from Point-DETR which encodes 3D positional information solely through a point encoder, we propose an explicit positional query initialization strategy to enhance the positional prior. Considering the low quality of pseudo labels at distant regions produced by the teacher model, we enhance the detector's perception by incorporating dense imagery data through a novel Cross-Modal Deformable RoI Fusion (D-RoI). Moreover, an innovative point-guided self-supervised learning technique is proposed to allow for fully exploiting point priors, even in student models. Extensive experiments on representative nuScenes dataset demonstrate our Point-DETR3D obtains significant improvements compared to previous works. Notably, with only 5% of labeled data, Point-DETR3D achieves over 90% performance of its fully supervised counterpart.
Hongzhi Gao, Lin Chen 0019, Jiaming Liu 0003, Shanghang Zhang, Feng Zhao 0004
AAAI1
2024 PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
abstract
Zaibin Zhang, Yongting Zhang, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zaibin Zhang, Yongting Zhang, Hongzhi Gao, Yu Qiao 0001, Lijun Wang 0001, Huchuan Lu, Feng Zhao 0004
ACL (1)5
2023 FouriDown: Factoring Down-Sampling into Shuffling and Superposing
abstract
Spatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employed in previous approaches. To overcome this limitation, we propose a novel down-sampling paradigm in the Fourier domain, abbreviated as FouriDown, which unifies existing down-sampling techniques. Drawing inspiration from the signal sampling theorem, we parameterize the non-parameter static weighting down-sampling operator as a learnable and context-adaptive operator within a unified Fourier function. Specifically, we organize the corresponding frequency positions of the 2D plane in a physically-closed manner within a single channel dimension. We then perform point-wise channel shuffling based on an indicator that determines whether a channel's signal frequency bin is susceptible to aliasing, ensuring the consistency of the weighting parameter learning. FouriDown, as a generic operator, comprises four key components: 2D discrete Fourier transform, context shuffling rules, Fourier weighting-adaptively superposing rules, and 2D inverse Fourier transform. These components can be easily integrated into existing image restoration networks. To demonstrate the efficacy of FouriDown, we conduct extensive experiments on image de-blurring and low-light image enhancement. The results consistently show that FouriDown can provide significant performance improvements. We will make the code publicly available to facilitate further exploration and application of FouriDown.
Qi Zhu 0010, Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Hongzhi Gao, Chongyi Li, Feng Zhao 0004
NeurIPS5