VLDB 2026 Research / reviewers in the wild / expert
Hongyu Pan
dblp:230/1429
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDSSM: Sparse Dual-domain State-Space Modeling for physics-prior-guided underwater image enhancement
Dingshuo Liu, Mingrui Kong, Hongyu Pan, Qingling Duan |
Pattern Recognit. | 3 |
| 2025 | Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving frameworks enable seamless integration of perception and planning but often rely on one-shot trajectory prediction, which may lead to unstable control and vulnerability to occlusions in single-frame perception. To address this, we propose the Momentum-Aware Driving (MomAD) framework, which introduces trajectory momentum and perception momentum to stabilize and refine trajectory predictions. MomAD comprises two core components: (1) Topological Trajectory Matching (TTM) employs Hausdorff Distance to select the optimal planning query that aligns with prior paths to ensure coherence; (2) Momentum Planning Interactor (MPI) cross-attends the selected planning query with historical queries to expand static and dynamic perception files. This enriched query, in turn, helps regenerate long-horizon trajectory and reduce collision risks. To mitigate noise arising from dynamic environments and detection errors, we introduce robust instance denoising during training, enabling the planning model to focus on critical signals and improve its robustness. We also propose a novel Trajectory Prediction Consistency (TPC) metric to quantitatively assess planning stability. Experiments on the nuScenes dataset demonstrate that MomAD achieves superior long-term consistency (≥ 3s) compared to SOTA methods. Moreover, evaluations on the curated Turning-nuScenes shows that MomAD reduces the collision rate by 26% and improves TPC by 0.97m (33.45%) over a 6s prediction horizon, while closed- loop on Bench2Drive demonstrates an up to 16.3% improvement in success rate. The source code is available at https://github.com/adept-thu/MomAD. Ziying Song, Caiyan Jia, Hongyu Pan, Shaoqing Xu, Lei Yang 0060, Yadan Luo |
CVPR | 4 |
| 2025 | Digital twin based intelligent control system on gas extraction from boreholes and experimental researchabstractIntelligent gas extraction in mines is a critical enabler for the realization of smart mine development. As an essential component of this process, the intelligent control of borehole gas extraction relies on real-time monitoring data from numerous extraction parameters, integrated with next-generation information technologies, to achieve objectives such as intelligent deployment of negative pressure, control of inefficient boreholes, and evaluate extraction effectiveness. This paper constructs an intelligent control system for borehole gas extraction based on a digital twin “Four-Dimensional” framework, enabling bidirectional mapping between physical entities and virtual digital twins through the integration of physical entities (PE) as foundational carriers, virtual entities (VE) as three-dimensional models, digital twin data (DD) as the control core, and services (SS) and connectivity (CN) as the methodological framework. Key technologies include developing a control model for the gas flow process from coal seam to borehole, processing multi-source extraction data using data fusion methods, constructing virtual twins with SolidWorks, and designing control schemes based on Model Predictive Control (MPC) algorithm. In order to verify the rationality and feasibility of the system, a pilot study on the digital twin for intelligent control of borehole gas extraction was carried out. The results show that the concentration of gas extraction after control rises significantly, and the gas extraction concentration of borehole 1# reaches the optimal range when the valve is opened to 75 %, while that of borehole 2# reaches the maximum concentration of the adjustable range when the valve is opened to 70 %, which meets the control target. Suinan He, Hongyu Pan, Tianjun Zhang, Xinshuang Cao, Yilun Xue |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Few-shot graph classification on cross-site scripting attacks detection
Hongyu Pan, Yong Fang 0002, Wenbo Guo 0011, Yijia Xu, Changhui Wang |
Comput. Secur. | 1 |
| 2024 | SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object DetectionabstractLiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12%$\text {AP}_{\text {3-D}}$on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released inhttps://github.com/liulin813/SparseDet.git. Ziying Song, Qiming Xia, Feiyang Jia, Caiyan Jia, Lei Yang 0060, Hongyu Pan |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | BE-STI: Spatial-Temporal Integrated Network for Class-agnostic Motion Prediction with Bidirectional EnhancementabstractDetermining the motion behavior of inexhaustible categories of traffic participants is critical for autonomous driving. In recent years, there has been a rising concern in performing class-agnostic motion prediction directly from the captured sensor data, like LiDAR point clouds or the combination of point clouds and images. Current motion prediction frameworks tend to perform joint semantic segmentation and motion prediction and face the trade-off between the performance of these two tasks. In this paper, we propose a novel Spatial-Temporal Integrated network with Bidirectional Enhancement, BE-STI, to improve the temporal motion prediction performance by spatial semantic features, which points out an efficient way to combine semantic segmentation and motion prediction. Specifically, we propose to enhance the spatial features of each individual point cloud with the similarity among temporal neighboring frames and enhance the global temporal features with the spatial difference among non-adjacent frames in a coarse-to-fine fashion. Extensive experiments on nuScenes and Waymo Open Dataset show that our proposed framework outperforms all state-of-the-art LiDAR-based and RGB+LiDAR-based methods with remarkable margins by using only point clouds as input.11The code will be released at https://github.com/be-sti/be-sti. Yunlong Wang 0009, Hongyu Pan, Yu-Huan Wu, Xin Zhan, Kun Jiang 0002, Diange Yang |
CVPR | 2 |
| 2022 | INT: Towards Infinite-Frames 3D Detection with an Efficient Framework
Jianyun Xu, Zhenwei Miao, Hongyu Pan, Peihan Hao, Zhengyang Sun, Xin Zhan |
ECCV (9) | 4 |
| 2022 | CPGNet: Cascade Point-Grid Fusion Network for Real-Time LiDAR Semantic SegmentationabstractLiDAR semantic segmentation essential for advanced autonomous driving is required to be accurate, fast, and easy-deployed on mobile platforms. Previous point-based or sparse voxel-based methods are far away from real-time applications since time-consuming neighbor searching or sparse 3D convolution are employed. Recent 2D projection-based methods, including range view and multi-view fusion, can run in real time, but suffer from lower accuracy due to information loss during the 2$D$projection. Besides, to improve the performance, previous methods usually adopt test time augmentation (TTA), which further slows down the inference process. To achieve a better speed-accuracy trade-off, we propose Cascade Point-Grid Fusion Network (CPGNet), which ensures both effectiveness and efficiency mainly by the following two techniques: 1) the novel Point-Grid (PG) fusion block extracts semantic features mainly on the 2D projected grid for efficiency, while summarizes both 2D and 3D features on 3D point for minimal information loss; 2) the proposed transformation consistency loss narrows the gap between the single-time model inference and TTA. The experiments on the SemanticKITTI and nuScenes benchmarks demonstrate that the CPGNet without ensemble models or TTA is comparable with the state-of-the-art RPVNet, while it runs 4.7 times faster. Hongyu Pan |
ICRA | 3 |
| 2021 | PVGNet: A Bottom-Up One-Stage 3D Object Detector With Integrated Multi-Level FeaturesabstractQuantization-based methods are widely used in LiDAR points 3D object detection for its efficiency in extracting context information. Unlike image where the context information is distributed evenly over the object, most LiDAR points are distributed along the object boundary, which means the boundary features are more critical in LiDAR points 3D detection. However, quantization inevitably introduces ambiguity during both the training and inference stages. To alleviate this problem, we propose a one-stage and voting-based 3D detector, named Point-Voxel-Grid Network (PVGNet). In particular, PVGNet extracts point, voxel and grid-level features in a unified backbone architecture and produces point-wise fusion features. It segments Li-DAR points into foreground and background, predicts a 3D bounding box for each foreground point, and performs group voting to get the final detection results. Moreover, we observe that instance-level point imbalance due to occlusion and observation distance also degrades the detection performance. A novel instance-aware focal loss is proposed to alleviate this problem and further improve the detection ability. We conduct experiments on the KITTI and Waymo datasets. Our proposed PVGNet outperforms previous state-of-the-art methods and ranks at the top of KITTI 3D/BEV detection leaderboards. Zhenwei Miao, Jikai Chen, Hongyu Pan, Ruiwen Zhang, Peihan Hao, Xin Zhan |
CVPR | 3 |
| 2021 | Deep Conditional Distribution Learning for Age EstimationabstractAge estimation is a challenging task not only because face appearance is affected by illumination, pose, and expression, but also because there exists age label ambiguity among different demographic groups. In this work, we first revisit different label distribution learning (LDL) based age estimation methods and propose a more general formulation, which can unify individual LDL-based age estimation methods, as well as the traditional regression, classification, and ranking based age estimation methods. Based on such a general formulation, we propose a novel deep conditional distribution learning (DCDL) method, which can flexibly leverage a varying number of auxiliary face attributes to achieve adaptive age-related feature learning and improve age estimation robustness against the challenges above. Experimental results on multiple age estimation datasets (MORPH II, AgeDB, FG-NET, MegaAge-Asian, CLAP2016, UTK-Face, and LFW+) show that the proposed approach outperforms the state-of-the-art age estimation methods by a large margin. In addition, the proposed approach can generalize well to other human attributes estimation tasks, like height, weight, and body mass index (BMI) estimation. Haomiao Sun, Hongyu Pan, Hu Han 0001, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Mean-Variance Loss for Deep Age Estimation From a FaceabstractAge estimation has wide applications in video surveillance, social networking, and human-computer interaction. Many of the published approaches simply treat age estimation as an exact age regression problem, and thus do not leverage a distribution's robustness in representing labels with ambiguity such as ages. In this paper, we propose a new loss function, called mean-variance loss, for robust age estimation via distribution learning. Specifically, the mean-variance loss consists of a mean loss, which penalizes difference between the mean of the estimated age distribution and the ground-truth age, and a variance loss, which penalizes the variance of the estimated age distribution to ensure a concentrated distribution. The proposed mean-variance loss and softmax loss are jointly embedded into Convolutional Neural Networks (CNNs) for age estimation. Experimental results on the FG-NET, MORPH Album II, CLAP2016, and AADB databases show that the proposed approach outperforms the state-of-the-art age estimation methods by a large margin, and generalizes well to image aesthetics assessment. Hongyu Pan, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2018 | Revised Contrastive Loss for Robust Age Estimation from FaceabstractAge estimation has broad applications in many fields, such as video surveillance, social networking, and human-computer interaction. Many of the existing approaches treat age estimation as a classification problem; however, the individual age values are not independent classes; they have an ordinal relationship. Classification loss such as softmax is not able to model such kind of relationship. In this paper, we propose a new loss, called revised contrastive loss, to model the ordinal relationship of individual ages. Specifically, the revised contrastive loss is proposed to penalize the distance between two face images in the feature space according to their age difference, which makes the learned features more discriminative for the age estimation task. We embed the proposed revised contrastive loss and softmax loss into a Convolutional Neural Network (CNN), and optimize the networks via Stochastic Gradient Descent (SGD) in an end-to-end fashion. Experimental results on a number of challenging face aging databases (FG-NET, MORPH Album II, and CLAP2016) show that the proposed approach outperforms the state-of-the-art methods by a large margin using a single model. Hongyu Pan, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ICPR | 1 |