VLDB 2026 Research / reviewers in the wild / expert
Peng Chen 0021
dblp:27/7017-21
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-8076-8989ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical Anonymous Two-Party Gradient Boosting Decision TreeabstractStructured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutually distrustful parties. High speed and interpretability make GBDTs popular in finance and healthcare, where neural networks may fall short. Enabling secure computation for GBDTs poses unique challenges, requiring secure record alignment for comparison. Relying on private set intersection (PSI) is a de facto approach. Mistaking PSI for a safety measure actually exposes which record identifiers (IDs) are shared between the datasets. Although circuit-PSI could help, it is costly for generic uses. New ideas are needed to efficiently train in a "dark forest". Aiming to hide the IDs, we initiate the study of anonymous GBDT training on split data held by two parties. Dual circuit-PSI in our design lets the parties alternate as receiver to run pick-then-sum over local features. Via oblivious programmable pseudorandom functions, we propagate circuit-PSI outputs as shared state across runs. Avoiding universal alignment, we resolve the neglected dilemma that ID hiding incurs a cost that scales with domain size. Next, we halve the cost of ciphertext packing used to convert single-instruction multiple-data homomorphic encryption from (ring) learning with errors in prior secure GBDT (Usenix Security' 23) and related secure machine-learning computations. Comparative experiments show our protocol remains competitive with leaky approaches in efficiency. Enabling ID-hiding aggregation, our techniques can extend to other vertically partitioned analytics. Minxin Du, Sherman S. M. Chow, Huangxun Chen, Huaming Rao, Danqing Huang, Peng Chen 0021 |
SP | 9 |
| 2026 | Multi-modal vehicle trajectory prediction via hierarchical attention and raster-vector maps encoding in unstructured road environments
Zhifa Chen, Peng Chen 0021, Songyue Yang, Rentao Sun, Guizhen Yu |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | CosineOpt: Optimization-Based Centralized Cooperative Speed Planning for Multiple CAVs Along Intersected Fixed PathsabstractThis paper focuses on cooperative speed planning for multiple connected and automated vehicles (CAVs) traversing along intersected fixed paths. Nominally, this task is formulated as an optimal control problem incorporating logical operators to represent collision-avoidance constraints. This formulation requires solving a mixed-integer nonlinear programming (MINLP) problem, while handling non-differentiable integer variables remains challenging for gradient-based solvers. Instead of solving the MINLP, we propose a cosine-based method, a novel geometric strategy for formulating collision-avoidance constraints between CAVs. Constructing such a geometric model introduces potential approximation errors, which are mitigated by fitted correction terms designed to compensate for geometric deviations and refine the distance calculation. We propose a simulation-based planner to provide the speed profile with the globally optimal passing order, serving as a warm start for the solver. A lightweight iterative optimization strategy is also adopted to enhance robustness. Additionally, we propose a fault-tolerant strategy to ensure both system safety and operational efficiency. Extensive simulation results verify the proposed method, and comparative experiments demonstrate its efficiency. Bai Li 0002, Peng Chen 0021, Guizhen Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | Dynamic OD Estimation Using Spatio-Temporal Hybrid Graph Convolutional Network With Asynchronous Multi-Source DataabstractReliable estimation of dynamic Origin-Destination (OD) matrices is essential for urban traffic management but remains a critical challenge. Emerging data sources, such as Connected Vehicle (CV) trajectories and Automatic Vehicle Identification (AVI) records, provide valuable insights for dynamic OD estimation. However, most existing approaches overlook the asynchronous nature of real-world data collection, limiting their applicability in complex urban environments. This study proposes a novel k-Nearest Neighbors Dynamic Adaptive Hybrid Graph Encoder (k-DAHGE) model that leverages the daily periodicity of travel demand to integrate asynchronous CV and AVI data for dynamic OD estimation. Firstly, we develop a spatio-temporal dependency modeling framework that addresses temporal inconsistencies among multi-source data by constructing dynamic OD relation graphs from historical CV-OD matrices. Then, we introduce a Dynamic Adaptive Hybrid Graph Encoder with learnable mixing parameters that dynamically fuses different graph convolution networks for multi-scale feature extraction. Experiments on a real-world urban road network demonstrate that our approach achieves a mean absolute error of 10.82 veh/30min for dynamic OD estimation, significantly outperforming popular neural networks and existing models. Further analyses under varying AVI coverage levels confirm the robustness and generalizability of the proposed approach even with limited AVI detection. Peng Chen 0021, Ying Guo 0008, Sheng Dong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2026 | Risk-Tolerant On-Site Dispatch for Autonomous Mining Truck Fleets With Uncertain Failure SignsabstractThis study proposes a risk-tolerant dispatch approach for a fleet of autonomous mining trucks in an open-pit mine, leveraging early signs to reduce the impact of potential failures that may or may not occur later. Unlike traditional methods that ignore these early signs or wait until a failure has actually happened, our approach proactively plans for both possible outcomes without relying on probabilities. We propose a Y-shaped solution structure composed of a shared trunk that covers the period before it becomes clear if the failure will occur and two separate branches that address the final scenarios. We formulate the dispatch problem as a mixed-integer linear program and solve it via Gurobi. To facilitate the solution process with Gurobi, an evolutionary algorithm is adopted to explore the solution space for a good initial guess. A high-performance discrete-event simulator is embedded in the cost function evaluation module of the evolutionary algorithm for quickly selecting qualified solution candidates. By integrating both the failure and non-failure scenarios into one unified plan, we avoid extreme risk-taking or undue conservatism, ensuring stable operational performance. Simulations and field trials at a real open-pit mine confirm that this risk-tolerant approach effectively manages failure risks when early signs are available. Rentao Sun, Guizhen Yu, Bin Zhou 0007, Peng Chen 0021, Bai Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | DataLab: A Unified Platform for LLM-Powered Business IntelligenceabstractBusiness intelligence (BI) transforms large volumes of data within modern organizations into actionable insights for informed decision-making. Recently, large language model (LLM)-based agents have streamlined the BI workflow by automatically performing task planning, reasoning, and actions in executable environments based on natural language (NL) queries. However, existing approaches primarily focus on individual BI tasks such as NL2SQL and NL2VIS. The fragmentation of tasks across different data roles and tools lead to inefficiencies and potential errors due to the iterative and collaborative nature of BI. In this paper, we introduce DataLab, a unified BI platform that integrates a one-stop LLM-based agent framework with an augmented computational notebook interface. DataLab supports various BI tasks for different data roles in data preparation, analysis, and visualization by seamlessly combining LLM assistance with user customization within a single environment. To achieve this unification, we design a domain knowledge incorporation module tailored for enterprise-specific BI tasks, an inter-agent communication mechanism to facilitate information sharing across the BI workflow, and a cell-based context management strategy to enhance context utilization efficiency in BI notebooks. Extensive experiments demonstrate that DataLab achieves state-of-the-art performance on various BI tasks across popular research benchmarks. Moreover, DataLab maintains high effectiveness and efficiency on real-world datasets from Tencent, achieving up to a 58.58% increase in accuracy and a 61.65 % reduction in token cost on enterprise-specific BI tasks. Luoxuan Weng, Yinghao Tang, Yingchaojie Feng, Zhuo Chang, Ruiqin Chen, Haozhe Feng, Chen Hou, Danqing Huang, Yang Li 0106, Huaming Rao, Canshi Wei, Xiuqi Huang, Minfeng Zhu 0001, Yuxin Ma 0001, Bin Cui 0001, Peng Chen 0021, Wei Chen 0001 |
ICDE | 20 |
| 2025 | SiriusBI: A Comprehensive LLM-powered Solution for Data Analytics in Business IntelligenceabstractWith the proliferation of Large Language Models (LLMs) in Business Intelligence (BI), existing solutions face critical challenges in industrial deployments: functionality deficiencies from legacy systems failing to meet evolving LLM-era user demands, interaction limitations from single-round SQL generation paradigms inadequate for multi-round clarification, and cost for domain adaptation arising from cross-domain methods migration. We present SiriusBI, a practical LLM-powered BI system addressing the challenges of industrial deployments through three key innovations: (a) An end-to-end architecture integrating multi-module coordination to overcome functionality gaps in legacy systems; (b) A multi-round dialogue with querying mechanism, consisting of semantic completion, knowledge-guided clarification, and proactive querying processes, to resolve interaction constraints in SQL generation; (c) A data-conditioned SQL generation method selection strategy that supports both an efficient one-step Fine-Tuning approach and a two-step method leveraging Semantic Intermediate Representation for low-cost cross-domain applications. Experiments on both real-world datasets and public benchmarks demonstrate the effectiveness of SiriusBI. User studies further confirm that SiriusBI enhances both productivity and user experience. As an independent service on Tencent's data platform, SiriusBI is deployed across finance, advertising, and cloud sectors, serving dozens of enterprise clients. It achieves over 93% accuracy in SQL generation and reduces data analysts' query time from minutes to seconds in real-world applications. Jie Jiang 0015, Haining Xie, Yu Shen 0003, Meng Lei, Yang Li 0106, Chunyou Li, Danqing Huang, Yinjun Wu, Wentao Zhang 0001, Bin Cui 0001, Peng Chen 0021 |
Proc. VLDB Endow. | 15 |
| 2025 | Real-Time Cooperative Trajectory Planning for Multiple CAVs at Unstructured Intersections: A Computational Optimal Control Approach
Bai Li 0002, Peng Chen 0021, Guizhen Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Hybrid Path Tracking Control for Autonomous Trucks: Integrating Pure Pursuit and Deep Reinforcement Learning With Adaptive Look-Ahead MechanismabstractPath tracking control is essential for ensuring the safe and efficient operation of autonomous trucks, but traditional methods often struggle with nonlinear vehicle dynamics. While deep reinforcement learning (DRL) approaches are model-free, they may lack the stability and interpretability required for reliable deployment. This study presents a hybrid control framework that combines Pure Pursuit (PP) with Proximal Policy Optimization (PPO) to enhance tracking accuracy and robustness. PP provides baseline stability and interpretability, while PPO refines control actions by optimizing policy gradients, ensuring better adaptability to nonlinear dynamics and complex driving conditions. An adaptive look-ahead mechanism, responsive to speed and curvature, dynamically adjusts preview distances using PPO-generated coefficients, facilitating early corrections during high-speed turns and enabling greater precision on sharp curves. A fusion training method, leveraging high-reward initialization and a decreasing learning rate, supports efficient exploration and stable convergence. The approach was validated in a high-fidelity simulation environment using PreScan, Simulink, and ROS, along with real-world experiments on a proportionally scaled intelligent vehicle chassis, demonstrating notable improvements in path tracking accuracy and robustness across varied path profiles. Zhixuan Han, Peng Chen 0021, Bin Zhou 0007, Guizhen Yu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | RailFusion: A Lidar-Camera Data Interaction Network for 3-D Railway Object DetectionabstractAccurate detection of 3D objects is vital to the perception of the railway environments for the safe operation autonomous trains, particularly given the complexity of railway environments and the challenges in detecting objects of variable sizes and of distant objects. This study introduces RailFusion, a LiDAR-Camera fusion network for integrating multi-modal features. RailFusion consists of two main modules: Cross-Domain Feature Extraction (CDFE) and Multi-Modal Fusion (MMF). Specifically, the CDFE module is designed with a novel feature extraction method to enhance cross-domain features interaction by utilizing LiDAR spatial depth and image semantic information. The MMF module uses deformable attention for aligning and fusing multi-modal features. Further to this, the channel normalization fusion is proposed to assign channel weights. Experimental results show that the mean average-precision (mAP) of our proposed RailFusion is 57.2%, which is 8.4% higher than the baseline 3D object detection network BEVFusion. Moreover, the results show that RailFusion is applicable to long-range detection as well as for detecting varying sized and short-range objects. All these indicate that RailFusion has the potential to be readily applicable in 3D object detection in railway environments. Guizhen Yu, Peng Chen 0021 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | WaveCRNet: Wavelet Transform-Guided Learning for Semantic Segmentation in Adverse Railway ScenesabstractSemantic segmentation is pivotal in autonomous train perception, significantly impacting the system’s intelligence and reliability. However, its performance in railway scenes is hindered by various challenges, including severe weather conditions, low-light situations, tunnel settings, and diverse and dynamic unstructured scenes. To address these challenges, this study proposes WaveCRNet, a novel architecture for real-time semantic segmentation in challenging conditions, simulating wavelet-constrained PID controller in feature and wavelet space. This study first designs an effective wavelet information enhancement algorithm using a differentiable wavelet transform to bridge the gap between the wavelet and feature information domains. Then, the wavelet-guided attention pag module (WAPM) is introduced to guide the learning and fusion of detailed features based on wavelet priors. Moreover, the traditional wavelet transforms are spectral aliasing and shift sensitivity. Inspired by dual-tree complex wavelet transform (DTCWT), the DTCWT-based channel reconstruction module (DCRM) is designed to assist the channel-based learning of boundary information from coarse to fine. Finally, the proposed architecture is evaluated by the public dataset RailSem19. The experimental results validate the consistent performance gains between accuracy and inference speed, improving by 63.7% mIoU and 87 FPS, surpassing those of the advanced methods for real-time segmentation. Zhihao Liao 0001, Pengcheng Wang 0003, Peng Chen 0021, Wenwen Luo |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Dual-Layer Path Planning for Unmanned Ground Vehicles Based on Probabilistic Roadmap and Proximal Policy OptimizationabstractAddressing the crucial challenge of autonomous navigation for unmanned ground vehicles (UGVs), this paper presents a dual-layer path planning method integrating Probabilistic Roadmap (PRM) and Proximal Policy Optimization (PPO). Combining global guidance with local optimization, this approach effectively mitigates the shortcomings of traditional path planning methods such as blindness and local optimality, thus enhancing the efficiency and feasibility of path planning. Specifically, we propose a PRM-RL dual-layer path planning framework that employs the PRM algorithm to generate sub-goals for guiding reinforcement learning exploration, thereby improving training efficiency. Simultaneously, we utilize the PPO algorithm to optimize paths, considering vehicle kinematics and introducing soft constraints to ensure smoother paths adaptable to diverse application scenarios. The superiority and practicality of our method are validated through ablation experiments and comparative experiments, offering a reliable path planning solution for autonomous navigation of UGVs. Zhixuan Han, Peng Chen 0021, Bin Zhou 0007, Guizhen Yu |
INDIN | 2 |
| 2024 | Event-Triggered Mechanism-Based MPC for Path-Tracking Control of Four-Wheel Steering VehiclesabstractIn this study, we tackle the path-tracking problem of a nonlinear four-wheel steering vehicle dynamics model subject to model mismatches and propose a model predictive control (MPC) algorithm based on an event-triggered mechanism (ET -MPC). The goal is to maintain closed-loop control performance while reducing the computational and communication burdens of traditional MPC. We introduce an ET -MPC framework utilizing a model-free reinforcement learning agent with proximal policy optimization (PPO). This agent interacts with the MPC system, progressively learning to determine the optimal event-triggered mechanism. To enhance exploration and training efficiency, we incorporate the Long Short-Term Memory (LSTM) technique into PPO. Experimental results show that the proposed ET -MPC framework, combined with reinforcement learning for reward optimization, demonstrates superior overall performance in path-tracking control of four-wheel steering vehicles. Guoyan Xu, Han Li 0007, Peng Chen 0021, Qi Xia 0002, Han Cai |
INDIN | 4 |
| 2024 | Dynamic Origin-Destination Flow Imputation Using Feature-Based Transfer LearningabstractReal-time and full-sample vehicle origin-destination (OD) information is essential for traffic management and control in urban road network. However, the low coverage of automatic vehicle identification (AVI) detection devices leads to difficulty in estimating OD. As an emerging traffic data, the trajectories of connected vehicles (CVs) can effectively provide information on their origin and destination. To this end, this paper presents a framework of an autoencoder network utilizing feature transfer to estimate urban dynamic OD based on the characteristics of two data sources. Specifically, a generative adversarial network is introduced to learn high-dimensional feature that is domain-invariant in two data domains. In addition, a pre-training fine-tuning approach is proposed to transfer knowledge pretrained from CV data to the limited AVI observation for OD imputation. Finally, the model was subjected to a real-world road network test. The results showed that for all OD flows the relative error was 11.23 vehicles/30 minutes, which outperformed baseline models, including popular neural networks and existing estimation models for multi-source data fusion. Furthermore, the model’s robustness to external factors, such as observation conditions and data quality, was examined. The results demonstrated that the model consistently delivers satisfactory estimation performance across a diverse range of conditions. Peng Chen 0021, Bin Zhou 0007, Guizhen Yu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Towards General and Efficient Online Tuning for SparkabstractThe distributed data analytic system - Spark is a common choice for processing massive volumes of heterogeneous data, while it is challenging to tune its parameters to achieve high performance. Recent studies try to employ auto-tuning techniques to solve this problem but suffer from three issues: limited functionality, high overhead, and inefficient search. In this paper, we present a general and efficient Spark tuning framework that can deal with the three issues simultaneously. First, we introduce a generalized tuning formulation, which can support multiple tuning goals and constraints conveniently, and a Bayesian optimization (BO) based solution to solve this generalized optimization problem. Second, to avoid high overhead from additional offline evaluations in existing methods, we propose to tune parameters along with the actual periodic executions of each job (i.e., online evaluations). To ensure safety during online job executions, we design a safe configuration acquisition method that models the safe region. Finally, three innovative techniques are leveraged to further accelerate the search process: adaptive sub-space generation, approximate gradient descent, and meta-learning method. We have implemented this framework as an independent cloud service, and applied it to the data platform in Tencent. The empirical results on both public benchmarks and large-scale production tasks demonstrate its superiority in terms of practicality, generality, and efficiency. Notably, this service saves an average of 57.00% memory cost and 34.93% CPU cost on 25K in-production tasks within 20 iterations, respectively. Yang Li 0106, Huaijun Jiang, Yu Shen 0003, Yide Fang, Danqing Huang, Xinyi Zhang 0002, Wentao Zhang 0001, Ce Zhang 0001, Peng Chen 0021, Bin Cui 0001 |
Proc. VLDB Endow. | 10 |
| 2022 | FarNet: An Attention-Aggregation Network for Long-Range Rail Track Point Cloud SegmentationabstractRail track segmentation is key to environmental perception of autonomous train. However, due to the complexity of railway track environment, critical issues such as the detection of rail tracks with different curvatures remain to be overcome. In this study, a novel architecture called FarNet is proposed for long-range railway track point cloud segmentation. The proposed FarNet is mainly divided into three parts, i.e., spherical projection, attention-aggregation network and results refinement. Specifically, spherical projection converts the LiDAR point cloud into a pseudo range image, and attention-aggregation network enables railway track detection using the pseudo range image. Furthermore, in the attention-aggregation network two components, i.e., spatial attention module and information aggregation module, are proposed to enhance the capability of rail track segmentation. Last, the results refinement helps further filter out the noise points after segmentation. Experimental results show that the proposed FarNet achieved 98.0% mean intersection-over-union (MIoU) and 98.9% mean pixel accuracy (MPA) for rail track segmentation. Guizhen Yu, Peng Chen 0021, Bin Zhou 0007, Songyue Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | A Particle Filter-Based Approach for Vehicle Trajectory Reconstruction Using Sparse Probe DataabstractTrajectory data collected from probe vehicles become increasingly important for urban traffic operation and management. However, current data tend to be sparse in time and space due to technical constraints or privacy concerns, which fail to provide a complete picture of traffic flow. This study proposes a particle filter (PF) based approach to reconstruct the vehicle trajectory for signalized arterial using sparse probe data. First, the arterial intersection is divided into multiple road cells and the estimation of cell travel time is formulated as a quadratic programming problem. Then, PF is applied to reconstruct the incomplete vehicle trajectory between consecutive updates. Specifically, to calculate and update the weight of initial particles, three measurability criteria are designed for importance sampling considering the structure of signalized arterial and the feature of vehicular updates, i.e., travel time adjustment accuracy, arterial link speed limit and travel time adjustment possibility. Last, NGSIM trajectory data are extracted at intervals to construct the sparse data, which are used to verify the effectiveness of the proposed method. The results show that reconstructed trajectories match closely with ground truth both at the single intersection and along the arterial with multiple intersections. Peng Chen 0021 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Reliable shortest path finding in stochastic time-dependent road network with spatial-temporal link correlations: A case study from Beijing
Peng Chen 0021, Rui Tong, Bin Yu 0018 |
Expert Syst. Appl. | 1 |
| 2020 | Cycle-Based End of Queue Estimation at Signalized Intersections Using Low-Penetration-Rate Vehicle TrajectoriesabstractQueue length is a crucial measure of intersection performance. Probe vehicles (PVs) with advanced sensors are capable of recording vehicle trajectories that can be used to estimate queue length, a technique of which has received considerable attention in the past decade. Noticeably, this technique usually requires high PV penetration rates (e.g., above 25%) in order to ensure estimation accuracy. Though the PVs are expected to increase, their penetration rate will still remain relatively low in the near future. Meanwhile, the initial queue length is another important factor that directly relates to queue dynamics at each cycle. However, most of the studies failed to adequately account for the effect of the initial queue on cyclic queue length estimation. To address the above challenges, this paper proposes a cycle-based end of queue estimation method using sampled vehicle trajectory data under relatively low penetration rates. Two major steps are involved: first, vehicle arrival process is modeled as a certain distribution in line with traffic conditions and an expectation maximum (EM) procedure is employed to estimate the arrival rate of each cycle; then, both ends of the queue and initial queue are estimated at each cycle based on shockwave theory. Microscopic traffic simulator VISSIM is utilized to examine the performance of the method. The experimental results reveal that the cycle-based end of the queue can be estimated with desirable accuracy in different scenarios, e.g., undersaturated, oversaturated, and queue spillback conditions. The comparison with the state-of-the-art methods further helps to verify the advantage of the method, especially under low-penetration-rate conditions. Henry X. Liu, Peng Chen 0021, Guizhen Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | The α-reliable path problem in stochastic road networks with link correlations: A moment-matching-based path finding algorithm
Peng Chen 0021, Rui Tong, Guangquan Lu |
Expert Syst. Appl. | 1 |