VLDB 2026 Research / reviewers in the wild / expert
Haocheng Peng
dblp:354/7020
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0009-7158-1002ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bidirectional Temporal-Aware Modeling with Multi-Scale Mixture-of-Experts for Multivariate Time Series ForecastingabstractRecent advances in deep learning have significantly boosted performance in multivariate time series forecasting (MTSF). While many existing approaches focus on capturing inter-variable (a.k.a. channel-wise) correlations to improve prediction accuracy, the temporal dimension, particularly its rich structural and contextual information, remains underexplored. In this paper, we propose BIM3, a novel framework that integrates BIdirectional temporal-aware modeling with Multi-Scale Mixture-of-Experts for MTSF. First, unlike existing methods that treat historical and future temporal information independently, we introduce a novel Timestamp Dual Cross-Attention Module, which employs a symmetric cross-attention mechanism to explicitly capture bidirectional temporal dependencies through timestamp interactions. Second, to address the complex and scale-varying temporal patterns commonly found in multivariate time series, we move beyond recent multi-scale forecasting models that share parameters across all channels and fail to capture channel-specific dynamics. Instead, we design a Multi-Scale Feature Extract Mixture-of-Experts module that adaptively routes time series to specialized experts based on their temporal characteristics. Extensive experiments on multiple real-world datasets show that BIM3 consistently outperforms state-of-the-art methods, highlighting its effectiveness in capturing both temporal structure and inter-variable diversity. Yifan Gao 0012, Boming Zhao, Haocheng Peng, Hujun Bao, Jiashu Zhao, Zhaopeng Cui |
CIKM | 3 |
| 2025 | Robust Robotic Assembly of Reusable, Rectangular BlocksabstractThis paper investigates the importance and design implications for use of rectangular blocks in collective robotic construction systems with distributed control. Specifically, we introduce an automated solver for optimizing the overlaps in user-specified structures; a new robot design capable of manipulating, fastening, and climbing over blocks as wide as the robot; detailed analysis of robot primitives and demonstration of rectilinear, curved, cantilever, and corbeled arch structures; and results from a physics simulator showing how overlaps improve structural integrity when the depositions are noisy. This work represents an important step towards efficient and versatile large-scale robotic construction. Zhongming Huang, Haocheng Peng, Shih-Ming Lin, Kirstin Petersen, Nils Napp |
IROS | 3 |
| 2025 | Decentralized Request Dispatch for Edge-Clouds: A Diffusion-Based Reinforcement Learning ParadigmabstractEdge-cloud systems have the potential to achieve ubiquitous computing by providing services in close proximity to users that submit service requests. The key challenge is how to efficiently orchestrate services and dispatch requests to satisfy the Quality of Service (QoS) requirements of users in dynamic edge-cloud environments. With the benefit of efficiently adapting to uncertainty, Reinforcement Learning (RL) based approaches are proposed to solve the request dispatch problem in edge-cloud environments. However, existing RL based approaches are often constrained by inexpressive policies that make highly suboptimal decisions in the field of request dispatch for edge-clouds. To enhance the effectiveness of RL in guaranteeing the QoS requirements of users, this paper presents D2Sched, a novel scheduling framework that represents the policy networks of Multi-Agent Deep Reinforcement Learning (MADRL) as diffusion models to generate request dispatch decisions. To improve the valid probability of generated dispatch decisions, D2Sched coordinates all agents for the resource competition among different requests by carefully considering the availability of system resources and latency targets of requests. Extensive experiments using synthetic and real traces demonstrate that D2Sched can improve the average system throughput under QoS requirements of users by up to 20.1% compared to representative baselines. Yaqiong Peng, Haocheng Peng |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | From Satellite to Ground: Satellite Assisted Visual Localization with Cross-view Semantic MatchingabstractOne of the key challenges of visual Simultaneous Localization and Mapping (SLAM) in large-scale environments is how to effectively use global localization to correct the cumulative errors from long-term tracking. This challenge presents itself in two main aspects: first, the difficulty for robots in revisiting previous locations to perform loop closure, and second, the considerable memory resources required to maintain point-cloud-based global maps. Recent solutions have resorted into neural networks, using satellite images as the references for ground-level localization. However, most of these methods merely provide cross-view patch-matching results, which leads to unfeasible in integration with the SLAM system. To address these issues, we present a semantic-based cross-view localization method. This approach combines semantic information with a reward and penalty mechanism, enabling us to obtain a global probability map and achieve precise 3-degree-of-freedom (3-DoF) localization. Based on that, we develop a SLAM system that capitalizes on satellite imagery for global localization. This strategy effectively bridges the gap between SLAM and real-world coordinates while also substantially reducing accumulated errors. Our experimental results demonstrate that our global localization method significantly outperforms existing satellite-based systems. Moreover, in scenarios where the robot struggles to find loop closures, employing our localization method improves the SLAM accuracy. Xiyue Guo, Haocheng Peng, Junjie Hu 0003, Hujun Bao, Guofeng Zhang 0001 |
ICRA | 2 |
| 2024 | PC-Planner: Physics-Constrained Self-Supervised Learning for Robust Neural Motion Planning with Shape-Aware Distance FunctionabstractMotion Planning (MP) is a critical challenge in robotics, especially pertinent with the burgeoning interest in embodied artificial intelligence. Traditional MP methods often struggle with high-dimensional complexities. Recently neural motion planners, particularly physics-informed neural planners based on the Eikonal equation, have been proposed to overcome the curse of dimensionality. However, these methods perform poorly in complex scenarios with shaped robots due to multiple solutions inherent in the Eikonal equation. To address these issues, this paper presents PC-Planner, a novel physics-constrained self-supervised learning framework for robot motion planning with various shapes in complex environments. To this end, we propose several physical constraints, including monotonic and optimal constraints, to stabilize the training process of the neural network with the Eikonal equation. Additionally, we introduce a novel shape-aware distance field that considers the robot's shape for efficient collision checking and Ground Truth (GT) speed computation. This field reduces the computational intensity, and facilitates adaptive motion planning at test time. Experiments in diverse scenarios with different robots demonstrate the superiority of the proposed method in efficiency and robustness for robot motion planning, particularly in complex environments. Xujie Shen, Haocheng Peng, Zesong Yang, Juzhan Xu, Hujun Bao, Ruizhen Hu, Zhaopeng Cui |
SIGGRAPH Asia | 2 |
| 2024 | InferFair: Towards QoS-aware scheduling for performance isolation guarantee in heterogeneous model serving systems
Yaqiong Peng, Haocheng Peng |
Future Gener. Comput. Syst. | 2 |
| 2024 | Serving DNN Inference With Fine-Grained Spatio-Temporal Sharing of GPU ServersabstractDeep Neural Networks(DNNs) are commonly deployed as online inference services. To meet interactive latency requirements of requests, DNN services require the use ofGraphics Processing Unit(GPU) to improve their responsiveness. The unique characteristics of inference workloads pose new challenges to manage GPU resources. First, the GPU scheduler needs to carefully manage requests to meet their latency targets. Second, a single inference task often underutilizes GPU resources. Third, the fluctuating patterns of inference workloads pose difficulties in determining the resources allocated to each DNN model. Therefore, it is critical for the GPU scheduler to maximize GPU utilization by collocating multiple DNN models without violating the latencyService-Level Objectives(SLOs) of requests. However, we find that existing works are not adequate for achieving this goal among latency-sensitive inference tasks. Hence, we propose FineST, a scheduling framework for serving DNNs with fine-grained spatio-temporal sharing of GPU inference servers. To maximize GPU utilization, FineST allocates intra-GPU computing resources from both spatial and temporal dimensions across DNNs in a cost-effective way, while predicting interference overheads under diverse consolidated executions for controlling SLO violation rates. Compared to a state-of-the-art work, FineST improves the peak throughput of serving heterogeneous DNNsby up to 64.7% under SLO constraints. Yaqiong Peng, Weiguo Gao, Haocheng Peng |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | DTCC: Multi-level dilated convolution with transformer for weakly-supervised crowd countingabstractCrowd counting provides an important foundation for public security and urban management. Due to the existence of small targets and large density variations in crowd images, crowd counting is a challenging task. Mainstream methods usually apply convolution neural networks (CNNs) to regress a density map, which requires annotations of individual persons and counts. Weakly-supervised methods can avoid detailed labeling and only require counts as annotations of images, but existing methods fail to achieve satisfactory performance because a global perspective field and multi-level information are usually ignored. We propose a weakly-supervised method, DTCC, which effectively combines multi-level dilated convolution and transformer methods to realize end-to-end crowd counting. Its main components include a recursive swin transformer and a multi-level dilated convolution regression head. The recursive swin transformer combines a pyramid visual transformer with a fine-tuned recursive pyramid structure to capture deep multi-level crowd features, including global features. The multi-level dilated convolution regression head includes multi-level dilated convolution and a linear regression head for the feature extraction module. This module can capture both low- and high-level features simultaneously to enhance the receptive field. In addition, two regression head fusion mechanisms realize dynamic and mean fusion counting. Experiments on four well-known benchmark crowd counting datasets (UCF_CC_50, ShanghaiTech, UCF_QNRF, and JHU-Crowd++) show that DTCC achieves results superior to other weakly-supervised methods and comparable to fully-supervised methods. Zhuangzhuang Miao, Yong Zhang 0029, Haocheng Peng |
Comput. Vis. Media | 4 |