EDBT 2026 Demo / reviewers in the wild / expert
Haoran Kong
dblp:306/9625
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing and Mitigating I/O Bottlenecks in LLM Inference on Disaggregated HPC Systems
Fanzi Zeng, Kenli Li, Haoran Kong, Longbao Dai |
APPT | 6 |
| 2026 | COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM TrainingabstractCollective communication is critical to scaling large language model (LLM) training across various parallelism strategies, including data, tensor, and pipeline parallelism on GPU clusters. However, as model sizes and training scales increase, communication overhead is emerging as a major performance bottleneck. While compression is a promising mitigation strategy, existing solutions often lack user-transparency, hinder deployment and extensibility, and are not co-designed with communication algorithms. To address these limitations, we present COCCL, a high-performance collective communication library built on top of NCCL. COCCL introduces a novel programming model that can easily integrate compression into communication workflows with flexible configurability. It features a suite of compression-aware collective algorithms and runtime overlap mechanisms that mitigate error propagation and reduce computational overhead. We integrate well-established compression techniques into COCCL and tune the compression configurations during 3D-parallel training on GPT and Qwen models with up to 7 billion parameters. Using the optimal configuration (COCCL-3D), we achieve 1.24× throughput improvement while maintaining training accuracy. Haoran Kong, Hairui Zhao 0002, Shengkai Lyu, Xingjian Tian, Liyang Zhao, Zhuohan Chen, Fakang Wang, Zizhong Chen, Zhan Wang 0003, Guangming Tan, Dingwen Tao |
PPoPP | 2 |
| 2025 | Accelerating the Martian Atmospheric Simulation of GoMars Model with Multi-GPUsabstractMars exploration is at the forefront of space science, which demands robust computational models to decipher its atmospheric dynamics. In this work, we present a significant advancement in computational efficiency for the GoPlanetMars (GoMars), a state-of-the-art Martian atmospheric model. By leveraging the parallel processing capabilities of Graphics Processing Units (GPUs), we accelerate the dynamic core of the GoMars model on multiple NVIDIA A800 GPUs. Through comprehensive performance analysis of GoMars, we optimize both parallel computation and communication patterns to leverage the computational power of multiple GPUs fully, achieving performance comparable to that of a thousand-core CPU cluster. Our evaluation results demonstrate that the GPU-accelerated GoMars model maintains the same level of precision as the native CPU-based implementation, while achieving a substantial speedup, making it a viable solution for high-performance Martian atmospheric simulations. Guofan Yu, Haoran Kong, Xin You 0001, Hailong Yang 0002, Shaokang Du, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
HPCC | 2 |
| 2025 | Cherry: Breaking the GPU Memory Wall for Large-Scale GNN Training via Micro-BatchingabstractGraph Neural Networks (GNNs) have shown remarkable performance across a variety of graph-related tasks.Recent efforts indicate that GNN performance can be enhanced through more sophisticated strategies, such as employing advanced aggregators, increasing aggregation depth, and utilizing larger sampling rates, etc.While these strategies yield promising results, it also incurs a significantly larger memory footprint that can easily surpass the GPU memory capacity.Micro-batching has emerged as a promising method to mitigate GPU memory bottleneck while preserving model accuracy.Nevertheless, integrating micro-batches into GNN Yan Wang 0022, Haoran Kong, Hao Chen 0002, Weile Jia, Dingwen Tao, Xin He 0054 |
ICS | 3 |
| 2025 | SA-MVSNet: Spatial-aware Multi-view Stereo Network with Attention Cost VolumeabstractDeep learning-based multi-view stereo (MVS) methods enable dense point cloud reconstruction in texture-rich areas. However, existing methods incur significant computational costs to capture pixel dependencies for complete reconstruction in low-texture regions. Additionally, discrete depth layers in occluded environments hinder the cost volume’s ability to model object information effectively. To address these issues, we propose a spatial-aware multi-view stereo network with attention cost volume, termed SA-MVSNet. The network introduces the pixel-driven spatial interaction (PDSI) module, which integrates the hierarchical spatial location enhancement mechanism (HSLE) and the spatial context aggregation mechanism (SCA). Leveraging an efficient parallel architecture, the PDSI module captures pixel-level spatial dependencies with the HSLE and strengthens global contextual information through the SCA. This design improves the network’s ability to represent features in low-texture regions while maintaining high inference efficiency. Furthermore, SA-MVSNet incorporates an attention weight generation branch that refines the cost volume by aggregating multi-scale depth cues, effectively mitigating the impact of occlusion. Experiments on the DTU dataset and the Tanks and Temples dataset show that our method outperforms other learning-based methods, achieving superior performance and strong generalization ability. Haoran Kong, Fanzi Zeng, Longbao Dai, Jingyang Hu, Jiang-hao Cai, Jianxia Chen, Ruihui Li, Hongbo Jiang 0001 |
IROS | 1 |
| 2025 | MDA: a multi-agent framework for data analysis task
Haoran Kong, Shen Gao |
Neural Comput. Appl. | 2 |
| 2025 | Throughput-Aware Cooperative Task Offloading in Dynamic Mobile Edge Computing SystemsabstractWith the commercialization of fifth-generation (5G) mobile communication technology and the rapid proliferation of mobile devices (MDs), demand for data computation is surging. This growth increases the reliance of MDs on low latency and high throughput. For this purpose, Mobile Edge Computing (MEC) enhances the user's data processing capability by offloading computation tasks to servers at the network edge. However, achieving high efficiency in task offloading is challenging due to factors such as decision complexity, network dynamics, and user data privacy protection. Additionally, energy causal constraints and the coupling between offloading proportions and resource distribution cannot be ignored. In this paper, we first establish a dynamic task offloading problem to optimize the long-term throughput of the system. Using perturbed Lyapunov optimization, we transform MD delay and energy threshold constraints into the stability control of corresponding virtual queues. Then, we propose the Lyapunov-guided federated deep reinforcement learning (DRL) online task offloading algorithm called LyFOTO, which combines a federated learning (FL) framework and an Actor-Critic (AC) model. Under favorable communication conditions, the LyFOTO algorithm adaptively boosts system throughput; under poorer conditions, it properly delays task offloading, without violating queue backlog constraints. Through mathematical analysis, we discuss the performance of the LyFOTO algorithm. Simulation experiments validate that LyFOTO effectively balances system throughput and device battery energy. Finally, Comparative results show that LyFOTO outperforms other benchmark algorithms in maximizing system throughput while ensuring task backlog and energy threshold constraints. Longbao Dai, Fanzi Zeng, Haoran Kong, Jiang-hao Cai, Hongbo Jiang 0001, Keqin Li 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Accelerating Big Data Application by Eliminating Redundancy on Hadoop ClusterabstractBig data applications are widely adopted to mine valuable information from a tremendous amount of industry data, which is commonly represented as a series of map-reduce operations. Among various map-reduce frameworks, Hadoop is most commonly adopted for data processing at large scale. Although Hadoop eases the development of highly scalable distributed big data applications, inefficient implementation due to poor coding practice and deep software abstractions can cause severe performance issues such as unreasonable slowdown, high response latency, and waste of computing resources, which can lead to unsatisfactory serving delay or significant maintenance cost. In this paper, we first categorize three common types of redundant patterns in big data applications. Then we propose a tool-assisted optimization workflow to detect the redundant patterns automatically, which profiles the application by sampling hardware performance monitoring units. Moreover, we present a profiling visualization method that can help to pinpoint the redundant codes. Based on these approaches, we optimize several big data applications by eliminating redundancies, yielding up to 14.8% performance improvement. Kelun Lei, Shaokang Du, Xin You 0001, Zhibo Xuan, Haoran Kong, Hailong Yang 0002, Jing Shang 0001, Zhiwen Xiao, Zhongzhi Luan, Depei Qian 0001 |
ICPADS | 5 |