EDBT 2026 Demo / reviewers in the wild / expert
Haoxu Li
dblp:229/8120
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum StatesabstractAI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. Among neural network quantum state (NNQS) approaches, the NNQS-SCI (Selected Configuration Interaction) method stands out as a state-of-the-art technique, recognized for its high accuracy and scalability. However, its application to larger systems is severely constrained by a hybrid CPU-GPU architecture. Specifically, centralized CPU-based global de-duplication creates a severe scalability barrier due to communication bottlenecks, while host-resident coupled-configuration generation induces prohibitive computational overheads. We introduce QiankunNet-cuSCI, a fully GPU-accelerated SCI framework designed to overcome these bottlenecks. It first integrates a distributed, load-balanced global de-duplication algorithm to minimize redundancy and communication overhead at scale. To address compute limitations, it employs specialized, fine-grained CUDA kernels for exact coupled configuration generation. Finally, to break the single-GPU memory barrier exposed by this full acceleration, it incorporates a GPU memory-centric runtime featuring GPU-side pooling, streaming mini-batches, and overlapped offloading. This design enables much larger configuration spaces and shifts the bottleneck from host-side limitations back to on-device inference. Our evaluation demonstrates that our work fundamentally expands the scale of solvable problems. On an NVIDIA A100 cluster with 64 GPUs, our work achieves up to 2.32 × end-to-end speedup over the highly-optimized NNQS-SCI baseline while preserving the same chemical accuracy. Furthermore, it demonstrates excellent distributed performance, maintaining over 90% parallel efficiency in strong scaling tests. Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao 0002, Haoxu Li, Ankang Feng, Wenjing Huang 0002, Yida Gu, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan |
HPDC | 5 |
| 2026 | CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingabstractAs training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes—substantially outperforming existing solutions. Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang 0001, Hairui Zhao 0002, Wenjing Huang 0002, Jinwu Yang, Yueyuan Zhou, Qian Zhao 0021, Haoxu Li, Zhan Wang 0003, Guangming Tan, Dingwen Tao |
PPoPP | 15 |
| 2025 | MANS: Efficient and Portable ANS Encoding for Multi-Byte Integer Data on CPUs and GPUsabstractLossless compression is a classic technique for reducing data storage and transmission requirements. Asymmetric Numeral Systems (ANS) is a high-throughput, high-ratio lossless compression algorithm, but it lacks effective support for multi-byte data and cross-platform compatibility. To address this issue, we propose an Adaptive Data Mapping (ADM) scheme, which maps multi-byte integer data into single-byte space based on the data’s characteristics, improving the compression ratio of ANS while maintaining low encoding redundancy. We also optimize the ADM algorithm and the ANS encoder for GPU and CPU architectures, respectively, and combine them to create an efficient and portable ANS encoding method for multi-byte integer data, called MANS. Experimental results show that MANS improves compression ratios by an average of 1.24 ×, achieves 870.27MB/s throughput on CPUs, and delivers up to 288.45 × and 135.86 × speedups on an NVIDIA A100 and an AMD MI210 GPU compared to the CPU version—demonstrating its efficiency and portability across platforms. Wenjing Huang 0002, Jinwu Yang, Shengquan Yin, Haoxu Li, Yida Gu, Xing Jing, Shiyuan Fu, Hao Hu 0015, Guangming Tan, Dingwen Tao |
SC | 4 |
| 2018 | Indoor visible light positioning combined with ellipse-based ACO-OFDMabstractVisible light communications have recently gained increasing attention as it is appealing for a wide range of applications such as indoor positioning. However, the light‐emitting diode‐based indoor positioning systems suffer from multipath distortion inside a room, leading to high positioning error. In order to mitigate the effect of multipath distortion of the optical channel, an indoor visible light positioning algorithm combined with ellipse‐based asymmetrically clipped optical orthogonal frequency division multiplexing (ACO‐OFDM) is proposed in this study, wherein the real‐valued output of orthogonal frequency division multiplexing is modulated onto an ellipse, and only the imaginary value from the complex point on the ellipse is transmitted. Moreover, a received‐signal‐strength technique is used to determine the location of the receiver, and the Trust‐region technique is employed to realise the 3D positioning. Simulation results demonstrate that this proposed algorithm can achieve high positioning accuracy, and the impact of different system parameters on the positioning accuracy is further investigated compared with the systems based on ACO‐OFDM. Haoxu Li, Rangzhong Wu |
IET Commun. | 1 |