EDBT 2026 Demo / reviewers in the wild / expert
Jintao Peng
dblp:245/0884
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007 |
Sci. China Inf. Sci. | 1 |
| 2024 | The Self-adaptive and Topology-aware MPI_Bcast leveraging Collective offload on Tianhe Express InterconnectabstractLarge parallel applications have heavily used MPI (Massage Passing Interface) collectives that support portable and efficient group communication operations. MPI_Bcast is one of the most commonly used MPI collectives that broadcast data to all processes of the communication domain. However, traditional software-based broadcast algorithms fail to fully utilize modern interconnection networks’ advanced features such as offloading collectives to the network hardware for efficient group communications. Besides, the semantic gap between MPI_Bcast and hardware multicast of underlying interconnects presents challenges for offload-based algorithms to accelerate MPI_Bcast for a wide range of message sizes.In this paper, we propose a hardware-software co-design MPI_Bcast by efficiently leveraging the NIC-based collective offload provided by Tianhe-express interconnect, which completely precludes the involvement of CPU to accelerate message broadcast. We detail this broadcast mechanism that can be adaptively tuned to offload MPI_Bcast operations from the CPU to the NIC for various message and system sizes. In addition, we further propose a topology-aware broadcast design in conjunction with this offload method to significantly reduce the broadcast latency by constructing the optimal global inter-node communication tree. We implement and evaluate the proposed Tianhe-Express Offload-based Broadcast (TOB) design on Tianhe-2A and Tianhe-EP supercomputers. Extensive experiments have been conducted to evaluate TOB performance at both microbenchmark and application levels. Our solution offers up to 4.94x significant performance speedup at the microbenchmark level over state-of-the-art MPI libraries. For the application-level evaluation, our technique accelerates scientific applications by a maximum speedup of 1.34x. Chongshan Liang, Jinbo Xu, Jintao Peng, Weixia Xu 0001, Jie Liu 0002, Zhiquan Lai, Sheng Ma |
IPDPS | 5 |
| 2023 | TH-Allreduce: Optimizing Small Data Allreduce Operation on Tianhe SystemabstractScaling up parallel applications can be challenging, especially when dealing with large volumes of data that need to be distributed across multiple nodes. In this paper, we explore the system architecture of Tianhe and propose a cutting-edge solution for optimizing global data communication. Our optimized small data blocking/non-blocking allreduce method (TH-Allreduce) is tailored for scientific applications like solving linear systems Ax=b, which often require massive data processing capabilities. To address intra-node communication challenges, we introduce the Ping-Pong Small Data Shared Memory (PPSDSM) framework, which utilizes ping-pong communication patterns to minimize Round-Trip Time (RTT) and reduce computational costs. We further present a latency-aware allreduce algorithm (PP-LA) based on PPSDSM that optimizes communication overheads and computational costs. For inter-node communication, we leverage the power of the Tianhe offloading engine to propose a topology-aware offloading allreduce method. Our experimental results show that our state-of-the-art library outperforms typical MPI implementations on different CPUs, achieving a speedup of 1.5-12x for intra-node allreduce and 1.32-3.34x for multi-node small data allreduce on the Tianhe Exascale Prototype Upgrade System at scale. These findings demonstrate that our proposed methods can significantly improve communication efficiency and scalability for distributed computing on Tianhe systems, opening up new avenues for a wide range of scientific applications. Jintao Peng, Lihua Chi |
ICPADS | 2 |
| 2023 | GLEX_Allreduce: Optimization for medium and small message of Allreduce on Tianhe systemabstractGlobal communication may affect the scalability of some parallel applications. Message Passing Interface (MPI) provides some commonly used collective communication Application Programming Interface(API). Allreduce is one of the APIs that is mostly used on parallel applications. Small message Allreduce is useful for dot products and solving linear systems. This paper proposes a medium and small message Allreduce for the Tianhe series system. For intra-node reduction/broadcast, this paper proposed a cache-aware tree and shared memory implementation. Besides, this paper proposes broadcast merge and cache line awareness methods to improve performance further. For inter-node communication, this paper proposes zero event Remote Direct Memory Access(RDMA) to avoid event overhead on the Tianhe system. In addition, a zero event immediate-data RDMA(Imm-RDMA) is used to optimize small message RDMA. Through experiments, for 16384 MPI processes, GLEX Allreduce achieves 2.4-5.1 times speedup compared to MPI. Compared to other collective communication libraries on Infiniband(IB) and Omni-Path, GLEX_Allreduce achieves similar or better performance. Jintao Peng, Jie Liu 0002, Liuhua Chi |
ICPADS | 2 |
| 2023 | Optimizing MPI Collectives on Shared Memory Multi-CoresabstractMessage Passing Interface (MPI) programs often experience performance slowdowns due to collective communication operations, like broadcasting and reductions. As modern CPUs integrate more processor cores, running multiple MPI processes on shared-memory machines to take advantage of hardware parallelism is becoming increasingly common. In this context, it is crucial to optimize MPI collective communications for shared-memory execution. However, existing MPI collective implementations on shared-memory systems have two primary drawbacks. The first is extensive redundant data movements when performing reduction collectives, and the second is the ineffective use of non-temporal instructions to optimize streamed data processing. To address these limitations, this paper proposes two optimization techniques that minimize data movements and enhance the use of non-temporal instructions. We evaluated our techniques by integrating them into the OpenMPI library and tested their performance using micro-benchmarks and real-world applications running on two multi-core clusters. Experimental results show that our approach significantly outperforms existing techniques, yielding a 1.2--6.4x performance improvement. Jintao Peng, Jianbin Fang, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Zheng Wang 0001 |
SC | 1 |
| 2022 | Reconstructing gene regulatory networks of biological function using differential equations of multilayer perceptronsabstractBACKGROUND: Building biological networks with a certain function is a challenge in systems biology. For the functionality of small (less than ten nodes) biological networks, most methods are implemented by exhausting all possible network topological spaces. This exhaustive approach is difficult to scale to large-scale biological networks. And regulatory relationships are complex and often nonlinear or non-monotonic, which makes inference using linear models challenging. RESULTS: In this paper, we propose a multi-layer perceptron-based differential equation method, which operates by training a fully connected neural network (NN) to simulate the transcription rate of genes in traditional differential equations. We verify whether the regulatory network constructed by the NN method can continue to achieve the expected biological function by verifying the degree of overlap between the regulatory network discovered by NN and the regulatory network constructed by the Hill function. And we validate our approach by adapting to noise signals, regulator knockout, and constructing large-scale gene regulatory networks using link-knockout techniques. We apply a real dataset (the mesoderm inducer Xenopus Brachyury expression) to construct the core topology of the gene regulatory network and find that Xbra is only strongly expressed at moderate levels of activin signaling. CONCLUSION: We have demonstrated from the results that this method has the ability to identify the underlying network topology and functional mechanisms, and can also be applied to larger and more complex gene network topologies. Guo Mao, Ruigeng Zeng, Jintao Peng, Ke Zuo, Zhengbin Pang, Jie Liu 0002 |
BMC Bioinform. | 3 |
| 2019 | Improving Performance of Batch Point-to-Point Communications by Active Contention Reduction Through Congestion-Avoiding Message Scheduling
Jintao Peng, Qingkai Liu |
ICA3PP (1) | 1 |