Peiheng Zhang

dblp:16/2545 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 37% Hardware accelerators and domain-specific architectures · 22% High-performance computing · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
bioinformatics accelerator
0.212016
Accelerating Irregular Computation in Massive Short Reads Mapping on FPGA Co-Processor · IEEE Trans. Parallel Distributed Syst. 2016
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA coprocessor
0.212016
Accelerating Irregular Computation in Massive Short Reads Mapping on FPGA Co-Processor · IEEE Trans. Parallel Distributed Syst. 2016
Reconfigurable computing and FPGAs › FPGA accelerator
short read mapping
0.212016
Accelerating Irregular Computation in Massive Short Reads Mapping on FPGA Co-Processor · IEEE Trans. Parallel Distributed Syst. 2016
Reconfigurable computing and FPGAs
FPGA accelerator
0.112012
A coarse-grained stream architecture for cryo-electron microscopy images 3D reconstruction · FPGA 2012
Hardware accelerators and domain-specific architectures
scientific computing accelerator
0.112012
A coarse-grained stream architecture for cryo-electron microscopy images 3D reconstruction · FPGA 2012
Processor architecture and microarchitecture › dataflow architecture
stream architecture
0.112012
A coarse-grained stream architecture for cryo-electron microscopy images 3D reconstruction · FPGA 2012
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing
0.112011
Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system · HPDC 2011
High-performance computing › scientific computing systems
cryo-EM 3D reconstruction
0.112011
Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system · HPDC 2011
Parallel and multicore computing › parallel programming models › structured parallelism
hierarchical parallelism
0.112011
Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system · HPDC 2011
High-performance computing
scientific computing systems
0.112011
Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system · HPDC 2011
Bioinformatics and computational biology
next-generation sequencing
0.112016
Accelerating Irregular Computation in Massive Short Reads Mapping on FPGA Co-Processor · IEEE Trans. Parallel Distributed Syst. 2016
Bioinformatics and computational biology › structural biology › cryo-electron microscopy
3d reconstruction
0.012012
A coarse-grained stream architecture for cryo-electron microscopy images 3D reconstruction · FPGA 2012
Bioinformatics and computational biology › structural biology
cryo-electron microscopy
0.012012
A coarse-grained stream architecture for cryo-electron microscopy images 3D reconstruction · FPGA 2012
Parallel and multicore computing › task scheduling
dynamic scheduling
0.012011
Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system · HPDC 2011
Parallel and multicore computing
parallel programming models
0.012011
Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system · HPDC 2011

Methods — techniques the papers use, named apart from their topics

word-level parallelism · 0.5scatter-gather memory · 0.5bit-level parallelism · 0.5stream architecture · 0.3pipelining · 0.3self-adaptive dynamic scheduling · 0.1OpenMP · 0.1CUDA · 0.1
YearPublicationVenuePosition
2024 ProXplore: A GPU-Enhanced Protein Discovery Engine
abstract
Language models have achieved unprecedented success in natural language processing tasks and have recently been adapted for biological sequences. However, GPUs still encounter significant performance bottlenecks when running BERT-style natural language models. Moreover, due to the vastly greater number of tokens in protein sequences compared to human languages (as illustrated in Figure 1), directly transferring language models from human languages to proteins results in significant changes in runtime behavior. Consequently, optimizations designed for short-input BERT models are less effective for protein language models.In this paper, we propose a novel GPU-based cross-layer optimization strategy. From an architectural perspective, our approach leverages intra- and inter-operator parallelism, multilevel data computation and warp parallelism to fully utilize GPU features. From an algorithmic perspective, we address the issue of varying lengths in protein input sequences and reduce memory overhead through padding removal. Experimental results demonstrate that ProXplore achieves a 1.8× increase in inference speed on the TAPE benchmark without a significant loss in prediction accuracy. This improvement effectively overcomes performance bottlenecks and provides substantial benefits for protein sequence analysis and other bioinformatics applications.
Zewen Sun, Peiheng Zhang, Xingzhuo Fu, Xueqi Li 0001
BIBM2
2024 CoPIM: A Collaborative Scheduling Framework for Commodity Processing-in-memory Systems
abstract
Processing in memory (PIM) is a promising paradigm to effectively alleviate the bottleneck of memory access for data-intensive applications. UPMEM is the first publicly-available real-world processing-in-memory (PIM) platform. However, current approaches mostly treat the CPU as the controller rather than considering the CPU-PIM system as a whole, thus failing to fully utilize the computational resources in modern CPUs. Additionally, the use of commodity PIM hardware as accelerators with independent memory addresses leads to significant data communication between the CPU and PIM, diminishing the performance benefits of reducing data movement. To address these challenges, we propose CoPIM, a scheduling framework that automates efficient workload scheduling. CoPIM introduces a fine-grained load modeling approach based on computational graphs and establishes a performance model to estimate workload performance and energy consumption. Our framework includes a scheduler that generates optimal scheduling parameters, enabling CPU-PIM collaborative processing. CoPIM enhances system parallelism and reduces data transfer between the CPU and PIM. We implement CoPIM on the UPMEM PIM system and demonstrate through experiments that our design achieves 4.0 x performance improvement on average when compared to general-purpose processors while reducing energy consumption by 56.6%.
Shunchen Shi, Xueqi Li 0001, Zhaowu Pan, Peiheng Zhang, Ninghui Sun
ICCD4
2024 Efficient Noninteractive Polynomial Commitment Scheme in the Discrete Logarithm Setting
abstract
Polynomial commitment schemes (PCSs) are fundamental components that can effectively solve the problems arising from the combination of Internet of Things and blockchain. These allow a committer to commit to a polynomial and then later evaluate the committed polynomial at an arbitrary challenge point along with a proof of valid, without revealing any additional information about the polynomial. Recent works have presented polynomial commitment schemes based on the discrete logarithm assumption. Their schemes do not require a trusted setup, and the verifier uses homomorphism to check the polynomial evaluation proofs. However, these schemes require two-party interactions and satisfy only special soundness and special honest verifier zero-knowledge, which are infeasible for some nonsimultaneous online or decentralized applications. In this article, we propose a novel PCS inspired by the idea of the Fiat–Shamir heuristic. Our scheme is noninteractive between the committer and the verifier. Instead of waiting for the challenge values from the verifier, the committer generates the values by accessing a random oracle. Moreover, it satisfies computational soundness and zero-knowledge by using a group operation to enhance the unpredictability of challenge values. We also propose a trapdoor commitment scheme to ensure the honest use of challenge values by the committers. Finally, we present the security and performance analysis of our scheme, which shows that our scheme is feasible with an acceptable time overhead.
Peiheng Zhang, Willy Susilo, Mingwu Zhang
IEEE Internet Things J.1
2023 TransCaller: An End-to-end Accelerated Transformer-based Nanopore Basecaller on GPUs
abstract
Basecalling is a crucial step in nanopore sequencing as it transforms the raw electrical signal obtained from the nanopore into a readable sequence. The accuracy of basecalling directly affects the quality and reliability of the sequenced data. Various algorithms and models are employed to perform accurate basecalling, with advancements in deep learning techniques. However, due to the difficulty of combining feature extraction capability, high parallelism and long sequence modeling capability of existing models, the accuracy and speed of basecalling model are still bottlenecks in data analysis.In this paper, we present TransCaller, an end-to-end accelerated transformer-based nanopore basecaller model. TransCaller comprises a low-parameter electrical signal sampler, a fully transformer-based encoder integrating multiple algorithm-specific modules, and a CTC decoder enhanced with a filter. To further refine and enhance the TransCaller model, we introduce a three-stage pyramid model structure and leverage knowledge distillation techniques to balance the accuracy and speed. Experimental results on 9 datasets demonstrate that TransCaller achieves an average accuracy of 92.95%, surpassing the state-of-the-art Bonito sup model by a margin of 0.65% to 2.34%. Moreover, the distilled TransCallermodel showcases a remarkable 2.97× acceleration in inference speed compared to the original model, effectively catering to the variegated requirements of distinct inference speeds and precision levels across a wide spectrum of scenarios.
Peiheng Zhang, Shunchen Shi, Xueqi Li 0001
BIBM2
2018 STrieGD: A Sampling Trie Indexed Compression Algorithm for Large-Scale Gene Data
Yanzhen Gao, Xiaozhen Bao, Peiheng Zhang
NPC6
2016 Accelerating Irregular Computation in Massive Short Reads Mapping on FPGA Co-Processor
abstract
Because there is an enormous amount of genomic data, next-generation sequencing (NGS) applications pose significant challenges to current computing systems. In this study, we investigate both algorithmic and architectural strategies to accelerate an NGS data analysis algorithm—short read mapping on commodity multi-core platform and customizable field programmable gate array (FPGA) co-processor architecture, respectively. A workload analysis reveals that conventional memory optimization is limited in its irregular computation of low arithmetic intensity and non-contiguous memory access pattern. To mitigate the inherent irregular computation in mapping, we have developed a FPGA co-processor based on Convey computer, which employs a scatter-gather memory mechanism that exploits both bit-level and word-level parallelism. The customized FPGA co-processor achieves a throughput of$947$Gbp per day, about$189$times higher than that of current mapping tools on single CPU core. Moreover, the co-processor's power efficiency is$29$times higher than that of a conventional 64-core multi-processor.
Guangming Tan, Peiheng Zhang, Ninghui Sun
IEEE Trans. Parallel Distributed Syst.4
2015 SuperDragon: A Heterogeneous Parallel System for Accelerating 3D Reconstruction of Cryo-Electron Microscopy Images
abstract
The data deluge in medical imaging processing requires faster and more efficient systems. Due to the advance in recent heterogeneous architecture, there has been a resurgence in research aimed at domain-specific accelerators. In this article, we develop an experimental system SuperDragon for evaluating acceleration of a single-particle Cryo-electron microscopy (Cryo-EM) 3D reconstruction package EMAN through a hybrid of CPU, GPU, and FPGA parallel architecture. Based on a comprehensive workload characterization, we exploit multigrained parallelism in the Cryo-EM 3D reconstruction algorithm and investigate a proper computational mapping to the underlying heterogeneous architecture. The package is restructured with task-level (MPI), thread-level (OpenMP), and data-level (GPU and FPGA) parallelism. Especially, the proposed FPGA accelerator is a stream architecture that emphasizes the importance of optimizing computing dominated data access patterns. Besides, the configurable computing streams are constructed by arranging the hardware modules and bypassing channels to form a linear deep pipeline. Compared to the multicore (six-core) program, the GPU and FPGA implementations achieve speedups of 8.4 and 2.25 times in execution time while improving power efficiency by factors of 7.2 and 14.2, respectively.
Guangming Tan, Wendi Wang 0002, Peiheng Zhang
ACM Trans. Reconfigurable Technol. Syst.4
2013 Write bandwidth optimization of online Erasure Code based cluster file system
abstract
As the data volume is growing from big to huge in many science labs and data centers, more and more data owners are willing to choose Erasure Code based storage to reduce the storage cost. However, online Erasure Code based cluster file systems still have not been applied widely because of write bottlenecks in data encoding and data placement. We proposed two optimizations to address them respectively. We propose a Partition Encoding policy to accelerate the encoding arithmetic through SIMD extensions and to overlap data encoding with data committing. We devise Adaptive Placement policy to provide incremental expansion and high availability, as well as good scalability. The experimental results in our prototype ECFS show that the aggregate write bandwidth can be improved by 42%, while keeping the storage in a more balanced state.
Zhigang Huo, Peiheng Zhang
CLUSTER6
2012 Accelerating Millions of Short Reads Mapping on a Heterogeneous Architecture with FPGA Accelerator
abstract
The explosion of Next Generation Sequencing (NGS) data with over one billion reads per day poses a great challenge to the capability of current computing systems. In this paper, we proposed a CPU-FPGA heterogeneous architecture for accelerating a short reads mapping algorithm, which was built upon the concept of hash-index. In particular, by extracting and mapping the most time-consuming and basic operations to specialized processing elements (PEs), our new algorithm is favorable to efficient acceleration on FPGAs. The proposed architecture is implemented and evaluated on a customized FPGA accelerator card with a Xilinx Virtex5 LX330 FPGA resided. Limited by available data transfer bandwidth, our NGS mapping accelerator, which operates at 175MHz, integrates up to 100 PEs. Compared to an Intel six-cores CPU, the speedup of our accelerator ranges from 22.2 times to 42.9 times.
Wendi Wang 0002, Bo Duan, Guangming Tan, Peiheng Zhang, Ninghui Sun
FCCM6
2012 A coarse-grained stream architecture for cryo-electron microscopy images 3D reconstruction
abstract
The wide acceptance of bioinformatics, medical imaging and multimedia applications, which have a data-centric favor to them, require more efficient and application-specific systems to be built. Due to the advances in modern FPGA technologies recently, there has been a resurgence in research aimed at accelerator design that leverages FPGAs to accelerate large-scale scientific applications. In this paper, we exploit this trend towards FPGA-based accelerator design and provide a proof-of-concept and comprehensive case study on FPGA-based accelerator design for a single-particle 3D reconstruction application in single-precision floating-point format. The proposed stream architecture is built by first offloading computing-intensive software kernels to dedicated hardware modules, which emphasizes the importance of optimizing computing dominated data access patterns. Then configurable computing streams are constructed by arranging the hardware modules and bypass channels to form a linear deep pipeline. The efficiency of the proposed stream architecture is justified by the reported 2.54 times speedup over a 4-cores CPU. In terms of power efficiency, our FPGA-based accelerator introduces a 7.33 and 3.4 times improvement over a 4-cores CPU and an up-to-date GPU device, respectively.
Wendi Wang 0002, Bo Duan, Guangming Tan, Peiheng Zhang, Ninghui Sun
FPGA6
2011 Floating-point mixed-radix FFT core generation for FPGA and comparison with GPU and CPU
abstract
Over the past decades, we noticed huge advances in FPGA technologies. The topic of floating-point accelerator on FPGA has gained renewed interests due to the increased device size and the emergence of fast hardware floating-point library. The popularity of FFT makes it easier to justify spending lots of effort doing detailed optimization. However, the ever increasing data size in some compelling application domains remains beyond the capability of existing FFT accelerators. The demand for more performance remains an active research topic. In this paper, leveraging structured description of FFT algorithms, we propose a FPGA-based FFT core generation framework, which emits Verilog HDL code given high-level algorithmic description and can handle radix-2 as well as prime-radix problem size. In particular, the proposed framework is optimized for 2D FFT and real FFT. The performance of our implementation is comparable with a commercial FFT IP. When compared with the latest results on GPU and CPU, measured in peak floating-point performance and energy efficiency, it shows that GPUs have outperformed FPGAs for FFT acceleration. However, we consider that FPGAs still have advantage in some situations.
Bo Duan, Wendi Wang 0002, Xingjian Li 0002, Peiheng Zhang, Ninghui Sun
FPT5
2011 Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous system
abstract
Heterogeneous architecture is becoming an important way to build a massive parallel computer system, i.e. the CPU-GPU heterogeneous systems ranked in Top500 list. However, it is a challenge to efficiently utilize massive parallelism of both applications and architectures on such heterogeneous systems. In this paper we present a practice on how to exploit and orchestrate parallelism at algorithm level to take advantage of underlying parallelism at architecture level. A potential Petaflops application -- cryo-EM 3D reconstruction is selected as an example. We exploit all possible parallelism in cryo-EM 3D reconstruction, and leverage a self-adaptive dynamic scheduling algorithm to create a proper parallelism mapping between the application and architecture. The parallelized programs are evaluated on a subsystem of Dawning Nebulae supercomputer, whose node is composed of two Intel six-core Xeon CPUs and one Nvidia Fermi GPU. The experiment confirms that hierarchical parallelism is an efficient pattern of parallel programming to utilize capabilities of both CPU and GPU in a heterogeneous system. The CUDA kernels run more than 3 times faster than the OpenMP parallelized ones using 12 cores (threads). Based on the GPU-only version, the hybrid CPU-GPU program further improves the whole application's performance by 30% on the average.
Linchuan Li, Xingjian Li 0002, Guangming Tan, Mingyu Chen 0001, Peiheng Zhang
HPDC5
2009 Short read DNA fragment anchoring algorithm
abstract
BACKGROUND: The emerging next-generation sequencing method based on PCR technology boosts genome sequencing speed considerably, the expense is also get decreased. It has been utilized to address a broad range of bioinformatics problems. Limited by reliable output sequence length of next-generation sequencing technologies, we are confined to study gene fragments with 30 - 50 bps in general and it is relatively shorter than traditional gene fragment length. Anchoring gene fragments in long reference sequence is an essential and prerequisite step for further assembly and analysis works. Due to the sheer number of fragments produced by next-generation sequencing technologies and the huge size of reference sequences, anchoring would rapidly becoming a computational bottleneck. RESULTS AND DISCUSSION: We compared algorithm efficiency on BLAT, SOAP and EMBF. The efficiency is defined as the count of total output results divided by time consumed to retrieve them. The data show that our algorithm EMBF have 3 - 4 times efficiency advantage over SOAP, and at least 150 times over BLAT. Moreover, when the reference sequence size is increased, the efficiency of SOAP will get degraded as far as 30%, while EMBF have preferable increasing tendency. CONCLUSION: In conclusion, we deem that EMBF is more suitable for short fragment anchoring problem where result completeness and accuracy is predominant and the reference sequences are relatively large.
Wendi Wang 0002, Peiheng Zhang, Xinchun Liu
BMC Bioinform.2
2007 Survey on index based homology search algorithms
Xianyang Jiang, Peiheng Zhang, Xinchun Liu, Stephen S.-T. Yau
J. Supercomput.2