Haoyu Cheng

dblp:164/0927 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Co-design of traffic-aware dynamic VC partitioning and congestion-aware routing in CPU-GPU heterogeneous NoCs
Juan Fang 0004, Haoyu Cheng, Yuening Wang, Juncheng Chen
J. Supercomput.3
2025 DRCD: a regional-contention-driven arbitration policy for CPU-GPU heterogeneous systems
abstract
In CPU–GPU heterogeneous systems, there exists intense resource contention between CPUs and GPUs. Traditional resource arbitration policies fail to account for the heterogeneity of cores, leading to inefficient network resource utilization for the CPU, which negatively impacts its performance. In heterogeneous networks, the degree of resource contention varies across different regions. This paper first uses reinforcement learning to analyze the message feature weights relied upon for resource arbitration in different network regions. To achieve more efficient resource allocation, a regional-contention-driven arbitration policy is proposed. The simulation results show that, compared to traditional arbitration policy, the overall network latency is reduced by 7.99%, and CPU performance is improved by 11.42%. Furthermore, a dynamic regional-contention-driven arbitration policy is proposed, which further reduces the overall network latency by 10.47% and increases CPU performance by 16.79% compared to traditional arbitration policy.
Juan Fang 0004, Haoyu Cheng, Yuening Wang, Ran Zhai
J. Supercomput.2
2024 Locate N' Rotate: Two-Stage Openable Part Detection with Foundation Model Priors
Siqi Li 0009, Xiaoxue Chen, Haoyu Cheng, Guyue Zhou, Hao Zhao 0002, Guanzhong Tian
ACCV (7)3
2023 DPBC-VCP: A Network-On-Chip Prioritization Mechanism Combined with VCP for CPU-GPU Heterogeneous Systems
abstract
When executing CPU and GPU applications in CPU-GPU heterogeneous systems, a common phenomenon arises where CPU applications performance is often interfered by GPU applications. This study substantiates this observation through an analysis of resource contention and identifies the limitations of the Virtual Channel Partitioning (VCP) approach in the crossbar switch allocation stage. In response to the resource contention problem in crossbar switch allocation stage, we propose a Probability-Based CPU-first Arbitration Strategy that enhances the priority of CPU packets in contention through specific probabilities. Furthermore, we introduce a Dynamic Probability-Based CPU-first Arbitration Strategy (DPBC) that dynamically selects probability values based on application execution phases to strike a balance between optimal CPU and GPU performance. Moreover, we combine this dynamic strategy with VCP to further enhance CPU performance, propose the DPBC-VCP method. The DPBC-VCP, a combination of network partitioning and prioritization techniques, yields an average enhancement of 48% in CPU performance compared to the baseline, with only a marginal 2.45% reduction in GPU performance.
Haoyu Cheng, Zhichao Wei, Huijing Yang
ICPADS2
2021 Real-time mapping of nanopore raw signals
abstract
MOTIVATION: Oxford Nanopore Technologies sequencing devices support adaptive sequencing, in which undesired reads can be ejected from a pore in real time. This feature allows targeted sequencing aided by computational methods for mapping partial reads, rather than complex library preparation protocols. However, existing mapping methods either require a computationally expensive base-calling procedure before using aligners to map partial reads or work well only on small genomes. RESULTS: In this work, we present a new streaming method that can map nanopore raw signals for real-time selective sequencing. Rather than converting read signals to bases, we propose to convert reference genomes to signals and fully operate in the signal space. Our method features a new way to index reference genomes using k-d trees, a novel seed selection strategy and a seed chaining algorithm tailored toward the current signal characteristics. We implemented the method as a tool Sigmap. Then we evaluated it on both simulated and real data and compared it to the state-of-the-art nanopore raw signal mapper Uncalled. Our results show that Sigmap yields comparable performance on mapping yeast simulated raw signals, and better mapping accuracy on mapping yeast real raw signals with a 4.4× speedup. Moreover, our method performed well on mapping raw signals to genomes of size >100 Mbp and correctly mapped 11.49% more real raw signals of green algae, which leads to a significantly higher F1-score (0.9354 versus 0.8660). AVAILABILITY AND IMPLEMENTATION: Sigmap code is accessible at https://github.com/haowenz/sigmap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haoran Li 0015, Haoyu Cheng, Kin Fai Au, Heng Li 0002, Srinivas Aluru
Bioinform.4
2019 BitMapper2: A GPU-Accelerated All-Mapper Based on the Sparse q-Gram Index
abstract
The explosive growth of next-generation sequencing (NGS) read datasets drives a need for new faster read mappers. One class of read mappers, called all-mappers, is designed to identify all mapping locations of each read. Many all-mappers have been developed over the past few years, but they are either time-consuming or memory-consuming. Here, we present BitMapper2, a GPU-accelerated read mapper that reports all mapping locations of NGS reads. To make full use of the parallel processing capability of GPUs, BitMapper2 proposes the sparse q-gram index, which reduces the memory requirement and the data transfer time between GPU and CPU. We also design the filtration part and the verification part of BitMapper2 specifically for the architecture of GPU. In addition, BitMapper2 is still time-efficient and memory-efficient even if there is no GPU available. Experiments show that BitMapper2 was significantly faster than the state-of-the-art all-mappers, while requiring less space.
Haoyu Cheng
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 FMtree: a fast locating algorithm of FM-indexes for genomic data
abstract
Motivation: As a fundamental task in bioinformatics, searching for massive short patterns over a long text has been accelerated by various compressed full-text indexes. These indexes are able to provide similar searching functionalities to classical indexes, e.g. suffix trees and suffix arrays, while requiring less space. For genomic data, a well-known family of compressed full-text indexes, called FM-indexes, presents unmatched performance in practice. One major drawback of FM-indexes is that their locating operations, which report all occurrence positions of patterns in a given text, are not efficient, especially for the patterns with many occurrences. Results: In this paper, we introduce a novel locating algorithm, FMtree, to fast retrieve all occurrence positions of any pattern via FM-indexes. When searching for a pattern over a given text, FMtree organizes the search space of the locating operation into a conceptual multiway tree. As a result, multiple occurrence positions of this pattern can be retrieved simultaneously by traversing the multiway tree. Compared with existing locating algorithms, our tree-based algorithm reduces large numbers of redundant operations and presents better data locality. Experimental results show that FMtree is usually one order of magnitude faster than the state-of-the-art algorithms, and still memory-efficient. Availability and implementation: FMtree is freely available at https://github.com/chhylp123/FMtree. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Haoyu Cheng, Ming Wu 0003
Bioinform.1
2015 BitMapper: an efficient all-mapper based on bit-vector computing
abstract
BACKGROUND: As the next-generation sequencing (NGS) technologies producing hundreds of millions of reads every day, a tremendous computational challenge is to map NGS reads to a given reference genome efficiently. However, existing methods of all-mappers, which aim at finding all mapping locations of each read, are very time consuming. The majority of existing all-mappers consist of 2 main parts, filtration and verification. This work significantly reduces verification time, which is the dominant part of the running time. RESULTS: An efficient all-mapper, BitMapper, is developed based on a new vectorized bit-vector algorithm, which simultaneously calculates the edit distance of one read to multiple locations in a given reference genome. Experimental results on both simulated and real data sets show that BitMapper is from several times to an order of magnitude faster than the current state-of-the-art all-mappers, while achieving higher sensitivity, i.e., better quality solutions. CONCLUSIONS: We present BitMapper, which is designed to return all mapping locations of raw reads containing indels as well as mismatches. BitMapper is implemented in C under a GPL license. Binaries are freely available at http://home.ustc.edu.cn/%7Echhy.
Haoyu Cheng, Huaipan Jiang, Jiaoyun Yang, Yi Shang
BMC Bioinform.1