Haichang Yao

dblp:226/6725 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2022
0000-0002-5751-960XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 FQCSpark: Efficient Spark-based Parallel Compression Algorithm for FASTQ Genome Sequences
abstract
The rapid development of Next-Generation Sequencing (NGS) technologies has posed serious challenges to the storage and transmission of genomic data, and the bioinformatics community urgently needs efficient genome compression algorithms to support genome analysis. The existing acceleration approaches for genome compression algorithms are mostly multithreading and limited to a single machine, which cannot adapt to the demand of large-scale genome compression in the distributed environment of cloud computing. In this paper, we propose a Spark-based efficient parallel compression algorithm for FASTQ genome sequences - FQCSpark. Experimental results show that FQCSpark outperforms existing algorithms with good compression ratios by several times in speed. This is due to the fine-grained degree of parallelism design and well-designed parallel operator flow. More importantly, the degree of parallelism design in this paper is also applicable to other algorithms which compress blocks independently. Meanwhile, FQCSpark provides good compression ratios, especially on the S.cerevisiae dataset, which is 15.4% better than the latest open-source genome compression tool - Genozip. FQCSpark is the first known Spark-based parallel compression algorithm for FASTQ genome sequences.
Yimu Ji 0001, Hu Fa, Haichang Yao, Mengxue Wu, Houzhi Fang, Shangdong Liu
CSCWD3
2022 SparkGC: Spark based genome compression for large collections of genomes
abstract
Since the completion of the Human Genome Project at the turn of the century, there has been an unprecedented proliferation of sequencing data. One of the consequences is that it becomes extremely difficult to store, backup, and migrate enormous amount of genomic datasets, not to mention they continue to expand as the cost of sequencing decreases. Herein, a much more efficient and scalable program to perform genome compression is required urgently. In this manuscript, we propose a new Apache Spark based Genome Compression method called SparkGC that can run efficiently and cost-effectively on a scalable computational cluster to compress large collections of genomes. SparkGC uses Spark's in-memory computation capabilities to reduce compression time by keeping data active in memory between the first-order and second-order compression. The evaluation shows that the compression ratio of SparkGC is better than the best state-of-the-art methods, at least better by 30%. The compression speed is also at least 3.8 times that of the best state-of-the-art methods on only one worker node and scales quite well with the number of nodes. SparkGC is of significant benefit to genomic data storage and transmission. The source code of SparkGC is publicly available at https://github.com/haichangyao/SparkGC .
Haichang Yao, GuangYong Hu, Shangdong Liu, Houzhi Fang, Yimu Ji 0001
BMC Bioinform.1
2022 DVO + LCLMF: A web service recommendation mechanism with QoS privacy preservation
abstract
Abstract QoS‐aware based web service recommendation is one of the crucial solutions to help users find high‐quality web services. To accurately predict the QoS values of candidate services, it is usually required to collect historical QoS data of users (QoS data for short). If these collected QoS data are improperly processed, QoS data privacy may be threatened. However, how to accurately predict the QoS values of candidate services while protecting QoS data privacy has not been well studied. In response to the situation, we propose a hybrid web service recommendation mechanism, which is divided into three parts. In the first part, the QoS data privacy preservation algorithm, which called DVO, is proposed based on keeping the cosine similarity of QoS data unchanged, that is, to realize the confusion of QoS data while ensuring the availability of QoS data remains unchanged. In the second part, a hybrid matrix factorization model based on location information and service features, which called LCLMF, is proposed to improve the accuracy of QoS values prediction. According to DVO and LCLMF, the DVO + LCLMF is designed in the third part, which can accurately predict QoS values while protecting QoS data privacy. The experimental results show that DVO + LCLMF can accurately predict the QoS values of candidate services on the basis of attaining QoS data privacy protection.
Yimu Ji 0001, Shangdong Liu, Fei Wu 0004, Haichang Yao, Jing He 0004, Yanlan Liu, Shuai You
Concurr. Comput. Pract. Exp.5
2022 Parallel compression for large collections of genomes
abstract
Summary With the development of genome sequencing technology, the cost of genome sequencing is continuously reducing, while the efficiency is increasing. Therefore, the amount of genomic data has been increasing exponentially, making the transmission and storage of genomic data an enormous challenge. Although many excellent genome compression algorithms have been proposed, an efficient compression algorithm for large collections of FASTA genomes, especially can be used in the distributed system of cloud computing, is still lacking. This article proposes two optimization schemes based on HRCM compression method. One is MtHRCM adopting multi‐thread parallel technology. The other is HadoopHRCM adopting distributed computing parallel technology. Experiments show that the schemes recognizably improve the compression speed of HRCM. Moreover, BSC algorithm instead of PPMD algorithm is used in the new schemes, the compression ratio is improved by 20% compared with HRCM. In addition, our proposed methods also perform well in robustness and scalability. The Java source codes of MtHRCM and HadoopHRCM can be freely downloaded from https://github.com/haicy/MtHRCM and https://github.com/haicy/HadoopHRCM .
Haichang Yao, Shangdong Liu, Yimu Ji 0001, GuangYong Hu, Ruchuan Wang 0001
Concurr. Comput. Pract. Exp.1
2022 Lane marking detection algorithm based on high-precision map and multisensor fusion
abstract
Summary In case of sharp road illumination changes, bad weather such as rain, snow or fog, wear or missing of the lane marking, the reflective water stain on the road surface, the shadow obstruction of the tree, and mixed lane markings and other signs, missing detection or wrong detection will occur for the traditional lane marking detection algorithm. In this manuscript, a lane marking detection algorithm based on high‐precision map and multisensor fusion is proposed. The basic principle of the algorithm is to use the centimeter‐level high‐precision positioning combined with high‐precision map data to complete the detection of lane markings. In the process of generating high‐precision maps or in the uncovered areas of high‐precision maps, LIDAR (LIght Detection And Ranging) is used to estimate the curvature of the road to assist in lane marking detection. The experimental results show that the algorithm has lower false detection rate in case of bad road conditions, and the algorithm is robust.
Haichang Yao, Shangdong Liu, Yimu Ji 0001, Guangyan Huang, Ruchuan Wang 0001
Concurr. Comput. Pract. Exp.1
2021 ACEA: A Queueing Model-Based Elastic Scaling Algorithm for Container Cluster
abstract
Elastic scaling is one of the techniques to deal with the sudden change of the number of tasks and the long average waiting time of tasks in the container cluster. The unreasonable resource supply may lead to the low comprehensive resource utilization rate of the cluster. Therefore, balancing the relationship between the average waiting time of tasks and the comprehensive resource utilization rate of the cluster based on the number of tasks is the key to elastic scaling. In this paper, an adaptive scaling algorithm based on the queuing model called ACEA is proposed. This algorithm uses the hybrid multiserver queuing model (M/M/s/K) to quantitatively describe the relationship among number of tasks, average waiting time of tasks, and comprehensive resource utilization rate of cluster and builds the cluster performance model, evaluation function, and quality of service (QoS) constraints. Particle swarm optimization (PSO) is used to search feasible solution space determined by the constraint relation of ACEA quickly, so as to improve the dynamic optimization performance and convergence timeliness of ACEA. The experimental results show that the algorithm can ensure the comprehensive resource utilization rate of the cluster while the average waiting time of tasks meets the requirement.
Yimu Ji 0001, Shangdong Liu, Haichang Yao, Shuai You
Wirel. Commun. Mob. Comput.4
2019 FastDRC: Fast and Scalable Genome Compression Based on Distributed and Parallel Processing
Yimu Ji 0001, Houzhi Fang, Haichang Yao, Jing He 0004, Shangdong Liu
ICA3PP (2)3
2019 Multi-Thread Concurrent Compression Algorithm for Genomic Big Data
abstract
At present, there are many excellent genome compression algorithms with high genome compression ratio. However, there is a lack of highly efficient compression algorithms for simultaneous compression of a large number of genomes. This manuscript presents an algorithm, which is called FastLNGC, for simultaneous compression of a large amount of genome data based on multi-thread concurrency. This algorithm is based on the LNGC (Large Number of Genomes Compressor) algorithm, and adopts multi-thread technology to achieve concurrent processing of genome data compression. A large number of experiments show that FastLNGC has better performance on compression of a large number of genes. The source code of FastLNGC is available at https://github.com/APandaThief/FastLNGC.
Yimu Ji 0001, Haichang Yao, Houzhi Fang, Shangdong Liu, Zhengyuan Xie, Kairui Wang
PDCAT3
2018 VC-TWJoin: A Stream Join Algorithm Based on Variable Update Cycle Time Window
abstract
Stream join is one of the key operations for real-time stream data query and calculation. In light of changeable velocity of stream data, traditional static stream join methods are not so adaptive that stream data computing performance will be affected. Based on the large quantity and constantly changing velocity of stream data, by considering traditional stream join algorithm, this paper proposes an optimized algorithm for variable update cycle stream based on time window (VC-TWJoin, Variable Cycle Time Window Join). For unsteady stream calculated in stream join, the optimal update cycle will be calculated to reduce the response time of stream join and improve join efficiency and real-time capability. Both theoretical analysis and experiments demonstrate that the algorithm is better than traditional join algorithms in terms of real-time capability, join response time and throughput.
Yimu Ji 0001, Shangdong Liu, Lili Lu, Xianbo Lang, Haichang Yao, Ruchuan Wang 0001
CSCWD5
2018 The Study on the Botnet and its Prevention Policies in the Internet of Things
abstract
With the rapid development of Internet of Things (IOT), IOT is more and more important. Also, it faces serious security issues. This paper analyzes Mirai's architecture. The core components are C & C server and Loader server that take charge of command and control, IOT equipments are in charge of broadcast and attack. Paper analyzes Botnet propagation model, Mirais infection attack procedure, impact factor and then proposes the corresponding anti-virus strategy.
Yimu Ji 0001, Shangdong Liu, Haichang Yao, Ruchuan Wang 0001
CSCWD4