Dinghua Li

dblp:162/5751 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence alignment
DNA-protein alignment
0.312018
AC-DIAMOND v1: accelerating large-scale DNA-protein alignment · Bioinform. 2018
Bioinformatics and computational biology
sequence alignment
0.312018
AC-DIAMOND v1: accelerating large-scale DNA-protein alignment · Bioinform. 2018
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly › de novo assembly
de bruijn graph assembly
0.212015
MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph · Bioinform. 2015
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.212015
MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph · Bioinform. 2015
Bioinformatics and computational biology › metagenomics
metagenomic assembly
0.212015
MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph · Bioinform. 2015
Bioinformatics and computational biology › multiple sequence alignment
large-scale sequence alignment
0.112018
AC-DIAMOND v1: accelerating large-scale DNA-protein alignment · Bioinform. 2018
Bioinformatics and computational biology › sequence analysis
read mapping
0.112018
AC-DIAMOND v1: accelerating large-scale DNA-protein alignment · Bioinform. 2018
Bioinformatics and computational biology
metagenomics
0.112015
MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph · Bioinform. 2015

Methods — techniques the papers use, named apart from their topics

compressed seed index · 0.3SIMD dynamic programming · 0.3succinct de bruijn graph · 0.2GPU acceleration · 0.2
YearPublicationVenuePosition
2018 BitFlow: Exploiting Vector Parallelism for Binary Neural Networks on CPU
abstract
Deep learning has revolutionized computer vision and other fields since its big bang in 2012. However, it is challenging to deploy Deep Neural Networks (DNNs) into real-world applications due to their high computational complexity. Binary Neural Networks (BNNs) dramatically reduce computational complexity by replacing most arithmetic operations with bitwise operations. Existing implementations of BNNs have been focusing on GPU or FPGA, and using the conventional image-to-column method that doesn't perform well for binary convolution due to low arithmetic intensity and unfriendly pattern for bitwise operations. We propose BitFlow, a gemm-operator-network three-level optimization framework for fully exploiting the computing power of BNNs on CPU. BitFlow features a new class of algorithm named PressedConv for efficient binary convolution using locality-aware layout and vector parallelism. We evaluate BitFlow with the VGG network. On a single core of Intel Xeon Phi, BitFlow obtains 1.8x speedup over unoptimized BNN implementations, and 11.5x speedup over counterpart full-precision DNNs. Over 64 cores, BitFlow enables BNNs to run 1.1x faster than counterpart full-precision DNNs on GPU (GTX 1080).
Jidong Zhai, Dinghua Li, Yifan Gong 0003, Yuhao Zhu 0001, Wei Liu 0143, Jiangming Jin
IPDPS3
2018 AC-DIAMOND v1: accelerating large-scale DNA-protein alignment
abstract
Summary: AC-DIAMOND (v1) is a DNA-protein alignment tool designed to tackle the efficiency challenge of aligning large amount of reads or contigs to protein databases. When compared with the previously most efficient method DIAMOND, AC-DIAMOND gains a 6- to 7-fold speed-up, while retaining a similar degree of sensitivity. The improvement is rooted at two aspects: first, using a compressed index of seeds with adaptive-length to speed-up the matching between query and reference sequences; second, adopting a compact form of dynamic programing to fully utilize the parallelism of the SIMD capability. Availability and implementation: Software source codes and binaries available at https://github.com/Maihj/AC-DIAMOND/. Supplementary information: Supplementary data are available at Bioinformatics online.
Huijun Mai, Dinghua Li, Henry C. M. Leung, Ruibang Luo, Chi-Kwong Wong, Hing-Fung Ting, Tak Wah Lam
Bioinform.3
2017 MegaGTA: a sensitive and accurate metagenomic gene-targeted assembler using iterative de Bruijn graphs
abstract
BACKGROUND: The recent release of the gene-targeted metagenomics assembler Xander has demonstrated that using the trained Hidden Markov Model (HMM) to guide the traversal of de Bruijn graph gives obvious advantage over other assembly methods. Xander, as a pilot study, indeed has a lot of room for improvement. Apart from its slow speed, Xander uses only 1 k-mer size for graph construction and whatever choice of k will compromise either sensitivity or accuracy. Xander uses a Bloom-filter representation of de Bruijn graph to achieve a lower memory footprint. Bloom filters bring in false positives, and it is not clear how this would impact the quality of assembly. Xander does not keep track of the multiplicity of k-mers, which would have been an effective way to differentiate between erroneous k-mers and correct k-mers. RESULTS: In this paper, we present a new gene-targeted assembler MegaGTA, which attempts to improve Xander in different aspects. Quality-wise, it utilizes iterative de Bruijn graphs to take full advantage of multiple k-mer sizes to make the best of both sensitivity and accuracy. Computation-wise, it employs succinct de Bruijn graphs (SdBG) to achieve low memory footprint and high speed (the latter is benefited from a highly efficient parallel algorithm for constructing SdBG). Unlike Bloom filters, an SdBG is an exact representation of a de Bruijn graph. It enables MegaGTA to avoid false-positive contigs and to easily incorporate the multiplicity of k-mers for building better HMM model. We have compared MegaGTA and Xander on an HMP-defined mock metagenomic dataset, and showed that MegaGTA excelled in both sensitivity and accuracy. On a large rhizosphere soil metagenomic sample (327Gbp), MegaGTA produced 9.7-19.3% more contigs than Xander, and these contigs were assigned to 10-25% more gene references. In our experiments, MegaGTA, depending on the number of k-mers used, is two to ten times faster than Xander. CONCLUSION: MegaGTA improves on the algorithm of Xander and achieves higher sensitivity, accuracy and speed. Moreover, it is capable of assembling gene sequences from ultra-large metagenomic datasets. Its source code is freely available at https://github.com/HKU-BAL/megagta .
Dinghua Li, Henry C. M. Leung, Ruibang Luo, Hing-Fung Ting, Tak Wah Lam
BMC Bioinform.1
2015 MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph
abstract
Abstract Summary: MEGAHIT is a NGS de novo assembler for assembling large and complex metagenomics data in a time- and cost-efficient manner. It finished assembling a soil metagenomics dataset with 252 Gbps in 44.1 and 99.6 h on a single computing node with and without a graphics processing unit, respectively. MEGAHIT assembles the data as a whole, i.e. no pre-processing like partitioning and normalization was needed. When compared with previous methods on assembling the soil data, MEGAHIT generated a three-time larger assembly, with longer contig N50 and average contig length; furthermore, 55.8% of the reads were aligned to the assembly, giving a fourfold improvement. Availability and implementation: The source code of MEGAHIT is freely available at https://github.com/voutcn/megahit under GPLv3 license. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Dinghua Li, Chi-Man Liu, Ruibang Luo, Kunihiko Sadakane, Tak Wah Lam
Bioinform.1