Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiahua He

dblp:90/5671 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
3since 2021 · last 2022
0000-0002-9075-4978ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 30% Memory systems · 25% High-performance computing · 25%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
structural bioinformatics
0.822020
Topology-independent and global protein structure alignment through an FFT-based algorithm · Bioinform. 2020
PepBDB: a comprehensive structural database of biological peptide-protein interactions · Bioinform. 2019
Bioinformatics and computational biology › protein structure prediction
protein-protein docking
0.612022
TRScore: a 3D RepVGG-based scoring method for ranking protein docking models · Bioinform. 2022
Bioinformatics and computational biology › molecular informatics › molecular modeling
scoring function
0.612022
TRScore: a 3D RepVGG-based scoring method for ranking protein docking models · Bioinform. 2022
Bioinformatics and computational biology › protein structure prediction
de novo structure prediction
0.512021
Full-length de novo protein structure determination from cryo-EM maps using deep learning · Bioinform. 2021
Bioinformatics and computational biology
structural biology
0.512021
Full-length de novo protein structure determination from cryo-EM maps using deep learning · Bioinform. 2021
Bioinformatics and computational biology › protein structure analysis
protein structure alignment
0.412020
Topology-independent and global protein structure alignment through an FFT-based algorithm · Bioinform. 2020
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular docking
0.412019
Protein-ensemble-RNA docking by efficient consideration of protein flexibility through homology models · Bioinform. 2019
Bioinformatics and computational biology › protein interaction
protein-peptide interaction
0.412019
PepBDB: a comprehensive structural database of biological peptide-protein interactions · Bioinform. 2019
Bioinformatics and computational biology
protein structure prediction
0.112021
Full-length de novo protein structure determination from cryo-EM maps using deep learning · Bioinform. 2021
Storage systems
flash and SSD
0.122010
DASH: a Recipe for a Flash-based Data Intensive Supercomputer · SC 2010
Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing · SC 2010
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA-protein interaction prediction
0.112019
Protein-ensemble-RNA docking by efficient consideration of protein flexibility through homology models · Bioinform. 2019
High-performance computing
data-intensive computing
0.112010
DASH: a Recipe for a Flash-based Data Intensive Supercomputer · SC 2010
Memory systems
non-volatile memory
0.112010
Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing · SC 2010
Performance modeling and evaluation
workload characterization
0.112008
Code coverage, performance approximation and automatic recognition of idioms in scientific applications · HPDC 2008
Operating systems › resource management › memory management
virtual memory
0.012010
Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing · SC 2010
Memory systems › shared memory
distributed shared memory
0.012010
DASH: a Recipe for a Flash-based Data Intensive Supercomputer · SC 2010
Storage systems › flash and SSD
solid-state drive
0.012010
Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing · SC 2010

Methods — techniques the papers use, named apart from their topics

voxelization · 0.6convolutional neural network · 0.63D RepVGG · 0.6densely connected convolutional networks · 0.5deep learning · 0.5FFT-based exhaustive search · 0.4scoring function · 0.4homology modeling · 0.4ensemble docking · 0.4docking and scoring data preparation · 0.4performance measurement · 0.2idiom benchmarking · 0.2code coverage analysis · 0.2performance evaluation · 0.1
YearPublicationVenuePosition
2022 TRScore: a 3D RepVGG-based scoring method for ranking protein docking models
abstract
MOTIVATION: Protein-protein interactions (PPI) play important roles in cellular activities. Due to the technical difficulty and high cost of experimental methods, there are considerable interests towards the development of computational approaches, such as protein docking, to decipher PPI patterns. One of the important and difficult aspects in protein docking is recognizing near-native conformations from a set of decoys, but unfortunately, traditional scoring functions still suffer from limited accuracy. Therefore, new scoring methods are pressingly needed in methodological and/or practical implications. RESULTS: We present a new deep learning-based scoring method for ranking protein-protein docking models based on a 3D RepVGG network, named TRScore. To recognize near-native conformations from a set of decoys, TRScore voxelizes the protein-protein interface into a 3D grid labeled by the number of atoms in different physicochemical classes. Benefiting from the deep convolutional RepVGG architecture, TRScore can effectively capture the subtle differences between energetically favorable near-native models and unfavorable non-native decoys without needing extra information. TRScore was extensively evaluated on diverse test sets including protein-protein docking benchmark 5.0 update set, DockGround decoy set, as well as realistic CAPRI decoy set and overall obtained a significant improvement over existing methods in cross-validation and independent evaluations. AVAILABILITY AND IMPLEMENTATION: Codes available at: https://github.com/BioinformaticsCSU/TRScore.
Linyuan Guo, Jiahua He, Peicong Lin, Sheng-You Huang, Jianxin Wang 0001
Bioinform.2
2021 EMNUSS: a deep learning framework for secondary structure annotation in cryo-EM maps
abstract
Cryo-electron microscopy (cryo-EM) has become one of important experimental methods in structure determination. However, despite the rapid growth in the number of deposited cryo-EM maps motivated by advances in microscopy instruments and image processing algorithms, building accurate structure models for cryo-EM maps remains a challenge. Protein secondary structure information, which can be extracted from EM maps, is beneficial for cryo-EM structure modeling. Here, we present a novel secondary structure annotation framework for cryo-EM maps at both intermediate and high resolutions, named EMNUSS. EMNUSS adopts a three-dimensional (3D) nested U-net architecture to assign secondary structures for EM maps. Tested on three diverse datasets including simulated maps, middle resolution experimental maps, and high-resolution experimental maps, EMNUSS demonstrated its accuracy and robustness in identifying the secondary structures for cyro-EM maps of various resolutions. The EMNUSS program is freely available at http://huanglab.phys.hust.edu.cn/EMNUSS.
Jiahua He, Sheng-You Huang
Briefings Bioinform.1
2021 Full-length de novo protein structure determination from cryo-EM maps using deep learning
abstract
MOTIVATION: Advances in microscopy instruments and image processing algorithms have led to an increasing number of Cryo-electron microscopy (cryo-EM) maps. However, building accurate models for the EM maps at 3-5 Å resolution remains a challenging and time-consuming process. With the rapid growth of deposited EM maps, there is an increasing gap between the maps and reconstructed/modeled three-dimensional (3D) structures. Therefore, automatic reconstruction of atomic-accuracy full-atom structures from EM maps is pressingly needed. RESULTS: We present a semi-automatic de novo structure determination method using a deep learning-based framework, named as DeepMM, which builds atomic-accuracy all-atom models from cryo-EM maps at near-atomic resolution. In our method, the main-chain and Cα positions as well as their amino acid and secondary structure types are predicted in the EM map using Densely Connected Convolutional Networks. DeepMM was extensively validated on 40 simulated maps at 5 Å resolution and 30 experimental maps at 2.6-4.8 Å resolution as well as an Electron Microscopy Data Bank-wide dataset of 2931 experimental maps at 2.6-4.9 Å resolution, and compared with state-of-the-art algorithms including RosettaES, MAINMAST and Phenix. Overall, our DeepMM algorithm obtained a significant improvement over existing methods in terms of both accuracy and coverage in building full-length protein structures on all test sets, demonstrating the efficacy and general applicability of DeepMM. AVAILABILITY AND IMPLEMENTATION: http://huanglab.phys.hust.edu.cn/DeepMM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiahua He, Sheng-You Huang
Bioinform.1
2020 Topology-independent and global protein structure alignment through an FFT-based algorithm
abstract
MOTIVATION: Protein structure alignment is one of the fundamental problems in computational structure biology. A variety of algorithms have been developed to address this important issue in the past decade. However, due to their heuristic nature, current structure alignment methods may suffer from suboptimal alignment and/or over-fragmentation and thus lead to a biologically wrong alignment in some cases. To overcome these limitations, we have developed an accurate topology-independent and global structure alignment method through an FFT-based exhaustive search algorithm, which is referred to as FTAlign. RESULTS: Our FTAlign algorithm was extensively tested on six commonly used datasets and compared with seven state-of-the-art structure alignment approaches, TMalign, DeepAlign, Kpax, 3DCOMB, MICAN, SPalignNS and CLICK. It was shown that FTAlign outperformed the other methods in reproducing manually curated alignments and obtained a high success rate of 96.7 and 90.0% on two gold-standard benchmarks, MALIDUP and MALISAM, respectively. Moreover, FTAlign also achieved the overall best performance in terms of biologically meaningful structure overlap (SO) and TMscore on both the sequential alignment test sets including MALIDUP, MALISAM and 64 difficult cases from HOMSTRAD, and the non-sequential sets including MALIDUP-NS, MALISAM-NS, 199 topology-different cases, where FTAlign especially showed more advantage for non-sequential alignment. Despite its global search feature, FTAlign is also computationally efficient and can normally complete a pairwise alignment within one second. AVAILABILITY AND IMPLEMENTATION: http://huanglab.phys.hust.edu.cn/ftalign/.
Zeyu Wen, Jiahua He, Sheng-You Huang
Bioinform.2
2019 Protein-ensemble-RNA docking by efficient consideration of protein flexibility through homology models
abstract
MOTIVATION: Given the importance of protein-ribonucleic acid (RNA) interactions in many biological processes, a variety of docking algorithms have been developed to predict the complex structure from individual protein and RNA partners in the past decade. However, due to the impact of molecular flexibility, the performance of current methods has hit a bottleneck in realistic unbound docking. Pushing the limit, we have proposed a protein-ensemble-RNA docking strategy to explicitly consider the protein flexibility in protein-RNA docking through an ensemble of multiple protein structures, which is referred to as MPRDock. Instead of taking conformations from MD simulations or experimental structures, we obtained the multiple structures of a protein by building models from its homologous templates in the Protein Data Bank (PDB). RESULTS: Our approach can not only avoid the reliability issue of structures from MD simulations but also circumvent the limited number of experimental structures for a target protein in the PDB. Tested on 68 unbound-bound and 18 unbound-unbound protein-RNA complexes, our MPRDock/DITScorePR considerably improved the docking performance and achieved a significantly higher success rate than single-protein rigid docking whether pseudo-unbound templates are included or not. Similar improvements were also observed when combining our ensemble docking strategy with other scoring functions. The present homology model-based ensemble docking approach will have a general application in molecular docking for other interactions. AVAILABILITY AND IMPLEMENTATION: http://huanglab.phys.hust.edu.cn/mprdock/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiahua He, Huanyu Tao, Sheng-You Huang
Bioinform.1
2019 PepBDB: a comprehensive structural database of biological peptide-protein interactions
abstract
Summary: A structural database of peptide-protein interactions is important for drug discovery targeting peptide-mediated interactions. Although some peptide databases, especially for special types of peptides, have been developed, a comprehensive database of cleaned peptide-protein complex structures is still not available. Such cleaned structures are valuable for docking and scoring studies in structure-based drug design. Here, we have developed PepBDB-a curated Peptide Binding DataBase of biological complex structures from the Protein Data Bank (PDB). PepBDB presents not only cleaned structures but also extensive information about biological peptide-protein interactions, and allows users to search the database with a variety of options and interactively visualize the search results. Availability and implementation: PepBDB is available at http://huanglab.phys.hust.edu.cn/pepbdb/.
Zeyu Wen, Jiahua He, Huanyu Tao, Sheng-You Huang
Bioinform.2
2011 Automatic Recognition of Performance Idioms in Scientific Applications
abstract
Basic data flow patterns that we call \textbf{performance idioms}, such as stream, transpose, reduction, random access and stencil, are common in scientific numerical applications. We hypothesize that a small number of idioms can cover most programming constructs that dominate the execution time of scientific codes and can be used to approximate the application performance. To check these hypotheses, we proposed an automatic idioms recognition method and implemented the method, based on the open source compiler Open64. With the NAS Parallel Benchmark (NPB) as a case study, the prototype system is about $90%$ accurate compared with idiom classification by a human expert. Our results showed that the above five idioms suffice to cover $100%$ of the six NPB codes (MG, CG, FT, BT, SP and LU). We also compared the performance of our idiom benchmarks with their corresponding instances in the NPB codes on two different platforms with different methods. The approximation accuracy is up to $96.6%$. The contribution is to show that a small set of idioms can cover more complex codes, that idioms can be recognized automatically, and that suitably defined idioms may approximate application performance.
Jiahua He, Allan Snavely, Rob F. Van der Wijngaart, Michael A. Frumkin
IPDPS1
2010 Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive Computing
abstract
Emerging storage technologies such as flash memories, phase-change memories, and spin-transfer torque memories are poised to close the enormous performance gap between disk-based storage and main memory. We evaluate several approaches to integrating these memories into computer systems by measuring their impact on IO-intensive, database, and memory-intensive applications. We explore several options for connecting solid-state storage to the host system and find that the memories deliver large gains in sequential and random access performance, but that different system organizations lead to different performance trade-offs. The memories provide substantial application-level gains as well, but overheads in the OS, file system, and application can limit performance. As a result, fully exploiting these memories' potential will require substantial changes to application and system software. Finally, paging to fast non-volatile memories is a viable option for some applications, providing an alternative to expensive, powerhungry DRAM for supporting scientific applications with large memory footprints.
Adrian M. Caulfield, Joel Coburn, Todor I. Mollov, Arup De, Ameen Akel, Jiahua He, Arun Jagatheesan, Rajesh K. Gupta 0001, Allan Snavely, Steven Swanson
SC6
2010 DASH: a Recipe for a Flash-based Data Intensive Supercomputer
abstract
Data intensive computing can be defined as computation involving large datasets and complicated I/O patterns. Data intensive computing is challenging because there is a five-orders-of-magnitude latency gap between main memory DRAM and spinning hard disks; the result is that an inordinate amount of time in data intensive computing is spent accessing data on disk. To address this problem we designed and built a prototype data intensive supercomputer named DASH that exploits flash-based Solid State Drive (SSD) technology and also virtually aggregated DRAM to fill the latency gap . DASH uses commodity parts including Intel® X25-E flash drives and distributed shared memory (DSM) software from ScaleMP®. The system is highly competitive with several commercial offerings by several metrics including achieved IOPS (input output operations per second), IOPS per dollar of system acquisition cost, IOPS per watt during operation, and IOPS per gigabyte (GB) of available storage. We present here an overview of the design of DASH, an analysis of its cost efficiency, then a detailed recipe for how we designed and tuned it for high data-performance, lastly show that running data-intensive scientific applications from graph theory, biology, and astronomy, we achieved as much as two orders-of- magnitude speedup compared to the same applications run on traditional architectures.
Jiahua He, Arun Jagatheesan, Sandeep K. S. Gupta, Jeffrey Bennett, Allan Snavely
SC1
2008 Code coverage, performance approximation and automatic recognition of idioms in scientific applications
abstract
Basic data flow patterns which we call idioms, such as stream, transpose, reduction, random access and stencil, are common in scientific numerical applications. We hypothesize that a small number of idioms can cover most programming constructs that dominate the execution time of scientific codes and can be used to approximate the application performance. In this paper, we start with a manual analysis of code coverage on the NAS Parallel Benchmark (NPB) and find that five idioms suffice to cover 100% of the NPB codes. We then compare the performance of our idiom benchmarks and their corresponding instances in different NPB codes on two different platforms and find that they differ by about 30%. To check the hypotheses with real applications further, we propose an automatic idioms recognition method, implement the method basing on the open source compiler Open64, and verify the prototype system with the previous manual analysis results.
Jiahua He, Allan Snavely, Rob F. Van der Wijngaart, Michael A. Frumkin
HPDC1