EDBT 2026 Demo / reviewers in the wild / expert
Ronghui You
dblp:157/4867
· DBLP profile ↗
14ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-7608-4867ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
4 papers |
Efficient and distributed learning · 67% Trustworthy machine learning · 12% Information extraction and text analysis · 9% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 63% Memory systems · 21% Embedded and real-time systems · 11% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 69% Indexing and storage engines · 31% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
immunoinformatics |
1.2 | 2 | 2023 | DeepMHCI: an anchor position-aware deep interaction model for accurate MHC-I peptide binding affinity prediction · Bioinform. 2023 DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity prediction · Bioinform. 2022 |
Bioinformatics and computational biology
biomedical text mining |
0.9 | 2 | 2021 | BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full text · Bioinform. 2021 FullMeSH: improving large-scale MeSH indexing with full text · Bioinform. 2020 |
Bioinformatics and computational biology › biomedical text mining
MeSH indexing |
0.9 | 2 | 2021 | BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full text · Bioinform. 2021 FullMeSH: improving large-scale MeSH indexing with full text · Bioinform. 2020 |
Machine learning › Efficient and distributed learning › federated learning › robust federated learning
byzantine-robust federated learning |
0.9 | 1 | 2025 | LiD-FL: Towards List-Decodable Federated Learning · AAAI 2025 |
Machine learning › Efficient and distributed learning
federated learning |
0.9 | 1 | 2025 | LiD-FL: Towards List-Decodable Federated Learning · AAAI 2025 |
Bioinformatics and computational biology › protein function prediction
gene ontology term prediction |
0.8 | 2 | 2021 | DeepGraphGO: graph neural network for large-scale, multispecies protein function prediction · Bioinform. 2021 GOLabeler: improving sequence-based large-scale protein function prediction by learning to rank · Bioinform. 2018 |
Bioinformatics and computational biology
protein function prediction |
0.8 | 2 | 2021 | DeepGraphGO: graph neural network for large-scale, multispecies protein function prediction · Bioinform. 2021 GOLabeler: improving sequence-based large-scale protein function prediction by learning to rank · Bioinform. 2018 |
Machine learning › Efficient and distributed learning › federated learning › model aggregation
byzantine-robust aggregation |
0.8 | 1 | 2024 | Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers · AAAI 2024 |
Algorithms and data structures
clustering |
0.8 | 1 | 2024 | Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers · AAAI 2024 |
Algorithms and data structures › clustering › robust clustering
clustering with outliers |
0.8 | 1 | 2024 | Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers · AAAI 2024 |
Bioinformatics and computational biology › immunoinformatics
neoantigen identification |
0.7 | 1 | 2023 | DeepMHCI: an anchor position-aware deep interaction model for accurate MHC-I peptide binding affinity prediction · Bioinform. 2023 |
Bioinformatics and computational biology › molecular property prediction
binding affinity prediction |
0.6 | 1 | 2022 | DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity prediction · Bioinform. 2022 |
Bioinformatics and computational biology › immunoinformatics
peptide-MHC binding prediction |
0.6 | 1 | 2022 | DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity prediction · Bioinform. 2022 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.6 | 1 | 2022 | DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity prediction · Bioinform. 2022 |
Parallel and multicore computing
parallel algorithms |
0.5 | 2 | 2017 | POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time Optimality · PPoPP 2017 Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency · PPoPP 2015 |
Natural language and speech › Information extraction and text analysis › text classification › multi-label text classification
extreme multi-label text classification |
0.4 | 1 | 2019 | AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text Classification · NeurIPS 2019 |
Memory systems › cache
cache-oblivious algorithms |
0.4 | 2 | 2017 | POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time Optimality · PPoPP 2017 Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency · PPoPP 2015 |
Parallel and multicore computing
space-time tradeoff |
0.3 | 1 | 2017 | POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time Optimality · PPoPP 2017 |
Machine learning › Trustworthy machine learning › robustness
adversarial attack |
0.3 | 1 | 2025 | LiD-FL: Towards List-Decodable Federated Learning · AAAI 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | LiD-FL: Towards List-Decodable Federated Learning · AAAI 2025 |
Information retrieval › document retrieval › domain-specific retrieval
biomedical information retrieval |
0.2 | 1 | 2016 | DeepMeSH: deep semantic representation for improving large-scale MeSH indexing · Bioinform. 2016 |
Indexing and storage engines › index construction
web-scale indexing |
0.2 | 1 | 2016 | DeepMeSH: deep semantic representation for improving large-scale MeSH indexing · Bioinform. 2016 |
Embedded and real-time systems › real-time scheduling
dependent task scheduling |
0.2 | 1 | 2015 | Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency · PPoPP 2015 |
Parallel and multicore computing › parallel algorithms
dynamic programming |
0.2 | 1 | 2015 | Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency · PPoPP 2015 |
Parallel and multicore computing
task scheduling |
0.2 | 1 | 2015 | Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency · PPoPP 2015 |
Machine learning › Graph learning
graph neural network |
0.1 | 1 | 2021 | DeepGraphGO: graph neural network for large-scale, multispecies protein function prediction · Bioinform. 2021 |
Information retrieval › retrieval models
neural retrieval |
0.1 | 1 | 2021 | BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full text · Bioinform. 2021 |
Information retrieval
retrieval models |
0.1 | 1 | 2021 | BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full text · Bioinform. 2021 |
Processor architecture and microarchitecture › chip multiprocessor
shared-memory multicore |
0.1 | 1 | 2017 | POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time Optimality · PPoPP 2017 |
Memory systems › cache
cache performance |
0.1 | 1 | 2015 | Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency · PPoPP 2015 |
Methods — techniques the papers use, named apart from their topics
learning to rank · 2.0deep learning · 1.7resilient aggregation · 1.51-mean clustering · 1.51-center clustering · 1.5convolutional neural network · 1.2transfer learning · 1.0multi-label classification · 1.0graph neural network · 1.0BERT · 1.0stochastic gradient descent · 0.9list decoding · 0.9position-wise gated layer · 0.7binding interaction convolution · 0.6attention-based convolutional neural network · 0.4probabilistic label tree · 0.4attention mechanism · 0.4space-time adaptive and reductive technique · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LiD-FL: Towards List-Decodable Federated LearningabstractFederated learning is often used in environments with many unverified participants. Therefore, federated learning under adversarial attacks receives significant attention. This paper proposes an algorithmic framework for list-decodable federated learning, where a central server maintains a list of models, with at least one guaranteed to perform well. The framework has no strict restriction on the fraction of honest clients, extending the applicability of Byzantine federated learning to the scenario with more than half adversaries. Assuming the variance of gradient noise in stochastic gradient descent is bounded, we prove a convergence theorem of our method for strongly convex and smooth losses. Experimental results, including image classification tasks with both convex and non-convex losses, demonstrate that the proposed algorithm can withstand the malicious majority under various attacks. Liren Shan, Ronghui You, Yuhao Yi |
AAAI | 4 |
| 2024 | Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with OutliersabstractByzantine machine learning has garnered considerable attention in light of the unpredictable faults that can occur in large-scale distributed learning systems. The key to secure resilience against Byzantine machines in distributed learning is resilient aggregation mechanisms. Although abundant resilient aggregation rules have been proposed, they are designed in ad-hoc manners, imposing extra barriers on comparing, analyzing, and improving the rules across performance criteria. This paper studies near-optimal aggregation rules using clustering in the presence of outliers. Our outlier-robust clustering approach utilizes geometric properties of the update vectors provided by workers. Our analysis show that constant approximations to the 1-center and 1-mean clustering problems with outliers provide near-optimal resilient aggregators for metric-based criteria, which have been proven to be crucial in the homogeneous and heterogeneous cases respectively. In addition, we discuss two contradicting types of attacks under which no single aggregation rule is guaranteed to improve upon the naive average. Based on the discussion, we propose a two-phase resilient aggregation framework. We run experiments for image classification using a non-convex loss function. The proposed algorithms outperform previously known aggregation rules by a large margin with both homogeneous and heterogeneous data distributions among non-faulty workers. Code and appendix are available at https://github.com/jerry907/AAAI24-RASHB. Yuhao Yi, Ronghui You |
AAAI | 2 |
| 2023 | DeepMHCI: an anchor position-aware deep interaction model for accurate MHC-I peptide binding affinity predictionabstractMOTIVATION: Computationally predicting major histocompatibility complex class I (MHC-I) peptide binding affinity is an important problem in immunological bioinformatics, which is also crucial for the identification of neoantigens for personalized therapeutic cancer vaccines. Recent cutting-edge deep learning-based methods for this problem cannot achieve satisfactory performance, especially for non-9-mer peptides. This is because such methods generate the input by simply concatenating the two given sequences: a peptide and (the pseudo sequence of) an MHC class I molecule, which cannot precisely capture the anchor positions of the MHC binding motif for the peptides with variable lengths. We thus developed an anchor position-aware and high-performance deep model, DeepMHCI, with a position-wise gated layer and a residual binding interaction convolution layer. This allows the model to control the information flow in peptides to be aware of anchor positions and model the interactions between peptides and the MHC pseudo (binding) sequence directly with multiple convolutional kernels. RESULTS: The performance of DeepMHCI has been thoroughly validated by extensive experiments on four benchmark datasets under various settings, such as 5-fold cross-validation, validation with the independent testing set, external HPV vaccine identification, and external CD8+ epitope identification. Experimental results with visualization of binding motifs demonstrate that DeepMHCI outperformed all competing methods, especially on non-9-mer peptides binding prediction. AVAILABILITY AND IMPLEMENTATION: DeepMHCI is publicly available at https://github.com/ZhuLab-Fudan/DeepMHCI. Ronghui You, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 2 |
| 2022 | DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity predictionabstractMOTIVATION: Computationally predicting major histocompatibility complex (MHC)-peptide binding affinity is an important problem in immunological bioinformatics. Recent cutting-edge deep learning-based methods for this problem are unable to achieve satisfactory performance for MHC class II molecules. This is because such methods generate the input by simply concatenating the two given sequences: (the estimated binding core of) a peptide and (the pseudo sequence of) an MHC class II molecule, ignoring biological knowledge behind the interactions of the two molecules. We thus propose a binding core-aware deep learning-based model, DeepMHCII, with a binding interaction convolution layer, which allows to integrate all potential binding cores (in a given peptide) with the MHC pseudo (binding) sequence, through modeling the interaction with multiple convolutional kernels. RESULTS: Extensive empirical experiments with four large-scale datasets demonstrate that DeepMHCII significantly outperformed four state-of-the-art methods under numerous settings, such as 5-fold cross-validation, leave one molecule out, validation with independent testing sets and binding core prediction. All these results and visualization of the predicted binding cores indicate the effectiveness of our model, DeepMHCII, and the importance of properly modeling biological facts in deep learning for high predictive performance and efficient knowledge discovery. AVAILABILITY AND IMPLEMENTATION: DeepMHCII is publicly available at https://github.com/yourh/DeepMHCII. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronghui You, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 1 |
| 2021 | BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full textabstractMOTIVATION: With the rapid increase of biomedical articles, large-scale automatic Medical Subject Headings (MeSH) indexing has become increasingly important. FullMeSH, the only method for large-scale MeSH indexing with full text, suffers from three major drawbacks: FullMeSH (i) uses Learning To Rank, which is time-consuming, (ii) can capture some pre-defined sections only in full text and (iii) ignores the whole MEDLINE database. RESULTS: We propose a computationally lighter, full text and deep-learning-based MeSH indexing method, BERTMeSH, which is flexible for section organization in full text. BERTMeSH has two technologies: (i) the state-of-the-art pre-trained deep contextual representation, Bidirectional Encoder Representations from Transformers (BERT), which makes BERTMeSH capture deep semantics of full text. (ii) A transfer learning strategy for using both full text in PubMed Central (PMC) and title and abstract (only and no full text) in MEDLINE, to take advantages of both. In our experiments, BERTMeSH was pre-trained with 3 million MEDLINE citations and trained on ∼1.5 million full texts in PMC. BERTMeSH outperformed various cutting-edge baselines. For example, for 20 K test articles of PMC, BERTMeSH achieved a Micro F-measure of 69.2%, which was 6.3% higher than FullMeSH with the difference being statistically significant. Also prediction of 20 K test articles needed 5 min by BERTMeSH, while it took more than 10 h by FullMeSH, proving the computational efficiency of BERTMeSH. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronghui You, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 1 |
| 2021 | DeepGraphGO: graph neural network for large-scale, multispecies protein function predictionabstractMOTIVATION: Automated function prediction (AFP) of proteins is a large-scale multi-label classification problem. Two limitations of most network-based methods for AFP are (i) a single model must be trained for each species and (ii) protein sequence information is totally ignored. These limitations cause weaker performance than sequence-based methods. Thus, the challenge is how to develop a powerful network-based method for AFP to overcome these limitations. RESULTS: We propose DeepGraphGO, an end-to-end, multispecies graph neural network-based method for AFP, which makes the most of both protein sequence and high-order protein network information. Our multispecies strategy allows one single model to be trained for all species, indicating a larger number of training samples than existing methods. Extensive experiments with a large-scale dataset show that DeepGraphGO outperforms a number of competing state-of-the-art methods significantly, including DeepGOPlus and three representative network-based methods: GeneMANIA, deepNF and clusDCA. We further confirm the effectiveness of our multispecies strategy and the advantage of DeepGraphGO over so-called difficult proteins. Finally, we integrate DeepGraphGO into the state-of-the-art ensemble method, NetGO, as a component and achieve a further performance improvement. AVAILABILITY AND IMPLEMENTATION: https://github.com/yourh/DeepGraphGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronghui You, Shuwei Yao, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 1 |
| 2020 | FullMeSH: improving large-scale MeSH indexing with full textabstractMOTIVATION: With the rapidly growing biomedical literature, automatically indexing biomedical articles by Medical Subject Heading (MeSH), namely MeSH indexing, has become increasingly important for facilitating hypothesis generation and knowledge discovery. Over the past years, many large-scale MeSH indexing approaches have been proposed, such as Medical Text Indexer, MeSHLabeler, DeepMeSH and MeSHProbeNet. However, the performance of these methods is hampered by using limited information, i.e. only the title and abstract of biomedical articles. RESULTS: We propose FullMeSH, a large-scale MeSH indexing method taking advantage of the recent increase in the availability of full text articles. Compared to DeepMeSH and other state-of-the-art methods, FullMeSH has three novelties: (i) Instead of using a full text as a whole, FullMeSH segments it into several sections with their normalized titles in order to distinguish their contributions to the overall performance. (ii) FullMeSH integrates the evidence from different sections in a 'learning to rank' framework by combining the sparse and deep semantic representations. (iii) FullMeSH trains an Attention-based Convolutional Neural Network for each section, which achieves better performance on infrequent MeSH headings. FullMeSH has been developed and empirically trained on the entire set of 1.4 million full-text articles in the PubMed Central Open Access subset. It achieved a Micro F-measure of 66.76% on a test set of 10 000 articles, which was 3.3% and 6.4% higher than DeepMeSH and MeSHLabeler, respectively. Furthermore, FullMeSH demonstrated an average improvement of 4.7% over DeepMeSH for indexing Check Tags, a set of most frequently indexed MeSH headings. AVAILABILITY AND IMPLEMENTATION: The software is available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Suyang Dai, Ronghui You, Zhiyong Lu, Xiaodi Huang 0001, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 2 |
| 2019 | DeepDock: Enhancing Ligand-protein Interaction Prediction by a Combination of Ligand and Structure InformationabstractThe prediction of precise protein-ligand binding activities can accelerate drug discovery by virtual screening-a computational technique that predicts whether a small molecule ligand is able to bind to a specific target. Thus, it is crucial to improve the performance of virtual screening. However, previous models for solving this problem are either ligand-based or structure-based. In this paper, we propose a universal deep neural network model called DeepDock that predicts protein-ligand interaction by using both ligand and structure information. Using the combination of two types of information, our model consists of embedding, convolution, max pooling, and fully-connected layers. In particular, different types of inputs are concatenated before being fed into the fully-connected layers. In the experiments, we compare our approach to the competing methods against two benchmark datasets under different settings. The experiment results have demonstrated that DeepDock can improve predictive performance by more than 4% on both DUD-E and MUV datasets in terms of AUPR. Zhirui Liao, Ronghui You, Xiaodi Huang 0001, Shanfeng Zhu |
BIBM | 2 |
| 2019 | AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text ClassificationabstractExtreme multi-label text classification (XMTC) is an important problem in the era of {\it big data}, for tagging a given text with the most relevant multiple labels from an extremely large-scale label set. XMTC can be found in many applications, such as item categorization, web page tagging, and news annotation. Traditionally most methods used bag-of-words (BOW) as inputs, ignoring word context as well as deep semantic information. Recent attempts to overcome the problems of BOW by deep learning still suffer from 1) failing to capture the important subtext for each label and 2) lack of scalability against the huge number of labels. We propose a new label tree-based deep learning model for XMTC, called AttentionXML, with two unique features: 1) a multi-label attention mechanism with raw text as input, which allows to capture the most relevant part of text to each label; and 2) a shallow and wide probabilistic label tree (PLT), which allows to handle millions of labels, especially for "tail labels". We empirically compared the performance of AttentionXML with those of eight state-of-the-art methods over six benchmark datasets, including Amazon-3M with around 3 million labels. AttentionXML outperformed all competing methods under all experimental settings. Experimental results also show that AttentionXML achieved the best performance against tail labels among label tree-based methods. The code and datasets are available at \url{http://github.com/yourh/AttentionXML} . Ronghui You, Suyang Dai, Hiroshi Mamitsuka, Shanfeng Zhu |
NeurIPS | 1 |
| 2018 | GOLabeler: improving sequence-based large-scale protein function prediction by learning to rankabstractMotivation: Gene Ontology (GO) has been widely used to annotate functions of proteins and understand their biological roles. Currently only <1% of >70 million proteins in UniProtKB have experimental GO annotations, implying the strong necessity of automated function prediction (AFP) of proteins, where AFP is a hard multilabel classification problem due to one protein with a diverse number of GO terms. Most of these proteins have only sequences as input information, indicating the importance of sequence-based AFP (SAFP: sequences are the only input). Furthermore, homology-based SAFP tools are competitive in AFP competitions, while they do not necessarily work well for so-called difficult proteins, which have <60% sequence identity to proteins with annotations already. Thus, the vital and challenging problem now is how to develop a method for SAFP, particularly for difficult proteins. Methods: The key of this method is to extract not only homology information but also diverse, deep-rooted information/evidence from sequence inputs and integrate them into a predictor in a both effective and efficient manner. We propose GOLabeler, which integrates five component classifiers, trained from different features, including GO term frequency, sequence alignment, amino acid trigram, domains and motifs, and biophysical properties, etc., in the framework of learning to rank (LTR), a paradigm of machine learning, especially powerful for multilabel classification. Results: The empirical results obtained by examining GOLabeler extensively and thoroughly by using large-scale datasets revealed numerous favorable aspects of GOLabeler, including significant performance advantage over state-of-the-art AFP methods. Availability and implementation: http://datamining-iip.fudan.edu.cn/golabeler. Supplementary information: Supplementary data are available at Bioinformatics online. Ronghui You, Yi Xiong 0002, Fengzhu Sun, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 1 |
| 2017 | DeepText2Go: Improving large-scale protein function prediction with deep semantic text representationabstractUniProtKB has collected more than 88 million protein sequences by July 2017. Less than 0.2% of these proteins, however, have added experimental GO annotations. To reduce this huge gap, automatic protein function prediction (AFP) becomes increasingly important. Results on CAFA (the Critical Assessment of protein Function Annotation algorithms) benchmark demonstrates that sequence homology based methods are highly competitive in AFP. One imperative issues will be incorporating other information sources other than sequence for AFP. In contrast to using BOW (bag of words) representation in traditional text-based AFP, we proposed a new method called DeepText2GO to improve large-scale AFP by using deep semantic text representation instead. Furthermore, DeepText2GO integrates both text-based and sequence homology-based methods through a consensus approach. Extensive experiments on the benchmark dataset extracted from UniProt/SwissProt have demonstrated that DeepText2GO significantly outperformed both text-based and sequence homology-based methods, validating its superiority. Ronghui You, Shanfeng Zhu |
BIBM | 1 |
| 2017 | POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time OptimalityabstractIt's important to hit a space-time balance for a real-world algorithm to achieve high performance on modern shared-memory multi-core or many-core systems. However, a large class of dynamic programs with more than $O(1)$ dependency achieve optimality either in space or time, but not both. In the literature, the problem is known as the fundamental space-time tradeoff. By exploiting properly on the runtime system, we show that our STAR (Space-Time Adaptive and Reductive) technique can help these dynamic programs to achieve sublinear parallel time bounds while still maintaining work-, space-, and cache-optimality in a processor- and cache-oblivious fashion. Ronghui You |
PPoPP | 2 |
| 2016 | DeepMeSH: deep semantic representation for improving large-scale MeSH indexingabstractMOTIVATION: Medical Subject Headings (MeSH) indexing, which is to assign a set of MeSH main headings to citations, is crucial for many important tasks in biomedical text mining and information retrieval. Large-scale MeSH indexing has two challenging aspects: the citation side and MeSH side. For the citation side, all existing methods, including Medical Text Indexer (MTI) by National Library of Medicine and the state-of-the-art method, MeSHLabeler, deal with text by bag-of-words, which cannot capture semantic and context-dependent information well. METHODS: We propose DeepMeSH that incorporates deep semantic information for large-scale MeSH indexing. It addresses the two challenges in both citation and MeSH sides. The citation side challenge is solved by a new deep semantic representation, D2V-TFIDF, which concatenates both sparse and dense semantic representations. The MeSH side challenge is solved by using the 'learning to rank' framework of MeSHLabeler, which integrates various types of evidence generated from the new semantic representation. RESULTS: DeepMeSH achieved a Micro F-measure of 0.6323, 2% higher than 0.6218 of MeSHLabeler and 12% higher than 0.5637 of MTI, for BioASQ3 challenge data with 6000 citations. AVAILABILITY AND IMPLEMENTATION: The software is available upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shengwen Peng, Ronghui You, Hongning Wang, ChengXiang Zhai, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 2 |
| 2015 | Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiencyabstractState-of-the-art cache-oblivious parallel algorithms for dynamic programming (DP) problems usually guarantee asymptotically optimal cache performance without any tuning of cache parameters, but they often fail to exploit the theoretically best parallelism at the same time. While these algorithms achieve cache-optimality through the use of a recursive divide-and-conquer (DAC) strategy, scheduling tasks at the granularity of task dependency introduces artificial dependencies in addition to those arising from the defining recurrence equations. We removed the artificial dependency by scheduling tasks ready for execution as soon as all its real dependency constraints are satisfied, while preserving the cache-optimality by inheriting the DAC strategy. We applied our approach to a set of widely known dynamic programming problems, such as Floyd-Warshall's All-Pairs Shortest Paths, Stencil, and LCS. Theoretical analyses show that our techniques improve the span of 2-way DAC-based Floyd Warshall's algorithm on an $n$ node graph from $Thn^2n$ to $Thn$, stencil computations on a $d$-dimensional hypercubic grid of width $w$ for $h$ time steps from $Th(d^2 h) w^ (d+2) - 1$ to $Thh$, and LCS on two sequences of length $n$ each from $Thn^_2 3$ to $Thn$. In each case, the total work and cache complexity remain asymptotically optimal. Experimental measurements exhibit a $3$ - $5$ times improvement in absolute running time, $10$ - $20$ times improvement in burdened span by Cilkview, and approximately the same L1/L2 cache misses by PAPI. Ronghui You, Haibin Kan, Jesmin Jahan Tithi, Pramod Ganapathi, Rezaul Alam Chowdhury |
PPoPP | 2 |