Raunak Shrestha

dblp:142/5059 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0002-1144-1413ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Mathematical optimization · 50% Algorithms and data structures · 50%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › cancer genomics
tumor evolution
0.412020
Identification of conserved evolutionary trajectories in tumors · Bioinform. 2020
Mathematical optimization
combinatorial optimization
0.412020
Identification of conserved evolutionary trajectories in tumors · Bioinform. 2020
Algorithms and data structures › computational biology
consensus trees
0.412020
Identification of conserved evolutionary trajectories in tumors · Bioinform. 2020
Data mining › dimensionality reduction
feature selection
0.312018
Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data · IEEE Trans. Knowl. Data Eng. 2018
Data mining › dimensionality reduction › feature selection
high-dimensional feature selection
0.312018
Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data · IEEE Trans. Knowl. Data Eng. 2018
Bioinformatics and computational biology
cancer genomics
0.212014
HIT'nDRIVE: Multi-driver Gene Prioritization Based on Hitting Time · RECOMB 2014
Bioinformatics and computational biology
phenotype classification
0.112018
Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data · IEEE Trans. Knowl. Data Eng. 2018

Methods — techniques the papers use, named apart from their topics

phylogenetic inference · 0.9combinatorial optimization · 0.9hilbert-schmidt independence criterion · 0.7HSIC Lasso · 0.7network analysis · 0.2hitting time · 0.2
YearPublicationVenuePosition
2020 Identification of conserved evolutionary trajectories in tumors
abstract
MOTIVATION: As multi-region, time-series and single-cell sequencing data become more widely available; it is becoming clear that certain tumors share evolutionary characteristics with others. In the last few years, several computational methods have been developed with the goal of inferring the subclonal composition and evolutionary history of tumors from tumor biopsy sequencing data. However, the phylogenetic trees that they report differ significantly between tumors (even those with similar characteristics). RESULTS: In this article, we present a novel combinatorial optimization method, CONETT, for detection of recurrent tumor evolution trajectories. Our method constructs a consensus tree of conserved evolutionary trajectories based on the information about temporal order of alteration events in a set of tumors. We apply our method to previously published datasets of 100 clear-cell renal cell carcinoma and 99 non-small-cell lung cancer patients and identify both conserved trajectories that were reported in the original studies, as well as new trajectories. AVAILABILITY AND IMPLEMENTATION: CONETT is implemented in C++ and available at https://github.com/ehodzic/CONETT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ermin Hodzic, Raunak Shrestha, Salem Malikic, Colin C. Collins, Kevin Litchfield, Samra Turajlic, Süleyman Cenk Sahinalp
Bioinform.2
2018 Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data
abstract
Machine learning methods are used to discover complex nonlinear relationships in biological and medical data. However, sophisticated learning models are computationally unfeasible for data with millions of features. Here, we introduce the first feature selection method for nonlinear learning problems that can scale up to large, ultra-high dimensional biological data. More specifically, we scale up the novel Hilbert-Schmidt Independence Criterion Lasso (HSIC Lasso) to handle millions of features with tens of thousand samples. The proposed method is guaranteed to find an optimal subset of maximally predictive features with minimal redundancy, yielding higher predictive power and improved interpretability. Its effectiveness is demonstrated through applications to classify phenotypes based on module expression in human prostate cancer patients and to detect enzymes among protein structures. We achieve high accuracy with as few as 20 out of one million features-a dimensionality reduction of 99.998 percent. Our algorithm can be implemented on commodity cloud computing platforms. The dramatic reduction of features may lead to the ubiquitous deployment of sophisticated prediction models in mobile health care applications.
Makoto Yamada, Jiliang Tang, Jose Lugo-Martinez, Ermin Hodzic, Raunak Shrestha, Avishek Saha, Hua Ouyang, Dawei Yin 0001, Hiroshi Mamitsuka, Süleyman Cenk Sahinalp, Predrag Radivojac, Filippo Menczer, Yi Chang 0001
IEEE Trans. Knowl. Data Eng.5
2014 HIT'nDRIVE: Multi-driver Gene Prioritization Based on Hitting Time
Raunak Shrestha, Ermin Hodzic, Jake Yeung, Kendric Wang, Thomas Sauerwald, Phuong Dao, Shawn Anderson, Himisha Beltran, Mark A. Rubin, Colin C. Collins, Gholamreza Haffari, Süleyman Cenk Sahinalp
RECOMB1