Melih Yilmaz

dblp:283/8331 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
proteomics
1.322024
A learned score function improves the power of mass spectrometry database search · Bioinform. 2024
De novo mass spectrometry peptide sequencing with a transformer model · ICML 2022
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.812024
A learned score function improves the power of mass spectrometry database search · Bioinform. 2024
Bioinformatics and computational biology › proteomics
peptide identification
0.812024
A learned score function improves the power of mass spectrometry database search · Bioinform. 2024
Bioinformatics and computational biology › proteomics
peptide-spectrum matching
0.812024
A learned score function improves the power of mass spectrometry database search · Bioinform. 2024
Bioinformatics and computational biology › proteomics › peptide sequencing
de novo peptide sequencing
0.612022
De novo mass spectrometry peptide sequencing with a transformer model · ICML 2022
Machine learning › Deep learning architectures and training
transformer
0.212022
De novo mass spectrometry peptide sequencing with a transformer model · ICML 2022

Methods — techniques the papers use, named apart from their topics

transformer · 1.1percolator · 0.8machine learning · 0.8de novo peptide sequencing · 0.8
YearPublicationVenuePosition
2024 A learned score function improves the power of mass spectrometry database search
abstract
MOTIVATION: One of the core problems in the analysis of protein tandem mass spectrometry data is the peptide assignment problem: determining, for each observed spectrum, the peptide sequence that was responsible for generating the spectrum. Two primary classes of methods are used to solve this problem: database search and de novo peptide sequencing. State-of-the-art methods for de novo sequencing use machine learning methods, whereas most database search engines use hand-designed score functions to evaluate the quality of a match between an observed spectrum and a candidate peptide from the database. We hypothesized that machine learning models for de novo sequencing implicitly learn a score function that captures the relationship between peptides and spectra, and thus may be re-purposed as a score function for database search. Because this score function is trained from massive amounts of mass spectrometry data, it could potentially outperform existing, hand-designed database search tools. RESULTS: To test this hypothesis, we re-engineered Casanovo, which has been shown to provide state-of-the-art de novo sequencing capabilities, to assign scores to given peptide-spectrum pairs. We then evaluated the statistical power of this Casanovo score function, Casanovo-DB, to detect peptides on a benchmark of three mass spectrometry runs from three different species. In addition, we show that re-scoring with the Percolator post-processor benefits Casanovo-DB more than other score functions, further increasing the number of detected peptides.
Varun Ananth, Justin Sanders, Melih Yilmaz, Sewoong Oh, William Stafford Noble
Bioinform.3
2022 De novo mass spectrometry peptide sequencing with a transformer model
abstract
Tandem mass spectrometry is the only high-throughput method for analyzing the protein content of complex biological samples and is thus the primary technology driving the growth of the field of proteomics. A key outstanding challenge in this field involves identifying the sequence of amino acids -the peptide- responsible for generating each observed spectrum, without making use of prior knowledge in the form of a peptide sequence database. Although various machine learning methods have been developed to address this de novo sequencing problem, challenges that arise when modeling tandem mass spectra have led to complex models that combine multiple neural networks and post-processing steps. We propose a simple yet powerful method for de novo peptide sequencing, Casanovo, that uses a transformer framework to map directly from a sequence of observed peaks (a mass spectrum) to a sequence of amino acids (a peptide). Our experiments show that Casanovo achieves state-of-the-art performance on a benchmark dataset using a standard cross-species evaluation framework which involves testing with spectra with never-before-seen peptide labels. Casanovo not only achieves superior performance but does so at a fraction of the model complexity and inference time required by other methods.
Melih Yilmaz, William Fondrie, Wout Bittremieux, Sewoong Oh, William Stafford Noble
ICML1
2020 Distinct Clusters of Patient reported outcome (PRO) trajectories among oncology patients receiving chemotherapy
Selen Bozkurt, Amee D. Azad, Melih Yilmaz, James D. Brooks, Douglas W. Blayney, Tina Hernandez-Boussard
AMIA3