VLDB 2026 Research / reviewers in the wild / expert
Jeffrey Yuan
dblp:32/2132
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0003-3855-1994ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Recommender systems · 73% Information retrieval · 27% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.6 | 2 | 2018 | Assembly of Long Error-Prone Reads Using Repeat Graphs · RECOMB 2018 Assembly of Long Error-Prone Reads Using de Bruijn Graphs · RECOMB 2016 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
long-read assembly |
0.6 | 2 | 2018 | Assembly of Long Error-Prone Reads Using Repeat Graphs · RECOMB 2018 Assembly of Long Error-Prone Reads Using de Bruijn Graphs · RECOMB 2016 |
Recommender systems › collaborative filtering
matrix factorization |
0.5 | 3 | 2013 | Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013 Focused matrix factorization for audience selection in display advertising · ICDE 2013 Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior · Proc. VLDB Endow. 2012 |
Recommender systems
collaborative filtering |
0.3 | 2 | 2013 | Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013 Focused matrix factorization for audience selection in display advertising · ICDE 2013 |
Graph algorithms and graph theory
graph algorithms |
0.3 | 1 | 2018 | Assembly of Long Error-Prone Reads Using Repeat Graphs · RECOMB 2018 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly › de novo assembly
de bruijn graph assembly |
0.2 | 1 | 2016 | Assembly of Long Error-Prone Reads Using de Bruijn Graphs · RECOMB 2016 |
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical bayesian model |
0.2 | 1 | 2013 | Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013 |
Recommender systems › content-based recommendation
attribute-based recommendation |
0.2 | 1 | 2013 | Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013 |
Information retrieval › online advertising
display advertising |
0.2 | 1 | 2013 | Focused matrix factorization for audience selection in display advertising · ICDE 2013 |
Recommender systems › user modeling
hierarchical preference modeling |
0.2 | 1 | 2013 | Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013 |
Recommender systems
data sparsity and cold-start |
0.1 | 1 | 2012 | Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior · Proc. VLDB Endow. 2012 |
Information retrieval › online advertising › sponsored search
query-ad matching |
0.1 | 1 | 2009 | Online expansion of rare queries for sponsored search · WWW 2009 |
Information retrieval › online advertising
sponsored search |
0.1 | 1 | 2009 | Online expansion of rare queries for sponsored search · WWW 2009 |
Bioinformatics and computational biology
sequence analysis |
0.1 | 2 | 2004 | Enhanced homology searching through genome reading frame predetermination · Bioinform. 2004 MULTICLUSTAL: a systematic method for surveying Clustal W alignment parameters · Bioinform. 1999 |
Parallel and multicore computing › parallel computing
parallel implementation |
0.0 | 1 | 2013 | Focused matrix factorization for audience selection in display advertising · ICDE 2013 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.0 | 1 | 2004 | Enhanced homology searching through genome reading frame predetermination · Bioinform. 2004 |
Recommender systems › sequential recommendation
temporal dynamics modeling |
0.0 | 1 | 2012 | Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior · Proc. VLDB Endow. 2012 |
Information retrieval
query understanding |
0.0 | 1 | 2009 | Online expansion of rare queries for sponsored search · WWW 2009 |
Bioinformatics and computational biology
multiple sequence alignment |
0.0 | 1 | 1999 | MULTICLUSTAL: a systematic method for surveying Clustal W alignment parameters · Bioinform. 1999 |
Methods — techniques the papers use, named apart from their topics
repeat graphs · 0.7error-prone read assembly · 0.7product taxonomy · 0.3matrix factorization · 0.3forward-filtering backward-smoothing · 0.3collaborative filtering · 0.3BPR ranking · 0.3de bruijn graph · 0.2parallel multi-core implementation · 0.1additive model · 0.1online query expansion · 0.1receiver operating characteristic analysis · 0.0FASTA · 0.0BLAST · 0.0parameter search automation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Assembly of Long Error-Prone Reads Using Repeat Graphs
Mikhail Kolmogorov, Jeffrey Yuan, Yu Lin 0001, Pavel A. Pevzner |
RECOMB | 2 |
| 2016 | Assembly of Long Error-Prone Reads Using de Bruijn Graphs
Yu Lin 0001, Max W. Shen, Jeffrey Yuan, Mark Chaisson, Pavel A. Pevzner |
RECOMB | 3 |
| 2013 | Focused matrix factorization for audience selection in display advertisingabstractAudience selection is a key problem in display advertising systems in which we need to select a list of users who are interested (i.e., most likely to buy) in an advertising campaign. The users' past feedback on this campaign can be leveraged to construct such a list using collaborative filtering techniques such as matrix factorization. However, the user-campaign interaction is typically extremely sparse, hence the conventional matrix factorization does not perform well. Moreover, simply combining the users feedback from all campaigns does not address this since it dilutes the focus on target campaign in consideration. To resolve these issues, we propose a novel focused matrix factorization model (FMF) which learns users' preferences towards the specific campaign products, while also exploiting the information about related products. We exploit the product taxonomy to discover related campaigns, and design models to discriminate between the users' interest towards campaign products and non-campaign products. We develop a parallel multi-core implementation of the FMF model and evaluate its performance over a real-world advertising dataset spanning more than a million products. Our experiments demonstrate the benefits of using our models over existing approaches. Bhargav Kanagal, Amr Ahmed 0001, Sandeep Pandey, Vanja Josifovski, Lluís Garcia Pueyo, Jeffrey Yuan |
ICDE | 6 |
| 2013 | Latent factor models with additive and hierarchically-smoothed user preferencesabstractItems in recommender systems are usually associated with annotated attributes: for e.g., brand and price for products; agency for news articles, etc. Such attributes are highly informative and must be exploited for accurate recommendation. While learning a user preference model over these attributes can result in an interpretable recommender system and can hands the cold start problem, it suffers from two major drawbacks: data sparsity and the inability to model random effects. On the other hand, latent-factor collaborative filtering models have shown great promise in recommender systems; however, its performance on rare items is poor. In this paper we propose a novel model LFUM, which provides the advantages of both of the above models. We learn user preferences (over the attributes) using a personalized Bayesian hierarchical model that uses a combination(additive model) of a globally learned preference model along with user-specific preferences. To combat data-sparsity, we smooth these preferences over the item-taxonomy using an efficient forward-filtering and backward-smoothing inference algorithm. Our inference algorithms can handle both discrete attributes (e.g., item brands) and continuous attributes (e.g., item prices). We combine the user preferences with the latent-factor models and train the resulting collaborative filtering system end-to-end using the successful BPR ranking algorithm. In our extensive experimental analysis, we show that our proposed model outperforms several commonly used baselines and we carry out an ablation study showing the benefits of each component of our model. Amr Ahmed 0001, Bhargav Kanagal, Sandeep Pandey, Vanja Josifovski, Lluís Garcia Pueyo, Jeffrey Yuan |
WSDM | 6 |
| 2012 | Supercharging Recommender Systems using Taxonomies for Learning User Purchase BehaviorabstractRecommender systems based on latent factor models have been effectively used for understanding user interests and predicting future actions. Such models work by projecting the users and items into a smaller dimensional space, thereby clustering similar users and items together and subsequently compute similarity between unknown user-item pairs. When user-item interactions are sparse ( sparsity problem) or when new items continuously appear ( cold start problem), these models perform poorly. In this paper, we exploit the combination of taxonomies and latent factor models to mitigate these issues and improve recommendation accuracy. We observe that taxonomies provide structure similar to that of a latent factor model: namely, it imposes human-labeled categories (clusters) over items. This leads to our proposed taxonomy-aware latent factor model (TF) which combines taxonomies and latent factors using additive models. We develop efficient algorithms to train the TF models, which scales to large number of users/items and develop scalable inference/recommendation algorithms by exploiting the structure of the taxonomy. In addition, we extend the TF model to account for the temporal dynamics of user interests using high-order Markov chains . To deal with large-scale data, we develop a parallel multi-core implementation of our TF model. We empirically evaluate the TF model for the task of predicting user purchases using a real-world shopping dataset spanning more than a million users and products. Our experiments demonstrate the benefits of using our TF models over existing approaches, in terms of both prediction accuracy and running time. Bhargav Kanagal, Amr Ahmed 0001, Sandeep Pandey, Vanja Josifovski, Jeffrey Yuan, Lluís Garcia Pueyo |
Proc. VLDB Endow. | 5 |
| 2009 | Online expansion of rare queries for sponsored searchabstractSponsored search systems are tasked with matching queries Andrei Z. Broder, Peter Ciccolo, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler, Lance Riedel, Jeffrey Yuan |
WWW | 7 |
| 2004 | Enhanced homology searching through genome reading frame predeterminationabstractMOTIVATION: Many bioinformatic approaches exist for finding novel genes within genomic sequence data. Traditionally, homology search-based methods are often the first approach employed in determining whether a novel gene exists that is similar to a known gene. Unfortunately, distantly related genes or motifs often are difficult to find using single query-based homology search algorithms against large sequence datasets such as the human genome. Therefore, the motivation behind this work was to develop an approach to enhance the sensitivity of traditional single query-based homology algorithms against genomic data without losing search selectivity. RESULTS: We demonstrate that by searching against a genome fragmented into all possible reading frames, the sensitivity of homology-based searches is enhanced without degrading its selectivity. Using the ETS-domain, bromodomain and acetyl-CoA acetyltransferase gene as queries, we were able to demonstrate that direct protein-protein searches using BLAST2P or FASTA3 against a human genome segmented among all possible reading frames and translated was substantially more sensitive than traditional protein-DNA searches against a raw genomic sequence using an application such as TBLAST2N. Receiver operating characteristic analysis was employed to demonstrate that the algorithms remained selective, while comparisons of the algorithms showed that the protein-protein searches were more sensitive in identifying hits. Therefore, through the overprediction of reading frames by this method and the increased sensitivity of protein-protein based homology search algorithms, a genome can be deeply mined, potentially finding hits overlooked by protein-DNA searches against raw genomic data. Jeffrey Yuan, Bruce Bush, Alex Elbrecht, Theresa Zhang, Richard Blevins |
Bioinform. | 1 |
| 1999 | MULTICLUSTAL: a systematic method for surveying Clustal W alignment parametersabstractMULTICLUSTAL is a Perl script designed to automate the process of alignment parameter choice for Clustal W with the goal of generating high quality multiple sequence alignments. Jeffrey Yuan, Angela Amend, Joseph A. Borkowski, Reynold DelMarco, Wendy Bailey, Guochun Xie, Richard Blevins |
Bioinform. | 1 |