Jeffrey Yuan

dblp:32/2132 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0003-3855-1994ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Recommender systems · 73% Information retrieval · 27%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.622018
Assembly of Long Error-Prone Reads Using Repeat Graphs · RECOMB 2018
Assembly of Long Error-Prone Reads Using de Bruijn Graphs · RECOMB 2016
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
long-read assembly
0.622018
Assembly of Long Error-Prone Reads Using Repeat Graphs · RECOMB 2018
Assembly of Long Error-Prone Reads Using de Bruijn Graphs · RECOMB 2016
Recommender systems › collaborative filtering
matrix factorization
0.532013
Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013
Focused matrix factorization for audience selection in display advertising · ICDE 2013
Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior · Proc. VLDB Endow. 2012
Recommender systems
collaborative filtering
0.322013
Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013
Focused matrix factorization for audience selection in display advertising · ICDE 2013
Graph algorithms and graph theory
graph algorithms
0.312018
Assembly of Long Error-Prone Reads Using Repeat Graphs · RECOMB 2018
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly › de novo assembly
de bruijn graph assembly
0.212016
Assembly of Long Error-Prone Reads Using de Bruijn Graphs · RECOMB 2016
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical bayesian model
0.212013
Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013
Recommender systems › content-based recommendation
attribute-based recommendation
0.212013
Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013
Information retrieval › online advertising
display advertising
0.212013
Focused matrix factorization for audience selection in display advertising · ICDE 2013
Recommender systems › user modeling
hierarchical preference modeling
0.212013
Latent factor models with additive and hierarchically-smoothed user preferences · WSDM 2013
Recommender systems
data sparsity and cold-start
0.112012
Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior · Proc. VLDB Endow. 2012
Information retrieval › online advertising › sponsored search
query-ad matching
0.112009
Online expansion of rare queries for sponsored search · WWW 2009
Information retrieval › online advertising
sponsored search
0.112009
Online expansion of rare queries for sponsored search · WWW 2009
Bioinformatics and computational biology
sequence analysis
0.122004
Enhanced homology searching through genome reading frame predetermination · Bioinform. 2004
MULTICLUSTAL: a systematic method for surveying Clustal W alignment parameters · Bioinform. 1999
Parallel and multicore computing › parallel computing
parallel implementation
0.012013
Focused matrix factorization for audience selection in display advertising · ICDE 2013
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.012004
Enhanced homology searching through genome reading frame predetermination · Bioinform. 2004
Recommender systems › sequential recommendation
temporal dynamics modeling
0.012012
Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior · Proc. VLDB Endow. 2012
Information retrieval
query understanding
0.012009
Online expansion of rare queries for sponsored search · WWW 2009
Bioinformatics and computational biology
multiple sequence alignment
0.011999
MULTICLUSTAL: a systematic method for surveying Clustal W alignment parameters · Bioinform. 1999

Methods — techniques the papers use, named apart from their topics

repeat graphs · 0.7error-prone read assembly · 0.7product taxonomy · 0.3matrix factorization · 0.3forward-filtering backward-smoothing · 0.3collaborative filtering · 0.3BPR ranking · 0.3de bruijn graph · 0.2parallel multi-core implementation · 0.1additive model · 0.1online query expansion · 0.1receiver operating characteristic analysis · 0.0FASTA · 0.0BLAST · 0.0parameter search automation · 0.0
YearPublicationVenuePosition
2018 Assembly of Long Error-Prone Reads Using Repeat Graphs
Mikhail Kolmogorov, Jeffrey Yuan, Yu Lin 0001, Pavel A. Pevzner
RECOMB2
2016 Assembly of Long Error-Prone Reads Using de Bruijn Graphs
Yu Lin 0001, Max W. Shen, Jeffrey Yuan, Mark Chaisson, Pavel A. Pevzner
RECOMB3
2013 Focused matrix factorization for audience selection in display advertising
abstract
Audience selection is a key problem in display advertising systems in which we need to select a list of users who are interested (i.e., most likely to buy) in an advertising campaign. The users' past feedback on this campaign can be leveraged to construct such a list using collaborative filtering techniques such as matrix factorization. However, the user-campaign interaction is typically extremely sparse, hence the conventional matrix factorization does not perform well. Moreover, simply combining the users feedback from all campaigns does not address this since it dilutes the focus on target campaign in consideration. To resolve these issues, we propose a novel focused matrix factorization model (FMF) which learns users' preferences towards the specific campaign products, while also exploiting the information about related products. We exploit the product taxonomy to discover related campaigns, and design models to discriminate between the users' interest towards campaign products and non-campaign products. We develop a parallel multi-core implementation of the FMF model and evaluate its performance over a real-world advertising dataset spanning more than a million products. Our experiments demonstrate the benefits of using our models over existing approaches.
Bhargav Kanagal, Amr Ahmed 0001, Sandeep Pandey, Vanja Josifovski, Lluís Garcia Pueyo, Jeffrey Yuan
ICDE6
2013 Latent factor models with additive and hierarchically-smoothed user preferences
abstract
Items in recommender systems are usually associated with annotated attributes: for e.g., brand and price for products; agency for news articles, etc. Such attributes are highly informative and must be exploited for accurate recommendation. While learning a user preference model over these attributes can result in an interpretable recommender system and can hands the cold start problem, it suffers from two major drawbacks: data sparsity and the inability to model random effects. On the other hand, latent-factor collaborative filtering models have shown great promise in recommender systems; however, its performance on rare items is poor. In this paper we propose a novel model LFUM, which provides the advantages of both of the above models. We learn user preferences (over the attributes) using a personalized Bayesian hierarchical model that uses a combination(additive model) of a globally learned preference model along with user-specific preferences. To combat data-sparsity, we smooth these preferences over the item-taxonomy using an efficient forward-filtering and backward-smoothing inference algorithm. Our inference algorithms can handle both discrete attributes (e.g., item brands) and continuous attributes (e.g., item prices). We combine the user preferences with the latent-factor models and train the resulting collaborative filtering system end-to-end using the successful BPR ranking algorithm. In our extensive experimental analysis, we show that our proposed model outperforms several commonly used baselines and we carry out an ablation study showing the benefits of each component of our model.
Amr Ahmed 0001, Bhargav Kanagal, Sandeep Pandey, Vanja Josifovski, Lluís Garcia Pueyo, Jeffrey Yuan
WSDM6
2012 Supercharging Recommender Systems using Taxonomies for Learning User Purchase Behavior
abstract
Recommender systems based on latent factor models have been effectively used for understanding user interests and predicting future actions. Such models work by projecting the users and items into a smaller dimensional space, thereby clustering similar users and items together and subsequently compute similarity between unknown user-item pairs. When user-item interactions are sparse ( sparsity problem) or when new items continuously appear ( cold start problem), these models perform poorly. In this paper, we exploit the combination of taxonomies and latent factor models to mitigate these issues and improve recommendation accuracy. We observe that taxonomies provide structure similar to that of a latent factor model: namely, it imposes human-labeled categories (clusters) over items. This leads to our proposed taxonomy-aware latent factor model (TF) which combines taxonomies and latent factors using additive models. We develop efficient algorithms to train the TF models, which scales to large number of users/items and develop scalable inference/recommendation algorithms by exploiting the structure of the taxonomy. In addition, we extend the TF model to account for the temporal dynamics of user interests using high-order Markov chains . To deal with large-scale data, we develop a parallel multi-core implementation of our TF model. We empirically evaluate the TF model for the task of predicting user purchases using a real-world shopping dataset spanning more than a million users and products. Our experiments demonstrate the benefits of using our TF models over existing approaches, in terms of both prediction accuracy and running time.
Bhargav Kanagal, Amr Ahmed 0001, Sandeep Pandey, Vanja Josifovski, Jeffrey Yuan, Lluís Garcia Pueyo
Proc. VLDB Endow.5
2009 Online expansion of rare queries for sponsored search
abstract
Sponsored search systems are tasked with matching queries
Andrei Z. Broder, Peter Ciccolo, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler, Lance Riedel, Jeffrey Yuan
WWW7
2004 Enhanced homology searching through genome reading frame predetermination
abstract
MOTIVATION: Many bioinformatic approaches exist for finding novel genes within genomic sequence data. Traditionally, homology search-based methods are often the first approach employed in determining whether a novel gene exists that is similar to a known gene. Unfortunately, distantly related genes or motifs often are difficult to find using single query-based homology search algorithms against large sequence datasets such as the human genome. Therefore, the motivation behind this work was to develop an approach to enhance the sensitivity of traditional single query-based homology algorithms against genomic data without losing search selectivity. RESULTS: We demonstrate that by searching against a genome fragmented into all possible reading frames, the sensitivity of homology-based searches is enhanced without degrading its selectivity. Using the ETS-domain, bromodomain and acetyl-CoA acetyltransferase gene as queries, we were able to demonstrate that direct protein-protein searches using BLAST2P or FASTA3 against a human genome segmented among all possible reading frames and translated was substantially more sensitive than traditional protein-DNA searches against a raw genomic sequence using an application such as TBLAST2N. Receiver operating characteristic analysis was employed to demonstrate that the algorithms remained selective, while comparisons of the algorithms showed that the protein-protein searches were more sensitive in identifying hits. Therefore, through the overprediction of reading frames by this method and the increased sensitivity of protein-protein based homology search algorithms, a genome can be deeply mined, potentially finding hits overlooked by protein-DNA searches against raw genomic data.
Jeffrey Yuan, Bruce Bush, Alex Elbrecht, Theresa Zhang, Richard Blevins
Bioinform.1
1999 MULTICLUSTAL: a systematic method for surveying Clustal W alignment parameters
abstract
MULTICLUSTAL is a Perl script designed to automate the process of alignment parameter choice for Clustal W with the goal of generating high quality multiple sequence alignments.
Jeffrey Yuan, Angela Amend, Joseph A. Borkowski, Reynold DelMarco, Wendy Bailey, Guochun Xie, Richard Blevins
Bioinform.1