VLDB 2026 Research / reviewers in the wild / expert
Wen-Yun Yang
dblp:53/864
· DBLP profile ↗
17ranked-venue papers
7as first author
6since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-authorArtificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Differentiable Retrieval Augmentation via Generative Language Modeling for E-commerce Query Intent ClassificationabstractRetrieval augmentation, which enhances downstream models by a knowledge retriever and an external corpus instead of by merely increasing the number of model parameters, has been successfully applied to many natural language processing(NLP) tasks such as text classification, question answering and so on. However, existing methods that separately or asynchronously train the retriever and downstream model mainly due to the non-differentiability between the two parts, usually lead to degraded performance compared to end-to-end joint training. In this paper, we propose Differentiable Retrieval Augmentation via Generative lANguage modeling(Dragan), to address this problem by a novel differentiable reformulation. We demonstrate the effectiveness of our proposed method on a challenging NLP task in e-commerce search, namely query intent classification. Both the experimental results and ablation study show that the proposed method significantly and reasonably improves the state-of-the-art baselines on both offline evaluation and online A/B test. Yunjiang Jiang, Yiming Qiu 0003, Han Zhang 0047, Wen-Yun Yang |
CIKM | 5 |
| 2022 | Pre-training Tasks for User Intent Detection and Embedding Retrieval in E-commerce SearchabstractBERT-style models pre-trained on the general corpus (e.g., Wikipedia) and fine-tuned on specific task corpus, have recently emerged as breakthrough techniques in many NLP tasks: question answering, text classification, sequence labeling and so on. However, this tech- nique may not always work, especially for two scenarios: a corpus that contains very different text from the general corpus Wikipedia, or a task that learns embedding spacial distribution for a specific purpose (e.g., approximate nearest neighbor search). In this paper, to tackle the above two scenarios that we have encountered in an industrial e-commerce search system, we propose customized and novel pre-training tasks for two critical modules: user intent detec- tion and semantic embedding retrieval. The customized pre-trained models after fine-tuning, being less than 10% of BERT-base's size in order to be feasible for cost-efficient CPU serving, significantly improve the other baseline models: 1) no pre-training model and 2) fine-tuned model from the official pre-trained BERT using general corpus, on both offline datasets and online system. We have open sourced our datasets 1 for the sake of reproducibility and future works. Yiming Qiu 0003, Han Zhang 0047, Jingwei Zhuo, Songlin Wang, Sulong Xu, Bo Long, Wen-Yun Yang |
CIKM | 10 |
| 2022 | Givens Coordinate Descent Methods for Rotation Matrix Learning in Trainable Embedding Indexes
Yunjiang Jiang, Han Zhang 0047, Yiming Qiu 0003, Bo Long, Wen-Yun Yang |
ICLR | 6 |
| 2021 | Query Rewriting via Cycle-Consistent Translation for E-Commerce SearchabstractNowadays e-commerce search has become an integral part of many people's shopping routines. One critical challenge in today's e-commerce search is the semantic matching problem where the relevant items may not contain the exact terms in the user query. In this paper, we propose a novel deep neural network based approach to query rewriting, in order to tackle this problem. Specifically, we formulate query rewriting into a cyclic machine translation problem to leverage abundant click log data. Then we introduce a novel cyclic consistent training algorithm in conjunction with state-of-the-art machine translation models to achieve the optimal performance in terms of query rewriting accuracy. In order to make it practical in industrial scenarios, we optimize the syntax tree construction to reduce computational cost and online serving latency. Offline experiments show that the proposed method is able to rewrite hard user queries into more standard queries that are more appropriate for the inverted index to retrieve. Comparing with human curated rule-based method, the proposed model significantly improves query rewriting diversity while maintaining good relevancy. Online A/B experiments show that it improves core e-commerce business metrics significantly. Since the summer of 2020, the proposed model has been launched into our search engine production, serving hundreds of millions of users. Yiming Qiu 0003, Kang Zhang 0005, Han Zhang 0047, Songlin Wang, Sulong Xu, Bo Long, Wen-Yun Yang |
ICDE | 8 |
| 2021 | SearchGCN: Powering Embedding Retrieval by Graph Convolution Networks for E-Commerce SearchabstractGraph convolution networks (GCN), which recently becomes new state-of-the-art method for graph node classification, recommendation and other applications, has not been successfully applied to industrial-scale search engine yet. In this proposal, we introduce our approach, namely SearchGCN, for embedding-based candidate retrieval in one of the largest e-commerce search engine in the world. Empirical studies demonstrate that SearchGCN learns better embedding representations than existing methods, especially for long tail queries and items. Thus, SearchGCN has been deployed into JD.com's search production since July 2020. Xinlin Xia, Han Zhang 0047, Songlin Wang, Sulong Xu, Bo Long, Wen-Yun Yang |
SIGIR | 8 |
| 2021 | Joint Learning of Deep Retrieval Model and Product Quantization based Embedding IndexabstractEmbedding index that enables fast approximate nearest neighbor(ANN) search, serves as an indispensable component for state-of-the-art deep retrieval systems. Traditional approaches, often separating the two steps of embedding learning and index building, incur additional indexing time and decayed retrieval accuracy. In this paper, we propose a novel method called Poeem, which stands for product quantization based embedding index jointly trained with deep retrieval model, to unify the two separate steps within an end-to-end training, by utilizing a few techniques including the gradient straight-through estimator, warm start strategy, optimal space decomposition and Givens rotation. Extensive experimental results show that the proposed method not only improves retrieval accuracy significantly but also reduces the indexing time to almost none. We have open sourced our approach for the sake of comparison and reproducibility. Han Zhang 0047, Hongwei Shen, Yiming Qiu 0003, Yunjiang Jiang, Songlin Wang, Sulong Xu, Bo Long, Wen-Yun Yang |
SIGIR | 9 |
| 2015 | Identification of causal genes for complex traitsabstractMOTIVATION: Although genome-wide association studies (GWAS) have identified thousands of variants associated with common diseases and complex traits, only a handful of these variants are validated to be causal. We consider 'causal variants' as variants which are responsible for the association signal at a locus. As opposed to association studies that benefit from linkage disequilibrium (LD), the main challenge in identifying causal variants at associated loci lies in distinguishing among the many closely correlated variants due to LD. This is particularly important for model organisms such as inbred mice, where LD extends much further than in human populations, resulting in large stretches of the genome with significantly associated variants. Furthermore, these model organisms are highly structured and require correction for population structure to remove potential spurious associations. RESULTS: In this work, we propose CAVIAR-Gene (CAusal Variants Identification in Associated Regions), a novel method that is able to operate across large LD regions of the genome while also correcting for population structure. A key feature of our approach is that it provides as output a minimally sized set of genes that captures the genes which harbor causal variants with probability ρ. Through extensive simulations, we demonstrate that our method not only speeds up computation, but also have an average of 10% higher recall rate compared with the existing approaches. We validate our method using a real mouse high-density lipoprotein data (HDL) and show that CAVIAR-Gene is able to identify Apoa2 (a gene known to harbor causal variants for HDL), while reducing the number of genes that need to be tested for functionality by a factor of 2. AVAILABILITY AND IMPLEMENTATION: Software is freely available for download at genetics.cs.ucla.edu/caviar. Farhad Hormozdiari, Gleb Kichaev, Wen-Yun Yang, Bogdan Pasaniuc, Eleazar Eskin |
Bioinform. | 3 |
| 2014 | A Spatial-Aware Haplotype Copying Model with Applications to Genotype Imputation
Wen-Yun Yang, Farhad Hormozdiari, Eleazar Eskin, Bogdan Pasaniuc |
RECOMB | 1 |
| 2013 | Leveraging reads that span multiple single nucleotide polymorphisms for haplotype inference from sequencing dataabstractMOTIVATION: Haplotypes, defined as the sequence of alleles on one chromosome, are crucial for many genetic analyses. As experimental determination of haplotypes is extremely expensive, haplotypes are traditionally inferred using computational approaches from genotype data, i.e. the mixture of the genetic information from both haplotypes. Best performing approaches for haplotype inference rely on Hidden Markov Models, with the underlying assumption that the haplotypes of a given individual can be represented as a mosaic of segments from other haplotypes in the same population. Such algorithms use this model to predict the most likely haplotypes that explain the observed genotype data conditional on reference panel of haplotypes. With rapid advances in short read sequencing technologies, sequencing is quickly establishing as a powerful approach for collecting genetic variation information. As opposed to traditional genotyping-array technologies that independently call genotypes at polymorphic sites, short read sequencing often collects haplotypic information; a read spanning more than one polymorphic locus (multi-single nucleotide polymorphic read) contains information on the haplotype from which the read originates. However, this information is generally ignored in existing approaches for haplotype phasing and genotype-calling from short read data. RESULTS: In this article, we propose a novel framework for haplotype inference from short read sequencing that leverages multi-single nucleotide polymorphic reads together with a reference panel of haplotypes. The basis of our approach is a new probabilistic model that finds the most likely haplotype segments from the reference panel to explain the short read sequencing data for a given individual. We devised an efficient sampling method within a probabilistic model to achieve superior performance than existing methods. Using simulated sequencing reads from real individual genotypes in the HapMap data and the 1000 Genomes projects, we show that our method is highly accurate and computationally efficient. Our haplotype predictions improve accuracy over the basic haplotype copying model by ∼20% with comparable computational time, and over another recently proposed approach Hap-SeqX by ∼10% with significantly reduced computational time and memory usage. AVAILABILITY: Publicly available software is available at http://genetics.cs.ucla.edu/harsh CONTACT: [email protected] or [email protected]. Wen-Yun Yang, Farhad Hormozdiari, Zhanyong Wang, Dan He 0001, Bogdan Pasaniuc, Eleazar Eskin |
Bioinform. | 1 |
| 2012 | CNVeM: Copy Number Variation Detection Using Uncertainty of Read Mapping
Zhanyong Wang, Farhad Hormozdiari, Wen-Yun Yang, Eran Halperin, Eleazar Eskin |
RECOMB | 3 |
| 2011 | A structural support vector method for extracting contexts and answers of questions from online forums
Yunbo Cao, Wen-Yun Yang, Chin-Yew Lin, Yong Yu 0001 |
Inf. Process. Manag. | 2 |
| 2011 | Incorporating cellular sorting structure for better prediction of protein subcellular locationsabstractThis article explores the interdependences between subcellular locations and incorporates them with support vector machines for prediction of protein subcellular localisation. Traditional prediction systems utilise a ‘flat’ structure of classifiers, such as the one-versus-all and one-versus-one schemes, with amino acid compositions to perform the prediction. Apart from those existing studies that ignore the interdependences between subcellular locations, we take advantage of a hierarchical structure to organise the subcellular locations and model their relationships. Here, we propose to use four kinds of hierarchical prediction methods and make comparative studies on three datasets. Experimental results show that three of the hierarchical models outperform the traditional ‘flat’ model in terms of tree loss values. In particular, one hierarchical model outperforms the traditional ‘flat’ model for all evaluation measures. Moreover, we gained some valuable insights into the sorting process by using hierarchical structures. Wen-Yun Yang, Bao-Liang Lu, James T. Kwok |
J. Exp. Theor. Artif. Intell. | 1 |
| 2010 | Spectral and Semidefinite Relaxation of the CLUHSIC AlgorithmabstractCLUHSIC is a recent clustering framework that unifies the geometric, spectral and statistical views of clustering.In this paper, we show that the recently proposed discriminative view of clustering, which includes the DIFFRAC and DisKmeans algorithms, can also be unified under the CLUH-SIC framework.Moreover, CLUHSIC involves integer programming and one has to resort to heuristics such as iterative local optimization.In this paper, we propose two relaxations that are much more disciplined.The first one uses spectral techniques while the second one is based on semidefinite programming (SDP).Experimental results on a number of structured clustering tasks show that the proposed method significantly outperforms existing optimization methods for CLUHSIC.Moreover, it can also be used in semi-supervised classification.Experiments on real-world protein subcellular localization data sets clearly demonstrate the ability of CLUHSIC in incorporating structural and evolutionary information. Wen-Yun Yang, James T. Kwok, Bao-Liang Lu |
SDM | 1 |
| 2009 | A Structural Support Vector Method for Extracting Contexts and Answers of Questions from Online Forums
Wen-Yun Yang, Yunbo Cao, Chin-Yew Lin |
EMNLP | 1 |
| 2008 | String Kernels with Feature Selection for SVM Protein Classification
Wen-Yun Yang, Bao-Liang Lu |
APBC | 1 |
| 2008 | Classification of Protein Sequences Based on Word Segmentation Methods
Yang Yang 0030, Bao-Liang Lu, Wen-Yun Yang |
APBC | 3 |
| 2006 | A Comparative Study on Feature Extraction from Protein Sequences for Subcellular Localization PredictionabstractOne of the central problems in computational biology is to identify the protein function in an automated and high-throughput fashion. A key step in this process is to predict subcellular compartment the protein belongs to, since the protein localization closely correlates with its function. A wide variety of methods for protein subcellular localization has been proposed over recent years. They fall into two categories, sequence-based and database-based. The first one is to extract useful features from amino acid sequences and strives to discover the principles behind protein localization process. The second one is more apt to conduct data mining from existing public annotation databases. This paper focuses on the sequence-based approach and exploits the discriminative ability contained in amino acid sequences for protein subcellular localization. By using support vector machines (SVMs) as predictors, we conducted comparisons among amino acid composition approach, amino acid tuple approach, voting scheme, and a new characteristic representation of proteins proposed in this paper. Our experiments are carried out on 7579 eukaryotic protein sequences from 12 subcellular locations. The highest accuracy, 82.8% across 5-fold cross validation, is obtained by voting scheme using five predictors. This is the best performance achieved on this dataset using sequence-based approach. Our experiments demonstrate that there are considerable potentials on improving prediction accuracy by exploiting protein sequences, which have not been fully utilized so far, and more explorations are still needed in this direction Wen-Yun Yang, Bao-Liang Lu, Yang Yang 0030 |
CIBCB | 1 |