Yuan Wang 0076

dblp:41/3241-76 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0002-4241-9905ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers
abstract
Yuan Wang, Xuyang Wu, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yuan Wang 0076, Xuyang Wu 0002, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang 0008
NAACL-HLT1
2024 Table Transformers for imputing textual attributes
abstract
Missing data in tabular dataset is a common issue as the performance of downstream tasks usually depends on the completeness of the training dataset. Previous missing data imputation methods focus on numeric and categorical columns, but we propose a novel end-to-end approach called Table Transformers for Imputing Textual Attributes (TTITA) based on the transformer to impute unstructured textual columns using other columns in the table. We conduct extensive experiments on three datasets, and our approach shows competitive performance outperforming baseline models such as recurrent neural networks and Llama2. The performance improvement is more significant when the target sequence has a longer length. Additionally, we incorporate multi-task learning to simultaneously impute for heterogeneous columns, boosting the performance for text imputation. We also qualitatively compare with ChatGPT for realistic applications. • Proposed TTITA to impute text attributes given other heterogeneous tabular columns. • Encoded inputs into a context vector for cross-attention in the transformer decoder. • Outperformed baseline models including the GRU and Llama2 on real-world datasets. • Incorporated multi-task learning for multi-column imputation and boosting performance. • Prepared the software as an open-source package for custom applications.
Ting-Ruen Wei, Yuan Wang 0076, Yoshitaka Inoue, Hsin-Tai Wu, Yi Fang 0008
Pattern Recognit. Lett.2
2024 A Unified Meta-Learning Framework for Fair Ranking With Curriculum Learning
abstract
In recent information retrieval systems, it is observed that the datasets used to train machine learning models can be biased, leading to systematic discrimination against certain demographic groups, which means the ranking utility of specific groups is often lower than others in a biased dataset. Training models on these datasets will further decrease the exposure of the minority groups. To address this problem, we propose a Meta Curriculum-based Fair Ranking framework (MCFR) which could alleviate the data bias issue through the weighted loss using gradient-based learning to learn. Specifically, we optimize a meta learner from a sampled dataset (meta-dataset), and meanwhile train a ranking model on the whole (biased) dataset. The meta-dataset is sampled with a curriculum learning scheduler to guide the meta learner's training to gradually mitigate the skewness towards biased attributes. The meta learner serves as a weighting function to make the ranking loss focus more on the minority group. We formulate the proposed MCFR as a bilevel optimization problem and solve it using gradients through gradients. Extensive experiments on real-world datasets demonstrate that our approach can be used as a generic framework to work with various ranking losses and fairness metrics.
Yuan Wang 0076, Zhiqiang Tao, Yi Fang 0008
IEEE Trans. Knowl. Data Eng.1
2023 An Empirical Study of Selection Bias in Pinterest Ads Retrieval
abstract
Data selection bias has been a long-lasting challenge in the machine learning domain, especially in multi-stage recommendation systems, where the distribution of labeled items for model training is very different from that of the actual candidates during inference time. This distribution shift is even more prominent in the context of online advertising where the user base is diverse and the platform contains a wide range of contents. In this paper, we first investigate the data selection bias in the upper funnel (Ads Retrieval) of Pinterest's multi-cascade ads ranking system. We then conduct comprehensive experiments to assess the performance of various state-of-the-art methods, including transfer learning, adversarial learning, and unsupervised domain adaptation. Moreover, we further introduce some modifications into the unsupervised domain adaptation and evaluate the performance of different variants of this modified method. Our online A/B experiments show that the modified version of unsupervised domain adaptation (MUDA) could provide the largest improvements to the performance of Pinterest's advertisement ranking system compared with other methods and the one used in current production.
Yuan Wang 0076, Peifeng Yin, Zhiqiang Tao, Hari Venkatesan, Jin Lai, Yi Fang 0008, PJ Xiao
KDD1
2022 A Meta-learning Approach to Fair Ranking
abstract
In recent years, the fairness in information retrieval (IR) system has received increasing research attention. While the data-driven ranking models achieve significant improvements over traditional methods, the dataset used to train such models is usually biased, which causes unfairness in the ranking models. For example, the collected imbalance dataset on the subject of the expert search usually leads to systematic discrimination on the specific demographic groups such as race, gender, etc, which further reduces the exposure for the minority group. To solve this problem, we propose a Meta-learning based Fair Ranking (MFR) model that could alleviate the data bias for protected groups through an automatically-weighted loss. Specifically, we adopt a meta-learning framework to explicitly train a meta-learner from an unbiased sampled dataset (meta-dataset), and simultaneously, train a listwise learning-to-rank (LTR) model on the whole (biased) dataset governed by "fair" loss weights. The meta-learner serves as a weighting function to make the ranking loss attend more on the minority group. To update the parameters of the weighting function and the ranking model, we formulate the proposed MFR as a bilevel optimization problem and solve it using the gradients through gradients. Experimental results on several real-world datasets demonstrate that the proposed method achieves a comparable ranking performance and significantly improves the fairness metric compared with state-of-the-art methods.
Yuan Wang 0076, Zhiqiang Tao, Yi Fang 0008
SIGIR1
2022 Learning user preferences through online conversations via personalized memory transfer
Nagaarchana Godavarthy, Yuan Wang 0076, Travis Ebesu, Un Suthee, Min Xie 0002, Yi Fang 0008
Inf. Retr. J.2
2021 Karaoke Key Recommendation Via Personalized Competence-Based Rating Prediction
abstract
Karaoke machines have become a popular choice for many people’s daily entertainment. In this paper, we address a novel task of recommending a suitable key for a user to sing a given song to meet his or her vocal competence, by proposing the Personalized Competence-based Rating Prediction (PCRP) model. Specifically, we learn the song embedding vectors from the sequences of songs’ notes, and then design a history encoder with recurrent units to extract users’ vocal information from the history rating records and utilize a rating decoder based on the Transformer. The experimental results on a real world karaoke rating dataset demonstrate the effectiveness of the proposed approach.
Yuan Wang 0076, Shigeki Tanaka, Keita Yokoyama, Hsin-Tai Wu, Yi Fang 0008
ICASSP1
2020 Leveraging an Efficient and Semantic Location Embedding to Seek New Ports of Bike Share Services
abstract
For short distance traveling in crowded urban areas, bike share services is becoming popular owing to the flexibility and convenience. To expand the service coverage, one of the key tasks is to seek new service ports, which requires to well understand the underlying features of the existing service ports. In this paper, we propose a new model, named for Efficient and Semantic Location Embedding (ESLE)1, which carries both geospatial and semantic information of the geo-locations. To generate ESLE, we first train a multi-label model with a deep Convolutional Neural Network (CNN) by feeding the static map-tile images and then extract location embedding vectors from the model. Compared to most recent relevant literature, ESLE is not only much cheaper in computation, but also easier to interpret via a systematic semantic analysis. Finally, we apply ESLE to seek new service ports for NTT DOCOMO’s bike share services operated in Japan. The initial results demonstrate the effectiveness of ESLE, and provide a few insights that might be difficult to discover by using the conventional approaches.
Yuan Wang 0076, Chenwei Wang 0001, Yinan Ling, Keita Yokoyama, Hsin-Tai Wu, Yi Fang 0008
IEEE BigData1