Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wei Yang 0017

dblp:03/1094-17 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0003-1266-048XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 100%
Artificial intelligence
5 papers
Deep learning architectures and training · 27% Information extraction and text analysis · 22% Representation and self-supervised learning · 16%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › retrieval models › neural retrieval
neural ranking model
1.232020
Capreolus: A Toolkit for End-to-End Neural Ad Hoc Retrieval · WSDM 2020
Critically Examining the "Neural Hype": Weak Baselines and the Additivity of Effectiveness Gains from Neural Ranking Models · SIGIR 2019
Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media Search · AAAI 2019
Information retrieval
retrieval models
0.822019
Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval · EMNLP/IJCNLP (1) 2019
Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media Search · AAAI 2019
Information retrieval › retrieval models
ad-hoc retrieval
0.522020
Capreolus: A Toolkit for End-to-End Neural Ad Hoc Retrieval · WSDM 2020
Critically Examining the "Neural Hype": Weak Baselines and the Additivity of Effectiveness Gains from Neural Ranking Models · SIGIR 2019
Natural language and speech › Machine translation
neural machine translation
0.512021
Optimizing Deeper Transformers on Small Datasets · ACL/IJCNLP (1) 2021
Machine learning › Deep learning architectures and training › neural network training
training on small datasets
0.512021
Optimizing Deeper Transformers on Small Datasets · ACL/IJCNLP (1) 2021
Machine learning › Deep learning architectures and training
transformer
0.512021
Optimizing Deeper Transformers on Small Datasets · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision
0.412020
Distant Supervision for Multi-Stage Fine-Tuning in Retrieval-Based Question Answering · WWW 2020
Natural language and speech › Question answering and dialogue systems
retrieval-based question answering
0.412020
Distant Supervision for Multi-Stage Fine-Tuning in Retrieval-Based Question Answering · WWW 2020
Information retrieval › evaluation › evaluation methodology
reproducibility
0.412020
Capreolus: A Toolkit for End-to-End Neural Ad Hoc Retrieval · WSDM 2020
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-domain retrieval
0.412019
Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval · EMNLP/IJCNLP (1) 2019
Information retrieval
evaluation
0.412019
Critically Examining the "Neural Hype": Weak Baselines and the Additivity of Effectiveness Gains from Neural Ranking Models · SIGIR 2019
Information retrieval › web search › web information retrieval › social media retrieval
microblog retrieval
0.412019
Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media Search · AAAI 2019
Information retrieval › retrieval models
neural retrieval
0.412019
Critically Examining the "Neural Hype": Weak Baselines and the Additivity of Effectiveness Gains from Neural Ranking Models · SIGIR 2019
Information retrieval › web search › web information retrieval
social media retrieval
0.412019
Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media Search · AAAI 2019
Machine learning › Representation and self-supervised learning › word representation
cross-domain word embedding
0.312017
A Simple Regularization-based Algorithm for Learning Cross-Domain Word Embeddings · EMNLP 2017
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.312017
A Simple Regularization-based Algorithm for Learning Cross-Domain Word Embeddings · EMNLP 2017

Methods — techniques the papers use, named apart from their topics

neural sentence modeling · 0.8query expansion · 0.4passage retrieval · 0.4neural network · 0.4distant supervision · 0.4data augmentation · 0.4BERT · 0.4syntactic structure · 0.4re-ranking · 0.4meta-analysis · 0.4hierarchical neural network · 0.4convolutional neural network · 0.4contextual embeddings · 0.4regularization · 0.3
YearPublicationVenuePosition
2021 Optimizing Deeper Transformers on Small Datasets
abstract
Peng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Simon J.D. Prince, Yanshuai Cao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dhruv Kumar 0005, Wei Yang 0017, Wenjie Zi, Keyi Tang, Chenyang Huang 0001, Jackie Chi Kit Cheung, Simon Prince, Yanshuai Cao
ACL/IJCNLP (1)3
2020 Capreolus: A Toolkit for End-to-End Neural Ad Hoc Retrieval
abstract
We present Capreolus, a toolkit designed to facilitate end-to-end it ad hoc retrieval experiments with neural networks by providing implementations of prominent neural ranking models within a common framework. Our toolkit adopts a standard reranking architecture via tight integration with the Anserini toolkit for candidate document generation using standard bag-of-words approaches. Using Capreolus, we are able to reproduce Yang et al.'s recent SIGIR 2019 finding that, in a reranking scenario on the test collection from the TREC 2004 Robust Track, many neural retrieval models do not significantly outperform a strong query expansion baseline. Furthermore, we find that this holds true for five additional models implemented in Capreolus. We describe the architecture and design of our toolkit, which includes a Web interface to facilitate comparisons between rankings returned by different models.
Andrew Yates, Siddhant Arora, Xinyu Zhang 0018, Wei Yang 0017, Kevin Martin Jose, Jimmy Lin
WSDM4
2020 Distant Supervision for Multi-Stage Fine-Tuning in Retrieval-Based Question Answering
abstract
We tackle the problem of question answering directly on a large document collection, combining simple “bag of words” passage retrieval with a BERT-based reader for extracting answer spans. In the context of this architecture, we present a data augmentation technique using distant supervision to automatically annotate paragraphs as either positive or negative examples to supplement existing training data, which are then used together to fine-tune BERT. We explore a number of details that are critical to achieving high accuracy in this setup: the proper sequencing of different datasets during fine-tuning, the balance between “difficult” vs. “easy” examples, and different approaches to gathering negative examples. Experimental results show that, with the appropriate settings, we can achieve large gains in effectiveness on two English and two Chinese QA datasets. We are able to achieve results at or near the state of the art without any modeling advances, which once again affirms the cliché “there’s no data like more data”.
Yuqing Xie 0001, Wei Yang 0017, Luchen Tan, Kun Xiong, Nicholas Jing Yuan, Baoxing Huai, Ming Li 0001, Jimmy Lin
WWW2
2019 Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media Search
abstract
Despite substantial interest in applications of neural networks to information retrieval, neural ranking models have mostly been applied to “standard” ad hoc retrieval tasks over web pages and newswire articles. This paper proposes MP-HCNN (Multi-Perspective Hierarchical Convolutional Neural Network), a novel neural ranking model specifically designed for ranking short social media posts. We identify document length, informal language, and heterogeneous relevance signals as features that distinguish documents in our domain, and present a model specifically designed with these characteristics in mind. Our model uses hierarchical convolutional layers to learn latent semantic soft-match relevance signals at the character, word, and phrase levels. A poolingbased similarity measurement layer integrates evidence from multiple types of matches between the query, the social media post, as well as URLs contained in the post. Extensive experiments using Twitter data from the TREC Microblog Tracks 2011–2014 show that our model significantly outperforms prior feature-based as well as existing neural ranking models. To our best knowledge, this paper presents the first substantial work tackling search over social media posts using neural ranking models. Our code and data are publicly available.1
Jinfeng Rao, Wei Yang 0017, Yuhao Zhang 0004, Ferhan Ture, Jimmy Lin
AAAI2
2019 Incorporating Contextual and Syntactic Structures Improves Semantic Similarity Modeling
abstract
Linqing Liu, Wei Yang, Jinfeng Rao, Raphael Tang, Jimmy Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Linqing Liu, Wei Yang 0017, Jinfeng Rao, Raphael Tang, Jimmy Lin
EMNLP/IJCNLP (1)2
2019 Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval
abstract
Zeynep Akkalyoncu Yilmaz, Wei Yang, Haotian Zhang, Jimmy Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zeynep Akkalyoncu Yilmaz, Wei Yang 0017, Haotian Zhang 0001, Jimmy Lin
EMNLP/IJCNLP (1)2
2019 Critically Examining the "Neural Hype": Weak Baselines and the Additivity of Effectiveness Gains from Neural Ranking Models
abstract
Is neural IR mostly hype? In a recent SIGIR Forum article, Lin expressed skepticism that neural ranking models were actually improving ad hoc retrieval effectiveness in limited data scenarios. He provided anecdotal evidence that authors of neural IR papers demonstrate "wins" by comparing against weak baselines. This paper provides a rigorous evaluation of those claims in two ways: First, we conducted a meta-analysis of papers that have reported experimental results on the TREC Robust04 test collection. We do not find evidence of an upward trend in effectiveness over time. In fact, the best reported results are from a decade ago and no recent neural approach comes close. Second, we applied five recent neural models to rerank the strong baselines that Lin used to make his arguments. A significant improvement was observed for one of the models, demonstrating additivity in gains. While there appears to be merit to neural IR approaches, at least some of the gains reported in the literature appear illusory.
Wei Yang 0017, Kuang Lu, Jimmy Lin
SIGIR1
2017 A Simple Regularization-based Algorithm for Learning Cross-Domain Word Embeddings
abstract
Learning word embeddings has received a significant amount of attention recently.Often, word embeddings are learned in an unsupervised manner from a large collection of text.The genre of the text typically plays an important role in the effectiveness of the resulting embeddings.How to effectively train word embedding models using data from different domains remains a problem that is underexplored.In this paper, we present a simple yet effective method for learning word embeddings based on text from different domains.We demonstrate the effectiveness of our approach through extensive experiments on various down-stream NLP tasks.
Wei Yang 0017, Wei Lu 0011, Vincent Wenchen Zheng
EMNLP1