VLDB 2026 Research / reviewers in the wild / expert
Wang Ling
dblp:91/7651
· DBLP profile ↗
30ranked-venue papers
11as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 11 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
17 papers |
Representation and self-supervised learning · 30% Language models and text generation · 28% Information extraction and text analysis · 14% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 54% Compilers and program optimization · 46% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model |
0.7 | 2 | 2019 | Variational Smoothing in Recurrent Neural Network Language Models · ICLR (Poster) 2019 Memory Architectures in Recurrent Neural Network Language Models · ICLR (Poster) 2018 |
Natural language and speech › Language models and text generation › decoding
adaptive decoding |
0.6 | 1 | 2022 | Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022 |
Natural language and speech › Language models and text generation
decoding |
0.6 | 1 | 2022 | Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022 |
Natural language and speech › Language models and text generation
text generation |
0.6 | 1 | 2022 | Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.5 | 3 | 2015 | Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015 Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces · ACL (1) 2015 Not All Contexts Are Created Equal: Better Word Representations with Variable Attention · EMNLP 2015 |
Machine learning › Representation and self-supervised learning › word representation
word representation learning |
0.5 | 2 | 2016 | Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning · ACL (1) 2016 Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces · ACL (1) 2015 |
Machine learning › Representation and self-supervised learning › representation learning
language representation learning |
0.4 | 1 | 2020 | A Mutual Information Maximization Perspective of Language Representation Learning · ICLR 2020 |
Machine learning › Representation and self-supervised learning
mutual information maximization |
0.4 | 1 | 2020 | A Mutual Information Maximization Perspective of Language Representation Learning · ICLR 2020 |
Machine learning › Representation and self-supervised learning
word representation |
0.4 | 2 | 2015 | Not All Contexts Are Created Equal: Better Word Representations with Variable Attention · EMNLP 2015 Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation · EMNLP 2015 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.4 | 1 | 2019 | Variational Smoothing in Recurrent Neural Network Language Models · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning
memory architectures |
0.3 | 1 | 2018 | Memory Architectures in Recurrent Neural Network Language Models · ICLR (Poster) 2018 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.3 | 1 | 2017 | Reference-Aware Language Models · EMNLP 2017 |
Natural language and speech › Language models and text generation
language modeling |
0.3 | 1 | 2017 | Reference-Aware Language Models · EMNLP 2017 |
Natural language and speech › Question answering and dialogue systems
math word problem solving |
0.3 | 1 | 2017 | Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems · ACL (1) 2017 |
Machine learning › Reinforcement learning
reinforcement learning for NLP |
0.3 | 1 | 2017 | Learning to Compose Words into Sentences with Reinforcement Learning · ICLR (Poster) 2017 |
Program synthesis and code generation
inductive program synthesis |
0.3 | 1 | 2017 | Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems · ACL (1) 2017 |
Machine learning › Learning paradigms
curriculum learning |
0.2 | 1 | 2016 | Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning · ACL (1) 2016 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.2 | 1 | 2016 | Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › semantic parsing
semi-supervised semantic parsing |
0.2 | 1 | 2016 | Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence transduction |
0.2 | 1 | 2016 | Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016 |
Machine learning › Representation and self-supervised learning › representation learning
sequential autoencoder |
0.2 | 1 | 2016 | Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016 |
Compilers and program optimization
code generation |
0.2 | 1 | 2016 | Latent Predictor Networks for Code Generation · ACL (1) 2016 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
intrinsic evaluation |
0.2 | 1 | 2015 | Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015 |
Computer vision › Vision and language
open-vocabulary models |
0.2 | 1 | 2015 | Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation · EMNLP 2015 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.2 | 1 | 2015 | Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
subspace alignment |
0.2 | 1 | 2015 | Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
subspace embedding |
0.2 | 1 | 2015 | Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.2 | 1 | 2015 | Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
transition-based dependency parsing |
0.2 | 1 | 2015 | Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015 |
Natural language and speech › Machine translation
parallel corpora |
0.2 | 1 | 2013 | Microblogs as Parallel Corpora · ACL (1) 2013 |
Methods — techniques the papers use, named apart from their topics
tree search · 0.6sequence-to-sequence · 0.6indirect supervision · 0.6mutual information estimation · 0.4contrastive learning · 0.4variational inference · 0.4smoothing · 0.4recurrent neural network · 0.3reinforcement learning · 0.3latent variable model · 0.3latent predictor networks · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards Dialogue Modeling Beyond TextabstractIn this paper, we model aspects of communication beyond the words that are said. Specifically, we aim to detect interruptions and active listening events, which are important elements in any dialogue. We build a dataset with fine-grained annotations for each category and train multimodal models that take into account all channels in a digital conversation, that is, the video, the audio, and the text. Our experiments show that multimodality is a necessary component in modeling the complexity of the non-textual components of the conversation as different artifacts require different modalities to capture effectively. Tongzi Wu, Wang Ling, Hojin Yang, Joana Veloso, Ruixin Huang, Norberto Guimaraes, Scott Sanner |
ICASSP | 3 |
| 2022 | Enabling Arbitrary Translation Objectives with Adaptive Tree Search
Wang Ling, Wojciech Stokowiec, Domenic Donato, Chris Dyer, Lei Yu 0008, Laurent Sartran, Austin Matthews |
ICLR | 1 |
| 2020 | A Mutual Information Maximization Perspective of Language Representation Learning
Lingpeng Kong, Cyprien de Masson d'Autume, Lei Yu 0008, Wang Ling, Zihang Dai, Dani Yogatama |
ICLR | 4 |
| 2020 | Better Document-Level Machine Translation with Bayes' RuleabstractWe show that Bayes’ rule provides an effective mechanism for creating document translation models that can be learned from only parallel sentences and monolingual documents a compelling benefit because parallel documents are not always available. In our formulation, the posterior probability of a candidate translation is the product of the unconditional (prior) probability of the candidate output document and the “reverse translation probability” of translating the candidate output back into the source language. Our proposed model uses a powerful autoregressive language model as the prior on target language documents, but it assumes that each sentence is translated independently from the target to the source language. Crucially, at test time, when a source document is observed, the document language model prior induces dependencies between the translations of the source sentences in the posterior. The model’s independence assumption not only enables efficient use of available data, but it additionally admits a practical left-to-right beam-search algorithm for carrying out inference. Experiments show that our model benefits from using cross-sentence context in the language model, and it outperforms existing document translation approaches. Lei Yu 0008, Laurent Sartran, Wojciech Stokowiec, Wang Ling, Lingpeng Kong, Phil Blunsom, Chris Dyer |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | Variational Smoothing in Recurrent Neural Network Language Models
Lingpeng Kong, Gábor Melis, Wang Ling, Lei Yu 0008, Dani Yogatama |
ICLR (Poster) | 3 |
| 2018 | Memory Architectures in Recurrent Neural Network Language Models
Dani Yogatama, Yishu Miao, Gábor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, Phil Blunsom |
ICLR (Poster) | 4 |
| 2017 | Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word ProblemsabstractSolving algebraic word problems requires executing a series of arithmetic operations-a program-to obtain a final answer.However, since programs can be arbitrarily complicated, inducing them directly from question-answer pairs is a formidable challenge.To make this task more feasible, we solve these problems by generating answer rationales, sequences of natural language and human-readable mathematical expressions that derive the final answer through a series of small steps.Although rationales do not explicitly specify programs, they provide a scaffolding for their structure via intermediate milestones.To evaluate our approach, we have created a new 100,000-sample dataset of questions, answers and rationales.Experimental results show that indirect supervision of program learning via answer rationales is a promising strategy for inducing arithmetic programs. Wang Ling, Dani Yogatama, Chris Dyer, Phil Blunsom |
ACL (1) | 1 |
| 2017 | Reference-Aware Language ModelsabstractWe propose a general class of language models that treat reference as discrete stochastic latent variables.This decision allows for the creation of entity mentions by accessing external databases of referents (required by, e.g., dialogue generation) or past internal state (required to explicitly model coreferentiality).Beyond simple copying, our coreference model can additionally refer to a referent using varied mention forms (e.g., a reference to "Jane" can be realized as "she"), a characteristic feature of reference in natural languages.Experiments on three representative applications show our model variants outperform models based on deterministic attention and standard language modeling baselines. Phil Blunsom, Chris Dyer, Wang Ling |
EMNLP | 4 |
| 2017 | Learning to Compose Words into Sentences with Reinforcement Learning
Dani Yogatama, Phil Blunsom, Chris Dyer, Edward Grefenstette, Wang Ling |
ICLR (Poster) | 5 |
| 2016 | Latent Predictor Networks for Code GenerationabstractWang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, Fumin Wang, Andrew Senior. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, Andrew W. Senior |
ACL (1) | 1 |
| 2016 | Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation LearningabstractWe use Bayesian optimization to learn curricula for word representation learning, optimizing performance on downstream tasks that depend on the learned representations as features.The curricula are modeled by a linear ranking function which is the scalar product of a learned weight vector and an engineered feature vector that characterizes the different aspects of the complexity of each instance in the training corpus.We show that learning the curriculum improves performance on a variety of downstream tasks over random orders and in comparison to the natural corpus order. Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Brian MacWhinney, Chris Dyer |
ACL (1) | 3 |
| 2016 | Semantic Parsing with Semi-Supervised Sequential AutoencodersabstractWe present a novel semi-supervised approach for sequence transduction and apply it to semantic parsing.The unsupervised component is based on a generative model in which latent sentences generate the unpaired logical forms.We apply this method to a number of semantic parsing tasks focusing on domains with limited access to labelled training data and extend those datasets with synthetically generated logical forms. Tomás Kociský, Gábor Melis, Edward Grefenstette, Chris Dyer, Wang Ling, Phil Blunsom, Karl Moritz Hermann |
EMNLP | 5 |
| 2016 | Neural Network-Based Abstract Generation for Opinions and ArgumentsabstractWe study the problem of generating abstractive summaries for opinionated text. We propose an attention-based neural network model that is able to absorb information from multiple text units to construct informative, concise, and fluent summaries. An importance-based sampling method is designed to allow the encoder to integrate information from an important subset of input. Automatic evaluation indicates that our system outperforms state-of-the-art abstractive and extractive summarization systems on two newly collected datasets of movie reviews and arguments. Our system summaries are also rated as more informative and grammatical in human evaluation. Lu Wang 0008, Wang Ling |
HLT-NAACL | 2 |
| 2016 | Mining Parallel Corpora from Sina Weibo and TwitterabstractMicroblogs such as Twitter, Facebook, and Sina Weibo (China's equivalent of Twitter) are a remarkable linguistic resource. In contrast to content from edited genres such as newswire, microblogs contain discussions of virtually every topic by numerous individuals in different languages and dialects and in different styles. In this work, we show that some microblog users post “self-translated” messages targeting audiences who speak different languages, either by writing the same message in multiple languages or by retweeting translations of their original posts in a second language. We introduce a method for finding and extracting this naturally occurring parallel data. Identifying the parallel content requires solving an alignment problem, and we give an optimally efficient dynamic programming algorithm for this. Using our method, we extract nearly 3M Chinese–English parallel segments from Sina Weibo using a targeted crawl of Weibo users who post in multiple languages. Additionally, from a random sample of Twitter, we obtain substantial amounts of parallel data in multiple language pairs. Evaluation is performed by assessing the accuracy of our extraction approach relative to a manual annotation as well as in terms of utility as training data for a Chinese–English machine translation system. Relative to traditional parallel data resources, the automatically extracted parallel data yield substantial translation quality improvements in translating microblog text and modest improvements in translating edited news content. Wang Ling, Luís Marujo, Chris Dyer, Alan W. Black, Isabel Trancoso |
Comput. Linguistics | 1 |
| 2016 | Exploring events and distributed representations of text in multi-document summarization
Luís Marujo, Wang Ling, Ricardo Ribeiro 0001, Anatole Gershman, Jaime G. Carbonell, David Martins de Matos, João Paulo da Silva Neto |
Knowl. Based Syst. | 2 |
| 2015 | Learning Word Representations from Scarce and Noisy Data with Embedding SubspacesabstractRamon F. Astudillo, Silvio Amir, Wang Ling, Mário Silva, Isabel Trancoso. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Ramón Fernandez Astudillo, Silvio Amir, Wang Ling, Mário J. Silva, Isabel Trancoso |
ACL (1) | 3 |
| 2015 | Transition-Based Dependency Parsing with Stack Long Short-Term MemoryabstractChris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith |
ACL (1) | 3 |
| 2015 | Finding Function in Form: Compositional Character Models for Open Vocabulary Word RepresentationabstractWang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, Tiago Luís |
EMNLP | 1 |
| 2015 | Not All Contexts Are Created Equal: Better Word Representations with Variable AttentionabstractWang Ling, Yulia Tsvetkov, Silvio Amir, Ramón Fermandez, Chris Dyer, Alan W Black, Isabel Trancoso, Chu-Cheng Lin. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramon Fermandez, Chris Dyer, Alan W. Black, Isabel Trancoso, Chu-Cheng Lin |
EMNLP | 1 |
| 2015 | Evaluation of Word Vector Representations by Subspace AlignmentabstractUnsupervisedly learned word vectors have proven to provide exceptionally effective features in many NLP tasks.Most common intrinsic evaluations of vector quality measure correlation with similarity judgments.However, these often correlate poorly with how well the learned representations perform as features in downstream evaluation tasks.We present QVEC-a computationally inexpensive intrinsic evaluation measure of the quality of word embeddings based on alignment to a matrix of features extracted from manually crafted lexical resources-that obtains strong correlation with performance of the vectors in a battery of downstream semantic evaluation tasks. 1 Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Guillaume Lample, Chris Dyer |
EMNLP | 3 |
| 2015 | Two/Too Simple Adaptations of Word2Vec for Syntax ProblemsabstractWang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso |
HLT-NAACL | 1 |
| 2015 | A linguistically motivated taxonomy for Machine Translation error analysis
Ângela Costa, Wang Ling, Tiago Luís, Rui Correia, Luísa Coheur |
Mach. Transl. | 2 |
| 2014 | Linguistic Evaluation of Support Verb Constructions by OpenLogos and Google Translate
Anabela Barreiro, Johanna Monti, Brigitte Orliac, Susanne Preuß, Kutz Arrieta, Wang Ling, Fernando Batista, Isabel Trancoso |
LREC | 6 |
| 2014 | Dual Subtitles as Parallel Corpora
Shikun Zhang, Wang Ling, Chris Dyer |
LREC | 2 |
| 2013 | Microblogs as Parallel Corpora
Wang Ling, Guang Xiang, Chris Dyer, Alan W. Black, Isabel Trancoso |
ACL (1) | 1 |
| 2013 | Paraphrasing 4 Microblog NormalizationabstractCompared to the edited genres that have played a central role in NLP research, microblog texts use a more informal register with nonstandard lexical items, abbreviations, and free orthographic variation.When confronted with such input, conventional text analysis tools often perform poorly.Normalization -replacing orthographically or lexically idiosyncratic forms with more standard variants -can improve performance.We propose a method for learning normalization rules from machine translations of a parallel corpus of microblog messages.To validate the utility of our approach, we evaluate extrinsically, showing that normalizing English tweets and then translating improves translation quality (compared to translating unnormalized text) using three standard web translation services as well as a phrase-based translation system trained on parallel microblog data. Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso |
EMNLP | 1 |
| 2012 | Overview of Computer-assisted Language Learning for European Portuguese at L2f
Thomas Pellegrini, Wang Ling, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede |
CSEDU (2) | 2 |
| 2012 | Entropy-based Pruning for Phrase-based Machine Translation
Wang Ling, João Graça, Isabel Trancoso, Alan W. Black |
EMNLP-CoNLL | 1 |
| 2011 | BP2EP - Adaptation of Brazilian Portuguese texts to European Portuguese
Luís Marujo, Nuno Grazina, Tiago Luís, Wang Ling, Luísa Coheur, Isabel Trancoso |
EAMT | 4 |
| 2011 | Discriminative Phrase-based Lexicalized Reordering Models using Weighted Reordering Graphs
Wang Ling, João Graça, David Martins de Matos, Isabel Trancoso, Alan W. Black |
IJCNLP | 1 |