Wang Ling

dblp:91/7651 · DBLP profile ↗
← Back
30ranked-venue papers
11as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 11 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Representation and self-supervised learning · 30% Language models and text generation · 28% Information extraction and text analysis · 14%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 54% Compilers and program optimization · 46%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model
0.722019
Variational Smoothing in Recurrent Neural Network Language Models · ICLR (Poster) 2019
Memory Architectures in Recurrent Neural Network Language Models · ICLR (Poster) 2018
Natural language and speech › Language models and text generation › decoding
adaptive decoding
0.612022
Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022
Natural language and speech › Language models and text generation
decoding
0.612022
Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022
Natural language and speech › Language models and text generation
text generation
0.612022
Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.532015
Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015
Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces · ACL (1) 2015
Not All Contexts Are Created Equal: Better Word Representations with Variable Attention · EMNLP 2015
Machine learning › Representation and self-supervised learning › word representation
word representation learning
0.522016
Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning · ACL (1) 2016
Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces · ACL (1) 2015
Machine learning › Representation and self-supervised learning › representation learning
language representation learning
0.412020
A Mutual Information Maximization Perspective of Language Representation Learning · ICLR 2020
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412020
A Mutual Information Maximization Perspective of Language Representation Learning · ICLR 2020
Machine learning › Representation and self-supervised learning
word representation
0.422015
Not All Contexts Are Created Equal: Better Word Representations with Variable Attention · EMNLP 2015
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation · EMNLP 2015
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.412019
Variational Smoothing in Recurrent Neural Network Language Models · ICLR (Poster) 2019
Machine learning › Reinforcement learning
memory architectures
0.312018
Memory Architectures in Recurrent Neural Network Language Models · ICLR (Poster) 2018
Natural language and speech › Information extraction and text analysis
coreference resolution
0.312017
Reference-Aware Language Models · EMNLP 2017
Natural language and speech › Language models and text generation
language modeling
0.312017
Reference-Aware Language Models · EMNLP 2017
Natural language and speech › Question answering and dialogue systems
math word problem solving
0.312017
Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems · ACL (1) 2017
Machine learning › Reinforcement learning
reinforcement learning for NLP
0.312017
Learning to Compose Words into Sentences with Reinforcement Learning · ICLR (Poster) 2017
Program synthesis and code generation
inductive program synthesis
0.312017
Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems · ACL (1) 2017
Machine learning › Learning paradigms
curriculum learning
0.212016
Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning · ACL (1) 2016
Natural language and speech › Information extraction and text analysis
semantic parsing
0.212016
Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016
Natural language and speech › Information extraction and text analysis › semantic parsing
semi-supervised semantic parsing
0.212016
Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence transduction
0.212016
Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016
Machine learning › Representation and self-supervised learning › representation learning
sequential autoencoder
0.212016
Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016
Compilers and program optimization
code generation
0.212016
Latent Predictor Networks for Code Generation · ACL (1) 2016
Machine learning › Representation and self-supervised learning › word representation › word embedding
intrinsic evaluation
0.212015
Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015
Computer vision › Vision and language
open-vocabulary models
0.212015
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation · EMNLP 2015
Machine learning › Deep learning architectures and training
recurrent neural network
0.212015
Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
subspace alignment
0.212015
Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
subspace embedding
0.212015
Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces · ACL (1) 2015
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.212015
Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
transition-based dependency parsing
0.212015
Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015
Natural language and speech › Machine translation
parallel corpora
0.212013
Microblogs as Parallel Corpora · ACL (1) 2013

Methods — techniques the papers use, named apart from their topics

tree search · 0.6sequence-to-sequence · 0.6indirect supervision · 0.6mutual information estimation · 0.4contrastive learning · 0.4variational inference · 0.4smoothing · 0.4recurrent neural network · 0.3reinforcement learning · 0.3latent variable model · 0.3latent predictor networks · 0.2
YearPublicationVenuePosition
2023 Towards Dialogue Modeling Beyond Text
abstract
In this paper, we model aspects of communication beyond the words that are said. Specifically, we aim to detect interruptions and active listening events, which are important elements in any dialogue. We build a dataset with fine-grained annotations for each category and train multimodal models that take into account all channels in a digital conversation, that is, the video, the audio, and the text. Our experiments show that multimodality is a necessary component in modeling the complexity of the non-textual components of the conversation as different artifacts require different modalities to capture effectively.
Tongzi Wu, Wang Ling, Hojin Yang, Joana Veloso, Ruixin Huang, Norberto Guimaraes, Scott Sanner
ICASSP3
2022 Enabling Arbitrary Translation Objectives with Adaptive Tree Search
Wang Ling, Wojciech Stokowiec, Domenic Donato, Chris Dyer, Lei Yu 0008, Laurent Sartran, Austin Matthews
ICLR1
2020 A Mutual Information Maximization Perspective of Language Representation Learning
Lingpeng Kong, Cyprien de Masson d'Autume, Lei Yu 0008, Wang Ling, Zihang Dai, Dani Yogatama
ICLR4
2020 Better Document-Level Machine Translation with Bayes' Rule
abstract
We show that Bayes’ rule provides an effective mechanism for creating document translation models that can be learned from only parallel sentences and monolingual documents a compelling benefit because parallel documents are not always available. In our formulation, the posterior probability of a candidate translation is the product of the unconditional (prior) probability of the candidate output document and the “reverse translation probability” of translating the candidate output back into the source language. Our proposed model uses a powerful autoregressive language model as the prior on target language documents, but it assumes that each sentence is translated independently from the target to the source language. Crucially, at test time, when a source document is observed, the document language model prior induces dependencies between the translations of the source sentences in the posterior. The model’s independence assumption not only enables efficient use of available data, but it additionally admits a practical left-to-right beam-search algorithm for carrying out inference. Experiments show that our model benefits from using cross-sentence context in the language model, and it outperforms existing document translation approaches.
Lei Yu 0008, Laurent Sartran, Wojciech Stokowiec, Wang Ling, Lingpeng Kong, Phil Blunsom, Chris Dyer
Trans. Assoc. Comput. Linguistics4
2019 Variational Smoothing in Recurrent Neural Network Language Models
Lingpeng Kong, Gábor Melis, Wang Ling, Lei Yu 0008, Dani Yogatama
ICLR (Poster)3
2018 Memory Architectures in Recurrent Neural Network Language Models
Dani Yogatama, Yishu Miao, Gábor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, Phil Blunsom
ICLR (Poster)4
2017 Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
abstract
Solving algebraic word problems requires executing a series of arithmetic operations-a program-to obtain a final answer.However, since programs can be arbitrarily complicated, inducing them directly from question-answer pairs is a formidable challenge.To make this task more feasible, we solve these problems by generating answer rationales, sequences of natural language and human-readable mathematical expressions that derive the final answer through a series of small steps.Although rationales do not explicitly specify programs, they provide a scaffolding for their structure via intermediate milestones.To evaluate our approach, we have created a new 100,000-sample dataset of questions, answers and rationales.Experimental results show that indirect supervision of program learning via answer rationales is a promising strategy for inducing arithmetic programs.
Wang Ling, Dani Yogatama, Chris Dyer, Phil Blunsom
ACL (1)1
2017 Reference-Aware Language Models
abstract
We propose a general class of language models that treat reference as discrete stochastic latent variables.This decision allows for the creation of entity mentions by accessing external databases of referents (required by, e.g., dialogue generation) or past internal state (required to explicitly model coreferentiality).Beyond simple copying, our coreference model can additionally refer to a referent using varied mention forms (e.g., a reference to "Jane" can be realized as "she"), a characteristic feature of reference in natural languages.Experiments on three representative applications show our model variants outperform models based on deterministic attention and standard language modeling baselines.
Phil Blunsom, Chris Dyer, Wang Ling
EMNLP4
2017 Learning to Compose Words into Sentences with Reinforcement Learning
Dani Yogatama, Phil Blunsom, Chris Dyer, Edward Grefenstette, Wang Ling
ICLR (Poster)5
2016 Latent Predictor Networks for Code Generation
abstract
Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, Fumin Wang, Andrew Senior. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, Andrew W. Senior
ACL (1)1
2016 Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning
abstract
We use Bayesian optimization to learn curricula for word representation learning, optimizing performance on downstream tasks that depend on the learned representations as features.The curricula are modeled by a linear ranking function which is the scalar product of a learned weight vector and an engineered feature vector that characterizes the different aspects of the complexity of each instance in the training corpus.We show that learning the curriculum improves performance on a variety of downstream tasks over random orders and in comparison to the natural corpus order.
Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Brian MacWhinney, Chris Dyer
ACL (1)3
2016 Semantic Parsing with Semi-Supervised Sequential Autoencoders
abstract
We present a novel semi-supervised approach for sequence transduction and apply it to semantic parsing.The unsupervised component is based on a generative model in which latent sentences generate the unpaired logical forms.We apply this method to a number of semantic parsing tasks focusing on domains with limited access to labelled training data and extend those datasets with synthetically generated logical forms.
Tomás Kociský, Gábor Melis, Edward Grefenstette, Chris Dyer, Wang Ling, Phil Blunsom, Karl Moritz Hermann
EMNLP5
2016 Neural Network-Based Abstract Generation for Opinions and Arguments
abstract
We study the problem of generating abstractive summaries for opinionated text. We propose an attention-based neural network model that is able to absorb information from multiple text units to construct informative, concise, and fluent summaries. An importance-based sampling method is designed to allow the encoder to integrate information from an important subset of input. Automatic evaluation indicates that our system outperforms state-of-the-art abstractive and extractive summarization systems on two newly collected datasets of movie reviews and arguments. Our system summaries are also rated as more informative and grammatical in human evaluation.
Lu Wang 0008, Wang Ling
HLT-NAACL2
2016 Mining Parallel Corpora from Sina Weibo and Twitter
abstract
Microblogs such as Twitter, Facebook, and Sina Weibo (China's equivalent of Twitter) are a remarkable linguistic resource. In contrast to content from edited genres such as newswire, microblogs contain discussions of virtually every topic by numerous individuals in different languages and dialects and in different styles. In this work, we show that some microblog users post “self-translated” messages targeting audiences who speak different languages, either by writing the same message in multiple languages or by retweeting translations of their original posts in a second language. We introduce a method for finding and extracting this naturally occurring parallel data. Identifying the parallel content requires solving an alignment problem, and we give an optimally efficient dynamic programming algorithm for this. Using our method, we extract nearly 3M Chinese–English parallel segments from Sina Weibo using a targeted crawl of Weibo users who post in multiple languages. Additionally, from a random sample of Twitter, we obtain substantial amounts of parallel data in multiple language pairs. Evaluation is performed by assessing the accuracy of our extraction approach relative to a manual annotation as well as in terms of utility as training data for a Chinese–English machine translation system. Relative to traditional parallel data resources, the automatically extracted parallel data yield substantial translation quality improvements in translating microblog text and modest improvements in translating edited news content.
Wang Ling, Luís Marujo, Chris Dyer, Alan W. Black, Isabel Trancoso
Comput. Linguistics1
2016 Exploring events and distributed representations of text in multi-document summarization
Luís Marujo, Wang Ling, Ricardo Ribeiro 0001, Anatole Gershman, Jaime G. Carbonell, David Martins de Matos, João Paulo da Silva Neto
Knowl. Based Syst.2
2015 Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces
abstract
Ramon F. Astudillo, Silvio Amir, Wang Ling, Mário Silva, Isabel Trancoso. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ramón Fernandez Astudillo, Silvio Amir, Wang Ling, Mário J. Silva, Isabel Trancoso
ACL (1)3
2015 Transition-Based Dependency Parsing with Stack Long Short-Term Memory
abstract
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith
ACL (1)3
2015 Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
abstract
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, Tiago Luís
EMNLP1
2015 Not All Contexts Are Created Equal: Better Word Representations with Variable Attention
abstract
Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramón Fermandez, Chris Dyer, Alan W Black, Isabel Trancoso, Chu-Cheng Lin. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramon Fermandez, Chris Dyer, Alan W. Black, Isabel Trancoso, Chu-Cheng Lin
EMNLP1
2015 Evaluation of Word Vector Representations by Subspace Alignment
abstract
Unsupervisedly learned word vectors have proven to provide exceptionally effective features in many NLP tasks.Most common intrinsic evaluations of vector quality measure correlation with similarity judgments.However, these often correlate poorly with how well the learned representations perform as features in downstream evaluation tasks.We present QVEC-a computationally inexpensive intrinsic evaluation measure of the quality of word embeddings based on alignment to a matrix of features extracted from manually crafted lexical resources-that obtains strong correlation with performance of the vectors in a battery of downstream semantic evaluation tasks. 1
Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Guillaume Lample, Chris Dyer
EMNLP3
2015 Two/Too Simple Adaptations of Word2Vec for Syntax Problems
abstract
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso
HLT-NAACL1
2015 A linguistically motivated taxonomy for Machine Translation error analysis
Ângela Costa, Wang Ling, Tiago Luís, Rui Correia, Luísa Coheur
Mach. Transl.2
2014 Linguistic Evaluation of Support Verb Constructions by OpenLogos and Google Translate
Anabela Barreiro, Johanna Monti, Brigitte Orliac, Susanne Preuß, Kutz Arrieta, Wang Ling, Fernando Batista, Isabel Trancoso
LREC6
2014 Dual Subtitles as Parallel Corpora
Shikun Zhang, Wang Ling, Chris Dyer
LREC2
2013 Microblogs as Parallel Corpora
Wang Ling, Guang Xiang, Chris Dyer, Alan W. Black, Isabel Trancoso
ACL (1)1
2013 Paraphrasing 4 Microblog Normalization
abstract
Compared to the edited genres that have played a central role in NLP research, microblog texts use a more informal register with nonstandard lexical items, abbreviations, and free orthographic variation.When confronted with such input, conventional text analysis tools often perform poorly.Normalization -replacing orthographically or lexically idiosyncratic forms with more standard variants -can improve performance.We propose a method for learning normalization rules from machine translations of a parallel corpus of microblog messages.To validate the utility of our approach, we evaluate extrinsically, showing that normalizing English tweets and then translating improves translation quality (compared to translating unnormalized text) using three standard web translation services as well as a phrase-based translation system trained on parallel microblog data.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso
EMNLP1
2012 Overview of Computer-assisted Language Learning for European Portuguese at L2f
Thomas Pellegrini, Wang Ling, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede
CSEDU (2)2
2012 Entropy-based Pruning for Phrase-based Machine Translation
Wang Ling, João Graça, Isabel Trancoso, Alan W. Black
EMNLP-CoNLL1
2011 BP2EP - Adaptation of Brazilian Portuguese texts to European Portuguese
Luís Marujo, Nuno Grazina, Tiago Luís, Wang Ling, Luísa Coheur, Isabel Trancoso
EAMT4
2011 Discriminative Phrase-based Lexicalized Reordering Models using Weighted Reordering Graphs
Wang Ling, João Graça, David Martins de Matos, Isabel Trancoso, Alan W. Black
IJCNLP1