VLDB 2026 Research / reviewers in the wild / expert
Ajay Nagesh
dblp:117/4073
· DBLP profile ↗
9ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Information extraction and text analysis · 74% Machine translation · 20% Knowledge representation and reasoning · 6% | |
| Human-computer interaction and pervasive computing
1 paper |
Learning and educational technologies · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › dataset construction
corpus filtering |
0.4 | 1 | 2020 | Parallel Corpus Filtering via Pre-trained Language Models · ACL 2020 |
Natural language and speech › Information extraction and text analysis › bootstrapping
bootstrapping for information extraction |
0.3 | 1 | 2018 | Visual Supervision in Bootstrapped Information Extraction · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › named entity recognition
named entity classification |
0.3 | 1 | 2018 | Visual Supervision in Bootstrapped Information Extraction · EMNLP 2018 |
Learning and educational technologies
active learning |
0.3 | 1 | 2018 | Visual Supervision in Bootstrapped Information Extraction · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.2 | 1 | 2014 | Noisy Or-based model for Relation Extraction using Distant Supervision · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.2 | 1 | 2014 | Noisy Or-based model for Relation Extraction using Distant Supervision · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2012 | Towards Efficient Named-Entity Rule Induction for Customizability · EMNLP-CoNLL 2012 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
rule learning |
0.1 | 1 | 2012 | Towards Efficient Named-Entity Rule Induction for Customizability · EMNLP-CoNLL 2012 |
Methods — techniques the papers use, named apart from their topics
scatterplot interface · 0.7embedding-based bootstrapping · 0.7pre-trained language model · 0.4GPT · 0.4BERT · 0.4soft constraints · 0.2noisy-OR model · 0.2integer linear programming · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Parallel Corpus Filtering via Pre-trained Language ModelsabstractWeb-crawled data provides a good source of parallel corpora for training machine translation models.It is automatically obtained, but extremely noisy, and recent work shows that neural machine translation systems are more sensitive to noise than traditional statistical machine translation methods.In this paper, we propose a novel approach to filter out noisy sentence pairs from web-crawled corpora via pre-trained language models.We measure sentence parallelism by leveraging the multilingual capability of BERT and use the Generative Pre-training (GPT) language model as a domain filter to balance data domains.We evaluate the proposed method on the WMT 2018 Parallel Corpus Filtering shared task, and on our own web-crawled Japanese-Chinese parallel corpus.Our method significantly outperforms baselines and achieves a new stateof-the-art.In an unsupervised setting, our method achieves comparable performance to the top-1 supervised method.We also evaluate on a web-crawled Japanese-Chinese parallel corpus that we make publicly available. Boliang Zhang, Ajay Nagesh, Kevin Knight |
ACL | 2 |
| 2018 | An Exploration of Three Lightly-supervised Representation Learning Approaches for Named Entity ClassificationabstractSeveral semi-supervised representation learning methods have been proposed recently that mitigate the drawbacks of traditional bootstrapping: they reduce the amount of semantic drift introduced by iterative approaches through one-shot learning; others address the sparsity of data through the learning of custom, dense representation for the information modeled. In this work, we are the first to adapt three of these methods, most of which have been originally proposed for image processing, to an information extraction task, specifically, named entity classification. Further, we perform a rigorous comparative analysis on two distinct datasets. Our analysis yields several important observations. First, all representation learning methods outperform state-of-the-art semi-supervised methods that do not rely on representation learning. To the best of our knowledge, we report the latest state-of-the-art results on the semi-supervised named entity classification task. Second, one-shot learning methods clearly outperform iterative representation learning approaches. Lastly, one of the best performers relies on the mean teacher framework (Tarvainen and Valpola, 2017), a simple teacher/student approach that is independent of the underlying task-specific model. Ajay Nagesh, Mihai Surdeanu |
COLING | 1 |
| 2018 | Visual Supervision in Bootstrapped Information ExtractionabstractWe challenge a common assumption in active learning, that a list-based interface populated by informative samples provides for efficient and effective data annotation.We show how a 2D scatterplot populated with diverse and representative samples can yield improved models given the same time budget.We consider this for bootstrapping-based information extraction, in particular named entity classification, where human and machine jointly label data.To enable effective data annotation in a scatterplot, we have developed an embeddingbased bootstrapping model that learns the distributional similarity of entities through the patterns that match them in a large data corpus, while being discriminative with respect to human-labeled and machine-promoted entities.We conducted a user study to assess the effectiveness of these different interfaces, and analyze bootstrapping performance in terms of human labeling accuracy, label quantity, and labeling consensus across multiple users.Our results suggest that supervision acquired from the scatterplot interface, despite being noisier, yields improvements in classification performance compared with the list interface, due to a larger quantity of supervision acquired. Matthew Berger, Ajay Nagesh, Joshua A. Levine, Mihai Surdeanu, Helen Zhang |
EMNLP | 2 |
| 2018 | Grounding Gradable Adjectives through Crowdsourcing
Rebecca Sharp, Mithun Paul, Ajay Nagesh, Dane Bell, Mihai Surdeanu |
LREC | 3 |
| 2015 | Optimizing Multivariate Performance Measures for Learning Relation Extraction ModelsabstractGholamreza Haffari, Ajay Nagesh, Ganesh Ramakrishnan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Gholamreza Haffari, Ajay Nagesh, Ganesh Ramakrishnan |
HLT-NAACL | 2 |
| 2015 | Exploring Relational Features and Learning under Distant Supervision for Information Extraction TasksabstractInformation Extraction (IE) has become an indispensable tool in our quest to handle the data deluge of the information age.IE can broadly be classified into Named-entity Recognition (NER) and Relation Extraction (RE).In this thesis, we view the task of IE as finding patterns in unstructured data, which can either take the form of features and/or be specified by constraints.In NER, we study the categorization of complex relational 1 features and outline methods to learn feature combinations through induction.We demonstrate the efficacy of induction techniques in learning : i) rules for the identification of named entities in text -the novelty is the application of induction techniques to learn in a very expressive declarative rule language ii) a richer sequence labeling model -enabling optimal learning of discriminative features.In RE, our investigations are in the paradigm of distant supervision, which facilitates the creation of large albeit noisy training data.We devise an inference framework in which constraints can be easily specified in learning relation extractors.In addition, we reformulate the learning objective in a max-margin framework.To the best of our knowledge, our formulation is the first to optimize multi-variate non-linear performance measures such as F β for a latent variable structure prediction task. Ajay Nagesh |
HLT-NAACL | 1 |
| 2014 | Noisy Or-based model for Relation Extraction using Distant SupervisionabstractDistant supervision, a paradigm of relation extraction where training data is created by aligning facts in a database with a large unannotated corpus, is an attractive approach for training relation extractors.Various models are proposed in recent literature to align the facts in the database to their mentions in the corpus.In this paper, we discuss and critically analyse a popular alignment strategy called the "at least one" heuristic.We provide a simple, yet effective relaxation to this strategy.We formulate the inference procedures in training as integer linear programming (ILP) problems and implement the relaxation to the "at least one " heuristic via a soft constraint in this formulation.Empirically, we demonstrate that this simple strategy leads to a better performance under certain settings over the existing approaches. Ajay Nagesh, Gholamreza Haffari, Ganesh Ramakrishnan |
EMNLP | 1 |
| 2012 | Towards Efficient Named-Entity Rule Induction for Customizability
Ajay Nagesh, Ganesh Ramakrishnan, Laura Chiticariu, Rajasekar Krishnamurthy, Ankush Dharkar, Pushpak Bhattacharyya |
EMNLP-CoNLL | 1 |
| 2012 | Probing the Space of Optimal Markov Logic Networks for Sequence Labeling
Naveen Nair, Ajay Nagesh, Ganesh Ramakrishnan |
ILP | 2 |