VLDB 2026 Research / reviewers in the wild / expert
Aditi Chaudhary
dblp:225/7684
· DBLP profile ↗
10ranked-venue papers
5as first author
7since 2021 · last 2024
0009-0002-4658-4040ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Information extraction and text analysis · 30% Machine translation · 29% Trustworthy machine learning · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
cross-modal retrieval |
0.8 | 1 | 2024 | WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.8 | 1 | 2024 | WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models · NeurIPS 2024 |
Information retrieval
evaluation |
0.8 | 1 | 2024 | WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models · NeurIPS 2024 |
Information retrieval › evaluation › test collection
retrieval benchmark |
0.8 | 1 | 2024 | WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models · NeurIPS 2024 |
Natural language and speech › Information extraction and text analysis › named entity recognition
low-resource named entity recognition |
0.7 | 2 | 2019 | A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers · EMNLP/IJCNLP (1) 2019 Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.7 | 2 | 2019 | A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers · EMNLP/IJCNLP (1) 2019 Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations · EMNLP 2018 |
Natural language and speech › Machine translation
idiom translation |
0.7 | 1 | 2023 | Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss Weighting · EMNLP 2023 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.7 | 1 | 2023 | Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss Weighting · EMNLP 2023 |
Machine learning › Trustworthy machine learning › interpretability
attention analysis |
0.5 | 1 | 2021 | Do Context-Aware Translation Models Pay the Right Attention? · ACL/IJCNLP (1) 2021 |
Natural language and speech › Machine translation › neural machine translation
context-aware neural machine translation |
0.5 | 1 | 2021 | Do Context-Aware Translation Models Pay the Right Attention? · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › morphological analysis
morphosyntactic analysis |
0.5 | 1 | 2021 | Evaluating the Morphosyntactic Well-formedness of Generated Texts · EMNLP (1) 2021 |
Natural language and speech › Machine translation
neural machine translation |
0.5 | 1 | 2021 | Do Context-Aware Translation Models Pay the Right Attention? · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.5 | 1 | 2021 | Evaluating the Morphosyntactic Well-formedness of Generated Texts · EMNLP (1) 2021 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.4 | 1 | 2019 | A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Machine translation
low-resource machine translation |
0.1 | 1 | 2018 | Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
zero-shot evaluation · 1.5vision-language model fine-tuning · 1.5retrieval augmentation · 0.7loss weighting · 0.7morphosyntactic well-formedness metrics · 0.5description generation · 0.5contrastive learning · 0.5contextual machine translation · 0.5attention analysis · 0.5morphological analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language ModelsabstractCross-modal (image-to-text and text-to-image) retrieval is an established task used in evaluation benchmarks to test the performance of vision-language models (VLMs). Several state-of-the-art VLMs (e.g. CLIP, BLIP-2) have achieved near-perfect performance on widely-used image-text retrieval benchmarks such as MSCOCO-Test-5K and Flickr30K-Test-1K. As a measure of out-of-distribution (OOD) generalization, prior works rely on zero-shot performance evaluated on one dataset (Flickr) using a VLM finetuned on another one (MSCOCO). We argue that such comparisons are insufficient to assess the OOD generalization capability of models due to high visual and linguistic similarity between the evaluation and finetuning datasets. To address this gap, we introduce WikiDO (drawn from Wikipedia Diversity Observatory), a novel cross-modal retrieval benchmark to assess the OOD generalization capabilities of pretrained VLMs. This consists of newly scraped 380K image-text pairs from Wikipedia with domain labels, a carefully curated, human-verified a)in-distribution (ID) test set (3K) and b) OOD test set (3K). The image-text pairs are very diverse in topics and geographical locations. We evaluate different VLMs of varying capacity on the \wikido benchmark; BLIP-2 achieves zero-shot performance of $R@1\approx66\%$ on the OOD test set, compared to $\approx$ $81\%$ on COCO and $\approx95\%$ on Flickr. When fine-tuned on WikiDO, the $R@1$ improvement is at most $\approx5\%$ on OOD instances compared to $\approx12\%$ on ID instances. We probe the VLMs with varying finetuning objectives and datasets of varying sizes to identify what aids OOD generalization the most. Our results confirm that WikiDO offers a strong cross-modal benchmark for current VLMs in specifically evaluating for OOD generalization. Our benchmark is hosted as a competition at https://kaggle.com/competitions/wikido24 with public access to dataset and code. Tankala Pavan Kalyan, Piyush Singh Pasi, Sahil Dharod, Azeem Motiwala, Preethi Jyothi, Aditi Chaudhary, Krishna Srinivasan |
NeurIPS | 6 |
| 2023 | Salient Span Masking for Temporal UnderstandingabstractSalient Span Masking (SSM) has shown itself to be an effective strategy to improve closedbook question answering performance.SSM extends general masked language model pretraining by creating additional unsupervised training sentences that mask a single entity or date span, thus oversampling factual information.Despite the success of this paradigm, the span types and sampling strategies are relatively arbitrary and not widely studied for other tasks.Thus, we investigate SSM from the perspective of temporal tasks, where learning a good representation of various temporal expressions is important.To that end, we introduce Temporal Span Masking (TSM) intermediate training.First, we find that SSM alone improves the downstream performance on three temporal tasks by an avg.+5.8 points.Further, we are able to achieve additional improvements (avg.+0.29 points) by adding the TSM task.These comprise the new best reported results on the targeted tasks.Our analysis suggests that the effectiveness of SSM stems from the sentences chosen in the training data rather than the mask choice: sentences with entities frequently also contain temporal expressions.Nonetheless, the additional targeted spans of TSM can still improve performance, especially in a zero-shot context. Jeremy R. Cole, Aditi Chaudhary, Bhuwan Dhingra, Partha Talukdar |
EACL | 2 |
| 2023 | Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss WeightingabstractIdioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts.Despite significant advances, machine translation systems still struggle to translate idiomatic expressions.We provide a simple characterization of idiomatic translation and related issues.This allows us to conduct a synthetic experiment revealing a tipping point at which transformer-based machine translation models correctly default to idiomatic translations.To expand multilingual resources, we compile a dataset of ∼ 4k natural sentences containing idiomatic expressions in French, Finnish, and Japanese.To improve translation of natural idioms, we introduce two straightforward yet effective techniques: the strategic upweighting of training loss on potentially idiomatic sentences, and using retrievalaugmented models.This not only improves the accuracy of a strong pretrained MT model on idiomatic sentences by up to 13% in absolute accuracy, but also holds potential benefits for non-idiomatic sentences.1 Emmy Liu, Aditi Chaudhary, Graham Neubig |
EMNLP | 2 |
| 2021 | Do Context-Aware Translation Models Pay the Right Attention?abstractKayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig |
ACL/IJCNLP (1) | 4 |
| 2021 | When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical SelectionabstractLearning fine-grained distinctions between vocabulary items is a key challenge in learning a new language.For example, the noun "wall" has different lexical manifestations in Spanish -"pared" refers to an indoor wall while "muro" refers to an outside wall.However, this variety of lexical distinction may not be obvious to non-native learners unless the distinction is explained in such a way.In this work, we present a method for automatically identifying fine-grained lexical distinctions, and extracting concise descriptions explaining these distinctions in a human-and machine-readable format.We confirm the quality of these extracted descriptions in a language learning setup for two languages, Spanish and Greek, where we use them to teach non-native speakers when to translate a given ambiguous word into its different possible translations.Code and data are publicly released here.1 Aditi Chaudhary, Kayo Yin, Antonios Anastasopoulos, Graham Neubig |
EMNLP (1) | 1 |
| 2021 | Evaluating the Morphosyntactic Well-formedness of Generated TextsabstractAdithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov |
EMNLP (1) | 4 |
| 2021 | Reducing Confusion in Active Learning for Part-Of-Speech Tagging
Aditi Chaudhary, Zaid Sheikh, Antonios Anastasopoulos, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Automatic Extraction of Rules Governing Morphological AgreementabstractAditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig |
EMNLP (1) | 1 |
| 2019 | A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity RecognizersabstractAditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, Jaime Carbonell. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Aditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, Jaime G. Carbonell |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Adapting Word Embeddings to New Languages with Morphological and Phonological Subword RepresentationsabstractMuch work in Natural Language Processing (NLP) has been for resource-rich languages, making generalization to new, less-resourced languages challenging.We present two approaches for improving generalization to lowresourced languages by adapting continuous word representations using linguistically motivated subword units: phonemes, morphemes and graphemes.Our method requires neither parallel corpora nor bilingual dictionaries and provides a significant gain in performance over previous methods relying on these resources.We demonstrate the effectiveness of our approaches on Named Entity Recognition for four languages, namely Uyghur, Turkish, Bengali and Hindi, of which Uyghur and Bengali are low resource languages, and also perform experiments on Machine Translation.Exploiting subwords with transfer learning gives us a boost of +15.2 NER F1 for Uyghur and +9.7 F1 for Bengali.We also show improvements in the monolingual setting where we achieve (avg.)+3 F1 and (avg.)+1.35 BLEU. Aditi Chaudhary, Chunting Zhou, Lori S. Levin, Graham Neubig, David R. Mortensen, Jaime G. Carbonell |
EMNLP | 1 |