VLDB 2026 Research / reviewers in the wild / expert
Anna Currey
dblp:185/5552
· DBLP profile ↗
8ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Machine translation · 37% Efficient and distributed learning · 28% Language models and text generation · 20% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 87% Recommender systems · 13% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
1.4 | 3 | 2023 | Pseudo-label Training and Model Inertia in Neural Machine Translation · ICLR 2023 Distilling Multiple Domains for Neural Machine Translation · EMNLP (1) 2020 Multi-Source Syntactic Neural Machine Translation · EMNLP 2018 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.9 | 1 | 2025 | Effective post-training embedding compression via temperature control in contrastive training · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
embedding compression |
0.9 | 1 | 2025 | Effective post-training embedding compression via temperature control in contrastive training · ICLR 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025 |
Natural language and speech › Language models and text generation
memory augmentation |
0.9 | 1 | 2025 | MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Effective post-training embedding compression via temperature control in contrastive training · ICLR 2025 |
Information retrieval › retrieval-augmented generation
memory retrieval |
0.9 | 1 | 2025 | MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025 |
Information retrieval
retrieval-augmented generation |
0.9 | 1 | 2025 | MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025 |
Natural language and speech › Machine translation
machine translation evaluation |
0.6 | 1 | 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation · EMNLP 2022 |
Natural language and speech › Machine translation
multi-domain neural machine translation |
0.4 | 1 | 2020 | Distilling Multiple Domains for Neural Machine Translation · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.3 | 2 | 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation · EMNLP 2022 GFST: Gender-Filtered Self-Training for More Accurate Gender in Translation · EMNLP (1) 2021 |
Recommender systems › interactive recommendation
conversational recommendation |
0.3 | 1 | 2025 | MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025 |
Natural language and speech › Machine translation
gender bias in machine translation |
0.2 | 1 | 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation · EMNLP 2022 |
Natural language and speech › Machine translation
domain adaptation for machine translation |
0.1 | 1 | 2020 | Distilling Multiple Domains for Neural Machine Translation · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 1 | 2018 | Multi-Source Syntactic Neural Machine Translation · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
autonomous memory augmentation · 1.7contrastive learning · 0.9pseudo-labeling · 0.7counterfactual evaluation · 0.6contextual evaluation · 0.6self-training · 0.5pseudo parallel corpora · 0.5back-translation · 0.5multi-domain training · 0.4knowledge distillation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MemInsight: Autonomous Memory Augmentation for LLM AgentsabstractLarge language model (LLM) agents have evolved to intelligently process information, make decisions, and interact with users or tools. A key capability is the integration of long-term memory capabilities, enabling these agents to draw upon historical interactions and knowledge. However, the growing memory size and need for semantic structuring pose significant challenges. In this work, we propose an autonomous memory augmentation approach, MemInsight, to enhance semantic data representation and retrieval mechanisms. By leveraging autonomous augmentation to historical interactions, LLM agents are shown to deliver more accurate and contextualized responses. We empirically validate the efficacy of our proposed approach in three task scenarios; conversational recommendation, question answering and event summarization. On the LLM-REDIAL dataset, MemInsight boosts persuasiveness of recommendations by up to 14%. Moreover, it outperforms a RAG baseline by 34% in recall for LoCoMo retrieval. Our empirical results show the potential of MemInsight to enhance the contextual performance of LLM agents across multiple tasks. Rana Salama, Jason Cai, Michelle Yuan, Anna Currey, Monica Sunkara, Yassine Benajiba |
EMNLP | 4 |
| 2025 | Effective post-training embedding compression via temperature control in contrastive trainingabstractFixed-size learned representations (dense representations, or embeddings) are widely used in many machine learning applications across language, vision or speech modalities. This paper investigates the role of the temperature parameter in contrastive training for text embeddings. We shed light on the impact this parameter has on the intrinsic dimensionality of the embedding spaces obtained, and show that lower intrinsic dimensionality is further correlated with effective compression of embeddings. We still observe a trade-off between absolute performance and effective compression and we propose temperature aggregation methods which reduce embedding size by an order of magnitude with minimal impact on quality. Georgiana Dinu, Corey D. Barrett, Miguel Romero Calvo, Anna Currey, Xing Niu 0001 |
ICLR | 5 |
| 2023 | Pseudo-label Training and Model Inertia in Neural Machine Translation
Benjamin Hsu, Anna Currey, Xing Niu 0001, Maria Nadejde, Georgiana Dinu |
ICLR | 2 |
| 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationabstractAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu, Benjamin Hsu, Georgiana Dinu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Anna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu 0001, Benjamin Hsu, Georgiana Dinu |
EMNLP | 1 |
| 2021 | GFST: Gender-Filtered Self-Training for More Accurate Gender in TranslationabstractTargeted evaluations have found that machine translation systems often output incorrect gender in translations, even when the gender is clear from context.Furthermore, these incorrectly gendered translations have the potential to reflect or amplify social biases.We propose gender-filtered self-training (GFST) to improve gender translation accuracy on unambiguously gendered inputs.Our GFST approach uses a source monolingual corpus and an initial model to generate gender-specific pseudo-parallel corpora which are then filtered and added to the training data.We evaluate GFST on translation from English into five languages, finding that it improves gender accuracy without damaging generic quality.We also show the viability of GFST on several experimental settings, including re-training from scratch, fine-tuning, controlling the gender balance of the data, forward translation, and back-translation. 1 Prafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana Dinu |
EMNLP (1) | 2 |
| 2020 | Distilling Multiple Domains for Neural Machine TranslationabstractNeural machine translation achieves impressive results in high-resource conditions, but performance often suffers when the input domain is low-resource.The standard practice of adapting a separate model for each domain of interest does not scale well in practice from both a quality perspective (brittleness under domain shift) as well as a cost perspective (added maintenance and inference complexity).In this paper, we propose a framework for training a single multi-domain neural machine translation model that is able to translate several domains without increasing inference time or memory usage.We show that this model can improve translation on both highand low-resource domains over strong multidomain baselines.In addition, our proposed model is effective when domain labels are unknown during training, as well as robust under noisy data conditions. Anna Currey, Prashant Mathur, Georgiana Dinu |
EMNLP (1) | 1 |
| 2018 | Multi-Source Syntactic Neural Machine TranslationabstractWe introduce a novel multi-source technique for incorporating source syntax into neural machine translation using linearized parses.This is achieved by employing separate encoders for the sequential and parsed versions of the same source sentence; the resulting representations are then combined using a hierarchical attention mechanism.The proposed model improves over both seq2seq and parsed baselines by over 1 BLEU on the WMT17 English→German task.Further analysis shows that our multi-source syntactic model is able to translate successfully without any parsed input, unlike standard parsed methods.In addition, performance does not deteriorate as much on long sentences as for the baselines. Anna Currey, Kenneth Heafield |
EMNLP | 1 |
| 2016 | Dynamic adjustment of language models for automatic speech recognition using word similarityabstractOut-of-vocabulary (OOV) words can pose a particular problem for automatic speech recognition (ASR) of broadcast news. The language models (LMs) of ASR systems are typically trained on static corpora, whereas new words (particularly new proper nouns) are continually introduced in the media. Additionally, such OOVs are often content-rich proper nouns that are vital to understanding the topic. In this work, we explore methods for dynamically adding OOVs to language models by adapting the n-gram language model used in our ASR system. We propose two strategies: the first relies on finding in-vocabulary (IV) words similar to the OOVs, where word embeddings are used to define similarity. Our second strategy leverages a small contemporary corpus to estimate OOV probabilities. The models we propose yield improvements in perplexity over the baseline; in addition, the corpus-based approach leads to a significant decrease in proper noun error rate over the baseline in recognition experiments. Anna Currey, Irina Illina, Dominique Fohr |
SLT | 1 |