VLDB 2026 Research / reviewers in the wild / expert
Daniel M. Cer
dblp:16/6461 · also Daniel Cer
· DBLP profile ↗
15ranked-venue papers
2as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Representation and self-supervised learning · 36% Transfer learning and domain adaptation · 27% Language models and text generation · 20% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
1.8 | 4 | 2022 | Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation · EMNLP 2022 Language-agnostic BERT Sentence Embedding · ACL (1) 2022 A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning › word representation
multilingual word embedding |
1.0 | 2 | 2022 | Language-agnostic BERT Sentence Embedding · ACL (1) 2022 Improving Multilingual Sentence Embedding using Bi-directional Dual Encoder with Additive Margin Softmax · IJCAI 2019 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
1.0 | 2 | 2022 | Language-agnostic BERT Sentence Embedding · ACL (1) 2022 Improving Multilingual Sentence Embedding using Bi-directional Dual Encoder with Additive Margin Softmax · IJCAI 2019 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.6 | 1 | 2022 | SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer · ACL (1) 2022 |
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
prompt-based transfer |
0.6 | 1 | 2022 | SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer · ACL (1) 2022 |
Natural language and speech › Language models and text generation
prompt tuning |
0.6 | 1 | 2022 | SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer · ACL (1) 2022 |
Natural language and speech › Language models and text generation › masked language modeling
conditional masked language model |
0.5 | 1 | 2021 | Universal Sentence Representation Learning with Conditional Masked Language Model · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning |
0.5 | 1 | 2021 | A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations · EMNLP (1) 2021 |
Machine learning › Trustworthy machine learning › debiasing
language bias mitigation |
0.5 | 1 | 2021 | A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
masked language modeling |
0.5 | 1 | 2021 | Universal Sentence Representation Learning with Conditional Masked Language Model · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
sentence representation learning |
0.5 | 1 | 2021 | Universal Sentence Representation Learning with Conditional Masked Language Model · EMNLP (1) 2021 |
Information retrieval
cross-language information retrieval |
0.4 | 1 | 2019 | Improving Multilingual Sentence Embedding using Bi-directional Dual Encoder with Additive Margin Softmax · IJCAI 2019 |
Natural language and speech › Language models and text generation
text summarization |
0.2 | 1 | 2022 | Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
bilingual word embedding |
0.2 | 1 | 2013 | Bilingual Word Embeddings for Phrase-Based Machine Translation · EMNLP 2013 |
Machine learning › Learning theory
online learning |
0.2 | 1 | 2013 | Fast and Adaptive Online Training of Feature-Rich Translation Models · ACL (1) 2013 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.2 | 1 | 2013 | Bilingual Word Embeddings for Phrase-Based Machine Translation · EMNLP 2013 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 1 | 2013 | Bilingual Word Embeddings for Phrase-Based Machine Translation · EMNLP 2013 |
Natural language and speech › Question answering and dialogue systems › multilingual question answering
cross-lingual question answering |
0.1 | 1 | 2021 | A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
dual encoder · 1.3translation language modeling · 0.6prompt tuning · 0.6parameter-efficient adaptation · 0.6masked language modeling · 0.6BERT · 0.6orthogonal projection · 0.5matrix factorization · 0.5geometric algebra · 0.5bitext retrieval · 0.5neural machine translation · 0.4additive margin softmax · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense RetrievalabstractNandan Thakur, Jianmo Ni, Gustavo Hernandez Abrego, John Wieting, Jimmy Lin, Daniel Cer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nandan Thakur, Jianmo Ni, Gustavo Hernández Ábrego, John Wieting, Jimmy Lin, Daniel M. Cer |
NAACL-HLT | 6 |
| 2022 | Language-agnostic BERT Sentence EmbeddingabstractWhile BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings have yet to be explored.We systematically investigate methods for learning multilingual sentence embeddings by combining the best methods for learning monolingual and cross-lingual representations including: masked language modeling (MLM), translation language modeling (TLM) (Conneau and Lample, 2019), dual encoder translation ranking (Guo et al., 2018), and additive margin softmax (Yang et al., 2019a).We show that introducing a pre-trained multilingual language model dramatically reduces the amount of parallel training data required to achieve good performance by 80%.Composing the best of these methods produces a model that achieves 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba, well above the 65.5% achieved by Artetxe and Schwenk (2019b), while still performing competitively on monolingual transfer learning benchmarks (Conneau and Kiela, 2018).Parallel data mined from CommonCrawl using our best model is shown to train competitive NMT models for en-zh and en-de.We publicly release our best multilingual sentence embedding model for 109+ languages at https://tfhub.dev/ google/LaBSE. Fangxiaoyu Feng, Yinfei Yang, Daniel M. Cer, Naveen Arivazhagan, Wei Wang 0236 |
ACL (1) | 3 |
| 2022 | SPoT: Better Frozen Model Adaptation through Soft Prompt TransferabstractThere has been growing interest in parameter-efficient methods to apply pre-trained language models to downstream tasks. Building on the Prompt Tuning approach of Lester et al. (2021), which learns task-specific soft prompts to condition a frozen pre-trained model to perform different tasks, we propose a novel prompt-based transfer learning approach called SPoT: Soft Prompt Transfer. SPoT first learns a prompt on one or more source tasks and then uses it to initialize the prompt for a target task. We show that SPoT significantly boosts the performance of Prompt Tuning across many tasks. More remarkably, across all model sizes, SPoT matches or outperforms standard Model Tuning (which fine-tunes all model parameters) on the SuperGLUE benchmark, while using up to 27,000× fewer task-specific parameters. To understand where SPoT is most effective, we conduct a large-scale study on task transferability with 26 NLP tasks in 160 combinations, and demonstrate that many tasks can benefit each other via prompt transfer. Finally, we propose an efficient retrieval approach that interprets task prompts as task embeddings to identify similar tasks and predict the most transferable source tasks for a novel target task. Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou, Daniel M. Cer |
ACL (1) | 5 |
| 2022 | Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual GenerationabstractIn this paper, we explore the challenging problem of performing a generative task in a target language when labeled data is only available in English, using summarization as a case study.We assume a strict setting with no access to parallel data or machine translation and find that common transfer learning approaches struggle in this setting, as a generative multilingual model fine-tuned purely on English catastrophically forgets how to generate non-English.Given the recent rise of parameter-efficient adaptation techniques, we conduct the first investigation into how one such method, prompt tuning (Lester et al., 2021), can overcome catastrophic forgetting to enable zero-shot cross-lingual generation.Our experiments show that parameter-efficient prompt tuning provides gains over standard fine-tuning when transferring between lessrelated languages, e.g., from English to Thai.However, a significant gap still remains between these methods and fully-supervised baselines.To improve cross-lingual transfer further, we explore several approaches, including: (1) mixing in unlabeled multilingual data, and (2) explicitly factoring prompts into recombinable language and task components.Our approaches can provide further quality gains, suggesting that robust zero-shot crosslingual generation is within reach.ING and standard MODELTUNING for zero-shot cross-lingual generation (XGEN).We show that increasing model scale and decreasing tunable parameter capacity are key for overcoming catastrophic forgetting on XGEN.• We propose WIKILINGUA-0, a challenging XGEN benchmark and an associated SP-ROUGE evaluation metric, which we hope will facilitate future work evaluating multilingual summarization.• We show that mixing in unsupervised multilingual data can boost XGEN performance, and are the first to combine this approach with PROMPTTUNING.• We propose "factorized prompts", a novel approach that can also help PROMPTTUNING overcome severe catastrophic forgetting.• To facilitate future work, we release our data, pretrained models Tu Vu, Aditya Barua, Brian Lester, Daniel M. Cer, Mohit Iyyer, Noah Constant |
EMNLP | 4 |
| 2021 | Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCOabstractBy supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning.Unfortunately, datasets have limited cross-modal associations: images are not paired with other images, captions are only paired with other captions of the same image, there are no negative associations and there are missing positive cross-modal associations.This undermines research into how inter-modality learning impacts intra-modality tasks.We address this gap with Crisscrossed Captions (CxC), an extension of the MS-COCO dataset with human semantic similarity judgments for 267,095 intra-and intermodality pairs.We report baseline results on CxC for strong existing unimodal and multimodal models.We also evaluate a multitask dual encoder trained on both image-caption and caption-caption pairs that crucially demonstrates CxC's value for measuring the influence of intra-and inter-modality learning. Zarana Parekh, Jason Baldridge, Daniel M. Cer, Austin Waters, Yinfei Yang |
EACL | 3 |
| 2021 | A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual RepresentationsabstractLanguage agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models.We explore this problem from a novel angle of geometric algebra and semantic space.A simple but highly effective method "Language Information Removal (LIR)" factors out language identity information from semantic related components in multilingual representations pre-trained on multi-monolingual data.A post-training and model-agnostic method, LIR only uses simple linear operations, e.g.matrix factorization and orthogonal projection.LIR reveals that for weak-alignment multilingual systems, the principal components of semantic spaces primarily encodes language identity information.We first evaluate the LIR on a cross-lingual question answer retrieval task (LAReQA), which requires the strong alignment for the multilingual embedding space.Experiment shows that LIR is highly effectively on this task, yielding almost 100% relative improvement in MAP for weakalignment models.We then evaluate the LIR on Amazon Reviews and XEVAL dataset, with the observation that removing language information is able to improve the cross-lingual transfer performance. Ziyi Yang 0011, Yinfei Yang, Daniel M. Cer, Eric Darve |
EMNLP (1) | 3 |
| 2021 | Universal Sentence Representation Learning with Conditional Masked Language ModelabstractThis paper presents a novel training method, Conditional Masked Language Modeling (CMLM), to effectively learn sentence representations on large scale unlabeled corpora.CMLM integrates sentence representation learning into MLM training by conditioning on the encoded vectors of adjacent sentences.Our English CMLM model achieves state-ofthe-art performance on SentEval (Conneau and Kiela, 2018), even outperforming models learned using supervised signals.As a fully unsupervised learning method, CMLM can be conveniently extended to a broad range of languages and domains.We find that a multilingual CMLM model co-trained with bitext retrieval (BR) and natural language inference (NLI) tasks outperforms the previous state-of-the-art multilingual models by a large margin, e.g.10% improvement upon baseline models on cross-lingual semantic search.We explore the same language bias of the learned representations, and propose a simple, post-training and model agnostic approach to remove the language identifying information from the representation while still retaining sentence semantics. Ziyi Yang 0011, Yinfei Yang, Daniel M. Cer, Jax Law, Eric Darve |
EMNLP (1) | 3 |
| 2019 | Improving Multilingual Sentence Embedding using Bi-directional Dual Encoder with Additive Margin SoftmaxabstractIn this paper, we present an approach to learn multilingual sentence embeddings using a bi-directional dual-encoder with additive margin softmax. The embeddings are able to achieve state-of-the-art results on the United Nations (UN) parallel corpus retrieval task. In all the languages tested, the system achieves P@1 of 86% or higher. We use pairs retrieved by our approach to train NMT models that achieve similar performance to models trained on gold pairs. We explore simple document-level embeddings constructed by averaging our sentence embeddings. On the UN document-level retrieval task, document embeddings achieve around 97% on P@1 for all experimented language pairs. Lastly, we evaluate the proposed model on the BUCC mining task. The learned embeddings with raw cosine similarity scores achieve competitive results compared to current state-of-the-art models, and with a second-stage scorer we achieve a new state-of-the-art level on this task. Yinfei Yang, Gustavo Hernández Ábrego, Steve Yuan, Mandy Guo, Qinlan Shen, Daniel M. Cer, Yun-Hsuan Sung, Brian Strope, Raymond Kurzweil |
IJCAI | 6 |
| 2013 | Fast and Adaptive Online Training of Feature-Rich Translation Models
Spence Green, Sida I. Wang, Daniel M. Cer, Christopher D. Manning |
ACL (1) | 3 |
| 2013 | Bilingual Word Embeddings for Phrase-Based Machine TranslationabstractWe introduce bilingual word embeddings: semantic embeddings associated across two languages in the context of neural language models.We propose a method to learn bilingual embeddings from a large unlabeled corpus, while utilizing MT word alignments to constrain translational equivalence.The new embeddings significantly out-perform baselines in word semantic similarity.A single semantic similarity feature induced with bilingual embeddings adds near half a BLEU point to the results of NIST08 Chinese-English machine translation task. Will Y. Zou, Richard Socher, Daniel M. Cer, Christopher D. Manning |
EMNLP | 3 |
| 2010 | Parsing to Stanford Dependencies: Trade-offs between Speed and Accuracy
Daniel M. Cer, Marie-Catherine de Marneffe, Daniel Jurafsky, Christopher D. Manning |
LREC | 1 |
| 2010 | The Best Lexical Metric for Phrase-Based Statistical MT System Optimization
Daniel M. Cer, Christopher D. Manning, Daniel Jurafsky |
HLT-NAACL | 1 |
| 2009 | Measuring machine translation quality as semantic equivalence: A metric based on entailment features
Sebastian Padó, Daniel M. Cer, Michel Galley, Daniel Jurafsky, Christopher D. Manning |
Mach. Transl. | 2 |
| 2006 | Learning to recognize features of valid textual entailments
Bill MacCartney, Trond Grenager, Marie-Catherine de Marneffe, Daniel M. Cer, Christopher D. Manning |
HLT-NAACL | 4 |
| 2005 | The detection of emphatic words using acoustic and lexical featuresabstractIn this study, we describe an automatic detector for prosodically salient or emphasized words in speech. Knowledge of whether a word is emphatic or not could improve Text-to-Speech synthesis as well as spoken language summarization. Previous work on emphasis detection has focused on the automatic recognition of pitch accents. Our model extends earlier research by automatically identifying emphatic pitch accents, a subset of pitch accents that mark special discourse functions with extreme degrees of salience. The overall best performance achieved by our system was 87.8 % correct, 8.0 % above baseline performance. The results of a feature selection algorithm show that the top-performing features in our models are primarily acoustic measures. Our work identifies important cues for emphasis in speech and shows that it is possible for an automated system to distinguish between two levels of perceived prominence in pitch accents with a high degree of accuracy. 1. Jason M. Brenier, Daniel M. Cer, Daniel Jurafsky |
INTERSPEECH | 2 |