Markus Dreyer

dblp:37/4227 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0002-1117-1070ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
abstract
Yukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra, Markus Dreyer, Venkatesh Saligrama. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra, Markus Dreyer, Venkatesh Saligrama
ACL (1)5
2024 REFINESUMM: Self-Refining MLLM for Generating a Multimodal Summarization Dataset
abstract
Multimodal Large Language Models (MLLMs) excel at synthesizing key information from diverse sources.However, generating accurate and faithful multimodal summaries is challenging, primarily due to the lack of appropriate multimodal datasets for fine-tuning that meaningfully integrate textual and visual modalities.To address this gap, we present a new dataset specifically designed for image-text multimodal summarization, harnessing the capabilities of state-of-the-art MLLMs.We generate summaries from Wikipedia sections and corresponding images and evaluate them across text-based, visual and multimodal dimensions, employing reference-free metrics.To refine the dataset, we: (1) filter the MLLM-generated summaries by training a critic model on human annotations and using its predictions to remove low-quality summaries; (2) fine-tune the MLLM with the filtered high-quality summaries; (3) use the fine-tuned model in turn to regenerate the summaries.This self-refinement process notably improves summary quality, as measured by human judgments and automatic multimodal metrics, resulting in a valuable dataset for multimodal summarization research.1 * Work done as an intern at Amazon AGI. 1 The dataset is publicly available at https://github. com/amazon-science/refinesumm.The Italian wall lizard or ruin lizard (Podarcis siculus, from the Greek meaning agile and feet) is a species of lizard in the family Lacertidae.P. siculus is native to Bosnia and Herzegovina, Croatia, France, Italy, Serbia, Montenegro, Slovenia, and Switzerland, but has also been introduced to Spain, ....
Vaidehi Patil, Leonardo F. R. Ribeiro, Mengwen Liu, Mohit Bansal, Markus Dreyer
ACL (1)5
2024 CCSum: A Large-Scale and High-Quality Dataset for Abstractive News Summarization
abstract
Xiang Jiang, Markus Dreyer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xiang Jiang 0001, Markus Dreyer
NAACL-HLT2
2023 Faithfulness-Aware Decoding Strategies for Abstractive Summarization
abstract
Despite significant progress in understanding and improving faithfulness in abstractive summarization, the question of how decoding strategies affect faithfulness is less studied.We present a systematic study of the effect of generation techniques such as beam search and nucleus sampling on faithfulness in abstractive summarization.We find a consistent trend where beam search with large beam sizes produces the most faithful summaries while nucleus sampling generates the least faithful ones.We propose two faithfulness-aware generation methods to further improve faithfulness over current generation techniques: (1) ranking candidates generated by beam search using automatic faithfulness metrics and (2) incorporating lookahead heuristics that produce a faithfulness score on the future summary.We show that both generation methods significantly improve faithfulness across two datasets as evaluated by four automatic faithfulness metrics and human evaluation.To reduce computational cost, we demonstrate a simple distillation approach that allows the model to generate faithful summaries with just greedy decoding.1
David Wan, Mengwen Liu, Kathy McKeown, Markus Dreyer, Mohit Bansal
EACL4
2023 Enhancing Multi-Document Summarization with Cross-Document Graph-based Information Extraction
abstract
Information extraction (IE) and summarization are closely related, both tasked with presenting a subset of the information contained in a natural language text. However, while IE extracts structural representations, summarization aims to abstract the most salient information into a generated text summary – thus potentially encountering the technical limitations of current text generation methods (e.g., hallucination). To mitigate this risk, this work uses structured IE graphs to enhance the abstractive summarization task. Specifically, we focus on improving Multi-Document Summarization (MDS) performance by using cross-document IE output, incorporating two novel components: (1) the use of auxiliary entity and event recognition systems to focus the summary generation model; (2) incorporating an alignment loss between IE nodes and their text spans to reduce inconsistencies between the IE graphs and text representations. Operationally, both the IE nodes and corresponding text spans are projected into the same embedding space and pairwise distance is minimized. Experimental results on multiple MDS benchmarks show that summaries generated by our model are more factually consistent with the source documents than baseline models while maintaining the same level of abstractiveness.
Heba Elfardy, Markus Dreyer, Kevin Small, Heng Ji 0001, Mohit Bansal
EACL3
2023 Background Summarization of Event Timelines
abstract
Generating concise summaries of news events is a challenging natural language processing task.While journalists often curate timelines to highlight key sub-events, newcomers to a news event face challenges in catching up on its historical context.In this paper, we address this need by introducing the task of background news summarization, which complements each timeline update with a background summary of relevant preceding events.We construct a dataset by merging existing timeline datasets and asking human annotators to write a background summary for each timestep of each news event.We establish strong baseline performance using state-of-the-art summarization systems and propose a query-focused variant to generate background summaries.To evaluate background summary quality, we present a question-answering-based evaluation metric, Background Utility Score (BUS), which measures the percentage of questions about a current event timestep that a background summary answers.Our experiments show the effectiveness of instruction fine-tuned systems such as Flan-T5, in addition to strong zero-shot performance using GPT-3.5. 1
Adithya Pratapa, Kevin Small, Markus Dreyer
EMNLP3
2023 Generating Summaries with Controllable Readability Levels
abstract
Readability refers to how easily a reader can understand a written text.Several factors affect the readability level, such as the complexity of the text, its subject matter, and the reader's background knowledge.Generating summaries based on different readability levels is critical for enabling knowledge consumption by diverse audiences.However, current text generation approaches lack refined control, resulting in texts that are not customized to readers' proficiency levels.In this work, we bridge this gap and study techniques to generate summaries at specified readability levels.Unlike previous methods that focus on a specific readability level (e.g., lay summarization), we generate summaries with fine-grained control over their readability.We develop three text generation techniques for controlling readability:(1) instruction-based readability control, (2) reinforcement learning to minimize the gap between requested and observed readability and (3) a decoding approach that uses lookahead to estimate the readability of upcoming decoding steps.We show that our generation methods significantly improve readability control on news summarization (CNN/DM dataset), as measured by various readability metrics and human judgement, establishing strong baselines for controllable readability in summarization. 1
Leonardo F. R. Ribeiro, Mohit Bansal, Markus Dreyer
EMNLP3
2023 On Conditional and Compositional Language Model Differentiable Prompting
abstract
Prompts have been shown to be an effective method to adapt a frozen Pretrained Language Model (PLM) to perform well on downstream tasks. Prompts can be represented by a human-engineered word sequence or by a learned continuous embedding. In this work, we investigate conditional and compositional differentiable prompting. We propose a new model, Prompt Production System (ProPS), which learns to transform task instructions or input metadata, into continuous prompts that elicit task-specific outputs from the PLM. Our model uses a modular network structure based on our neural formulation of Production Systems, which allows the model to learn discrete rules -- neural functions that learn to specialize in transforming particular prompt input patterns, making it suitable for compositional transfer learning and few-shot learning. We present extensive empirical and theoretical analysis and show that ProPS consistently surpasses other PLM adaptation techniques, and often improves upon fully fine-tuned models, on compositional generalization tasks, controllable summarization and multilingual translation, while needing fewer trainable parameters.
Jonathan Pilault, Mohit Bansal, Markus Dreyer
IJCAI4
2022 FactGraph: Evaluating Factuality in Summarization with Semantic Graph Representations
abstract
Leonardo Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, Mohit Bansal. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Leonardo F. R. Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, Mohit Bansal
NAACL-HLT4
2021 Efficiently Summarizing Text and Graph Encodings of Multi-Document Clusters
abstract
Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Sujith Ravi, Markus Dreyer. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Sujith Ravi, Markus Dreyer
NAACL-HLT5
2019 Multi-Task Networks with Universe, Group, and Task Feature Learning
abstract
We present methods for multi-task learning that take advantage of natural groupings of related tasks.Task groups may be defined along known properties of the tasks, such as task domain or language.Such task groups represent supervised information at the inter-task level and can be encoded into the model.We investigate two variants of neural network architectures that accomplish this, learning different feature spaces at the levels of individual tasks, task groups, as well as the universe of all tasks: (1) parallel architectures encode each input simultaneously into feature spaces at different levels; (2) serial architectures encode each input successively into feature spaces at different levels in the task hierarchy.We demonstrate the methods on natural language understanding (NLU) tasks, where a grouping of tasks into different task domains leads to improved performance on ATIS, Snips, and a large inhouse dataset.* This work was done while Shiva Pentyala was interning at Amazon Alexa.
Shiva Pentyala, Mengwen Liu, Markus Dreyer
ACL (1)3
2017 Zero-Shot Learning Across Heterogeneous Overlapping Domains
Anjishnu Kumar, Pavankumar Reddy Muddireddy, Markus Dreyer, Björn Hoffmeister
INTERSPEECH3
2016 LatticeRnn: Recurrent Neural Networks Over Lattices
Faisal Ladhak, Ankur Gandhe, Markus Dreyer, Lambert Mathias, Ariya Rastrow, Björn Hoffmeister
INTERSPEECH3
2015 APRO: All-Pairs Ranking Optimization for MT Tuning
abstract
We present APRO, a new method for machine translation tuning that can handle large feature sets. As opposed to other popular methods (e.g., MERT, MIRA, PRO), which involve randomness and require multiple runs to obtain a reliable result, APRO gives the same result on any run, given initial feature weights. APRO follows the pairwise ranking approach of PRO (Hopkins and May, 2011), but instead of ranking a small sampled subset of pairs from thekbest list, APRO efficiently ranks all pairs. By obviating the need for manually determined sampling settings, we obtain more reliable results. APRO converges more quickly than PRO and gives similar or better translation results.
Markus Dreyer, Yuanzhe Dong
HLT-NAACL1
2015 hyp: A Toolkit for Representing, Manipulating, and Optimizing Hypergraphs
abstract
We present hyp, an open-source toolkit for the representation, manipulation, and optimization of weighted directed hypergraphs. hyp provides compose, project, invert functionality, k-best path algorithms, the inside and outside algorithms, and more. Finite-state machines are modeled as a special case of directed hypergraphs. hyp consists of a C++ API, as well as a command line tool, and is available for download at github.com/sdl-research/hyp.
Markus Dreyer, Jonathan Graehl
HLT-NAACL1
2012 HyTER: Meaning-Equivalent Semantics for Translation Evaluation
Markus Dreyer, Daniel Marcu
HLT-NAACL1
2011 Discovering Morphological Paradigms from Plain Text Using a Dirichlet Process Mixture Model
Markus Dreyer, Jason Eisner
EMNLP1
2011 Hill climbing on speech lattices: A new rescoring framework
abstract
We describe a new approach for rescoring speech lattices - with long-span language models or wide-context acoustic models - that does not entail computationally intensive lattice expansion or limited rescoring of only an N-best list. We view the set of word-sequences in a lattice as a discrete space equipped with the edit-distance metric, and develop a hill climbing technique to start with, say, the 1-best hypothesis under the lattice-generating model(s) and iteratively search a local neighborhood for the highest-scoring hypothesis under the rescoring model(s); such neighborhoods are efficiently constructed via finite state techniques. We demonstrate empirically that to achieve the same reduction in error rate using a better estimated, higher order language model, our technique evaluates fewer utterance-length hypotheses than conventional N-best rescoring by two orders of magnitude. For the same number of hypotheses evaluated, our technique results in a significantly lower error rate.
Ariya Rastrow, Markus Dreyer, Abhinav Sethy, Sanjeev Khudanpur, Bhuvana Ramabhadran, Mark Dredze
ICASSP2
2009 Graphical Models over Multiple Strings
Markus Dreyer, Jason Eisner
EMNLP1
2008 Latent-Variable Modeling of String Transductions with Finite-State Methods
Markus Dreyer, Jason Smith 0006, Jason Eisner
EMNLP1
2007 Exploiting prosody for PCFGs with latent annotations
abstract
We propose novel methods for integrating prosody in syntax using generative models. By adopting a grammar whose constituents have latent annotations, the influence of prosody on syntax can be learned from data. In one method, prosody is utilized to seed the latent annotations of a grammar which is then refined using EM iterations. In an orthogonal approach, we integrate prosody into grammar more explicitly using a model that jointly observes words and associated prosody. We evaluate the two methods by parsing speech data from the Switchboard corpus. The results are compared against baseline results from a model that does not use prosody. The experiments show that prosody improves a grammar in terms of accuracy as well as the parsimonious use of parameters. 1.
Markus Dreyer, Izhak Shafran
INTERSPEECH1
2006 Vine Parsing and Minimum Risk Reranking for Speed and Precision
Markus Dreyer, David A. Smith, Noah A. Smith
CoNLL1
2006 Better Informed Training of Latent Syntactic Features
Markus Dreyer, Jason Eisner
EMNLP1