VLDB 2026 Research / reviewers in the wild / expert
Marjorie Freedman
dblp:93/4232
· DBLP profile ↗
16ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-3060-4273ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Question answering and dialogue systems · 30% Language models and text generation · 22% Probabilistic and Bayesian machine learning · 21% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference |
0.9 | 1 | 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo |
0.9 | 1 | 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025 |
Natural language and speech › Language models and text generation
text generation |
0.9 | 1 | 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
counterfactual question answering |
0.7 | 1 | 2023 | ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
personalized dialogue generation |
0.7 | 1 | 2023 | RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation · ACL (1) 2023 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.7 | 1 | 2023 | RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation · ACL (1) 2023 |
Computer vision › Vision and language
video-language understanding |
0.7 | 1 | 2023 | ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos · EMNLP 2023 |
Computer vision › Video understanding and tracking
video question answering |
0.7 | 1 | 2023 | ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
closed-book question answering |
0.5 | 1 | 2021 | Perhaps PTLMs Should Go to School - A Task to Assess Open Book and Closed Book QA · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
open book question answering |
0.5 | 1 | 2021 | Perhaps PTLMs Should Go to School - A Task to Assess Open Book and Closed Book QA · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.1 | 1 | 2011 | Extreme Extraction - Machine Reading in a Week · EMNLP 2011 |
Natural language and speech › Information extraction and text analysis › coreference resolution
cross-document coreference resolution |
0.1 | 1 | 2008 | Who is Who and What is What: Experiments in Cross-Document Co-Reference · EMNLP 2008 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.0 | 1 | 2011 | Extreme Extraction - Machine Reading in a Week · EMNLP 2011 |
Methods — techniques the papers use, named apart from their topics
sequential monte carlo · 0.9probabilistic conditioning · 0.9video question answering · 0.7hierarchical transformer retriever · 0.7context-aware prefix encoder · 0.7retrieval · 0.5pre-trained language model · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte CarloabstractA wide range of LM applications require generating text that conforms to syntactic or semantic constraints. Imposing such constraints can be naturally framed as _probabilistic conditioning_, but exact generation from the resulting distribution—which can differ substantially from the LM’s base distribution—is generally intractable. In this work,
we develop an architecture for controlled LM generation based on sequential Monte Carlo (SMC). Our SMC framework allows us to flexibly incorporate domain- and problem-specific constraints at inference time, and efficiently reallocate computational resources in light of new information during the course of generation. By comparing to a number of alternatives and ablations on four challenging domains---Python code generation for data science, text-to-SQL, goal inference, and molecule synthesis—we demonstrate that, with little overhead, our approach allows small open-source language models to outperform models over 8$\times$ larger, as well as closed-source, fine-tuned ones.
In support of the probabilistic perspective, we show that these performance improvements are driven by better approximation to the posterior distribution.
[Our system](https://github.com/probcomp/genlm-control) builds on the framework of Lew et al. (2023) and integrates with its _language model probabilistic programming language_, giving users a simple, programmable way to apply SMC to a broad variety of controlled generation problems. João Loula, Benjamin LeBrun, Benjamin Lipkin, Clemente Pasti, Gabriel Grand, Tianyu Liu 0004, Yahya Emara, Marjorie Freedman, Jason Eisner, Ryan Cotterell, Vikash Mansinghka 0001, Alexander K. Lew, Tim Vieira, Timothy J. O'Donnell |
ICLR | 9 |
| 2023 | RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response GenerationabstractEndowing chatbots with a consistent persona is essential to an engaging conversation, yet it remains an unresolved challenge.In this work, we propose a new retrieval-enhanced approach for personalized response generation.Specifically, we design a hierarchical transformer retriever trained on dialogue domain data to perform personalized retrieval and a context-aware prefix encoder that fuses the retrieved information to the decoder more effectively.Extensive experiments on a real-world dataset demonstrate the effectiveness of our model at generating more fluent and personalized responses.We quantitatively evaluate our model's performance under a suite of human and automatic metrics and find it to be superior compared to state-of-the-art baselines on English Reddit conversations. 1 Hyundong Cho, Marjorie Freedman, Xuezhe Ma, Jonathan May |
ACL (1) | 3 |
| 2023 | ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosabstractTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Chandra, Marjorie Freedman, Ralph Weischedel, Nanyun Peng. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Te-Lin Wu, Zi-Yi Dou, Nischal Reddy Chandra, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng 0001 |
EMNLP | 6 |
| 2022 | Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional ManualsabstractTe-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph Weischedel, Nanyun Peng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Te-Lin Wu, Alexander Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng 0001 |
ACL (1) | 4 |
| 2021 | A Grounded Approach to Modeling Generic Knowledge Acquisition
Deniz Beser, Joe Cecil 0002, Marjorie Freedman, Jacob A. Lichtefeld, Mitchell P. Marcus, Sarah R. B. Payne, Charles Yang 0001 |
CogSci | 3 |
| 2021 | Grounding Word Learning Across Situations
Ryan Gabbard, Jacob A. Lichtefeld, Deniz Beser, Joe Cecil 0002, Mitchell P. Marcus, Sarah R. B. Payne, Charles Yang 0001, Marjorie Freedman |
CogSci | 8 |
| 2021 | Perhaps PTLMs Should Go to School - A Task to Assess Open Book and Closed Book QAabstractOur goal is to deliver a new task and leaderboard to stimulate research on question answering and pre-trained language models (PTLMs) to understand a significant instructional document, e.g., an introductory college textbook or a manual.PTLMs have shown great success in many question-answering tasks, given significant supervised training, but much less so in zero-shot settings.We propose a new task that includes two college-level introductory texts in the social sciences (American Government 2e) and humanities (U.S. History), hundreds of true/false statements based on review questions written by the textbook authors, validation/development tests based on the first eight chapters of the textbooks, blind tests based on the remaining textbook chapters, and baseline results given state-of-the-art PTLMs.Since the questions are balanced, random performance should be ~50%.T5, fine-tuned with BoolQ achieves the same performance, suggesting that the textbook's content is not prerepresented in the PTLM.Taking the exam closed book, but having read the textbook (i.e., adding the textbook to T5's pre-training), yields at best minor improvement (56%), suggesting that the PTLM may not have "understood" the textbook (or perhaps misunderstood the questions).Performance is better (~60%) when the exam is taken open-book (i.e., allowing the machine to automatically retrieve a paragraph and use it to answer the question). Manuel R. Ciosici, Joe Cecil 0002, Alex Hedges, Marjorie Freedman, Ralph M. Weischedel |
EMNLP (1) | 5 |
| 2019 | Contextualized Cross-Lingual Event Trigger Extraction with Minimal ResourcesabstractEvent trigger extraction is an information extraction task of practical utility, yet it is challenging due to the difficulty of disambiguating word sense meaning. Previous approaches rely extensively on hand-crafted language-specific features and are applied mainly to English for which annotated datasets and Natural Language Processing (NLP) tools are available. However, the availability of such resources varies from one language to another. Recently, contextualized Bidirectional Encoder Representations from Transformers (BERT) models have established state-of-the-art performance for a variety of NLP tasks. However, there has not been much effort in exploring language transfer using BERT for event extraction. In this work, we treat event trigger extraction as a sequence tagging problem and propose a cross-lingual framework for training it without any hand-crafted features. We experiment with different flavors of transfer learning from high-resourced to low-resourced languages and compare the performance of different multilingual embeddings for event trigger extraction. Our results show that training in a multilingual setting outperforms language-specific models for both English and Chinese. Our work is the first to experiment with two event architecture variants in a cross-lingual setting, to show the effectiveness of contextualized embeddings obtained using BERT, and to explore and analyze its performance on Arabic. Meryem M'hamdi, Marjorie Freedman, Jonathan May |
CoNLL | 2 |
| 2018 | When ACE met KBP: End-to-End Evaluation of Knowledge Base Population with Component-level Annotation
Bonan Min, Marjorie Freedman, Roger Bock, Ralph M. Weischedel |
LREC | 2 |
| 2018 | Combining rule-based and statistical mechanisms for low-resource named entity recognition
Ryan Gabbard, Jay DeYoung, Constantine Lignos, Marjorie Freedman, Ralph M. Weischedel |
Mach. Transl. | 4 |
| 2017 | Probabilistic Inference for Cold Start Knowledge Base Population with Prior World KnowledgeabstractBuilding knowledge bases (KB) automatically from text corpora is crucial for many applications such as question answering and web search.The problem is very challenging and has been divided into sub-problems such as mention and named entity recognition, entity linking and relation extraction.However, combining these components has shown to be under-constrained and often produces KBs with supersize entities and common-sense errors in relations (a person has multiple birthdates).The errors are difficult to resolve solely with IE tools but become obvious with world knowledge at the corpus level.By analyzing Freebase and a large text collection, we found that per-relation cardinality and the popularity of entities follow the power-law distribution favoring flat long tails with lowfrequency instances.We present a probabilistic joint inference algorithm to incorporate this world knowledge during KB construction.Our approach yields stateof-the-art performance on the TAC Cold Start task, and 42% and 19.4% relative improvements in F1 over our baseline on Cold Start hop-1 and all-hop queries respectively. Bonan Min, Marjorie Freedman, Talya Meltzer |
EACL (1) | 2 |
| 2017 | Learning Transferable Representation for Bilingual Relation Extraction via Convolutional Neural NetworksabstractTypically, relation extraction models are trained to extract instances of a relation ontology using only training data from a single language. However, the concepts represented by the relation ontology (e.g. ResidesIn, EmployeeOf) are language independent. The numbers of annotated examples available for a given ontology vary between languages. For example, there are far fewer annotated examples in Spanish and Japanese than English and Chinese. Furthermore, using only language-specific training data results in the need to manually annotate equivalently large amounts of training for each new language a system encounters. We propose a deep neural network to learn transferable, discriminative bilingual representation. Experiments on the ACE 2005 multilingual training corpus demonstrate that the joint training process results in significant improvement in relation classification performance over the monolingual counterparts. The learnt representation is discriminative and transferable between languages. When using 10% (25K English words, or 30K Chinese characters) of the training data, our approach results in doubling F1 compared to a monolingual baseline. We achieve comparable performance to the monolingual system trained with 250K English words (or 300K Chinese characters) With 50% of training data. Bonan Min, Zhuolin Jiang, Marjorie Freedman, Ralph M. Weischedel |
IJCNLP(1) | 3 |
| 2014 | Researching persons & organizations: AWAKE: From text to an entity-centric knowledge baseabstractWe describe a pilot experiment building a capability to automatically read documents, develop a knowledge base, support analytics, and visualize the information found. The capability allows someone researching a topic of interest of focus on analysis and synthesis rather than on reading. We show how information from multiple modalities (speech, text, structured databases) and multiple approaches (ontology driven and open information extraction) can be fused to create a resource about both previously known and novel entities. We describe an extensible framework for language understanding tools that allows for scalability, plug-and-play of alternative components, and incorporation of additional input streams, including video, images, and foreign language text. Elizabeth Boschee, Marjorie Freedman, Saurabh Khanwalkar, Amit Srivastava, Ralph M. Weischedel |
IEEE BigData | 2 |
| 2011 | Extreme Extraction - Machine Reading in a Week
Marjorie Freedman, Lance A. Ramshaw, Elizabeth Boschee, Ryan Gabbard, Gary Kratkiewicz, Nicolas Ward, Ralph M. Weischedel |
EMNLP | 1 |
| 2008 | Who is Who and What is What: Experiments in Cross-Document Co-Reference
Alex Baron, Marjorie Freedman |
EMNLP | 2 |
| 2007 | Selecting on-topic sentences from natural language corporaabstractWe describe a system that examines input sentences with re-spect to arbitrary topics formulated as natural language expres-sions. It extracts predicate-argument structures from text in-tervals and links them into semantically organized proposition trees. By instantiating trees constructed for topic descriptions in trees representing input sentences or parts thereof, we are able to assess degree of “topicality ” for each sentence. The presented strategy was used in the BBN distillation system for the GALE Year 1 evaluation and achieved outstanding results compared to other systems and human participants. Index Terms: machine learning, question answering, topicality. 1. Michael Levit, Elizabeth Boschee, Marjorie Freedman |
INTERSPEECH | 3 |