VLDB 2026 Research / reviewers in the wild / expert
Daniel Khashabi
dblp:71/10515
· DBLP profile ↗
58ranked-venue papers
10as first author
40since 2021 · last 2026
0009-0009-7664-2230ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 9 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table GenerationabstractLiterature review tables are essential for summarizing and comparing collections of scientific papers.In this paper, we study automatic generation of such tables from a pool of papers to satisfy a user's information need.Building on recent work (Newman et al., 2024), we move beyond oracle settings by (i) simulating wellspecified yet schema-agnostic user demands that avoid leaking gold column names or values, (ii) explicitly modeling retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators, and (iii) introducing a lightweight, annotation-free, utilization-oriented evaluation that decomposes utility (schema coverage, unary cell fidelity, pairwise relational consistency) and measures paper selection via a two-way QA procedure (gold→system and system→gold) with recall, precision, and F1.To support reproducible evaluation, we introduce ARXIV2TABLE, a benchmark of 1,957 tables referencing 7,158 papers, with human-verified distractors and rewritten, schema-agnostic user demands.We also develop an iterative, batch-based generation method that co-refines paper filtering and schema over multiple rounds.We validate the evaluation protocol with human audits and cross-evaluator checks.Extensive experiments show that our method consistently improves over strong baselines, while absolute scores remain modest, underscoring the task's difficulty.Our data and code is available at https://github.com Weiqi Wang 0001, Jiefu Ou, Yangqiu Song, Benjamin Van Durme, Daniel Khashabi |
ACL (1) | 5 |
| 2026 | Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
Debashish Chakraborty, Eugene Yang 0001, Daniel Khashabi, Dawn J. Lawrie, Kevin Duh |
ECIR (2) | 3 |
| 2026 | Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context ProblemsabstractAbstract Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effectiveness for long-context problems? We challenge the assumption that SC’s benefits generalize to long-context settings, where LLMs often struggle with position bias—the systematic over-reliance on specific context regions—which hinders their ability to utilize information effectively from all parts of their context. Through comprehensive experimentation with varying state-of-the-art models, tasks, and SC formulations, we find that SC not only fails to improve but actively degrades performance on long-context tasks. This degradation is driven by persistent position bias, which worsens with longer context lengths and smaller model sizes but remains invariant to prompt format or task type. Unlike short-context tasks, where SC diversifies reasoning paths, long-context SC amplifies positional errors. These comprehensive results provide valuable insight into the limitations of current LLMs in long-context understanding and highlight the need for more sophisticated approaches. Adam Byerly, Daniel Khashabi |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated ResponsesabstractCan LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment. Dongwei Jiang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi |
AAAI | 6 |
| 2025 | WorldAPIs: The World Is Worth How Many APIs? A Thought ExperimentabstractAI systems make decisions in physical environments through primitive actions or affordances that are accessed via API calls. While deploying AI agents in the real world involves numerous high-level actions, existing embodied simulators offer a limited set of domain-salient APIs. This naturally brings up the questions: how many primitive actions (APIs) are needed for a versatile embodied agent, and how should they look like? We explore this via a thought experiment: assuming that wikiHow tutorials cover a wide variety of human-written tasks, what is the space of APIs needed to cover these instructions? We propose a framework to iteratively induce new APIs by grounding wikiHow instruction to situated agent policies. Inspired by recent successes in large language models (LLMs) for embodied planning, we propose a few-shot prompting to steer GPT-4 to generate Pythonic programs as agent policies and bootstrap a universe of APIs by 1) reusing a seed set of APIs; and then 2) fabricate new API calls when necessary. The focus of this thought experiment is on defining these APIs rather than their excitability. We apply the proposed pipeline on instructions from wiki- How tutorials. On a small fraction (0.5%) of tutorials, we induce an action space of 300+ APIs necessary for capturing the rich variety of tasks in the physical world. A detailed automatic and human analysis of the induction output reveals that the proposed pipeline enables effective reuse and creation of APIs. Moreover, a manual review revealed that existing simulators support only a small subset of the induced APIs (9 of the top 50 frequent APIs), motivating the development of action-rich embodied environments. Jiefu Ou, Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi |
AAAI | 4 |
| 2025 | RATIONALYST: Pre-training Process-Supervision for Improving ReasoningabstractDongwei Jiang, Guoxuan Wang, Yining Lu, Andrew Wang, Jingyu Zhang, Chuyu Liu, Benjamin Van Durme, Daniel Khashabi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Dongwei Jiang, Guoxuan Wang, Yining Lu, Chuyu Liu, Benjamin Van Durme, Daniel Khashabi |
ACL (1) | 8 |
| 2025 | Evaluating the Evaluators: Are readability metrics good measures of readability?abstractPlain Language Summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences.In this paper, we conduct a thorough survey of PLS literature, and identify that the current standard practice for readability evaluation is to use traditional readability metrics, such as Flesch-Kincaid Grade Level (FKGL).However, despite proven utility in other fields, these metrics have not been compared to human readability judgments in PLS.We evaluate 8 readability metrics and show that most correlate poorly with human judgments, including the most popular metric, FKGL.We then show that Language Models (LMs) are better judges of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments.Extending our analysis to PLS datasets, which contain summaries aimed at non-expert audiences, we find that LMs better capture deeper measures of readability, such as required background knowledge, and lead to different conclusions than the traditional metrics.Based on these findings, we offer recommendations for best practices in the evaluation of plain language summaries.We release our analysis code and survey data. Isabel Cachola, Daniel Khashabi, Mark Dredze |
EMNLP | 2 |
| 2025 | ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution CiphersabstractRecent works have suggested that In-Context Learning (ICL) operates in dual modes, i.e. task retrieval (remember learned patterns from pre-training) and task learning (inference-time "learning" from demonstrations).However, disentangling these the two modes remains a challenging goal.We introduce ICL CIPHERS, a class of task reformulations based on substitution ciphers borrowed from classic cryptography.In this approach, a subset of tokens in the in-context inputs are substituted with other (irrelevant) tokens, rendering English sentences less comprehensible to human eye.However, by design, there is a latent, fixed pattern to this substitution, making it reversible.This bijective (reversible) cipher ensures that the task remains a well-defined task in some abstract sense, despite the transformations.It is a curious question if LLMs can solve tasks reformulated by ICL CIPHERS with a BIJEC-TIVE mapping, which requires "deciphering" the latent cipher.We show that LLMs are better at solving tasks reformulated by ICL CIPHERS with BIJECTIVE mappings than the NON-BIJECTIVE (irreversible) baseline, providing a novel approach to quantify "learning" in ICL.While this gap is small, it is consistent across the board on four datasets and six models.Finally, our interpretability analysis shows evidence that LLMs can internally decode ciphered inputs. 1 Zhouxiang Fang, Aayush Mishra, Muhan Gao, Daniel Khashabi |
EMNLP | 5 |
| 2025 | Certified Mitigation of Worst-Case LLM Copyright InfringementabstractThe exposure of large language models (LLMs) to copyrighted material during pre-training raises concerns about unintentional copyright infringement post deployment.This has driven the development of "copyright takedown" methods-post-training approaches aimed at preventing models from generating content substantially similar to copyrighted ones.While current mitigation approaches are somewhat effective for average-case risks, we demonstrate that they overlook worst-case copyright risks exhibited by the existence of long, verbatim quotes from copyrighted sources.We propose BLOOMSCRUB, a remarkably simple yet highly effective inference-time approach that provides certified copyright takedown.Our method repeatedly interleaves quote detection with rewriting techniques to transform potentially infringing segments.By leveraging efficient data sketches (Bloom filters), our approach enables scalable copyright screeningeven for large-scale real-world corpora.When quotes beyond a length threshold cannot be removed, the system can abstain from responding, offering certified risk reduction.Experimental results show that BLOOMSCRUB reduces infringement risk, preserves utility, and accommodates different levels of enforcement stringency with adaptive abstention.Our results suggest that lightweight, inference-time methods can be surprisingly effective for copyright prevention. Jiacan Yu, Marc Marone, Benjamin Van Durme, Daniel Khashabi |
EMNLP | 5 |
| 2025 | GenEx: Generating an Explorable WorldabstractUnderstanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step toward this goal by introducing *GenEx*, a system capable of planning complex embodied world exploration, guided by its generative imagination that forms expectations about the surrounding environments. *GenEx* generates high-quality, continuous 360-degree virtual environments, achieving robust loop consistency and active 3D mapping over extended trajectories. Leveraging generative imagination, GPT-assisted agents can undertake complex embodied tasks, including goal-agnostic exploration and goal-driven navigation. Agents utilize imagined observations to update their beliefs, simulate potential outcomes, and enhance their decision-making. Training on the synthetic urban dataset *GenEx-DB* and evaluation on *GenEx-EQA* demonstrate that our approach significantly improves agents' planning capabilities, providing a transformative platform toward intelligent, imaginative embodied exploration. Taiming Lu, Tianmin Shu, Alan L. Yuille, Daniel Khashabi, Jieneng Chen |
ICLR | 4 |
| 2025 | Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety RequirementsabstractThe current paradigm for safety alignment of large language models (LLMs) follows a _one-size-fits-all_ approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In addition, users may have diverse safety needs, making a model with _static_ safety standards too restrictive to be useful, as well as too costly to be re-aligned.
We propose _Controllable Safety Alignment_ (CoSA), a framework designed to adapt models to diverse safety requirements without re-training. Instead of aligning a fixed model, we align models to follow _safety configs_—free-form natural language descriptions of the desired safety behaviors—that are provided as part of the system prompt. To adjust model safety behavior, authorized users only need to modify such safety configs at inference time. To enable that, we propose CoSAlign, a data-centric method for aligning LLMs to easily adapt to diverse safety configs. Furthermore, we devise a novel controllability evaluation protocol that considers both helpfulness and configured safety, summarizing them into CoSA-Score, and construct CoSApien, a _human-authored_ benchmark that consists of real-world LLM use cases with diverse safety requirements and corresponding evaluation prompts. We show that CoSAlign leads to substantial gains of controllability over strong baselines including in-context alignment. Our framework encourages better representation and adaptation to pluralistic human values in LLMs, and thereby increasing their practicality. Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, Benjamin Van Durme |
ICLR | 4 |
| 2025 | SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference LearningabstractAligning language models with human preferences relies on pairwise preference datasets. While some studies suggest that on-policy data consistently outperforms off-policy data for preference learning, others indicate that the advantages of on-policy data are task-dependent, highlighting the need for a systematic exploration of their interplay. In this work, we show that on-policy and off-policy data offer complementary strengths: on-policy data is particularly effective for reasoning tasks like math and coding, while off-policy data performs better on subjective tasks such as creative writing and making personal recommendations. Guided by these findings, we introduce SimpleMix, an approach to combine the complementary strengths of on-policy and off-policy preference learning by simply mixing these two data sources. Our empirical results across diverse tasks and benchmarks demonstrate that SimpleMix substantially improves language model alignment. Specifically, SimpleMix improves upon on-policy DPO and off-policy DPO by an average of 6.03 on Alpaca Eval 2.0. Moreover, it surpasses prior approaches that are much more complex in combining on- and off-policy data, such as HyPO and DPO-Mix-P, by an average of 3.05. These findings validate the effectiveness and efficiency of SimpleMix for enhancing preference-based alignment. Tianjian Li, Daniel Khashabi |
ICML | 2 |
| 2025 | Upsample or Upweight? Balanced Training on Heavily Imbalanced DatasetsabstractTianjian Li, Haoran Xu, Weiting Tan, Kenton Murray, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tianjian Li, Weiting Tan, Kenton Murray, Daniel Khashabi |
NAACL (Long Papers) | 5 |
| 2025 | Benchmarking Language Model Creativity: A Case Study on Code GenerationabstractYining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, Sanjeev Khudanpur, Meng Jiang, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, Sanjeev Khudanpur, Meng Jiang 0001, Daniel Khashabi |
NAACL (Long Papers) | 7 |
| 2025 | TurkingBench: A Challenge Benchmark for Web AgentsabstractKevin Xu, Yeganeh Kordi, Tanay Nayak, Adi Asija, Yizhong Wang, Kate Sanders, Adam Byerly, Jingyu Zhang, Benjamin Van Durme, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kevin Xu, Yeganeh Kordi, Tanay Nayak, Adi Asija, Yizhong Wang, Kate Sanders 0002, Adam Byerly, Benjamin Van Durme, Daniel Khashabi |
NAACL (Long Papers) | 10 |
| 2025 | Verifiable by Design: Aligning Language Models to Quote from Pre-Training DataabstractJingyu Zhang, Marc Marone, Tianjian Li, Benjamin Van Durme, Daniel Khashabi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Marc Marone, Tianjian Li, Benjamin Van Durme, Daniel Khashabi |
NAACL (Long Papers) | 5 |
| 2025 | FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External FeedbackabstractRecent studies have shown LLMs possess some ability to improve their responses when given external feedback. However, it remains unclear how effectively and thoroughly these models can incorporate extrinsic feedback. In an ideal scenario, if LLMs receive near-perfect and complete feedback, we would expect them to fully integrate the feedback and reach correct solutions. In this paper, we systematically investigate LLMs’ ability to incorporate feedback by designing a controlled experimental environment. For each problem, a solver model attempts a solution, then a feedback generator with access to near-complete ground-truth answers produces targeted feedback, after which the solver tries again. We evaluate this pipeline across a diverse range of tasks, including math reasoning, knowledge reasoning, scientific reasoning, and general multi-domain evaluations with state-of-the-art language models including Claude 3.7 with extended thinking. Surprisingly, even under these near-ideal conditions, solver models consistently show resistance to feedback, a limitation that we term FEEDBACK FRICTION. To mitigate this limitation, we experiment with sampling-based strategies like progressive temperature increases and explicit rejection of previously attempted incorrect answers, which yield improvements but still fail to help models achieve target performance. We analyze FEEDBACK FRICTION and find that models’ confidence on specific questions, measured by semantic entropy, predicts feedback resistance: high-confidence predictions remain resistant to external correction. We hope that highlighting this issue in LLMs will help future research in self-improvement. Dongwei Jiang, Nicholas Andrews, Daniel Khashabi |
NeurIPS | 5 |
| 2024 | RORA: Robust Free-Text Rationale EvaluationabstractZhengping Jiang, Yining Lu, Hanjie Chen, Daniel Khashabi, Benjamin Van Durme, Anqi Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhengping Jiang, Yining Lu, Daniel Khashabi, Benjamin Van Durme, Anqi Liu 0001 |
ACL (1) | 4 |
| 2024 | GEAR: Augmenting Language Models with Generalizable and Efficient Tool ResolutionabstractAugmenting large language models (LLM) to use external tools enhances their performance across a variety of tasks.However, prior works over-rely on task-specific demonstration of tool use that limits their generalizability and computational cost due to making many calls to large-scale LLMs.We introduce GEAR, a computationally efficient query-tool grounding algorithm that is generalizable to various tasks that require tool use while not relying on taskspecific demonstrations.GEAR achieves better efficiency by delegating tool grounding and execution to small language models (SLM) and LLM, respectively; while leveraging semantic and pattern-based evaluation at both question and answer levels for generalizable tool grounding.We evaluate GEAR on 14 datasets across 6 downstream tasks, demonstrating its strong generalizability to novel tasks, tools and different SLMs.Despite offering more efficiency, GEAR achieves higher precision in tool grounding compared to prior strategies using LLM prompting, thus improving downstream accuracy at a reduced computational cost.For example, we demonstrate that GEAR-augmented GPT-J and GPT-3 outperform counterpart toolaugmented baselines because of better tool use. Yining Lu, Haoping Yu, Daniel Khashabi |
EACL (1) | 3 |
| 2024 | "According to . . . ": Prompting Language Models Improves Quoting from Pre-Training DataabstractOrion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, Benjamin Van Durme. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Orion Weller, Marc Marone, Nathaniel Weir, Dawn J. Lawrie, Daniel Khashabi, Benjamin Van Durme |
EACL (1) | 5 |
| 2024 | AnaloBench: Benchmarking the Identification of Abstract and Long-context AnalogiesabstractXiao Ye, Andrew Wang, Jacob Choi, Yining Lu, Shreya Sharma, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, Daniel Khashabi. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jacob Choi, Yining Lu, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, Daniel Khashabi |
EMNLP | 9 |
| 2024 | Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsabstractText generation models are notoriously vulnerable to errors in the training data. With the wide-spread availability of massive amounts of web-crawled data becoming more commonplace, how can we enhance the robustness of models trained on a massive amount of noisy web-crawled text? In our work, we propose Error Norm Truncation (ENT), a robust enhancement method to the standard training objective that truncates noisy data. Compared to methods that only uses the negative log-likelihood loss to estimate data quality, our method provides a more accurate estimation by considering the distribution of non-target tokens, which is often overlooked by previous work. Through comprehensive experiments across language modeling, machine translation, and text summarization, we show that equipping text generation models with ENT improves generation quality over standard training and previous soft and hard truncation methods. Furthermore, we show that our method improves the robustness of models against two of the most detrimental types of noise in machine translation, resulting in an increase of more than 2 BLEU points over the MLE baseline when up to 50\% of noise is added to the data. Tianjian Li, Philipp Koehn, Daniel Khashabi, Kenton Murray |
ICLR | 4 |
| 2024 | The Trickle-down Impact of Reward Inconsistency on RLHFabstractStandard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations. A notable subject that is understudied is the (in-)consistency of RMs --- whether they can recognize the semantic changes to different prompts and
appropriately adapt their reward assignments
--- and their impact on the downstream RLHF model.
In this paper, we visit a series of research questions relevant to RM inconsistency:
(1) How can we measure the consistency of reward models?
(2) How consistent are the existing RMs and how can we improve them?
(3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training?
We propose **Contrast Instruction** -- a benchmarking strategy for the consistency of RM.
Each example in **Contrast Instruction** features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on \contrast{} compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques **ConvexDA** and **RewardFusion**, which enhance reward consistency
through extrapolation during the RM training and inference stage, respectively.
We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process. Lingfeng Shen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu 0001 |
ICLR | 7 |
| 2024 | Position: Do pretrained Transformers Learn In-Context by Gradient Descent?abstractThe emergence of In-Context Learning (ICL) in LLMs remains a remarkable phenomenon that is partially understood. To explain ICL, recent studies have created theoretical connections to Gradient Descent (GD). We ask, do such connections hold up in actual pre-trained language models? We highlight the limiting assumptions in prior works that make their setup considerably different from the practical setup in which language models are trained. For example, their experimental verification uses ICL objective (training models explicitly for ICL), which differs from the emergent ICL in the wild. Furthermore, the theoretical hand-constructed weights used in these studies have properties that don’t match those of real LLMs. We also look for evidence in real models. We observe that ICL and GD have different sensitivity to the order in which they observe demonstrations. Finally, we probe and compare the ICL vs. GD hypothesis in a natural setting. We conduct comprehensive empirical analyses on language models pre-trained on natural data (LLaMa-7B). Our comparisons of three performance metrics highlight the inconsistent behavior of ICL and GD as a function of various factors such as datasets, models, and the number of demonstrations. We observe that ICL and GD modify the output distribution of language models differently. These results indicate that the equivalence between ICL and GD remains an open hypothesis and calls for further studies. Lingfeng Shen, Aayush Mishra, Daniel Khashabi |
ICML | 3 |
| 2024 | SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text GenerationabstractAbe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Abe Bohan Hou, Tianxing He, Yichen Wang 0002, Yung-Sung Chuang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov |
NAACL-HLT | 9 |
| 2024 | Efficient Large Multi-modal Models via Visual Context CompressionabstractWhile significant advancements have been made in compressed representations for text embeddings in large language models (LLMs), the compression of visual tokens in multi-modal LLMs (MLLMs) has remained a largely overlooked area. In this work, we present the study on the analysis of redundancy concerning visual tokens and efficient training within these models. Our initial experiments
show that eliminating up to 70% of visual tokens at the testing stage by simply average pooling only leads to a minimal 3% reduction in visual question answering accuracy on the GQA benchmark, indicating significant redundancy in visual context. Addressing this, we introduce Visual Context Compressor, which reduces the number of visual tokens to enhance training and inference efficiency without sacrificing performance. To minimize information loss caused by the compression on visual tokens while maintaining training efficiency, we develop LLaVolta as a light and staged training scheme that incorporates stage-wise visual context compression to progressively compress the visual tokens from heavily to lightly compression during training, yielding no loss of information when testing. Extensive experiments demonstrate that our approach enhances the performance of MLLMs in both image-language and video-language understanding, while also significantly cutting training costs and improving inference efficiency. Jieneng Chen, Luoxin Ye, Ju He, Daniel Khashabi, Alan L. Yuille |
NeurIPS | 5 |
| 2024 | DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech TranslationabstractNon-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate text data. Although NATs generate high-quality outputs and offer faster inference than autoregressive models, they tend to produce incoherent and repetitive results due to complex data distribution (e.g., acoustic and linguistic variations in speech). In this work, we introduce DiffNorm, a diffusion-based normalization strategy that simplifies data distributions for training NAT models. After training with a self-supervised noise estimation objective, DiffNorm constructs normalized target data by denoising synthetically corrupted speech features. Additionally, we propose to regularize NATs with classifier-free guidance, improving model robustness and translation quality by randomly dropping out source information during training. Our strategies result in a notable improvement of about $+7$ ASR-BLEU for English-Spanish (En-Es) translation and $+2$ ASR-BLEU for English-French (En-Fr) on the CVSS benchmark, while attaining over $14\times$ speedup for En-Es and $5 \times$ speedup for En-Fr translations compared to autoregressive baselines. Weiting Tan, Lingfeng Shen, Daniel Khashabi, Philipp Koehn |
NeurIPS | 4 |
| 2023 | When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesabstractAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 5 |
| 2023 | Self-Instruct: Aligning Language Models with Self-Generated InstructionsabstractYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 6 |
| 2023 | Generating Sequences by Learning to Self-Correct
Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, Yejin Choi 0001 |
ICLR | 6 |
| 2022 | Cross-Task Generalization via Natural Language Crowdsourcing InstructionsabstractHumans (e.g., crowdworkers) have a remarkable ability in solving different tasks, by simply reading textual instructions that define them and looking at a few examples.Despite the success of the conventional supervised learning on individual datasets, such models often struggle with generalization across tasks (e.g., a question-answering system cannot solve classification tasks).A long-standing challenge in AI is to build a model that learns a new task by understanding the humanreadable instructions that define it.To study this, we introduce NATURAL INSTRUCTIONS, a dataset of 61 distinct tasks, their humanauthored instructions, and 193k task instances (input-output pairs).The instructions are obtained from crowdsourcing instructions used to create existing NLP datasets and mapped to a unified schema.Using this meta-dataset, we measure cross-task generalization by training models on seen tasks and measuring generalization to the remaining unseen ones.We adopt generative pre-trained language models to encode task-specific instructions along with input and generate task output.Our results indicate that models benefit from instructions when evaluated in terms of generalization to unseen tasks (19% better for models utilizing instructions).These models, however, are far behind an estimated performance upperbound, indicating significant room for more progress in this direction.1 Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi |
ACL (1) | 2 |
| 2022 | ProsocialDialog: A Prosocial Backbone for Conversational AgentsabstractMost existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.To address this issue, we introduce PROSOCIALDIALOG, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms.Covering diverse unethical, problematic, biased, and toxic situations, PROSOCIALDIALOG contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-ofthumb, RoTs).Created via a human-AI collaborative framework, PROSOCIALDIALOG consists of 58K dialogues, with 331K utterances, 160K unique RoTs, and 497K dialogue safety labels accompanied by free-form rationales.With this dataset, we introduce a dialogue safety detection module, Canary, capable of generating RoTs given conversational context, and a socially-informed dialogue agent, Prost.Empirical results show that Prost generates more socially acceptable dialogues compared to other state-of-the-art language and dialogue models in both in-domain and out-of-domain settings.Additionally, Canary effectively guides off-the-shelf language models to generate significantly more prosocial responses.Our work highlights the promise and importance of creating and steering conversational AI to be socially responsible. Hyunwoo Kim 0002, Youngjae Yu, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi 0001, Maarten Sap |
EMNLP | 5 |
| 2022 | GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationabstractDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, Daniel Weld. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi 0001, Noah A. Smith, Daniel S. Weld |
EMNLP | 1 |
| 2022 | Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous PromptsabstractDaniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson 0001, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh 0001, Yejin Choi 0001 |
NAACL-HLT | 1 |
| 2022 | NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead HeuristicsabstractXiming Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah Smith, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ximing Lu, Sean Welleck, Peter West, Jungo Kasai, Daniel Khashabi, Ronan Le Bras 0001, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, Yejin Choi 0001 |
NAACL-HLT | 6 |
| 2022 | Time Waits for No One! Analysis and Challenges of Temporal MisalignmentabstractKelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, Noah A. Smith |
NAACL-HLT | 2 |
| 2022 | COLD Decoding: Energy-based Constrained Text Generation with Langevin DynamicsabstractMany applications of text generation require incorporating different constraints to control the semantics or style of generated text. These constraints can be hard (e.g., ensuring certain keywords are included in the output) and soft (e.g., contextualizing the output with the left- or right-hand context). In this paper, we present Energy-based Constrained Decoding with Langevin Dynamics (COLD), a decoding framework which unifies constrained generation as specifying constraints through an energy function, then performing efficient differentiable reasoning over the constraints through gradient-based sampling. COLD decoding is a flexible framework that can be applied directly to off-the-shelf left-to-right language models without the need for any task-specific fine-tuning, as demonstrated through three challenging text generation applications: lexically-constrained generation, abductive reasoning, and counterfactual reasoning. Our experiments on these constrained generation tasks point to the effectiveness of our approach, both in terms of automatic and human evaluation. Lianhui Qin, Sean Welleck, Daniel Khashabi, Yejin Choi 0001 |
NeurIPS | 3 |
| 2021 | Text Modular Networks: Learning to Decompose Tasks in the Language of Existing ModelsabstractTushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, Ashish Sabharwal. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tushar Khot, Daniel Khashabi, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
NAACL-HLT | 2 |
| 2021 | Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning StrategiesabstractAbstract A key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce StrategyQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies. We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, StrategyQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in StrategyQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of ∼ 66%. Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | ParsiNLU: A Suite of Language Understanding Challenges for PersianabstractAbstract Despite the progress made in recent years in addressing natural language understanding (NLU) challenges, the majority of this progress remains to be concentrated on resource-rich languages like English. This work focuses on Persian language, one of the widely spoken languages in the world, and yet there are few NLU datasets available for this language. The availability of high-quality evaluation datasets is a necessity for reliable assessment of the progress on different NLU tasks and domains. We introduce ParsiNLU, the first benchmark in Persian language that includes a range of language understanding tasks—reading comprehension, textual entailment, and so on. These datasets are collected in a multitude of ways, often involving manual annotations by native speakers. This results in over 14.5k new instances across 6 distinct NLU tasks. Additionally, we present the first results on state-of-the-art monolingual and multilingual pre-trained language models on this benchmark and compare them with human performance, which provides valuable insights into our ability to tackle natural language understanding challenges in Persian. We hope ParsiNLU fosters further research and advances in Persian language understanding.1 Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, Mozhdeh Gheini, Arman Kabiri, Rabeeh Karimi Mahabadi, Omid Memarrast, Ahmadreza Mosallanezhad, Erfan Noury, Shahab Raji, Mohammad Sadegh Rasooli, Sepideh Sadeghi, Erfan Sadeqi Azer, Niloofar Safi Samghabadi, Mahsa Shafaei, Saber Sheybani, Ali Tazarv, Yadollah Yaghoobzadeh |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess HypothesesabstractEmpirical research in Natural Language Processing (NLP) has adopted a narrow set of principles for assessing hypotheses, relying mainly on p-value computation, which suffers from several known issues.While alternative proposals have been well-debated and adopted in other fields, they remain rarely discussed or used within the NLP community.We address this gap by contrasting various hypothesis assessment techniques, especially those not commonly used in the field (such as evaluations based on Bayesian inference).Since these statistical techniques differ in the hypotheses they can support, we argue that practitioners should first decide their target hypothesis before choosing an assessment method.This is crucial because common fallacies, misconceptions, and misinterpretation surrounding hypothesis assessment methods often stem from a discrepancy between what one would like to claim versus what the method used actually assesses.Our survey reveals that these issues are omnipresent in the NLP research community.As a step forward, we provide best practices and guidelines tailored towards NLP research, as well as an easy-to-use package called HyBayes for Bayesian assessment of hypotheses, 1 complementing existing tools. Erfan Sadeqi Azer, Daniel Khashabi, Ashish Sabharwal, Dan Roth 0001 |
ACL | 2 |
| 2020 | Temporal Common Sense Acquisition with Minimal SupervisionabstractTemporal common sense (e.g., duration and frequency of events) is crucial for understanding natural language.However, its acquisition is challenging, partly because such information is often not expressed explicitly in text, and human annotation on such concepts is costly.This work proposes a novel sequence modeling approach that exploits explicit and implicit mentions of temporal common sense, extracted from a large corpus, to build TACOLM, 1 a temporal common sense language model.Our method is shown to give quality predictions of various dimensions of temporal common sense (on UDST and a newly collected dataset from Real-News).It also produces representations of events for relevant tasks such as duration comparison, parent-child relations, event coreference and temporal QA (on TimeBank, HiEVE and MCTACO) that are better than using the standard BERT.Thus, it will be an important component of temporal NLP. Ben Zhou, Qiang Ning, Daniel Khashabi, Dan Roth 0001 |
ACL | 3 |
| 2020 | More Bang for Your Buck: Natural Perturbation for Robust Question AnsweringabstractDeep learning models for linguistic tasks require large training datasets, which are expensive to create.As an alternative to the traditional approach of creating new instances by repeating the process of creating one instance, we propose doing so by first collecting a set of seed examples and then applying humandriven natural perturbations (as opposed to rule-based machine perturbations), which often change the gold label as well.Such perturbations have the advantage of being relatively easier (and hence cheaper) to create than writing out completely new examples.Further, they help address the issue that even models achieving human-level scores on NLP datasets are known to be considerably sensitive to small changes in input.To evaluate the idea, we consider a recent question-answering dataset (BOOLQ) and study our approach as a function of the perturbation cost ratio, the relative cost of perturbing an existing question vs. creating a new one from scratch.We find that when natural perturbations are moderately cheaper to create (cost ratio under 60%), it is more effective to use them for training BOOLQ models: such models exhibit 9% higher robustness and 4.5% stronger generalization, while retaining performance on the original BOOLQ dataset. Daniel Khashabi, Tushar Khot, Ashish Sabharwal |
EMNLP (1) | 1 |
| 2020 | TransOMCS: From Linguistic Graphs to Commonsense KnowledgeabstractCommonsense knowledge acquisition is a key problem for artificial intelligence. Conventional methods of acquiring commonsense knowledge generally require laborious and costly human annotations, which are not feasible on a large scale. In this paper, we explore a practical way of mining commonsense knowledge from linguistic graphs, with the goal of transferring cheap knowledge obtained with linguistic patterns into expensive commonsense knowledge. The result is a conversion of ASER [Zhang et al., 2020], a large-scale selectional preference knowledge resource, into TransOMCS, of the same representation as ConceptNet [Liu and Singh, 2004] but two orders of magnitude larger. Experimental results demonstrate the transferability of linguistic knowledge to commonsense knowledge and the effectiveness of the proposed approach in terms of quantity, novelty, and quality. TransOMCS is publicly available at: https://github.com/HKUST-KnowComp/TransOMCS. Hongming Zhang 0009, Daniel Khashabi, Yangqiu Song, Dan Roth 0001 |
IJCAI | 2 |
| 2019 | "Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense UnderstandingabstractBen Zhou, Daniel Khashabi, Qiang Ning, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ben Zhou, Daniel Khashabi, Qiang Ning, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Question Answering as Global Reasoning Over Semantic AbstractionsabstractWe propose a novel method for exploiting the semantic structure of text to answer multiple-choice questions. The approach is especially suitable for domains that require reasoning over a diverse set of linguistic constructs but have limited training data. To address these challenges, we present the first system, to the best of our knowledge, that reasons over a wide range of semantic abstractions of the text, which are derived using off-the-shelf, general-purpose, pre-trained natural language modules such as semantic role labelers, coreference resolvers, and dependency parsers. Representing multiple abstractions as a family of graphs, we translate question answering (QA) into a search for an optimal subgraph that satisfies certain global and local properties. This formulation generalizes several prior structured QA systems. Our system, SEMANTICILP, demonstrates strong performance on two domains simultaneously. In particular, on a collection of challenging science QA datasets, it outperforms various state-of-the-art approaches, including neural models, broad coverage information retrieval, and specialized techniques using structured knowledge bases, by 2%-6%. Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
AAAI | 1 |
| 2018 | Zero-Shot Open Entity Typing as Type-Compatible GroundingabstractThe problem of entity-typing has been studied predominantly in supervised learning fashion, mostly with task-specific annotations (for coarse types) and sometimes with distant supervision (for fine types).While such approaches have strong performance within datasets, they often lack the flexibility to transfer across text genres and to generalize to new type taxonomies.In this work we propose a zero-shot entity typing approach that requires no annotated data and can flexibly identify newly defined types.Given a type taxonomy defined as Boolean functions of FREEBASE "types", we ground a given mention to a set of type-compatible Wikipedia entries and then infer the target mention's types using an inference algorithm that makes use of the types of these entries.We evaluate our system on a broad range of datasets, including standard fine-grained and coarse-grained entity typing datasets, and also a dataset in the biological domain.Our system is shown to be competitive with state-of-theart supervised NER systems and outperforms them on out-of-domain datasets.We also show that our system significantly outperforms other zero-shot fine typing systems. Ben Zhou, Daniel Khashabi, Chen-Tse Tsai, Dan Roth 0001 |
EMNLP | 2 |
| 2018 | CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001 |
LREC | 1 |
| 2018 | Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple SentencesabstractDaniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Daniel Khashabi, Snigdha Chaturvedi, Michael Roth 0001, Shyam Upadhyay, Dan Roth 0001 |
NAACL-HLT | 1 |
| 2017 | Learning What is Essential in QuestionsabstractQuestion answering (QA) systems are easily distracted by irrelevant or redundant words in questions, especially when faced with long or multi-sentence questions in difficult domains. This paper introduces and studies the notion of essential question terms with the goal of improving such QA solvers. We illustrate the importance of essential question terms by showing that humans' ability to answer questions drops significantly when essential terms are eliminated from questions.We then develop a classifier that reliably (90% mean average precision) identifies and ranks essential terms in questions. Finally, we use the classifier to demonstrate that the notion of question term essentiality allows state-of-the-art QA solver for elementary-level science questions to make better and more informed decisions,improving performance by up to 5%.We also introduce a new dataset of over 2,200 crowd-sourced essential terms annotated science questions. Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
CoNLL | 1 |
| 2016 | Combining Retrieval, Statistics, and Inference to Answer Elementary Science QuestionsabstractWhat capabilities are required for an AI system to pass standard 4th Grade Science Tests? Previous work has examined the use of Markov Logic Networks (MLNs) to represent the requisite background knowledge and interpret test questions, but did not improve upon an information retrieval (IR) baseline. In this paper, we describe an alternative approach that operates at three levels of representation and reasoning: information retrieval, corpus statistics, and simple inference over a semi-automatically constructed knowledge base, to achieve substantially improved results. We evaluate the methods on six years of unseen, unedited exam questions from the NY Regents Science Exam (using only non-diagram, multiple choice questions), and show that our overall system’s score is 71.3%, an improvement of 23.8% (absolute) over the MLN-based method described in previous work. We conclude with a detailed analysis, illustrating the complementary strengths of each method in the ensemble. Our datasets are being released to enable further research. Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter D. Turney, Daniel Khashabi |
AAAI | 7 |
| 2016 | Better call Saul: Flexible Programming for Learning and Inference in NLPabstractWe present a novel way for designing complex joint inference and learning models using Saul (Kordjamshidi et al., 2015), a recently-introduced declarative learning-based programming language (DeLBP). We enrich Saul with components that are necessary for a broad range of learning based Natural Language Processing tasks at various levels of granularity. We illustrate these advances using three different, well-known NLP problems, and show how these generic learning and inference modules can directly exploit Saul’s graph-based data representation. These properties allow the programmer to easily switch between different model formulations and configurations, and consider various kinds of dependencies and correlations among variables of interest with minimal programming effort. We argue that Saul provides an extremely useful paradigm both for the design of advanced NLP systems and for supporting advanced research in NLP. Parisa Kordjamshidi, Daniel Khashabi, Christos Christodoulopoulos 0001, Bhargav Mangipudi, Sameer Singh 0001, Dan Roth 0001 |
COLING | 2 |
| 2016 | Question Answering via Integer Programming over Semi-Structured Knowledge
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni, Dan Roth 0001 |
IJCAI | 1 |
| 2016 | EDISON: Feature Extraction for NLP, Simplified
Mark Sammons, Christos Christodoulopoulos 0001, Parisa Kordjamshidi, Daniel Khashabi, Vivek Srikumar, Dan Roth 0001 |
LREC | 4 |
| 2015 | Solving Hard Coreference ProblemsabstractCoreference resolution is a key problem in natural language understanding that still escapes reliable solutions.One fundamental difficulty has been that of resolving instances involving pronouns since they often require deep language understanding and use of background knowledge.In this paper we propose an algorithmic solution that involves a new representation for the knowledge required to address hard coreference problems, along with a constrained optimization framework that uses this knowledge in coreference decision making.Our representation, Predicate Schemas, is instantiated with knowledge acquired in an unsupervised way, and is compiled automatically into constraints that impact the coreference decision.We present a general coreference resolution system that significantly improves state-of-the-art performance on hard, Winograd-style, pronoun resolution cases, while still performing at the stateof-the-art level on standard coreference resolution datasets. Haoruo Peng, Daniel Khashabi, Dan Roth 0001 |
HLT-NAACL | 2 |
| 2015 | Online Learning with Adversarial DelaysabstractWe study the performance of standard online learning algorithms when the feedback is delayed by an adversary. We show that \texttt{online-gradient-descent} and \texttt{follow-the-perturbed-leader} achieve regret $O(\sqrt{D})$ in the delayed setting, where $D$ is the sum of delays of each round's feedback. This bound collapses to an optimal $O(\sqrt{T})$ bound in the usual setting of no delays (where $D = T$). Our main contribution is to show that standard algorithms for online learning already have simple regret bounds in the most general setting of delayed feedback, making adjustments to the analysis and not to the algorithms themselves. Our results help affirm and clarify the success of recent algorithms in optimization and machine learning that operate in a delayed feedback model. Kent Quanrud, Daniel Khashabi |
NIPS | 2 |
| 2014 | Joint Demosaicing and Denoising via Learned Nonparametric Random FieldsabstractWe introduce a machine learning approach to demosaicing, the reconstruction of color images from incomplete color filter array samples. There are two challenges to overcome by a demosaicing method: 1) it needs to model and respect the statistics of natural images in order to reconstruct natural looking images and 2) it should be able to perform well in the presence of noise. To facilitate an objective assessment of current methods, we introduce a public ground truth data set of natural images suitable for research in image demosaicing and denoising. We then use this large data set to develop a machine learning approach to demosaicing. Our proposed method addresses both demosaicing challenges by learning a statistical model of images and noise from hundreds of natural images. The resulting model performs simultaneous demosaicing and denoising. We show that the machine learning approach has a number of benefits: 1) the model is trained to directly optimize a user-specified performance measure such as peak signal-to-noise ratio (PSNR) or structural similarity; 2) we can handle novel color filter array layouts by retraining the model on such layouts; and 3) it outperforms the previous state-of-the-art, in some setups by 0.7-dB PSNR, faithfully reconstructing edges, textures, and smooth areas. Our results demonstrate that in demosaicing and related imaging applications, discriminatively trained machine learning models have the potential for peak performance at comparatively low engineering effort. Daniel Khashabi, Sebastian Nowozin, Jeremy Jancsary, Andrew W. Fitzgibbon |
IEEE Trans. Image Process. | 1 |
| 2011 | Adaptive tiled Neural NetworksabstractIn this paper, a novel function approximation approach based on a combination of conventional Neural Networks and tile coding approximators is proposed. The proposed approach can maintain the desired features of both approaches whiles eliminates the deficiencies of each method. The combination will reduce the sharpness of tile coding. It will also provide an easy way to adjust the accuracy/complexity of the approximation according to the function being approximated (adaptive tiling) and the subspace used on. In this algorithm, it is possible to construct the approximator with specified and various approximation accuracies in different subspaces. This feature enables us to allocate an arbitrary accuracy/complexity wherever a more accurate approximation is needed. Finally simulation studies are presented to show the efficiency of and applicability of the proposed approach. Mohammad Nokhbeh-Zaeem, Daniel Khashabi, Heidar Ali Talebi, Shiva Navabi, Faramarz Vaziri |
SMC | 2 |