Semih Yavuz

dblp:129/1217 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
abstract
The large language model (LLM)-as-judge paradigm has been used to meet the demand for a cheap, reliable, and fast evaluation of model outputs during AI system development and post-deployment monitoring. While judge models—LLMs finetuned to specialize in assessing and critiquing model outputs—have been touted as general purpose evaluators, they are typically evaluated only on non-contextual scenarios, such as instruction following. The omission of contextual settings—those where external information is used as context to generate an output—is surprising given the increasing prevalence of retrieval-augmented generation (RAG) and summarization use cases. Contextual assessment is uniquely challenging, as evaluation often depends on practitioner priorities, leading to conditional evaluation criteria (e.g., comparing responses based on factuality and then considering completeness if they are equally factual). To address the gap, we propose ContextualJudgeBench, a judge benchmark with 2,000 challenging response pairs across eight splits inspired by real-world contextual evaluation scenarios. We build our benchmark with a multi-pronged data construction pipeline that leverages both existing human annotations and model-based perturbations. Our comprehensive study across 11 judge models and 7 general purpose models, reveals that the contextual information and assessment criteria present a significant challenge to even state-of-the-art models. For example, o1, the best-performing model, barely reaches 55% consistent accuracy.
Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, Shafiq R. Joty
ACL (1)4
2025 VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
abstract
Embedding models play a crucial role in a variety of downstream tasks, including semantic similarity, information retrieval, and clustering. While there has been a surge of interest in developing universal text embedding models that generalize across tasks (e.g., MTEB), progress in learning universal multimodal embedding models has been comparatively slow, despite their importance and practical applications. In this work, we explore the potential of building universal multimodal embeddings capable of handling a broad range of downstream tasks. Our contributions are twofold: (1) we propose MMEB (Massive Multimodal Embedding Benchmark), which covers four meta-tasks (classification, visual question answering, multimodal retrieval, and visual grounding) and 36 datasets, including 20 training datasets and 16 evaluation datasets spanning both in-distribution and out-of-distribution tasks, and (2) VLM2Vec (Vision-Language Model → Vector), a contrastive training framework that transforms any vision-language model into an embedding model through contrastive training on MMEB. Unlike previous models such as CLIP and BLIP, which encode text and images independently without task-specific guidance, VLM2Vec can process any combination of images and text while incorporating task instructions to generate a fixed-dimensional vector. We develop a series of VLM2Vec models based on state-of-the-art VLMs, including Phi-3.5-V, LLaVA-1.6, and Qwen2-VL, and evaluate them on MMEB’s benchmark. With LoRA tuning, VLM2Vec achieves a 10% to 20% improvement over existing multimodal embedding models on MMEB’s evaluation sets. Our findings reveal that VLMs are surprisingly strong embedding models.
Ziyan Jiang, Xinyi Yang 0002, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen
ICLR4
2025 Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
abstract
Contrastive learning (CL) is a prevalent technique for training embedding models, which pulls semantically similar examples (positives) closer in the representation space while pushing dissimilar ones (negatives) further apart. A key source of negatives are "in-batch" examples, i.e., positives from other examples in the batch. Effectiveness of such models is hence strongly influenced by the size and quality of training batches. In this work, we propose *Breaking the Batch Barrier* (B3), a novel batch construction strategy designed to curate high-quality batches for CL. Our approach begins by using a pretrained teacher embedding model to rank all examples in the dataset, from which a sparse similarity graph is constructed. A community detection algorithm is then applied to this graph to identify clusters of examples that serve as strong negatives for one another. The clusters are then used to construct batches that are rich in in-batch negatives. Empirical results on the MMEB multimodal embedding benchmark (36 tasks) demonstrate that our method sets a new state of the art, outperforming previous best methods by +1.3 and +2.9 points at the 7B and 2B model scales, respectively. Notably, models trained with B3 surpass existing state-of-the-art results even with a batch size as small as 64, which is 4–16× smaller than that required by other methods. Moreover, experiments show that B3 generalizes well across domains and tasks, maintaining strong performance even when trained with considerably weaker teachers.
Raghuveer Thirukovalluru, Ye Liu 0006, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen, Bhuwan Dhingra
NeurIPS7
2024 FOLIO: Natural Language Reasoning with First-Order Logic
abstract
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev
EMNLP27
2024 Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
abstract
Large Language Models (LLMs) have demonstrated a powerful ability for text generation.However, achieving optimal results with a given prompt or instruction can be challenging, especially for billion-sized models.Additionally, undesired behaviors such as toxicity or hallucinations can manifest.While much larger models (e.g., ChatGPT) may demonstrate strength in mitigating these issues, there is still no guarantee of complete prevention.In this work, we propose formalizing text generation as a future-constrained generation problem to minimize undesirable behaviors and enforce faithfulness to instructions.The estimation of future constraint satisfaction, accomplished using LLMs, guides the text generation process.Our extensive experiments demonstrate the effectiveness of the proposed approach across three distinct text generation tasks: keywordconstrained generation (Lin et al., 2020), toxicity reduction (Gehman et al., 2020), and factual correctness in question-answering (Gao et al., 2023). 1
Lifu Tu, Semih Yavuz, Jin Qu, Jiacheng Xu 0001, Caiming Xiong, Yingbo Zhou 0002
EMNLP2
2024 L2CEval: Evaluating Language-to-Code Generation Capabilities of Large Language Models
abstract
Abstract Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs. Despite promising results, there is a notable lack of a comprehensive evaluation of these models’ language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning, and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition, we assess confidence calibration, and conduct human evaluations to identify typical failures across different tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We release the evaluation framework1 and all model outputs, hoping to lay the groundwork for further future research. All future evaluations (e.g., LLaMA-3, StarCoder2, etc) will be updated on the project website: https://l2c-eval.github.io/.
Ansong Ni, Yilun Zhao 0001, Martin Riddell, Troy Feng, Stephen Yin, Ye Liu 0006, Semih Yavuz, Caiming Xiong, Shafiq R. Joty, Yingbo Zhou 0002, Dragomir R. Radev, Arman Cohan
Trans. Assoc. Comput. Linguistics9
2023 Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors
abstract
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu 0001, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, Greg Durrett
ACL (1)6
2022 Modeling Multi-hop Question Answering as Single Sequence Prediction
abstract
Fusion-in-decoder (FID) (Izacard and Grave, 2021) is a generative question answering (QA) model that leverages passage retrieval with a pre-trained transformer and pushed the state of the art on single-hop QA.However, the complexity of multi-hop QA hinders the effectiveness of the generative QA approach.In this work, we propose a simple generative approach (PATHFID) that extends the task beyond just answer generation by explicitly modeling the reasoning process to resolve the answer for multihop questions.By linearizing the hierarchical reasoning path of supporting passages, their key sentences, and finally the factoid answer, we cast the problem as a single sequence prediction task.To facilitate complex reasoning with multiple clues, we further extend the unified flat representation of multiple input documents by encoding cross-passage interactions.Our extensive experiments demonstrate that PATHFID leads to strong performance gains on two multihop QA datasets: HotpotQA and IIRC.Besides the performance gains, PATHFID is more interpretable, which in turn yields answers that are more faithfully grounded to the supporting passages and facts compared to the baseline FID model.
Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Nitish Shirish Keskar, Caiming Xiong
ACL (1)1
2022 RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering
abstract
Existing KBQA approaches, despite achieving strong performance on i.i.d.test data, often struggle in generalizing to questions involving unseen KB schema items.Prior rankingbased approaches have shown some success in generalization, but suffer from the coverage issue.We present RnG-KBQA, a Rank-and-Generate approach for KBQA, which remedies the coverage issue with a generation model while preserving a strong generalization capability.Our approach first uses a contrastive ranker to rank a set of candidate logical forms obtained by searching over the knowledge graph.It then introduces a tailored generation model conditioned on the question and the top-ranked candidates to compose the final logical form.We achieve new state-ofthe-art results on GRAILQA and WEBQSP datasets.In particular, our method surpasses the prior state-of-the-art by a large margin on the GRAILQA leaderboard.In addition, RnG-KBQA outperforms all prior approaches on the popular WEBQSP benchmark, even including the ones that use the oracle entity linking.The experimental results demonstrate the effectiveness of the interplay between ranking and generation, which leads to the superior performance of our proposed approach across all settings with especially strong improvements in zero-shot generalization. 1 * Work done during internship at Salesforce Research. 1 Code available at https://github.com/salesforce/rng-kbqa.
Xi Ye 0003, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Caiming Xiong
ACL (1)2
2022 Uni-Parser: Unified Semantic Parser for Question Answering on Knowledge Base and Database
abstract
Parsing natural language questions into executable logical forms is a useful and interpretable way to perform question answering on structured data such as knowledge bases (KB) or databases (DB).However, existing approaches on semantic parsing cannot adapt to both modalities, as they suffer from the exponential growth of the logical form candidates and can hardly generalize to unseen data.In this work, we propose Uni-Parser, a unified semantic parser for question answering (QA) on both KB and DB.We introduce the primitive (relation and entity in KB, and table name, column name and cell value in DB) as an essential element in our framework.The number of primitives grows linearly with the number of retrieved relations in KB and DB, preventing us from dealing with exponential logic form candidates.We leverage the generator to predict final logical forms by altering and composing topranked primitives with different operations (e.g.select, where, count).With sufficiently pruned search space by a contrastive primitive ranker, the generator is empowered to capture the composition of primitives enhancing its generalization ability.We achieve competitive results on multiple KB and DB QA benchmarks more efficiently, especially in the compositional and zero-shot settings.
Ye Liu 0006, Semih Yavuz, Dragomir R. Radev, Caiming Xiong, Yingbo Zhou 0002
EMNLP2
2021 Unsupervised Paraphrasing with Pretrained Language Models
abstract
Paraphrase generation has benefited extensively from recent progress in the designing of training objectives and model architectures.However, previous explorations have largely focused on supervised methods, which require a large amount of labeled data that is costly to collect.To address this drawback, we adopt a transfer learning approach and propose a training pipeline that enables pre-trained language models to generate high-quality paraphrases in an unsupervised setting.Our recipe consists of task-adaptation, self-supervision, and a novel decoding algorithm named Dynamic Blocking (DB).To enforce a surface form dissimilar from the input, whenever the language model emits a token contained in the source sequence, DB prevents the model from outputting the subsequent source token for the next generation step.We show with automatic and human evaluations that our approach achieves state-of-the-art performance on both the Quora Question Pair (QQP) and the ParaNMT datasets and is robust to domain shift between the two datasets of distinct distributions.We also demonstrate that our model transfers to paraphrasing in other languages without any additional finetuning.
Semih Yavuz, Yingbo Zhou 0002, Nitish Shirish Keskar, Huan Wang 0016, Caiming Xiong
EMNLP (1)2
2021 CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers
Semih Yavuz, Kazuma Hashimoto, Jia Li 0015, Nazneen Fatema Rajani, Xifeng Yan, Yingbo Zhou 0002, Caiming Xiong
ICLR2
2020 Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging
abstract
The concept of Dialogue Act (DA) is universal across different task-oriented dialogue domains -the act of "request" carries the same speaker intention whether it is for restaurant reservation or flight booking.However, DA taggers trained on one domain do not generalize well to other domains, which leaves us with the expensive need for a large amount of annotated data in the target domain.In this work, we investigate how to better adapt DA taggers to desired target domains with only unlabeled data.We propose MASKAUGMENT, a controllable mechanism that augments text input by leveraging the pre-trained MASK token from BERT model.Inspired by consistency regularization, we use MASKAUGMENT to introduce an unsupervised teacher-student learning scheme to examine the domain adaptation of DA taggers.Our extensive experiments on the Simulated Dialogue (GSim) and Schema-Guided Dialogue (SGD) datasets show that MASKAUGMENT is useful in improving the cross-domain generalization for DA tagging.
Semih Yavuz, Kazuma Hashimoto, Wenhao Liu 0003, Nitish Shirish Keskar, Richard Socher, Caiming Xiong
EMNLP (1)1
2020 A Simple Language Model for Task-Oriented Dialogue
abstract
Task-oriented dialogue is often decomposed into three tasks: understanding user input, deciding actions, and generating a response. While such decomposition might suggest a dedicated model for each sub-task, we find a simple, unified approach leads to state-of-the-art performance on the MultiWOZ dataset. SimpleTOD is a simple approach to task-oriented dialogue that uses a single, causal language model trained on all sub-tasks recast as a single sequence prediction problem. This allows SimpleTOD to fully leverage transfer learning from pre-trained, open domain, causal language models such as GPT-2. SimpleTOD improves over the prior state-of-the-art in joint goal accuracy for dialogue state tracking, and our analysis reveals robustness to noisy annotations in this setting. SimpleTOD also improves the main metrics used to evaluate action decisions and response generation in an end-to-end setting: inform rate by 8.1 points, success rate by 9.7 points, and combined score by 7.2 points.
Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, Richard Socher
NeurIPS4
2019 Monotonic Infinite Lookback Attention for Simultaneous Machine Translation
abstract
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, Colin Raffel. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li 0133, Colin Raffel
ACL (1)5
2019 Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset
abstract
Bill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, Andy Cedilnik. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
William J. Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, Andy Cedilnik
EMNLP/IJCNLP (1)7
2019 HierCon: Hierarchical Organization of Technical Documents Based on Concepts
abstract
In this work we study the hierarchical organization of technical documents, where given a set of documents and a hierarchy of categories, the goal is to assign documents to their corresponding categories. Unlike prior work on supervised hierarchical document categorization that relies on large amount of labeled training data, which is expensive to obtain in closed technical domain and tends to stale as new knowledge emerges, we study this problem in a weak supervision setting, by leveraging semantic information from concepts. The core idea is to project both documents and categories into a common concept embedding space, where their fine-grained similarity can be easily and effectively computed. Experiments over real-world datasets from the subject of computer science, physics & mathematics, and medicine demonstrated the superior performance of our approach over a wide range of state of the art baseline approaches.
Keqian Li, Semih Yavuz, Hanwen Zha, Yu Su 0001, Xifeng Yan
ICDM3
2019 Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
abstract
Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans.Video question answering is a specific scenario of such AI-human interaction where an agent generates a natural language response to a question regarding the video of a dynamic scene.Incorporating features from multiple modalities, which often provide supplementary information, is one of the challenging aspects of video question answering.Furthermore, a question often concerns only a small segment of the video, hence encoding the entire video sequence using a recurrent neural network is not computationally efficient.Our proposed question-guided video representation module efficiently generates the token-level video summary guided by each word in the question.The learned representations are then fused with the question to generate the answer.Through empirical evaluation on the Audio Visual Scene-aware Dialog (AVSD) dataset (Alamri et al., 2019a), our proposed models in single-turn and multiturn question answering achieve state-of-theart performance on several automatic natural language generation evaluation metrics.
Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz, Dilek Hakkani-Tür, Jindong Chen, Ian Lane
SIGdial3
2019 DeepCopy: Grounded Response Generation with Hierarchical Pointer Networks
abstract
Recent advances in neural sequence-tosequence models have led to promising results for several language generation-based tasks, including dialogue response generation, summarization, and machine translation.However, these models are known to have several problems, especially in the context of chit-chat based dialogue systems: they tend to generate short and dull responses that are often too generic.Furthermore, these models do not ground conversational responses on knowledge and facts, resulting in turns that are not accurate, informative and engaging for the users.In this paper, we propose and experiment with a series of response generation models that aim to serve in the general scenario where in addition to the dialogue context, relevant unstructured external knowledge in the form of text is also assumed to be available for models to harness.Our proposed approach extends pointer-generator networks (See et al., 2017) by allowing the decoder to hierarchically attend and copy from external knowledge in addition to the dialogue context.We empirically show the effectiveness of the proposed model compared to several baselines including (Ghazvininejad et al., 2018;Zhang et al., 2018) through both automatic evaluation metrics and human evaluation on CONVAI2 dataset.
Semih Yavuz, Abhinav Rastogi, Guan-Lin Chao, Dilek Hakkani-Tür
SIGdial1
2018 DialSQL: Dialogue Based Structured Query Generation
abstract
The recent advance in deep learning and semantic parsing has significantly improved the translation accuracy of natural language questions to structured queries.However, further improvement of the existing approaches turns out to be quite challenging.Rather than solely relying on algorithmic innovations, in this work, we introduce DialSQL, a dialoguebased structured query generation framework that leverages human intelligence to boost the performance of existing algorithms via user interaction.DialSQL is capable of identifying potential errors in a generated SQL query and asking users for validation via simple multi-choice questions.User feedback is then leveraged to revise the query.We design a generic simulator to bootstrap synthetic training dialogues and evaluate the performance of DialSQL on the WikiSQL dataset.Using SQLNet as a black box query generation tool, DialSQL improves its performance from 61.3% to 69.0% using only 2.4 validation questions per dialogue.
Izzeddin Gur, Semih Yavuz, Yu Su 0001, Xifeng Yan
ACL (1)2
2018 CaLcs: Continuously Approximating Longest Common Subsequence for Sequence Level Optimization
abstract
Maximum-likelihood estimation (MLE) is one of the most widely used approaches for training structured prediction models for textgeneration based natural language processing applications.However, besides exposure bias, models trained with MLE suffer from wrong objective problem where they are trained to maximize the word-level correct next step prediction, but are evaluated with respect to sequence-level discrete metrics such as ROUGE and BLEU.Several variants of policy-gradient methods address some of these problems by optimizing for final discrete evaluation metrics and showing improvements over MLE training for downstream tasks like text summarization and machine translation.However, policy-gradient methods suffers from high sample variance, making the training process very difficult and unstable.In this paper, we present an alternative direction towards mitigating this problem by introducing a new objective (CALCS) based on a differentiable surrogate of longest common subsequence (LCS) measure that captures sequence-level structure similarity.Experimental results on abstractive summarization and machine translation validate the effectiveness of the proposed approach.
Semih Yavuz, Chung-Cheng Chiu, Patrick Nguyen
EMNLP1
2018 What It Takes to Achieve 100 Percent Condition Accuracy on WikiSQL
abstract
WikiSQL is a newly released dataset for studying the natural language sequence to SQL translation problem.The SQL queries in Wik-iSQL are simple: Each involves one relation and does not have any join operation.Despite of its simplicity, none of the publicly reported structured query generation models can achieve an accuracy beyond 62%, which is still far from enough for practical use.In this paper, we ask two questions, "Why is the accuracy still low for such simple queries?" and "What does it take to achieve 100% accuracy on WikiSQL?"To limit the scope of our study, we focus on the WHERE clause in SQL.The answers will help us gain insights about the directions we should explore in order to further improve the translation accuracy.We will then investigate alternative solutions to realize the potential ceiling performance on WikiSQL.Our proposed solution can reach up to 88.6% condition accuracy on the WikiSQL dataset.
Semih Yavuz, Izzeddin Gur, Yu Su 0001, Xifeng Yan
EMNLP1
2018 Global Relation Embedding for Relation Extraction
abstract
Yu Su, Honglei Liu, Semih Yavuz, Izzeddin Gür, Huan Sun, Xifeng Yan. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Yu Su 0001, Honglei Liu 0001, Semih Yavuz, Izzeddin Gur, Huan Sun 0001, Xifeng Yan
NAACL-HLT3
2017 Recovering Question Answering Errors via Query Revision
abstract
The existing factoid QA systems often lack a post-inspection component that can help models recover from their own mistakes.In this work, we propose to crosscheck the corresponding KB relations behind the predicted answers and identify potential inconsistencies.Instead of developing a new model that accepts evidences collected from these relations, we choose to plug them back to the original questions directly and check if the revised question makes sense or not.A bidirectional LSTM is applied to encode revised questions.We develop a scoring mechanism over the revised question encodings to refine the predictions of a base QA system.This approach can improve the F 1 score of STAGG (Yih et al., 2015), one of the leading QA systems, from 52.5% to 53.9% on WE-BQUESTIONS data.
Semih Yavuz, Izzeddin Gur, Yu Su 0001, Xifeng Yan
EMNLP1
2016 Improving Semantic Parsing via Answer Type Inference
abstract
In this work, we show the possibility of inferring the answer type before solving a factoid question and leveraging the type information to improve semantic parsing.By replacing the topic entity in a question with its type, we are able to generate an abstract form of the question, whose answer corresponds to the answer type of the original question.A bidirectional LSTM model is built to train over the abstract form of questions and infer their answer types.It is also observed that if we convert a question into a statement form, our LSTM model achieves better accuracy.Using the predicted type information to rerank the logical forms returned by AgendaIL, one of the leading semantic parsers, we are able to improve the F1-score from 49.7% to 52.6% on the WE-BQUESTIONS data.
Semih Yavuz, Izzeddin Gur, Yu Su 0001, Mudhakar Srivatsa, Xifeng Yan
EMNLP1
2013 Capacity region of multi-resolution streaming in peer-to-peer networks
abstract
We consider multi-resolution streaming in fully-connected peer-to-peer networks, where transmission rates are constrained by arbitrarily specified upload capacities of the source and peers. We fully characterize the capacity region of rate vectors achievable with arbitrary coding, where an achievable rate vector describes a vector of throughputs of the different resolutions that can be supported by the network. We then prove that all rate vectors in the capacity region can be achieved using pure routing strategies. This shows that coding has no capacity advantage over routing in this scenario.
Batuhan Karagöz, Semih Yavuz, Tracey Ho, Michelle Effros
ITW2