VLDB 2026 Research / reviewers in the wild / expert
Siva Reddy
dblp:64/8153
· DBLP profile ↗
52ranked-venue papers
6as first author
34since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 6 first-author · 34 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Warmup Generations: A Task-Agnostic Approach for Guiding Sequence-to-Sequence Learning with Unsupervised Initial State GenerationabstractTraditional supervised fine-tuning (SFT) strategies for sequence-to-sequence tasks often train models to directly generate the target output. Recent work has shown that guiding models with intermediate steps—such as keywords, outlines, or reasoning chains—can significantly improve performance, coherence, and interpretability. However, these methods often depend on predefined intermediate formats and annotated data, limiting their scalability and generalizability. In this work, we introduce a task-agnostic framework that enables models to generate intermediate “warmup” sequences. These warmup sequences, serving as an initial state for subsequent generation, are optimized to enhance the probability of generating the target sequence without relying on external supervision or human-designed structures. Drawing inspiration from reinforcement learning principles, our method iteratively refines these intermediate steps to maximize their contribution to the final output, similar to reward-driven optimization in reinforcement learning with human feedback. Experimental results across tasks such as translation, summarization, and multi-choice question answering for logical reasoning show that our approach outperforms traditional SFT methods, and offers a scalable and flexible solution for sequence-to-sequence tasks. Senyu Li, Zipeng Sun, Jiayi Wang 0010, Pontus Stenetorp, Siva Reddy, David Ifeoluwa Adelani |
ACL (1) | 6 |
| 2025 | WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code GenerationabstractRabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Rabiul Awal, Mahsa Massoud, Aarash Feizi, Suyuchen Wang, Christopher Joseph Pal, Aishwarya Agrawal, David Vázquez 0001, Siva Reddy, Juan A. Rodríguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar |
EMNLP | 9 |
| 2025 | REARANK: Reasoning Re-ranking Agent via Reinforcement LearningabstractWe present REARANK, a large language model (LLM)-based listwise reasoning reranking agent.REARANK explicitly reasons before reranking, significantly improving both performance and interpretability.Leveraging reinforcement learning and data augmentation, REARANK achieves substantial improvements over baseline models across popular information retrieval benchmarks, notably requiring only 179 annotated samples.Built on top of Qwen2.5-7B,our REARANK-7B demonstrates performance comparable to GPT-4 on both indomain and out-of-domain benchmarks and even surpasses GPT-4 on reasoning-intensive BRIGHT benchmarks.These results underscore the effectiveness of our approach and highlight how reinforcement learning can enhance LLM reasoning capabilities in reranking. Bo Wang 0084, Xipeng Qiu, Siva Reddy, Aishwarya Agrawal |
EMNLP | 4 |
| 2025 | VinePPO: Refining Credit Assignment in RL Training of LLMsabstractLarge language models (LLMs) are increasingly applied to complex reasoning tasks that require executing several complex steps before receiving any reward. Properly assigning credit to these steps is essential for enhancing model performance. Proximal Policy Optimization (PPO), a common reinforcement learning (RL) algorithm used for LLM finetuning, employs value networks to tackle credit assignment. However, recent approaches achieve strong results without it, raising questions about the efficacy of value networks in practice. In this work, we systematically evaluate the efficacy of value networks and reveal their significant shortcomings in reasoning-heavy LLM tasks, showing that they often produce poor estimate of expected return and barely outperform a random baseline when comparing alternative steps. This motivates our key question: Can improved credit assignment enhance RL training for LLMs? To address this, we propose VinePPO, a straightforward approach that leverages the flexibility of language environments to compute unbiased Monte Carlo-based estimates. Our method consistently outperforms PPO and other baselines across MATH and GSM8K datasets in less wall-clock time (up to 3.0x). Crucially, it achieves higher test accuracy for a given training accuracy, capturing more generalization signal per sample. These results emphasize the importance of accurate credit assignment in RL training of LLM. Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni, Siva Reddy, Aaron C. Courville, Nicolas Le Roux |
ICML | 5 |
| 2025 | SafeArena: Evaluating the Safety of Autonomous Web AgentsabstractLLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, a benchmark focused on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories—misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents. Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stanczak, Siva Reddy |
ICML | 9 |
| 2025 | Language Models Largely Exhibit Human-like Constituent Ordering PreferencesabstractAda Tur, Gaurav Kamath, Siva Reddy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ada Defne Tur, Gaurav Kamath, Siva Reddy |
NAACL (Long Papers) | 3 |
| 2025 | The Promise of RL for Autoregressive Image EditingabstractWhile image generation techniques are now capable of producing high-quality images that respect prompts which span multiple sentences, the task of text-guided image editing remains a challenge. Even edit requests that consist of only a few words often fail to be executed correctly. We explore three strategies to enhance performance on a wide range of image editing tasks: supervised fine-tuning (SFT), reinforcement learning (RL), and Chain-of-Thought (CoT) reasoning. In order to study all these components in one consistent framework, we adopt an autoregressive multimodal model that processes textual and visual tokens in a unified manner.
We find RL combined with a large multi-modal LLM verifier to be the most effective of these strategies.
As a result, we release **EARL**: **E**diting with **A**utoregression and **RL**, a strong RL-based image editing model that performs competitively on a diverse range of edits compared to strong baselines, despite using much less training data. Thus, EARL pushes the frontier of autoregressive multimodal models on image editing. We release our code, training data, and trained models at [https://github.com/mair-lab/EARL](https://github.com/mair-lab/EARL). Saba Ahmadi, Rabiul Awal, Ankur Sikarwar, Amirhossein Kazemnejad, Ge Ya Luo, Juan A. Rodríguez, Sai Rajeswar, Siva Reddy, Christopher Joseph Pal, Benno Krojer, Aishwarya Agrawal |
NeurIPS | 8 |
| 2024 | Benchmarking Vision Language Models for Cultural UnderstandingabstractShravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Van Steenkiste, Lisa Anne Hendricks, Karolina Stanczak, Aishwarya Agrawal. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd van Steenkiste, Lisa Anne Hendricks, Karolina Stanczak, Aishwarya Agrawal |
EMNLP | 4 |
| 2024 | WebLINX: Real-World Website Navigation with Multi-Turn DialogueabstractWe propose the problem of conversational web navigation, where a digital agent controls a web browser and follows user instructions to solve real-world tasks in a multi-turn dialogue fashion. To support this problem, we introduce WEBLINX - a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation. Our benchmark covers a broad range of patterns on over 150 real-world websites and can be used to train and evaluate agents in diverse scenarios. Due to the magnitude of information present, Large Language Models (LLMs) cannot process entire web pages in real-time. To solve this bottleneck, we design a retrieval-inspired model that efficiently prunes HTML pages by ranking relevant elements. We use the selected elements, along with screenshots and action history, to assess a variety of models for their ability to replicate human behavior when navigating the web. Our experiments span from small text-only to proprietary multimodal LLMs. We find that smaller finetuned decoders surpass the best zero-shot LLMs (including GPT-4V), but also larger finetuned multimodal models which were explicitly pretrained on screenshots. However, all finetuned models struggle to generalize to unseen websites. Our findings highlight the need for large multimodal models that can generalize to novel settings. Xing Han Lù, Zdenek Kasner, Siva Reddy |
ICML | 3 |
| 2024 | Faithfulness Measurable Masked Language ModelsabstractA common approach to explaining NLP models is to use importance measures that express which tokens are important for a prediction. Unfortunately, such explanations are often wrong despite being persuasive. Therefore, it is essential to measure their faithfulness. One such metric is if tokens are truly important, then masking them should result in worse model performance. However, token masking introduces out-of-distribution issues, and existing solutions that address this are computationally expensive and employ proxy models. Furthermore, other metrics are very limited in scope. This work proposes an inherently faithfulness measurable model that addresses these challenges. This is achieved using a novel fine-tuning method that incorporates masking, such that masking tokens become in-distribution by design. This differs from existing approaches, which are completely model-agnostic but are inapplicable in practice. We demonstrate the generality of our approach by applying it to 16 different datasets and validate it using statistical in-distribution tests. The faithfulness is then measured with 9 different importance measures. Because masking is in-distribution, importance measures that themselves use masking become consistently more faithful. Additionally, because the model makes faithfulness cheap to measure, we can optimize explanations towards maximal faithfulness; thus, our model becomes indirectly inherently explainable. Andreas Madsen, Siva Reddy, Sarath Chandar |
ICML | 2 |
| 2024 | Evaluating In-Context Learning of Libraries for Code GenerationabstractArkil Patel, Siva Reddy, Dzmitry Bahdanau, Pradeep Dasigi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Arkil Patel, Siva Reddy, Dzmitry Bahdanau, Pradeep Dasigi |
NAACL-HLT | 2 |
| 2024 | Learning Action and Reasoning-Centric Image Editing from Videos and SimulationabstractAn image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current general instruction-guided editing models have significant shortcomings with action and reasoning-centric edits.Object, attribute or stylistic changes can be learned from visually static datasets. On the other hand, high-quality data for action and reasoning-centric edits is scarce and has to come from entirely different sources that cover e.g. physical dynamics, temporality and spatial reasoning.To this end, we meticulously curate the AURORA Dataset (Action-Reasoning-Object-Attribute), a collection of high-quality training data, human-annotated and curated from videos and simulation engines.We focus on a key aspect of quality training data: triplets (source image, prompt, target image) contain a single meaningful visual change described by the prompt, i.e., truly minimal changes between source and target images.To demonstrate the value of our dataset, we evaluate an AURORA-finetuned model on a new expert-curated benchmark (AURORA-Bench) covering 8 diverse editing tasks.Our model significantly outperforms previous editing models as judged by human raters.For automatic evaluations, we find important flaws in previous metrics and caution their use for semantically hard editing tasks.Instead, we propose a new automatic metric that focuses on discriminative understanding.We hope that our efforts : (1) curating a quality training dataset and an evaluation benchmark, (2) developing critical evaluations, and (3) releasing a state-of-the-art model, will fuel further progress on general image editing. Benno Krojer, Dheeraj Vattikonda, Luis Lara, Varun Jampani, Eva Portelance, Christopher Joseph Pal, Siva Reddy |
NeurIPS | 7 |
| 2024 | Evaluating Correctness and Faithfulness of Instruction-Following Models for Question AnsweringabstractAbstract Instruction-following models are attractive alternatives to fine-tuned approaches for question answering (QA). By simply prepending relevant documents and an instruction to their input, these models can be adapted to various information domains and tasks without additional training. However, these models tend to produce verbose responses with supplementary information, which makes traditional QA metrics like exact match (EM) and F1 unreliable for accurately quantifying model performance. In this work, we evaluate instruction-following models along two fronts: 1) how well they satisfy user’s information need (correctness), and 2) whether they disseminate information supported by the provided knowledge (faithfulness). Guided by human evaluation and analysis, we highlight the shortcomings of traditional metrics for both correctness and faithfulness and propose simple token-overlap metrics that correlate highly with human judgments. Our analysis reveals that for correctness, instruction-following models perform comparably to models specifically fine-tuned for that task. However, they struggle to accurately judge the relevance of the provided knowledge and often hallucinate in their responses. We hope our work encourages more holistic evaluation of instruction-following models for QA. Our code and human annotation data is available at https://github.com/McGill-NLP/instruct-qa. Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lù, Nicholas Meade, Siva Reddy |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | Scope Ambiguities in Large Language ModelsabstractAbstract Sentences containing multiple semantic operators with overlapping scope often create ambiguities in interpretation, known as scope ambiguities. These ambiguities offer rich insights into the interaction between semantic structure and world knowledge in language processing. Despite this, there has been little research into how modern large language models treat them. In this paper, we investigate how different versions of certain autoregressive language models—GPT-2, GPT-3/3.5, Llama 2, and GPT-4—treat scope ambiguous sentences, and compare this with human judgments. We introduce novel datasets that contain a joint total of almost 1,000 unique scope-ambiguous sentences, containing interactions between a range of semantic operators, and annotated for human judgments. Using these datasets, we find evidence that several models (i) are sensitive to the meaning ambiguity in these sentences, in a way that patterns well with human judgments, and (ii) can successfully identify human-preferred readings at a high level of accuracy (over 90% in some cases).1 Gaurav Kamath, Sebastian Schuster 0001, Sowmya Vajjala, Siva Reddy |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine IntentsabstractWe introduce the StatCan Dialogue Dataset 1 consisting of 19,379 conversation turns between agents working at Statistics Canada and online users looking for published data tables.The conversations stem from genuine intents, are held in English or French, and lead to agents retrieving one of over 5000 complex data tables.Based on this dataset, we propose two tasks: (1) automatic retrieval of relevant tables based on a on-going conversation, and (2) automatic generation of appropriate agent responses at each turn.We investigate the difficulty of each task by establishing strong baselines.Our experiments on a temporal data split reveal that all models struggle to generalize to future conversations, as we observe a significant drop in performance across both tasks when we move from the validation to the test set.In addition, we find that response generation models struggle to decide when to return a table.Considering that the tasks pose significant challenges to existing models, we encourage the community to develop models for our task, which can be directly used to help knowledge workers find relevant tables for live chat users.† Work done as visiting researcher at ServiceNow Research 1 Website: mcgill-nlp.github.io/statcan-dialogue-datasetU1: Hi, I'm looking to obtain quarterly data in regards to GDP grow (Canada), BC Housing STarts, Canada Oil Price/BBL A1: Hello, my name is Xing Han Lù, Siva Reddy, Harm de Vries |
EACL | 2 |
| 2023 | Combining Parameter-efficient Modules for Task-level GeneralisationabstractA modular design encourages neural models to disentangle and recombine different facets of knowledge to generalise more systematically to new tasks.In this work, we assume that each task is associated with a subset of latent skills from an (arbitrary size) inventory.In turn, each skill corresponds to a parameter-efficient (sparse / low-rank) model adapter.By jointly learning adapters and a routing function that allocates skills to each task, the full network is instantiated as the average of the parameters of active skills.We propose several inductive biases that encourage re-usage and composition of the skills, including variable-size skill allocation and a dual-speed learning rate.We evaluate our latent-skill model in two main settings: 1) multitask reinforcement learning for instruction following on 8 levels of the BabyAI platform; and 2) few-shot fine-tuning of language models on 160 NLP tasks of the CrossFit benchmark.We find that the modular design of our network enhances sample efficiency in reinforcement learning and few-shot generalisation in supervised learning, compared to a series of baselines.These include models where parameters are fully shared, task-specific, or conditionally generated (HyperFormer), as well as sparse mixture-of-experts (Task-MoE). Edoardo Maria Ponti, Alessandro Sordoni, Yoshua Bengio, Siva Reddy |
EACL | 4 |
| 2023 | Syntactic Substitutability as Unsupervised Dependency SyntaxabstractSyntax is a latent hierarchical structure which underpins the robust and compositional nature of human language.In this work, we explore the hypothesis that syntactic dependencies can be represented in language model attention distributions and propose a new method to induce these structures theory-agnostically.Instead of modeling syntactic relations as defined by annotation schemata, we model a more general property implicit in the definition of dependency relations, syntactic substitutability.This property captures the fact that words at either end of a dependency can be substituted with words from the same category.Substitutions can be used to generate a set of syntactically invariant sentences whose representations are then used for parsing.We show that increasing the number of substitutions used improves parsing accuracy on natural data.On longdistance subject-verb agreement constructions, our method achieves 79.5% recall compared to 8.9% using a previous method.Our method also provides improvements when transferred to a different parsing setup, demonstrating that it generalizes. Jasper Jian, Siva Reddy |
EMNLP | 2 |
| 2023 | MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsabstractHumans possess a remarkable ability to assign novel interpretations to linguistic expressions, enabling them to learn new words and understand community-specific connotations.However, Large Language Models (LLMs) have a knowledge cutoff and are costly to finetune repeatedly.Therefore, it is crucial for LLMs to learn novel interpretations in-context.In this paper, we systematically analyse the ability of LLMs to acquire novel interpretations using incontext learning.To facilitate our study, we introduce MAGNIFICO, an evaluation suite implemented within a text-to-SQL semantic parsing framework that incorporates diverse tokens and prompt settings to simulate real-world complexity.Experimental results on MAGNIFICO demonstrate that LLMs exhibit a surprisingly robust capacity for comprehending novel interpretations from natural language descriptions as well as from discussions within long conversations.Nevertheless, our findings also highlight the need for further improvements, particularly when interpreting unfamiliar words or when composing multiple novel interpretations simultaneously in the same example.Additionally, our analysis uncovers the semantic predispositions in LLMs and reveals the impact of recency bias for information presented in long contexts. Arkil Patel, Satwik Bhattamishra, Siva Reddy, Dzmitry Bahdanau |
EMNLP | 3 |
| 2023 | The Impact of Positional Encoding on Length Generalization in TransformersabstractLength generalization, the ability to generalize from small training context sizes to larger ones, is a critical challenge in the development of Transformer-based language models. Positional encoding (PE) has been identified as a major factor influencing length generalization, but the exact impact of different PE schemes on extrapolation in downstream tasks remains unclear. In this paper, we conduct a systematic empirical study comparing the length generalization performance of decoder-only Transformers with five different position encoding approaches including Absolute Position Embedding (APE), T5's Relative PE, ALiBi, and Rotary, in addition to Transformers without positional encoding (NoPE). Our evaluation encompasses a battery of reasoning and mathematical tasks. Our findings reveal that the most commonly used positional encoding methods, such as ALiBi, Rotary, and APE, are not well suited for length generalization in downstream tasks. More importantly, NoPE outperforms other explicit positional encoding methods while requiring no additional computation. We theoretically demonstrate that NoPE can represent both absolute and relative PEs, but when trained with SGD, it mostly resembles T5's relative PE attention patterns. Finally, we find that scratchpad is not always helpful to solve length generalization and its format highly impacts the model's performance. Overall, our work suggests that explicit position embeddings are not essential for decoder-only Transformers to generalize well to longer sequences. Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Siva Reddy |
NeurIPS | 5 |
| 2023 | Are Diffusion Models Vision-And-Language Reasoners?abstractText-conditioned image generation models have recently shown immense qualitative success using denoising diffusion processes. However, unlike discriminative vision-and-language models, it is a non-trivial task to subject these diffusion-based generative models to automatic fine-grained quantitative evaluation of high-level phenomena such as compositionality.
Towards this goal, we perform two innovations. First, we transform diffusion-based models (in our case, Stable Diffusion) for any image-text matching (ITM) task using a novel method called DiffusionITM.
Second, we introduce the Generative-Discriminative Evaluation Benchmark (GDBench) benchmark with 7 complex vision-and-language tasks, bias evaluation and detailed analysis.
We find that Stable Diffusion + DiffusionITM is competitive on many tasks and outperforms CLIP on compositional tasks like like CLEVR and Winoground.
We further boost its compositional performance with a transfer setup by fine-tuning on MS-COCO while retaining generative capabilities.
We also measure the stereotypical bias in diffusion models, and find that Stable Diffusion 2.1 is, for the most part, less biased than Stable Diffusion 1.5.
Overall, our results point in an exciting direction bringing discriminative and generative model evaluation closer. We will release code and benchmark setup soon. Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Christopher Joseph Pal, Siva Reddy |
NeurIPS | 5 |
| 2022 | Compositional Generalization in Dependency ParsingabstractCompositionality-the ability to combine familiar units like words into novel phrases and sentences-has been the focus of intense interest in artificial intelligence in recent years.To test compositional generalization in semantic parsing, Keysers et al. (2020) introduced Compositional Freebase Queries (CFQ).This dataset maximizes the similarity between the test and train distributions over primitive units, like words, while maximizing the compound divergence-the dissimilarity between test and train distributions over larger structures, like phrases.Dependency parsing, however, lacks a compositional generalization benchmark.In this work, we introduce a gold-standard set of dependency parses for CFQ, and use this to analyze the behavior of a state-of-the art dependency parser (Qi et al., 2020) on the CFQ dataset.We find that increasing compound divergence degrades dependency parsing performance, although not as dramatically as semantic parsing performance.Additionally, we find the performance of the dependency parser does not uniformly degrade relative to compound divergence, and the parser performs differently on different splits with the same compound divergence.We explore a number of hypotheses for what causes the non-uniform degradation in dependency parsing performance, and identify a number of syntactic structures that drive the dependency parser's lower performance on the most challenging splits. Emily Goodwin, Siva Reddy, Timothy J. O'Donnell, Dzmitry Bahdanau |
ACL (1) | 2 |
| 2022 | Image Retrieval from Contextual DescriptionsabstractBenno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Ponti, Siva Reddy. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Maria Ponti, Siva Reddy |
ACL (1) | 6 |
| 2022 | An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsabstractRecent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on.This has attracted attention to developing techniques that mitigate such biases.In this work, we perform an empirical survey of five recently proposed bias mitigation techniques: Counterfactual Data Augmentation (CDA), Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebias.We quantify the effectiveness of each technique using three intrinsic bias benchmarks while also measuring the impact of these techniques on a model's language modeling ability, as well as its performance on downstream NLU tasks.We experimentally find that: (1) Self-Debias is the strongest debiasing technique, obtaining improved scores on all bias benchmarks; (2) Current debiasing techniques perform less consistently when mitigating non-gender biases; And (3) improvements on bias benchmarks such as StereoSet and CrowS-Pairs by using debiasing strategies are often accompanied by a decrease in language modeling ability, making it difficult to determine whether the bias mitigation was effective.1 Nicholas Meade, Elinor Poole-Dayan, Siva Reddy |
ACL (1) | 3 |
| 2022 | IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesabstractReliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community. Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic |
ICML | 4 |
| 2022 | On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models?abstractNouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, Siva Reddy. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nouha Dziri, Sivan Milton, Mo Yu, Osmar R. Zaïane, Siva Reddy |
NAACL-HLT | 5 |
| 2022 | TopiOCQA: Open-domain Conversational Question Answering with Topic SwitchingabstractAbstract In a conversational question answering scenario, a questioner seeks to extract information about a topic through a series of interdependent questions and answers. As the conversation progresses, they may switch to related topics, a phenomenon commonly observed in information-seeking search sessions. However, current datasets for conversational question answering are limiting in two ways: 1) they do not contain topic switches; and 2) they assume the reference text for the conversation is given, that is, the setting is not open-domain. We introduce TopiOCQA (pronounced Tapioca), an open-domain conversational dataset with topic switches based on Wikipedia. TopiOCQA contains 3,920 conversations with information-seeking questions and free-form answers. On average, a conversation in our dataset spans 13 question-answer turns and involves four topics (documents). TopiOCQA poses a challenging test-bed for models, where efficient retrieval is required on multiple turns of the same conversation, in conjunction with constructing valid responses using conversational history. We evaluate several baselines, by combining state-of-the-art document retrieval methods with neural reader models. Our best model achieves F1 of 55.8, falling short of human performance by 14.2 points, indicating the difficulty of our dataset. Our dataset and code are available at https://mcgill-nlp.github.io/topiocqa. Vaibhav Adlakha, Shehzaad Dhuliawala, Kaheer Suleman, Harm de Vries, Siva Reddy |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | FaithDial: A Faithful Benchmark for Information-Seeking DialogueabstractAbstract The goal of information-seeking dialogue is to respond to seeker queries with natural language utterances that are grounded on knowledge sources. However, dialogue systems often produce unsupported utterances, a phenomenon known as hallucination. To mitigate this behavior, we adopt a data-centric solution and create FaithDial, a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark. We observe that FaithDial is more faithful than WoW while also maintaining engaging conversations. We show that FaithDial can serve as training signal for: i) a hallucination critic, which discriminates whether an utterance is faithful or not, and boosts the performance by 12.8 F1 score on the BEGIN benchmark compared to existing datasets for dialogue coherence; ii) high-quality dialogue generation. We benchmark a series of state-of-the-art models and propose an auxiliary contrastive objective that achieves the highest level of faithfulness and abstractiveness based on several automated metrics. Further, we find that the benefits of FaithDial generalize to zero-shot transfer on other datasets, such as CMU-Dog and TopicalChat. Finally, human evaluation reveals that responses generated by models trained on FaithDial are perceived as more interpretable, cooperative, and engaging. Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar R. Zaïane, Mo Yu, Edoardo Maria Ponti, Siva Reddy |
Trans. Assoc. Comput. Linguistics | 7 |
| 2021 | StereoSet: Measuring stereotypical bias in pretrained language modelsabstractMoin Nadeem, Anna Bethke, Siva Reddy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Moin Nadeem, Anna Bethke, Siva Reddy |
ACL/IJCNLP (1) | 3 |
| 2021 | Visually Grounded Reasoning across Languages and CulturesabstractThe design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems. Fangyu Liu 0001, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott |
EMNLP (1) | 4 |
| 2021 | Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring ExpressionsabstractNeural module networks (NMN) are a popular approach for grounding visual referring expressions.Prior implementations of NMN use pre-defined and fixed textual inputs in their module instantiation.This necessitates a large number of modules as they lack the ability to share weights and exploit associations between similar textual contexts (e.g."dark cube on the left" vs. "black cube on the left").In this work, we address these limitations and evaluate the impact of contextual clues in improving the performance of NMN models.First, we address the problem of fixed textual inputs by parameterizing the module arguments.This substantially reduce the number of modules in NMN by up to 75% without any loss in performance.Next we propose a method to contextualize our parameterized model to enhance the module's capacity in exploiting the visiolinguistic associations.Our model outperforms the state-of-the-art NMN model on CLEVR-Ref+ dataset with +8.1% improvement in accuracy on the single-referent test set and +4.3% on the full test set.Additionally, we demonstrate that contextualization provides +11.2% and +1.7% improvements in accuracy over prior NMN models on CLO-SURE and NLVR2.We further evaluate the impact of our contextualization by constructing a contrast set for CLEVR-Ref+, which we call CC-Ref+.We significantly outperform the baselines by as much as +10.4% absolute accuracy on CC-Ref+, illustrating the generalization skills of our approach.Our dataset is publicly available at https://github.com/ McGill-NLP/contextual-nmn. Arjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu, Siva Reddy |
EMNLP (1) | 5 |
| 2021 | Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage RetrievalabstractIn this work, we introduce back-training, an alternative to self-training for unsupervised domain adaptation (UDA) from source to target domain.While self-training generates synthetic training data where natural inputs are aligned with noisy outputs, back-training results in natural outputs aligned with noisy inputs.This significantly reduces the gap between the target domain and synthetic data distribution, and reduces model overfitting to the source domain.We run UDA experiments on question generation and passage retrieval from the Natural Questions domain to machine learning and biomedical domains.We find that back-training vastly outperforms selftraining by a mean improvement of 7.8 BLEU-4 points on generation, and 17.6% top-20 retrieval accuracy across both domains.We further propose consistency filters to remove low-quality synthetic data before training.We also release a new domain-adaptation dataset-MLQuestions containing 35K unaligned questions, 50K unaligned passages, and 3K aligned question-passage pairs. Devang Kulshreshtha, Robert Belfer, Iulian Serban, Siva Reddy |
EMNLP (1) | 4 |
| 2021 | Understanding by Understanding Not: Modeling Negation in Language ModelsabstractArian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, Aaron Courville. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Seyed Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R. Devon Hjelm, Alessandro Sordoni, Aaron C. Courville |
NAACL-HLT | 2 |
| 2021 | Explicitly Modeling Syntax in Language Models with Incremental Parsing and a Dynamic OracleabstractYikang Shen, Shawn Tan, Alessandro Sordoni, Siva Reddy, Aaron Courville. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yikang Shen, Shawn Tan, Alessandro Sordoni, Siva Reddy, Aaron C. Courville |
NAACL-HLT | 4 |
| 2021 | End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question AnsweringabstractWe present an end-to-end differentiable training method for retrieval-augmented open-domain question answering systems that combine information from multiple retrieved documents when generating answers. We model retrieval decisions as latent variables over sets of relevant documents. Since marginalizing over sets of retrieved documents is computationally hard, we approximate this using an expectation-maximization algorithm. We iteratively estimate the value of our latent variable (the set of relevant documents for a given question) and then use this estimate to update the retriever and reader parameters. We hypothesize that such end-to-end training allows training signals to flow to the reader and then to the retriever better than staged-wise training. This results in a retriever that is able to select more relevant documents for a question and a reader that is trained on more accurate documents to generate an answer. Experiments on three benchmark datasets demonstrate that our proposed method outperforms all existing approaches of comparable size by 2-3% absolute exact match points, achieving new state-of-the-art results. Our results also demonstrate the feasibility of learning to retrieve to improve answer generation without explicit supervision of retrieval decisions. Devendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer, Dani Yogatama |
NeurIPS | 2 |
| 2020 | Words Aren't Enough, Their Order Matters: On the Robustness of Grounding Visual Referring ExpressionsabstractVisual referring expression recognition is a challenging task that requires natural language understanding in the context of an image.We critically examine RefCOCOg, a standard benchmark for this task, using a human study and show that 83.7% of test instances do not require reasoning on linguistic structure, i.e., words are enough to identify the target object, the word order doesn't matter.To measure the true progress of existing models, we split the test set into two sets, one which requires reasoning on linguistic structure and the other which doesn't.Additionally, we create an out-of-distribution dataset Ref-Adv by asking crowdworkers to perturb in-domain examples such that the target object changes.Using these datasets, we empirically show that existing methods fail to exploit linguistic structure and are 12% to 23% lower in performance than the established progress for this task.We also propose two methods, one based on contrastive learning and the other based on multi-task learning, to increase the robustness of ViLBERT, the current state-ofthe-art model for this task.Our datasets are publicly Arjun R. Akula, Spandana Gella, Yaser Al-Onaizan, Song-Chun Zhu, Siva Reddy |
ACL | 5 |
| 2020 | Measuring Systematic Generalization in Neural Proof Generation with TransformersabstractWe are interested in understanding how well Transformer language models (TLMs) can perform reasoning tasks when trained on knowledge encoded in the form of natural language. We investigate their systematic generalization abilities on a logical reasoning task in natural language, which involves reasoning over relationships between entities grounded in first-order logical proofs. Specifically, we perform soft theorem-proving by leveraging TLMs to generate natural language proofs. We test the generated proofs for logical consistency, along with the accuracy of the final inference. We observe length-generalization issues when evaluated on longer-than-trained sequences. However, we observe TLMs improve their generalization performance after being exposed to longer, exhaustive proofs. In addition, we discover that TLMs are able to generalize better using backward-chaining proofs compared to their forward-chaining counterparts, while they find it easier to generate forward chaining proofs. We observe that models that are not trained to generate proofs are better at generalizing to problems based on longer proofs. This suggests that Transformers have efficient internal reasoning strategies that are harder to interpret. These results highlight the systematic generalization behavior of TLMs in the context of logical reasoning, and we believe this work motivates deeper inspection of their underlying reasoning strategies. Nicolas Angelard-Gontier, Koustuv Sinha, Siva Reddy, Christopher Joseph Pal |
NeurIPS | 3 |
| 2019 | Learning an Executable Neural Semantic ParserabstractThis article describes a neural semantic parser that maps natural language utterances onto logical forms that can be executed against a task-specific environment, such as a knowledge base or a database, to produce a response. The parser generates tree-structured logical forms with a transition-based approach, combining a generic tree-generation algorithm with domain-general grammar defined by the logical language. The generation process is modeled by structured recurrent neural networks, which provide a rich encoding of the sentential context and generation history for making predictions. To tackle mismatches between natural language and logical form tokens, various attention mechanisms are explored. Finally, we consider different training settings for the neural semantic parser, including fully supervised training where annotated logical forms are given, weakly supervised training where denotations are provided, and distant supervision where only unlabeled sentences and a knowledge base are available. Experiments across a wide range of data sets demonstrate the effectiveness of our parser. Jianpeng Cheng 0001, Siva Reddy, Vijay A. Saraswat, Mirella Lapata |
Comput. Linguistics | 2 |
| 2019 | CoQA: A Conversational Question Answering ChallengeabstractHumans gather information through conversations involving a series of interconnected questions and answers. For machines to assist in information gathering, it is therefore essential to enable them to answer conversational questions. We introduce CoQA, a novel dataset for building Conversational Question Answering systems. Our dataset contains 127k questions with answers, obtained from 8k conversations about text passages from seven diverse domains. The questions are conversational, and the answers are free-form text with their corresponding evidence highlighted in the passage. We analyze CoQA in depth and show that conversational questions have challenging phenomena not present in existing reading comprehension datasets (e.g., coreference and pragmatic reasoning). We evaluate strong dialogue and reading comprehension models on CoQA. The best system obtains an F1 score of 65.4%, which is 23.4 points behind human performance (88.8%), indicating that there is ample room for improvement. We present CoQA as a challenge to the community at https://stanfordnlp.github.io/coqa . Siva Reddy, Danqi Chen 0001, Christopher D. Manning |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | Learning Typed Entailment Graphs with Global Soft ConstraintsabstractThis paper presents a new method for learning typed entailment graphs from text. We extract predicate-argument structures from multiple-source news corpora, and compute local distributional similarity scores to learn entailments between predicates with typed arguments (e.g., person contracted disease). Previous work has used transitivity constraints to improve local decisions, but these constraints are intractable on large graphs. We instead propose a scalable method that learns globally consistent similarity scores based on new soft constraints that consider both the structures across typed entailment graphs and inside each graph. Learning takes only a few hours to run over 100K predicates and our results show large improvements over local similarity scores on two entailment data sets. We further show improvements over paraphrases and entailments from the Paraphrase Database, and prior state-of-the-art entailment graphs. We show that the entailment graphs improve performance in a downstream task. Mohammad Javad Hosseini, Nathanael Chambers, Siva Reddy, Xavier R. Holt, Shay B. Cohen, Mark Johnson 0001, Mark Steedman |
Trans. Assoc. Comput. Linguistics | 3 |
| 2017 | Learning Structured Natural Language Representations for Semantic ParsingabstractWe introduce a neural semantic parser which is interpretable and scalable.Our model converts natural language utterances to intermediate, domain-general natural language representations in the form of predicate-argument structures, which are induced with a transition system and subsequently mapped to target domains.The semantic parser is trained end-to-end using annotated logical forms or their denotations.We achieve the state of the art on SPADES and GRAPHQUESTIONS and obtain competitive results on GEO-QUERY and WEBQUESTIONS.The induced predicate-argument structures shed light on the types of representations useful for semantic parsing and how these are different from linguistically motivated ones. 1 Jianpeng Cheng 0001, Siva Reddy, Vijay A. Saraswat, Mirella Lapata |
ACL (1) | 2 |
| 2017 | Learning to Paraphrase for Question AnsweringabstractQuestion answering (QA) systems are sensitive to the many different ways natural language expresses the same information need.In this paper we turn to paraphrases as a means of capturing this knowledge and present a general framework which learns felicitous paraphrases for various QA tasks.Our method is trained end-toend using question-answer pairs as a supervision signal.A question and its paraphrases serve as input to a neural scoring model which assigns higher weights to linguistic expressions most likely to yield correct answers.We evaluate our approach on QA over Freebase and answer sentence selection.Experimental results on three datasets show that our framework consistently improves performance, achieving competitive results despite the use of simple QA models. Li Dong 0004, Jonathan Mallinson, Siva Reddy, Mirella Lapata |
EMNLP | 3 |
| 2017 | Universal Semantic ParsingabstractUniversal Dependencies (UD) offer a uniform cross-lingual syntactic representation, with the aim of advancing multilingual applications.Recent work shows that semantic parsing can be accomplished by transforming syntactic dependencies to logical forms.However, this work is limited to English, and cannot process dependency graphs, which allow handling complex phenomena such as control.In this work, we introduce UDEPLAMBDA, a semantic interface for UD, which maps natural language to logical forms in an almost language-independent fashion and can process dependency graphs.We perform experiments on question answering against Freebase and provide German and Spanish translations of the WebQuestions and GraphQuestions datasets to facilitate multilingual evaluation.Results show that UDEPLAMBDA outperforms strong baselines across languages and datasets.For English, it achieves a 4.9 F 1 point improvement over the state-of-the-art on Graph-Questions.ENTITY ⇒ λx.word(x a ); e.g.Oscar ⇒ λx.Oscar(x a ) EVENT ⇒ λx.word(x e ); e.g. won ⇒ λx.won(x e ) FUNCTIONAL ⇒ λx.TRUE; e.g. an ⇒ λx.TRUE COPY ⇒ λ f gx.∃y.f (x) ∧ g(y) ∧ rel(x, y) e.g.nsubj, dobj, nmod, advmod INVERT ⇒ λ f gx.∃y.f (x) ∧ g(y) ∧ rel i (y, x) e.g.amod, acl MERGE ⇒ λ f gx.f (x) ∧ g(x) e.g.compound, appos, amod, acl HEAD ⇒ λ f gx.f (x) e.g.case, punct, aux, mark . Siva Reddy, Oscar Täckström, Slav Petrov, Mark Steedman, Mirella Lapata |
EMNLP | 1 |
| 2016 | Question Answering on Freebase via Relation Extraction and Textual EvidenceabstractExisting knowledge-based question answering systems often rely on small annotated training data.While shallow methods like relation extraction are robust to data scarcity, they are less expressive than the deep meaning representation methods like semantic parsing, thereby failing at answering questions involving multiple constraints.Here we alleviate this problem by empowering a relation extraction method with additional evidence from Wikipedia.We first present a neural network based relation extractor to retrieve the candidate answers from Freebase, and then infer over Wikipedia to validate these answers.Experiments on the WebQuestions question answering dataset show that our method achieves an F 1 of 53.3%, a substantial improvement over the state-of-the-art. Kun Xu 0005, Siva Reddy, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001 |
ACL (1) | 2 |
| 2016 | Evaluating Induced CCG Parsers on Grounded Semantic ParsingabstractWe compare the effectiveness of four different syntactic CCG parsers for a semantic slotfilling task to explore how much syntactic supervision is required for downstream semantic analysis.This extrinsic, task-based evaluation also provides a unique window into the semantics captured (or missed) by unsupervised grammar induction systems. Yonatan Bisk, Siva Reddy, John Blitzer, Julia Hockenmaier, Mark Steedman |
EMNLP | 2 |
| 2016 | Paraphrase Generation from Latent-Variable PCFGs for Semantic ParsingabstractOne of the limitations of semantic parsing approaches to open-domain question answering is the lexicosyntactic gap between natural language questions and knowledge base entries -there are many ways to ask a question, all with the same answer.In this paper we propose to bridge this gap by generating paraphrases of the input question with the goal that at least one of them will be correctly mapped to a knowledge-base query.We introduce a novel grammar model for paraphrase generation that does not require any sentence-aligned paraphrase corpus.Our key idea is to leverage the flexibility and scalability of latent-variable probabilistic context-free grammars to sample paraphrases.We do an extrinsic evaluation of our paraphrases by plugging them into a semantic parser for Freebase.Our evaluation experiments on the WebQuestions benchmark dataset show that the performance of the semantic parser improves over strong baselines. Shashi Narayan, Siva Reddy, Shay B. Cohen |
INLG | 2 |
| 2016 | Assessing Relative Sentence Complexity using an Incremental CCG ParserabstractGiven a pair of sentences, we present computational models to assess if one sentence is simpler to read than the other.While existing models explored the usage of phrase structure features using a non-incremental parser, experimental evidence suggests that the human language processor works incrementally.We empirically evaluate if syntactic features from an incremental CCG parser are more useful than features from a non-incremental phrase structure parser.Our evaluation on Simple and Standard Wikipedia sentence pairs suggests that incremental CCG features are indeed more useful than phrase structure features achieving 0.44 points gain in performance.Incremental CCG parser also gives significant improvements in speed (12 times faster) in comparison to the phrase structure parser.Furthermore, with the addition of psycholinguistic features, we achieve the strongest result to date reported on this task. Bharat Ram Ambati, Siva Reddy, Mark Steedman |
HLT-NAACL | 2 |
| 2016 | Transforming Dependency Structures to Logical Forms for Semantic ParsingabstractThe strongly typed syntax of grammar formalisms such as CCG, TAG, LFG and HPSG offers a synchronous framework for deriving syntactic structures and semantic logical forms. In contrast—partly due to the lack of a strong type system—dependency structures are easy to annotate and have become a widely used form of syntactic analysis for many languages. However, the lack of a type system makes a formal mechanism for deriving logical forms from dependency structures challenging. We address this by introducing a robust system based on the lambda calculus for deriving neo-Davidsonian logical forms from dependency trees. These logical forms are then used for semantic parsing of natural language to Freebase. Experiments on the Free917 and Web-Questions datasets show that our representation is superior to the original dependency trees and that it outperforms a CCG-based representation on this task. Compared to prior work, we obtain the strongest result to date on Free917 and competitive results on WebQuestions. Siva Reddy, Oscar Täckström, Michael Collins 0001, Tom Kwiatkowski, Dipanjan Das 0001, Mark Steedman, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Large-scale Semantic Parsing without Question-Answer PairsabstractIn this paper we introduce a novel semantic parsing approach to query Freebase in natural language without requiring manual annotations or question-answer pairs. Our key insight is to represent natural language via semantic graphs whose topology shares many commonalities with Freebase. Given this representation, we conceptualize semantic parsing as a graph matching problem. Our model converts sentences to semantic graphs using CCG and subsequently grounds them to Freebase guided by denotations as a form of weak supervision. Evaluation experiments on a subset of the Free917 and WebQuestions benchmark datasets show our semantic parser improves over the state of the art. Siva Reddy, Mirella Lapata, Mark Steedman |
Trans. Assoc. Comput. Linguistics | 1 |
| 2012 | Word Sketches for Turkish
Bharat Ram Ambati, Siva Reddy, Adam Kilgarriff |
LREC | 2 |
| 2011 | Dynamic and Static Prototype Vectors for Semantic Composition
Siva Reddy, Ioannis P. Klapaftis, Diana McCarthy, Suresh Manandhar |
IJCNLP | 1 |
| 2011 | An Empirical Study on Compositionality in Compound Nouns
Siva Reddy, Diana McCarthy, Suresh Manandhar |
IJCNLP | 1 |
| 2010 | A Corpus Factory for Many Languages
Adam Kilgarriff, Siva Reddy, Jan Pomikálek, P. V. S. Avinesh |
LREC | 2 |