VLDB 2026 Research / reviewers in the wild / expert
Harish Tayyar Madabushi
dblp:186/7335
· DBLP profile ↗
16ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-5260-3653ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMsabstractHaritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi, Iryna Gurevych. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haritz Puerto, Tilek Chubakov, Xiaodan Zhu 0001, Harish Tayyar Madabushi, Iryna Gurevych |
ACL (1) | 4 |
| 2025 | UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency AssessmentabstractJoseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds 0001, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi |
EMNLP | 19 |
| 2024 | Are Emergent Abilities in Large Language Models just In-Context Learning?abstractSheng Lu, Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi, Iryna Gurevych. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi, Iryna Gurevych |
ACL (1) | 4 |
| 2024 | A Construction Grammar Corpus of Varying Schematicity: A Dataset for the Evaluation of Abstractions in Language ModelsabstractLarge Language Models (LLMs) have been developed without a theoretical framework, yet we posit that evaluating and improving LLMs will benefit from the development of theoretical frameworks that enable comparison of the structures of human language and the model of language built up by LLMs through the processing of text. In service of this goal, we develop the Construction Grammar Schematicity (“CoGS”) corpus of 10 distinct English constructions, where the constructions vary with respect to schematicity, or in other words the level to which constructional slots require specific, fixed lexical items, or can be filled with a variety of elements that fulfill a particular semantic role of the slot. Our corpus constructions are carefully curated to range from substantive, frozen constructions (e.g., Let-alone) to entirely schematic constructions (e.g., Resultative). The corpus was collected to allow us to probe LLMs for constructional information at varying levels of abstraction. We present our own probing experiments using this corpus, which clearly demonstrate that even the largest LLMs are limited to more substantive constructions and do not exhibit recognition of the similarity of purely schematic constructions. We publicly release our dataset, prompts, and associated model responses. Claire Bonial, Harish Tayyar Madabushi |
LREC/COLING | 2 |
| 2024 | Pre-Trained Language Models Represent Some Geographic Populations Better than OthersabstractThis paper measures the skew in how well two families of LLMs represent diverse geographic populations. A spatial probing task is used with geo-referenced corpora to measure the degree to which pre-trained language models from the OPT and BLOOM series represent diverse populations around the world. Results show that these models perform much better for some populations than others. In particular, populations across the US and the UK are represented quite well while those in South and Southeast Asia are poorly represented. Analysis shows that both families of models largely share the same skew across populations. At the same time, this skew cannot be fully explained by sociolinguistic factors, economic factors, or geographic factors. The basic conclusion from this analysis is that pre-trained models do not equally represent the world’s population: there is a strong skew towards specific geographic populations. This finding challenges the idea that a single model can be used for all populations. Jonathan Dunn, Benjamin Adams, Harish Tayyar Madabushi |
LREC/COLING | 3 |
| 2024 | Code-Mixed Probes Show How Pre-Trained Models Generalise on Code-Switched TextabstractCode-switching is a prevalent linguistic phenomenon in which multilingual individuals seamlessly alternate between languages. Despite its widespread use online and recent research trends in this area, research in code-switching presents unique challenges, primarily stemming from the scarcity of labelled data and available resources. In this study we investigate how pre-trained Language Models handle code-switched text in three dimensions: a) the ability of PLMs to detect code-switched text, b) variations in the structural information that PLMs utilise to capture code-switched text, and c) the consistency of semantic information representation in code-switched text. To conduct a systematic and controlled evaluation of the language models in question, we create a novel dataset of well-formed naturalistic code-switched text along with parallel translations into the source languages. Our findings reveal that pre-trained language models are effective in generalising to code-switched text, shedding light on abilities of these models to generalise representations to CS corpora. We release all our code and data, including the novel corpus, at https://github.com/francesita/code-mixed-probes. Frances Adriana Laureano De Leon, Harish Tayyar Madabushi, Mark Lee 0001 |
LREC/COLING | 2 |
| 2024 | LongEval: Longitudinal Evaluation of Model Performance at CLEF 2024
Rabab Alkhalifa, Hsuvas Borkakoty, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Tobias Fink, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, David Iommi, Maria Liakata, Harish Tayyar Madabushi, Pablo Medina-Alias, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga |
ECIR (6) | 12 |
| 2024 | Standardize: Aligning Language Models with Expert-Defined Standards for Content GenerationabstractDomain experts across engineering, healthcare, and education follow strict standards for producing quality content such as technical manuals, medication instructions, and children's reading materials.However, current works in controllable text generation have yet to explore using these standards as references for control.Towards this end, we introduce STANDARD-IZE, a retrieval-style in-context learning-based framework to guide large language models to align with expert-defined standards.Focusing on English language standards in the education domain as a use case, we consider the Common European Framework of Reference for Languages (CEFR) and Common Core Standards (CCS) for the task of open-ended content generation.Our findings show that models can gain 45% to 100% increase in precise accuracy across open and commercial LLMs evaluated, demonstrating that the use of knowledge artifacts extracted from standards and integrating them in the generation process can effectively guide models to produce better standardaligned content. 1 Joseph Marvin Imperial, Gail Forey, Harish Tayyar Madabushi |
EMNLP | 3 |
| 2023 | LongEval: Longitudinal Evaluation of Model Performance at CLEF 2023
Rabab Alkhalifa, Iman Munire Bilal, Hsuvas Borkakoty, José Camacho-Collados, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, Elena Kochkina, Maria Liakata, Daniel Loureiro, Harish Tayyar Madabushi, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga |
ECIR (3) | 14 |
| 2022 | Improving Tokenisation by Alternative Treatment of SpacesabstractTokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text.Existing algorithms have problems, often producing tokenisations of limited linguistic validity and representing equivalent strings differently depending on their position within a word.We hypothesise that these problems hinder the ability of transformer-based models to handle complex words, and suggest that these problems are a result of allowing tokens to include spaces.We thus experiment with an alternative tokenisation approach where spaces are always treated as individual tokens.Specifically, we apply this modification to the BPE and Unigram algorithms.We find that our modified algorithms lead to improved performance on downstream NLP tasks that involve handling complex words, whilst having no detrimental effect on performance in general natural language understanding tasks.Intrinsically, we find that our modified algorithms give more morphologically correct tokenisations, in particular when handling prefixes.Given the results of our experiments, we advocate for always treating spaces as individual tokens as an improved tokenisation method. Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline Villavicencio |
EMNLP | 2 |
| 2022 | Abstraction not Memory: BERT and the English Article SystemabstractHarish Tayyar Madabushi, Dagmar Divjak, Petar Milin. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Harish Tayyar Madabushi, Dagmar Divjak, Petar Milin |
NAACL-HLT | 1 |
| 2021 | Learned Construction Grammars Converge Across Registers Given Increased ExposureabstractThis paper measures the impact of increased exposure on whether learned construction grammars converge onto shared representations when trained on data from different registers.Register influences the frequency of constructions, with some structures common in formal but not informal usage.We expect that a grammar induction algorithm exposed to different registers will acquire different constructions.To what degree does increased exposure lead to the convergence of register-specific grammars?The experiments in this paper simulate language learning in 12 languages (half Germanic and half Romance) with corpora representing three registers (Twitter, Wikipedia, Web).These simulations are repeated with increasing amounts of exposure, from 100k to 2 million words, to measure the impact of exposure on the convergence of grammars.The results show that increased exposure does lead to converging grammars across all languages.In addition, a shared core of register-universal constructions remains constant across increasing amounts of exposure. Jonathan Dunn, Harish Tayyar Madabushi |
CoNLL | 2 |
| 2020 | CxGBERT: BERT meets Construction GrammarabstractWhile lexico-semantic elements no doubt capture a large amount of linguistic information, it has been argued that they do not capture all information contained in text.This assumption is central to constructionist approaches to language which argue that language consists of constructions, learned pairings of a form and a function or meaning that are either frequent or have a meaning that cannot be predicted from its component parts.BERT's training objectives give it access to a tremendous amount of lexico-semantic information, and while BERTology has shown that BERT captures certain important linguistic dimensions, there have been no studies exploring the extent to which BERT might have access to constructional information.In this work we design several probes and conduct extensive experiments to answer this question.Our results allow us to conclude that BERT does indeed have access to a significant amount of information, much of which linguists typically call constructional information.The impact of this observation is potentially far-reaching as it provides insights into what deep learning methods learn from text, while also showing that information contained in constructions is redundantly encoded in lexicosemantics. Harish Tayyar Madabushi, Laurence Romain, Dagmar Divjak, Petar Milin |
COLING | 1 |
| 2020 | Multi-class Hierarchical Question Classification for Multiple Choice Science ExamsabstractPrior work has demonstrated that question classification (QC), recognizing the problem domain of a question, can help answer it more accurately. However, developing strong QC algorithms has been hindered by the limited size and complexity of annotated data available. To address this, we present the largest challenge dataset for QC, containing 7,787 science exam questions paired with detailed classification labels from a fine-grained hierarchical taxonomy of 406 problem domains. We then show that a BERT-based model trained on this dataset achieves a large (+0.12 MAP) gain compared with previous methods, while also achieving state-of-the-art performance on benchmark open-domain and biomedical QC datasets. Finally, we show that using this model’s predictions of question topic significantly improves the accuracy of a question answering system by +1.7% P@1, with substantial future gains possible as QC performance improves. Dongfang Xu, Peter A. Jansen, Jaycie Martin, Zhengnan Xie, Vikas Yadav, Harish Tayyar Madabushi, Oyvind Tafjord, Peter Clark |
LREC | 6 |
| 2018 | Integrating Question Classification and Deep Learning for improved Answer SelectionabstractWe present a system for Answer Selection that integrates fine-grained Question Classification with a Deep Learning model designed for Answer Selection. We detail the necessary changes to the Question Classification taxonomy and system, the creation of a new Entity Identification system and methods of highlighting entities to achieve this objective. Our experiments show that Question Classes are a strong signal to Deep Learning models for Answer Selection, and enable us to outperform the current state of the art in all variations of our experiments except one. In the best configuration, our MRR and MAP scores outperform the current state of the art by between 3 and 5 points on both versions of the TREC Answer Selection test set, a standard dataset for this task. Harish Tayyar Madabushi, Mark Lee 0001, John A. Barnden |
COLING | 1 |
| 2016 | High Accuracy Rule-based Question Classification using Question Syntax and SemanticsabstractWe present in this paper a purely rule-based system for Question Classification which we divide into two parts: The first is the extraction of relevant words from a question by use of its structure, and the second is the classification of questions based on rules that associate these words to Concepts. We achieve an accuracy of 97.2%, close to a 6 point improvement over the previous State of the Art of 91.6%. Additionally, we believe that machine learning algorithms can be applied on top of this method to further improve accuracy. Harish Tayyar Madabushi, Mark Lee 0001 |
COLING | 1 |