VLDB 2026 Research / reviewers in the wild / expert
Abbas Ghaddar
dblp:184/8839
· DBLP profile ↗
17ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMsabstractPost-training hybridization of large language models (LLMs) often replaces quadratic selfattention with sliding-window attention (SWA) to reduce KV cache usage and improve latency.Existing hybridization schemes are typically defined either at the layer level (e.g., interleaving) or at the head level via static rankings from local to global.Layer-level schemes ignore that local and global dependencies are routed through heads within the same layer, while static head-level rankings suffer from entanglement: a head's local/global behavior can change after hybridization.We propose BOSCH, Black-box Binary Optimization for Short-context Head Selection, a training-free method that formulates the problem as a Large Neighborhood Search and decomposes it into three subproblems: (i) layer-importance detection via small-budget black-box probes, (ii) adaptive per-layer SWA-ratio assignment based on these sensitivities, and (iii) grouped headlevel optimization within ratio buckets.Extensive experiments on 4 LLMs ranging from 1.7B to 30B parameters, across 4 SWA ratios, show that BOSCH consistently outperforms layerlevel heuristics and 6 strong static head-level methods, with larger gains at higher SWA ratios.Under continual pretraining, BOSCH recover original long-context performance faster and to a higher level.Analysis of the selected heads reveals substantial turnover for BOSCH across different SWA ratios, underscoring the importance of performing head-level selection for each target ratio rather than relying on fixed locality rankings. Abbas Ghaddar, Ivan Kobyzev, Boxing Chen, Yufei Cui |
ACL (1) | 1 |
| 2025 | Integral Transformer: Denoising Attention, Not Too Much Not Too LittleabstractSoftmax self-attention often assigns disproportionate weight to semantically uninformative tokens such as special tokens and punctuation, a phenomenon known as attention noise.While recent methods like Cog Attention and the Differential Transformer have addressed this by introducing negative attention scores, they risk discarding useful information.In this paper, we propose the Integral Transformer, a novel selfattention mechanism that denoises attention by integrating signals sampled from the logit distribution.Our approach mitigates noise while preserving the contributions of special tokens critical for model performance.Extensive experiments demonstrate that our model outperforms vanilla, Cog, and Differential attention variants on well-established knowledge and reasoning language benchmarks.Moreover, our analysis reveals that employing vanilla self-attention in the lower Transformer layers enhances performance and that the Integral Transformer effectively balances attention distributions and reduces rank collapse in upper layers. Ivan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing Chen |
EMNLP | 2 |
| 2024 | EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering SystemsabstractMohammad Dehghan, Mohammad Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang, Xiaoguang Li, Jianye Hao, Qun Liu, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mohammad Dehghan, Mohammad Ali Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang 0001, Jianye Hao, Qun Liu 0001, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh |
ACL (1) | 6 |
| 2024 | CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational SearchabstractIn this paper, we study how open-source large language models (LLMs) can be effectively deployed for improving query rewriting in conversational search, especially for ambiguous queries.We introduce CHIQ, a two-step method that leverages the capabilities of LLMs to resolve ambiguities in the conversation history before query rewriting.This approach contrasts with prior studies that predominantly use closed-source LLMs to directly generate search queries from conversation history.We demonstrate on five well-established benchmarks that CHIQ leads to state-of-the-art results across most settings, showing highly competitive performances with systems leveraging closed-source LLMs.Our study provides a first step towards leveraging open-source LLMs in conversational search, as a competitive alternative to the prevailing reliance on commercial LLMs for query rewriting. Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh, Boxing Chen, Qun Liu 0001, Jian-Yun Nie |
EMNLP | 2 |
| 2023 | On the utility of enhancing BERT syntactic bias with Token Reordering PretrainingabstractYassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Phillippe Langlais, Prasanna Parthasarathi. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. Yassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Philippe Langlais, Prasanna Parthasarathi |
CoNLL | 3 |
| 2022 | CILDA: Contrastive Data Augmentation Using Intermediate Layer Knowledge DistillationabstractKnowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by leveraging Contrastive Learning, Intermediate Layer Distillation, Data Augmentation, and Adversarial Training. In this work, we propose a learning-based data augmentation technique tailored for knowledge distillation, called CILDA. To the best of our knowledge, this is the first time that intermediate layer representations of the main task are used in improving the quality of augmented samples. More precisely, we introduce an augmentation technique for KD based on intermediate layer matching using contrastive loss to improve masked adversarial data augmentation. CILDA outperforms existing state-of-the-art KD approaches on the GLUE benchmark, as well as in an out-of-domain evaluation. Md. Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar, Khalil Bibi, Philippe Langlais, Pascal Poupart |
COLING | 3 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 1 |
| 2022 | Inferring a spatial code of cell-cell interactions across a whole animal bodyabstractCell-cell interactions shape cellular function and ultimately organismal phenotype. Interacting cells can sense their mutual distance using combinations of ligand-receptor pairs, suggesting the existence of a spatial code, i.e., signals encoding spatial properties of cellular organization. However, this code driving and sustaining the spatial organization of cells remains to be elucidated. Here we present a computational framework to infer the spatial code underlying cell-cell interactions from the transcriptomes of the cell types across the whole body of a multicellular organism. As core of this framework, we introduce our tool cell2cell, which uses the coexpression of ligand-receptor pairs to compute the potential for intercellular interactions, and we test it across the Caenorhabditis elegans' body. Leveraging a 3D atlas of C. elegans' cells, we also implement a genetic algorithm to identify the ligand-receptor pairs most informative of the spatial organization of cells across the whole body. Validating the spatial code extracted with this strategy, the resulting intercellular distances are negatively correlated with the inferred cell-cell interactions. Furthermore, for selected cell-cell and ligand-receptor pairs, we experimentally confirm the communicatory behavior inferred with cell2cell and the genetic algorithm. Thus, our framework helps identify a code that predicts the spatial organization of cells across a whole-animal body. Erick Armingol, Abbas Ghaddar, Chintan J. Joshi, Hratch Baghdassarian, Isaac Shamie, Hsuan-Lin Her, Samuel Berhanu, Anushka Dar, Fabiola Rodriguez-Armstrong, Olivia Yang, Eyleen J. O'Rourke, Nathan E. Lewis |
PLoS Comput. Biol. | 2 |
| 2021 | Towards Zero-Shot Knowledge Distillation for Natural Language ProcessingabstractKnowledge distillation (KD) is a common knowledge transfer algorithm used for model compression across a variety of deep learning based natural language processing (NLP) solutions.In its regular manifestations, KD requires access to the teacher's training data for knowledge transfer to the student network.However, privacy concerns, data regulations and proprietary reasons may prevent access to such data.We present, to the best of our knowledge, the first work on Zero-shot Knowledge Distillation for NLP, where the student learns from the much larger teacher without any task specific data.Our solution combines out-ofdomain data and adversarial training to learn the teacher's output distribution.We investigate six tasks from the GLUE benchmark and demonstrate that we can achieve between 75% and 92% of the teacher's classification score (accuracy or F1) while compressing the model 30 times.* Work done during an internship at Huawei Noah's Ark Lab. Ahmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi Rezagholizadeh |
EMNLP (1) | 3 |
| 2021 | Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationabstractIntermediate layer matching is shown as an effective approach for improving knowledge distillation (KD).However, this technique applies matching in the hidden spaces of two different networks (i.e.student and teacher), which lacks clear interpretability.Moreover, intermediate layer KD cannot easily deal with other problems such as layer mapping search and architecture mismatch (i.e. it requires the teacher and student to be of the same model type).To tackle the aforementioned problems all together, we propose Universal-KD to match intermediate layers of the teacher and the student in the output space (by adding pseudo classifiers on intermediate layers) via the attention-based layer projection.By doing this, our unified approach has three merits: (i) it can be flexibly combined with current intermediate layer distillation techniques to improve their results (ii) the pseudo classifiers of the teacher can be deployed instead of extra expensive teacher assistant networks to address the capacity gap problem in KD which is a common issue when the gap between the size of the teacher and student networks becomes too large; (iii) it can be used in cross-architecture intermediate layer KD.We did comprehensive experiments in distilling BERT-base into BERT-4, RoBERTa-large into DistilRoBERTa and BERT-base into CNN and LSTM-based models.Results on the GLUE tasks show that our approach is able to outperform other KD techniques. Yimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar, Ali Ghodsi 0001 |
EMNLP (1) | 3 |
| 2021 | Context-aware Adversarial Training for Name Regularity Bias in Named Entity RecognitionabstractAbstract In this work, we examine the ability of NER models to use contextual information when predicting the type of an ambiguous entity. We introduce NRB, a new testbed carefully designed to diagnose Name Regularity Bias of NER models. Our results indicate that all state-of-the-art models we tested show such a bias; BERT fine-tuned models significantly outperforming feature-based (LSTM-CRF) ones on NRB, despite having comparable (sometimes lower) performance on standard benchmarks. To mitigate this bias, we propose a novel model-agnostic training method that adds learnable adversarial noise to some entity mentions, thus enforcing models to focus more strongly on the contextual signal, leading to significant gains on NRB. Combining it with two other training strategies, data augmentation and parameter freezing, leads to further gains. Abbas Ghaddar, Philippe Langlais, Ahmad Rashid, Mehdi Rezagholizadeh |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | SEDAR: a Large Scale French-English Financial Domain Parallel CorpusabstractThis paper describes the acquisition, preprocessing and characteristics of SEDAR, a large scale English-French parallel corpus for the financial domain. Our extensive experiments on machine translation show that SEDAR is essential to obtain good performance on finance. We observe a large gain in the performance of machine translation systems trained on SEDAR when tested on finance, which makes SEDAR suitable to study domain adaptation for neural machine translation. The first release of the corpus comprises 8.6 million high quality sentence pairs that are publicly available for research at https://github.com/autorite/sedar-bitext. Abbas Ghaddar, Philippe Langlais |
LREC | 1 |
| 2018 | Robust Lexical Features for Improved Neural Network Named-Entity RecognitionabstractNeural network approaches to Named-Entity Recognition reduce the need for carefully hand-crafted features. While some features do remain in state-of-the-art systems, lexical features have been mostly discarded, with the exception of gazetteers. In this work, we show that this is unfair: lexical features are actually quite useful. We propose to embed words and entity types into a low-dimensional vector space we train from annotated data produced by distant supervision thanks to Wikipedia. From this, we compute — offline — a feature vector representing each word. When used with a vanilla recurrent neural network model, this representation yields substantial improvements. We establish a new state-of-the-art F1 score of 87.95 on ONTONOTES 5.0, while matching state-of-the-art performance with a F1 score of 91.73 on the over-studied CONLL-2003 dataset. Abbas Ghaddar, Philippe Langlais |
COLING | 1 |
| 2018 | Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus
Abbas Ghaddar, Philippe Langlais |
LREC | 1 |
| 2017 | WiNER: A Wikipedia Annotated Corpus for Named Entity RecognitionabstractWe revisit the idea of mining Wikipedia in order to generate named-entity annotations. We propose a new methodology that we applied to English Wikipedia to build WiNER, a large, high quality, annotated corpus. We evaluate its usefulness on 6 NER tasks, comparing 4 popular state-of-the art approaches. We show that LSTM-CRF is the approach that benefits the most from our corpus. We report impressive gains with this model when using a small portion of WiNER on top of the CONLL training material. Last, we propose a simple but efficient method for exploiting the full range of WiNER, leading to further improvements. Abbas Ghaddar, Philippe Langlais |
IJCNLP(1) | 1 |
| 2016 | Coreference in Wikipedia: Main Concept ResolutionabstractWikipedia is a resource of choice exploited in many NLP applications, yet we are not aware of recent attempts to adapt coreference resolution to this resource. In this work, we revisit a seldom studied task which consists in identifying in a Wikipedia article all the mentions of the main concept being described. We show that by exploiting the Wikipedia markup of a document, as well as links to external knowledge bases such as Freebase, we can acquire useful information on entities that helps to classify mentions as coreferent or not. We designed a classifier which drastically outperforms fair baselines built on top of state-of-the-art coreference resolution systems. We also measure the benefits of this classifier in a full coreference resolution pipeline applied to Wikipedia texts. Abbas Ghaddar, Philippe Langlais |
CoNLL | 1 |
| 2016 | WikiCoref: An English Coreference-annotated Corpus of Wikipedia Articles
Abbas Ghaddar, Philippe Langlais |
LREC | 1 |