Muhammad Khalifa

dblp:246/4401 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 On Many-Shot In-Context Learning for Long-Context Evaluation
abstract
Many-shot in-context learning (ICL) has emerged as a unique setup to both utilize and test the ability of large language models to handle long context.This paper delves into long-context language model (LCLM) evaluation through many-shot ICL.We first ask: what types of ICL tasks benefit from additional demonstrations, and how effective are they in evaluating LCLMs?We find that classification and summarization tasks show performance improvements with additional demonstrations, while translation and reasoning tasks do not exhibit clear trends.Next, we investigate the extent to which different tasks necessitate retrieval versus global context understanding.We develop metrics to categorize ICL tasks into two groups: (i) similar-sample learning (SSL): tasks where retrieval of the most similar examples is sufficient for good performance, and (ii) all-sample learning (ASL): tasks that necessitate a deeper comprehension of all examples in the prompt.Lastly, we introduce a new many-shot ICL benchmark built on existing ICL tasks, MANYICLBENCH, to characterize model's ability on both fronts and benchmark 12 LCLMs using MANYICLBENCH.We find that while state-of-the-art models demonstrate good performance up to 64k tokens in SSL tasks, many models experience significant performance drops at only 16k tokens in ASL tasks.
Kaijian Zou, Muhammad Khalifa, Lu Wang 0008
ACL (1)2
2025 MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
abstract
We introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies.Unlike prior work, e.g., AI Scientist, which evaluates the end-to-end agentic pipeline by using LLM-as-a-judge, MLRC-Bench measures the key steps of proposing and implementing novel research methods and evaluates them with rigorous protocol and objective metrics.Our curated suite of 7 competition tasks reveals significant challenges for LLM agents. Even the best-performing tested agent (gemini-exp-1206 under MLAB) closes only 9.3% of the gap between baseline and top human participant scores.Furthermore, our analysis reveals a misalignment between the LLM-judged innovation and their actual performance on cutting-edge ML research problems. MLRC-Bench is a dynamic benchmark, which is designed to continually grow with new ML competitions to encourage rigorous and objective evaluations of AI’s research capabilities. Our leaderboard and code are publicly available at https://huggingface.co/spaces/launch/MLRC_Bench.
Muhammad Khalifa, Shitanshu Bhushan, Grant D. Murphy, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, Lu Wang 0008
NeurIPS2
2024 Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models
abstract
Recent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that “it’s all been solved.” Not surprisingly, this has, in turn, made many NLP researchers – especially those at the beginning of their careers – worry about what NLP research area they should focus on. Has it all been solved, or what remaining questions can we work on regardless of LLMs? To address this question, this paper compiles NLP research directions rich for exploration. We identify fourteen different research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. While we identify many research areas, many others exist; we do not cover areas currently addressed by LLMs, but where LLMs lag behind in performance or those focused on LLM development. We welcome suggestions for other research directions to include: https://bit.ly/nlp-era-llm.
Oana Ignat, Zhijing Jin 0001, Artem Abzaliev, Laura Biester, Santiago Castro, Naihao Deng, Xinyi Gao 0004, Aylin Gunal, Jacky He, Ashkan Kazemi, Muhammad Khalifa, Namho Koh, Andrew Lee 0001, Siyang Liu 0003, Do June Min, Shinka Mori, Joan Nwatu, Verónica Pérez-Rosas, Zekun Wang 0002, Winston Wu, Rada Mihalcea
LREC/COLING11
2024 LitCab: Lightweight Language Model Calibration over Short- and Long-form Responses
abstract
A model is considered well-calibrated when its probability estimate aligns with the actual likelihood of the output being correct. Calibrating language models (LMs) is crucial, as it plays a vital role in detecting and mitigating hallucinations of LMs as well as building more trustworthy models. However, standard calibration techniques may not be suited for LM calibration. For instance, post-processing methods such as temperature scaling do not reorder the candidate generations. On the other hand, training-based methods require fine-tuning the entire model, which is impractical for LMs of large scale. We present LitCab, a lightweight calibration mechanism consisting of a single linear layer that takes the input text representation and predicts a bias term, which is then added to the LM output logits. LitCab improves model calibration by only adding < 2% of the original model parameters. For evaluation, we construct CaT, a benchmark consisting of eight text generation tasks, covering responses ranging from short phrases to paragraphs. We test LitCab with Llama2-7B, where it improves calibration across all tasks, reducing the average ECE score by as large as 30%. We further conduct a comprehensive evaluation with multiple popular open-sourced LMs from GPT and LLaMA families, yielding the following key findings: (i) Larger models within the same family exhibit better calibration on tasks with short generation tasks, but not necessarily for longer ones. (ii) GPT-family models show superior calibration compared to LLaMA, Llama2, and Vicuna models, despite having much fewer parameters. (iii) Fine-tuning pretrained model (e.g., LLaMA) with samples of limited purpose (e.g., conversations) may lead to worse calibration, highlighting the importance of fine-tuning setups for calibrating LMs.
Muhammad Khalifa, Lu Wang 0008
ICLR2
2024 Learning to Reason via Program Generation, Emulation, and Search
abstract
Program synthesis with language models (LMs) has unlocked a large set of reasoning abilities; code-tuned LMs have proven adept at generating programs that solve a wide variety of algorithmic symbolic manipulation tasks (e.g. word concatenation). However, not all reasoning tasks are easily expressible as code, e.g. tasks involving commonsense reasoning, moral decision-making, and sarcasm understanding. Our goal is to extend a LM’s program synthesis skills to such tasks and evaluate the results via pseudo-programs, namely Python programs where some leaf function calls are left undefined. To that end, we propose, Code Generation and Emulated EXecution (COGEX). COGEX works by (1) training LMs to generate pseudo-programs and (2) teaching them to emulate their generated program’s execution, including those leaf functions, allowing the LM’s knowledge to fill in the execution gaps; and (3) using them to search over many programs to find an optimal one. To adapt the COGEX model to a new task, we introduce a method for performing program search to find a single program whose pseudo-execution yields optimal performance when applied to all the instances of a given dataset. We show that our approach yields large improvements compared to standard in-context learning approaches on a battery of tasks, both algorithmic and soft reasoning. This result thus demonstrates that code synthesis can be applied to a much broader class of problems than previously considered.
Nathaniel Weir, Muhammad Khalifa, Linlu Qiu, Orion Weller, Peter Clark
NeurIPS2
2023 Few-shot Reranking for Multi-hop QA via Language Model Prompting
abstract
We study few-shot reranking for multi-hop QA (MQA) with open-domain questions.To alleviate the need for a large number of labeled question-document pairs for retriever training, we propose PROMPTRANK, which relies on language model prompting for multi-hop path reranking.PROMPTRANK first constructs an instruction-based prompt that includes a candidate document path and then computes the relevance score between a given question and the path based on the conditional likelihood of the question given the path prompt according to a language model.PROMPTRANK yields strong retrieval performance on HotpotQA with only 128 training examples compared to state-of-theart methods trained on thousands of examples -73.6 recall@10 by PROMPTRANK vs. 77.8 by PathRetriever (Asai et al., 2020) and 77.5 by multi-hop dense retrieval (Xiong et al., 2021).
Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang 0008
ACL (1)1
2023 Merging Generated and Retrieved Knowledge for Open-Domain QA
abstract
Open-domain question answering (QA) systems are often built with retrieval modules.However, retrieving passages from a given source is known to suffer from insufficient knowledge coverage.Alternatively, prompting large language models (LLMs) to generate contextual passages based on their parametric knowledge has been shown to improve QA performance.Yet, LLMs tend to "hallucinate" content that conflicts with the retrieved knowledge.Based on the intuition that answers supported by both sources are more likely to be correct, we propose COMBO, a Compatibility-Oriented knowledge Merging for Better Open-domain QA framework, to effectively leverage the two sources of information.Concretely, we match LLM-generated passages with retrieved counterparts into compatible pairs, based on discriminators trained with silver compatibility labels.Then a Fusionin-Decoder-based (Izacard and Grave, 2021b) reader model handles passage pairs to arrive at the final answer.Experiments show that COMBO outperforms competitive baselines on three out of four tested open-domain QA benchmarks.Further analysis reveals that our proposed framework demonstrates greater efficacy in scenarios with a higher degree of knowledge conflicts.
Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang 0008
EMNLP2
2021 Self-Training Pre-Trained Language Models for Zero- and Few-Shot Multi-Dialectal Arabic Sequence Labeling
abstract
A sufficient amount of annotated data is usually required to fine-tune pre-trained language models for downstream tasks.Unfortunately, attaining labeled data can be costly, especially for multiple language varieties and dialects.We propose to self-train pre-trained language models in zero-and few-shot scenarios to improve performance on data-scarce varieties using only resources from data-rich ones.We demonstrate the utility of our approach in the context of Arabic sequence labeling by using a language model fine-tuned on Modern Standard Arabic (MSA) only to predict named entities (NE) and part-of-speech (POS) tags on several dialectal Arabic (DA) varieties.We show that self-training is indeed powerful, improving zero-shot MSA-to-DA transfer by as large as 10% F 1 (NER) and 2% accuracy (POS tagging).We acquire even better performance in few-shot scenarios with limited amounts of labeled data.We conduct an ablation study and show that the performance boost observed directly results from training data augmentation possible with DA examples via self-training.This opens up opportunities for developing DA models exploiting only MSA resources.Our approach can also be extended to other languages and tasks. 1
Muhammad Khalifa, Muhammad Abdul-Mageed, Khaled Shaalan
EACL1
2021 A Bag of Tricks for Dialogue Summarization
abstract
Dialogue summarization comes with its own peculiar challenges as opposed to news or scientific articles summarization.In this work, we explore four different challenges of the task: handling and differentiating parts of the dialogue belonging to multiple speakers, negation understanding, reasoning about the situation, and informal language understanding.Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data.Our experiments show that our proposed techniques indeed improve summarization performance, outperforming strong baselines.
Muhammad Khalifa, Miguel Ballesteros, Kathy McKeown
EMNLP (1)1
2021 A Distributional Approach to Controlled Text Generation
Muhammad Khalifa, Hady ElSahar, Marc Dymetman
ICLR1
2021 Extracting Synonyms from Bilingual Dictionaries
abstract
We present our progress in developing a novel algorithm to extract synonyms from bilingual dictionaries.Identification and usage of synonyms play a significant role in improving the performance of information access applications.The idea is to construct a translation graph from translation pairs, then to extract and consolidate cyclic paths to form bilingual sets of synonyms.The initial evaluation of this algorithm illustrates promising results in extracting Arabic-English bilingual synonyms.In the evaluation, we first converted the synsets in the Arabic WordNet into translation pairs (i.e., losing word-sense memberships).Next, we applied our algorithm to rebuild these synsets.We compared the original and extracted synsets obtaining an F-Measure of 82.3% and 82.1% for Arabic and English synsets extraction, respectively.
Mustafa Jarrar, Eman Karajah, Muhammad Khalifa, Khaled Shaalan
GWC3
2019 Character convolutions for Arabic Named Entity Recognition with Long Short-Term Memory Networks
Muhammad Khalifa, Khaled Shaalan
Comput. Speech Lang.1