VLDB 2026 Research / reviewers in the wild / expert
P. P. Manakul
dblp:243/6654 · also Potsawee Manakul
· DBLP profile ↗
15ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0001-7108-8626ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mind the Gap: Static and Interactive Evaluations of Large Audio ModelsabstractMinzhi Li, William Held, Michael J. Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, Diyi Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Minzhi Li, William Barr Held, Michael J. Ryan, Kunat Pipatanakul, P. P. Manakul, Diyi Yang |
ACL (1) | 5 |
| 2025 | SkillAggregation: Reference-free LLM-Dependent AggregationabstractLarge Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments.Single LLMs were used initially, however, recent work suggests using multiple LLMs as judges yields improved performance.An important step in exploiting multiple judgements is the combination stage, aggregation.Existing methods in NLP either assign equal weight to all LLM judgments or are designed for specific tasks such as hallucination detection.This work focuses on aggregating predictions from multiple systems where no reference labels are available.A new method called SkillAggregation is proposed, which learns to combine estimates from LLM judges without needing additional data or ground truth.It extends the Crowdlayer aggregation method, developed for image classification, to exploit the judge estimates during inference.The approach is compared to a range of standard aggregation methods on HaluEval-Dialogue, TruthfulQA and Chatbot Arena tasks.SkillAggregation outperforms Crowdlayer on all tasks, and yields the best performance over all approaches on the majority of tasks. 1 Guangzhi Sun, Anmol Kagrecha, P. P. Manakul, Philip C. Woodland, Mark J. F. Gales |
ACL (1) | 3 |
| 2025 | Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?abstractUnlearning has emerged as a critical capability for large language models (LLMs) to support data privacy, regulatory compliance, and ethical AI deployment.Recent techniques often rely on obfuscation by injecting incorrect or irrelevant information to suppress knowledge.Such methods effectively constitute knowledge addition rather than true removal, often leaving models vulnerable to probing.In this paper, we formally distinguish unlearning from obfuscation and introduce a probing-based evaluation framework to assess whether existing approaches genuinely remove targeted information.Moreover, we propose DF-MCQ, a novel unlearning method that flattens the model predictive distribution over automatically generated multiple-choice questions using KLdivergence, effectively removing knowledge about target individuals and triggering appropriate refusal behaviour.Experimental results demonstrate that DF-MCQ achieves unlearning with over 90% refusal rate and a random choice-level uncertainty that is much higher than obfuscation on probing questions. 1 Guangzhi Sun, P. P. Manakul, Xiao Zhan, Mark J. F. Gales |
EMNLP | 2 |
| 2025 | Prior Prompt Engineering for Reinforcement Fine-TuningabstractThis paper investigates prior prompt engineering (pPE) in the context of reinforcement finetuning (RFT), where language models (LMs) are incentivized to exhibit behaviors that maximize performance through reward signals.While existing RFT research has primarily focused on algorithms, reward shaping, and data curation, the design of the prior prompt-the instructions prepended to queries during training to elicit behaviors such as step-by-step reasoning-remains underexplored.We investigate whether different pPE approaches can guide LMs to internalize distinct behaviors after RFT.Inspired by inference-time prompt engineering (iPE), we translate five representative iPE strategies-reasoning, planning, codebased reasoning, knowledge recall, and nullexample utilization-into corresponding pPE approaches.We experiment with Qwen2.5-7B using each of the pPE approaches, then evaluate performance on in-domain and out-of-domain benchmz arks (e.g., AIME2024, HumanEval+, and GPQA-Diamond).Our results show that all pPE-trained models surpass their iPE-prompted counterparts, with the null-example pPE approach achieving the largest average performance gain and the highest improvement on AIME2024 and GPQA-Diamond, surpassing the commonly used reasoning approach.Furthermore, by adapting a behavior-classification framework, we demonstrate that different pPE strategies instill distinct behavioral styles in the resulting models.These findings position pPE as a powerful yet understudied axis for RFT. Pittawat Taveekitworachai, P. P. Manakul, Sarana Nutanong, Kunat Pipatanakul |
EMNLP | 2 |
| 2025 | Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
P. P. Manakul, Guangzhi Sun, Warit Sirichotedumrong, Kasima Tharnpipitchai, Kunat Pipatanakul |
INTERSPEECH | 1 |
| 2024 | LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language ModelsabstractCurrent developments in large language models (LLMs) have enabled impressive zero-shot capabilities across various natural language tasks.An interesting application of these systems is in the automated assessment of natural language generation (NLG), a highly challenging area with great practical benefit.In this paper, we explore two options for exploiting the emergent abilities of LLMs for zero-shot NLG assessment: absolute score prediction, and comparative assessment which uses relative comparisons between pairs of candidates.Though comparative assessment has not been extensively studied in NLG assessment, we note that humans often find it more intuitive to compare two options rather than scoring each one independently.This work examines comparative assessment from multiple perspectives: performance compared to absolute grading; positional biases in the prompt; and efficient ranking in terms of the number of comparisons.We illustrate that LLM comparative assessment is a simple, general and effective approach for NLG assessment.For moderatesized open-source LLMs, such as FlanT5 and Llama2-chat, comparative assessment is superior to prompt scoring, and in many cases can achieve performance competitive with state-ofthe-art methods.Additionally, we demonstrate that LLMs often exhibit strong positional biases when making pairwise comparisons, and we propose debiasing methods that can further improve performance. Adian Liusie, P. P. Manakul, Mark J. F. Gales |
EACL (1) | 2 |
| 2024 | An Empirical Study of Multilingual Reasoning Distillation for Question AnsweringabstractPatomporn Payoungkhamdee, Peerat Limkonchotiwat, Jinheon Baek, Potsawee Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Patomporn Payoungkhamdee, Peerat Limkonchotiwat, Jinheon Baek, P. P. Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong |
EMNLP | 4 |
| 2024 | Efficient Overshadowed Entity Disambiguation by Mitigating Shortcut LearningabstractPanuthep Tasawong, Peerat Limkonchotiwat, Potsawee Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Panuthep Tasawong, Peerat Limkonchotiwat, P. P. Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong |
EMNLP | 3 |
| 2023 | SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsabstractGenerative Large Language Models (LLMs) such as GPT-3 are capable of generating highly fluent responses to a wide variety of user prompts.However, LLMs are known to hallucinate facts and make non-factual statements which can undermine trust in their output.Existing fact-checking approaches either require access to the output probability distribution (which may not be available for systems such as ChatGPT) or external databases that are interfaced via separate, often complex, modules.In this work, we propose "SelfCheckGPT", a simple sampling-based approach that can be used to fact-check the responses of black-box models in a zero-resource fashion, i.e. without an external database.SelfCheckGPT leverages the simple idea that if an LLM has knowledge of a given concept, sampled responses are likely to be similar and contain consistent facts.However, for hallucinated facts, stochastically sampled responses are likely to diverge and contradict one another.We investigate this approach by using GPT-3 to generate passages about individuals from the WikiBio dataset, and manually annotate the factuality of the generated passages.We demonstrate that SelfCheck-GPT can: i) detect non-factual and factual sentences; and ii) rank passages in terms of factuality.We compare our approach to several baselines and show that our approach has considerably higher AUC-PR scores in sentence-level hallucination detection and higher correlation scores in passage-level factuality assessment compared to grey-box methods. P. P. Manakul, Adian Liusie, Mark J. F. Gales |
EMNLP | 1 |
| 2023 | MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in SummarizationabstractPotsawee Manakul, Adian Liusie, Mark Gales. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. P. P. Manakul, Adian Liusie, Mark J. F. Gales |
IJCNLP (1) | 1 |
| 2021 | Long-Span Summarization via Local Attention and Content SelectionabstractPotsawee Manakul, Mark Gales. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. P. P. Manakul, Mark J. F. Gales |
ACL/IJCNLP (1) | 1 |
| 2021 | Sparsity and Sentence Structure in Encoder-Decoder Attention of Summarization SystemsabstractTransformer models have achieved state-ofthe-art results in a wide range of NLP tasks including summarization.Training and inference using large transformer models can be computationally expensive.Previous work has focused on one important bottleneck, the quadratic self-attention mechanism in the encoder.Modified encoder architectures such as LED or LoBART use local attention patterns to address this problem for summarization.In contrast, this work focuses on the transformer's encoder-decoder attention mechanism.The cost of this attention becomes more significant in inference or training approaches that require model-generated histories.First, we examine the complexity of the encoder-decoder attention.We demonstrate empirically that there is a sparse sentence structure in document summarization that can be exploited by constraining the attention mechanism to a subset of input sentences, whilst maintaining system performance.Second, we propose a modified architecture that selects the subset of sentences to constrain the encoder-decoder attention.Experiments are carried out on abstractive summarization tasks, including CNN/DailyMail, XSum, Spotify Podcast, and arXiv. 1 P. P. Manakul, Mark J. F. Gales |
EMNLP (1) | 1 |
| 2020 | Abstractive Spoken Document Summarization Using Hierarchical Model with Multi-Stage Attention Diversity OptimizationabstractAbstractive summarization is a standard task for written documents, such as news articles. Applying summarization schemes to spoken documents is more challenging, especially in situations involving human interactions, such as meetings. Here, utterances tend not to form complete sentences and sometimes contain little information. Moreover, speech disfluencies will be present as well as recognition errors for automated systems. For current attention-based sequence-to-sequence summarization systems, these additional challenges can yield a poor attention distribution over the spoken document words and utterances, impacting performance. In this work, we propose a multi-stage method based on a hierarchical encoder-decoder model to explicitly model utterance-level attention distribution at training time; and enforce diversity at inference time using a unigram diversity term. Furthermore, multitask learning tasks including dialogue act classification and extractive summarization are incorporated. The performance of the system is evaluated on the AMI meeting corpus. The inclusion of both training and inference diversity terms improves performance, outperforming current state-of-the-art systems in terms of ROUGE scores. Additionally, the impact of ASR errors, as well as performance on the multitask learning tasks, is evaluated. P. P. Manakul, Mark J. F. Gales |
INTERSPEECH | 1 |
| 2019 | Automatic Grammatical Error Detection of Non-native Spoken Learner EnglishabstractAutomatic language assessment and learning systems are required to support the global growth in English language learning. They need to be able to provide reliable and meaningful feedback to help learners develop their skills. This paper considers the question of detecting "grammatical" errors in non-native spoken English as a first step to providing feedback on a learner's use of the language. A state-of-the-art deep learning based grammatical error detection (GED) system designed for written texts is investigated on free speaking tasks across the full range of proficiency grades with a mix of first languages (L1s). This presents a number of challenges. Free speech contains disfluencies that disrupt the spoken language flow but are not grammatical errors. The lower the level of the learner the more these both will occur which makes the underlying task of automatic transcription harder. The baseline written GED system is seen to perform less well on manually transcribed spoken language. When the GED model is fine-tuned to free speech data from the target domain the spoken system is able to match the written performance. Given the current state-of-the-art in ASR, however, and the ability to detect disfluencies grammatical error feedback from automated transcriptions remains a challenge. Kate M. Knill, Mark J. F. Gales, P. P. Manakul, Andrew Caines |
ICASSP | 3 |
| 2019 | Impact of ASR Performance on Spoken Grammatical Error DetectionabstractComputer assisted language learning (CALL) systems aidlearners to monitor their progress by providing scoring andfeedback on language assessment tasks. Free speaking tests al-low assessment of what a learner has said, as well as how theysaid it. For these tasks, Automatic Speech Recognition (ASR)is required to generate transcriptions of a candidate’s responses,the quality of these transcriptions is crucial to provide reliablefeedback in downstream processes. This paper considers theimpact of ASR performance on Grammatical Error Detection(GED) for free speaking tasks, as an example of providing feed-back on a learner’s use of English. The performance of an ad-vanced deep-learning based GED system, initially trained onwritten corpora, is used to evaluate the influence of ASR errors.One consequence of these errors is that grammatical errors canresult from incorrect transcriptions as well as learner errors, thismay yield confusing feedback. To mitigate the effect of theseerrors, and reduce erroneous feedback, ASR confidence scoresare incorporated into the GED system. By additionally adaptingthe written text GED system to the speech domain, using ASRtranscriptions, significant gains in performance can be achieved.Analysis of the GED performance for different grammatical er-ror types and across grade is also presented. Yiting Lu, Mark J. F. Gales, Kate M. Knill, P. P. Manakul, Yu Wang 0027 |
INTERSPEECH | 4 |