VLDB 2026 Research / reviewers in the wild / expert
Pittawat Taveekitworachai
dblp:344/1703
· DBLP profile ↗
16ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-6824-2634ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Role of Large Language Model-Generated Stories in the Narrative Experience of Serious Visual Novel GamesabstractThis study examines the impact of Large Language Model-generated narratives in a climate-change-themed Visual Novel, comparing two versions: First, the story is generated using thematic keywords in the prompts, and second, the story is generated without keywords. Fifty participants (21 female, 29 male) completed the study. Results showed that participants in the group without thematic keywords had higher levels of narrative engageability score, as measured by the Narrative Engageability Scale, than those with thematic keywords. This indicated that the ability to engage with the story was stronger in the group without keywords. However, when assessing the narrative experience using the Game User Experience Satisfaction Scale, both groups reported similar levels of satisfaction, suggesting that while the ability to engage with the narrative differed between groups, the overall narrative experience was mainly the same. These findings suggested that thematic keywords in prompts significantly impacted participants’ narrative experience of the game. Mustafa Can Gursesli, Mury F. Dewantoro, Xiao You, Ege Anbar, Pittawat Taveekitworachai, Febri Abdullah, Pietro Tarchi, Mirko Duradoni, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas |
Int. J. Hum. Comput. Interact. | 6 |
| 2025 | Prior Prompt Engineering for Reinforcement Fine-TuningabstractThis paper investigates prior prompt engineering (pPE) in the context of reinforcement finetuning (RFT), where language models (LMs) are incentivized to exhibit behaviors that maximize performance through reward signals.While existing RFT research has primarily focused on algorithms, reward shaping, and data curation, the design of the prior prompt-the instructions prepended to queries during training to elicit behaviors such as step-by-step reasoning-remains underexplored.We investigate whether different pPE approaches can guide LMs to internalize distinct behaviors after RFT.Inspired by inference-time prompt engineering (iPE), we translate five representative iPE strategies-reasoning, planning, codebased reasoning, knowledge recall, and nullexample utilization-into corresponding pPE approaches.We experiment with Qwen2.5-7B using each of the pPE approaches, then evaluate performance on in-domain and out-of-domain benchmz arks (e.g., AIME2024, HumanEval+, and GPQA-Diamond).Our results show that all pPE-trained models surpass their iPE-prompted counterparts, with the null-example pPE approach achieving the largest average performance gain and the highest improvement on AIME2024 and GPQA-Diamond, surpassing the commonly used reasoning approach.Furthermore, by adapting a behavior-classification framework, we demonstrate that different pPE strategies instill distinct behavioral styles in the resulting models.These findings position pPE as a powerful yet understudied axis for RFT. Pittawat Taveekitworachai, P. P. Manakul, Sarana Nutanong, Kunat Pipatanakul |
EMNLP | 1 |
| 2025 | BenchING: A Benchmark for Evaluating Large Language Models in Following Structured Output Format Instruction in Text-Based Narrative Game TasksabstractIn this article, we present BenchING, a new benchmark for evaluating large language models (LLMs) on their ability to follow structured output format instructions in text-based procedural content generation (PCG) tasks. The ability to condition LLMs to output in specified formats proves useful, as downstream components in LLM-integrated games often require structured outputs for exchanging information. However, there is a gap in evaluating this aspect of LLMs, especially in narrative PCG tasks, making it difficult to select LLMs and design games or applications integrating these LLMs. To demonstrate the potential of our benchmark, we evaluate nine LLMs for their ability to generate parseable formatted outputs using five selected text-based PCG tasks. We report on the performance of these LLMs on these tasks. In addition, we categorize more detailed error types and propose solutions by utilizing LLMs to fix these errors. We also conduct a scaling study, investigating an emergent point of LLMs for their ability to fix malformed formatted content using eight quantized LLMs with varying original sizes from 0.62 to 72.3 B. Furthermore, we perform a qualitative study to assess the quality of the generated content. We make our source code and raw data available for future research. Pittawat Taveekitworachai, Mury F. Dewantoro, Pratch Suntichaikul, Ruck Thawonmas |
IEEE Trans. Games | 1 |
| 2024 | ChatGPT4PCG 2 Competition: Prompt Engineering for Science Birds Level GenerationabstractThis paper presents the second ChatGPT4PCG competition at the 2024 IEEE Conference on Games. In this edition of the competition, we follow the first edition, but make several improvements and changes. We introduce a new evaluation metric along with allowing a more flexible format for participants’ submissions and making several improvements to the evaluation pipeline. Continuing from the first edition, we aim to foster and explore the realm of prompt engineering (PE) for procedural content generation (PCG). While the first competition saw success, it was hindered by various limitations; we aim to mitigate these limitations in this edition. We introduce diversity as a new metric to discourage submissions aimed at producing repetitive structures. Furthermore, we allow submission of a Python program instead of a prompt text file for greater flexibility in implementing advanced PE approaches, which may require control flow, including conditions and iterations. We also make several improvements to the evaluation pipeline with a better classifier for similarity evaluation and better-performing function signatures. We thoroughly evaluate the effectiveness of the new metric and the improved classifier. Additionally, we perform an ablation study to select a function signature to instruct ChatGPT for level generation. Finally, we provide implementation examples of various PE techniques in Python and evaluate their preliminary performance. We hope this competition serves as a resource and platform for learning about PE and PCG in general1.1Source code and raw data: https://github.com/chatgpt4pcg/experiments2024 Pittawat Taveekitworachai, Febri Abdullah, Mury F. Dewantoro, Pratch Suntichaikul, Ruck Thawonmas, Julian Togelius, Jochen Renz |
CoG | 1 |
| 2024 | Assessing Inherent Biases Following Prompt Compression of Large Language Models for Game Story GenerationabstractThis paper investigates how prompt compression, a technique to reduce the number of tokens in the prompt while maintaining prompt performance, affects inherent biases in large language models (LLMs) for the story ending of the game story generation task. Previous studies have explored inherent biases in LLMs and found an innate inclination of LLMs towards generating positive-ending stories. While prompt compression is known to retain task performance and utilize fewer tokens in the prompt, we explore a different perspective on how prompt compression could affect inherent biases in LLMs. We follow existing studies’ approach in evaluating story ending biases of six LLMs comparing uncompressed and compressed prompts. We find that prompt compression does not affect story generation from positive-ending story synopses, to which these LLMs are inclined. However, it is not the same for negative-ending story synopses: prompt compression either makes the LLMs generate a higher amount of negative-ending stories or not at all. We also notice that the classification of other types of story endings, other than those specified in the prompt, aligns with an existing study. We recommend game developers and future studies to always perform empirical tests on prompt compression, as it is not straightforward and may greatly alter model behaviors. Pittawat Taveekitworachai, Kantinan Plupattanakit, Ruck Thawonmas |
CoG | 1 |
| 2024 | Towards LLM4PCG: A Preliminary Evaluation of Open-Weight Large Language Models Beyond ChatGPT4PCGabstractThis paper presents an initial step towards general evaluations of open-weight large language models (LLMs) using the ChatGPT4PCG platform, a Science Birds level generation challenge designed to evaluate LLMs on the complex task of generating stable, English-character-resembling, and diverse levels. While ChatGPT4PCG competitions have their own merit in providing a comprehensive platform for evaluating ChatGPT on complex tasks, the competitions focus solely on ChatGPT is rather limiting considering the fact that there are many available choices of open-weight LLMs. We report 13 LLMs from five model families of various properties in their design choices and sizes to evaluate on a modified ChatGPT4PCG 2 competition platform. We observe that the scaling law holds in general, but the inherent capabilities of LLMs due to their pre-training and architecture choices also play an equal role. We open-source the modification of the ChatGPT4PCG platform to support future research on evaluating LLMs in this area1.1https://github.com/Pittawat2542/llm4pcg-python and https://github.com/Pittawat2542/llm4pcg-experiment Pittawat Taveekitworachai, Pratch Suntichaikul, Ruck Thawonmas |
CoG | 1 |
| 2024 | Null-Shot Prompting: Rethinking Prompting Large Language Models With HallucinationabstractThis paper investigates an interesting phenomenon where we observe performance increases in large language models (LLMs) when providing a prompt that causes and exploits hallucination.We propose null-shot prompting, a counter-intuitive approach where we deliberately instruct LLMs to reference a null, nonexistent, section.We evaluate null-shot prompting across a variety of tasks, including arithmetic reasoning, commonsense reasoning, and reading comprehension.Notably, we observe a substantial increase in performance in arithmetic reasoning tasks for various models, with up to a 44.62% increase compared to a baseline in one model.Additional experiments on more complex mathematical problem-solving and hallucination detection benchmarks also reveal similar benefits from this approach.Furthermore, we explore the effects of combining reasoning, which typically mitigates hallucination, with hallucination within the prompt and find several cases of performance improvements.We hope this paper stimulates further interest, investigation, and discussion on how hallucination in prompts may not only affect LLMs but, in certain cases, enhance their performance. Pittawat Taveekitworachai, Febri Abdullah, Ruck Thawonmas |
EMNLP | 1 |
| 2024 | Dungeons, Dragons, and Emotions: A Preliminary Study of Player Sentiment in LLM-driven TTRPGsabstractIn this paper, we present a Tabletop Role-Playing game (TTRPG) driven by ChatGPT. Prompts are employed to instruct ChatGPT to act as Game Masters (GMs). In crafting each prompt to integrate a distinctive role, three roles denoted as Role 1, Role 2, and Role 3, are established. Subsequently, we perform pre-game and post-game emotional assessments employing the Positive and Negative Affect Schedule (PANAS) questionnaire to scrutinize players’ emotional dynamics throughout the gaming experience. Upon analyzing the collected data, we observe that Role 1 and Role 2 affect players’ positive emotions. Notably, Role 2 exhibits the most pronounced influence on players’ positive emotions. Our findings demonstrate that a TTRPG GM powered by ChatGPT can significantly enhance players’ positive emotions. This leads us to recognize that TTRPG GM powered by ChatGPT plays a positive role in enhancing the mental well-being of specific populations. Xiao You, Pittawat Taveekitworachai, Mustafa Can Gursesli, Ruck Thawonmas |
FDG | 2 |
| 2024 | Don't Do That! Reverse Role Prompting Helps Large Language Models Stay in Personality Traits
Pittawat Taveekitworachai, Mustafa Can Gursesli, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas |
ICIDS (1) | 2 |
| 2024 | Speed Up! Cost-Effective Large Language Model for ADAS Via Knowledge DistillationabstractThis paper presents a cost-effective approach to utilizing large language models (LLMs) as part of advanced driver-assistance systems (ADAS) through a knowledge-distilled model for driving assessment. LLMs have recently been employed across various domains. However, due to their size, they require sufficient computing infrastructure for deployment and ample time for generation. These characteristics make LLMs challenging to integrate into applications requiring real-time feedback, including ADAS. An existing study employed a vector database containing responses generated from an LLM to act as a surrogate model. However, this approach is limited when handling out-of-distribution (OOD) scenarios, which LLMs excel at. We propose a novel approach that utilizes a distilled model obtained from an established knowledge distillation technique to perform as a surrogate model for a target LLM, offering high resilience in handling OOD situations with substantially faster inference time. To assess the performance of the proposed approach, we also introduce a new dataset for driving scenarios and situations (DriveSSD), containing 124,248 records. Additionally, we augment randomly selected 12,425 records, 10% of our DriveSSD, with text embeddings generated from an embedding model. We distill the model using 10,000 augmented records and test all approaches on the remaining 2,425 records. We find that the distilled model introduced in this study has better performance across metrics, with half of the inference time used by the previous approach. We make our source code and data publicly available1. Pittawat Taveekitworachai, Pratch Suntichaikul, Chakarida Nukoolkit, Ruck Thawonmas |
IV | 1 |
| 2024 | The First ChatGPT4PCG CompetitionabstractThis study summarizes the first ChatGPT4PCG competition held at the 2023 IEEE Conference on Games. The goal of the competition is to explore emergent abilities of publicly available LLMs in performing complex tasks related to procedural content generation, specifically physics-based level generation for Angry Bird-like games. Participants are tasked with submitting their prompts for ChatGPT to generate Angry Birds-like game structures that resemble English uppercase characters. A structure is a collection of stacked game objects comprising a part of an entire Angry Birds-like level. A prompt is an input for large language models (LLMs) including ChatGPT. Two evaluation metrics, i.e., stability and similarity, are used to evaluate the submitted prompts. Stability measures the sturdiness of a structure to withstand in-game gravity, while similarity measures a structure's resemblance to the target character. With such evaluation, participants are challenged not only to produce character-like but also stable structures by utilizing prompt engineering techniques. Finally, the competition's results are discussed to provide valuable insights for future studies and competitions. Febri Abdullah, Pittawat Taveekitworachai, Mury F. Dewantoro, Ruck Thawonmas, Julian Togelius, Jochen Renz |
IEEE Trans. Games | 2 |
| 2023 | ChatGPT4PCG Competition: Character-like Level Generation for Science BirdsabstractThis paper presents the first ChatGPT4PCG Competition at the 2023 IEEE Conference on Games. The objective of this competition is for participants to create effective prompts for ChatGPT–enabling it to generate Science Birds levels with high stability and character-like qualities–fully using their creativity as well as prompt engineering skills. ChatGPT is a conversational agent developed by OpenAI. Science Birds is selected as the competition platform because designing an Angry Birds-like level is not a trivial task due to the in-game gravity; the quality of the levels is determined by their stability. To lower the entry barrier to the competition, we limit the task to the generation of capitalized English alphabetical characters. We also allow only a single prompt to be used for generating all the characters. Here, the quality of the generated levels is determined by their stability and similarity to the given characters. A sample prompt is provided to participants for their reference. An experiment is conducted to determine the effectiveness of several modified versions of this sample prompt on level stability and similarity by testing them on several characters. To the best of our knowledge, we believe that ChatGPT4PCG is the first competition of its kind and hope to inspire enthusiasm for prompt engineering in procedural content generation. Pittawat Taveekitworachai, Febri Abdullah, Mury F. Dewantoro, Ruck Thawonmas, Julian Togelius, Jochen Renz |
CoG | 1 |
| 2023 | The Chronicles of ChatGPT: Generating and Evaluating Visual Novel Narratives on Climate Change Through ChatGPT
Mustafa Can Gursesli, Pittawat Taveekitworachai, Febri Abdullah, Mury F. Dewantoro, Antonio Lanatà, Andrea Guazzini, Van Khôi Lê, Adrien Villars, Ruck Thawonmas |
ICIDS (2) | 2 |
| 2023 | Analyzing Audience Comments: Improving Interactive Narrative with ChatGPT
Xiao You, Pittawat Taveekitworachai, Ruck Thawonmas |
ICIDS (2) | 4 |
| 2023 | What Is Waiting for Us at the End? Inherent Biases of Game Story Endings in Large Language Models
Pittawat Taveekitworachai, Febri Abdullah, Mustafa Can Gursesli, Mury F. Dewantoro, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas |
ICIDS (2) | 1 |
| 2023 | Breaking Bad: Unraveling Influences and Risks of User Inputs to ChatGPT for Game Story Generation
Pittawat Taveekitworachai, Febri Abdullah, Mustafa Can Gursesli, Mury F. Dewantoro, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas |
ICIDS (2) | 1 |