Mury F. Dewantoro

dblp:282/7189 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0002-1967-8607ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 The Role of Large Language Model-Generated Stories in the Narrative Experience of Serious Visual Novel Games
abstract
This study examines the impact of Large Language Model-generated narratives in a climate-change-themed Visual Novel, comparing two versions: First, the story is generated using thematic keywords in the prompts, and second, the story is generated without keywords. Fifty participants (21 female, 29 male) completed the study. Results showed that participants in the group without thematic keywords had higher levels of narrative engageability score, as measured by the Narrative Engageability Scale, than those with thematic keywords. This indicated that the ability to engage with the story was stronger in the group without keywords. However, when assessing the narrative experience using the Game User Experience Satisfaction Scale, both groups reported similar levels of satisfaction, suggesting that while the ability to engage with the narrative differed between groups, the overall narrative experience was mainly the same. These findings suggested that thematic keywords in prompts significantly impacted participants’ narrative experience of the game.
Mustafa Can Gursesli, Mury F. Dewantoro, Xiao You, Ege Anbar, Pittawat Taveekitworachai, Febri Abdullah, Pietro Tarchi, Mirko Duradoni, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas
Int. J. Hum. Comput. Interact.3
2026 Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage
Ibrahim Khan, Mury F. Dewantoro, Wenwen Ouyang, Ruck Thawonmas
IEEE Trans. Games3
2025 Can Multimodal LLMs Reason About Stability? An Exploratory Study with Insights from the LLMs4PCG Challenge
abstract
This study investigates the extent to which multimodal large language models (MLLMs) demonstrate physical reasoning capabilities in dynamic, visually grounded environments. We evaluate whether using images as context in a prompt can enhance the ability of MLLMs to simulate and predict the outcomes of physical interactions. Using Science Birds, an Angry Birds-like physicsbased platform, we design a suite of tasks that probe core competencies in binary, comparative stability, and forward simulation using visual and textual inputs. Our evaluation shows that MLLMs can perform visual reasoning tasks with quantifiable accuracy, especially when making predictions based on image-rich input. However, their performance varies considerably depending on the specific model architecture and the type of input modality used. These findings highlight that incorporating both visual and textual data is crucial for accurate physical inference. Nonetheless, the evaluated MLLMs still have substantial limitations in structured visual reasoning tasks. By systematically analyzing the strengths and weaknesses of different models, our work provides practical guidance for advancing MLLM-based physical reasoning and supports the development of future benchmarks and competitions in this area, including contributions to the LLMs4PCG challenge. We make our source code and raw data available for future research.11https://anonymous.4open.science/r/cog2025-game-physics-eval/.
Mury F. Dewantoro, Febri Abdullah, Ibrahim Khan, Ruck Thawonmas, Wenwen Ouyang
CoG1
2025 BenchING: A Benchmark for Evaluating Large Language Models in Following Structured Output Format Instruction in Text-Based Narrative Game Tasks
abstract
In this article, we present BenchING, a new benchmark for evaluating large language models (LLMs) on their ability to follow structured output format instructions in text-based procedural content generation (PCG) tasks. The ability to condition LLMs to output in specified formats proves useful, as downstream components in LLM-integrated games often require structured outputs for exchanging information. However, there is a gap in evaluating this aspect of LLMs, especially in narrative PCG tasks, making it difficult to select LLMs and design games or applications integrating these LLMs. To demonstrate the potential of our benchmark, we evaluate nine LLMs for their ability to generate parseable formatted outputs using five selected text-based PCG tasks. We report on the performance of these LLMs on these tasks. In addition, we categorize more detailed error types and propose solutions by utilizing LLMs to fix these errors. We also conduct a scaling study, investigating an emergent point of LLMs for their ability to fix malformed formatted content using eight quantized LLMs with varying original sizes from 0.62 to 72.3 B. Furthermore, we perform a qualitative study to assess the quality of the generated content. We make our source code and raw data available for future research.
Pittawat Taveekitworachai, Mury F. Dewantoro, Pratch Suntichaikul, Ruck Thawonmas
IEEE Trans. Games2
2024 ChatGPT4PCG 2 Competition: Prompt Engineering for Science Birds Level Generation
abstract
This paper presents the second ChatGPT4PCG competition at the 2024 IEEE Conference on Games. In this edition of the competition, we follow the first edition, but make several improvements and changes. We introduce a new evaluation metric along with allowing a more flexible format for participants’ submissions and making several improvements to the evaluation pipeline. Continuing from the first edition, we aim to foster and explore the realm of prompt engineering (PE) for procedural content generation (PCG). While the first competition saw success, it was hindered by various limitations; we aim to mitigate these limitations in this edition. We introduce diversity as a new metric to discourage submissions aimed at producing repetitive structures. Furthermore, we allow submission of a Python program instead of a prompt text file for greater flexibility in implementing advanced PE approaches, which may require control flow, including conditions and iterations. We also make several improvements to the evaluation pipeline with a better classifier for similarity evaluation and better-performing function signatures. We thoroughly evaluate the effectiveness of the new metric and the improved classifier. Additionally, we perform an ablation study to select a function signature to instruct ChatGPT for level generation. Finally, we provide implementation examples of various PE techniques in Python and evaluate their preliminary performance. We hope this competition serves as a resource and platform for learning about PE and PCG in general1.1Source code and raw data: https://github.com/chatgpt4pcg/experiments2024
Pittawat Taveekitworachai, Febri Abdullah, Mury F. Dewantoro, Pratch Suntichaikul, Ruck Thawonmas, Julian Togelius, Jochen Renz
CoG3
2024 The First ChatGPT4PCG Competition
abstract
This study summarizes the first ChatGPT4PCG competition held at the 2023 IEEE Conference on Games. The goal of the competition is to explore emergent abilities of publicly available LLMs in performing complex tasks related to procedural content generation, specifically physics-based level generation for Angry Bird-like games. Participants are tasked with submitting their prompts for ChatGPT to generate Angry Birds-like game structures that resemble English uppercase characters. A structure is a collection of stacked game objects comprising a part of an entire Angry Birds-like level. A prompt is an input for large language models (LLMs) including ChatGPT. Two evaluation metrics, i.e., stability and similarity, are used to evaluate the submitted prompts. Stability measures the sturdiness of a structure to withstand in-game gravity, while similarity measures a structure's resemblance to the target character. With such evaluation, participants are challenged not only to produce character-like but also stable structures by utilizing prompt engineering techniques. Finally, the competition's results are discussed to provide valuable insights for future studies and competitions.
Febri Abdullah, Pittawat Taveekitworachai, Mury F. Dewantoro, Ruck Thawonmas, Julian Togelius, Jochen Renz
IEEE Trans. Games3
2023 ChatGPT4PCG Competition: Character-like Level Generation for Science Birds
abstract
This paper presents the first ChatGPT4PCG Competition at the 2023 IEEE Conference on Games. The objective of this competition is for participants to create effective prompts for ChatGPT–enabling it to generate Science Birds levels with high stability and character-like qualities–fully using their creativity as well as prompt engineering skills. ChatGPT is a conversational agent developed by OpenAI. Science Birds is selected as the competition platform because designing an Angry Birds-like level is not a trivial task due to the in-game gravity; the quality of the levels is determined by their stability. To lower the entry barrier to the competition, we limit the task to the generation of capitalized English alphabetical characters. We also allow only a single prompt to be used for generating all the characters. Here, the quality of the generated levels is determined by their stability and similarity to the given characters. A sample prompt is provided to participants for their reference. An experiment is conducted to determine the effectiveness of several modified versions of this sample prompt on level stability and similarity by testing them on several characters. To the best of our knowledge, we believe that ChatGPT4PCG is the first competition of its kind and hope to inspire enthusiasm for prompt engineering in procedural content generation.
Pittawat Taveekitworachai, Febri Abdullah, Mury F. Dewantoro, Ruck Thawonmas, Julian Togelius, Jochen Renz
CoG3
2023 The Chronicles of ChatGPT: Generating and Evaluating Visual Novel Narratives on Climate Change Through ChatGPT
Mustafa Can Gursesli, Pittawat Taveekitworachai, Febri Abdullah, Mury F. Dewantoro, Antonio Lanatà, Andrea Guazzini, Van Khôi Lê, Adrien Villars, Ruck Thawonmas
ICIDS (2)4
2023 What Is Waiting for Us at the End? Inherent Biases of Game Story Endings in Large Language Models
Pittawat Taveekitworachai, Febri Abdullah, Mustafa Can Gursesli, Mury F. Dewantoro, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas
ICIDS (2)4
2023 Breaking Bad: Unraveling Influences and Risks of User Inputs to ChatGPT for Game Story Generation
Pittawat Taveekitworachai, Febri Abdullah, Mustafa Can Gursesli, Mury F. Dewantoro, Antonio Lanatà, Andrea Guazzini, Ruck Thawonmas
ICIDS (2)4
2022 Science Birds Gameplay With a Smile Interface to Promote the Spectator's Emotion
abstract
This demo paper presents a smile interface to promote the spectator’s emotion. Here, a smile interface is defined as a human-computer interaction system that detects the user’s smile. A past study suggested that watching Science Birds gameplay featuring Rube Goldberg Machine mechanisms with a domino effect can promote the spectator’s emotion. Science Birds is a clone version of Angry Birds for research purposes. Furthermore, other studies suggested that playing Science Birds with a smile interface has the benefit of promoting the player’s emotion. However, such benefit from the spectator’s perspective is yet to be investigated. The presented interface will be used in such investigation in the future.
Febri Abdullah, Mury F. Dewantoro, Ruck Thawonmas, Fitra Abdurrachman Bachtiar
CoG2