VLDB 2026 Research / reviewers in the wild / expert
Agnieszka Mensfelt
dblp:169/4197
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-2385-2017ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards a Common Framework for AutoformalizationabstractAutoformalization has emerged as a term referring to the automation of formalization in the context of the formalization of mathematics using interactive theorem provers (proof assistants). Its rapid development has been driven by progress in deep learning, especially large language models (LLMs). More recently, usage of the term has expanded beyond mathematics to describe tasks that involve translating natural language input into verifiable logical representations. At the same time, a growing body of research explores using LLMs to translate informal language into formal representations for reasoning, planning, and knowledge representation, but without explicitly referring to this process as autoformalization. As a result, despite addressing similar tasks, the largely independent development of these research areas has limited opportunities for shared methodologies, benchmarks, and theoretical frameworks that could accelerate progress. Our goal is to review - explicit or implicit - instances of what can be considered autoformalization and to propose a unified framework, encouraging cross-pollination between different fields to advance the development of next generation AI systems. Agnieszka Mensfelt, David Tena Cucala, Santiago Franco, Angeliki Koutsoukou-Argyraki, Vince Trencsenyi, Kostas Stathis |
AAAI | 1 |
| 2025 | Generative Agents for Multi-Agent Autoformalization of Interaction ScenariosabstractMulti-agent simulations are a versatile tool for exploring interactions among natural and artificial agents, but their development typically demands domain expertise and manual effort. This work introduces the Generative Agents for Multi-Agent Autoformalization (GAMA) framework, which automates the formalization of interaction scenarios in simulations using agents augmented with large language models (LLMs). To demonstrate the application of GAMA, we use natural language descriptions of game-theoretic scenarios representing social interactions, and we autoformalize them into executable logic programs defining game rules, with syntactic correctness enforced through a solver-based validation. To ensure runtime validity, an iterative, tournament-based procedure tests the generated rules and strategies, followed by exact semantic validation when ground truth outcomes are available. In experiments with 110 natural language descriptions across five 2 × 2 simultaneous-move games, GAMA achieves 100% syntactic and 76.5% semantic correctness with Claude 3.5 Sonnet, and 99.82% syntactic and 77% semantic correctness with GPT-4o. The framework also shows high semantic accuracy in autoformalizing agents’ strategies. Agnieszka Mensfelt, Kostas Stathis, Vince Trencsenyi |
ECAI | 1 |
| 2025 | The Influence of Human-Inspired Agentic Sophistication in LLM-Driven Strategic ReasonersabstractThe rapid rise of large language models (LLMs) has shifted artificial intelligence (AI) research toward agentic systems, motivating the use of weaker and more flexible notions of agency. However, this shift raises key questions about the extent to which LLM-based agents replicate human strategic reasoning, particularly in game-theoretic settings. In this context, we examine the role of agentic sophistication in shaping artificial reasoners’ performance by evaluating three agent designs: a simple game-theoretic model, an unstructured LLM-as-agent model, and an LLM integrated into a traditional agentic framework. Using guessing games as a testbed, we benchmarked these agents against human participants across general reasoning patterns and individual role-based objectives. Furthermore, we introduced obfuscated game scenarios to assess agents’ ability to generalise beyond training distributions. Our analysis, covering over 2000 reasoning samples across 25 agent configurations, shows that human-inspired cognitive structures can enhance LLM agents’ alignment with human strategic behaviour. Still, the relationship between agentic design complexity and human-likeness is non-linear, highlighting a critical dependence on underlying LLM capabilities and suggesting limits to simple architectural augmentation. Vince Trencsenyi, Agnieszka Mensfelt, Kostas Stathis |
ECAI | 2 |
| 2025 | Enhancing Quality-Diversity Optimization Through Domain-Specific Dissimilarity as Crowding DistanceabstractQuality-diversity algorithms aim to simultaneously optimize solution performance and maintain diversity within a population. In this paper, we explore the use of NSGA-II as a quality-diversity algorithm for the evolutionary design of 3D structures, modifying its crowding distance calculation to utilize dissimilarity measures. While NSGA-II is widely employed for multi-objective optimization, its use of fitness for calculating crowding distance may not be the most effective for tasks requiring solution diversity. We propose leveraging both genetic and phenotypic dissimilarity metrics to improve diversity management. To evaluate this approach, we compare the standard NSGA-II using fitness-based crowding distance and Diversity-Enhancing NSGA-II (DE-NSGA-II) using various combinations of dissimilarity-based metrics for crowding distance and diversity scores. Experiments are conducted using two distinct genetic representations on two optimization tasks: height of the center of gravity of passive structures and velocity of active structures. Results demonstrate the potential of dissimilarity-based crowding distance to enhance the diversity and overall quality of solutions in complex evolutionary design tasks. Maciej Komosinski, Agnieszka Mensfelt |
GECCO | 2 |
| 2025 | Towards Logically Sound Natural Language Reasoning with Logic-Enhanced Language Model AgentsabstractLarge language models (LLMs) are increasingly explored as general-purpose reasoners, particularly in agentic contexts. However, their outputs remain prone to mathematical and logical errors. This is especially challenging in open-ended tasks, where unstructured outputs lack explicit ground truth and may contain subtle inconsistencies. To address this issue, we propose Logic-Enhanced Language Model Agents (LELMA), a framework that integrates LLMs with formal logic to enable validation and refinement of natural language reasoning. LELMA comprises three components: an LLM-Reasoner, an LLM-Translator, and a Solver, and employs autoformalization to translate reasoning into logic representations, which are then used to assess logical validity. Using game-theoretic scenarios such as the Prisoner's Dilemma as testbeds, we highlight the limitations of both less capable (Gemini 1.0 Pro) and advanced (GPT4o) models in generating logically sound reasoning. LELMA achieves high accuracy in error detection and improves reasoning correctness via self-refinement, particularly in GPT-4o. The study also highlights challenges in autoformalization accuracy and in evaluation of inherently ambiguous open-ended reasoning tasks. Agnieszka Mensfelt, Kostas Stathis, Vince Trencsenyi |
ICTAI | 1 |
| 2025 | Approximating Human Strategic Reasoning with LLM-Enhanced Recursive Reasoners Leveraging Multi-agent Hypergames
Vince Trencsenyi, Agnieszka Mensfelt, Kostas Stathis |
MABS | 2 |
| 2024 | Distance-Targeting Mutation Operator for Evolutionary Design of 3D StructuresabstractEvolutionary design of 3D structures - an automated design by the methods of evolutionary algorithms - is a hard optimization problem. One of the contributing factors is a complex genotype-to-phenotype mapping often associated with the genetic representations of the designs. In such case, the genetic operators may exhibit low locality, i.e., a small change introduced in a genotype may result in a significant change in the phenotype and its fitness, hampering the search process. To overcome this challenge in evolutionary design, we introduce the Distance-Targeting Mutation Operator (DTM). The aim of this operator is to create offspring whose distance to the parent solution, according to a selected dissimilarity measure, approximates a predefined value. We compare the performance of the DTM operator to the performance of the mutation operator without parent-offspring distance control in a series of evolutionary experiments. We use different genetic representations, dissimilarity measures, and optimization goals, including velocity and height of active and passive 3D structures. The introduced DTM operator outperforms the standard one in terms of best fitness in most of the considered cases. Maciej Komosinski, Agnieszka Mensfelt |
GECCO | 2 |
| 2024 | Explaining Teleo-reactive Strategic BehaviourabstractGame-theoretic simulations are a powerful tool for exploring strategic interactions and guiding decision-making in fields ranging from business to economics and politics. However, simulations of such complex systems are often difficult to understand intuitively and debug. In this context, particularly challenging and hard to detect are logical (intentional) errors resulting from faulty logic in an agent's strategy. To address this challenge, we develop a framework for question-based explanations, enhancing the understanding of agents' behaviour. We focus on teleo-reactive agents in game-theoretical simulations, specifically in tournament and evolutionary contexts. Our approach centres on trace-based explanations, utilising behavioural logs to identify the steps leading to particular outcomes. We formally describe explanation templates for “why” and “why not” question types, linking them to agents' goals, beliefs, and condition-action rules. Furthermore, we provide formal definitions of the answers to these questions, linking them to output templates. The methodology is demonstrated through example dialogue scenarios, showing how these explanations can improve debugging efficiency by offering high-level insights into agents' behaviours. Nausheen Saba Shahid, Agnieszka Mensfelt, Kostas Stathis |
ICTAI | 2 |
| 2021 | Automated Development of Latent Representations for Optimization of Sequences Using AutoencodersabstractIn this paper, we propose an automated method for the development of new representations of sequences. For this purpose, we introduce a two-way mapping from variable length sequence representations to a latent representation modelled as the bottleneck of an LSTM (long short-term memory) autoencoder. Desirable properties of such mappings include smooth fitness landscapes for optimization problems and better evolvability. This work explores the capabilities of such latent encodings in the context of optimization of 3D structures. Various improvements are adopted that include manipulating the autoencoder architecture and its training procedure. The results of evolutionary algorithms that use different variants of automatically developed encodings are compared. Piotr Kaszuba, Maciej Komosinski, Agnieszka Mensfelt |
CEC | 3 |
| 2019 | A Flexible Dissimilarity Measure for Active and Passive 3D Structures and Its Application in the Fitness-Distance Analysis
Maciej Komosinski, Agnieszka Mensfelt |
EvoApplications | 2 |