VLDB 2026 Research / reviewers in the wild / expert
Tiziano Santilli
dblp:345/7061
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0005-9505-3846ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CIAO - Code In Architecture Out - Automated Software Architecture Documentation with Large Language ModelsabstractSoftware architecture documentation is essential for system comprehension, yet it is often unavailable or incomplete. While recent LLM-based techniques can generate documentation from code, they typically address local artifacts rather than producing coherent, system-level architectural descriptions. This paper presents a structured process for automatically generating system-level architectural documentation directly from GitHub repositories using Large Language Models. The process, called CIAO (Code In Architecture Out), defines an LLM-based work-flow that takes a repository as input and produces system-level architectural documentation following a template derived from ISO/IEC/IEEE 42010, SEI Views & Beyond, and the C4 model. The resulting documentation can be directly added to the target repository. We evaluated the process through a study with 22 developers, each reviewing the documentation generated for a repository they had contributed to. The evaluation shows that developers generally perceive the produced documentation as valuable, comprehensible, and broadly accurate with respect to the source code, while also highlighting limitations in diagram quality, high-level context modeling, and deployment views. We also assessed the operational cost of the process, finding that generating a complete architectural document requires only a few minutes and is inexpensive to run. Overall, the results indicate that a structured, standards-oriented approach can effectively guide LLMs in producing system-level architectural documentation that is both usable and cost-effective. Tiziano Santilli, Domenico Amalfitano, Anna Rita Fasolino, Patrizio Pelliccione |
ICSA | 2 |
| 2026 | A decontextualized LLM-based safeguard technique for automated jailbreak mitigationabstractContext: Large Language Models (LLMs) are increasingly deployed in high-risk settings, where harmful or unethical outputs remain a risk. Adversarial prompting (“jailbreaks”) can circumvent default safeguards. Emerging regulation (e.g., the EU AI Act) demands proactive controls that verify outputs before delivery. Objectives: We present and evaluate D-SHIELD, a plugin-based safeguard that separates generation from validation via a stateless, decontextualized validator. Objectives are to assess alignment with expert judgments, evaluate end-to-end mitigation on publicly sourced jailbreaks, compare with representative plugin-based defenses, and examine a lightweight configuration optimized for cost without reducing protection. Methods: D-SHIELD routes candidate responses from the user-facing LLM to a secondary, decontextualized LLM operating in isolation (no prompt or conversation context) to classify each response based on indications of prohibited content derived from the EU AI Act, The General-Purpose AI Code of Practice, GDPR, and provider policies. This decontextualized design intentionally prevents prompt contamination, adversarial framing, and conversational drift from influencing the validation decision, addressing key weaknesses of context-aware validators. We create an expert-labeled dataset from designed jailbreaks for direct comparison with the decontextualized validator’s classification. We then embed the validator in a working prototype and evaluate on publicly sourced jailbreaks. Finally, we conduct a comparative study against baseline jailbreak-mitigation techniques and analyze a lightweight guard variant. Results: The decontextualized validator closely aligns with expert decisions, especially for explicit harms, while adopting a conservative stance on borderline cases. In prototype evaluation on publicly sourced jailbreaks, the safeguard blocked most harmful responses. Compared with baselines, D-SHIELD yields fewer successful attacks under a common benchmark. The lightweight variant delivers comparable protection at markedly lower cost. Conclusion: Decontextualized, output-level validation provides an effective, regulation-aligned solution for LLM safety. Restricting the validator to the generated text complements input-level defenses and supports practical deployment, particularly in a lightweight configuration. Tiziano Santilli, Domenico Amalfitano, Anna Rita Fasolino, Patrizio Pelliccione |
Inf. Softw. Technol. | 1 |
| 2026 | LLM-Assisted Reinforcement Learning for Affective Game Adaptation EICS029abstractAdaptive games increasingly combine player modeling, affect sensing, and runtime control, but evidence for how large language models (LLMs) and Reinforcement Learning (RL) could participate in such loops remains limited. We present an affect-aware game adaptation architecture that combines a warm-started tabular Q-learning controller with bounded LLM-assisted action suggestion inside a MAPE-K loop. The LLM is not allowed to act freely: it can only recommend one action from a fixed, predefined adaptation menu, while reinforcement learning remains responsible for value updates and policy improvement. We implemented the architecture in CubeWars, a top-down shooter instrumented with browser-side facial-expression sensing and four concrete adaptation levers: difficulty, zombie color, background audio, and modal gameplay prompts. We evaluated three within-subject conditions with 46 participants: a static baseline, RL-only adaptation, and LLM-assisted RL adaptation. The strongest objective result is that both adaptive conditions improved absolute progression outcomes relative to the static baseline, whereas RL-only and LLM-assisted RL did not differ significantly on normalized objective throughput measures. The clearest hybrid-specific benefits were experiential: the LLM-assisted condition produced higher engagement, stronger absorption, and greater awareness of adaptation than the RL-only condition. For affect, the hybrid condition showed the highest mean value on an exploratory reward-relevant emotion composite, but the pairwise hybrid-versus-RL-only contrast did not survive Bonferroni correction; we therefore interpret the affective findings primarily through their per-emotion decomposition and event-level analyses. The paper contributes: 1) a pattern for integrating constrained LLM advice into a self-adaptive interactive system without replacing the underlying RL controller; 2) an artifact-grounded implementation of affect-driven runtime game adaptation in a playable game; and 3) an empirical evaluation showing that, in the present evidence, the main added value of the hybrid design lies in perceived responsiveness and player experience. Mahyar Tourchi Moghaddam, Tiziano Santilli, Mina Alipour |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2026 | Continuous Behavioral Synthesis for Adaptive Health Dashboards: An LLM-Mediated Architecture Integrating Explicit Preference, Spatial Reorganization, and Attention Allocation Signals EICS027abstractThe engineering of adaptive user interfaces has traditionally relied on either rule-based systems encoding designer intuitions about user needs or machine learning approaches requiring substantial historical data before achieving effective personalization. We present a technical architecture that leverages Large Language Models as behavioral synthesis engines to enable immediate adaptation from sparse, heterogeneous user signals. Our system integrates three distinct behavioral channels, i) explicit micro-feedback on individual interface elements, ii) spatial priority inferred from manual widget reorganization through drag-and-drop interaction, iii) and attentional investment measured through dwell time during hover events, within a structured prompt engineering framework that continuously regenerates dashboard layouts while maintaining explanatory coherence. The architecture addresses the technical challenge of translating low-level interaction patterns into high-level design decisions through a layered prompt construction methodology that separates temporal context determination, behavioral signal extraction, explicit preference enforcement, and user profile synthesis. The approach combines manually specified behavioral interpretations and temporal heuristics with LLM-mediated synthesis, enabling the reconciliation of multiple simultaneous signals that would be difficult to encode through explicit rules alone. We demonstrate the system through an instantiation in the personal health monitoring domain, including an analytical evaluation of adaptation behavior across multiple scenarios and a working implementation managing fourteen distinct health metrics across seven widget visualization modalities. The evaluation compares profile-driven initialization, multi-signal behavioral adaptation, and presents the resulting interfaces through representative post-adaptation screenshots. The analytical evaluation shows that the system preserves explicit user constraints, keeps user-prioritized metrics in prominent positions, expands widgets that receive sustained attention. The technical contribution comprises the multi-modal behavioral aggregation strategy, the structured LLM prompt engineering approach for maintaining design consistency across regeneration cycles, and the explainability generation mechanism that exposes adaptation rationale to end users. Our work provides a reproducible engineering approach for building LLM-powered adaptive interfaces that can be generalized beyond health dashboards to any domain requiring continuous interface personalization from heterogeneous user behavior. A current limitation is that, while the system can infer and act on behavioral signals, it does not yet incorporate mechanisms to independently verify whether adaptations improve user experience without additional feedback signals. Tiziano Santilli, Mina Alipour, Mahyar Tourchi Moghaddam |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2025 | Automated Software Architecture Design Recovery from Source Code Using LLMs
Domenico Amalfitano, Tiziano Santilli, Patrizio Pelliccione, Anna Rita Fasolino |
ECSA | 3 |
| 2024 | Characterizing Software Architectural Metrics for Continuous Compliance in the Automotive DomainabstractThe software of critical systems, such as automotive, is increasingly required to change and evolve after production. In the automotive domain, this is a consequence of self-driving and connected cars, which continuously collect data from the field that is then exploited to produce safer and more advanced and reliable versions of the used algorithms or AI modules. Consequently, there exists a need for techniques and tools to facilitate incremental and Continuous Compliance with safety and security standards. This paper focuses on software architectural metrics that can be used for Continuous Compliance in the automotive domain. Our initial stride involved a literature review to find metrics capable of assessing software architectures. Subsequently, in collaboration with architecture, safety, and security experts in the automotive domain, we proposed a framework defining the characteristics these metrics must possess for continuous evaluation of software architectural compliance. The framework was used to characterize 48 metrics gathered from the literature review and to associate them with a score expressing their suitability to be used in software architecture Continuous Compliance processes. Domenico Amalfitano, Anna Rita Fasolino, Patrizio Pelliccione, Tiziano Santilli |
ICSA | 5 |
| 2024 | We're Drifting Apart: Architectural Drift from the Developers' PerspectiveabstractDespite the recognized importance of software architecture, it is common that the implementation diverges from the intended architecture over time. This phenomenon is referred to as architectural drift. In the past decades, mainly technical solutions and tools have been developed to detect and address architectural inconsistencies and drift. There is still a lack of evidence from the perspective of developers and a lack of best practices to manage drift. This mixed-methods study relies on interviews with 11 developers and a survey answered by 63 developers from different companies and domains. We analyzed the data by dividing developers into senior and junior to see the different perspectives based on work experience. We found that juniors tend to rely more on documentation, while seniors have a more experience-related approach. We identified practices that developers use to mitigate drift, including defining clear responsibilities, setting best practices, and maintaining reliable documentation. Finally, we designed and evaluated guidelines to help developers to face architectural drift. Emilie Anthony, Astrid Berntsson, Tiziano Santilli, Rebekka Wohlrab |
ICSA | 3 |
| 2023 | Quality Metrics in Software ArchitectureabstractThe importance of software architecture is largely recognized also in iterative and agile development settings. However, it is quite complex to provide evidence that an architecture is of good quality and that the architectural decisions are appropriate, correct, or optimal. Architecture evaluation aims at showing and providing confidence that design decisions contribute to fulfilling the stakeholder concerns. Some architecture evaluation methods are scenario-based and aim at balancing many potentially conflicting quality attributes. Other works focus on a specific quality attribute and provide metrics to measure it.In this paper we survey the state of the art in metrics for evaluating quality attributes of architectures. The elicited metrics are organized into a catalog, which associates them with the specific quality attributes they aim to measure. We contribute also an MDE framework that generates web views facilitating the analysis of architectures. In this way, researchers and practitioners can easily retrieve the metrics that are appropriate to their specific needs. The catalog of metrics and quality attributes is released to the research community and open to contributions from experts and practitioners. Samira Silva, Adiel Tuyishime, Tiziano Santilli, Patrizio Pelliccione, Ludovico Iovino |
ICSA | 3 |