Dimitrios P. Panagoulias

dblp:303/7199 · DBLP profile ↗
← Back
17ranked-venue papers
14as first author
17since 2021 · last 2025
0000-0002-9421-141XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 11 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MemoryBERT: A Systematic Framework for Memory Identification
abstract
This paper introduces a systematic framework for the identification of memory statements in human-AI interactions. We present a novel benchmark dataset of 4000 labeled statements spanning five types of memory (episodic, semantic, spatial, emotional, associative) and a non-memory category, generated through a controlled GPT-4o and Faker-based pipeline with thematic randomization. The dataset was balanced across memory categories to ensure equitable class representation, while domain contexts (e.g., medical, educational, personal) were grouped by span identifiers to maintain contextual consistency without affecting label stratification. A representative, categorybalanced subset was reviewed and validated by board-certified psychiatrists to confirm semantic and cognitive plausibility, enhancing dataset reliability. On this corpus, we fine-tuned a RoBERTa-based transformer backbone, resulting in a specialized model we term MemoryBERT, which achieved perfect accuracy across all classes. Unlike prior studies that presuppose reliable memory categorization, our contribution establishes the first dedicated backbone for systematic memory identification in AI systems. As this work represents a feasibility study and preliminary stage of a broader research effort, future work will extend expert validation, expand the dataset with humanauthored and adversarial samples, and benchmark additional transformer architectures on the same classification task.
Dimitrios P. Panagoulias, Persephone Papatheodosiou, Anastasios Bonakis, Dimitris G. Dikeos, Maria Virvou, George A. Tsihrintzis
BIBE1
2025 A Simulation Framework for Battery Optimization Under Extreme Energy Imbalance Using Machine Learning Forecasting
abstract
The increasing integration of renewable sources introduces volatility and uncertainty into energy systems, particularly under extremely low production and/or extremely high consumption events. To address this, we present BASTION.ev (Bayesian Simulation & Forecasting for Extreme Events), a modular and scalable framework for intelligent energy management, which combines Extreme Value Analysis, Machine Learningbased forecasting, and simulation-driven optimization. We opt for grid search instead of Bayesian Optimization to simulate and optimize control strategies due to its lower implementation complexity and transparent evaluation across parameter combinations. Specifically, our proposed framework ingests time-series energy data to engineer temporal and contextual features to support predictive modeling. Extreme Value Analysis is used to define critical thresholds using generalized Pareto distributions, allowing the classification of operational states into extreme low, normal, and extreme high imbalance scenarios. These states are forecasted using tree-based multi-class classifiers across multiple time horizons (e.g., 1 to 72 hours ahead). The simulation of battery behavior is then evaluated under forecasted imbalance conditions, with control strategies (charge/discharge rates) optimized through a Bayesian Optimization to finetune parameters in order to maximize energy efficiency and, subsequently, economic return. Although based on mock-up data, simulations reveal that high feed-in profits are achieved with aggressive charging and minimal discharging. As a use case in this work, we focused our simulation on a 72-hour ahead forecasting window, where, despite class imbalance, the model achieved a macro F1-score of 0.63 and an overall accuracy of 91 %. The predicted imbalance states drive a simulation engine that evaluates energy system responses across control strategies. To optimize performance, we employ grid search as a baseline, over manually predefined charge/discharge configurations. Using mock-up data, the top configuration (charge$=15 \text{kWh}$, discharge$=1 \text{kWh}$) yielded a simulated profit of € 420.3, with only 66 kWh imported and 4335 kWh fed into the grid. Then, we use Bayesian Optimization for parameter finetuning and results improve greatly, Yielding € 584.00 simulated profit with 0 kWh imported energy for a fully self-sufficient PV production and 5840 kWh surplus fed into the grid with a charge rate of 20 kWh and a discharge rate of 0.1 kWh.
Dimitrios P. Panagoulias, Elissaios Sarmas, Vangelis Marinakis, Maria Virvou, George A. Tsihrintzis
ICTAI1
2025 Trustworthy Integration of Generative AI and LLMs in the Energy Digital Spine: Epistemic Risk Models and Trust Performance Indicators
abstract
This paper introduces Energy-VERITIES, a framework for modeling epistemic risk and trust towards the integration of generative AI and large language models (LLMs) within the energy digital spine ecosystem. It defines epistemic risk zones using a set-theoretic Venn model over AI outputs ($\boldsymbol{A}$), expert ($E$) and non-expert ($N$) beliefs, internet sources ($D \_I$), and verified domain facts ($D$), with a critical risk zone formalized as (D_I$\cap A \cap E \cap N$) - D, representing hallucinated outputs that appear credible due to overlapping AI, expert, and stakeholder beliefs but diverge from verified domain facts. To assess energy stakeholderAI interaction, we propose a set of Trust Performance Indicators (TPIs) that quantify epistemic outcomes from the user's perspective, extending the confusion matrix and the previously devised VIRTSI model, which characterizes dynamic human trust states through a finite state automaton informed by user validation behavior. These include, for example, MJO-HRAR (Misconceptually Justified Overtrust), where users accept false outputs after flawed validation, and BOV-HRAR (Blind Overtrust), where no validation occurs. An empirical study of ChatGPT's energy advice in Greece illustrates the framework's diagnostic value, revealing that 39 % of validated responses were still false, over 30 % of hallucinations were accepted without validation, and correct answers were often dismissed. These findings indicate the need for improvements in interface-level explainability, validation support, and targeted AI literacy for energy stakeholders. The Energy-VERITIES framework offers a foundation for identifying epistemic risks and guiding the trustworthy integration of generative AI and LLMs in high-stakes domains, like energy.
Maria Virvou, George A. Tsihrintzis, Vangelis Marinakis, Elissaios Sarmas, Dimitrios P. Panagoulias, Evangelia-Aikaterini Tsichrintzi
ICTAI5
2025 An OSCE-Inspired Framework for Fine-Tuning Multimodal LLMs in Medical Diagnosis
abstract
We present COGNET-MD-X-REQ: a COGnitive NETwork Evaluation Toolkit that identifies precise fine-tuning needs for Large Language Models (LLMs) in multimodal medical diagnosis. COGNET-MD-X-REQ is a novel, two-step evaluation framework, inspired by Objective Structured Clinical Examinations (OSCEs), that enhances LLM applicability and precision. Our framework integrates IoT-driven data retrieval with structured interaction evaluation and domain-specific analysis. Leveraging Image-Metadata Analysis, Named Entity Recognition, and Knowledge Graphs, COGNET-MD-X-REQ identifies weak performance areas by analyzing multimodal model outputs across medical subdomains. We selected GPT-4V to use in the evaluation of our proposed framework on a set of publicly available image-based MCQs in General Pathology. The model achieved 84% accuracy, with notable weaknesses in cardiovascular conditions like atherosclerosis. The framework pinpointed domain-specific deficiencies, enabling targeted fine-tuning. The evaluation leads to the conclusion that our proposed COGNET-MD-X-REQ framework introduces a precision-focused, iterative approach to fine-tuning LLMs by dynamically identifying under-performing areas or knowledge domains that can be related to organs or/and specific conditions and health states. This method reduces reliance on broad retraining, making it suitable for resource-sensitive and safety-critical domains such as medical diagnostics.
Dimitrios P. Panagoulias, Anastasios P. Palamidas, Maria Virvou, George A. Tsihrintzis
KES1
2025 Truth in Trees: A Multilayered Linguistic Framework for Automated Fact-Checking with AI-Driven Reasoning
abstract
In an era of information overload, distinguishing fact from misinformation is a critical challenge. We introduce AlethaNet, a multilayered linguistic framework for automated fact-checking that integrates syntactic, semantic, and pragmatic analysis with contextual metadata. Our approach combines feature extraction, machine learning (XGBoost), and SHAP-based interpretability with a dynamic credibility scoring system. This score is computed using a weighted formula incorporating historical truthfulness, speaker metadata (e.g., job title, party, platform), and AI-inferred reasoning via DeepSeek R1—a generative LLM shown to approximate expert evaluations on the LIAR dataset. Unlike static, binary classification models, AlethaNet frames the problem as truthful vs. misleading, offering a more nuanced detection of misinformation. Experiments demonstrate improved accuracy (75.65%) when incorporating credibility scores. DeepSeek R1 extends the framework’s ability to assess previously unseen sources, while SHAP ensures transparency by highlighting the most influential features. AlethaNet thus presents a scalable, explainable, and autonomous method for adaptive credibility assessment, advancing AI-driven misinformation detection.
Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis
KES1
2025 Which AI and How Trusted by Human Energy Stakeholders? Comparing Generative AI ChatGPT With Domain Specific AI-ENERGIA-SYS
abstract
Artificial Intelligence (AI) has been growing significantly recently, based on machine learning, deep learning and the latest Generative AI and Large Language Models (LLMs), like ChatGPT. All AI tools are promising to improve decision-making in many domains, including complex ones, such as energy. As such, there have been energy-specific AI systems that incorporate domain expertise, in addition to general-purpose LLMs that integrate knowledge on a plethora of domains, including energy, based on their training from the internet. However, AI systems can produce errors due to their probabilistic nature, and many users, like energy stakeholders, lack the training to use them effectively. Which AI do we mean and how well is it trusted. It is worth exploring differences in AI systems.By employing the VIRTSI model, this paper compares trust states of energy stakeholders on domain specific AI systems versus general-purpose generative AI. For the purposes of the comparison, we use a domain specific AI-based Energy system entitled AI-ENERGIA-SYS that has been previously developed versus ChatGPT 4.0 by OpenAI. VIRTSI is a rigorous computational model for the dynamics of human trust states, spanning from overtrust to distrust, through user modelling and quantifies the efficiency of the interaction in VIRTSI-adapted confusion matrices. The findings reveal that ChatGPT is persuasive and user-friendly to stakeholders, but its lack of domain awareness and explainability leads to overtrust, risking decision quality. In contrast, AI-ENERGIA-SYS, though more complex, supports better trust calibration, meaning that, when trusted, it is usually correct, but it is not trusted so frequently as ChatGPT. The results suggest that future combinations of energy-specific AI systems with general-purpose generative AI could provide improved trust dynamics and more effective decision support.
George A. Tsihrintzis, Elissaios Sarmas, Vangelis Marinakis, Dimitrios P. Panagoulias, Evangelia-Aikaterini Tsichrintzi, Maria Virvou
SMC4
2025 A framework for evaluation and requirement extraction for fine-tuning of Large Language Models in multimodal medical diagnosis
abstract
Objective: Large language models constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. In this study, we propose a novel evaluation framework for extraction of fine-tuning requirements based on the Objective Structured Clinical Examinations (OSCE) that can increase LLM potential and applicability. Methods: We developed an OSCE based evaluation meta-framework leveraging IoT-based data retrieval with a two-step approach designed to analyze and guide improvement of LLMs in multimodal medical diagnosis: (1) structured interaction evaluation and (2) domain-specific analysis of extracted data. Using Image-Metadata Analysis (IMA), Named Entity Recognition (NER), and Knowledge Graphs (KG), this framework identifies image domains, extracts relevant entities, and assesses connections in KGs. These methods collectively reveal areas for improvement, guiding fine-tuning to enhance diagnostic accuracy and contextual understanding in medical applications. Results: Using this paradigm, (1) we evaluate the correctness and accuracy of generated medical diagnosis with publicly available multimodal-multiple-choice-questions in the vast domain of General Pathology and (2) proceed to the domain-specific analysis. We identify and visualize the model performance across specific organs, diseases, and pathological themes, detecting areas of lower accuracy, such as in cardiovascular conditions like atherosclerosis. This targeted approach enables precision-focused fine-tuning, applying additional data to specific weaknesses rather than a broad, generalized tuning across all pathology. Contributions: Our framework’s primary contribution is its OSCE-inspired ability to dynamically identify and target under-performing areas within a broad domain, enhancing fine-tuning efficiency and diagnostic accuracy in a resource-effective and iterative manner removing the dependency on bulk adjustments, making it particularly suitable for sensitive applications where precision and resource efficiency are essential, such as in medical diagnostics.
Dimitrios P. Panagoulias, Anastasios P. Palamidas, Maria Virvou, George A. Tsihrintzis
Knowl. Based Syst.1
2024 Memory and Schema in Human-Generative Artificial Intelligence Interactions
abstract
In this paper, we explore memory and Schema through the lens of mathematical representation. By defining themes related to cognitive processes, we propose a compression algorithm that measures and retains important information of previous (historic) user-Generative AI interaction. Time and memory are interconnected via a decay mechanism, where context gets more abstract as time passes. At the same time, memories with thematic resemblance can be structured into Schemas using set theory, based on rules influenced by a maturity threshold. This threshold is determined by factors such as criticality, emotional-user interaction, and mass, which is the measure of the summation of related memories. Our approach reconstructs memories into reusable parameters and formulates Schemas, offering a foundation for more personalized and cost-effective human-GAI interactions.
Dimitrios P. Panagoulias, Persephone Papatheodosiou, Anastasios Bonakis, Dimitrios Dikeos, Maria Virvou, George A. Tsihrintzis
ICTAI1
2024 Knowledge Space reduction via Sequential Language Model Integration
abstract
In Large Language Models (LLMs), such as GPT, BERT, Mistral or others, “reducing the domain space” for text generation involves limiting the range of content that can be utilized for generating responses. This approach aims at enhancing the relevance and precision of the text produced. While various strategies exist to achieve this, this work explores Sequential Language Model Integration (SLMI), which mirrors the organization and distribution of knowledge across different fields of expertise. More specifically, SLMI is the technique of linking multiple LLMs (LLM-Chains) in a systematic manner. In this paper, we refer to a process of choosing, linking and connecting LLMs with other services (often to complete a generative task, invoke external functions and machine learning services, or tackle problems) as “Large Language Models as a Service”. We outline the development and evaluation process of an SLMI methodology to refine response accuracy. Focusing on the medical field, we also establish a framework for knowledge reduction based on “knowledge paths”, analogous to the distinct specializations within medicine. We apply this framework to a dermatology case study and utilize our evaluation pipeline to assess the results. Reducing the knowledge domain from medicine in general down to dermatology, we tested our methodology and found gains regarding accuracy and diagnostic improvement, as well as a reduction in costs regarding total tokens generated.
Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis
KES1
2024 Leveraging Artificial Intelligence for personalised insomnia-sleep calibration via the Big Five Personality Traits
abstract
This paper introduces Morpheas, an AI-empowered sleep evaluation and calibration system that leverages state-of-the-art technologies, like Large Language Models (LLMs) and Named Entity Recognition (NER). Morpheas integrates Sequential Language Model Integration (SLMI) workflows to simulate the initial steps of sleep disorder diagnosis, utilizing the GPT-4 engine enhanced with Rules of Conduct. Using SLMI and medical and psychological diagnostic tools, we propose a novel multi-step personalisation methodology for creating a gradation system for the improvement of patient-AI interactions. To test and showcase this personalisation approach, we simulate keeping a sleep diary for diagnosing insomnia using a trait-based personalised LLM aimed at addressing sleep concerns. For this purpose, we apply the Big Five Personality Traits (BFPT) where LLMs are again used to extract the responses and facilitate the patient throughout the process, providing guidance and explanation. We then extract the entities from these interactions with NER, in order to identify patterns, provide an explainability basis for patients and effectively customize Cognitive Behavioral Therapy for Insomnia (CBTi), whether through one-on-one sessions or via digital platforms.
Persephone Papatheodosiou, Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis, Anastasios Bonakis, Dimitrios Dikeos
KES2
2024 A novel framework for artificial intelligence explainability via the Technology Acceptance Model and Rapid Estimate of Adult Literacy in Medicine using machine learning
Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis
Expert Syst. Appl.1
2023 Evaluation of ChatGPT-supported diagnosis, staging and treatment planning for the case of lung cancer
abstract
In this paper, we evaluate the validity, accuracy, usefulness, and specificity of medical diagnoses related to lung cancer and its staging provided by ChatGPT based on symptoms described by humans. The evaluation is grounded on three main pillars: the validity and accuracy of answers in relation to context and associated references. The specificity and usefulness of the information for both doctors and patients. The economic value added to the healthcare system, determined by several weighted factors derived from the provided answers. The system’s responses are expected to return proposed diagnoses and diagnostic steps, ranked by probability and importance. A specialist conducts the review process.
Dimitrios P. Panagoulias, Filippos A. Palamidas, Maria Virvou, George A. Tsihrintzis
AICCSA1
2023 An Empirical Study Concerning the Impact of Perceived Usefulness and Ease of Use on the Adoption of AI-Empowered Medical Applications
abstract
In this paper, we explore multiple theoretical frameworks to understand and predict user behavior concerning the adoption of innovative, AI-empowered technologies in healthcare. Specifically, our research centers on evaluating the potential adoption rate of AI-empowered medical applications among physicians. To provide empirical support for our investigation, we carried out a comprehensive study employing questionnaires that were disseminated to a practicing medical doctors and medical students. Our methodological framework incorporates two key theories: the Technology Acceptance Model (TAM) and the Diffusion of Innovation Theory (DOI). Utilizing these theories allows us to examine critical factors that influence physicians' willingness to adopt new technologies, such as perceived ease of use and perceived usefulness. Through a nuanced understanding of doctors' perceptions and attitudes toward AI, our research aims to craft targeted strategies that could enhance the rate of adoption for these cutting-edge medical technologies. The overarching goal is to accelerate the integration of AI applications into clinical practice, thereby improving healthcare outcomes and operational efficiencies.
Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis
BIBE1
2023 Rule-Augmented Artificial Intelligence-empowered Systems for Medical Diagnosis using Large Language Models
abstract
In this paper, we investigate the enhancement of Artificial Intelligence (AI) technologies in healthcare and the better understanding of medical literature with the use of Large Language Models (LLMs) and Natural Language Processing (NLP). Specifically, we introduce a rule-augmented AI-empowered system which incorporates a rule-based decision system, the ChatGPT application programming interface (API), and other external machine learning and analytical APIs to offer diagnostic suggestions to patients. The complexities of patient healthcare experiences, including doctor-patient interactions, understanding levels, treatment procedures, and preventive care, are considered. We illustrate how a diagnostic process typically integrates various strategies depending on various factors. To digitize the greatest portion of the process, we propose and illustrate the use of LLMs for humanizing the communication process and investigating ways to reduce burdens and costs in primary healthcare. We also outline a theoretical decision model for evaluating the use of technological components from external sources versus building them from scratch. The paper is structured into sections detailing background theories and context, our proposed and implemented rule-augmented AI-empowered system, as well as a system test in a corresponding use case. Finally, the paper key findings are presented, which contribute valuable insights for future work in this field.
Dimitrios P. Panagoulias, Filippos A. Palamidas, Maria Virvou, George A. Tsihrintzis
ICTAI1
2023 Tailored Explainability in Medical Artificial Intelligence-empowered Applications: Personalisation via the Technology Acceptance Model
abstract
The great momentum of Artificial Intelligence-empowered applications makes the requirement for detailed and tailored explainability frameworks more crucial. This is particularly evident in the medical domain, where validation of methodologies and outcomes is very important to the adoption of such systems. The depth and the level of understanding of Artificial Intelligence-related concepts is a significant design parameter and necessitates a systemic approach to ensure that a proper level of transparency is incorporated in an Artificial Intelligence-empowered application. In this paper, we propose a novel and generalised approach for the analysis of user requirements and abilities in relation to Artificial Intelligence-empowered applications. Specifically, we use the Technology Acceptance Model (TAM), as a technical methodology to measure the user perception of usefulness and usability of a technology and, subsequently, identify the corresponding depths of explainablity requirements. As a result, we design a layered and personalised explainability framework that may increase adoption rates of domain-specific Artificial Intelligence-empowered technologies.
Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis
ICTAI1
2022 A microservices-based iterative development approach for usable, reliable and explainable A.I.-infused medical applications using R.U.P
abstract
The Rational Unified Process (R.U.P.) is an iterative Software Engineering Process, that ensures alignment between engineers and stakeholders through optimized and detailed partinionalised steps within predefined constraints [1]. The independently deployable services that are components of distributed systems are called microservices. They are part of applications and are easier to manage and scale. Each microservice serve different purpose and has a unique responsibility making it easier to understand manage and collaborate on. Medical applications are created to assist in treatment, disease prevention and health optimization. However patients' needs and abilities vary and patients' requirements and definitions on the usability aspect of an application are different and should be acknowledged for the application to be successful and for the patients/users to benefit from it. In this study the development of an A.I.-infused medical application, is outlined through the R.U.P. methodology and built using microservices, where the patients' needs and requirements are at the center of continuous development process focused on improvements.
Dimitrios P. Panagoulias, Maria Virvou, George A. Tsihrintzis
ICTAI1
2021 Biomarker-based deep learning for personalized nutrition
abstract
In the digital era, disease diagnosis and patient management is taking a decisive leap forward to dynamically link personalized medical decision making, disease recognition and management with the massive individuality and uniqueness of the human body. Individual needs are recognised and associated with unique individual demands and tastes via automated processes and the pre-processing power of complex and accurate recom-mender systems. The gigantic pool of biometrics and biomarkers are used to predict outcomes and identify patterns in patients in an individualized manner. In this paper, we continue and improve upon recent previous work of ours [1] and develop software systems for personalized nutrition based on biomarkers and deep learning algorithms. Evaluation on real data demonstrates the high performance of our approach.
Dimitrios P. Panagoulias, Dionisios N. Sotiropoulos, George A. Tsihrintzis
ICTAI1