Dumitru Roman

dblp:49/7031 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LUMEN: Enhancing IoT System Observability with Multi-Agent Large Language Models and Knowledge Graphs
abstract
The rapid expansion of Internet of Things (IoT) systems has transformed industries through real-time monitoring and automation, generating vast and heterogeneous data streams. As IoT networks expand, the increasing volume and diversity of data, spanning real-time telemetry, device logs, and historical records, complicate the management of IoT systems, including system monitoring, analysis, and reasoning. To address this challenge, we introduce LUMEN (Large Language Models as Unified Multi-Agent Systems for IoT ENhancement), a novel approach combining multi-agent Large Language Models (LLMs), knowledge graphs, and heterogeneous databases to enable cognitive digital twins for IoT observability. LUMEN models IoT systems as knowledge graphs, capturing device relationships and metadata while monitoring data is stored in time-series or object databases. Specialized LLM-based agents collaborate dynamically to analyze IoT systems and explain the findings in natural language, generating and executing analysis code when necessary. Integrated with off-the-shelf network monitoring tools, LUMEN facilitates semantic reasoning and human-in-the-loop collaboration, delivering adaptive insights across diverse data contexts. Two industrial case studies demonstrate the ability of LUMEN to automate analysis workflows, enhance system adaptability, and provide interpretable analytics. This work advances IoT observability by integrating LLMs, semantic intelligence, and explainable analytics into a scalable and adaptive solution using a multi-agent architecture for complex IoT systems.
Adela-Aniela Nedisan, Arda Goknil, Dumitru Roman, Ahmet Soylu
ACM Trans. Internet Things4
2025 Positioning LLM-Enabled Agents as Legal Compliance Aides for Data Pipelines
Adela-Aniela Nedisan, Nikolay Nikolov, Carl-Henrik Lien, Arda Goknil, Sagar Sen, Ahmet Soylu, Dumitru Roman
RuleML+RR8
2024 European Data Science Day: KDD-2024 Special Day
abstract
The European Data Science Day offers a full day focused exclusively on innovative KDD-relevant research and development projects from national and regional funding programs, as well as corporate, start-up, and nonprofit channels.The idea is to bring together a diverse community of researchers in Data Science, Machine Learning, Language Technologies, and Knowledge Discovery, as well as partnerships in the social and physical sciences/arts, to showcase the state-of-the-art in research and applications.
Dunja Mladenic, Dumitru Roman
KDD2
2024 Cloud storage cost: a taxonomy and survey
abstract
Abstract Cloud service providers offer application providers with virtually infinite storage and computing resources, while providing cost-efficiency and various other quality of service (QoS) properties through a storage-as-a-service (StaaS) approach. Organizations also use multi-cloud or hybrid solutions by combining multiple public and/or private cloud service providers to avoid vendor lock-in, achieve high availability and performance, and optimise cost. Indeed cost is one of the important factors for organizations while adopting cloud storage; however, cloud storage providers offer complex pricing policies, including the actual storage cost and the cost related to additional services (e.g., network usage cost). In this article, we provide a detailed taxonomy of cloud storage cost and a taxonomy of other QoS elements, such as network performance, availability, and reliability. We also discuss various cost trade-offs, including storage and computation, storage and cache, and storage and network. Finally, we provide a cost comparison across different storage providers under different contexts and a set of user scenarios to demonstrate the complexity of cost structure and discuss existing literature for cloud storage selection and cost optimization. We aim that the work presented in this article will provide decision-makers and researchers focusing on cloud storage selection for data placement, cost modelling, and cost optimization with a better understanding and insights regarding the elements contributing to the storage cost and this complex problem domain.
Akif Quddus Khan, Mihhail Matskin, Radu Prodan, Christoph Bussler, Dumitru Roman, Ahmet Soylu
World Wide Web (WWW)5
2023 TRANSQLATION: TRANsformer-based SQL RecommendATION
abstract
The exponential growth of data production emphasizes the importance of database management systems (DBMS) for managing vast amounts of data. However, the complexity of writing Structured Query Language (SQL) queries requires a diverse range of skills, which can be a challenge for many users. Different approaches are proposed to address this challenge by aiding SQL users in mitigating their skill gaps. One of these approaches is to design recommendation systems that provide several suggestions to users for writing their next SQL queries. Despite the availability of such recommendation systems, they often have several limitations, such as lacking sequence-awareness, session-awareness, and context-awareness. In this paper, we propose TRANSQLATION, a session-aware and sequence-aware recommendation system that recommends the fragments of the subsequent SQL query in a user session. We demonstrate that TRANSQLATION outperforms existing works by achieving, on average, 22% more recommendation accuracy when having a large amount of data and is still effective even when training data is limited. We further demonstrate that considering contextual similarity is a critical aspect that can enhance the accuracy and relevance of recommendations in query recommendation systems.
Shirin Tahmasebi, Amir Hossein Payberah, Ahmet Soylu, Dumitru Roman, Mihhail Matskin
IEEE Big Data4
2023 Towards Graph-based Cloud Cost Modelling and Optimisation
abstract
Cloud computing has become an increasingly popular choice for businesses and individuals due to its flexibility, scalability, and convenience; however, the rising cost of cloud resources has become a significant concern for many. The pay-per-use model used in cloud computing means that costs can accumulate quickly, and the lack of visibility and control can result in unexpected expenses. The cost structure becomes even more complicated when dealing with hybrid or multi-cloud environments. For businesses, the cost of cloud computing can be a significant portion of their IT budget, and any savings can lead to better financial stability and competitiveness. In this respect, it is essential to manage cloud costs effectively. This requires a deep understanding of current resource utilization, forecasting future needs, and optimising resource utilization to control costs. To address this challenge, new tools and techniques are being developed to provide more visibility and control over cloud computing costs. In this respect, this paper explores a graph-based solution for modelling cost elements and cloud resources and potential ways to solve the resulting constraint problem of cost optimisation. We primarily consider utilization, cost, performance, and availability in this context. Such an approach will eventually help organizations make informed decisions about cloud resource placement and manage the costs of software applications and data workflows deployed in single, hybrid, or multi-cloud environments.
Akif Quddus Khan, Nikolay Nikolov, Mihhail Matskin, Radu Prodan, Christoph Bussler, Dumitru Roman, Ahmet Soylu
COMPSAC6
2023 ContrastNER: Contrastive-based Prompt Tuning for Few-shot NER
abstract
Prompt-based language models have produced encouraging results in numerous applications, including Named Entity Recognition (NER) tasks. NER aims to identify entities in a sentence and provide their types. However, the strong performance of most available NER approaches is heavily dependent on the design of discrete prompts and a verbalizer to map the model-predicted outputs to entity categories, which are complicated undertakings. To address these challenges, we present ContrastNER, a prompt-based NER framework that employs both discrete and continuous tokens in prompts and uses a contrastive learning approach to learn the continuous prompts and forecast entity types. The experimental results demonstrate that ContrastNER obtains competitive performance to the state-of-the-art NER methods in high-resource settings and outperforms the state-of-the-art models in low-resource circumstances without requiring extensive manual prompt engineering and verbalizer design.
Amirhossein Layegh, Amir Hossein Payberah, Ahmet Soylu, Dumitru Roman, Mihhail Matskin
COMPSAC4
2023 The Graph-Massivizer Approach Toward a European Sustainable Data Center Digital Twin
abstract
Modeling and understanding an expensive next-generation data center operating at a sustainable exascale performance remains a challenge yet to solve. The paper presents the approach taken by the Graph-Massivizer project, funded by the European Union, towards a sustainable data center, targeting a massive graph representation and analysis of its digital twin. We introduce five interoperable open-source tools that support this undertaking, creating an automated, sustainable loop of graph creation, analytics, optimization, sustainable resource management, and operation, emphasizing state-of-the-art progress. We plan to employ the tools for designing a massive data center graph, representing a digital twin describing spatial, semantic, and temporal relationships between the monitoring metrics, hardware nodes, cooling equipment, and jobs. The project aims to strengthen Bologna Technopole as a leading European supercomputing and big data hub offering sustainable green computing for improved societally relevant science throughput.
Martin Molan, Junaid Ahmed Khan, Andrea Bartolini, Roberta Turra, Giorgio Pedrazzi, Michael Cochez, Alexandru Iosup, Dumitru Roman, Joze M. Rozanec, Ana Lucia Varbanescu, Radu Prodan
COMPSAC8
2023 A Taxonomy for Cloud Storage Cost
Akif Quddus Khan, Nikolay Nikolov, Mihhail Matskin, Radu Prodan, Christoph Bussler, Dumitru Roman, Ahmet Soylu
MEDES6
2023 Scaling Data Science Solutions with Semantics and Machine Learning: Bosch Case
Baifan Zhou, Nikolay Nikolov, Zhuoxun Zheng, Xianghui Luo, Ognjen Savkovic, Dumitru Roman, Ahmet Soylu, Evgeny Kharlamov
ISWC6
2022 Dataclouddsl: Textual and Visual Presentation of Big Data Pipelines
abstract
This paper describes the DATACLOUDDSL language and the DEF-PIPE tool for describing Big Data pipelines. DAT-ACLOUDDSL has both a textual and a visual form and supports requirements obtained both from analyzing existing data pipeline specification tools and from interviews with relevant industrial actors. Particularly, DATACLOUDDSL supports (i) separation of concerns between design and run-time issues, (ii) reuse of previously developed pipeline steps and pipelines in designing new pipelines, (iii) flexible data transfer between pipelines steps and containerization of pipelines and pipeline steps, and (iv) integration of description and simulation components in Big Data pipeline orchestration systems. Additionally, it provides an interface to the discovery and deployment tools of the DataCloud toolbox.
Shirin Tahmasebi, Amirhossein Layegh, Nikolay Nikolov, Amir Hossein Payberah, Khoa Dinh, Vlado Mitrovic, Dumitru Roman, Mihhail Matskin
COMPSAC7
2022 SIM-PIPE DryRunner: An approach for testing container-based big data pipelines and generating simulation data
abstract
Big data pipelines are becoming increasingly vital in a wide range of data intensive application domains such as digital healthcare, telecommunication, and manufacturing for efficiently processing data. Data pipelines in such domains are complex and dynamic and involve a number of data processing steps that are deployed on heterogeneous computing resources under the realm of the Edge-Cloud paradigm. The processes of testing and simulating big data pipelines on heterogeneous resources need to be able to accurately represent this complexity. However, since big data processing is heavily resource-intensive, it makes testing and simulation based on historical execution data impractical. In this paper, we introduce the SIM - PIPE Dry Runner approach - a dry run approach that deploys a big data pipeline step by step in an isolated environment and executes it with sample data; this approach could be used for testing big data pipelines and realising practical simulations using existing simulators.
Aleena Thomas, Nikolay Nikolov, Antoine Pultier, Dumitru Roman, Brian Elvesæter, Ahmet Soylu
COMPSAC4
2022 TranSQL: A Transformer-based Model for Classifying SQL Queries
abstract
Domain-Specific Languages (DSL) are becoming popular in various fields as they enable domain experts to focus on domain-specific concepts rather than software-specific ones. Many domain experts usually reuse their previously-written scripts for writing new ones; however, to make this process straightforward, there is a need for techniques that can enable domain experts to find existing relevant scripts easily. One fundamental component of such a technique is a model for identifying similar DSL scripts. Nevertheless, the inherent nature of DSLs and lack of data makes building such a model challenging. Hence, in this work, we propose TRANSQL, a transformer-based model for classifying DSL scripts based on their similarities, considering their few-shot context. We build TRANSQL using BERT and GPT-3, two performant language models. Our experiments focus on SQL as one of the most commonly-used DSLs. The experiment results reveal that the BERT-based TRANSQL cannot perform well for DSLs since they need extensive data for the fine-tuning phase. However, the GPT-based TRANSQL gives markedly better and more promising results.
Shirin Tahmasebi, Amir Hossein Payberah, Ahmet Soylu, Dumitru Roman, Mihhail Matskin
ICMLA4
2021 Big Data Pipelines on the Computing Continuum: Ecosystem and Use Cases Overview
abstract
Organisations possess and continuously generate huge amounts of static and stream data, especially with the proliferation of Internet of Things technologies. Collected but unused data, i.e., Dark Data, mean loss in value creation potential. In this respect, the concept of Computing Continuum extends the traditional more centralised Cloud Computing paradigm with Fog and Edge Computing in order to ensure low latency pre-processing and filtering close to the data sources. However, there are still major challenges to be addressed, in particular related to management of various phases of Big Data processing on the Computing Continuum. In this paper, we set forth an ecosystem for Big Data pipelines in the Computing Continuum and introduce five relevant real-life example use cases in the context of the proposed ecosystem.
Dumitru Roman, Nikolay Nikolov, Ahmet Soylu, Brian Elvesæter, Radu Prodan, Dragi Kimovski, Andrea Marrella, Francesco Leotta, Mihhail Matskin, Ioannis Ledakis 0001, Konstantinos Theodosiou, Anthony Simonet, Fernando Perales, Evgeny Kharlamov, Alexandre Ulisses, Arnor Solberg, Raffaele Ceccarelli
ISCC1
2021 Locality-Aware Workflow Orchestration for Big Data
abstract
The development of the Edge computing paradigm shifts data processing from centralised infrastructures to heterogeneous and geographically distributed infrastructure. Such a paradigm requires data processing solutions that consider data locality in order to reduce the performance penalties from data transfers between remote (in network terms) data centres. However, existing Big Data processing solutions have limited support for handling data locality and are inefficient in processing small and frequent events specific to Edge environments. This paper proposes a novel architecture and a proof-of-concept implementation for software container-centric Big Data workflow orchestration that puts data locality at the forefront. Our solution considers any available data locality information by default, leverages long-lived containers to execute workflow steps, and handles the interaction with different data sources through containers. We compare our system with Argo workflow and show significant performance improvements in terms of speed of execution for processing units of data using our data locality aware Big Data workflow approach.
Andrei-Alin Corodescu, Nikolay Nikolov, Akif Quddus Khan, Ahmet Soylu, Mihhail Matskin, Amir Hossein Payberah, Dumitru Roman
MEDES7
2020 Scalable Execution of Big Data Workflows using Software Containers
abstract
Big Data processing involves handling large and complex data sets, incorporating different tools and frameworks as well as other processes that help organisations make sense of their data collected from various sources. This set of operations, referred to as Big Data workflows, require taking advantage of the elasticity of cloud infrastructures for scalability. In this paper, we present the design and prototype implementation of a Big Data workflow approach based on the use of software container technologies and message-oriented middleware (MOM) to enable highly scalable workflow execution. The approach is demonstrated in a use case together with a set of experiments that demonstrate the practical applicability of the proposed approach for the scalable execution of Big Data workflows. Furthermore, we present a scalability comparison of our proposed approach with that of Argo Workflows - one of the most prominent tools in the area of Big Data workflows.
Yared Dejene Dessalk, Nikolay Nikolov, Mihhail Matskin, Ahmet Soylu, Dumitru Roman
MEDES5
2020 Enhancing Public Procurement in the European Union Through Constructing and Exploiting an Integrated Knowledge Graph
Ahmet Soylu, Óscar Corcho, Brian Elvesæter, Carlos Badenes-Olmedo, Francisco Yedro Martínez, Matej Kovacic, Matej Posinkovic, Ian Makgill, Chris Taggart, Elena Simperl, Till C. Lech, Dumitru Roman
ISWC (2)12
2019 Semantically-Enabled Optimization of Digital Marketing Campaigns
Vincenzo Cutrona, Flavio De Paoli, Aljaz Kosmerlj, Nikolay Nikolov, Matteo Palmonari, Fernando Perales, Dumitru Roman
ISWC (2)7
2015 WSMO-Lite and hRESTS: Lightweight semantic annotations for Web services and RESTful APIs
Dumitru Roman, Jacek Kopecký, Tomas Vitvar, John Domingue, Dieter Fensel
J. Web Semant.1
2009 Stream Reasoning: A Survey and Further Research Directions
Gulay Ünel, Dumitru Roman
FQAS2
2008 WSMO Choreography: From Abstract State Machines to Concurrent Transaction Logic
Dumitru Roman, Michael Kifer, Dieter Fensel
ESWC1
2008 Semantic Web Service Choreography: Contracting and Enactment
Dumitru Roman, Michael Kifer
ISWC1
2007 A Multi-criteria Service Ranking Approach Based on Non-Functional Properties Rules Evaluation
Ioan Toma, Dumitru Roman, Dieter Fensel, Brahmananda Sapkota, Juan Miguel Gómez 0001
ICSOC2
2007 Reasoning about the Behavior of Semantic Web Services with Concurrent Transaction Logic
Dumitru Roman, Michael Kifer
VLDB1
2006 On Describing, Analyzing, and Executing Complex Behavior of Services
abstract
This paper gives a high level overview of a framework for describing, analyzing, and executing complex behavior of services in the context of Service Oriented Computing. Having as a starting point the commonly used patterns in workflow specifications, we propose an extension to incorporate temporal constraints, and a way to root them on a logic for transaction composition (Concurrent Transaction Logic) and on a methodology for modelling systems (Abstract State Machines). We motivate our choices and explain the potential benefits of our framework. Finally, we propose concrete steps for future research in this area.
Dumitru Roman, Ioan Toma, Dieter Fensel
ICSEA1
2005 Peer-to-Peer Technology Usage in Web Service Discovery and Matchmaking
Brahmananda Sapkota, Laurentiu Vasiliu, Ioan Toma, Dumitru Roman, Christoph Bussler
WISE4
2003 FPGA-Based Hardware/Software CoDesign of an Expert System Shell
Aurel Netin, Dumitru Roman, Octavian Cret, Kalman Pusztai, Lucia Vacariu
FPL2