Mili Orucevic

dblp:385/8378 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0003-7995-1334ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond the Code: The Value of Practicing and Evaluating Technical Debt Management
abstract
Background: Effective management of Technical Debt (TD) is essential for improving software quality and sustainability. However, there is a lack of evidence about the benefits of TD management practices.
Mili Orucevic, Maren Maritsdatter Kruke, Antonio Martini 0001
TechDebt@ICSE1
2026 QualiTagger: automating software quality categorization in issue trackers
abstract
Abstract Managing software quality is crucial for maintainable software systems, yet understanding how quality attributes are discussed within the informal, noisy context of issue trackers remains a significant challenge. Current automatic approaches for categorizing quality concerns often falter, as they are typically validated on small domain-specific datasets and are ill-equipped to parse the conversational language of developer discourse. This hinders effective prioritization and technical debt management. This paper introduces QualiTagger, an automated approach for classifying seven distinct software quality attributes from issue tracker text, and QualiDataSet, a novel, curated dataset of over 700,000 labeled GitHub issues that underpins this work. We demonstrate that an ensemble of specialized binary classifiers, built upon the DistilRoBERTa architecture, significantly outperforms a single multiclass model and shows superior or comparable efficacy to a general-purpose Large Language Model (GPT-4o) for this task. The model’s real-world applicability is further validated through an industrial case study at Visma focusing on security issues) and a user study with software engineering students. Our evaluation confirms that QualiTagger achieves high classification accuracy and, crucially, generalizes effectively to previously unseen (Out-of-Distribution) projects, a key indicator of its practical utility. By providing both a robust classification tool and a large-scale public dataset, this research enables a more nuanced, data-driven understanding of how software quality is managed in practice, offering valuable insights for project management and future empirical software engineering research
Karthik Shivashankar, Rafael Capilla, Maren Maritsdatter Kruke, Mili Orucevic, Antonio Martini 0001
Softw. Qual. J.4
2025 Combining Insights from Multiple Tools to Manage Technical Debt in Industrial C# Projects
abstract
Technical Debt (TD) is a critical challenge in software development, leading to increased maintenance costs and reduced software quality over time. While considerable research has focused on identifying and managing TD in Java projects, studies on. NET (C#) projects remain limited. Additionally, existing approaches often rely on a single tool for TD detection, overlooking the benefits of combining multiple tools. In this paper, we analyze the effectiveness of Arcan, CodeScene, Designite, and DV8 on four industrial C#. NET 8 software products to address these research gaps. To validate and enrich our findings, we conducted online seminars and interviews with developers, architects, and managers involved in these projects, gathering practitioner insights on TD relevance and tool effectiveness. By leveraging complementary tools and practitioner feedback, we uncover different types of TD, including code-level, design, architectural, and knowledge debt. Our findings highlight each tool's strengths and limitations and demonstrate how integrating their outputs with expert input provides a more comprehensive and actionable TD assessment. Based on these insights, we propose a conceptual model for prioritizing and managing TD, offering guidance for practitioners.
Simeon Tverdal, Phu Hong Nguyen, Arda Goknil, Antonio Martini 0001, Merve Astekin, Mili Orucevic, Maren Maritsdatter Kruke, Håvard Stranden
ICSME6
2025 DebtQuest: Discover Technical Debt Management Issues with Survey Visualization
abstract
While Technical Debt (TD) manifest in systems as code-related symptoms, it is often caused by management-related issues. However, existing TD visualization tools are primarily code-based, and limited in addressing organizational factors. In this context, we introduce DebtQuest: an interactive, survey-based visualization tool designed to help assess an organization’s TD management practices. DebtQuest enables stakeholders to explore insights, pinpoint issues, perform benchmarking, and compare changes across time. Complementing the capabilities of code-based tools to discover TD symptoms, DebtQuest supports the discovery of underlying organizational causes. DebtQuest’s design is grounded in task analysis and a targeted review of relevant visualization techniques. It was refined through real-world use and evaluated by industry stakeholders. The contribution of this work is an accessible and flexible visualization design that extends recent research in TD management, revealing organizational insights that are difficult to discover using code-based tools.
Marius Irgens, Mili Orucevic, Mads Boye, Jan Henrik Gundelsby, Antonio Martini 0001
VISSOFT2
2025 BEACon-TD: Classifying Technical Debt and its types across diverse software projects issues using transformers
abstract
Technical Debt (TD) identification in software projects issues is crucial for maintaining code quality, reducing long-term maintenance costs, and improving overall project health. This study advances TD identification in issues tracker using transformer-based models, addressing the critical need for accurate and efficient TD identification in large-scale software development. Our methodology employs multiple binary classifiers for TD and its type, combined through ensemble learning , to enhance accuracy and robustness in detecting various forms of TD. We train and evaluate these models on a comprehensive dataset from GitHub Archive Issues (2015–2024), supplemented with industrial data validation. We demonstrate that in-project fine-tuned transformer models significantly outperform task-specific fine-tuned models in TD classification, highlighting the importance of project-specific context in accurate TD identification. Our research also reveals the superiority of specialized binary classifiers over multi-class models for TD and its type identification, enabling more targeted debt resolution strategies. A comparative analysis shows that the smaller DistilRoBERTa model is more effective than larger language models like GPTs for TD classification tasks , especially after fine-tuning, offering insights into efficient model selection for specific TD detection tasks. The study also assesses generalization capabilities using metrics such as MCC, AUC ROC, Recall, and F1 score, focusing on model effectiveness, fine-tuning impact, and relative performance . By validating our approach on out-of-distribution and real-world industrial datasets, we ensure practical applicability, addressing the diverse nature of software projects. This research significantly enhances TD detection and offers a more nuanced understanding of TD types, contributing to improved software maintenance strategies in both academic and industrial settings. The release of our curated dataset aims to stimulate further advancements in TD classification research , ultimately enhancing software project outcomes and development practices by enabling early TD identification and management.
Karthik Shivashankar, Mili Orucevic, Maren Maritsdatter Kruke, Antonio Martini 0001
J. Syst. Softw.2