Karthik Shivashankar

dblp:37/11127 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
10since 2021 · last 2026
0009-0001-8508-2978ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TD-Suite: All Batteries Included Framework for Technical Debt Classification
abstract
Abstract In Agile software development, maintaining velocity requires the continuous management of Technical Debt (TD). However, the rapid iteration cycles inherent to Agile often obscure debt accumulation, making manual identification in issue trackers prohibitively expensive. To address this, we present TD-Suite, a comprehensive framework engineered to automate the classification of technical debt. It leverages state-of-the-art transformer models to analyze textual artifacts, such as developer discussions in issue reports, where subtle indicators of debt often lie hidden. TD-Suite provides a seamless end-to-end pipeline suitable for Agile ML Engineering, managing everything from initial data ingestion and rigorous preprocessing to model training, thorough evaluation, and final inference. It supports both binary classification (debt or no debt) and granular categorization—identifying code, architecture, design, or documentation debt—enabling Agile teams to formulate targeted refactoring strategies. To ensure robustness on real-world, imbalanced datasets, TD-Suite incorporates k-fold cross-validation, early stopping, and class weighting strategies. The framework explicitly integrates the tracking and reporting of carbon emissions associated with model training. Furthermore, it features a user-friendly Gradio web interface within a Docker container, simplifying integration into DevOps pipelines and democratizing access for practitioners without deep ML expertise.
Karthik Shivashankar, Antonio Martini 0001
XP1
2026 QualiTagger: automating software quality categorization in issue trackers
abstract
Abstract Managing software quality is crucial for maintainable software systems, yet understanding how quality attributes are discussed within the informal, noisy context of issue trackers remains a significant challenge. Current automatic approaches for categorizing quality concerns often falter, as they are typically validated on small domain-specific datasets and are ill-equipped to parse the conversational language of developer discourse. This hinders effective prioritization and technical debt management. This paper introduces QualiTagger, an automated approach for classifying seven distinct software quality attributes from issue tracker text, and QualiDataSet, a novel, curated dataset of over 700,000 labeled GitHub issues that underpins this work. We demonstrate that an ensemble of specialized binary classifiers, built upon the DistilRoBERTa architecture, significantly outperforms a single multiclass model and shows superior or comparable efficacy to a general-purpose Large Language Model (GPT-4o) for this task. The model’s real-world applicability is further validated through an industrial case study at Visma focusing on security issues) and a user study with software engineering students. Our evaluation confirms that QualiTagger achieves high classification accuracy and, crucially, generalizes effectively to previously unseen (Out-of-Distribution) projects, a key indicator of its practical utility. By providing both a robust classification tool and a large-scale public dataset, this research enables a more nuanced, data-driven understanding of how software quality is managed in practice, offering valuable insights for project management and future empirical software engineering research
Karthik Shivashankar, Rafael Capilla, Maren Maritsdatter Kruke, Mili Orucevic, Antonio Martini 0001
Softw. Qual. J.1
2025 MLScent: A Tool for Anti-Pattern Detection in ML Projects
abstract
Machine learning (ML) codebases face unprecedented challenges in maintaining code quality and sustainability as their complexity grows exponentially. While traditional code smell detection tools exist, they fail to address ML-specific issues that can significantly impact model performance, reproducibility, and maintainability. This paper introduces MLScent, a novel static analysis tool that leverages sophisticated Abstract Syntax Tree (AST) analysis to detect anti-patterns and code smells specific to ML projects. MLScent implements 76 distinct detectors across major ML frameworks including TensorFlow (13 detectors), PyTorch (12 detectors), Scikit-learn (9 detectors), and Hugging Face (10 detectors), along with data science libraries like Pandas and NumPy (8 detectors each). The tool's architecture also integrates general ML smell detection (16 detectors), and specialized analysis for data preprocessing and model training workflows. Our evaluation demonstrates MLScent's effectiveness through both quantitative classification metrics and qualitative assessment via user studies feedback with ML practitioners. Results show high accuracy in identifying framework-specific anti-patterns, data handling issues, and general ML code smells across real-world projects.
Karthik Shivashankar, Antonio Martini 0001
CAIN1
2025 PyExamine: A Comprehensive, Un-Opinionated Smell Detection Tool for Python
abstract
The growth of Python adoption across diverse domains has led to increasingly complex codebases, presenting challenges in maintaining code quality. While numerous tools attempt to address these challenges, they often fall short in providing comprehensive analysis capabilities or fail to consider Pythonspecific contexts. PyExamine addresses these critical limitations through an approach to code smell detection that operates across multiple levels of analysis. PyExamine architecture enables detailed examination of code quality through three distinct but interconnected layers: architectural patterns, structural relationships, and code-level implementations. This approach allows for the detection and analysis of 49 distinct metrics, providing developers with an understanding of their codebase’s health. The metrics span across all levels of code organization, from high-level architectural concerns to granular implementation details. Through evaluation on 7 diverse projects, PyExamine achieved detection accuracy rates: $91.4 \%$ for code-level smells, $89.3 \%$ for structural smells, and $80.6 \%$ for architectural smells. These results were further validated through extensive user feedback and expert evaluations, confirming PyExamine’s capability to identify potential issues across all levels of code organization with high recall accuracy. In additional to this, we have also used PyExamine to analysis the prevalence of different type of smells, across 183 diverse Python projects ranging from small utilities to large-scale enterprise applications. PyExamine’s distinctive combination of comprehensive analysis, Python-specific detection, and high customizability makes it a valuable asset for both individual developers and large teams seeking to enhance their code quality practices.
Karthik Shivashankar, Antonio Martini 0001
MSR1
2025 Enhancing Python Code Maintainability Through Large Language Model-Based Approaches
Karthik Shivashankar, Antonio Martini 0001
PROFES1
2025 BEACon-TD: Classifying Technical Debt and its types across diverse software projects issues using transformers
abstract
Technical Debt (TD) identification in software projects issues is crucial for maintaining code quality, reducing long-term maintenance costs, and improving overall project health. This study advances TD identification in issues tracker using transformer-based models, addressing the critical need for accurate and efficient TD identification in large-scale software development. Our methodology employs multiple binary classifiers for TD and its type, combined through ensemble learning , to enhance accuracy and robustness in detecting various forms of TD. We train and evaluate these models on a comprehensive dataset from GitHub Archive Issues (2015–2024), supplemented with industrial data validation. We demonstrate that in-project fine-tuned transformer models significantly outperform task-specific fine-tuned models in TD classification, highlighting the importance of project-specific context in accurate TD identification. Our research also reveals the superiority of specialized binary classifiers over multi-class models for TD and its type identification, enabling more targeted debt resolution strategies. A comparative analysis shows that the smaller DistilRoBERTa model is more effective than larger language models like GPTs for TD classification tasks , especially after fine-tuning, offering insights into efficient model selection for specific TD detection tasks. The study also assesses generalization capabilities using metrics such as MCC, AUC ROC, Recall, and F1 score, focusing on model effectiveness, fine-tuning impact, and relative performance . By validating our approach on out-of-distribution and real-world industrial datasets, we ensure practical applicability, addressing the diverse nature of software projects. This research significantly enhances TD detection and offers a more nuanced understanding of TD types, contributing to improved software maintenance strategies in both academic and industrial settings. The release of our curated dataset aims to stimulate further advancements in TD classification research , ultimately enhancing software project outcomes and development practices by enabling early TD identification and management.
Karthik Shivashankar, Mili Orucevic, Maren Maritsdatter Kruke, Antonio Martini 0001
J. Syst. Softw.1
2025 Enhancing Task Prioritization in Software Development Issues Tracking System
abstract
ABSTRACT Modern software development faces a critical bottleneck in manually prioritizing the overwhelming volume of issues generated in platforms like Jira and GitHub. This labor‐intensive process leads to delays, increased costs, inconsistent handling, and developer burnout, worsened by the lack of standardized priority labels. This paper investigates the potential of automated issue priority classification using state‐of‐the‐art Transformer models to alleviate this burden. We evaluate the performance of models like BERT, DeBERTa, and ModernBERT, comparing them against general large language models (LLMs) such as GPT‐3.5, Qwen2.5‐3B and Llama‐3.2‐3B, using curated datasets derived from public Jira and GitHub repositories. Our research addresses the effectiveness of these models for their generalization capabilities on out‐of‐distribution projects, the impact of fine‐tuning, and performs a detailed performance comparison across different priority levels and model types. Results demonstrate that Transformer models, particularly ModernBERT, achieve high classification performance (e.g., accuracy > 81%), significantly outperforming the evaluated general LLMs (accuracy 75%) for this specific task. We find that binary classification is more effective than multilabel approaches, models generalize well to unseen projects, and performance is further enhanced by fine‐tuning. Key contributions include the provision of cleaned, labeled datasets and a comprehensive evaluation confirming the viability and benefits of using specialized Transformer models for automated issue priority suggestion, offering a path to improved efficiency and resource allocation in software development workflows.
Karthik Shivashankar, Kristian Marison Haugerud, Antonio Martini 0001
J. Softw. Evol. Process.1
2024 Towards Enhancing Task Prioritization in Software Development Through Transformer-Based Issues Classification
Kristian Marison Haugerud, Karthik Shivashankar, Antonio Martini 0001
PROFES2
2023 Technical Debt Classification in Issue Trackers using Natural Language Processing based on Transformers
abstract
Background: Technical Debt (TD) needs to be controlled and tracked during software development. Support to automatically track TD in issue trackers is limited. Aim: We explore the usage of a large dataset of developer-labeled TD issues in combination with cutting-edge Natural Language Processing (NLP) approaches to automatically classify TD in issue trackers. Method: We mine and analyze more than 160GB of textual data from GitHub projects, collecting over 55,600 TD issues and consolidating them into a large dataset (GTD dataset). We use such datasets to train and test Transformer ML models. Then we test the model’s generalization ability by testing them on six unseen projects. Finally, we re-train the models including part of the TD issues from the target project to test their adaptability. Results and conclusion: (i) We create and release the GTD dataset, a comprehensive dataset including TD issues from 6,401 public repositories with various contexts; (ii) By training Transformers using the GTD dataset, we achieve performance metrics that are promising; (iii) Our results are a significant step forward towards supporting the automatic classification of TD in issue trackers, especially when the models are adapted to the context of unseen projects after fine-tuning.
Daniel Skryseth, Karthik Shivashankar, Ildikó Pilán, Antonio Martini 0001
TechDebt@ICSE2
2022 Maintainability Challenges in ML: A Systematic Literature Review
abstract
Background: As Machine Learning (ML) advances rapidly in many fields, it is being adopted by academics and businesses alike. However, ML has a number of different challenges in terms of maintenance not found in traditional software projects. Identifying what causes these maintainability challenges can help mitigate them early and continue delivering value in the long run without degrading ML performance. Aim: This study aims to identify and synthesise the maintainability challenges in different stages of the ML workflow and understand how these stages are interdependent and impact each other’s maintainability. Method: Using a systematic literature review, we screened more than 13000 papers, then selected and qualitatively analysed 56 of them. Results: (i) a catalogue of maintainability challenges in different stages of Data Engineering, Model Engineering workflows and the current challenges when building ML systems are discussed; (ii) a map of 13 maintainability challenges to different interdependent stages of ML that impact the overall workflow; (iii) Provided insights to developers of ML tools and researchers. Conclusions: In this study, practitioners and organisations will learn about maintainability challenges and their impact at different stages of ML workflow. This will enable them to avoid pitfalls and help to build a maintainable ML system. The implications and challenges will also serve as a basis for future research to strengthen our understanding of the ML system’s maintainability.
Karthik Shivashankar, Antonio Martini 0001
SEAA1