Aiswarya Raj Munappy

dblp:253/6366 · also Aiswarya Munappy, Aiswarya Raj, M. Aiswarya Raj · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
4since 2021 · last 2023
0000-0002-1333-5825ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 8 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Maturity Assessment Model for Industrial Data Pipelines
abstract
Data pipelines can be defined as a complex chain of interconnected activities that starts with a data source and ends in a data sink. They can process data in multiple formats from various data sources with minimal human intervention, speed up data life cycle operations, and enhance productivity in data-driven organizations. As a result, companies place a high value on strengthening the maturity of their data pipelines. The available literature, on the other hand, is significantly insufficient in terms of providing a comprehensive roadmap to guide companies in assessing the maturity of their data pipelines. Therefore, this case study focuses on developing a data pipeline maturity assessment model that can evaluate the maturity of data pipelines in a staged manner from maturity level 1 to maturity level 5. We conducted empirical research in order to develop the maturity assessment model on the basis of five different determinants to address the specific needs of each data pipeline maturity level. Accordingly, it aims to support organizations in assessing their current data pipeline maturity, determining challenges at each stage, and preparing an extensive roadmap and suggestions for data pipeline maturity improvement. In future work, we plan to employ the maturity model in different companies as a case study to evaluate its applicability and usefulness.
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson
APSEC1
2022 Customer Support In The Era of Continuous Deployment: A Software-Intensive Embedded Systems Case Study
abstract
Supporting customers after they acquire the prod-uct is essential for companies producing and selling software-intensive embedded systems products. Generally, customer sup-port is the first interaction point between the product users and the product vendor. Customer support is often engaged with answering customers' questions, troubleshooting, fault identification, and fixing product faults. While continuous deployment advocates for closer cooperation between the ones operating the software and the ones developing it, the means of such collaboration in general and the role of customer support, in particular, has not been addressed in the context of software-intensive embedded systems. Therefore, to better understand the impact that continuous deployment has on customer support and the role customer support should play in this context, we conducted a case study at a multinational company developing and selling telecommunications networks infrastructure. We focused on the 4th and 5th Generation (4G and 5G) Radio Access Networks (RAN) products, which can be considered a high volume product as they cover more than 80% of the world's population. Our study reveals that customer support needs to transition from a transaction-based and passive function triggered by customer support requests, to take an active role characterized by being proactive and preemptive to cope with the shorter operational time of a software version introduced by continuous deployment. In addition, customer support plays an essential role in making the feedback actionable by aggregating and consolidating feedback data to the R&D organization.
Anas Dakkak, Aiswarya Raj Munappy, Jan Bosch, Helena Olsson
COMPSAC2
2022 Data management for production quality deep learning models: Challenges and solutions
abstract
Deep learning (DL) based software systems are difficult to develop and maintain in industrial settings due to several challenges. Data management is one of the most prominent challenges which complicates DL in industrial deployments. DL models are data-hungry and require high-quality data. Therefore, the volume, variety, velocity, and quality of data cannot be compromised. This study aims to explore the data management challenges encountered by practitioners developing systems with DL components, identify the potential solutions from the literature and validate the solutions through a multiple case study. We identified 20 data management challenges experienced by DL practitioners through a multiple interpretive case study. Further, we identified 48 articles through a systematic literature review that discuss the solutions for the data management challenges. With the second round of multiple case study, we show that many of these solutions have limitations and are not used in practice due to a combination of four factors: high cost, lack of skill-set and infrastructure, inability to solve the problem completely, and incompatibility with certain DL use cases. Thus, data management for data-intensive DL models in production is complicated. Although the DL technology has achieved very promising results, there is still a significant need for further research in the field of data management to build high-quality datasets and streams that can be used for building production-ready DL systems. Furthermore, we have classified the data management challenges into four categories based on the availability of the solutions.
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson, Anders Arpteg, Björn Brinne
J. Syst. Softw.1
2021 On the Impact of ML use cases on Industrial Data Pipelines
abstract
The impact of the Artificial Intelligence revolution is undoubtedly substantial in our society, life, firms, and employment. With data being a critical element, organizations are working towards obtaining high-quality data to train their AI models. Although data, data management, and data pipelines are part of industrial practice even before the introduction of ML models, the significance of data increased further with the advent of ML models, which force data pipeline developers to go beyond the traditional focus on data quality. The objective of this study is to analyze the impact of ML use cases on data pipelines. We assume that the data pipelines that serve ML models are given more importance compared to the conventional data pipelines. We report on a study that we conducted by observing software teams at three companies as they develop both conventional(Non-ML) data pipelines and data pipelines that serve ML-based applications. We study six data pipelines from three companies and categorize them based on their criticality and purpose. Further, we identify the determinants that can be used to compare the development and maintenance of these data pipelines. Finally, we map these factors in a two-dimensional space to illustrate their importance on a scale of low, moderate, and high.
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson, Anders Jansson 0002
APSEC1
2020 Towards Automated Detection of Data Pipeline Faults
abstract
Data pipelines play an important role throughout the data management process. It automates the steps ranging from data generation to data reception thereby reducing the human intervention. A failure or fault in a single step of a data pipeline has cascading effects that might result in hours of manual intervention and clean-up. Data pipeline failure due to faults at different stages of data pipelines is a common challenge that eventually leads to significant performance degradation of data-intensive systems. To ensure early detection of these faults and to increase the quality of the data products, continuous monitoring and fault detection mechanism should be included in the data pipeline. In this study, we have explored the need for incorporating automated fault detection mechanisms and mitigation strategies at different stages of the data pipeline. Further, we identified faults at different stages of the data pipeline and possible mitigation strategies that can be adopted for reducing the impact of data pipeline faults thereby improving the quality of data products. The idea of incorporating fault detection and mitigation strategies is validated by realizing a small part of the data pipeline using action research in the analytics team at a large software-intensive organization within the telecommunication domain.
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson, Tian J. Wang
APSEC1
2020 Modelling Data Pipelines
abstract
The following topics are dealt with: software development management; software quality; software engineering; software maintenance; project management; formal specification; public domain software; learning (artificial intelligence); software prototyping; software architecture.
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson, Tian J. Wang
SEAA1
2020 From Ad-Hoc Data Analytics to DataOps
abstract
The collection of high-quality data provides a key competitive advantage to companies in their decision-making process. It helps to understand customer behavior and enables the usage and deployment of new technologies based on machine learning. However, the process from collecting the data, to clean and process it to be used by data scientists and applications is often manual, non-optimized and error-prone. This increases the time that the data takes to deliver value for the business. To reduce this time companies are looking into automation and validation of the data processes. Data processes are the operational side of data analytic workflow.
Aiswarya Raj Munappy, David Issa Mattos, Jan Bosch, Helena Olsson, Anas Dakkak
ICSSP1
2020 Data Pipeline Management in Practice: Challenges and Opportunities
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson
PROFES1
2020 Large-scale machine learning systems in real-world industrial settings: A review of challenges and solutions
Lucy Ellen Lwakatare, Aiswarya Raj Munappy, Ivica Crnkovic, Jan Bosch, Helena Olsson
Inf. Softw. Technol.2
2019 Data Management Challenges for Deep Learning
abstract
Deep learning is one of the most exciting and fast-growing techniques in Artificial Intelligence. The unique capacity of deep learning models to automatically learn patterns from the data differentiates it from other machine learning techniques. Deep learning is responsible for a significant number of recent breakthroughs in AI. However, deep learning models are highly dependent on the underlying data. So, consistency, accuracy, and completeness of data is essential for a deep learning model. Thus, data management principles and practices need to be adopted throughout the development process of deep learning models. The objective of this study is to identify and categorise data management challenges faced by practitioners in different stages of end-to-end development. In this paper, a case study approach is employed to explore the data management issues faced by practitioners across various domains when they use real-world data for training and deploying deep learning models. Our case study is intended to provide valuable insights to the deep learning community as well as for data scientists to guide discussion and future research in applied deep learning with real-world data.
Aiswarya Raj Munappy, Jan Bosch, Helena Olsson, Anders Arpteg, Björn Brinne
SEAA1
2019 A Taxonomy of Software Engineering Challenges for Machine Learning Systems: An Empirical Investigation
abstract
Abstract Artificial intelligence enabled systems have been an inevitable part of everyday life. However, efficient software engineering principles and processes need to be considered and extended when developing AI- enabled systems. The objective of this study is to identify and classify software engineering challenges that are faced by different companies when developing software-intensive systems that incorporate machine learning components. Using case study approach, we explored the development of machine learning systems from six different companies across various domains and identified main software engineering challenges. The challenges are mapped into a proposed taxonomy that depicts the evolution of use of ML components in software-intensive system in industrial settings. Our study provides insights to software engineering community and research to guide discussions and future research into applied machine learning.
Lucy Ellen Lwakatare, Aiswarya Raj Munappy, Jan Bosch, Helena Olsson, Ivica Crnkovic
XP2