Subhajit Datta

dblp:01/4987 · DBLP profile ↗
← Back
19ranked-venue papers
12as first author
6since 2021 · last 2026
0000-0001-9161-7951ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2026 Does developer familiarity hasten bug resolution? A causal inference perspective
Reshma Roychoudhuri, Subhajit Datta, Subhashis Majumder
Empir. Softw. Eng.3
2022 Litmus Test for Linus' Law: A Structural Equation Modeling Based Approach
abstract
“Many eyeballs make all bugs shallow” - referred to as Linus’ Law was framed by Raymond. The more is the number of developers working together on similar bugs for their resolution, the more easily and quickly will they get resolved. In this paper, we will be analyzing an open source dataset of 1000+ Android bugs, owned by 70 developers. Our results indicate that, for an arbitrary developer, the stronger is the developer network with other developers working in a similar environment, the more likely it is that they come across the same bugs during the development time. If they get access to the solution of the known bugs in the early phase of the development, then less time will be required for bug resolution and hence, in turn the quality of the software can be enhanced at a faster pace. We have done SEM analysis using Lavaan and have provided significant statistical evidence in support of our results.
Reshmi Maulik, Subhajit Datta, Subhashis Majumder
EASE2
2021 Degree doesn't Matter: Identifying the Drivers of Interaction in Software Development Ecosystems
abstract
Large scale software development ecosystems represent one of the most complex human enterprises. In such settings, developers are embedded in a web of shared concerns, responsibilities, and objectives at individual and collective levels. A deep understanding of the factors that influence developers to connect with one another is crucial in appreciating the challenges of such ecosystems as well as formulating strategies to overcome those challenges. We use real world data from multiple software development ecosystems to construct developer interaction networks and examine the mechanisms of such network formation using statistical models to identify developer attributes that have maximal influence on whether and how developers connect with one another. Our results challenge the conventional wisdom on the importance of particular developer attributes in their interaction practices, and offer useful insights for individual developers, project managers, and organizational decision-makers.
Ishita Bardhan, Subhajit Datta, Subhashis Majumder
APSEC2
2021 Links do Matter: Understanding the Drivers of Developer Interactions in Software Ecosystems
abstract
Studies of collaborating individuals engaged in collective enterprises usually focus on the individuals, rather than the links supporting their interaction. Accordingly, large scale software development ecosystems have also been examined primarily in terms of developer engagement. We posit that communication links between developers play a central role in the sustenance and effectiveness of such ecosystems. In this paper, we investigate whether and how developer attributes relate to the importance of the communication channels between them. We present a technique using 2nd order Markov models to extract features of interest of the links and apply the technique on data from a real-world project. Our statistical models - developed on records involving 900+ software developers, exchanging 20,000+ comments, across 500 units of work - offer surprising insights on factors associated with link importance, even after controlling for known effects. These results inform a deeper appreciation of the importance of links in large scale software development along with a number of practical implications.
Subhajit Datta, Amrita Bhattacharjee, Subhashis Majumder
ICSME1
2021 Clustering, Separation, and Connection: A Tale of Three Characteristics
abstract
In large and complex software development ecosystems, developers collaborate in multiple dimensions. How characteristics of such collaboration vary over time can offer insights about the dynamics of large scale software development. In this paper we have constructed networks of developers who co-comment on and co-change units of code, and analysed the patterns of variation of clustering, connection, and separation between such developers over time, using development data from a large open source system. Though clustering, connection, and separation essentially represent facets of the same collaboration activities, we found them exhibiting distinct time-varying characteristics. The variation of clustering indicate that developers congregate more closely towards the beginning of the project when the architecture of the system is not yet stable and they need to reach out to one another to fulfil their collective responsibilities. However, separation between developers shows a quick rise and then a saturation around a particular value. Developer connection continues rise throughout the observation period. Using time-series analysis, our results allow us to derive insights on the evolutionary trends in large scale software development and can inform the tuning of tools and process towards effective team assembly and governance.
Subhajit Datta, Aniruddha Mysore, Haziqshah Wira, Santonu Sarkar
ICSME1
2021 Understanding the relation between repeat developer interactions and bug resolution times in large open source ecosystems: A multisystem study
abstract
Abstract Large‐scale software systems are being increasingly built by distributed teams of developers who interact across geographies and time zones. Ensuring smooth knowledge transfer and the percolation of skills within and across such teams remain key challenges for organizations. Towards addressing this challenge, organizations often grapple with questions around whether and how repeat collaborations between members of a team relate to outcomes of important activities. In the context of this paper, the word ‘repeat interaction’ does not imply a greater number of interactions; it refers to repeat interaction between a pair of developers who have collaborated before. In this paper, we empirically examine such a question using real‐world data from three diverse development ecosystems, collectively involving 400,000+ units of work and 600,000+ comments exchanged between numerous developers. Our statistical models consistently establish a counter‐intuitive relation between repeat developer interaction and bug resolution times. Our experimental results show that more instances of repeat developer interactions over bug fixing are associated with more time taken for the bugs to be fixed. Given the expanse and variety of the underlying data, our results offer an unexpected set of insights on a key dynamic of collaboration in software development ecosystems. We discuss how these insights can influence the practice of large‐scale software development at individual, team and organizational levels.
Subhajit Datta, Reshma Roychoudhuri, Subhashis Majumder
J. Softw. Evol. Process.1
2019 Influence, Information and Team Outcomes in Large Scale Software Development
abstract
It is widely perceived that the egalitarian ecosystems of large scale open source software development foster effective team outcomes. In this study, we question this conventional wisdom by examining whether and how the centralization of information and influence in a software development team relate to the quality of the team's work products. Analyzing data from more than a hundred real world projects that include development activities over close to a decade, involving 2000+ developers, who collectively resolve more than two hundred thousand defects through discussions covering more than six hundred thousand comments, we arrive at statistically significant evidence indicating that concentration of information and influence in the developer communication networks of the projects are associated with the quality of a team's work products, even after controlling for various factors related to levels of developer engagement. Our results suggest that merely facilitating easy interaction between team members may not be sufficient to enhance team outcomes. The design of efficient collaborative development environments, and devising tools and processes for team assembly and governance can be informed by our results.
Subhajit Datta
APSEC1
2018 How does developer interaction relate to software quality? an examination of product development data
Subhajit Datta
Empir. Softw. Eng.1
2017 On negative results when using sentiment analysis tools for software engineering research
abstract
Recent years have seen an increasing attention to social aspects of software engineering, including studies of emotions and sentiments experienced and expressed by the software developers. Most of these studies reuse existing sentiment analysis tools such as SentiStrength and NLTK. However, these tools have been trained on product reviews and movie reviews and, therefore, their results might not be applicable in the software engineering domain. In this paper we study whether the sentiment analysis tools agree with the sentiment recognized by human evaluators (as reported in an earlier study) as well as with each other. Furthermore, we evaluate the impact of the choice of a sentiment analysis tool on software engineering studies by conducting a simple study of differences in issue resolution times for positive, negative and neutral texts. We repeat the study for seven datasets (issue trackers and Stack Overflow questions) and different sentiment analysis tools and observe that the disagreement between the tools can lead to diverging conclusions. Finally, we perform two replications of previously published studies and observe that the results of those studies cannot be confirmed when a different sentiment analysis tool is used.
Robbert Jongeling, Proshanta Sarkar, Subhajit Datta, Alexander Serebrenik
Empir. Softw. Eng.3
2017 The Habits of Highly Effective Researchers: An Empirical Study
abstract
Interest in the habits of influential individuals cuts across domains. As researchers, we are intrigued why few attain significant eminence in their fields, whereas many operate in obscurity. An empirical examination of this question has been made possible by the recent availability of large scale publication data. In this paper, we use information from the AMiner Paper Citation and Author Collaboration Networks to discern factors that relate to the impact of influential researchers across five domains in the computing discipline. We propose and apply a novel algorithm to identify influential vertices in co-authorship networks built from total corpora of 1,00,000+ papers and 72,000+ authors over a span of more than 50 years. The results from our study indicate that the impact of these influential researchers relate to a variety of factors. Surprisingly, we find evidence across the domains that higher impact is associated with lower levels of collaboration, and authority.
Subhajit Datta, Partha Basuchowdhuri, Surajit Acharya, Subhashis Majumder
IEEE Trans. Big Data1
2016 How Long Will This Live? Discovering the Lifespans of Software Engineering Ideas
abstract
We all want to be associated with long lasting ideas; as originators, or at least, expositors. For a tyro researcher or a seasoned veteran, knowing how long an idea will remain interesting in the community is critical in choosing and pursuing research threads. In the physical sciences, the notion of half-life is often evoked to quantify decaying intensity. In this paper, we study a corpus of 19,000+ papers written by 21,000+ authors across 16 software engineering publication venues from 1975 to 2010, to empirically determine the half-life of software engineering research topics. In the absence of any consistent and well-accepted methodology for associating research topics to a publication, we have used natural language processing techniques to semi-automatically identify and associate a set of topics with a paper. We adapted measures of half-life already existing in the bibliometric context for our study, and also defined a new measure based on publication and citation counts. We find evidence that some of the identified research topics show a mean half-life of close to 15 years, and there are topics with sustaining interest in the community. We report the methodology of our study in this paper, as well as the implications and utility of our results.
Subhajit Datta, Santonu Sarkar, A. S. M. Sajeev
IEEE Trans. Big Data1
2015 The Importance of Being Isolated: An Empirical Study on Chromium Reviews
abstract
As large scale software development has become more collaborative, and software teams more globally distributed, several studies have explored how developer interaction influences software development outcomes. The emphasis so far has been largely on outcomes like defect count, the time to close modification requests etc. In the paper, we examine data from the Chromium project to understand how different aspects of developer discussion relate to the closure time of reviews. On the basis of analyzing reviews discussed by 2000+ developers, our results indicate that quicker closure of reviews owned by a developer relates to higher reception of information and insights from peers. However, we also find evidence that higher engagement in collaboration by a developer is associated with slower closure of the reviews she owns. Within the scope of our study, these results lead us to conclude that peer review of code may have a distinct dynamic that is facilitated by developers working in relative isolation.
Subhajit Datta, Devarshi Bhatt, Proshanta Sarkar, Santonu Sarkar
ESEM1
2015 Choosing your weapons: On sentiment analysis tools for software engineering research
abstract
Recent years have seen an increasing attention to social aspects of software engineering, including studies of emotions and sentiments experienced and expressed by the software developers. Most of these studies reuse existing sentiment analysis tools such as SentiStrength and NLTK. However, these tools have been trained on product reviews and movie reviews and, therefore, their results might not be applicable in the software engineering domain. In this paper we study whether the sentiment analysis tools agree with the sentiment recognized by human evaluators (as reported in an earlier study) as well as with each other. Furthermore, we evaluate the impact of the choice of a sentiment analysis tool on software engineering studies by conducting a simple study of differences in issue resolution times for positive, negative and neutral texts. We repeat the study for seven datasets (issue trackers and Stack Overflow questions) and different sentiment analysis tools and observe that the disagreement between the tools can lead to contradictory conclusions.
Robbert Jongeling, Subhajit Datta, Alexander Serebrenik
ICSME2
2014 Does latitude hurt while longitude kills? geographical and temporal separation in a large scale software development project
abstract
Distributed software development allows firms to leverage cost advantages and place work near centers of competency. This distribution comes at a cost -- distributed teams face challenges from differing cultures, skill levels, and a lack of shared working hours. In this paper we examine whether and how geographic and temporal separation in a large scale distributed software development influences developer interactions. We mine the work item trackers for a large commercial software project with a globally distributed development team. We examine both the time to respond and the propensity of individuals to respond and find that when taken together, geographic distance has little effect, while temporal separation has a significant negative impact on the time to respond. However, both have little impact on the social network of individuals in the organization. These results suggest that while temporally distributed teams do communicate, it is at a slower rate, and firms may wish to locate partner teams in similar time zones for maximal performance.
Patrick Wagstrom, Subhajit Datta
ICSE2
2014 How Many Eyeballs Does a Bug Need? An Empirical Validation of Linus' Law
Subhajit Datta, Proshanta Sarkar, Sutirtha Das, Sonu Sreshtha, Prasanth Lade, Subhashis Majumder
XP1
2013 Factors Influencing Research Contributions and Researcher Interactions in Software Engineering: An Empirical Study
abstract
Research into software engineering (SE) education is largely concentrated on teaching and learning issues in coursework programs. This paper, in contrast, provides a meta analysis of research publications in software engineering to help with research education in SE. Studying publication patterns in a discipline will assist research students and supervisors gain a deeper understanding of how successful research has occurred in the discipline. We present results from a large scale empirical study covering over three and a half decades of software engineering research publications. We identify how different factors of publishing relate to the number of papers published as well as citations received for a researcher, and how the most successful researchers collaborate and co-cite one another. Our results show that authors with high publication rates do not concentrate on a few selected venues to publish, researchers with high publication rates behave differently from researchers of high citation rates (with the latter group co-authoring and citing their peers to a much lesser extent than the former), and collaborators citing each other's works is not a significant phenomenon in SE research.
Subhajit Datta, A. S. M. Sajeev, Santonu Sarkar, Nishant Kumar 0002
APSEC (1)1
2013 Introducing Programmers to Pair Programming: A Controlled Experiment
A. S. M. Sajeev, Subhajit Datta
XP2
2012 Talk versus work: characteristics of developer collaboration on the jazz platform
abstract
IBM's Jazz initiative offers a state-of-the-art collaborative development environment (CDE) facilitating developer interactions around interdependent units of work. In this paper, we analyze development data across two versions of a major IBM product developed on the Jazz platform, covering in total 19 months of development activity, including 17,000+ work items and 61,000+ comments made by more than 190 developers in 35 locations. By examining the relation between developer talk and work, we find evidence that developers maintain a reasonably high level of connectivity with peer developers with whom they share work dependencies, but the span of a developer's communication goes much beyond the known dependencies of his/her work items. Using multiple linear regression models, we find that the number of defects owned by a developer is impacted by the number of other developers (s)he is connected through talk, his/her interpersonal influence in the network of work dependencies, the number of work items (s)he comments on, and the number work items (s)he owns. These effects are maintained even after controlling for workload, role, work dependency, and connection related factors. We discuss the implications of our results for collaborative software development and project governance.
Subhajit Datta, Renuka Sindhgatta, Bikram Sengupta
OOPSLA1
2008 COMP-REF: A Technique to Guide the Delegation of Responsibilities to Components in Software Systems
Subhajit Datta, Robert A. van Engelen
FASE1