Renuka Sindhgatta

dblp:s/RenukaSindhgatta · also S. R. Renuka · DBLP profile ↗
← Back
45ranked-venue papers
14as first author
16since 2021 · last 2026
0000-0001-7533-533XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 12 first-author · 3 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
abstract
Recent advancements in Large Language Models (LLMs) has lead to the development of agents capable of complex reasoning and interaction with external tools. In enterprise contexts, the effective use of such tools that are often enabled by application programming interfaces (APIs) is hindered by poor documentation, complex input or output schema, and large number of operations. These challenges make tool selection difficult and reduce the accuracy of payload formation upto 25%. We propose ACE, an automated tool creation and enrichment framework that transforms enterprise APIs into LLM-compatible tools. ACE (i) generates enriched tool specifications with parameter descriptions and examples to improve selection and invocation accuracy, and (ii) incorporates a dynamic shortlisting mechanism that filters relevant tools at runtime, reducing prompt complexity while maintaining scalability. We validate our framework on both proprietary and open-source APIs and demonstrate its integration with agentic frameworks. To the best of our knowledge, ACE is the first end-to-end framework that automates the creation, enrichment, and dynamic selection of enterprise API tools for LLM agents.
Prerna Agarwal, Soujanya Soni, Rohith D. Vallam, Renuka Sindhgatta, Sameep Mehta
AAAI5
2026 DFAgent: From Natural Language Data Interactions to Reusable Agent-Ready Tools
abstract
We present DataFoundry Agent (DFAgent), a system that forges reusable, agent-ready tools from interactive data exploration, quality, and remediation tasks. Users engage with data through natural-language prompts for operations that include inspection, transformation, and visualization. These interactions automatically generate executable code snippets that are logged. From these snippets, DFAgent acts as a foundry, synthesizing a governed catalog of enriched tools exposed via the Model Context Protocol (MCP). In this way, user-derived logic for all data operations is transformed into standardized, composable tools without reimplementation. We demonstrate how diverse interactions accumulate into a reusable toolset, highlighting a paradigm that unifies natural language interaction, executable code generation, and tool foundry processes for agentic data systems.
Neelamadhav Gantayat, Renuka Sindhgatta, Sambit Ghosh, Sameep Mehta, Soujanya Soni
AAAI2
2026 ToolSmith: A Multi-Agent Framework for Enterprise Tool Creation
abstract
Although LLMs can generate tools for generic domains and tasks, they struggle with enterprise-related domains that involve proprietary APIs and data schemas. We present ToolSmith, a framework for autonomously generating and validating agent-compatible tools. Given an API specification and a Tool Specification Requirement (TSR), ToolSmith produces a tool function and verifies it through a closed-loop process: it creates natural language (NL) tests and executes the tool in a secure agent sandbox for validation. For state-changing tools, ToolSmith confirms outcomes by querying the API with parameters derived from the NL tests. If the tool fails to produce the desired output, ToolSmith generates diagnostic feedback to iteratively regenerate it. By ensuring both functional correctness and agent compatibility, ToolSmith enables reliable automation of enterprise workflows.
Purna Chandra Sekhar Vakudavathu, Kushal Mukherjee, Jayachandu Bandlamudi, Renuka Sindhgatta, Sameep Mehta
AAAI4
2025 Developing guidelines for functionally-grounded evaluation of explainable artificial intelligence using tabular data
abstract
Explainable Artificial Intelligence (XAI) techniques are used to provide transparency to complex, opaque predictive models. However, these techniques are often designed for image and text data, and it is unclear how fit-for-purpose they are when applied to tabular data. As XAI techniques are rarely evaluated in the context of tabular data, the applicability of existing evaluation criteria and methods are also unclear and needs re-examination. For example, some works suggest that evaluation methods may unduly influence the evaluation results when using tabular data. This lack of clarity on evaluation procedures can lead to reduced transparency and ineffective use of XAI techniques in real world settings. In this study, we examine literature on XAI evaluation to derive guidelines on functionally-grounded assessment of local, post hoc XAI techniques. We identify 20 evaluation criteria and associated evaluation methods, and derive guidelines on when and how each criterion should be evaluated. We also identify key research gaps to be addressed by future work. Our study contributes to the body of knowledge on XAI evaluation through in-depth examination of functionally-grounded XAI evaluation protocols, and has laid the groundwork for future research on XAI evaluation.
Mythreyi Velmurugan, Chun Ouyang 0001, Yue Xu 0001, Renuka Sindhgatta, Bemali Wickramanayake, Catarina Moreira
Eng. Appl. Artif. Intell.4
2024 Multi-Stage Prompting for Next Best Agent Recommendations in Adaptive Workflows
abstract
Traditional business processes such as loan processing, order processing, or procurement have a series of steps that are pre-defined at design and executed by enterprise systems. Recent advancements in new-age businesses, however, focus on having adaptive and ad-hoc processes by stitching together a set of functions or steps enabled through autonomous agents. Further, to enable business users to execute a flexible set of steps, there have been works on providing a conversational interface to interact and execute automation. Often, it is necessary to guide the user through the set of possible steps in the process (or workflow). Existing work on recommending the next agent to run relies on historical data. However, with changing workflows and new automation constantly getting added, it is important to provide recommendations without historical data. Additionally, hand-crafted recommendation rules do not scale. The adaptive workflow being a combination of structured and unstructured information, makes it harder to mine. Hence, in this work, we leverage Large Language Models (LLMs) to combine process knowledge with the meta-data of agents to discover NBAs specifically at cold-start. We propose a multi-stage approach that uses existing process knowledge and agent meta-data information to prompt LLM and recommend meaningful next best agent (NBA) based on user utterances.
Prerna Agarwal, Harshit Dave, Jayachandu Bandlamudi, Renuka Sindhgatta, Kushal Mukherjee
AAAI4
2024 Building Conversational Artifacts to Enable Digital Assistant for APIs and RPAs
abstract
In the realm of business automation, digital assistants/chatbots are emerging as the primary method for making automation software accessible to users in various business sectors. Access to automation primarily occurs through APIs and RPAs. To effectively convert APIs and RPAs into chatbots on a larger scale, it is crucial to establish an automated process for generating data and training models that can recognize user intentions, identify questions for conversational slot filling, and provide recommendations for subsequent actions. In this paper, we present a technique for enhancing and generating natural language conversational artifacts from API specifications using large language models (LLMs). The goal is to utilize LLMs in the "build" phase to assist humans in creating skills for digital assistants. As a result, the system doesn't need to rely on LLMs during conversations with business users, leading to efficient deployment. Experimental results highlight the effectiveness of our proposed approach. Our system is deployed in the IBM Watson Orchestrate product for general availability.
Jayachandu Bandlamudi, Kushal Mukherjee, Prerna Agarwal, Ritwik Chaudhuri, Rakesh Pimplikar, Sampath Dechu, Alex Straley, Anbumunee Ponniah, Renuka Sindhgatta
AAAI9
2024 AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data
abstract
The efficiency of business processes relies on business key performance indicators (Biz-KPIs), that can be negatively impacted by IT failures. Business and IT Observability (BizITObs) data fuses both Biz-KPIs and IT event channels together as multivariate time series data. Forecasting Biz-KPIs in advance can enhance efficiency and revenue through proactive corrective measures. However, BizITObs data generally exhibit both useful and noisy inter-channel interactions between Biz-KPIs and IT events that need to be effectively decoupled. This leads to suboptimal forecasting performance when existing multivariate forecasting models are employed. To address this, we introduce AutoMixer, a time-series Foundation Model (FM) approach, grounded on the novel technique of channel-compressed pretrain and finetune workflows. AutoMixer leverages an AutoEncoder for channel-compressed pretraining and integrates it with the advanced TSMixer model for multivariate time series forecasting. This fusion greatly enhances the potency of TSMixer for accurate forecasts and also generalizes well across several downstream tasks. Through detailed experiments and dashboard analytics, we show AutoMixer's capability to consistently improve the Biz-KPI's forecasting accuracy (by 11-15%) which directly translates to actionable business insights.
Santosh Palaskar, Vijay Ekambaram, Arindam Jati, Neelamadhav Gantayat, Avirup Saha, Seema Nagar, Nam H. Nguyen, Pankaj Dayama 0001, Renuka Sindhgatta, Prateeti Mohapatra, Jayant Kalagnanam, Nandyala Hemachandra, Narayan Rangaraj
AAAI9
2024 Sequential API Function Calling Using GraphQL Schema
abstract
Avirup Saha, Lakshmi Mandal, Balaji Ganesan, Sambit Ghosh, Renuka Sindhgatta, Carlos Eberhardt, Dan Debrunner, Sameep Mehta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Avirup Saha, Lakshmi Mandal, Balaji Ganesan, Sambit Ghosh, Renuka Sindhgatta, Carlos Eberhardt, Dan Debrunner, Sameep Mehta
EMNLP5
2024 LLM-powered GraphQL Generator for Data Retrieval
Balaji Ganesan, Sambit Ghosh, Nitin Gupta 0005, Manish Kesarwani, Sameep Mehta, Renuka Sindhgatta
IJCAI6
2024 Multi-Objective Evolutionary Search for Optimal Robotic Process Automation Architectures
abstract
Robotic Process Automation (RPA) design and implementation requires an architecture which facilitates the seamless transition between human agents, robotic agents, and intelligent agents to automate information acquisition tasks and decision-making tasks. Coordination of those agents must consider various factors, such as efficiency of a resource when completing tasks, the quality of completed complex tasks, and the cost of the used resources. This article proposes a novel approach for generating an optimal architecture based on distinct types of resources, including human agents, intelligent agents, and robotic agents. An optimal architecture is the optimal enactment of process instances executed by a combination of human and automation agents based on their characteristics. The architecture provides a set of resources and their characteristics that are tailored to meet multiple objectives for process execution. The proposed approach is validated through an empirical evaluation based on a real-world business process. An empirical evaluation demonstrates that, given equal computational time, our approach outperforms conventional constraint optimization ILOG CPLEX (Manual 1987).
Geeta Mahala, Renuka Sindhgatta, Khanh Hoa Dam, Aditya Ghose
IEEE Trans. Serv. Comput.2
2023 Towards Hybrid Automation by Bootstrapping Conversational Interfaces for IT Operation Tasks
abstract
Process automation has evolved from end-to-end automation of repetitive process branches to hybrid automation where bots perform some activities and humans serve other activities. In the context of knowledge-intensive processes such as IT operations, implementing hybrid automation is a natural choice where robots can perform certain mundane functions, with humans taking over the decision of when and which IT systems need to act. Recently, ChatOps, which refers to conversation-driven collaboration for IT operations, has rapidly accelerated efficiency by providing a cross-organization and cross-domain platform to resolve and manage issues as soon as possible. Hence, providing a natural language interface to bots is a logical progression to enable collaboration between humans and bots. This work presents a no-code approach to provide a conversational interface that enables human workers to collaborate with bots executing automation scripts. The bots identify the intent of users' requests and automatically orchestrate one or more relevant automation tasks to serve the request. We further detail our process of mining the conversations between humans and bots to monitor performance and identify the scope for improvement in service quality.
Jayachandu Bandlamudi, Kushal Mukherjee, Prerna Agarwal, Sampath Dechu, Siyu Huo, Vatche Isahagian, Vinod Muthusamy, Naveen Purushothaman, Renuka Sindhgatta
AAAI9
2023 Editorial: recent advances in process analytics
Paolo Ceravolo, Claudio Di Ciccio, Chiara Di Francescomarino, María Teresa Gómez-López, Fabrizio Maria Maggi, Renuka Sindhgatta
J. Intell. Inf. Syst.6
2022 Building interpretable models for business process prediction using shared and specialised attention mechanisms
Bemali Wickramanayake, Zhipeng He 0002, Chun Ouyang 0001, Catarina Moreira, Yue Xu 0001, Renuka Sindhgatta
Knowl. Based Syst.6
2021 DeepProcess: Supporting Business Process Execution Using a MANN-Based Recommender System
Muhammad Asjad Khan, Hung Le 0002, Kien Do, Truyen Tran 0001, Aditya Ghose, Khanh Hoa Dam, Renuka Sindhgatta
ICSOC7
2021 Evaluating Stability of Post-hoc Explanations for Business Process Predictions
Mythreyi Velmurugan, Chun Ouyang 0001, Catarina Moreira, Renuka Sindhgatta
ICSOC4
2021 LINDA-BN: An interpretable probabilistic approach for demystifying black-box predictive models
Catarina Moreira, Yu-Liang Chou, Mythreyi Velmurugan, Chun Ouyang 0001, Renuka Sindhgatta, Peter Bruza
Decis. Support Syst.5
2020 Exploring Interpretable Predictive Models for Business Processes
Renuka Sindhgatta, Catarina Moreira, Chun Ouyang 0001, Alistair Barros
BPM1
2020 Co-destruction Patterns in Crowdsourcing - Formal/Technical Paper
Reihaneh Bidar, Arthur H. M. ter Hofstede, Renuka Sindhgatta
CAiSE3
2020 Resource-Based Adaptive Robotic Process Automation - Formal/Technical Paper
Renuka Sindhgatta, Arthur H. M. ter Hofstede, Aditya Ghose
CAiSE1
2020 Designing Optimal Robotic Process Automation Architectures
Geeta Mahala, Renuka Sindhgatta, Khanh Hoa Dam, Aditya Ghose
ICSOC2
2020 Exploring Interpretability for Predictive Process Analytics
Renuka Sindhgatta, Chun Ouyang 0001, Catarina Moreira
ICSOC1
2019 Targeted Example Generation for Compilation Errors
abstract
We present TEGCER, an automated feedback tool for novice programmers. TEGCER uses supervised classification to match compilation errors in new code submissions with relevant pre-existing errors, submitted by other students before. The dense neural network used to perform this classification task is trained on 15000+ error-repair code examples. The proposed model yields a test set classification Pred@3 accuracy of 97.7% across 212 error category labels. Using this model as its base, TEGCER presents students with the closest relevant examples of solutions for their specific error on demand. A large scale (N>230) usability study shows that students who use TEGCER are able to resolve errors more than 25% faster on average than students being assisted by human tutors.
Umair Z. Ahmed, Renuka Sindhgatta, Nisheeth Srivastava, Amey Karkare
ASE2
2018 Balancing Human Efforts and Performance of Student Response Analyzer in Dialog-Based Tutors
Tejas I. Dhamecha, Smit Marvaniya, Swarnadeep Saha, Renuka Sindhgatta, Bikram Sengupta
AIED (1)4
2018 Sentence Level or Token Level Features for Automatic Short Answer Grading?: Use Both
Swarnadeep Saha, Tejas I. Dhamecha, Smit Marvaniya, Renuka Sindhgatta, Bikram Sengupta
AIED (1)4
2018 Creating Scoring Rubric from Representative Student Answers for Improved Short Answer Grading
abstract
Automatic short answer grading remains one of the key challenges of any dialog-based tutoring system due to the variability in the student answers. Typically, each question may have no or few expert authored exemplary answers which make it difficult to (1) generalize to all correct ways of answering the question, or (2) represent answers which are either partially correct or incorrect. In this paper, we propose an affinity propagation based clustering technique to obtain class-specific representative answers from the graded student answers. Our novelty lies in formulating the Scoring Rubric by incorporating class-specific representatives obtained after proposed clustering, selecting, and ranking of graded student answers. We experiment with baseline as well as stateof-the-art sentence-embedding based features to demonstrate the feature-agnostic utility of class-specific representative answers. Experimental evaluations on our large-scale industry dataset and a benchmarking dataset show that the Scoring Rubric significantly improves the classification performance of short answer grading.
Smit Marvaniya, Swarnadeep Saha, Tejas I. Dhamecha, Peter W. Foltz, Renuka Sindhgatta, Bikram Sengupta
CIKM5
2018 Augmenting Classrooms with AI for Personalized Education
abstract
Intelligent tutoring systems (ITS) have been a topic of great interest for about five decades. Over the years, ITS research has leveraged AI advancements, and has also helped push the boundaries of AI capabilities with grounded usage scenarios. Using ITSs along with classroom instruction to augment traditional teaching is a canonical example of how humans and machines can work together to solve problems that are otherwise overwhelming and non-scalable individually. The experiences of personalized learning created by (1) seamless orchestration of human decision-making at few critical points with (2) scalability of cognitive capabilities using AI systems can drive increased student engagement leading to improved learning outcomes. By considering two particular use-cases of early childhood learning and higher education, we discuss the challenges involved in designing these complex human-centric systems. These systems integrate technologies involving interactivity, dialog, automated question generation, and learning analytics.
Ravi Kokku, Sharad Sundararajan, Renuka Sindhgatta, Satya V. Nitta, Bikram Sengupta
ICASSP4
2018 Impact of Tutor Errors on Student Engagement in a Dialog Based Intelligent Tutoring System
Shazia Afzal, Vinay Shashidhar, Renuka Sindhgatta, Bikram Sengupta
ITS3
2017 Inferring Frequently Asked Questions from Student Question Answering Forums
Renuka Sindhgatta, Smit Marvaniya, Tejas I. Dhamecha, Bikram Sengupta
EDM1
2017 Intelligent Math Tutor: Problem-Based Approach to Create Cognizance
abstract
Mathematical word problems (or story problems) allow students to apply their mathematical problem solving ability to other subjects and real-world situations. Word problems build higher-order thinking, critical problem-solving, and reasoning skills. Generally solving a word problem is associated with mathematical modeling of a real word situation or a concept of another subject which is embedded in the problem. Manually creating word problems require knowledge of other topics a student is learning in parallel. Besides this, modeling mathematics with some other dissociated concept is a time-consuming and labor-intensive task. Due to lack of this integrated knowledge of other topics being taught, the substantive breadth of word problems is often very narrow and is limited to very few concepts. To address this limitation, we built a tool called Intelligent Math Tutor (IMT), which automatically generates mathematical word problems such that teachings from other subjects from a given curriculum can also be incorporated. Our tool thus widens the scope of word problems and uses this problem-solving based approach to indirectly create cognizance in its students. To the best of our knowledge, our tool is the first of its kind tool which explicitly blends knowledge from multiple dissociated subjects and uses it to enhance the cognizance of its learners.
Monika Gupta 0002, Neelamadhav Gantayat, Renuka Sindhgatta
L@S3
2016 Context-Aware Analysis of Past Process Executions to Aid Resource Allocation Decisions
Renuka Sindhgatta, Aditya Ghose, Khanh Hoa Dam
CAiSE1
2016 Context-Aware Recommendation of Task Allocations in Service Systems
Renuka Sindhgatta, Aditya Ghose, Khanh Hoa Dam
ICSOC1
2015 Analyzing Resource Behavior to Aid Task Assignment in Service Systems
Renuka Sindhgatta, Aditya Ghose, Gargi Dasgupta
ICSOC1
2014 Analysis of Operational Data for Expertise Aware Staffing
Renuka Sindhgatta, Gargi Dasgupta, Aditya Ghose
BPM1
2014 An Extended Agent Based Model for Service Delivery Optimization
Mohammadreza Mohagheghian, Renuka Sindhgatta, Aditya Ghose
PRIMA2
2013 Accelerating Collaboration in Task Assignment Using a Socially Enhanced Resource Model
Shivali Agarwal, Renuka Sindhgatta, Juhnyoung Lee
BPM3
2013 Does One-Size-Fit-All Suffice for Service Delivery Clients?
Shivali Agarwal, Renuka Sindhgatta, Gargi Dasgupta
ICSOC2
2013 Behavioral Analysis of Service Delivery Models
Gargi Dasgupta, Renuka Sindhgatta, Shivali Agarwal
ICSOC2
2013 Interleaving Execution into Model Driven Service Design
abstract
Business Process driven Service Oriented Architecture (SOA) allows for designing services that execute (or realize) atomic tasks of business processes. When developing and designing SOA applications using Model Driven Development (MDD), business processes and services are represented using specifications such as Unified Modeling Language (UML). UML based models help in specifying structural properties of services and the behavioral properties of a service composition, executing a business process. In a typical design scenario, architects develop process models (comprising of tasks) and service design independently and establish associations between tasks and the designed services. Service compositions representing business processes are verified with the designed services. Verification at design time is a manual activity and relies on the static information captured in these UML models. In this paper we describe a model-based approach that enables executing service compositions at design time. The process model and service model is transformed to an executable UML model and UML Action language (UAL) code fragments are automatically generated. UAL enables unambiguous specification of the behavior of service operations and their compositions. The generated executable model enables a precise behavior analysis of the process realization. We demonstrate the use of this approach by reporting on a reference model with 17 business processes and 127 business tasks realized using 13 Service interfaces.
Renuka Sindhgatta
ICWS1
2012 SmartDispatch: enabling efficient ticket dispatch in an IT service environment
abstract
In an IT service delivery environment, the speedy dispatch of a ticket to the correct resolution group is the crucial first step in the problem resolution process. The size and complexity of such environments make the dispatch decision challenging, and incorrect routing by a human dispatcher can lead to significant delays that degrade customer satisfaction, and also have adverse financial implications for both the customer and the IT vendor. In this paper, we present SmartDispatch, a learning-based tool that seeks to automate the process of ticket dispatch while maintaining high accuracy levels. SmartDispatch comes with two classification approaches - the well-known SVM method, and a discriminative term-based approach that we designed to address some of the issues in SVM classification that were empirically observed. Using a combination of these approaches, SmartDispatch is able to automate the dispatch of a ticket to the correct resolution group for a large share of the tickets, while for the rest, it is able to suggest a short list of 3-5 groups that contain the correct resolution group with a high probability. Empirical evaluation of SmartDispatch on data from 3 large service engagement projects in IBM demonstrate the efficacy and practical utility of the approach.
Shivali Agarwal, Renuka Sindhgatta, Bikram Sengupta
KDD2
2012 Talk versus work: characteristics of developer collaboration on the jazz platform
abstract
IBM's Jazz initiative offers a state-of-the-art collaborative development environment (CDE) facilitating developer interactions around interdependent units of work. In this paper, we analyze development data across two versions of a major IBM product developed on the Jazz platform, covering in total 19 months of development activity, including 17,000+ work items and 61,000+ comments made by more than 190 developers in 35 locations. By examining the relation between developer talk and work, we find evidence that developers maintain a reasonably high level of connectivity with peer developers with whom they share work dependencies, but the span of a developer's communication goes much beyond the known dependencies of his/her work items. Using multiple linear regression models, we find that the number of defects owned by a developer is impacted by the number of other developers (s)he is connected through talk, his/her interpersonal influence in the network of work dependencies, the number of work items (s)he comments on, and the number work items (s)he owns. These effects are maintained even after controlling for workload, role, work dependency, and connection related factors. We discuss the implications of our results for collaborative software development and project governance.
Subhajit Datta, Renuka Sindhgatta, Bikram Sengupta
OOPSLA2
2010 Timesheet assistant: mining and reporting developer effort
abstract
Timesheets are an important instrument used to track time spent by team members in a software project on the tasks assigned to them. In a typical project, developers fill timesheets manually on a periodic basis. This is often tedious, time consuming and error prone. Over or under reporting of time spent on tasks causes errors in billing development costs to customers and wrong estimation baselines for future work, which can have serious business consequences. In order to assist developers in filling their timesheets accurately, we present a tool called Timesheet Assistant (TA) that non-intrusively mines developer activities and uses statistical analysis on historical data to estimate the actual effort the developer may have spent on individual assigned tasks. TA further helps the developer or project manager by presenting the details of the activities along with effort data so that the effort may be seen in the context of the actual work performed. We report on an empirical study of TA in a software maintenance project at IBM that provides preliminary validation of its feasibility and usefulness. Some of the limitations of the TA approach and possible ways to address those are also discussed.
Renuka Sindhgatta, Nanjangud C. Narendra, Bikram Sengupta, Karthik Visweswariah, Arthur G. Ryman
ASE1
2008 Identifying domain expertise of developers from source code
abstract
We are interested in identifying the domain expertise of developers of a software system. A developer gains expertise on the code base as well as the domain of the software system he/she develops. This information forms a useful input in allocating software implementation tasks to developers. Domain concepts represented by the system are discovered by taking into account the linguistic information available in the source code. The vocabulary contained in source code as identifiers such as class, method, variable names and comments are extracted. Concepts present in the code base are identified and grouped based on a well known text processing hypothesis - words are similar to the extent to which they share similar words. The developer's association with the source code and the concepts it represents is arrived at using the version repository information. In this line, the analysis first derives documents from source code by discarding all the programming language constructs. KMeans clustering is further used to cluster documents and extract closely related concepts. The key concepts present in the documents authored by the developer determine his/her domain expertise. To validate our approach we apply it on large software systems, two of which are presented in detail in this paper.
Renuka Sindhgatta
KDD1
2007 Identifying Software Decompositions by Applying Transaction Clustering on Source Code
abstract
Majority of the software clustering algorithms use structural dependencies to decompose large software systems. While these techniques have merit, they do not always match the decompositions generated by experts who often group software entities based on their purpose. This paper presents an approach to identifying decompositions of a software system, based on the joint participation of software entities in realizing the functionality of the system. Software transactions representing units of functionality are extracted from the source code. Transactions are clustered based on the commonality of software entities used in the transactions. Our approach also assesses the use of check-in data from configuration management system for software clustering where software entities that are modified or updated together form a software transaction. We introduce CoST, a clustering tool that uses Transaction Clustering to identify software decompositions. We apply CoST to three large software systems. The results indicate that this approach produces groupings that come close to decompositions prepared by experts.
Renuka Sindhgatta, Krishnakumar Pooloth
COMPSAC (1)1
2006 Using an information retrieval system to retrieve source code samples
abstract
Software developers often face steep learning curves in using a new framework, library, or new versions of frameworks for developing their piece of software. In large organizations, developers learn and explore use of frameworks, rarely realizing, several peers may have already explored the same. A tool that helps locate samples of code, demonstrating use of frameworks or libraries would provide benefits of reuse, improved code quality and faster development. This paper describes an approach for locating common samples of source code from a repository by providing extensions to an information retrieval system. The approach improves the existing approaches in two ways. First, it provides the scalability of an information retrieval system, supporting search over thousands of source code files of an organization. Second, it provides more specific search on source code by preprocessing source code files and understanding elements of the code as opposed to considering code as plain text.
Renuka Sindhgatta
ICSE1
2005 Functional and Non-functional Requirements Specification for Enterprise Applications
Renuka Sindhgatta, Srinivas Thonse
PROFES1