Parag C. Pendharkar

dblp:29/3030 · DBLP profile ↗
← Back
41ranked-venue papers
33as first author
4since 2021 · last 2023
0000-0001-8120-1995ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 23 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 A Radial Basis Function Neural Network for Stochastic Frontier Analyses of General Multivariate Production and Cost Functions
Parag C. Pendharkar
Neural Process. Lett.1
2022 On the analysis of power law distribution in software component sizes
abstract
Abstract Component‐based software development (CBSD) is an active area of research. Ascertaining the quality of components is important for overall software quality assurance in CBSD. One of the important metrics for measuring defects, analyzability, efforts, and cost in CBSD is component size. The paper presents an analytical model based on maximization of Tsallis entropy to obtain closed form expression for component size distribution (maximum Tsallis entropy component size distribution, MTECSD) in steady state. It is found that the component size distribution follows power law asymptotically. A procedure based on generalized Jensen–Shannon measure is developed to estimate model parameters. A detailed analysis of many popular probability distributions along with MTECSD is carried out on many diverse real data sets of component‐based softwares. The analysis reveals that lognormal and MTECSD distributions fit well to component sizes in many software conforming the presence of power law behavior. The software whose component size distributions are described by MTECSD are in equilibrium implying that new defects in these software systems occur occasionally. Power law behavior in component sizes also imply high variation leading to difficulty in software analyzability. The precise knowledge of component size distribution also provides an alternative method to compute efforts and cost estimates by modified COCOMO model.
Shachi Sharma, Parag C. Pendharkar
J. Softw. Evol. Process.2
2022 Learning Component Size Distributions for Software Cost Estimation: Models Based on Arithmetic and Shifted Geometric Means Rules
abstract
Understanding software size distribution is critical to software cost estimation using COCOMO model and design of reliable production function model. This paper proposes and validates a theoretical framework based on the maximization of Shannon entropy to learn component size distribution of software systems when partial information about the moments is given. Specification of appropriate moment constraints either in the form of shifted geometric mean or arithmetic mean or both geometric and arithmetic means are considered. The models are validated using 30 real datasets. The analysis reveals that software systems where component sizes depict power-law behavior are governed by shifted geometric mean whereas those systems in which component size distribution shows exponential behavior are described by arithmetic mean. Another type of software system is also considered where the component size distribution is found to depict gamma distribution. Such systems are characterized by specification of both arithmetic and geometric means. The study underlines that the use of modern object-oriented programming languages adheres to power-law distribution indicating the existence of team synergies leading to substantial containment of software costs when compared to the use of traditional procedural programming languages.
Shachi Sharma, Parag C. Pendharkar, Karmeshu
IEEE Trans. Software Eng.2
2021 Quantitative software project management with mixed data: A comparison of radial, nonradial, and ensemble data envelopment analysis models
abstract
Abstract Data envelopment analysis (DEA) models are often used for benchmarking software projects, and traditional DEA models only allow for the use of continuous variables. This study considers the use of DEA for datasets with mixed continuous and discrete variables to ranking software projects. It uses the existing radial DEA model and extends the nonradial DEA model to allow for the use of mixed variables. Further efficiency scores from the two DEA models are averaged in an ensemble DEA score. Using three real‐world software engineering datasets, this study finds that the nonradial DEA and the ensemble DEA models have better discriminating power (lower tied efficiency scores) to rank software projects, and the radial DEA model generates more general ranking distribution (higher entropy) of normalized efficiency scores. The choice between selecting radial and nonradial DEA models for ranking software projects appears to depend on the extent to which managers want to introduce bias into the efficiency score ranking distribution. Radial models appear to have a lower bias than nonradial models. The ensemble DEA model appears to be the best performing DEA model for datasets containing two or more discrete and continuous variables.
Parag C. Pendharkar, James A. Rodger
J. Softw. Evol. Process.1
2018 Trading financial indices with reinforcement learning agents
Parag C. Pendharkar, Patrick Cusatis
Expert Syst. Appl.1
2018 A hybrid genetic algorithm and DEA approach for multi-criteria fixed cost allocation
Parag C. Pendharkar
Soft Comput.1
2017 Bayesian posterior misclassification error risk distributions for ensemble classifiers
Parag C. Pendharkar
Eng. Appl. Artif. Intell.1
2017 An Examination of Determinants of Software Testing and Project Management Effort
abstract
Software estimation research has primarily focused on software effort involved in direct software development. As more and more organizations buy instead of building software, more effort is spent on software testing and project management. In this empirical study, the effect of program duration, computer platform, and software development tool (SDT) on program testing effort and project management effort is studied. The study results point to program duration and software tool as significant determinants of testing and management effort. Computer platform, however, does not have an effect on testing and management effort. Furthermore, the mean testing effort for third generation (3G) development environment was significantly higher than the mean testing effort for fourth generation (4G) environments that used IDE. In addition, the management effort for 4G environment projects without the use of IDE was lower than nonprogramming report generation projects.
Girish H. Subramanian, Parag C. Pendharkar, Dinesh R. Pai
J. Comput. Inf. Syst.2
2015 Linear models for cost-sensitive classification
abstract
Abstract In this paper, we investigate the performance of statistical, mathematical programming and heuristic linear models for cost‐sensitive classification. In particular, we use five cost‐sensitive techniques including Fisher's discriminant analysis (DA), asymmetric misclassification cost mixed integer programming (AMC‐MIP), cost‐sensitive support vector machine (CS‐SVM), a hybrid support vector machine and mixed integer programming (SVMIP) and heuristic cost‐sensitive genetic algorithm (CGA) techniques. Using simulated datasets of varying group overlaps, data distributions and class biases, and real‐world datasets from financial and medical domains, we compare the performances of our five techniques based on overall holdout sample misclassification cost. The results of our experiments on simulated datasets indicate that when group overlap is low and data distribution is exponential, DA appears to provide superior performance. For all other situations with simulated datasets, CS‐SVM provides superior performance. In case of real‐world datasets from financial domain, CGA and AMC‐MIP hold a slight edge over the two SVM‐based classifiers. However, for medical domains with mixed continuous and discrete attributes, SVM classifiers perform better than heuristic (CGA) and AMC‐MIP classifiers. The SVMIP model is the most computationally inefficient model and poor performing model.
Parag C. Pendharkar
Expert Syst. J. Knowl. Eng.1
2015 Ensemble based point and confidence interval forecasting in software engineering
Parag C. Pendharkar
Expert Syst. Appl.1
2014 A misclassification cost risk bound based on hybrid particle swarm optimization heuristic
Parag C. Pendharkar
Expert Syst. Appl.1
2013 Scatter search based interactive multi-criteria optimization of fuzzy objectives for coal production planning
Parag C. Pendharkar
Eng. Appl. Artif. Intell.1
2013 A maximum-margin genetic algorithm for misclassification cost minimizing feature selection problem
Parag C. Pendharkar
Expert Syst. Appl.1
2013 A Normalized Probabilistic Expectation-Maximization Neural Network for Minimizing Bayesian Misclassification Cost Risk
Parag C. Pendharkar
Neural Process. Lett.1
2012 A distributed problem-solving framework for probabilistic software effort estimation
abstract
Abstract Distributed problem‐solving (DPS) systems use a framework of human organizational notions and principles of intelligent systems to solve complex problems. Human organizational notions are used to decompose a complex problem into sub‐problems that can be solved using intelligent systems. The solutions of these sub‐problems are combined to solve the original complex problem. In this paper, we propose a DPS system for probabilistic estimation of software development effort. Using a real‐world software engineering dataset, we compare the performance of the DPS system with a neural network (NN) and show that the performance of the DPS system is equal to or better than that of the NN with the additional benefits of modularity, probabilistic estimates, greater interpretability, flexibility and capability to handle incomplete input data.
Parag C. Pendharkar, James A. Rodger
Expert Syst. J. Knowl. Eng.1
2012 Game theoretical applications for multi-agent systems
Parag C. Pendharkar
Expert Syst. Appl.1
2012 DEA based data preprocessing for maximum decisional efficiency linear case valuation models
Parag C. Pendharkar, Marvin D. Troutt
Expert Syst. Appl.1
2012 Fuzzy classification using the data envelopment analysis
Parag C. Pendharkar
Knowl. Based Syst.1
2010 Exhaustive and heuristic search approaches for learning a software defect prediction model
Parag C. Pendharkar
Eng. Appl. Artif. Intell.1
2010 Probabilistic estimation of software size and effort
Parag C. Pendharkar
Expert Syst. Appl.1
2009 Genetic algorithm based neural network approaches for predicting churn in cellular wireless network services
Parag C. Pendharkar
Expert Syst. Appl.1
2008 A threshold varying bisection method for cost sensitive learning in neural networks
Parag C. Pendharkar
Expert Syst. Appl.1
2008 An empirical study of the Cobb-Douglas production function properties of software development effort
Parag C. Pendharkar, James A. Rodger, Girish H. Subramanian
Inf. Softw. Technol.1
2007 The theory and experiments of designing cooperative intelligent systems
Parag C. Pendharkar
Decis. Support Syst.1
2007 A field study of database communication issues peculiar to users of a voice activated medical tracking application
James A. Rodger, Parag C. Pendharkar
Decis. Support Syst.2
2007 A comparison of gradient ascent, gradient descent and genetic-algorithm-based artificial neural networks for the binary classification problem
abstract
Abstract: We compare log maximum likelihood gradient ascent, root‐mean‐square error minimizing gradient descent and genetic‐algorithm‐based artificial neural network procedures for a binary classification problem. We use simulated data and real‐world data sets, and four different performance metrics of correct classification, sensitivity, specificity and reliability for our comparisons. Our experiments indicate that a genetic‐algorithm‐based artificial neural network that maximizes the total number of correct classifications generally fares well for the binary classification problem. However, if the training data set contains inconsistent decisions or noise then the log maximum likelihood maximizing gradient ascent may be the best classification approach to use. The root‐mean‐square minimizing gradient descent approach appears to overfit training data and has the lowest reliability among the approaches considered for our research. At the end of the paper, we provide a few guidelines, including computational complexity, for selection of an appropriate technique for a given binary classification problem.
Parag C. Pendharkar
Expert Syst. J. Knowl. Eng.1
2006 A Multi-Agent Distributed Channel Allocation Approach for Wireless Networks
abstract
We propose a multi-agent approach for distributed channel allocation (MA-DCA) in mobile cellular networks. Our approach assumes that each cell in a cellular network works as an agent that negotiates its bandwidth (channel) requirements with its neighbors so that the probability of its call drops is minimized, and channel reuse constraint is satisfied. Using simulations, we compare our MA-DCA approach with simple fixed channel allocation and dynamic channel borrowing approaches and illustrate that the MA-DCA is a superior approach for channel allocation.
Parag C. Pendharkar
VTC Fall1
2006 An empirical study of the effect of complexity, platform, and program type on software development effort of business applications
Girish H. Subramanian, Parag C. Pendharkar, Mary Wallace
Empir. Softw. Eng.2
2005 A threshold varying bisection method for cost sensitive learning in neural networks
abstract
We propose a bisection method for varying classification threshold value for cost sensitive neural network learning. Using simulated data and different cost asymmetries, we test the proposed threshold varying bisection method and compare it with the traditional fixed-threshold method based neural network learning. The results of our experiments illustrate that the proposed threshold varying bisection method performs better than the traditional fixed-threshold method.
Parag C. Pendharkar
IJCNN1
2005 Hybrid approaches for classification under information acquisition cost constraint
Parag C. Pendharkar
Decis. Support Syst.1
2005 A Data Envelopment Analysis-Based Approach for Data Preprocessing
abstract
In this paper, we show how the data envelopment analysis (DEA) model might be useful to screen training data so a subset of examples that satisfy monotonicity property can be identified. Using real-world health care and software engineering data, managerial monotonicity assumption, and artificial neural network (ANN) as a forecasting model, we illustrate that DEA-based data screening of training data improves forecasting accuracy of an ANN.
Parag C. Pendharkar
IEEE Trans. Knowl. Data Eng.1
2005 A Probabilistic Model for Predicting Software Development Effort
abstract
Recently, Bayesian probabilistic models have been used for predicting software development effort. One of the reasons for the interest in the use of Bayesian probabilistic models, when compared to traditional point forecast estimation models, is that Bayesian models provide tools for risk estimation and allow decision-makers to combine historical data with subjective expert estimates. In this paper, we use a Bayesian network model and illustrate how a belief updating procedure can be used to incorporate decision-making risks. We develop a causal model from the literature and, using a data set of 33 real-world software projects, we illustrate how decision-making risks can be incorporated in the Bayesian networks. We compare the predictive performance of the Bayesian model with popular nonparametric neural-network and regression tree forecasting models and show that the Bayesian model is a competitive model for forecasting software development effort.
Parag C. Pendharkar, Girish H. Subramanian, James A. Rodger
IEEE Trans. Software Eng.1
2004 An exploratory study of object-oriented software component size determinants and the application of regression tree forecasting models
Parag C. Pendharkar
Inf. Manag.1
2004 Editorial
Parag C. Pendharkar
Int. J. Hum. Comput. Stud.1
2004 A field study of the impact of gender and user's technical experience on the performance of voice-activated medical tracking application
James A. Rodger, Parag C. Pendharkar
Int. J. Hum. Comput. Stud.2
2004 Human-computer interaction issues for mobile computing in a variable work context
Judy York, Parag C. Pendharkar
Int. J. Hum. Comput. Stud.2
2003 A Probabilistic Model for Predicting Software Development Effort
Parag C. Pendharkar, Girish H. Subramanian, James A. Rodger
ICCSA (2)1
2003 Technical efficiency-based selection of learning cases to improve forecasting accuracy of neural networks under monotonicity assumption
Parag C. Pendharkar, James A. Rodger
Decis. Support Syst.1
2002 Connectionist Models for Learning, Discovering, and Forecasting Software Effort: An Empirical Study
Parag C. Pendharkar, Girish H. Subramanian
J. Comput. Inf. Syst.1
2001 Development and Testing of an Instrument for Measuring the User Evaluations of Information Technology in Health Care
Parag C. Pendharkar, Mehdi Khosrow-Pour, James A. Rodger
J. Comput. Inf. Syst.1
2001 Mobile Computing at the Department of Defense
abstract
This paper is designed to relate the rationale used by the Department of Defense, to utilize Telemedicine, to meet increasing global crises, and for the U.S. military to find ways to more effectively manage manpower and time. A mobile Telemedicine package has been developed by the Department of Defense (DOD) to collect and transmit near-real-time, far-forward medical data and to assess how this improved capability enhances medical management of the battlespace. Telemedicine has been successful in resolving uncertain organizational and technological military deficiencies and in improving medical communications and information management. The deployable, mobile Teams are the centerpieces of this Telemedicine package. These teams have the capability of inserting essential networking and communications capabilities into austere theaters and establishing an immediate means for enhancing health protection, collaborative planning, situational awareness, and strategic decision-making.
James A. Rodger, Parag C. Pendharkar, Mehdi Khosrow-Pour
J. Database Manag.2