James Miller 0001

dblp:44/7043-1 · DBLP profile ↗
← Back
108ranked-venue papers
15as first author
12since 2021 · last 2026
0000-0001-5095-3000ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 74 · 14 first-author · 7 since 2021Databases, data management, data science and information retrieval · 15 · 3 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Computer networks · 3Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CARE: Context Aware Root Cause Identification Using Distributed Traces and Profiling Metrics
abstract
Root cause localization in microservices is challenging due to intricate service dependencies and the high volume and heterogeneity of collected monitoring data, which add complexity to the analysis. Conventional methods often overlook nuanced propagation patterns and contextual interactions among services, and they are limited in leveraging multi-source observability data for comprehensive root cause identification. This study introduces CARE, a context-aware, spectrum-analysis-based approach that integrates multi-source observability data and employs network analysis to prioritize the contextual significance of components in propagating anomalies across individual services, service communities, and requests. CARE’s weighted spectrum analysis leverages these prioritized contexts to pinpoint underlying performance issues. Evaluations on 224 cases from the TrainTicket benchmark and a real-world Internet service provider’s production system demonstrate CARE’s substantial accuracy gains, with top-1 accuracy of 72%-89% and top-5 accuracy of 84%-99% for single root causes, outperforming baselines by 8%-41%. CARE also shows significant improvements in dual root cause identification, exceeding baseline performance by 18%-37%, all while maintaining efficient resource usage, establishing CARE as a robust and resource-effective solution for root cause localization in complex microservice environments.
Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj, James Miller 0001
IEEE Trans. Software Eng.4
2025 Refining software defect prediction through attentive neural models for code understanding
Mona Nashaat, James Miller 0001
J. Syst. Softw.2
2024 ServiceAnomaly: An anomaly detection approach in microservices using distributed traces and profiling metrics
Mahsa Panahandeh, Abdelwahab Hamou-Lhadj, Mohammad Hamdaqa, James Miller 0001
J. Syst. Softw.4
2024 Towards Efficient Fine-Tuning of Language Models With Organizational Data for Automated Software Review
abstract
Large language models like BERT and GPT possess significant capabilities and potential impacts across various applications. Software engineers often use these models for code-related tasks, including generating, debugging, and summarizing code. Nevertheless, large language models still have several flaws, including model hallucination. (e.g., generating erroneous code and producing outdated and inaccurate programs) and the substantial computational resources and energy required for training and fine-tuning. To tackle these challenges, we propose CodeMentor, a framework for few-shot learning to train large language models with the data available within the organization. We employ the framework to train a language model for code review activities, such as code refinement and review generation. The framework utilizes heuristic rules and weak supervision techniques to leverage available data, such as previous review comments, issue reports, and related code updates. Then, the framework employs the constructed dataset to fine-tune LLMs for code review tasks. Additionally, the framework integrates domain expertise by employing reinforcement learning with human feedback. This allows domain experts to assess the generated code and enhance the model performance. Also, to assess the performance of the proposed model, we evaluate it with four state-of-the-art techniques in various code review tasks. The experimental results attest that CodeMentor enhances the performance in all tasks compared to the state-of-the-art approaches, with an improvement of up to 22.3%, 43.4%, and 24.3% in code quality estimation, review generation, and bug report summarization tasks, respectively.
Mona Nashaat, James Miller 0001
IEEE Trans. Software Eng.2
2022 An End-to-End Trainable Deep Convolutional Neuro-Fuzzy Classifier
abstract
A key challenge in artificial intelligence is the well-known tradeoff between the interpretability of an algorithm, and its accuracy. Designing interpretable, highly accurate AI models is considered essential to broad acceptance of AI technology, and is the focus of the eXplainable Artificial Intelligence (XAI) community. We report on the design of a new deep neural network that achieves improved interpretability without sacrificing accuracy. Our design is a hybrid deep learning algorithm based in part upon fuzzy logic, which performs as accurately as existing convolutional neural networks. The network is an end-to-end trainable deep convolutional network, which replaces the final dense layers (the classifier component) with a modified ANFIS. We exploit the transparency of fuzzy logic by deriving explanations, in the form of saliency maps, based on the fuzzy rules learned in the ANFIS component.
Mojtaba Yeganejou, Ryan Kluzinski, Scott Dick, James Miller 0001
FUZZ-IEEE4
2022 Automatically inferring user behavior models in large-scale web applications
Saeedeh Sajjadi-Ghaem-Maghami, Seyedeh Sepideh Emam, James Miller 0001
Inf. Softw. Technol.3
2022 Integrated-Block: A New Combination Model to Improve Web Page Segmentation
abstract
Context: Web page segmentation methods have been used for different purposes such as web page classification and content analysis. These methods categorize a web page into different blocks, where each block contains similar components. Objective: The goal of this paper is to propose a new segmentation approach that semantically segments web pages into integrated blocks and obtains high segmentation accuracy. Method: In this paper, we propose a new segmentation model that semantically segments web pages into integrated blocks, where (1) it merges web page content into basic-blocks by simulating human perception using Gestalt laws of grouping; and, (2) it utilizes semantic text similarity to identify similar blocks and regroup these similar basic-blocks as integrated blocks. Results: To verify the accuracy of our approach, we (1) applied it to three datasets, (2) compared it with the five existing state-of-the-art algorithms. The results show that our approach outperforms all the five comparison methods in terms of precision, recall, F-1 score, and ARI. Conclusion: In this paper, we propose a new segmentation model and apply it to three datasets to (1) generate basic-blocks by simulating human perception to segment a web page, (2) identify semantically related blocks and regroup them as an integrated block, and (3) address limitations found in existing approaches.
Saeedeh Sajjadi-Ghaem-Maghami, James Miller 0001
J. Web Eng.2
2022 Semi-Supervised Ensemble Learning for Dealing with Inaccurate and Incomplete Supervision
abstract
In real-world tasks, obtaining a large set of noise-free data can be prohibitively expensive. Therefore, recent research tries to enable machine learning to work with weakly supervised datasets, such as inaccurate or incomplete data. However, the previous literature treats each type of weak supervision individually, although, in most cases, different types of weak supervision tend to occur simultaneously. Therefore, in this article, we present Smart MEnDR, a Classification Model that applies Ensemble Learning and Data-driven Rectification to deal with inaccurate and incomplete supervised datasets. The model first applies a preliminary phase of ensemble learning in which the noisy data points are detected while exploiting the unlabelled data. The phase employs a semi-supervised technique with maximum likelihood estimation to decide on the disagreement rate. Second, the proposed approach applies an iterative meta-learning step to tackle the problem of knowing which points should be made correct to improve the performance of the final classifier. To evaluate the proposed framework, we report the classification performance, noise detection, and the labelling accuracy of the proposed method against state-of-the-art techniques. The experimental results demonstrate the effectiveness of the proposed framework in detecting noise, providing correct labels, and attaining high classification performance.
Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader
ACM Trans. Knowl. Discov. Data3
2022 Interpretation of Structural Preservation in Low-Dimensional Embeddings
abstract
Despite being commonly used in big-data analytics; the outcome of dimensionality reduction remains a black-box to most of its users. Understanding the quality of a low-dimensional embedding is important as not only it enables trust in the transformed data, but it can also help to select the most appropriate dimensionality reduction algorithm in a given scenario. As existing research primarily focuses on the visual exploration of embeddings, there is still a need for enhancing interpretability of such algorithms. To bridge this gap, we propose two novel interactive explanation techniques for low-dimensional embeddings obtained fromanydimensionality reduction algorithm. The first technique LAPS produces a local approximation of the neighborhood structure to generate interpretable explanations on the preserved locality for a single instance. The second method GAPS explains the retained global structure of a high-dimensional dataset in its embedding, by combining non-redundant local-approximations from a coarse discretization of the projection space. We demonstrate the applicability of the proposed techniques using 16 real-life tabular, text, image, and audio datasets. Our extensive experimental evaluation shows the utility of the proposed techniques in interpreting the quality of low-dimensional embeddings, as well as with selecting the most suitable dimensionality reduction algorithm for any given dataset.
Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader
IEEE Trans. Knowl. Data Eng.3
2022 VisExPreS: A Visual Interactive Toolkit for User-Driven Evaluations of Embeddings
abstract
Although popularly used in big-data analytics, dimensionality reduction is a complex, black-box technique whose outcome is difficult to interpret and evaluate. In recent years, a number of quantitative and visual methods have been proposed for analyzing low-dimensional embeddings. On the one hand, quantitative methods associate numeric identifiers to qualitative characteristics of these embeddings; and, on the other hand, visual techniques allow users to interactively explore these embeddings and make decisions. However, in the former case, users do not have control over the analysis, while in the latter case. assessment decisions are entirely dependent on the user's perception and expertise. In order to bridge the gap between the two, in this article, we present VisExPreS, a visual interactive toolkit that enables a user-driven assessment of low-dimensional embeddings. VisExPreS is based on three novel techniques namely PG-LAPS, PG-GAPS, and RepSubset, that generate interpretable explanations of the preserved local and global structures in embeddings. In the first two techniques, the VisExPreS system proactively guides users during every step of the analysis. We demonstrate the utility of VisExPreS in interpreting, analyzing, and evaluating embeddings from different dimensionality reduction algorithms using multiple case studies and an extensive user study.
Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader
IEEE Trans. Vis. Comput. Graph.3
2021 A New Semantic Approach to Improve Webpage Segmentation
abstract
Webpage analysis is carried out for various purposes such as webpage segmentation. The goal of webpage segmentation is to divide a page into blocks that have similar elements. A fusion approach that combines different analyses is required in order to obtain high segmentation accuracy. In this paper, we propose a new fusion model for webpage segmentation, where we (1) merge webpage content into basic-blocks by simulating human perception; and, (2) identify similar blocks using semantic text similarity and regroup these similar blocks as fusion blocks. This approach is applied to three public datasets and evaluated by comparing with state-of-the-art algorithms. The results characterize that our proposed approach outperforms other existing webpage segmentation methods, in terms of accuracy.
Saeedeh Sajjadi-Ghaem-Maghami, James Miller 0001
J. Web Eng.2
2021 Context-Based Evaluation of Dimensionality Reduction Algorithms - Experiments and Statistical Significance Analysis
abstract
Dimensionality reduction is a commonly used technique in data analytics. Reducing the dimensionality of datasets helps not only with managing their analytical complexity but also with removing redundancy. Over the years, several such algorithms have been proposed with their aims ranging from generating simple linear projections to complex non-linear transformations of the input data. Subsequently, researchers have defined several quality metrics in order to evaluate the performances of different algorithms. Hence, given a plethora of dimensionality reduction algorithms and metrics for their quality analysis, there is a long-existing need for guidelines on how to select the most appropriate algorithm in a given scenario. In order to bridge this gap, in this article, we have compiled 12 state-of-the-art quality metrics and categorized them into 5 identified analytical contexts. Furthermore, we assessed 15 most popular dimensionality reduction algorithms on the chosen quality metrics using a large-scale and systematic experimental study. Later, using a set of robust non-parametric statistical tests, we assessed the generalizability of our evaluation on 40 real-world datasets. Finally, based on our results, we present practitioners’ guidelines for the selection of an appropriate dimensionally reduction algorithm in the present analytical contexts.
Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader
ACM Trans. Knowl. Discov. Data3
2020 Interpretable Deep Convolutional Fuzzy Classifier
abstract
While deep learning has proven to be a powerful new tool for modeling and predicting a wide variety of complex phenomena, those models remain incomprehensible black boxes. This is a critical impediment to the widespread deployment of deep learning technology, as decades of research have found that users simply will not trust (i.e., make decisions based on) a model whose solutions cannot be explained. Fuzzy systems, on the other hand, are by design much more easily understood. In this article, we propose to create more comprehensible deep networks by hybridizing them with fuzzy logic. Our proposed architecture first employs a convolutional neural network as an automated feature extractor and then performs a fuzzy clustering in the derived feature space. After hardening the clusters, we employ Rocchio's algorithm to classify the data points. Experiments on three datasets show that the automated feature extraction substantially improves the accuracy of the fuzzy classifier, and while the substitution of a fuzzy classifier slightly decreases the network's performance, we are able to introduce an effective interpretation mechanism.
Mojtaba Yeganejou, Scott Dick, James Miller 0001
IEEE Trans. Fuzzy Syst.3
2019 The current state of software license renewals in the I.T. industry
Aindrila Ghosh, Mona Nashaat, James Miller 0001
Inf. Softw. Technol.3
2019 M-Lean: An end-to-end development framework for predictive models in B2B scenarios
Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader, Chad Marston
Inf. Softw. Technol.3
2019 A Pure Visual Approach for Automatically Extracting and Aligning Structured Web Data
abstract
Database-driven websites and the amount of data stored in their databases are growing enormously. Web databases retrieve relevant information in response to users’ queries; the retrieved information is encoded in dynamically generated web pages as structured data records. Identifying and extracting retrieved data records is a fundamental task for many applications, such as competitive intelligence and comparison shopping. This task is challenging due to the complex underlying structure of such web pages and the existence of irrelevant information. Numerous approaches have been introduced to address this problem, but most of them are HTML-dependent solutions that may no longer be functional with the continuous development of HTML. Although a few vision-based techniques have been introduced, various issues exist that inhibit their performance. To overcome this, we propose a novel visual approach, i.e., programming-language-independent, for automatically extracting structured web data. The proposed approach makes full use of the natural human tendency of visual object perception and the Gestalt laws of grouping. The extraction system consists of two tasks: (1) data record extraction, where we apply three of the Gestalt laws (i.e., laws of continuity, proximity, and similarity), which are used to group the adjacently aligned visually similar data records on a web page; and (2) data item extraction and alignment, where we employ the Gestalt law of similarity, which is utilized to group the visually identical data items. Our experiments upon large-scale test sets show that the proposed system is highly effective and outperforms the two state-of-art vision-based approaches, ViDE and rExtractor. The experiments produce an average F1 score of 86.02%, which is approximately 55% and 36% better than that of ViDE and rExtractor for data record extraction, respectively; and an average F1 score of 86.19%, which is approximately 39% better than that of ViDE for data item extraction.
Fadwa Estuka, James Miller 0001
ACM Trans. Internet Techn.2
2019 Integrative Double Kaizen Loop (IDKL): Towards a Culture of Continuous Learning and Sustainable Improvements for Software Organizations
abstract
In the past decades, software organizations have been relying on implementing process improvement methods to advance quality, productivity, and predictability of their development and maintenance efforts. However, these methods have proven to be challenging to implement in many situations, and when implemented, their benefits are often not sustained. Commonly, the workforce requires guidance during the initial deployment, but what happens after the guidance stops? Why do not traditional improvement methods deliver the desired results? And, how do we maintain the improvements when they are realized? In response to these questions, we have combined social and organizational learning methods with Lean's continuous improvement philosophy, Kaizen, which has resulted in an IDKL model that has successfully promoted continuous learning and improvement. The IDKL has evolved through a real-life project with an industrial partner; the study employed ethnographic action research with 231 participants and had lasted for almost 3 years. The IDKL requires employees to continuously apply small improvements to the daily routines of the work-procedures. The small improvements by themselves are unobtrusive. However, the IDKL has helped the industrial partner to implant continuous improvement as a daily habit. This has led to realizing sustainable and noticeable improvements. The findings show that on average, Lead Time has dropped by 46 percent, Process Cycle Efficiency has increased by 137 percent, First-Pass Process Yield has increased by 27 percent, and Customer Satisfaction has increased by 25 percent.
Osama Al-Baik, James Miller 0001
IEEE Trans. Software Eng.2
2018 Hybridization of Active Learning and Data Programming for Labeling Large Industrial Datasets
abstract
Modern machine learning (ML) models are being used heavily in business domains to build effective decision support systems. As a primary requirement, supervised ML models need large labeled datasets. However, obtaining a high volume of labeled training data is both expensive and time-consuming. Researchers have proposed several labeling approaches to avoid manual labeling efforts. Active learning (AL) and Data Programming (DP) are two state-of-the-art techniques used to label datasets. Nevertheless, both approaches have their strengths and weaknesses. For example, AL is computationally expensive to apply on large industrial datasets; and labels generated by DP are often inaccurate and difficult to interpret. To address these challenges, in this paper, we propose a novel hybrid method that integrates the scalability of DP with the user engagement and accuracy of AL. The proposed approach aims at optimizing the labeling process by applying DP to generate initial noisy training data and then use AL to query the user to label only those points that maximize the accuracy of the final labels with a minimum annotation cost. To evaluate the proposed approach, we have used five open source datasets and a real-world business dataset of 1.5 million records. We use traditional active learning and data programming techniques as baselines to compare the performance and annotation cost of our proposed approach. The results show that the proposed method can achieve higher labeling accuracy than data programming. It also can minimize the labeling cost in real-world business scenarios, while delivering a comparable level of performance (accuracy) with active learning.
Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader, Chad Marston, Jean-François Puget
IEEE BigData3
2018 Black-box tree test case generation through diversity
Ali Shahbazi, Mahsa Panahandeh, James Miller 0001
Autom. Softw. Eng.3
2018 Cross-Browser Differences Detection Based on an Empirical Metric for Web Page Visual Similarity
abstract
This article aims to develop a method to detect visual differences introduced into web pages when they are rendered in different browsers. To achieve this goal, we propose an empirical visual similarity metric by mimicking human mechanisms of perception. The Gestalt laws of grouping are translated into a computer compatible rule set. A block tree is then parsed by the rules for similarity calculation. During the translation of the Gestalt laws, experiments are performed to obtain metrics for proximity, color similarity, and image similarity. After a validation experiment, the empirical metric is employed to detect cross-browser differences. Experiments and case studies on the world’s most popular web pages provide positive results for this methodology.
James Miller 0001
ACM Trans. Internet Techn.2
2018 Inferring Extended Probabilistic Finite-State Automaton Models from Software Executions
abstract
Behavioral models are useful tools in understanding how programs work. Although several inference approaches have been introduced to generate extended finite-state automatons from software execution traces, they suffer from accuracy, flexibility, and decidability issues. In this article, we apply a hybrid technique to use both reinforcement learning and stochastic modeling to generate an extended probabilistic finite state automaton from software traces. Our approach—ReHMM (Reinforcement learning-based Hidden Markov Modelling)—is able to address the problems of inflexibility and un-decidability reported in other state-of-the-art approaches. Experimental results indicate that ReHMM outperforms other inference algorithms.
Seyedeh Sepideh Emam, James Miller 0001
ACM Trans. Softw. Eng. Methodol.2
2018 A comprehensive review of tools for exploratory analysis of tabular industrial datasets
abstract
Exploratory data analysis plays a major role in obtaining insights from data. Over the last two decades, researchers have proposed several visual data exploration tools that can assist with each step of the analysis process. Nevertheless, in recent years, data analysis requirements have changed significantly. With constantly increasing size and types of data to be analyzed, scalability and analysis duration are now among the primary concerns of researchers. Moreover, in order to minimize the analysis cost, businesses are in need of data analysis tools that can be used with limited analytical knowledge. To address these challenges, traditional data exploration tools have evolved within the last few years. In this paper, with an in-depth analysis of an industrial tabular dataset, we identify a set of additional exploratory requirements for large datasets. Later, we present a comprehensive survey of the recent advancements in the emerging field of exploratory data analysis. We investigate 50 academic and non-academic visual data exploration tools with respect to their utility in the six fundamental steps of the exploratory data analysis process. We also examine the extent to which these modern data exploration tools fulfill the additional requirements for analyzing large datasets. Finally, we identify and present a set of research opportunities in the field of visual exploratory data analysis.
Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader, Chad Marston
Vis. Informatics3
2017 Detecting Security Vulnerabilities in Object-Oriented PHP Programs
abstract
PHP is one of the most popular web development tools in use today. A major concern though is the improper and insecure uses of the language by application developers, motivating the development of various static analyses that detect security vulnerabilities in PHP programs. However, many of these approaches do not handle recent, important PHP features such as object orientation, which greatly limits the use of such approaches in practice. In this paper, we present OOPIXY, a security analysis tool that extends the PHP security analyzer PIXY to support reasoning about object-oriented features in PHP applications. Our empirical evaluation shows that OOPIXY detects 88% of security vulnerabilities found in micro benchmarks. When used on real-world PHP applications, OOPIXY detects security vulnerabilities that could not be detected using state-of-the-art tools, retaining a high level of precision. We have contacted the maintainers of those applications, and two applications' development teams verified the correctness of our findings. They are currently working on fixing the bugs that lead to those vulnerabilities.
Mona Nashaat, Karim Ali 0001, James Miller 0001
SCAM3
2017 A Fuzzy Recommender System for Public Library Catalogs
abstract
Recommendation engines are one of the “discovery” products built into integrated library systems. These are a subclass of enterprise systems designed specifically for public and research libraries that incorporate an electronic card catalogue, circulation and inventory management, personnel and payroll systems, etc. The system vendors offer customizations for different contexts of specific library systems, but cannot create a bespoke solution for every customer. Our partner, an Edmonton-area company, is filling this gap for a consortium of rural libraries in Alberta by creating a mobile app that interfaces with their electronic card catalog. Rural libraries are generally smaller than major urban public libraries, meaning that their holdings are limited overall, and within any given genre. This poses a severe problem for traditional collaborative-filtering recommender algorithms, as the item sets for recommendations are limited by supply rather than by readers’ interests. The library's relatively small clientele also limits the item sets available for comparison. To deal with this ongoing “cold-start” problem, we propose a hybridization of collaborative filtering with a content filter using a fuzzy taste vector. Experiments on two benchmark recommender data sets show that this approach is at least as accurate as existing fuzzy recommenders and is particularly effective on sparse data sets.
Jason M. Morawski, Torin Stepan, Scott Dick, James Miller 0001
Int. J. Intell. Syst.4
2017 A comparative study of the performance of local feature-based pattern recognition algorithms
Narges Roshanbin, James Miller 0001
Pattern Anal. Appl.2
2017 Using π-calculus for Formal Modeling and Verification of WS-CDL Choreographies
abstract
Service-Oriented applications are realized by composing and aggregating existing web services. Orchestration and Choreography are two interaction models for building SOA applications and several standards exist to capture and describe such interactions. The Web Service Choreography Description Language (WS-CDL) is a standard for modeling choreographies. In this paper, we propose a calculus (Chor-calculus) for formal modeling of WS-CDL and we use this language for generating WS-CDL programs. This approach enables the static verification of choreographies using existing pi-calculus model-checker tools and sets the ground for enabling the runtime monitoring of choreographies for behavioral correctness. We validate the calculus for its expressiveness by evaluating the language support for representing workflow, and service interaction, patterns. We demonstrate the use of the HAL toolkit to verify the correctness properties of choreographies.
Adel Khaled, James Miller 0001
IEEE Trans. Serv. Comput.2
2016 ADAMAS: Interweaving unicode and color to enhance CAPTCHA security
Narges Roshanbin, James Miller 0001
Future Gener. Comput. Syst.2
2016 An investigation of implicit features in compression-based learning for comparing webpages
Teh-Chung Chen, Torin Stepan, Scott Dick, James Miller 0001
Pattern Anal. Appl.4
2016 Incorporating Spatial, Temporal, and Social Context in Recommendations for Location-Based Social Networks
abstract
Location-based social networks (LBSNs) such as Foursquare, Brightkite, and Gowalla are a growing area where recommendation algorithms find a practical application. With an ever-increasing variety of venues to choose from deciding on a destination can be overwhelming. Recommenders aid their users in the decision-making process by providing a list of locations likely to be relevant to the user's needs and interests. Traditional collaborative filtering algorithms consider relationships between users and locations, finding users to be similar only if their location histories overlap. However, the availability of spatial, temporal, and social information in an LBSN offers an opportunity to improve the quality of a recommendation engine. Social network data allows us to connect users who can directly influence each other's decisions. Temporal data allows us to account for the drifting preferences of users, giving more weight to recent location visits over historical selections, and taking advantages of repetitive behaviors. Spatial information allows us to focus recommendations on locations close to the user, keeping our recommendations relevant as a user travels. We introduce a novel approach to incorporate spatial, temporal, and social context into a traditional collaborative filtering algorithm. We evaluate our method on data sets collected from three LBSNs, and demonstrate that our approach is at the least competitive with existing state-of-the-art location recommenders.
Torin Stepan, Jason M. Morawski, Scott Dick, James Miller 0001
IEEE Trans. Comput. Soc. Syst.4
2016 Black-Box String Test Case Generation through a Multi-Objective Optimization
abstract
String test cases are required by many real-world applications to identify defects and security risks. Random Testing (RT) is a low cost and easy to implement testing approach to generate strings. However, its effectiveness is not satisfactory. In this research, black-box string test case generation methods are investigated. Two objective functions are introduced to produce effective test cases. The diversity of the test cases is the first objective, where it can be measured through string distance functions. The second objective is guiding the string length distribution into a Benford distribution based on the hypothesis that the population of strings is right-skewed within its range. When both objectives are applied via a multi-objective optimization algorithm, superior string test sets are produced. An empirical study is performed with several real-world programs indicating that the generated string test cases outperform test cases generated by other methods.
Ali Shahbazi, James Miller 0001
IEEE Trans. Software Eng.2
2016 Identifying semantic blocks in Web pages using Gestalt laws of grouping
James Miller 0001
World Wide Web2
2015 A New Webpage Classification Model Based on Visual Information Using Gestalt Laws of Grouping
James Miller 0001
WISE (2)2
2015 The kanban approach, between agility and leanness: a systematic review
Osama Al-Baik, James Miller 0001
Empir. Softw. Eng.2
2015 Test Case Prioritization Using Extended Digraphs
abstract
Although many test case prioritization techniques exist, their performance is far from perfect. Hence, we propose a new fault-based test case prioritization technique to promote fault-revealing test cases inmodel-based testing(MBT) procedures. We seek to improve the fault detection rate—a measure of how fast a test suite is able to detect faults during testing—in scenarios such as regression testing. We propose an extended digraph model as the basis of this new technique. The model is realized using a novelreinforcement-learning(RL)- and hidden-Markov-model (HMM)-based technique which is able to prioritize test cases for regression testing objectives. We present a method to initialize and train an HMM based upon RL concepts applied to an application's digraph model. The model prioritizes test cases based upon forward probabilities, a new test case prioritization approach. In addition, we also propose an alternative approach to prioritizing test cases according to the amount of change they cause in applications. To evaluate the effectiveness of the proposed techniques, we perform experiments ongraphical user interface(GUI)-based applications and compare the results with state-of-the-art test case prioritization approaches. The experimental results show that the proposed technique is able to detect faults early within test runs.
Seyedeh Sepideh Emam, James Miller 0001
ACM Trans. Softw. Eng. Methodol.2
2014 Towards Real Time Contextual Advertising
Abhimanyu Panwar, Iosif-Viorel Onut, James Miller 0001
WISE (2)3
2014 Waste identification and elimination in information technology organizations
Osama Al-Baik, James Miller 0001
Empir. Softw. Eng.2
2014 An empirical investigation of Web session workloads: Can self-similarity be explained by deterministic chaos?
Scott Dick, Omolbanin Yazdanbakhsh, Xiuli Tang, Toan Huynh, James Miller 0001
Inf. Process. Manag.5
2014 An Anti-Phishing System Employing Diffused Information
abstract
The phishing scam and its variants are estimated to cost victims billions of dollars per year. Researchers have responded with a number of anti-phishing systems, based either on blacklists or on heuristics. The former cannot cope with the churn of phishing sites, while the latter usually employ decision rules that are not congruent to human perception. We propose a novel heuristic anti-phishing system that explicitly employs gestalt and decision theory concepts to model perceptual similarity. Our system is evaluated on three corpora contrasting legitimate Web sites with real-world phishing scams. The proposed system’s performance was equal or superior to current best-of-breed systems. We further analyze current anti-phishing warnings from the perspective of warning theory, and propose a new warning design employing our Gestalt approach.
Teh-Chung Chen, Torin Stepan, Scott Dick, James Miller 0001
ACM Trans. Inf. Syst. Secur.4
2014 Extended Subtree: A New SimilarityFunction for Tree Structured Data
abstract
Although several distance or similarity functions for trees have been introduced, their performance is not always satisfactory in many applications, ranging from document clustering to natural language processing. This research proposes a new similarity function for trees, namely Extended Subtree (EST), where a new subtree mapping is proposed. EST generalizes the edit base distances by providing new rules for subtree mapping. Further, the new approach seeks to resolve the problems and limitations of previous approaches. Extensive evaluation frameworks are developed to evaluate the performance of the new approach against previous proposals. Clustering and classification case studies utilizing three real-world and one synthetic labeled data sets are performed to provide an unbiased evaluation where different distance functions are investigated. The experimental results demonstrate the superior performance of the proposed distance function. In addition, an empirical runtime analysis demonstrates that the new approach is one of the best tree distance functions in terms of runtime efficiency.
Ali Shahbazi, James Miller 0001
IEEE Trans. Knowl. Data Eng.2
2014 Automated cookie collection testing
abstract
Cookies are used by over 80% of Web applications utilizing dynamic Web application frameworks. Applications deploying cookies must be rigorously verified to ensure that the application is robust and secure. Given the intense time-to-market pressures faced by modern Web applications, testing strategies that are low cost and automatable are required. Automated Cookie Collection Testing (CCT) is presented, and is empirically demonstrated to be a low-cost and highly effective automated testing solution for modern Web applications. Automatable test oracles and evaluation metrics specifically designed for Web applications are presented, and are shown to be significant diagnostic tests. Automated CCT is shown to detect faults within five real-world Web applications. A case study of over 580 test results for a single application is presented demonstrating that automated CCT is an effective testing strategy. Moreover, CCT is found to detect security bugs in a Web application released into full production.
Andrew F. Tappenden, James Miller 0001
ACM Trans. Softw. Eng. Methodol.2
2013 Deficient documentation detection: a methodology to locate deficient project documentation using topic analysis
abstract
A project's documentation is the primary source of information for developers using that project. With hundreds of thousands of programming-related questions posted on programming Q&A websites, such as Stack Overflow, we question whether the developer-written documentation provides enough guidance for programmers. In this study, we wanted to know if there are any topics which are inadequately covered by the project documentation. We combined questions from Stack Overflow and documentation from the PHP and Python projects. Then, we applied topic analysis to this data using latent Dirichlet allocation (LDA), and found topics in Stack Overflow that did not overlap the project documentation. We successfully located topics that had deficient project documentation. We also found topics in need of tutorial documentation that were outside of the scope of the PHP or Python projects, such as MySQL and HTML.
Joshua Charles Campbell, Chenlei Zhang, Abram Hindle, James Miller 0001
MSR5
2013 Securing web-clients with instrumented code and dynamic runtime monitoring
Ejike Ofuonye, James Miller 0001
J. Syst. Softw.2
2013 Refactoring legacy AJAX applications to improve the efficiency of the data exchange component
Ming Ying 0003, James Miller 0001
J. Syst. Softw.2
2013 A Survey and Analysis of Current CAPTCHA Approaches
Narges Roshanbin, James Miller 0001
J. Web Eng.2
2013 Centroidal Voronoi Tessellations - A New Approach to Random Testing
abstract
Although Random Testing (RT) is low cost and straightforward, its effectiveness is not satisfactory. To increase the effectiveness of RT, researchers have developed Adaptive Random Testing (ART) and Quasi-Random Testing (QRT) methods which attempt to maximize the test case coverage of the input domain. This paper proposes the use of Centroidal Voronoi Tessellations (CVT) to address this problem. Accordingly, a test case generation method, namely, Random Border CVT (RBCVT), is proposed which can enhance the previous RT methods to improve their coverage of the input space. The generated test cases by the other methods act as the input to the RBCVT algorithm and the output is an improved set of test cases. Therefore, RBCVT is not an independent method and is considered as an add-on to the previous methods. An extensive simulation study and a mutant-based software testing investigation have been performed to demonstrate the effectiveness of RBCVT against the ART and QRT methods. Results from the experimental frameworks demonstrate that RBCVT outperforms previous methods. In addition, a novel search algorithm has been incorporated into RBCVT reducing the order of computational complexity of the new approach. To further analyze the RBCVT method, randomness analysis was undertaken demonstrating that RBCVT has the same characteristics as ART methods in this regard.
Ali Shahbazi, Andrew F. Tappenden, James Miller 0001
IEEE Trans. Software Eng.3
2012 Constructing high quality use case models: a systematic review of current practices
Mohamed El-Attar 0001, James Miller 0001
Requir. Eng.2
2011 Finding Homoglyphs - A Step towards Detecting Unicode-Based Visual Spoofing Attacks
Narges Roshanbin, James Miller 0001
WISE2
2010 An empirical investigation into open source web applications' implementation vulnerabilities
Toan Huynh, James Miller 0001
Empir. Softw. Eng.2
2010 Practical Elimination of External Interaction Vulnerabilities in Web Applications
James Miller 0001, Toan Huynh
J. Web Eng.1
2010 Investigating the Distributional Property of the Session Workload
James Miller 0001, Toan Huynh
J. Web Eng.1
2010 Developing comprehensive acceptance tests from use cases and robustness diagrams
Mohamed El-Attar 0001, James Miller 0001
Requir. Eng.2
2010 Improving the quality of use case models using antipatterns
Mohamed El-Attar 0001, James Miller 0001
Softw. Syst. Model.2
2010 Detecting visually similar Web pages: Application to phishing detection
abstract
We propose a novel approach for detecting visual similarity between two Web pages. The proposed approach applies Gestalt theory and considers a Web page as a single indivisible entity. The concept of supersignals, as a realization of Gestalt principles, supports our contention that Web pages must be treated as indivisible entities. We objectify, and directly compare, these indivisible supersignals using algorithmic complexity theory. We illustrate our approach by applying it to the problem of detecting phishing scams. Via a large-scale, real-world case study, we demonstrate that 1) our approach effectively detects similar Web pages; and 2) it accuractely distinguishes legitimate and phishing pages.
Teh-Chung Chen, Scott Dick, James Miller 0001
ACM Trans. Internet Techn.3
2009 A subject-based empirical evaluation of SSUCD's performance in reducing inconsistencies in use case models
Mohamed El-Attar 0001, James Miller 0001
Empir. Softw. Eng.2
2009 Another viewpoint on "evaluating web software reliability based on workload and failure data extracted from server logs"
Toan Huynh, James Miller 0001
Empir. Softw. Eng.2
2009 Do You Know Where Your Data Is? A Study of the Effect of Enforcement Strategies on Privacy Policies
abstract
Numerous countries around the world have enacted privacy-protection legislation, in an effort to protect their citizens and instill confidence in the valuable business-to-consumer E-commerce industry. These laws will be most effective if and when they establish a standard of practice that consumers can use as a guideline for the future behavior of e-commerce vendors. However, while privacy-protection laws share many similarities, the enforcement mechanisms supporting them vary hugely. Furthermore, it is unclear which (if any) of these mechanisms are effective in promoting a standard of practice that fits with the social norms of those countries. We present a large-scale empirical study of the role of legal enforcement in standardizing privacy protection on the Internet. Our study is based on an automated analysis of documents posted on the 100,000 most popular websites (as ranked by Alexa.com). We find that legal frameworks have had little success in creating standard practices for privacy-sensitive actions.
Ian Reay, Patricia Beatty, Scott Dick, James Miller 0001
Int. J. Inf. Secur. Priv.4
2009 Empirical observations on the session timeout threshold
Toan Huynh, James Miller 0001
Inf. Process. Manag.2
2009 An analysis of privacy signals on the World Wide Web: Past, present and future
Ian Reay, Scott Dick, James Miller 0001
Inf. Sci.3
2009 A Scalable Testing Framework for Location-Based Services
Andrew F. Tappenden, James Miller 0001, Michael R. Smith 0001
J. Comput. Sci. Technol.3
2009 A Survey of Cookie Technology Adoption Amongst Nations
Andrew F. Tappenden, James Miller 0001
J. Web Eng.2
2009 A Novel Evolutionary Approach for Adaptive Random Testing
abstract
Random testing is a low cost strategy that can be applied to a wide range of testing problems. While the cost and straightforward application of random testing are appealing, these benefits must be evaluated against the reduced effectiveness due to the generality of the approach. Recently, a number of novel techniques, coined Adaptive Random Testing, have sought to increase the effectiveness of random testing by attempting to maximize the testing coverage of the input domain. This paper presents the novel application of an evolutionary search algorithm to this problem. The results of an extensive simulation study are presented in which the evolutionary approach is compared against the Fixed Size Candidate Set (FSCS), Restricted Random Testing (RRT), quasi-random testing using the Sobol sequence (Sobol), and random testing (RT) methods. The evolutionary approach was found to be superior to FSCS, RRT, Sobol, and RT amongst block patterns, the arena in which FSCS, and RRT have demonstrated the most appreciable gains in testing effectiveness. The results among fault patterns with increased complexity were shown to be similar to those of FSCS, and RRT; and showed a modest improvement over Sobol, and RT. A comparison of the asymptotic and empirical runtimes of the evolutionary search algorithm, and the other testing approaches, was also considered, providing further evidence that the application of an evolutionary search algorithm is feasible, and within the same order of time complexity as the other adaptive random testing approaches.
Andrew F. Tappenden, James Miller 0001
IEEE Trans. Reliab.2
2009 A large-scale empirical study of P3P privacy policies: Stated actions vs. legal obligations
abstract
Numerous studies over the past ten years have shown that concern for personal privacy is a major impediment to the growth of e-commerce. These concerns are so serious that most if not all consumer watchdog groups have called for some form of privacy protection for Internet users. In response, many nations around the world, including all European Union nations, Canada, Japan, and Australia, have enacted national legislation establishing mandatory safeguards for personal privacy. However, recent evidence indicates that Web sites might not be adhering to the requirements of this legislation. The goal of this study is to examine the posted privacy policies of Web sites, and compare these statements to the legal mandates under which the Web sites operate. We harvested all available P3P (Platform for Privacy Preferences Protocol) documents from the 100,000 most popular Web sites (over 3,000 full policies, and another 3,000 compact policies). This allows us to undertake an automated analysis of adherence to legal mandates on Web sites that most impact the average Internet user. Our findings show that Web sites generally do not even claim to follow all the privacy-protection mandates in their legal jurisdiction (we do not examine actual practice, only posted policies). Furthermore, this general statement appears to be true for every jurisdiction with privacy laws and any significant number of P3P policies, including European Union nations, Canada, Australia, and Web sites in the USA Safe Harbor program.
Ian Reay, Scott Dick, James Miller 0001
ACM Trans. Web3
2009 Cookies: A deployment study and the testing implications
abstract
The results of an extensive investigation of cookie deployment amongst 100,000 Internet sites are presented. Cookie deployment is found to be approaching universal levels and hence there exists an associated need for relevant Web and software engineering processes, specifically testing strategies which actively consider cookies. The semi-automated investigation demonstrates that over two-thirds of the sites studied deploy cookies. The investigation specifically examines the use of first-party, third-party, sessional, and persistent cookies within Web-based applications, identifying the presence of a P3P policy and dynamic Web technologies as major predictors of cookie usage. The results are juxtaposed with the lack of testing strategies present in the literature. A number of real-world examples, including two case studies are presented, further accentuating the need for comprehensive testing strategies for Web-based applications. The use of antirandom test case generation is explored with respect to the testing issues discussed. Finally, a number of seeding vectors are presented, providing a basis for testing cookies within Web-based applications.
Andrew F. Tappenden, James Miller 0001
ACM Trans. Web2
2008 A Three-Tiered Testing Strategy for Cookies
abstract
Cookies, the HTTP state management mechanism, are the backbone of many web applications. Despite a high adoption rate, cookies have remained virtually unexplored by the academic community. This paper presents an EBNF grammatical definition and a three- tiered testing strategy for cookies. The testing strategy builds upon anti-random and grammar-based methodologies examining cookies from three perspectives: cookies collections, individual cookie transformations and application-specific test-case generation. The collection of cookies maintained within a user-agent are explored in light of the anti-random test- suite reduction techniques and the grammatical definition of a cookie, culminating in the definition of a number of seeding test-vectors providing the basis for a scalable test-suite. A number of distinct grammatically correct cookie transformations are presented, providing further scalability to the proposed testing strategy. Finally a discussion of application-specific cookie transformations is presented, with focus upon the security and reliability concerns of modern web applications.
Andrew F. Tappenden, James Miller 0001
ICST2
2008 E-RACE, A Hardware-Assisted Approach to Lockset-Based Data Race Detection for Embedded Products
abstract
Limited research exists for identifying data races under the specific characteristics found in embedded systems. E-RACE is a new style of data-race identification tool which directly utilizes specialized hardware capabilities to monitor the flow of data and instructions. Compared to existing data race analysis approaches, the hardware-assisted E-RACE tool has advantages of recognizing data-race issues without requiring extensive software code instrumentation. The tool is integrated into an Embedded Unit Testing Driven Development Framework to encourage the construction of testable code and early identification of data-races.
Lily Huang, Michael R. Smith 0001, Albert Tran, James Miller 0001
ISSRE4
2008 Resolving JavaScript Vulnerabilities in the Browser Runtime
abstract
The volume of Web based malware on the Internet keeps rising despite huge investments on Web security. JavaScript, the dominant scripting language for Web applications, is the primary channel for most of these attacks. In this paper, we describe research into the design and implementation of new Web client protection system based on code instrumentation techniques. This system combines traditional static analysis techniques with a dynamic HTML, CSS and JavaScript code runtime monitoring agent to offer an efficient, easily deployable, policy driven framework for improved user protection. Rewriting and runtime monitoring are based on providing safe equivalents of JavaScript code constructs known to contain in securities and hence exploitable by malicious Web applications. As a demonstration of the practical capabilities of our framework, we also include a case study attack and empirical analysis of some of its various aspects across 1000 home pages belonging to the most popular web sites on the Internet.
Ejike Ofuonye, James Miller 0001
ISSRE2
2008 A Hardware-Assisted Tool for Fast, Full Code Coverage Analysis
abstract
Software reliability can be improved by using code coverage analysis to ensure that all statements are executed at least once during the testing process. When full code coverage information is obtained through software code instrumentation, high runtime performance overheads are incurred. Techniques that perform deferred or selective code instrumentation have shown success in reducing run-time overheads; however, the execution profile remains distorted. Techniques have been proposed that use internal processor hardware during the data gathering process, e.g. program counter logging. These approaches have been shown to reduce overheads; but currently trade swift execution for sparse code coverage. By combining the branch-vector hardware designed for debugging modern embedded processors with on-demand code coverage analysis, we have developed a new tool which provides full code coverage, while minimizing performance distortions. Experimental results show a performance impact of only 8 - 12%, while still providing 100% code coverage information.
Albert Tran, Michael R. Smith 0001, James Miller 0001
ISSRE3
2008 Triangulation as a basis for knowledge discovery in software engineering
James Miller 0001
Empir. Softw. Eng.1
2008 On the possibilities of (pseudo-) software cloning from external interactions
Marek Z. Reformat, Xinwei Chai, James Miller 0001
Soft Comput.3
2008 Producing robust use case diagrams via reverse engineering of use case descriptions
Mohamed El-Attar 0001, James Miller 0001
Softw. Syst. Model.2
2007 A TDD approach to introducing students to embedded programming
abstract
Learning embedded programming is a highly demanding exercise. The beginner is bombarded with complexity from the start: embedded production based around a myriad of C++ constructs with low-level elements integrated onto ever more complicated processor architectures. The picture is further compounded by tool support having unfamiliar roles and appearances from previous student experiences. This demanding situation often has the student bewildered; seeking for "a crutch" or the simplest way forward regardless of the overall consequences. To control this potentially chaotic picture, the instructor needs to introduce devices to combat this complexity. We argue that test driven development (TDD) should become the instructor's principal weapon in this fight. Reasons for this belief combined with our, and the students', experiences with this novel approach are discussed.
James Miller 0001, Michael R. Smith 0001
ITiCSE1
2007 A practical approach to testing GUI systems
Toan Huynh, Marek Z. Reformat, James Miller 0001
Empir. Softw. Eng.4
2007 Empirical evaluation of optimization algorithms when used in goal-oriented automated test data generation techniques
Mohamed El-Attar 0001, Marek Z. Reformat, James Miller 0001
Empir. Softw. Eng.4
2007 A Survey and Analysis of the P3P Protocol's Agents, Adoption, Maintenance, and Future
abstract
In this paper, we survey the adoption of the platform for privacy preferences protocol (P3P) on Internet Web sites to determine if P3P is a growing or stagnant technology. We conducted a pilot survey in February 2005 and our full survey in November 2005. We compare the results from these two surveys and the previous (July 2003) survey of P3P adoption. In general, we find that P3P adoption is stagnant, and errors in P3P documents are a regular occurrence. In addition, very little maintenance of P3P policies is apparent. These observations call into question P3P's viability as an online privacy-enhancing technology. Our survey exceeds other previous surveys in our use of both detailed statistical analysis and scope; our February pilot survey analyzed more than 23,000 unique Web sites, and our full survey in November 2005 analyzed more than 100,000 unique Web sites.
Ian Reay, Patricia Beatty, Scott Dick, James Miller 0001
IEEE Trans. Dependable Secur. Comput.4
2006 Matching Antipatterns to Improve the Quality of Use Case Models
abstract
Use case modeling is an effective technique used to capture functional requirements. Use case models are mainly composed of textual descriptions written in natural language and simple diagrams that adhere to a few syntactic rules. This simplicity can be deceptive as many modelers create use case models that are incorrect, inconsistent, and ambiguous and contain restrictive design decisions. In this paper, a new methodology is described that utilizes antipatterns to detect potentially defective areas in use case models. This paper introduces the tool ARBIUM, which will support the proposed technique and aid analysts to improve the quality of their models. ARBIUM presents a framework that will allow developers to define their own antipatterns using OCL and textual descriptions. The proposed approach and tool are applied to a distributed biodiversity database use case model to demonstrate its feasibility. Our results indicate that they can improve the overall clarity and precision of use case models
Mohamed El-Attar 0001, James Miller 0001
RE2
2006 AGADUC: Towards a More Precise Presentation of Functional Requirement in Use Case Mod
abstract
Use case (UC) models describe functional requirements as a set of interactions between a software system and its environment. In essence, UC descriptions state a set of workflows that would allow a system's user to benefit from its services. It is critical that designers have a common and precise understanding of what these workflows are. Otherwise they are in danger of building the 'wrong' system. Traditionally, UC descriptions are authored using natural language, which as shown in this article, proves to be a poor vehicle, and insufficient, to describe the underlying workflows. Simply, the inherit ambiguity in natural language leads to misinterpretations and misunderstandings. UC diagrams do not provide any information about the dependencies between workflows spanning several UCs. In this paper, we present the process AGADUC, which systematically generate activity-like diagrams that represent the embedded workflows in the UC textual descriptions. A GADUC provides a great deal of information regarding how UCs are dependent on each other, without the need to iterate through several pages of UC descriptions. Using activity-like diagram ensures that all stakeholders have a precise and consistent understanding of the workflows. A case study conducted on a simplified Library case is presented and have shown that AGADUC overcomes many limitations in traditional UC models. The featured tool AREUCD automates the AGADUC process and it is demonstrated within the case study
Mohamed El-Attar 0001, James Miller 0001
SERA2
2006 Making Fit / FitNesse Appropriate for Biomedical Engineering Research
Michael R. Smith 0001, Adam Geras, James Miller 0001
XP4
2006 Extending the Embedded System E-TDDunit Test Driven Development Tool for Development of a Real Time Video Security System Prototype
Steven Daeninck, Michael R. Smith 0001, James Miller 0001, Linda Ko
XP3
2006 Configuring Hybrid Agile-Traditional Software Processes
Adam Geras, Michael R. Smith 0001, James Miller 0001
XP3
2006 Automatic test data generation using genetic algorithm and program dependence graphs
James Miller 0001, Marek Z. Reformat, Howard Zhang
Inf. Softw. Technol.1
2005 Testing the Semantics of W3C XML Schema
abstract
The XML schema language is becoming the preferred means of defining and validating highly structured XML instance documents. We have extended the conventional mutation method to be applicable for W3C XML schemas. In this paper a technique for using mutation analysis to test the semantic correctness of W3C XML schemas is presented. We introduce a mutation analysis model and a set of W3C XML schema (XSD) mutation operators that can be used to detect faults involving name-spaces, user-defined types, and inheritance. Preliminary evaluation of our technique shows that it is effectiveness to test the semantics of W3C XML schema documents.
Jian Bing Li, James Miller 0001
COMPSAC (1)2
2005 A Survey of Test Notations and Tools for Customer Testing
Adam Geras, James Miller 0001, Michael R. Smith 0001, James Love
XP2
2005 E-TDD - Embedded Test Driven Development a Tool for Hardware-Software Co-design Projects
Michael R. Smith 0001, Andrew K. C. Kwan, Alan Martin, James Miller 0001
XP4
2005 Agile Testing of Location Based Services
Andrew F. Tappenden, Adam Geras, Michael R. Smith 0001, James Miller 0001
XP5
2005 Replicating software engineering experiments: a poisoned chalice or the Holy Grail
James Miller 0001
Inf. Softw. Technol.1
2004 Self-assessment of performance in software inspection processes
Zhichao Yin, Alastair Dunsmore, James Miller 0001
Inf. Softw. Technol.3
2004 Statistical significance testing--a panacea for software technology experiments?
James Miller 0001
J. Syst. Softw.1
2004 Multi-Project Management in Software Engineering Using Simulation Modelling
Bengee Lee, James Miller 0001
Softw. Qual. J.2
2004 A Cognitive-Based Mechanism for Constructing Software Inspection Teams
abstract
Software inspection is well-known as an effective means of defect detection. Nevertheless, recent research has suggested that the technique requires further development to optimize the inspection process. As the process is inherently group-based, one approach to improving performance is to attempt to minimize the commonality within the process and the group. This work proposes an approach to add diversity into the process by using a cognitively-based team selection mechanism. The paper argues that a team with diverse information processing strategies, as defined by the selection mechanism, maximize the number of different defects discovered.
James Miller 0001, Zhichao Yin
IEEE Trans. Software Eng.1
2003 Experiments in Automatic Programming for General Purposes
abstract
Although the generation and application of software clones is relatively unexplored, it is believed that this is a fundamental technology that can have many different applications within a software engineering environment. For example, software clones could be used in software fault tolerance. Clearly, for these clones to be usable, their production needs to be automated. An interesting approach to this automatic production or generation problem is the application of evolutionary-based genetic programming (GP). Using the paradigms of best fit, selection, crossover and mutation a number of clones, satisfying specific requirements, can be automatically generated. In general, GP is a flexible and powerful algorithm suitable for solving variety of different problems. The paper presents the results of studies that have been conducted in order to answer questions related to feasibility of using GP for clone generation: what features of GP are important? What works and what does not? How GP can be "tuned" for the problem? The results have been used to draw a set of suggestions and conclusions that indicate possible usability of GP-based approach to automatic generation of clones.
Marek Z. Reformat, Xinwei Chai, James Miller 0001
ICTAI3
2003 Practical assessment of the models for identification of defect-prone classes in object-oriented commercial systems using design metrics
Giancarlo Succi, Witold Pedrycz, Milorad Stefanovic, James Miller 0001
J. Syst. Softw.4
2000 Applying meta-analytical procedures to software engineering experiments
James Miller 0001
J. Syst. Softw.1
1999 A Comparison of Computer Support Systems for Software Inspection
Fraser MacDonald, James Miller 0001
Autom. Softw. Eng.2
1999 Multi-method research: An empirical investigation of object-oriented technology
Murray Wood, John W. Daly, James Miller 0001, Marc Roper
J. Syst. Softw.3
1999 Estimating the Number of Remaining Defects after Inspection
abstract
An essential component of all software inspection processes is a well-founded decision about continuing or stopping the current process. This decision should be based upon directly relevant quantitative information — the number of defects remaining in the artefact. This quantity can be estimated by the use of capture–recapture methods. Several software engineering papers have explored this topic, but the question of which capture–recapture technique is best still remains unresolved. This paper attempts to shed further light upon this question. After reviewing the relevant capture–recapture models and previous evaluations within a software engineering context, the paper proceeds to evaluate the models by using data collected from subject-based experiments on software inspection. The experiments used artefacts where the number of defects was known and hence a direct measure of the accuracy of the various capture–recapture techniques was available. The paper reports that most heterogeneity models show, in general, superior performance — especially the Jackknife estimator. However, the paper concludes that further work is required to correct the limitations of the current models if reliable estimates are to be achieved. Copyright © 1999 John Wiley & Sons, Ltd.
James Miller 0001
Softw. Test. Verification Reliab.1
1998 ASSISTing Exit Decisions in Software Inspection
abstract
Software inspection is a valuable technique for detecting defects in the products of software development. One avenue of research within inspection concerns the development of computer support. It is hoped that such support will provide even greater benefits when applying inspection. A number of prototype systems have been developed by researchers, yet these suffer from some fundamental limitations. One of the most serious of these concerns is the lack of facilities to monitor the process, and to provide the moderator with quantitative information on the performance of the process. The paper begins by briefly outlining the measurement component that has been introduced into the system, and especially discusses an approach to estimating the effectiveness of the current process.
James Miller 0001, Fraser MacDonald
ASE1
1998 Further Experiences with Scenarios and Checklists
James Miller 0001, Murray Wood, Marc Roper
Empir. Softw. Eng.1
1997 Statistical power and its subcomponents - missing and misunderstood concepts in empirical software engineering research
James Miller 0001, John W. Daly, Murray Wood, Marc Roper, Andrew Brooks
Inf. Softw. Technol.1
1997 An empirical evaluation of defect detection techniques
Marc Roper, Murray Wood, James Miller 0001
Inf. Softw. Technol.3
1997 A Software Inspection Process Definition Language and Prototype Support Tool
abstract
Software inspection is a widely used method for finding defects in all types of software development documents. Many process variations exist, each designed for use under certain circumstances or to address some perceived deficiency in other methods. A desirable attribute of inspection is rigour, allowing the use of historical data to predict future performance and to suggest process improvements. Recent work in tool support for inspection is designed to tackle the issue of enforcing rigorous inspection, but these tools concentrate on enforcing a single, usually proprietary, method. This paper investigates existing inspection methods and derives a generic inspection process which can be used to describe any of these methods. This process is then used to determine a notation for describing any inspection process, which can consequently be used as input to an inspection support tool, allowing the support of any inspection method. The paper also demonstrates a system which uses the language to provide support for multiple inspection processes, and describes other desirable features of such a tool. © 1997 by John Wiley & Sons, Ltd.
Fraser MacDonald, James Miller 0001
Softw. Test. Verification Reliab.2
1996 Automating the Software Inspection Process
Fraser MacDonald, James Miller 0001, Andrew Brooks, Marc Roper, Murray Wood
Autom. Softw. Eng.2
1996 Evaluating inheritance depth on the maintainability of object-oriented software
John W. Daly, Andrew Brooks, James Miller 0001, Marc Roper, Murray Wood
Empir. Softw. Eng.3
1996 Applying Inspection to Object-Oriented Code
abstract
The benefits of the object-oriented paradigm are widely cited. At the same time, inspection is deemed to be the most cost-effective means of detecting defects in software products. Why then, is there no published experience, let alone quantitative data, on the application of inspection to object-oriented systems? This paper describes the facilities of the object-oriented paradigm and the issues that these raise when inspecting object-oriented code. Several problems are caused by the disparity between the static code structure and its dynamic runtime behaviour. The large number of small methods in object-oriented systems can also cause problems. This is followed by a description of three areas which may help mitigate problems found. Firstly, the use of various programming techniques may assist in making object-oriented code easier to inspect. Secondly, improved program documentation can help the inspector understand the code under inspection. Finally, tool support can help the inspector analyse the dynamic behaviour of the code. In conclusion, while both the object-oriented paradigm and inspection provide excellent benefits on their own, combining the two may be a difficult exercise, requiring extensive support if it is to be successful.
Fraser MacDonald, James Miller 0001, Andrew Brooks, Marc Roper, Murray Wood
Softw. Test. Verification Reliab.2
1995 A Survey of Experiences amongst Object-Oriented Practitioners
abstract
The object-oriented paradigm is becoming increasingly popular as a result of expert opinion and anecdotal evidence and not on the basis of sound empirical data. The questionnaire survey was undertaken as part of a programme of research to validate unsupported claims about the paradigm. The questionnaire follows structured interviews of experienced object-oriented developers with the intention of confirming the findings on a wider practitioner group. It was posted to relevant electronic newsgroups and to members of an object-oriented (postal) mailing list. The survey received 167 responses to the electronic questionnaire and 119 responses (30% response rate) to the postal version. Results show that respondents are of the view that: (i) the object-oriented paradigm has advantages over other paradigms in terms of ease of analysis and design, programmer productivity, software reuse, and ease of maintenance; (ii) inheritance can introduce difficulties when trying to understand object-oriented software; (iii) missing design documentation and poor or inappropriate design are prevalent problems; (iv) maintenance causes degradation of object-oriented software, but less frequently than conventional software; (v) C++ has many deficiencies in comparison to other purer object-oriented languages.
John W. Daly, James Miller 0001, Andrew Brooks, Marc Roper, Murray Wood
APSEC2
1995 The effect of inheritance on the maintainability of object-oriented software: an empirical study
abstract
The empirical study was undertaken as part of a programme of research to explore unsupported claims about the object-oriented paradigm: a series of experiments tested the effect of inheritance on the maintainability of object-oriented software. Subjects were asked to modify object-oriented software with a hierarchy of 3 levels of inheritance depth and equivalent object-based software with no inheritance. The collected timing data showed that subjects maintaining object-oriented software using inheritance performed the modification tasks, on average, approximately 20% quicker than those maintaining equivalent object-based software with no inheritance. An initial inductive analysis revealed that 2 out of 3 subjects performed faster when maintaining the object-oriented software with inheritance. The findings are sufficiently important that attempts to verify the results should be made by independent researchers. Subsequent studies should seek to scale up the findings to the maintenance of more complex software by professional programmers.
John W. Daly, Andrew Brooks, James Miller 0001, Marc Roper, Murray Wood
ICSM3
1995 Towards a benchmark for the evaluation of software testing techniques
James Miller 0001, Marc Roper, Murray Wood, Andrew Brooks
Inf. Softw. Technol.1
1994 Applying object-oriented construction to fault tolerant systems
abstract
This paper investigates the application of object-oriented construction to fault tolerant systems. The resulting system provides traditional fault tolerance within objects, but also a new form of fault tolerance between objects: object diversity. Object diversity extends current practice by integrating diversity in two directions: data and algorithm. This resulting form will allow increased diversity to be incorporated within fault tolerant systems. Further benefits are derived from the use of the inheritance hierarchy as a natural source of redundant components. As class libraries (both general and application specific) grow, more and more "free" redundant components will become available, yielding increasing savings on production costs.>
James Miller 0001, Murray Wood, Andrew Brooks, Marc Roper
APSEC1
1994 Verification of Results in Software Maintenance Through External Replication
abstract
Empirical studies carried out to help understand the problems of software maintenance are widely held to be of value. A view perhaps less widely recognised within the software engineering domain is that experiments should be replicated both internally and externally to validate the results and build up a cohesive body of knowledge. This paper presents the external replication findings of an experiment which tested the benefits to maintenance of using modular code against nonmodular (monolithic) code. The results of our replication were strikingly different from those of the original which showed that a modular program could be maintained significantly faster than an equivalent monolithic version. An inductive analysis, undertaken to investigate the reasons for this, uncovered evidence of an ability effect and the suggestion that the experiment may have been too artificial.>
John W. Daly, Andrew Brooks, James Miller 0001, Marc Roper, Murray Wood
ICSM3