Miriam A. M. Capretz

dblp:c/MiriamAMCapretz · also Miriam Akemi Manabe Capretz · DBLP profile ↗
← Back
55ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-1380-971XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 16 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 2 since 2021Computer networks · 6 · 1 since 2021Systems, architecture and hardware · 5 · 2 since 2021Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 3Security and privacy · 1
YearPublicationVenuePosition
2025 Robust Smartphone Screen Integration with Deep Learning for Virtual Reality Pass-through
abstract
Virtual reality is revolutionizing immersive experiences, yet the seamless integration of everyday devices, such as smartphones, remains a challenging frontier. This paper presents a method for dynamically integrating a live smartphone screen into the virtual reality pass-through view. Our approach leverages a lightweight convolutional neural network (CNN) to detect the smartphone in real time, accurately determining its position, scale, and orientation within the camera feed. By synchronizing live screen capture with inertial sensor data, our system computes precise affine transformations that ensure the overlay remains perfectly aligned with its physical counterpart, even during rapid movements. Experimental evaluations demonstrate that our solution achieves an average of 29.2 detection runs per second, delivering a stable and high-fidelity integration without the need for additional hardware. Evaluations also showed a 0.44 increase in detection F1 score during live comparisons to the baseline alternative, which does not use deep learning. This dynamic overlay enhances visual clarity and interaction in virtual reality and bridges the gap between virtual and real-world interfaces, empowering users to access notifications, messages, and productivity applications seamlessly.
Lucas Hartman, Ethan Pigou, Nicholas Strzelczyk, Santiago Gomez-Rosero, Miriam A. M. Capretz
COMPSAC5
2025 Synthetic Fouling Image Data Generation for Heat Exchanger Predictive Maintenance
abstract
Regular maintenance of heating, ventilation, and air conditioning (HVAC) systems is vital for building sustainability and energy efficiency, as neglect can lead to performance degradation, increased energy consumption, and reduced equipment lifespan. Fouling, the accumulation of unwanted material on surfaces, is a significant issue in HVAC heat exchangers, leading to efficiency losses and increased operational costs. While preventive maintenance is commonly employed, predictive maintenance offers a more effective approach, particularly when detecting fouling that requires visual data. However, the rarity of anomalies in such data complicates the training of deep learning models. This paper introduces an approach for synthetic fouling data generation for predictive maintenance designed to simulate fouling on a heat exchanger in a cooling tower. A synthetic image dataset is generated with two distinct types: continuous growth scenarios for training purposes and scheduled maintenance scenarios for evaluating model performance. Resulting in a collection of 4,260 images over a 60-day period. The effectiveness of the dataset is validated through experiments where a U-Net model is trained for semantic segmentation in fouling detection, achieving an average F1 score of 0.8595 across the test scenarios. These results confirm that the synthetic dataset is effective for training convolutional neural networks for fouling detection, laying the groundwork for integrating such models into a broader predictive maintenance strategy for cooling towers.
Nicholas Strzelczyk, Santiago Gomez-Rosero, Amanda O. Timotheo, Leonardo de M. Honório, Miriam A. M. Capretz
COMPSAC5
2025 Prospect utility with hyperbolic tangent function
abstract
One branch of safety in reinforcement learning is through integrating risk sensitivity within the Markov Decision Process framework. The objective is to mitigate low-probability events that could lead to severe negative outcomes. Eliminating such risky events is usually done by incorporating a utility function on the expected return; therefore, reshaping the reward structures according to the risk levels associated with different outcomes. The temporal difference learning algorithm can be modified with a utility to capture risk. Notably, such utility functions are either convex or concave depending on the desired risk behavior. Given the outcome space and depending on the risk-sensitivity mode, concave utilities may promote risk-averse behavior and convex utilities may encourage risk-seeking strategies. Such function structure is demonstrated in Prospect Theory, and this motivates a novel formulation using the hyperbolic tangent function called PTanh. Using PTanh, experiments are performed to assess the effect of the diminishing marginal property on the risk-averse policies. It is concluded that there is a correlation between the marginal and selecting the risk-averse parameters. The marginals influence the effectiveness of the averse policies. When the marginals are considered, PTanh can demonstrate better results in terms of a ratio of average reward per prohibited state rate. Furthermore, using empirical evidence, the policy experiments shown with PTanh generalize to other utilities of the Prospect Shape. • Introduces PTanh utility function based on hyperbolic tangent for risk-sensitive RL. • Analyzes impact of diminishing marginals on risk-averse policy effectiveness. • Compares PTanh integrated with Q-learning and Expected SARSA against their base algorithms. • Demonstrates correlation between marginal effects and risk-averse policy parameters.
Patrick Adjei, Santiago Gomez-Rosero, Miriam A. M. Capretz
Int. J. Approx. Reason.3
2023 Multi-Factor Edge-Weighting with Reinforcement Learning for Load Balancing of Electric Vehicle Charging Stations
abstract
The number of electric vehicle (EV) owners continues to grow at a rate that makes it increasingly difficult to avoid overloading the capacity of charging stations. Existing solutions to balance this load primarily focus on solving parts of the problem and fail to consider a more holistic approach that can solve the issue of balancing the load across multiple stations by handling EV routing. This paper proposes a novel solution that routes EVs while considering both travel times and peak loads at charging stations. The approach integrates a reinforcement learning algorithm with a path-finding algorithm and a custom-built simulation environment. Compared with baseline methods, the solution showed improved performance in minimizing peak power loads across charging stations while increasing individual trip duration by less than 5%. This approach has the potential to significantly improve the efficiency of EV charging by reducing peak loads at charging stations.
Lucas Hartman, Santiago Gomez-Rosero, Patrick Adjei, Miriam A. M. Capretz
ICMLA4
2023 Scheduling Electric Vehicle Charging for Grid Load Balancing
abstract
In recent years, electric vehicles (EVs) have been widely adopted because of their environmental benefits. However, the increasing volume of EVs poses capacity issues for grid operators as simultaneously charging many EVs may result in grid instabilities. Scheduling EV charging for grid load balancing has a potential to prevent load peaks caused by simultaneous EV charging and contribute to balance of supply and demand. This paper proposes a user-preference-based scheduling approach to minimize costs for the user while balancing grid loads. The EV owners benefit by charging when the electricity cost is lower, but still within the user-defined preferred charging periods. On the other hand, the approach reduces the pressure on the grid by balancing the grid load. Two methods, the greedy algorithm and nonlinear programming, are considered along with users' charging preferences and durations. For scheduling small numbers of charging activities, the nonlinear programming model achieves better load balancing than the greedy algorithm; however, for scheduling medium to large numbers of charging activities, the greedy algorithm has a clear advantage in terms of time complexity.
Zhixin Han, Katarina Grolinger, Miriam A. M. Capretz, Syed Mir
IECON3
2023 Dynamic Successor Features for transfer learning and guided exploration
abstract
The Successor Feature framework for Reinforcement Learning algorithms improves task transfer by decomposing the learned state–action value function. The decomposition involves two components, one that captures future-expected state features and the other that models the task-related reward structure. However, successful transfer between tasks depends heavily on how the reward function changes, possibly leading to failure of the original Successor Feature formulation. This paper proposes the Dynamic Successor Feature framework, DynSF, by extending the mathematical formulation of the original Successor Feature framework to center around a learned state-transition model. Under this formulation, the state-transition model dynamically induces the acting policy. The flexibility of DynSF also extends to the architecture, requiring only a state-transition model and a small vector of parameters. This architecture provides immense flexibility in the choice of the model used to learn the state-transition model. The DynSF framework is evaluated and compared to other baseline algorithms through several experiments in a continuous grid world environment, a robotic Reacher, and pixels in the Doom environment.
Norman L. Tasfi, Eder Santana, Luisa H. B. Liboni, Miriam A. M. Capretz
Knowl. Based Syst.4
2022 An edge-fog architecture for distributed 3D reconstruction
Vinicius F. Vidal, Leonardo de M. Honório, Milena F. Pinto, Mario A. R. Dantas, Maria J. R. Aguiar, Miriam A. M. Capretz
Future Gener. Comput. Syst.6
2021 Dynamic Planning Networks
abstract
We introduce Dynamic Planning Networks (DPN), a novel architecture for deep reinforcement learning, that combines model-based and model-free aspects for online planning. Our architecture learns to dynamically construct plans using a learned state-transition model by selecting and traversing between simulated states and actions to maximize information before acting. DPN learns to efficiently form plans by expanding a single action-conditional state transition at a time instead of exhaustively evaluating each action, reducing the number of state-transitions used during planning. We observe emergent planning patterns in our agent, including classical search methods such as breadth-first and depth-first search. DPN shows improved performance over existing baselines across multiple axes.
Norman L. Tasfi, Miriam A. M. Capretz
IJCNN2
2021 Blockchain for Collaborative Businesses
abstract
Abstract Blockchain applications have continuously improved ever since its first debut on cryptocurrency. From then on, its uses have branched out from the financial realm, finding their way into numerous industries such as health, environmental, and governmental. Businesses are starting to take advantage of the intrinsic traits that made blockchain so notorious into their operations, such as security, integrity, and transparency. Blockchain’s versatility allows companies to cooperate on a shielded environment with business partners safely. This paper details how permissioned blockchain networks can accommodate collaborative business models securely to provide thriving business alliances. Examples of cooperative business models and business relationship orientation are described here, as well as how they generate value when paired with permissioned blockchain networks - a more business-oriented variety of blockchain. To support this study’s endeavors, business use cases are presented to highlight how simple it is to put in place a permissioned blockchain to businesses to achieve tighter bonds with business partners. The use cases contain particular goals that enterprises seek to accomplish by partnering up with other companies, as well as how, with the employment of blockchain, they can attain them.
Augusto R. C. Bedin, Miriam A. M. Capretz, Syed Mir
Mob. Networks Appl.2
2020 Predicting Residential Energy Consumption Using Wavelet Decomposition With Deep Neural Network
abstract
Electricity consumption is accelerating due to economic and population growth. Hence, energy consumption prediction is becoming vital for overall consumption management and infrastructure planning. Recent advances in smart electric meter technology are making high-resolution energy consumption data available. However, many parameters influencing energy consumption are not typically monitored for residential buildings. Therefore, this study's main objective is to develop a data-driven energy consumption forecasting model (next-hour consumption) for residential houses solely based on analyzing electricity consumption data. This research proposes a deep neural network architecture that combines stationary wavelet transform features and convolutional neural networks. The proposed approach utilizes automatically extracted features from smart-meter readings by applying wavelet decomposition, convolution, and pooling operations. This study's findings have demonstrated the advantage of integrating wavelet features with convolutional neural networks to improve forecasting accuracy while automating feature extraction.
Dagimawi D. Eneyew, Miriam A. M. Capretz, Girma T. Bitsuamlak, Syed Mir
ICMLA2
2020 Deep neural network for load forecasting centred on architecture evolution
abstract
Nowadays, electricity demand forecasting is critical for electric utility companies. Accurate residential load forecasting plays an essential role as an individual component for integrated areas such as neighborhood load consumption. Short-term load forecasting can help electric utility companies reduce waste because electric power is expensive to store. This paper proposes a novel method to evolve deep neural networks for time series forecasting applied to residential load forecasting. The approach centres its efforts on the neural network architecture during the evolution. Then, the model weights are adjusted using an evolutionary optimization technique to tune the model performance automatically. Experimental results on a large dataset containing hourly load consumption of a residence in London, Ontario shows that the performance of unadjusted weights architecture is comparable to other state-of-the-art approaches. Furthermore, when the architecture weights are adjusted the model accuracy surpassed the state-of-the-art method called LSTM one shot by 3.0%.
Santiago Gomez-Rosero, Miriam A. M. Capretz, Syed Mir
ICMLA2
2020 Noisy Importance Sampling Actor-Critic: An Off-Policy Actor-Critic With Experience Replay
abstract
This paper presents Noisy Importance Sampling Actor-Critic (NISAC), a set of empirically validated modifications to the advantage actor-critic algorithm (A2C), allowing off-policy reinforcement learning and increased performance. NISAC uses additive action space noise, aggressive truncation of importance sample weights, and large batchsizes. We see that additive noise drastically changes how off-sample experience is weighted for policy updates. The modified algorithm achieves an increase in convergence speed and sample efficiency compared to both the on-policy actor-critic A2C and the importance weighted off-policy actor-critic algorithm. In comparison to state-of-the-art (SOTA) methods, such as actor-critic with experience replay (ACER), NISAC nears the performance on several of the tested environments while training 40% faster and being significantly easier to implement. The effectiveness of NISAC is demonstrated against existing on-policy and off-policy actor-critic algorithms on a subset of the Atari domain.
Norman L. Tasfi, Miriam A. M. Capretz
IJCNN2
2019 PLDSD: Personalized Linked Data Semantic Distance for LOD-Based Recommender Systems
abstract
A vast amount of data that can be easily read by machines have been published in freely accessible and interconnected datasets, creating the so-called Linked Open Data cloud. This phenomenon has opened opportunities for the development of semantic applications, including recommender systems. In this paper, we propose Personalized Linked Data Semantic Distance (PLDSD), a novel similarity measure for linked data that personalizes the RDF graph by adding weights to the edges, based on previous user's choices. Thus, our approach has the purpose of minimizing the sparsity problem by ranking the best features for a particular user, and also, of solving the item cold-start problem, since the feature ranking task is based on features shared between old items and the new item. We evaluate PLDSD in the context of a LOD-based Recommender System using mixed data from DBpedia and MovieLens, and the experimental results indicate better accuracy of recommendations compared to a non-personalized baseline similarity method.
Gabriela Oliveira Mota Da Silva, Frederico Araújo Durão, Miriam A. M. Capretz
iiWAS3
2019 Assets Predictive Maintenance Using Convolutional Neural Networks
abstract
Predictive Maintenance (PdM) performs maintenance based on the asset's health status indicators. Sensors can measure an unusual pattern of these indicators, such as an increased motor's vibration level or higher energy consumption, and, in most cases, failures are preceded by an unusual pattern of these measurements. Convolutional Neural Network (CNN) is a Machine Learning technique capable of extracting data representation. This paper presents a CNN framework to tackle assets predictive maintenance problem and a method to transform 1-dimensional (1-D) data into an image-like representation (2-D). A data transformation step is very important to make the use of CNN feasible. To evaluate the proposed framework two datasets were obtained from fans, with distinct electrical pattern, from a building at Western University. The data was preprocessed, transformed in a image-like representation and fed to a tuned classifier. The results presented by the CNN-PdM framework showed that the combination of CNN with the proposed data transformation method outperformed traditional machine learning techniques (Random Forest, Support Vector Machine, and Multi-Layer Perceptron). The model created by the CNN-PdM framework achieved accuracy rates as high as 98% for one of the datasets and 95% for the other.
Willamos Silva, Miriam A. M. Capretz
SNPD2
2019 An approach for SDN traffic monitoring based on big data techniques
abstract
Software-defined networking overcomes the limitations of traditional networks by splitting the control plane from the data plane. The logic of the network is moved to a component called the controller that manages devices in the data plane. To implement this architecture, it has become the norm to use the OpenFlow (OF) protocol, which defines several counters maintained by network devices. These counters are the starting point for Traffic Engineering (TE) activities. TE monitors several network parameters, including network bandwidth utilization. A great challenge for TE is to collect and generate statistics about bandwidth utilization for monitoring and traffic analysis activities. This becomes even more challenging if fine-grained monitoring is required. Network management tasks such as network provisioning, capacity planning, load balancing, and anomaly detection can benefit from this fine-grained monitoring. Because the counters are updated for every packet that crosses the switch, they must be retrieved in a streaming fashion. This scenario suggests the use of Big Data streaming techniques to collect and process counter values. Therefore, this paper proposes an approach based on a fine-grained Big Data monitoring method to collect and generate traffic statistics using counter values. This research work can significantly leverage TE. The approach can provide a more detailed view of network resource utilization because it can deliver individual and aggregated statistical analyses of bandwidth consumption. Experimental results show the effectiveness of the proposed method.
Wander Queiroz, Miriam A. M. Capretz, Mario A. R. Dantas
J. Netw. Comput. Appl.2
2018 Forecasting Residential Energy Consumption: Single Household Perspective
abstract
With the development of smart electricity metering technologies, huge amounts of consumption data can be retrieved on a daily and hourly basis. Energy consumption forecasting facilitates electricity demand management and utilities load planning. Most studies have been focussed on commercial customers or residential building-level energy consumption, or have used behavioral and occupancy sensor data to characterize an individual household's electrical consumption. This study has analyzed energy consumption at single household level using smart meter data to improve residential energy services and gain insights into planning demand response programs. Electricity consumption for anonymous individual households has been predicted using a Support Vector Regression (SVR) modelling with both daily and hourly data granularity. The electricity usage data set for 2014 to 2016 was obtained from a Canadian utility company. Exploratory data analysis (EDA) was used for data visualization and feature selection. The analysis presented here demonstrates that forecasting residential energy consumption for individual households is feasible, but the accuracy is highly dependable on household behaviour variability.
Xiaoou Monica Zhang, Katarina Grolinger, Miriam A. M. Capretz, Luke Seewald
ICMLA3
2018 Evaluating Iris Scanning Technology to Link Data Related to Homelessness and Other Disadvantaged Populations with Mental Illness and Addiction
Cheryl Forchuk, Lorie Donelle, Miriam A. M. Capretz, Fatima Bukair, John Kok
ICOST3
2018 Evaluation of Cache for Bandwidth Optimization in ICN Through Software-Defined Networks
abstract
Traffic reduction in network segments through cache implementations has become a major research topic due to the exponential increase in data requests through the network. Even with high-speed connections, the conventional model still depends on point-to-point communication between two systems. Throughout the world, more connected devices are, accessing services and obtaining information. To support this activity, servers must have massive storage to support creation, retrieval, updated and deletion of large amounts of data. Therefore, in studies of Information Centric Networks (ICN), this model has been widely discussed as the new content distribution model for the Internet. To provide improved network management many approaches are using software-defined networks (SDN) to develop flexible content-based networks. This paper proposes to use cache replication for ICN through SDN to avoid duplicated requests in the same connection. The redundant cache reduced the bandwidth consumed by duplicated requests, from a maximum of 3.20 Gbps to 2.07 Gbps, reducing the bandwidth consumption by 11.3%.
Erick Nascimento 0001, Douglas Dyllon Jeronimo de Macedo, Edward D. Moreno, Luis C. E. Bona, Miriam A. M. Capretz
ISCC5
2018 (CF)2 architecture: contextual collaborative filtering
abstract
Recommender systems have dramatically changed the way we consume content. Internet applications rely on these systems to help users navigate among the ever-increasing number of choices available. However, most current systems ignore the fact that user preferences can change according to context, resulting in recommendations that do not fit user interests. This research addresses these issues by proposing the $$({ CF})^2$$ architecture, which uses local learning techniques to embed contextual awareness into collaborative filtering models. The proposed architecture is demonstrated on two large-scale case studies involving over 130 million and over 7 million unique samples, respectively. Results show that contextual models trained with a small fraction of the data provided similar accuracy to collaborative filtering models trained with the complete dataset. Moreover, the impact of taking into account context in real-world datasets has been demonstrated by higher accuracy of context-based models in comparison to random selection models.
Dennis Bachmann, Katarina Grolinger, Hany F. ElYamany, Wilson A. Higashino, Miriam A. M. Capretz, Majid Fekri, Bala Gopalakrishnan
Inf. Retr. J.5
2016 Collective contextual anomaly detection framework for smart buildings
abstract
Buildings are responsible for a significant amount of total global energy consumption and as a result account for a substantial portion of overall carbon emissions. Moreover, buildings have a great potential for helping to meet energy efficiency targets. Hence, energy saving goals that target buildings can have a significant contribution in reducing environmental impact. Today's smart buildings achieve energy efficiency by monitoring energy usage with the aim of detecting and diagnosing abnormal energy consumption behaviour. This research proposes a generic collective contextual anomaly detection (CCAD) framework that uses sliding window approach and integrates historic sensor data along with generated and contextual features to train an autoencoder to recognize normal consumption patterns. Subsequently, by determining a threshold that optimizes sensitivity and specificity, the framework identifies abnormal consumption behaviour. The research compares two models trained with different features using real-world data provided by Powersmiths, located in Brampton, Ontario, Canada.
Daniel B. Araya, Katarina Grolinger, Hany F. ElYamany, Miriam A. M. Capretz, Girma T. Bitsuamlak
IJCNN4
2016 CEPSim: Modelling and simulation of Complex Event Processing systems in cloud environments
abstract
The emergence of Big Data has had profound impacts on how data are stored and processed. As technologies created to process continuous streams of data with low latency, Complex Event Processing (CEP) and Stream Processing (SP) have often been related to the Big Data velocity dimension and used in this context. Many modern CEP and SP systems leverage cloud environments to provide the low latency and scalability required by Big Data applications, yet validating these systems at the required scale is a research problem per se . Cloud computing simulators have been used as a tool to facilitate reproducible and repeatable experiments in clouds. Nevertheless, existing simulators are mostly based on simple application and simulation models that are not appropriate for CEP or for SP. This article presents CEPSim , a simulator for CEP and SP systems in cloud environments. CEPSim proposes a query model based on Directed Acyclic Graphs (DAGs) and introduces a simulation algorithm based on a novel abstraction called event sets. CEPSim is highly customizable and can be used to analyse the performance and scalability of user-defined queries and to evaluate the effects of various query processing strategies. Experimental results show that CEPSim can simulate existing systems in large Big Data scenarios with accuracy and precision.
Wilson A. Higashino, Miriam A. M. Capretz, Luiz Fernando Bittencourt
Future Gener. Comput. Syst.2
2016 Attributed Graph Rewriting for Complex Event Processing Self-Management
abstract
The use of Complex Event Processing (CEP) and Stream Processing (SP) systems to process high-volume, high-velocity Big Data has renewed interest in procedures for managing these systems. In particular, self-management and adaptation of runtime platforms have been common research themes, as most of these systems run under dynamic conditions. Nevertheless, the research landscape in this area is still young and fragmented. Most research is performed in the context of specific systems, and it is difficult to generalize the results obtained to other contexts. To enable generic and reusable CEP/SP system management procedures and self-management policies, this research introduces the Attributed Graph Rewriting for Complex Event Processing Management ( AGeCEP ) formalism. AGeCEP represents queries in a language- and technology-agnostic fashion using attributed graphs. Query reconfiguration capabilities are expressed through standardized attributes, which are defined based on a novel classification of CEP query operators. By leveraging this representation, AGeCEP also proposes graph rewriting rules to define consistent reconfigurations of queries. To demonstrate AGeCEP feasibility, this research has used it to design an autonomic manager and to define a selected set of self-management policies. Finally, experiments demonstrate that AGeCEP can indeed be used to develop algorithms that can be integrated into diverse CEP systems.
Wilson A. Higashino, Cédric Eichler, Miriam A. M. Capretz, Luiz Fernando Bittencourt, Thierry Monteil 0001
ACM Trans. Auton. Adapt. Syst.3
2015 A Generalized Service Replication Process in Distributed Environments
abstract
Paper presented May 2015. Also published in Proceedings of the 5th International Conference on Cloud Computing and Services Science, pages 186-193; ISBN 978-989-758-104-5.
Hany F. ElYamany, Marwa F. Mohamed, Katarina Grolinger, Miriam A. M. Capretz
CLOSER4
2015 Predicting Energy Demand Peak Using M5 Model Trees
abstract
Predicting energy demand peak is a key factor for reducing energy demand and electricity bills for commercial customers. Features influencing energy demand are many and complex, such as occupant behaviours and temperature. Feature selection can decrease prediction model complexity without sacrificing performance. In this paper, features were selected based on their multiple linear regression correlation coefficients. This paper discusses the capabilities of M5 model trees in energy demand prediction for commercial buildings. M5 model trees are similar to regression trees, however they are more suitable for continuous prediction problems. The M5 model tree prediction was developed based on a selected feature set including sensor energy demand readings, day of the week, season, humidity, and weather conditions (sunny, rain, etc.). The performance of the M5 model tree was evaluated by comparing it to the support vector regression (SVR) and artificial neural networks (ANN) models. The M5 model tree outperformed the SVR and ANN models with a mean absolute error (MAE) of 8.94 compared to 10.02 and 12.04 for the SVR and ANN models respectively.
Sara S. Abdelkader, Katarina Grolinger, Miriam A. M. Capretz
ICMLA3
2015 MLaaS: Machine Learning as a Service
abstract
The demand for knowledge extraction has been increasing. With the growing amount of data being generated by global data sources (e.g., social media and mobile apps) and the popularization of context-specific data (e.g., the Internet of Things), companies and researchers need to connect all these data and extract valuable information. Machine learning has been gaining much attention in data mining, leveraging the birth of new solutions. This paper proposes an architecture to create a flexible and scalable machine learning as a service. An open source solution was implemented and presented. As a case study, a forecast of electricity demand was generated using real-world sensor and weather data by running different algorithms at the same time.
Mauro Ribeiro, Katarina Grolinger, Miriam A. M. Capretz
ICMLA3
2015 Fine-grained filtering to provide access control for data providing services within collaborative environments
abstract
Summary A data providing service (DPS) in service‐oriented architecture is tasked only with the retrieval of data that are annotated over a domain ontology. One particular motivating application of DPSs is their use within collaborative environments. An important characteristic for the enterprises of such a collaborative environment is the ability to employ data sharing with one another. A major concern in this situation is the protection of each enterprise's privacy while still permitting data sharing. One potential solution is to provide filtered data through access control. This work describes how to implement access control through fine‐grained filtering of DPS response messages; it is accomplished using a filtering ontology and relations between the domain ontology of DPS and the proposed filtering ontology. Therefore, enterprises can write enterprise‐specific access control policies referencing a common filtering ontology defined within a collaborative environment, enabling access control‐based data sharing within the environment. This work additionally illustrates the implementation of our general solution to data providing web services, interpreted by an eXtensible Access Control Markup Language‐based access control framework. The implementation is further evaluated in a case study of real world data, provided by a health research institute in London, Canada. Copyright © 2013 John Wiley & Sons, Ltd.
Kevin P. Brown, Michael A. Hayes, David S. Allison, Miriam A. M. Capretz, Margaret Sazio, Rupinder Mann
Concurr. Comput. Pract. Exp.4
2014 Service Evolution Patterns
abstract
Service evolution is the process of maintaining and evolving existing Web services to cater for new requirements and technological changes. In this paper, a service evolution model is proposed to analyze service dependencies, identify changes on services and estimate impact on consumers that will use new versions of these services. Based on the proposed service evolution model, four service evolution patterns are described: compatibility, transition, split-map, and merge-map. These proposed patterns provide reusable templates to encourage well-defined service evolution while minimizing issues that arise otherwise. They can be applied in the service evolution scenario where a single service is used by many, possibly unknown, consumers' applications. In such a scenario, providers evolve their services independently from consumers, which might cause unexpected errors and incur unpredicted impact on the dependent consumers' applications. Therefore, providers can use these patterns to estimate the impact that changes to be introduced to their services may cause on their consumers, and to allow consumers smoothly migrate to the newest version of the service.
Shuying Wang, Wilson A. Higashino, Michael A. Hayes, Miriam A. M. Capretz
ICWS4
2014 Energy-aware resource selection on opportunistic grids
abstract
Energy consumption has been a constant concern for high-performance computing (HPC). Recently, this concern has gained attention from the research community, which is aiming to reduce its costs. The performance gain in such an environment is usually proportional to cost. Examples of such environments are computational grids, which are used in the academic and enterprise domains. On the other hand, one way of obtaining high-performance computing with low-cost investment is by using opportunistic grids, which have become a viable alternative to super-computers and dedicated clusters. This paper proposes an energy-aware resource-selection algorithm to reduce energy consumption in opportunistic grids. The proposed algorithm takes into consideration resource status as well as actions to be taken before allocation to calculate energy consumption. Experimental analysis conducted in this study, taking into account network traffic and node status, shows that a more efficient resource-selection outcome can be obtained, leading to reduced energy consumption. Tests demonstrate an energy-consumption reduction of around 9.5% compared to a commonly used approach.
Izaias De Faria, Mario A. R. Dantas, Miriam A. M. Capretz
ISCC3
2014 Challenges for MapReduce in Big Data
abstract
In the Big Data community, MapReduce has been seen as one of the key enabling approaches for meeting continuously increasing demands on computing resources imposed by massive data sets. The reason for this is the high scalability of the MapReduce paradigm which allows for massively parallel and distributed execution over a large number of computing nodes. This paper identifies MapReduce issues and challenges in handling Big Data with the objective of providing an overview of the field, facilitating better planning and management of Big Data projects, and identifying opportunities for future research in this field. The identified challenges are grouped into four main categories corresponding to Big Data tasks types: data storage (relational databases and NoSQL stores), Big Data analytics (machine learning and interactive analytics), online processing, and security and privacy. Moreover, current efforts aimed at improving and extending MapReduce to address identified challenges are presented. Consequently, by identifying issues and challenges MapReduce faces when handling Big Data, this study encourages future Big Data research.
Katarina Grolinger, Michael A. Hayes, Wilson A. Higashino, Alexandra L'Heureux, David S. Allison, Miriam A. M. Capretz
SERVICES6
2014 Network and Energy-Aware Resource Selection Model for Opportunistic Grids
abstract
Due to increasing hardware capacity, computing grids have been handling and processing more data. This has led to higher amount of energy being consumed by grids, hence the necessity for strategies to reduce their energy consumption. Scheduling is a process carried out to define in which node tasks will be executed in the grid. This process can significantly impact the global system performance, including energy consumption. This paper focuses on a scheduling model for opportunistic grids that considers network traffic, distance between input files and execution node as well as the execution node status. The model was tested in a simulated environment created using Green Cloud. The simulation results of this model compared to a usual approach show a total power consumption savings of 7.10%.
Izaias De Faria, Mario A. R. Dantas, Miriam A. M. Capretz, Wilson A. Higashino
WETICE3
2014 Evaluation of Particle Swarm Optimization Applied to Grid Scheduling
abstract
The problem of scheduling independent users' jobs to resources in Grid Computing systems is of paramount importance. This problem is known to be NP-hard, and many techniques have been proposed to solve it, such as heuristics, genetic algorithms (GA), and, more recently, particle swarm optimization (PSO). This article aims to use PSO to solve grid scheduling problems, and compare it with other techniques. It is shown that many often-overlooked implementation details can have a huge impact on the performance of the method. In addition, experiments also show that the PSO has a tendency to stagnate around local minima in high-dimensional input problems. Therefore, this work also proposes a novel hybrid PSO-GA method that aims to increase swarm diversity when a stagnation condition is detected. The method is evaluated and compared with other PSO formulations, the results show that the new method can successfully improve the scheduling solution.
Wilson A. Higashino, Miriam A. M. Capretz, Maria Beatriz Felgar de Toledo
WETICE2
2014 Query Analyzer and Manager for Complex Event Processing as a Service
abstract
Complex Event Processing (CEP) is a set of tools and techniques that can be used to obtain insights from high-volume, high-velocity continuous streams of events. CEP-based systems have been adopted in many situations that require prompt establishment of system diagnostics and execution of reaction plans, such as in monitoring of complex systems. This article describes the Query Analyzer and Manager (QAM) module, a first effort toward the development of a CEP as a Service (CEPaaS) system. This module is responsible for analyzing user-defined CEP queries and for managing their execution in distributed cloud-based environments. Using a language-agnostic internal query representation, QAM has a modular design that enables its adoption by virtually any CEP system.
Wilson A. Higashino, Cédric Eichler, Miriam A. M. Capretz, Thierry Monteil 0001, Maria Beatriz Felgar de Toledo, Patricia Stolf
WETICE3
2014 Integration of business process modeling and Web services: a survey
Katarina Grolinger, Miriam A. M. Capretz, Americo Cunha, Saïd Tazi 0001
Serv. Oriented Comput. Appl.2
2013 An analysis of replication and retrieval of medical image data using a database management system and a distributed file system
abstract
This paper presents a research study consisting of a comparison analysis of replication and retrieval of medical image data which use the Digital Imaging and Communications in Medicine (DICOM) standard. These data are stored in a relational database management system (RDBMS) and in the Hierarchical Data Format (HDF) using a distributed file system as a data backend. The importance of this work was measured by verification of elapsed-time reduction for medical image data replication and retrieval in a real telemedicine environment.
Elias Amaral Santos, Eros Comunello, Douglas Dyllon Jeronimo de Macedo, Miriam A. M. Capretz, Thiago Coelho Prado, Mario A. R. Dantas
ISCC4
2013 ODEP-DPS: Ontology-driven engineering process for the collaborative development of semantic data providing services
Kevin P. Brown, Miriam A. M. Capretz
Inf. Softw. Technol.2
2012 Ontology-based Representation of Simulation Models
Katarina Grolinger, Miriam A. M. Capretz, José R. Martí, Krishan D. Srivastava
SEKE2
2011 A Multi-layered Approach for the Declarative Development of Data Providing Services
abstract
Data Providing Services (DPSs) have the sole purpose of retrieving data from existing sources according to their input parameters while also providing a semantic description of the data they provide using a parametrized view over a domain ontology. A layered model of viewing DPSs is proposed consisting of the data acquisition, syntactic and semantic layers. It is shown that by defining all three layers, a DPS may be generated and managed exclusively by its declarative definition. This will increase the agility and efficiency with which DPSs may be deployed and managed. As a development model, a set of reusable messages are created, these messages are to be semantically annotated using a view over the domain ontology and are syntactically represented such that they may be exported to XML Schema. These messages are used within the DPS definition where their views over the domain ontology are parametrized and the data acquisition layer is defined to acquire data from the source.
Kevin P. Brown, Miriam A. M. Capretz
ICWS2
2011 Trust bootstrapping services and service providers
abstract
Trust is an important factor for successful online communication. Trust has been used as a criterion for service selection. Most trust and reputation studies assume a system where trust and reputations already exist. However, it is important to initialize trust rates for new services, which have no rating history, the so-called trust bootstrapping process. Trust bootstrapping assists the requestors in their service selection decision. Trust bootstrapping is the initial step in trust building process. Trust bootstrapping is important for reliable interaction with services and service providers that are new to the system. This paper proposes an approach for trust bootstrapping services and service providers. The proposed solution follows the trust principles and addresses a number of trust challenges in the literature. Experiment study is conducted and the results are analysed.
Zainab M. Aljazzaf, Miriam A. M. Capretz, Mark Perry 0001
PST2
2011 From Glossaries to Ontologies: Disaster Management Domain(S)
Katarina Grolinger, Kevin P. Brown, Miriam A. M. Capretz
SEKE3
2011 Dependency and Entropy Based Impact Analysis for Service-Oriented System Evolution
abstract
Service-oriented architecture (SOA) has become a widely accepted approach that provides a flexible IT infrastructure in order to deal with the increasing pace of business changes and global competition. However, the evolution of SOA challenges traditional research methodologies and motivates us to explore new approaches for analyzing and evaluating the effects of change. In order to identify the basic principles that analyze the SOA evolution, this work will utilize dependency analysis and information entropy. Dependency analysis is applied in order to derive information from service dependencies to measure the relative importance of the service in a service-oriented system. In addition, we combine information entropy with dependency analysis to measure the service and system entropy. Consequently, we provide a quantificational approach based on dependency analysis and information entropy to estimate the change effects on a service-oriented system.
Shuying Wang, Miriam A. M. Capretz
Web Intelligence2
2011 A unit test approach for database schema evolution
Katarina Grolinger, Miriam A. M. Capretz
Inf. Softw. Technol.2
2010 An infrastructure for executing WS-BPEL workflows in a Cluster of Clusters
abstract
The use of heterogeneous distributed systems, such as Clusters of Clusters (CoC) and computational grids, enables the development of complex computational systems. However some of these systems require services to be easily composed and executed while maintaining the dependencies among them; both the services and the dependencies can be represented as workflows. Moreover, Web Services have been adopted in distributed systems, and WS-BPEL is the standard for Web Services composition. Accordingly, the objective of this work is to enable the execution of WS-BPEL workflows in a CoC. In order to achieve this execution, an extension to the WS-BPEL language is presented that includes the specification of Quality of Services (QoS), along with the required computational resources for service execution in CoCs. Additionally, this paper presents an infrastructure that enables the execution of such workflows in CoCs. Finally, a case study using SHARCNET as a CoC is discussed in order to evaluate the proposed approach.
Thiago A. Lechuga, Maria Beatriz Felgar de Toledo, Miriam A. M. Capretz
ISCC3
2010 Privacy Protection Mechanisms for Web Service Technology
abstract
The successful use of Web service technology in areas such as healthcare and government depends on its support to privacy preservation. As there is currently no privacy standard for Web services, several solutions have recently been proposed in the literature to deal with privacy in Web services. However, there is no solution that provides a suitable mechanism to describe privacy properties in Web services. When loosely-coupled components are involved, such as in Web service environments, a rich description of components is needed to determine whether they can interact in a manner that preserves privacy. The goal of this paper is to present privacy protection mechanisms for Web services, which use policies defined in the Web Services Policy Framework (WS-Policy) and an ontology defined in the Web Ontology Language (OWL) in order to support Web service interactions with suitable privacy preservation levels.
David S. Allison, Miriam A. M. Capretz, Maria Beatriz Felgar de Toledo
SERA3
2010 Trust in Web Services
abstract
Trust is an important factor to predict the behaviour of a Web Service and as a criterion for Web Service selection. Although considerable research has been performed in the offline and online worlds, analysis of trust in the Web Services environment has been limited. Most trust studies in Web Services are focused on trust establishment without identifying and considering the main trust definition components and trust principles. Thus, this paper presents trust definition and trust principles based on the exploration of the trust literature in the offline and online worlds and Web Services. The trust definition and principles form a basis to establish trust in Web Services.
Zainab M. Aljazzaf, Mark Perry 0001, Miriam A. M. Capretz
SERVICES3
2010 Moving from SaaS Applications towards SOA Services
abstract
This paper presents a brief introduction of Software as a Service (SaaS) and Service Oriented Architecture (SOA). Specifically, the paper introduces a five-step model to show how SaaS can be offered as SOA services. Furthermore, a real-life scenario is provided to demonstrate the benefits of using the proposed model.
Ali Bou Nassif, Miriam A. M. Capretz
SERVICES2
2010 Security Protocols in Service-Oriented Architecture
abstract
In this paper, a comprehensive Quality of Security Service (QoSS) model for addressing security within a Service-Oriented Architecture (SOA) is proposed. We define a detailed SOA security model that supports and incorporates a number of networking security techniques and protocols. It utilizes symmetric keys, public keys and hash functions techniques, in order to provide different levels of QoSS agreements to satisfy the requirements of both the services providers and requesters. These levels are based on core networking security requirements such as Mutual Authentication, Session keys, Anonymity, and Perfect forward Secrecy. In addition, the proposed model forms a strong line of defense against Replay, Man-in-the-Middle, and Denial-of-Services attacks.
Abdelkader H. Ouda, David S. Allison, Miriam A. M. Capretz
SERVICES3
2010 A Reference Ontology based Approach for Service Oriented Ontology Management
Shuying Wang, Jinghui Lu, Miriam A. M. Capretz
WEBIST (2)3
2010 Intelligent security and access control framework for service-oriented architecture
Hany F. ElYamany, Miriam A. M. Capretz, David S. Allison
Inf. Softw. Technol.2
2009 A Dependency Impact Analysis Model for Web Services Evolution
abstract
As many software systems have been turned as Web services, the evolutionary changes of Web services are becoming an important issue. To understand the way in which the change affects the services, we must ascertain parts of the system that will be effected by the change and examine them for additional impacts. In this paper, we propose an impact analysis model based on service dependency. In particular, the service dependency graph model, service dependency and the relation matrix are examined. Based on the shift and calculation of the matrix, the dependency and impact of the service evolution can be analyzed and its quantity can be ascertained. Furthermore, we also represent an approach for service change annotation and for service evolution process. Overall, these works provide a foundation for the automatic management, control, and evaluation of service evolution.
Shuying Wang, Miriam A. M. Capretz
ICWS2
2009 QoSS Policies within SOA
abstract
In this work, we present a metadata for Quality of Security Service (QoSS) for Service-Oriented Architecture, which supports the description of authentication, authorization and privacy features. The metadata is encapsulated by a QoSS service in order to assist the service consumer and provider to achieve a QoSS contract meeting both of their security requirements. This contract performs as an enforced policy for managing the interactions between the parties.
Hany F. ElYamany, Miriam A. M. Capretz, David S. Allison, Maria Beatriz Felgar de Toledo
Web Intelligence2
2006 A Multi-Agent Framework for Testing Distributed Systems
abstract
Software testing is a very expensive and time consuming process. It can account for up to 50% of the total cost of the software development. Distributed systems make software testing a daunting task. The research described in this paper investigates a novel multi-agent framework for testing 3-tier distributed systems. This paper describes the framework architecture as well as the communication mechanism among agents in the architecture. Web-based application is examined as a case study to validate the proposed framework. The framework is considered as a step forward to automate testing for distributed systems in order to enhance their reliability within an acceptable range of cost and time.
Hany F. ElYamany, Miriam A. M. Capretz, Luiz Fernando Capretz
COMPSAC (2)2
1995 A model for negotiation among agents based on the transaction analysis theory
abstract
In the research work presented, interactions among the various types of agents in autonomous decentralized systems are based on the theory of transaction analysis (TA) of psychology. Human beings are considered one particular type of autonomous agent. Hence, the agent's structure as well as interactions among them are defined and performed much the same as human behavior and attitudes. Cooperation among agents is negotiated by exchanging strokes. After the cooperation among agents is established, each of them has a specific role to perform in order to reach a particular goal. The paper presents the definition of the inner structure of an agent along with the negotiation process occurred among them until the cooperation is established.>
Zixue Cheng, Miriam A. M. Capretz, Minetada Osano
ISADS2
1994 The object-oriented paradigm for software evolution
abstract
In this research work object-oriented concepts are applied to a software maintenance method named COMFORM (configuration management formalization for maintenance). COMFORM is composed of several phases to provide the necessary guidance to maintain existing software systems. One of the aims of the method is redocumentation by keeping the maintenance history and information related to the software modules being maintained. Redocumentation is obtained by filling in pre-defined forms which go hand in hand with the phases of the method. Each form is mapped by a class definition. The concepts of classes and inheritance are used along with software maintenance to help the system manager in the creation and administration of such forms. Version control of the predefined forms is carried out either by reusing common parts of these forms or by defining new subclasses from them. As a consequence, a generic implementation of the method is achieved, thus maintenance support for a wide range of existing software systems with various profiles is accomplished.>
Miriam A. M. Capretz, Luiz Fernando Capretz
COMPSAC1
1994 Software configuration management issues in the maintenance of existing systems
abstract
Abstract The application of the Software Configuration Management (SCM) discipline during the maintenance process of an existing poorly documented software system is vital to bring it under control. Incremental documentation, the activity of building up the software documentation whilst it is examined during the maintenance process, has a key role in such a process. The COMFORM (COnfiguration Management FORmalization for Maintenance) system provides a framework for aiding the maintenance process through the application of the SCM discipline. In this way reliable documentation of an existing software system is obtained incrementally whilst maintaining it. Our goal in the design of the COMFORM system is to define a method for the maintenance process through the use of forms representing each phase of a software maintenance model. This approach allows traceability among the software representations throughout the software maintenance process. In this paper we describe the foregoing steps of the COMFORM system's method towards its formalization.
Miriam A. M. Capretz, Malcolm Munro
J. Softw. Maintenance Res. Pract.1
1992 COMFORM-a software maintenance method based on the software configuration management discipline
abstract
COMFORM (configuration management formalisation for maintenance), a method which provides guidelines and procedures for carrying out a variety of activities performed during the maintenance process, is discussed. The proposed method accommodates a change control framework around which the software configuration management discipline is applied. The aim is to exert control over an existing software system while simultaneously incrementally redocumenting it.>
Miriam A. M. Capretz, Malcolm Munro
ICSM1