Andrew McCarren

dblp:168/5000 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-7297-0984ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Artificial intelligence and machine learning · 4Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Programming Language Selection in Software Engineering: Results from an Adapted MLR Focused on Go, Haskell, Python and Rust
Zoe Collins, Luigi Di Paolo, Cathal O'Grady, Niall Ryan, Gerard Marks, Murat Yilmaz 0001, Paul M. Clarke, Andrew McCarren
EuroSPI (1)8
2025 Inner Source, Outer Source, Low Code, and No Code: Pros, Cons and Contexts - Results from an MLR
Georgijs Pitkevics, Lakshita Dubey, Chee Hin Choa, Yana Koleva, Murat Yilmaz 0001, Paul M. Clarke, Andrew McCarren
EuroSPI (1)7
2025 A Review of Estimation in Software Engineering
Kevin James Tomescu, Niamh Gowran, Lorena Gomez, Eoin Delahunty, Andrew McCarren, Gerard Marks, Murat Yilmaz 0001, Richard Messnarz, Paul M. Clarke
EuroSPI (1)5
2025 Current AI-Based Software Engineering, Strengths and Weaknesses - Results from a MLR
Jed Walshe, Robert Maloney, Evun Grant, Carlos Conde, Gerard Marks, Murat Yilmaz 0001, Richard Messnarz, Paul M. Clarke, Andrew McCarren
EuroSPI (1)9
2025 Exploring the trie of rules: a fast data structure for the representation of association rules
abstract
Association rule mining techniques can generate a large volume of sequential data when implemented on transactional databases. Extracting insights from a large set of association rules has been found to be a challenging process. When examining a ruleset, the fundamental question is how to summarise and represent meaningful mined knowledge efficiently. Many algorithms and strategies have been developed to address issue of knowledge extraction; however, the effectiveness of this process can be limited by the data structures. A better data structure can sufficiently affect the speed of the knowledge extraction process. This paper proposes a novel data structure, called the Trie of rules, for storing a ruleset that is generated by association rule mining. The resulting data structure is a prefix-tree graph structure made of pre-mined rules . This graph stores the rules as paths within the prefix-tree in a way that similar rules overlay each other. Each node in the tree represents a rule where a consequent is this node, and an antecedent is a path from this node to the root of the tree. The evaluation showed that the proposed representation technique shows significant value. It compresses a ruleset with no data loss and benefits in terms of time for basic operations such as searching for a specific rule, which is the base for many knowledge discovery methods. Moreover, our method demonstrated a significant improvement in graph traversal time compared to traditional data structures.
Mikhail Kudriavtsev, Vuong M. Ngo, Mark Roantree, Marija Bezbradica, Andrew McCarren
J. Intell. Inf. Syst.5
2024 Investigating Systems Modernisation: Approaches, Challenges and Risks
Gareth Hogan, Patricija Shalkauskaite, Mengte Zhu, Martin Derwin, Murat Yilmaz 0001, Andrew McCarren, Paul M. Clarke
EuroSPI (1)6
2024 Comparative Analysis of Methods for Performing a Side-Channel Video-Fingerprinting Attack
abstract
In an era of increasing digital privacy concerns, side-channel attacks pose a risk to the privacy of many streaming service clients. With the attack outlined in this study, it is possible to detect what YouTube video a user is watching despite network traffic encryption. An attacker can achieve this by performing a side-channel attack that targets the byte sizes of the Hypertext Transfer Protocol Secure (HTTPS) requests transmitted between the user’s browser and the YouTube server. This study investigates the feasibility of detecting YouTube videos using HTTPS-packet-related network data captured from application logs. Convolutional Neural Networks (CNN), logistic regression, and Dynamic Time Warping (DTW) were used for identifying fingerprinted YouTube videos with side-channel data in an open-world scenario, where most of the examples the model is presented with are unknown. Our research shows that it is possible to identify YouTube videos using network traces with an F1-score of 80% in an open-world scenario using a CNN approach. We further show that it is possible to distinguish known and unknown videos with a recall of 92% in the same open-world scenario. We also show that similar results could not be attained using logistic regression or DTW proposed by other authors.
Deborah Djon, Darragh Connaughton, Geoff Hamilton, Andrew McCarren
ISCC4
2023 Investigating Sources and Effects of Bias in AI-Based Systems - Results from an MLR
Caoimhe De Buitlear, Ailbhe Byrne, Eric McEvoy, Abasse Camara, Murat Yilmaz 0001, Andrew McCarren, Paul M. Clarke
EuroSPI (1)6
2023 Decomposition of Monolith Applications Into Microservices Architectures: A Systematic Review
abstract
Microservices architecture has gained significant traction, in part owing to its potential to deliver scalable, robust, agile, and failure-resilient software products. Consequently, many companies that use large and complex software systems are actively looking for automated solutions to decompose their monolith applications into microservices. This paper rigorously examines 35 research papers selected from well-known databases using a Systematic Literature Review (SLR) protocol and snowballing method, extracting data to answer the research questions, and presents the following four contributions. First, the Monolith to Microservices Decomposition Framework (M2MDF) which identifies the major phases and key elements of decomposition. Second, a detailed analysis of existing decomposition approaches, tools and methods. Third, we identify the metrics and datasets used to evaluate and validate monolith to microservice decomposition processes. Fourth, we propose areas for future research. Overall, the findings suggest that monolith decomposition into microservices remains at an early stage and there is an absence of methods for combining static, dynamic, and evolutionary data. Insufficient tool support is also in evidence. Furthermore, standardised metrics, datasets, and baselines have yet to be established. These findings can assist practitioners seeking to understand the various dimensions of monolith decomposition and the community's current capabilities in that endeavour. The findings are also of value to researchers looking to identify areas to further extend research in the monolith decomposition space.
Yalemisew M. Abgaz, Andrew McCarren, Peter Elger, David Solan, Neil Lapuz, Marin Bivol, Glenn Jackson, Murat Yilmaz 0001, Jim Buckley, Paul M. Clarke
IEEE Trans. Software Eng.2
2020 A Multivocal Literature Review of Function-as-a-Service (FaaS) Infrastructures and Implications for Software Developers
Jake Grogan, Connor Mulready, James McDermott, Martynas Urbanavicius, Murat Yilmaz 0001, Yalemisew M. Abgaz, Andrew McCarren, Silvana Togneri MacMahon, Vahid Garousi, Peter Elger, Paul M. Clarke
EuroSPI7
2019 Representative Sample Extraction from Web Data Streams
Michael Scriney, Congcong Xing, Andrew McCarren, Mark Roantree
DEXA (1)3
2019 A Method for Automated Transformation and Validation of Online Datasets
abstract
While using online datasets for machine learning is commonplace today, the quality of these datasets impacts on the performance of prediction algorithms. One method for improving the semantics of new data sources is to map these sources to a common data model or ontology. While semantic and structural heterogeneities must still be resolved, this provides a well established approach to providing clean datasets, suitable for machine learning and analysis. However, when there is a requirement for a close to real time usage of online data, a method for dynamic Extract-Transform-Load of new sources data must be developed. In this work, we present a framework for integrating online and enterprise data sources, in close to real time, to provide datasets for machine learning and predictive algorithms. An exhaustive evaluation compares a human built data transformation process with our system's machine generated ETL process, with very favourable results, illustrating the value and impact of an automated approach.
Suzanne McCarthy, Andrew McCarren, Mark Roantree
EDOC2
2019 Automating Data Mart Construction from Semi-structured Data Sources
abstract
The global food and agricultural industry has a total market value of USD 8 trillion in 2016, and decision makers in the Agri sector require appropriate tools and up-to-date information to make predictions across a range of products and areas. Traditionally, these requirements are met with information processed into a data warehouse and data marts constructed for analyses. Increasingly however, data are coming from outside the enterprise and often in unprocessed forms. As these sources are outside the control of companies, they are prone to change and new sources may appear. In these cases, the process of accommodating these sources can be costly and very time consuming. To automate this process, what is required is a sufficiently robust extract–transform–load process; external sources are mapped to some form of ontology, and an integration process to merge the specific data sources. In this paper, we present an approach to automating the integration of data sources in an Agri environment, where new sources are examined before an attempt to merge them with existing data marts. Our validation uses a case study of real world Agri data to demonstrate the robustness of our approach and the efficiency of materializing data marts.
Michael Scriney, Suzanne McCarthy, Andrew McCarren, Paolo Cappellari, Mark Roantree
Comput. J.3
2018 Multistep-ahead Prediction: A Comparison of Analytical and Algorithmic Approaches
Fouad Bahrpeyma, Mark Roantree, Andrew McCarren
DaWaK3
2018 Combining Web and Enterprise Data for Lightweight Data Mart Construction
Suzanne McCarthy, Andrew McCarren, Mark Roantree
DEXA (2)2
2017 Detecting Feature Interactions in Agricultural Trade Data Using a Deep Neural Network
Jim O'Donoghue, Mark Roantree, Andrew McCarren
DaWaK3
2016 Variable interactions in risk factors for dementia
abstract
Current estimates predict 1 in 3 people born today will develop dementia, suggesting a major impact on future population health. As such, research needs to connect specialist clinicians, data scientists and the general public. The In-MINDD project seeks to address this through the provision of a Profiler, a socio-technical information system connecting all three groups. The public interact, providing raw data; data scientists develop and refine prediction algorithms; and clinicians use in-built services to inform decisions. Common across these groups are Risk Factors, used for dementia-free survival prediction. Risk interactions could greatly inform prediction but determining these interactions is a problem underpinned by massive numbers of possible combinations. Our research employs a machine learning approach to automatically select best performing hyperparameters for prediction and learns variable interactions in a non-linear survival-analysis paradigm. Demonstrating effectiveness, we evaluate this approach using longitudinal data with a relatively small sample size.
Jim O'Donoghue, Mark Roantree, Andrew McCarren
RCIS3