Amanda S. Barnard

dblp:10/8716 · also Amanda Susan Barnard · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-4784-2382ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BITLUME: Precision-Flexible Photonic Computing for Ultra-Fast and Energy-Efficient DNN Acceleration
abstract
As deep learning expands across emerging domains, computational demands are pushing traditional electronic accelerators to their limits. Silicon photonics has emerged as a promising technology for accelerating deep learning workloads, but precision remains a challenge due to noise and non-idealities. In this paper, we present BITLUME, a novel photonic computing unit that enables multiplications beyond 8-bit precision through a precision-flexible scheme. We further propose an optimized round-truncation algorithm and data mapping strategy for BITLUME to reduce optoelectronic conversions, enhance data reuse, and maintain computational accuracy. A hybrid optoelectronic architecture integrating BITLUME is developed and validated using a prototype built with FPGA, RF, and photonic components, achieving 3.7× lower end-to-end latency than the A100 GPU in dot product. Simulations of training seven DNN models at FP32 show that BITLUME achieves up to 3.35× and 10.78× speedup, and 1.53× and 4.12× energy savings, compared to the state-of-the-art photonic accelerator and A100 GPU, respectively.
Chengpeng Xia, Haibo Zhang 0001, Hao Zhang 0058, Yawen Chen 0001, Amanda S. Barnard
ICCAD5
2025 FedShapleX: Shapley Value Driven Context-Aware Model-Heterogeneous Federated Learning
abstract
Model Heterogeneous Federated Learning(MHFL) builds on traditional Federated Learning (FL) to better leverage the knowledge and data distributed across hardware-heterogeneous devices. Among various heterogeneous FL approaches, the Partially Training (PT)-based methods are one of the most promising approaches, which extract submodels from the global model for local training. However, existing state-of-the-art(SOTA) methods lack effective guidance for updating the global model, making it challenging to handle the Non-IID data distribution and maintain generalization across clients. To guide the update of the global model to mitigate the impact of Non-IID data and enhance the generalization of the global model, we proposed FedShapleX: Shapley Value Driven Context-Aware Submodel Extraction for Model-Heterogeneous Federated Learning. In this work, we first proposed a Parameter-based Class-Specific Shapley Value (PCSV), which quantifies each client’s class-specific contribution to the global model, providing a measure of how effectively the local knowledge is utilized. Leveraging the contribution assessment, we further develop a Reinforcement Learning-aided Large Neighbourhood Search Algorithm (RL-LNS) algorithm, which optimizes the submodel extraction scheme based on context-aware contribution information, thereby guiding the global model update more effectively. Leveraging the actor-critic scheme, the RL-LNS combines the strengths of Large Neighbourhood Search (LNS) and Reinforcement Learning (RL), improving the LNS’s search efficiency while simplifying the design of RL policies. To validate the RL-LNS, we have compared the FedShaplex against the state-of-the-art (SOTA) partial training-based approach MHFL, the global model performance, and its average accuracy on clients’ datasets.
Jifeng Chen, Haibo Zhang 0001, Amanda S. Barnard
ICDCS3
2025 ROCKET: An RNS-based Photonic Accelerator for High-Precision and Energy-Efficient DNN Training
abstract
In recent years, the rapid development of Deep Neural Networks (DNNs) has posed significant challenges in terms of training duration and costs. High-frequency, low-power photonic computing has emerged as a highly promising solution. However, the substantial cost of data conversion and the limitations introduced by noise in photonic devices continue to hinder the realization of high-precision and energy-efficient DNN training. To address this challenge, we propose a novel photonic accelerator, ROCKET, based on the Residue Number System (RNS). RNS is based on modular arithmetic and enables support for high-precision computation through parallel multi-path low-precision operations. First, we leverage specialized lookup tables to enable high-throughput, low-latency conversions between high-precision and low-precision numerical representations. Next, we design a low-power photonic accelerator architecture utilizing intensity modulators, which minimizes the number of computational components while maximizing data reuse. Subsequently, we propose a hybrid photonic-electronic pipelined dataflow to maximize parallelism within the photonic-electronic computation path. Finally, we develop a high-frequency (4.096 GHz) hybrid photonic-electronic prototype using FPGA, Radio Frequency (RF), and photonic components to validate the feasibility of the ROCKET. Our large-scale simulations on seven mainstream DNN models show that, compared to the A100 GPU, TPU v4, and the state-of-the-art photonic accelerator Mirage, ROCKET achieves speedups of 33×, 243×, and 198×, respectively, while saving energy by factors of 64×, 204×, and 142×.
Hao Zhang 0058, Haibo Zhang 0001, Chengpeng Xia, Zhiyi Huang 0001, Yawen Chen 0001, Amanda S. Barnard
ICS6
2025 Contribution-Driven Personalization for Model Heterogeneous Federated Learning
abstract
To address the challenges of hardware heterogeneity in Federated Learning (FL), several model-heterogeneous FL schemes have been proposed based on the traditional model-homogeneous approaches. Among the state-of-the-art (SOTA) model-heterogeneous FL approaches, the Partial Training (PT) approach is considered one of the most promising approaches, where submodels are extracted from the global model for local training. However, existing studies focus on either the submodel extraction scheme or the creation of personalized submodels for each client, which lack global model updating or introduce high computational complexity. This can result in poor adaptability, especially in edge computing environments with Non-IID data distribution. In this paper, we presented CDPFL, Contribution-Driven Personalization for Model Heterogeneous Federated Learning, in which the contributions made by the local clients to the global model are evaluated using the Shapley Value. Using the contribution information, Gate Recurrent Unit (GRU) is then used to determine the weight of each client in the next round of model aggregation. In this way, CDPFL is capable of controlling the update of the global model based on the contribution information. To evaluate CDPFL, we compare it against the SOTA PT-based methods. Experimental results show that our approach achieves an improvement of up to 10.17% in global model accuracy under high data heterogeneity scenarios and consistently outperforms all baselines in both high and low heterogeneity scenarios.
Jifeng Chen, Haibo Zhang 0001, Amanda S. Barnard
IJCNN3
2024 Exploring the cloud of feature interaction scores in a Rashomon set
abstract
Interactions among features are central to understanding the behavior of machine learning models. Recent research has made significant strides in detecting and quantifying feature interactions in single predictive models. However, we argue that the feature interactions extracted from a single pre-specified model may not be trustworthy since: *a well-trained predictive model may not preserve the true feature interactions and there exist multiple well-performing predictive models that differ in feature interaction strengths*. Thus, we recommend exploring feature interaction strengths in a model class of approximately equally accurate predictive models. In this work, we introduce the feature interaction score (FIS) in the context of a Rashomon set, representing a collection of models that achieve similar accuracy on a given task. We propose a general and practical algorithm to calculate the FIS in the model class. We demonstrate the properties of the FIS via synthetic data and draw connections to other areas of statistics. Additionally, we introduce a Halo plot for visualizing the feature interaction variance in high-dimensional space and a swarm plot for analyzing FIS in a Rashomon set. Experiments with recidivism prediction and image classification illustrate how feature interactions can vary dramatically in importance for similarly accurate predictive models. Our results suggest that the proposed FIS can provide valuable insights into the nature of feature interactions in machine learning models.
Sichao Li, Quanling Deng, Amanda S. Barnard
ICLR4
2023 Shapley Based Residual Decomposition for Instance Analysis
abstract
In this paper, we introduce the idea of decomposing the residuals of regression with respect to the data instances instead of features. This allows us to determine the effects of each individual instance on the model and each other, and in doing so makes for a model-agnostic method of identifying instances of interest. In doing so, we can also determine the appropriateness of the model and data in the wider context of a given study. The paper focuses on the possible applications that such a framework brings to the relatively unexplored field of instance analysis in the context of Explainable AI tasks.
Tommy Liu, Amanda S. Barnard
ICML2
2023 Variance Tolerance Factors For Interpreting All Neural Networks
abstract
Black box models only provide results for deep learning tasks, and lack informative details about how these results were obtained. Knowing how input variables are related to outputs, in addition to why they are related, can be critical to translating predictions into laboratory experiments, or defending a model prediction under scrutiny. In this paper, we propose a general theory that defines a variance tolerance factor (VTF) inspired by influence function, to interpret features in the context of black box neural networks by ranking the importance of features, and construct a novel architecture consisting of a base model and feature model to explore the feature importance in a Rashomon set that contains all well-performing neural networks. Two feature importance ranking methods in the Rashomon set and a feature selection method based on the VTF are created and explored. A thorough evaluation on synthetic and benchmark datasets is provided, and the method is applied to two real world examples predicting the formation of noncrystalline gold nanoparticles and the chemical toxicity 1793 aromatic compounds exposed to a protozoan ciliate for 40 hours.
Sichao Li, Amanda S. Barnard
IJCNN2
2023 A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
abstract
The GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that minimises the multi-thread GEMM runtime.We present a proof-of-concept approach to building an Architecture and Data-Structure Aware Linear Algebra (ADSALA) software library that uses machine learning to optimise the runtime performance of BLAS routines. More specifically, our method uses a machine learning model on-the-fly to automatically select the optimal number of threads for a given GEMM task based on the collected training data. Test results on two different HPC node architectures, one based on a two-socket Intel Cascade Lake and the other on a two-socket AMD Zen 3, revealed a 25 to 40 per cent speedup compared to traditional GEMM implementations in BLAS when using GEMM of memory usage within 100 MB.
Yufan Xia, Marco De La Pierre, Amanda S. Barnard, Giuseppe M. J. Barca
IPDPS3
2023 Explainable discovery of disease biomarkers: The case of ovarian cancer to illustrate the best practice in machine learning and Shapley analysis
abstract
OBJECTIVE: Ovarian cancer is a significant health issue with lasting impacts on the community. Despite recent advances in surgical, chemotherapeutic and radiotherapeutic interventions, they have had only marginal impacts due to an inability to identify biomarkers at an early stage. Biomarker discovery is challenging, yet essential for improving drug discovery and clinical care. Machine learning (ML) techniques are invaluable for recognising complex patterns in biomarkers compared to conventional methods, yet they can lack physical insights into diagnosis. eXplainable Artificial Intelligence (XAI) is capable of providing deeper insights into the decision-making of complex ML algorithms increasing their applicability. We aim to introduce best practice for combining ML and XAI techniques for biomarker validation tasks. METHODS: We focused on classification tasks and a game theoretic approach based on Shapley values to build and evaluate models and visualise results. We described the workflow and apply the pipeline in a case study using the CDAS PLCO Ovarian Biomarkers dataset to demonstrate the potential for accuracy and utility. RESULTS: The case study results demonstrate the efficacy of the ML pipeline, its consistency, and advantages compared to conventional statistical approaches. CONCLUSION: The resulting guidelines provide a general framework for practical application of XAI in medical research that can inform clinicians and validate and explain cancer biomarkers.
Weitong Huang, Hanna Suominen, Tommy Liu, Gregory Rice, Carlos Salomon, Amanda S. Barnard
J. Biomed. Informatics6