EDBT 2026 Demo / reviewers in the wild / expert
Javier Fernández-Marqués
dblp:171/7908
· DBLP profile ↗
17ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-3747-6523ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards active participant centric vertical federated learning: Some representations may be all you need
Jon Irureta, Jon Imaz, Aizea Lojo, Javier Fernández-Marqués, Iñigo Perona, Marco González |
Knowl. Based Syst. | 4 |
| 2025 | FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language ModelsabstractLarge Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications. Yan Gao 0016, Massimo Roberto Scamarcia, Javier Fernández-Marqués, Mohammad Naseri, Chong Shen Ng, Dimitris Stripelis, Zexi Li 0001, Tao Shen 0002, Jiamu Bai, Daoyuan Chen, Zikai Zhang 0003, Rui Hu 0005, Inseo Song, Kangyoon Lee, Hong Jia, Ting Dang, Zheyuan Liu 0002, Daniel J. Beutel, Lingjuan Lyu, Nicholas D. Lane |
NeurIPS | 3 |
| 2025 | PQA: Exploring the Potential of Product Quantization in DNN Hardware AccelerationabstractConventional multiply-accumulate (MAC) operations have long dominated computation time for deep neural networks (DNNs), especially convolutional neural networks (CNNs). Recently, product quantization (PQ) has been applied to these workloads, replacing MACs with memory lookups to pre-computed dot products. To better understand the efficiency tradeoffs of product-quantized DNNs (PQ-DNNs), we create a custom hardware accelerator to parallelize and accelerate nearest-neighbor search and dot-product lookups. Additionally, we perform an empirical study to investigate the efficiency–accuracy tradeoffs of different PQ parameterizations and training methods. We identify PQ configurations that improve performance-per-area for ResNet20 by up to 3.1×, even when compared to a highly optimized conventional DNN accelerator, with similar improvements on two additional compact DNNs. When comparing to recent PQ solutions, we outperform prior work by 4× in terms of performance-per-area with a 0.6% accuracy degradation. Finally, we reduce the bitwidth of PQ operations to investigate the impact on both hardware efficiency and accuracy. With only 2–6-bit precision on three compact DNNs, we were able to maintain DNN accuracy eliminating the need for DSPs. Ahmed F. AbouElhamayed, Angela Cui, Javier Fernández-Marqués, Nicholas D. Lane, Mohamed S. Abdelfattah |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | Recurrent Early Exits for Federated Learning with Heterogeneous ClientsabstractFederated learning (FL) has enabled distributed learning of a model across multiple clients in a privacy-preserving manner. One of the main challenges of FL is to accommodate clients with varying hardware capacities; clients have differing compute and memory requirements. To tackle this challenge, recent state-of-the-art approaches leverage the use of early exits. Nonetheless, these approaches fall short of mitigating the challenges of joint learning multiple exit classifiers, often relying on hand-picked heuristic solutions for knowledge distillation among classifiers and/or utilizing additional layers for weaker classifiers. In this work, instead of utilizing multiple classifiers, we propose a recurrent early exit approach named ReeFL that fuses features from different sub-models into a single shared classifier. Specifically, we use a transformer-based early-exit module shared among sub-models to i) better exploit multi-layer feature representations for task-specific prediction and ii) modulate the feature representation of the backbone model for subsequent predictions. We additionally present a per-client self-distillation approach where the best sub-model is automatically selected as the teacher of the other sub-models at each client. Our experiments on standard image and speech classification benchmarks across various emerging federated fine-tuning baselines demonstrate ReeFL effectiveness over previous works. Royson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li 0001, Stefanos Laskaridis, Lukasz Dudziak, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
ICML | 2 |
| 2024 | Privacy-Preserving Federated Learning using Flower FrameworkabstractAI projects often face the challenge of limited access to meaningful amounts of training data. In traditional approaches, collecting data in a central location can be problematic, especially in industry settings with sensitive and distributed data. However, there is a solution -"moving the computation to the data" through Federated Learning. Mohammad Naseri, Javier Fernández-Marqués, Yan Gao 0016 |
KDD | 2 |
| 2023 | A First Look into the Carbon Footprint of Federated LearningabstractDespite impressive results, deep learning-based technologies also raise severe privacy and environmental concerns induced by the training procedure often conducted in data centers. In response, alternatives to centralized training such as Federated Learning (FL) have emerged. FL is now starting to be deployed at a global scale by companies that must adhere to new legal demands and policies originating from governments and social groups advocating for privacy protection. However, the potential environmental impact related to FL remains unclear and unexplored. This article offers the first-ever systematic study of the carbon footprint of FL. We propose a rigorous model to quantify the carbon footprint, hence facilitating the investigation of the relationship between FL design and carbon emissions. We also compare the carbon footprint of FL to traditional centralized learning. Our findings show that, depending on the configuration, FL can emit up to two orders of magnitude more carbon than centralized training. However, in certain settings, it can be comparable to centralized learning due to the reduced energy consumption of embedded devices. Finally, we highlight and connect the results to the future challenges and trends in FL to reduce its environmental impact, including algorithms efficiency, hardware capabilities, and stronger industry transparency. Xinchi Qiu, Titouan Parcollet, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Yan Gao 0016, Daniel J. Beutel, Taner Topal, Akhil Mathur, Nicholas D. Lane |
J. Mach. Learn. Res. | 3 |
| 2023 | Mitigating Memory Wall Effects in CNN Engines with On-the-Fly Weights GenerationabstractThe unprecedented accuracy of convolutional neural networks (CNNs) across a broad range of AI tasks has led to their widespread deployment in mobile and embedded settings. In a pursuit for high-performance and energy-efficient inference, significant research effort has been invested in the design of field-programmable gate array (FPGA)–based CNN accelerators. In this context, single computation engines constitute a popular design approach that enables the deployment of diverse models without the overhead of fabric reconfiguration. Nevertheless, this flexibility often comes with significantly degraded performance on memory-bound layers and resource underutilisation due to the suboptimal mapping of certain layers on the engine’s fixed configuration. In this work, we investigate the implications in terms of CNN engine design for a class of models that introduce a pre-convolution stage to decompress the weights at runtime. We refer to these approaches as on-the-fly . This article presents unzipFPGA, a novel CNN inference system that counteracts the limitations of existing CNN engines. The proposed framework comprises a novel CNN hardware architecture that introduces a weights generator module that enables the on-chip on-the-fly generation of weights, alleviating the negative impact of limited bandwidth on memory-bound layers. We further enhance unzipFPGA with an automated hardware-aware methodology that tailors the weights generation mechanism to the target CNN-device pair, leading to an improved accuracy–performance balance. Finally, we introduce an input selective processing element (PE) design that balances the load between PEs in suboptimally mapped layers. Quantitative evaluation shows that the proposed framework yields hardware designs that achieve an average of 2.57× performance efficiency gain over highly optimised GPU designs for the same power constraints and up to 3.94× higher performance density over a diverse range of state-of-the-art FPGA-based CNN accelerators. Stylianos I. Venieris, Javier Fernández-Marqués, Nicholas D. Lane |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | End-to-End Speech Recognition from Federated Acoustic ModelsabstractTraining Automatic Speech Recognition (ASR) models under federated learning (FL) settings has attracted a lot of attention recently. However, the FL scenarios often presented in the literature are artificial and fail to capture the complexity of real FL systems. In this paper, we construct a challenging and realistic ASR federated experimental setup consisting of clients with heterogeneous data distributions using the French and Italian sets of the CommonVoice dataset, a large heterogeneous dataset containing thousands of different speakers, acoustic environments and noises. We present the first empirical study on an attention-based sequence-to-sequence End-to-End (E2E) ASR model with three aggregation weighting strategies – standard FedAvg, loss-based aggregation and a novel word error rate (WER)-based aggregation, compared in two realistic FL scenarios: cross-silo with 10 clients and cross-device with 2K and 4K clients. This 4K cross-device ASR experiment is the largest ever performed. Our first-of-its-kind analysis on E2E ASR from heterogeneous and realistic federated acoustic models provides the foundations for future research and development of realistic FL ASR applications. Yan Gao 0016, Titouan Parcollet, Mohamed Salah Zaïem, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Daniel J. Beutel, Nicholas D. Lane |
ICASSP | 4 |
| 2022 | ZeroFL: Efficient On-Device Training for Federated Learning with Local Sparsity
Xinchi Qiu, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Yan Gao 0016, Titouan Parcollet, Nicholas D. Lane |
ICLR | 2 |
| 2022 | Federated Self-supervised Speech Representations: Are We There Yet?
Yan Gao 0016, Javier Fernández-Marqués, Titouan Parcollet, Abhinav Mehrotra, Nicholas D. Lane |
INTERSPEECH | 2 |
| 2022 | Match to Win: Analysing Sequences Lengths for Efficient Self-Supervised Learning in Speech and AudioabstractSelf-supervised learning (SSL) has proven vital in speech and audio-related applications. The paradigm trains a general model on unlabeled data that can later be used to solve specific downstream tasks. This type of model is costly to train as it requires manipulating long input sequences that can only be handled by powerful centralised servers. Surprisingly, despite many attempts to increase training efficiency through model compression, the effects of truncating input sequence lengths to reduce computation have not been studied. In this paper, we provide the first empirical study of SSL pre-training for different specified sequence lengths and link this to various downstream tasks. We find that training on short sequences can dramatically reduce resource costs while retaining a satisfactory performance for all tasks. This simple one-line change would promote the migration of SSL training from data centres to user-end edge devices for more realistic and personalised applications. Yan Gao 0016, Javier Fernández-Marqués, Titouan Parcollet, Pedro Porto Buarque de Gusmão, Nicholas D. Lane |
SLT | 2 |
| 2021 | unzipFPGA: Enhancing FPGA-based CNN Engines with On-the-Fly Weights GenerationabstractSingle computation engines have become a popular design choice for FPGA-based convolutional neural networks (CNNs) enabling the deployment of diverse models without fabric reconfiguration. This flexibility, however, often comes with significantly reduced performance on memory-bound layers and resource underutilisation due to suboptimal mapping of certain layers on the engine's fixed configuration. In this work, we investigate the implications in terms of CNN engine design for a class of models that introduce a pre-convolution stage to decompress the weights at run time. We refer to these approaches as on-the-fly. To minimise the negative impact of limited bandwidth on memory-bound layers, we present a novel hardware component that enables the on-chip on-the-fly generation of weights. We further introduce an input selective processing element (PE) design that balances the load between PEs on suboptimally mapped layers. Finally, we present unzipFPGA, a framework to train on-the-fly models and traverse the design space to select the highest performing CNN engine configuration. Quantitative evaluation shows that unzipFPGA yields an average speedup of 2.14× and 71% over optimised status-quo and pruned CNN engines under constrained bandwidth and up to 3.69× higher performance density over the state-of-the-art FPGA-based CNN accelerators. Stylianos I. Venieris, Javier Fernández-Marqués, Nicholas D. Lane |
FCCM | 2 |
| 2021 | Degree-Quant: Quantization-Aware Training for Graph Neural Networks
Shyam A. Tailor, Javier Fernández-Marqués, Nicholas D. Lane |
ICLR | 2 |
| 2019 | An Empirical study of Binary Neural Networks' Optimisation
Milad Alizadeh, Javier Fernández-Marqués, Nicholas D. Lane, Yarin Gal |
ICLR (Poster) | 2 |
| 2018 | Deterministic Binary Filters for Convolutional Neural NetworksabstractWe propose Deterministic Binary Filters, an approach to Convolutional Neural Networks that learns weighting coefficients of predefined orthogonal binary basis instead of the conventional approach of learning directly the convolutional filters. This approach results in model architectures with significantly fewer parameters (4x to 16x) and smaller model sizes (32x due to the use of binary rather than floating point precision). We show our deterministic filter design can be integrated into well-known network architectures (such as ResNet and SqueezeNet) with as little as 2% loss of accuracy (under datasets like CIFAR-10). Under ImageNet, they result in 3x model size reduction compared to sub-megabyte binary networks while reaching comparable accuracy levels. Vincent W. S. Tseng, Sourav Bhattacharya, Javier Fernández-Marqués, Milad Alizadeh, Catherine Tong, Nicholas D. Lane |
IJCAI | 3 |
| 2018 | Deterministic binary filters for keyword spotting applicationsabstractWe present a binary architecture with 60% fewer parameters and 50% fewer operations during inference compared to the current state of the art for keyword spotting (KWS) applications at the cost of 3.4% accuracy. We construct convolutional filters on-the-fly using orthogonal binary codes and results in a compact architecture that would fit in devices with less than 30kB of memory. Javier Fernández-Marqués, Vincent W. S. Tseng, Sourav Bhattacharya, Nicholas D. Lane |
MobiSys | 1 |
| 2015 | Quantification of the 3D collagen network geometry in confocal reflection microscopyabstractThe geometry of 3D collagen networks is a key factor that influences the behavior of live cells within extracellular matrices. This paper presents a hybrid, two-step method for fully automatic quantification of the 3D collagen network geometry at fiber resolution in confocal reflection microscopy images. First, a coarse binary mask of the entire network is obtained using steerable filtering and local Otsu thresholding. Second, individual collagen fibers are reconstructed by tracing maximum ridges in the Euclidean distance map of the binary mask. The proposed method, validated in a novel framework using 3D collagen gels with various collagen and Matrigel concentrations, reveal that Matrigel affects the collagen network geometry by decreasing the network pore size while preserving the fiber length and fiber persistence length. Martin Maska, Cristina Ederra, Javier Fernández-Marqués, Arrate Muñoz-Barrutia, Michal Kozubek 0001, Carlos Ortiz-de-Solorzano |
ICIP | 3 |