VLDB 2026 Research / reviewers in the wild / expert
Pedro Porto Buarque de Gusmão
dblp:88/10808 · also Pedro P. B. de Gusmao
· DBLP profile ↗
22ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0002-7072-9898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Semi-Supervised Federated Learning with Limited Labeled Data via Adaptive Batchsize and Pseudo LabelingabstractFederated learning (FL) is a distributed learning method that leverages numerous edge devices for training while protecting data privacy. However, most of the data produced by distributed edge devices are unlabeled, leading to the emergence of semi-supervised federated learning (SSFL) to address this issue. Although state-of-the-art approaches perform well when labeled data is sufficient, it shows severe performance degradation and training becomes more challenging as labeled data becomes scarce. In this paper, we improve performance in environments with limited labeled data by dynamically adjusting batch sizes and applying an Adaptive Threshold (AT). Additionally, we propose methods to resolve issues arising when adopting Adaptive Thresholds method of the centralized approach and investigate their limitations. Byoungjun Park, Pedro Porto Buarque de Gusmão, Minhoe Kim |
CCNC | 2 |
| 2024 | FedRepOpt: Gradient Re-parametrized Optimizers in Federated Learning
Kin Wai Lau, Yasar Abbas Ur Rehman, Pedro Porto Buarque de Gusmão, Lai-Man Po |
ACCV (8) | 3 |
| 2023 | L-DAWA: Layer-wise Divergence Aware Weight Aggregation in Federated Self-Supervised Visual Representation LearningabstractThe ubiquity of camera-enabled devices has led to large amounts of unlabeled image data being produced at the edge. The integration of self-supervised learning (SSL) and federated learning (FL) into one coherent system can potentially offer data privacy guarantees while also advancing the quality and robustness of the learned visual representations without needing to move data around. However, client bias and divergence during FL aggregation caused by data heterogeneity limits the performance of learned visual representations on downstream tasks. In this paper, we propose a new aggregation strategy termed Layer-wise Divergence Aware Weight Aggregation (L-DAWA) to mitigate the influence of client bias and divergence during FL aggregation. The proposed method aggregates weights at the layer-level according to the measure of angular divergence between the clients’ model and the global model. Extensive experiments with cross-silo and cross-device settings on CIFAR-10/100 and Tiny ImageNet datasets demonstrate that our methods are effective and obtain new SOTA performance on both contrastive and non-contrastive SSL approaches. Yasar Abbas Ur Rehman, Yan Gao 0016, Pedro Porto Buarque de Gusmão, Mina Alibeigi, Nicholas D. Lane |
ICCV | 3 |
| 2023 | FedVal: Different good or different bad in federated learning
Viktor Valadi, Xinchi Qiu, Pedro Porto Buarque de Gusmão, Nicholas D. Lane, Mina Alibeigi |
USENIX Security Symposium | 3 |
| 2023 | Decentralized Training of 3D Lane Detection with Automatic Labeling Using HD MapsabstractTo have competent 3D lane detection for real-world driving, a massive amount of data from all over the world is needed, but data collection and manual annotation are costly and time-consuming. The diversity of data collected by developmental cars might still be limited compared to the data collected by a large fleet of customer cars. Federated learning enables training models on edge without transferring data out of devices. However, training supervised learning tasks at the edge is directly tied to having access to high-quality labels, which is limited at the edge.In this paper, we propose a fully automatic method to generate 3D lane labels at the edge using a pre-recorded HD map to enable the federated training of the 3D lane detection model. As a reference, a semi-automatic method is applied for creating a 3D-lane dataset used as ground truth. Our experimental results show that the model can achieve comparable performance when training on the same dataset in both a centralized and a decentralized manner. And the models trained on semi-automatic labeled datasets slightly outperform those trained on fully-automatically labeled datasets. This study shows that a well-performing 3D lane detection model can be trained in a supervised and fully decentralized manner, and most importantly, data privacy at the edge is guaranteed. Yadong Mao, Zhuqi Xiao, Che-Tsung Lin, Pedro Porto Buarque de Gusmão, Nicholas D. Lane, Christopher Zach, Mina Alibeigi |
VTC2023-Spring | 4 |
| 2023 | A First Look into the Carbon Footprint of Federated LearningabstractDespite impressive results, deep learning-based technologies also raise severe privacy and environmental concerns induced by the training procedure often conducted in data centers. In response, alternatives to centralized training such as Federated Learning (FL) have emerged. FL is now starting to be deployed at a global scale by companies that must adhere to new legal demands and policies originating from governments and social groups advocating for privacy protection. However, the potential environmental impact related to FL remains unclear and unexplored. This article offers the first-ever systematic study of the carbon footprint of FL. We propose a rigorous model to quantify the carbon footprint, hence facilitating the investigation of the relationship between FL design and carbon emissions. We also compare the carbon footprint of FL to traditional centralized learning. Our findings show that, depending on the configuration, FL can emit up to two orders of magnitude more carbon than centralized training. However, in certain settings, it can be comparable to centralized learning due to the reduced energy consumption of embedded devices. Finally, we highlight and connect the results to the future challenges and trends in FL to reduce its environmental impact, including algorithms efficiency, hardware capabilities, and stronger industry transparency. Xinchi Qiu, Titouan Parcollet, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Yan Gao 0016, Daniel J. Beutel, Taner Topal, Akhil Mathur, Nicholas D. Lane |
J. Mach. Learn. Res. | 4 |
| 2022 | Federated Self-supervised Learning for Video Understanding
Yasar Abbas Ur Rehman, Yan Gao 0016, Pedro Porto Buarque de Gusmão, Nicholas D. Lane |
ECCV (31) | 4 |
| 2022 | End-to-End Speech Recognition from Federated Acoustic ModelsabstractTraining Automatic Speech Recognition (ASR) models under federated learning (FL) settings has attracted a lot of attention recently. However, the FL scenarios often presented in the literature are artificial and fail to capture the complexity of real FL systems. In this paper, we construct a challenging and realistic ASR federated experimental setup consisting of clients with heterogeneous data distributions using the French and Italian sets of the CommonVoice dataset, a large heterogeneous dataset containing thousands of different speakers, acoustic environments and noises. We present the first empirical study on an attention-based sequence-to-sequence End-to-End (E2E) ASR model with three aggregation weighting strategies – standard FedAvg, loss-based aggregation and a novel word error rate (WER)-based aggregation, compared in two realistic FL scenarios: cross-silo with 10 clients and cross-device with 2K and 4K clients. This 4K cross-device ASR experiment is the largest ever performed. Our first-of-its-kind analysis on E2E ASR from heterogeneous and realistic federated acoustic models provides the foundations for future research and development of realistic FL ASR applications. Yan Gao 0016, Titouan Parcollet, Mohamed Salah Zaïem, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Daniel J. Beutel, Nicholas D. Lane |
ICASSP | 5 |
| 2022 | ZeroFL: Efficient On-Device Training for Federated Learning with Local Sparsity
Xinchi Qiu, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Yan Gao 0016, Titouan Parcollet, Nicholas D. Lane |
ICLR | 3 |
| 2022 | Match to Win: Analysing Sequences Lengths for Efficient Self-Supervised Learning in Speech and AudioabstractSelf-supervised learning (SSL) has proven vital in speech and audio-related applications. The paradigm trains a general model on unlabeled data that can later be used to solve specific downstream tasks. This type of model is costly to train as it requires manipulating long input sequences that can only be handled by powerful centralised servers. Surprisingly, despite many attempts to increase training efficiency through model compression, the effects of truncating input sequence lengths to reduce computation have not been studied. In this paper, we provide the first empirical study of SSL pre-training for different specified sequence lengths and link this to various downstream tasks. We find that training on short sequences can dramatically reduce resource costs while retaining a satisfactory performance for all tasks. This simple one-line change would promote the migration of SSL training from data centres to user-end edge devices for more realistic and personalised applications. Yan Gao 0016, Javier Fernández-Marqués, Titouan Parcollet, Pedro Porto Buarque de Gusmão, Nicholas D. Lane |
SLT | 4 |
| 2022 | SelfVIO: Self-supervised deep monocular Visual-Inertial Odometry and depth estimationabstractIn the last decade, numerous supervised deep learning approaches have been proposed for visual-inertial odometry (VIO) and depth map estimation, which require large amounts of labelled data. To overcome the data limitation, self-supervised learning has emerged as a promising alternative that exploits constraints such as geometric and photometric consistency in the scene. In this study, we present a novel self-supervised deep learning-based VIO and depth map recovery approach (SelfVIO) using adversarial training and self-adaptive visual-inertial sensor fusion. SelfVIO learns the joint estimation of 6 degrees-of-freedom (6-DoF) ego-motion and a depth map of the scene from unlabelled monocular RGB image sequences and inertial measurement unit (IMU) readings. The proposed approach is able to perform VIO without requiring IMU intrinsic parameters and/or extrinsic calibration between IMU and the camera. We provide comprehensive quantitative and qualitative evaluations of the proposed framework and compare its performance with state-of-the-art VIO, VO, and visual simultaneous localization and mapping (VSLAM) approaches on the KITTI, EuRoC and Cityscapes datasets. Detailed comparisons prove that SelfVIO outperforms state-of-the-art VIO approaches in terms of pose estimation and depth recovery, making it a promising approach among existing methods in the literature. Yasin Almalioglu, Mehmet Turan, Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Andrew Markham, Agathoniki Trigoni |
Neural Networks | 4 |
| 2022 | Graph-Based Thermal-Inertial SLAM With Probabilistic Neural NetworksabstractSimultaneous localization and mapping (SLAM) system typically employs vision-based sensors to observe the surrounding environment. However, the performance of such systems highly depends on the ambient illumination conditions. In scenarios with adverse visibility or in the presence of airborne particulates (e.g., smoke, dust, etc.), alternative modalities such as those based on thermal imaging and inertial sensors are more promising. In this article, we propose the first complete thermal–inertial SLAM system that combines neural abstraction in the SLAM front end with robust pose-graph optimization in the SLAM back end. We model the sensor abstraction in the front end by employing probabilistic deep learning parameterized by mixture density networks (MDNs). Our key strategies to successfully model this encoding from thermal imagery are the usage of normalized 14-b radiometric data, the incorporation of hallucinated visual (RGB) features, and the inclusion of feature selection to estimate the MDN parameters. To enable a full SLAM system, we also design an efficient global image descriptor that is able to detect loop closures from thermal embedding vectors. We performed extensive experiments and analysis using three datasets, namely self-collected ground robot and hand-held data taken in indoor environment, and one public dataset (SubT-tunnel) collected in underground tunnel. Finally, we demonstrate that an accurate thermal–inertial SLAM system can be realized in conditions of both benign and adverse visibility. Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Pedro Porto Buarque de Gusmão, Bing Wang 0013, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Robotics | 3 |
| 2021 | RadarLoc: Learning to Relocalize in FMCW RadarabstractRelocalization is a fundamental task in the field of robotics and computer vision. There is considerable work in the field of deep camera relocalization, which directly estimates poses from raw images. However, learning-based methods have not yet been applied to the radar sensory data. In this work, we investigate how to exploit deep learning to predict global poses from Emerging Frequency-Modulated Continuous Wave (FMCW) radar scans. Specifically, we propose a novel end-to-end neural network with self-attention, termed RadarLoc, which is able to estimate 6-DoF global poses directly. We also propose to improve the localization performance by utilizing geometric constraints between radar scans. We validate our approach on the recently released challenging outdoor dataset Oxford Radar RobotCar. Comprehensive experiments demonstrate that the proposed method outperforms radar-based localization and deep camera relocalization methods by a significant margin. Wei Wang 0226, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Andrew Markham, Agathoniki Trigoni |
ICRA | 2 |
| 2020 | milliEgo: single-chip mmWave radar aided egomotion estimation via deep sensor fusionabstractRobust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently dominated by optical techniques e.g., visual-inertial odometry these suffer from challenges with scene illumination or featureless surfaces. As an alternative, we propose milliEgo, a novel deep-learning approach to robust egomotion estimation which exploits the capabilities of low-cost mm Wave radar. Although mmWave radar has a fundamental advantage over monocular cameras of being metric i.e., providing absolute scale or depth, current single chip solutions have limited and sparse imaging resolution, making existing point-cloud registration techniques brittle. We propose a new architecture that is optimized for solving this challenging pose transformation problem. Secondly, to robustly fuse mmWave pose estimates with additional sensors, e.g. inertial or visual sensors we introduce a mixed attention approach to deep fusion. Through extensive experiments, we demonstrate our proposed system is able to achieve 1.3% 3D error drift and generalizes well to unseen environments. We also show that the neural architecture can be made highly efficient and suitable for real-time embedded applications. Xiaoxuan Lu 0001, Muhamad Risqi Utama Saputra, Peijun Zhao, Yasin Almalioglu, Pedro Porto Buarque de Gusmão, Changhao Chen, Ke Sun 0012, Agathoniki Trigoni, Andrew Markham |
SenSys | 5 |
| 2019 | Map-aided Navigation for Emergency SearchesabstractReal-time positioning of emergency personnel has been an active research topic for many years. However, studies on how to improve navigation accuracy by using prior information on the idiosyncratic motion characteristics of firefighters are scarce. This paper presents an algorithm for generating pseudo observations of position and orientation based on standard search patterns used by fire-fighters. The iterative closest point algorithm is used to compare walking trajectories estimated from inertial odometry with search patterns generated from digital maps. The resulting fitting errors are then used to integrate the pseudo observations into a map-aided navigation filter. Specifically, we present a sequential Monte Carlo solution where the pattern comparison is used to both update particle weights and create new particle samples. Experimental results involving professional firefighters demonstrate that the proposed pseudo observations can achieve a stable localization error of about one meter, and offer increased robustness in the presence of map errors. Johan Wahlström, Pedro Porto Buarque de Gusmão, Andrew Markham, Agathoniki Trigoni |
DCOSS | 2 |
| 2019 | Distilling Knowledge From a Deep Pose Regressor NetworkabstractThis paper presents a novel method to distill knowledge from a deep pose regressor network for efficient Visual Odometry (VO). Standard distillation relies on ''dark knowledge'' for successful knowledge transfer. As this knowledge is not available in pose regression and the teacher prediction is not always accurate, we propose to emphasize the knowledge transfer only when we trust the teacher. We achieve this by using teacher loss as a confidence score which places variable relative importance on the teacher prediction. We inject this confidence score to the main training task via Attentive Imitation Loss (AIL) and when learning the intermediate representation of the teacher through Attentive Hint Training (AHT) approach. To the best of our knowledge, this is the first work which successfully distill the knowledge from a deep pose regression network. Our evaluation on the KITTI and Malaga dataset shows that we can keep the student prediction close to the teacher with up to 92.95% parameter reduction and 2.12x faster in computation time. Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Yasin Almalioglu, Andrew Markham, Agathoniki Trigoni |
ICCV | 2 |
| 2019 | GANVO: Unsupervised Deep Monocular Visual Odometry and Depth Estimation with Generative Adversarial NetworksabstractIn the last decade, supervised deep learning approaches have been extensively employed in visual odometry (VO) applications, which is not feasible in environments where labelled data is not abundant. On the other hand, unsupervised deep learning approaches for localization and mapping in unknown environments from unlabelled data have received comparatively less attention in VO research. In this study, we propose a generative unsupervised learning framework that predicts 6-DoF pose camera motion and monocular depth map of the scene from unlabelled RGB image sequences, using deep convolutional Generative Adversarial Networks (GANs). We create a supervisory signal by warping view sequences and assigning the re-projection minimization to the objective loss function that is adopted in multi-view pose estimation and single-view depth generation network. Detailed quantitative and qualitative evaluations of the proposed framework on the KITTI [1] and Cityscapes [2] datasets show that the proposed method outperforms both existing traditional and unsupervised deep VO methods providing better results for both pose estimation and depth recovery. Yasin Almalioglu, Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Andrew Markham, Agathoniki Trigoni |
ICRA | 3 |
| 2019 | Learning Monocular Visual Odometry through Geometry-Aware Curriculum LearningabstractInspired by the cognitive process of humans and animals, Curriculum Learning (CL) trains a model by gradually increasing the difficulty of the training data. In this paper, we study whether CL can be applied to complex geometry problems like estimating monocular Visual Odometry (VO). Unlike existing CL approaches, we present a novel CL strategy for learning the geometry of monocular VO by gradually making the learning objective more difficult during training. To this end, we propose a novel geometry-aware objective function by jointly optimizing relative and composite transformations over small windows via bounded pose regression loss. A cascade optical flow network followed by recurrent network with a differentiable windowed composition layer, termed CL-VO, is devised to learn the proposed objective. Evaluation on three real-world datasets shows superior performance of CL-VO over state-of-the-art feature-based and learning-based VO. Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
ICRA | 2 |
| 2019 | DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural NetworkabstractOdometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches have led to more accurate and robust VO systems. However, they have not been well applied to point cloud data yet. In this work, we investigate how to exploit deep learning to estimate point cloud odometry (PCO), which may serve as a critical component in point cloud-based downstream tasks or learning-based systems. Specifically, we propose a novel end-to-end deep parallel neural network called DeepPCO, which can estimate the 6-DOF poses using consecutive point clouds. It consists of two parallel sub-networks to estimate 3D translation and orientation respectively rather than a single neural network. We validate our approach on KITTI Visual Odometry/SLAM benchmark dataset with different baselines. Experiments demonstrate that the proposed approach achieves good performance in terms of pose accuracy. Wei Wang 0226, Muhamad Risqi Utama Saputra, Peijun Zhao, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IROS | 4 |
| 2016 | Gabor filter based image representation for object classificationabstractData representation plays an important role in a classifier's accuracy. A given dataset may lead to better results by simply applying a change of basis while keeping the original number of parameters. In this paper, Gabor Filter based image representation has been exploited for object classification. First, Gabor filter based convolution is computed for features extraction, then down-sampling is performed and features are normalized to zero mean and unit variance. This image representation having discriminative visual patterns is used for training of object classifier in Matlab Neural Toolbox. Performance of this proposed image representation is examined on two real world image datasets CIFAR and MNIST and results show that data representation using Gabor can provide good classification without increasing the number of trainable parameters. Finally, this approach is compared to different configurations of Convolutional Neural Network having trainable parameters to verify the validity of proposed image representation. Syed Tahir Hussain Rizvi, Gianpiero Cabodi, Pedro Porto Buarque de Gusmão, Gianluca Francini |
CoDIT | 3 |
| 2015 | Loop detection in robotic navigation using MPEG CDVSabstractThe choice for image descriptor in a visual navigation system is not straightforward. Descriptors must be distinctive enough to allow for correct localization while still offering low matching complexity and short descriptor size for real-time applications. MPEG Compact Descriptor for Visual Search is a low complexity image descriptor that offers several levels of compromises between descriptor distinctiveness and size. In this work we describe how these trade-offs can be used for efficient loop-detection in a typical indoor environment. Pedro Porto Buarque de Gusmão, Stefano Rosa, Enrico Magli, Skjalg Lepsøy, Gianluca Francini |
MMSP | 1 |
| 2011 | Statistical modelling of outliers for fast visual searchabstractThe matching of keypoints present in two images is an uncertain process in which many matches may be incorrect. The statistical properties of the log distance ratio for pairs of in correct matches are distinctly different from the properties of that for correct matches. Based on a statistical model, we propose a goodness-of-fit test in order to establish whether two images contain views of the same object. This technique can be used as a fast geometric consistency check for visual search. Skjalg Lepsøy, Gianluca Francini, Giovanni Cordara, Pedro Porto Buarque de Gusmão |
ICME | 4 |