VLDB 2026 Research / reviewers in the wild / expert
Pedro Henrique Exenberger Becker
dblp:210/3763
· DBLP profile ↗
4ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-5947-1913ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Caravan: A Hardware/Software Co-Design for Efficient SIMD Neighbor Search on Point CloudsabstractNeighbor search is the backbone task for point cloud processing, which is widely employed in current 3D computer vision applications.To be efficient, neighbor search relies on spatial data structures such as k-d trees, to prune the search space.We found that consecutive neighbor search queries in point cloud processing are often similar, visiting k-d tree nodes with considerable resemblance.In this work, we show how to leverage this observation to effectively exploit the available CPU Vector Processing Unit (VPU) to cheaply speed up neighbor search.We devise our solution with a hardware/software co-design called Caravan.At the software level, Caravan-SW exploits this search similarity, gathering consecutive queries to search for their neighbors in parallel with SIMD instructions.Yet, when the navigation of queries diverges, particularly in the deeper levels of the k-d tree, Caravan-SW faces sparsity and VPU lanes are underutilized.We tackle this with Caravan-HW, adding two new instructions that re-index valid vector elements and allow fast operand shuffling and dense SIMD operations to take place, suppressing the hard-to-predict runtime sparsity of Caravan-SW.With AVX512, Caravan-SW speeds up neighbor search by 4.05× (1.85× end-to-end) during point cloud segmentation in a commodity CPU.With the additional Caravan-HW support, the leaf processing part of neighbor search can be further optimized, boosting speed up to 5.19× (1.97× end-to-end), with minimal area costs of 0.032 mm2.Our programmable and minimally intrusive solution has end-to-end benefits comparable to accelerators. Pedro Henrique Exenberger Becker, Franyell Silfa, José-María Arnau, Antonio González 0001 |
ISCA | 1 |
| 2023 | K-D Bonsai: ISA-Extensions to Compress K-D Trees for Autonomous Driving TasksabstractAutonomous Driving (AD) systems extensively manipulate 3D point clouds for object detection and vehicle localization. Thereby, efficient processing of 3D point clouds is crucial in these systems. In this work we propose K-D Bonsai, a technique to cut down memory usage during radius search, a critical building block of point cloud processing. K-D Bonsai exploits value similarity in the data structure that holds the point cloud (a k-d tree) to compress the data in memory. K-D Bonsai further compresses the data using a reduced floating-point representation, exploiting the physically limited range of point cloud values. For easy integration into nowadays systems, we implement K-D Bonsai through Bonsai-extensions, a small set of new CPU instructions to compress, decompress, and operate on points. To maintain baseline safety levels, we carefully craft the Bonsai-extensions to detect precision loss due to compression, allowing re-computation in full precision to take place if necessary. Therefore, K-D Bonsai reduces data movement, improving performance and energy efficiency, while guaranteeing baseline accuracy and programmability. We evaluate K-D Bonsai over the euclidean cluster task of Autoware.ai, a state-of-the-art software stack for AD. We achieve an average of 9.26% improvement in end-to-end latency, 12.19% in tail latency, and a reduction of 10.84% in energy consumption. Differently from expensive accelerators proposed in related work, K-D Bonsai improves radius search with minimal area increase (0.36%). Pedro Henrique Exenberger Becker, José-María Arnau, Antonio González 0001 |
ISCA | 1 |
| 2021 | Improving multitask performance and energy consumption with partial-ISA multicores
Jeckson Dellagostin Souza, Pedro Henrique Exenberger Becker, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 2 |
| 2020 | Tuning the ISA for increased heterogeneous computation in MPSoCsabstractHeterogeneous MPSoCs are crucial to meeting energy efficiency and performance, given their combination of cores and accelerators. In this work, we propose a novel technique for MPSoCs design, increasing their specialization and task-parallelism within a given area and power budget. By removing the microarchitectural support of costly ISA extensions (e.g., FP, SIMD, crypto) from a few cores (transforming them into PartialISA Cores), we make room to add extra (full and simpler) inorder cores and hardware accelerators. While applications must migrate from Partial-ISA cores when they need the removed ISA support, they also execute at lower power consumption during their ISA-extension-free phases, since partial cores have much simpler datapaths compared to their full-ISA counterparts. On top of it, the additional cores and accelerators increase task-level parallelism and make the MPSoC more suitable for application-specific scenarios. We show the effectiveness of our approach by composing different MPSoCs in distinct execution scenarios, using the FP instructions and RISC-V ISA as a case study. To support our system, we also propose two scheduling policies, performance- and energy-oriented, to coordinate the execution of this novel design. For the former policy, we achieve 2.8× speedup for a neural network road sign detection, 1.53x speedup for a video-streaming app, and 1.2x speedup for a taskparallel scenario, consuming 68%, 75%, and 33% less energy, respectively. For the energy-oriented policy, partial-ISA reduces energy consumption by 29% over a highly efficient baseline, with increased performance. Pedro Henrique Exenberger Becker, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
DATE | 1 |