Leandro de Souza Rosa

dblp:129/9405 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
1since 2021 · last 2022
0000-0003-3457-9164ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Image and video processing · 87% Computational photography and imaging · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 67% Reconfigurable computing and FPGAs · 33%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › feature detection
corner detection
0.612022
luvHarris: A Practical Corner Detector for Event-Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Image and video processing
feature detection
0.612022
luvHarris: A Practical Corner Detector for Event-Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Electronic design automation
high-level synthesis
0.412019
Scaling Up Modulo Scheduling for High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining
0.412019
Scaling Up Modulo Scheduling for High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Reconfigurable computing and FPGAs
modulo scheduling
0.412019
Scaling Up Modulo Scheduling for High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Computational photography and imaging
event-based vision
0.212022
luvHarris: A Practical Corner Detector for Event-Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2022

Methods — techniques the papers use, named apart from their topics

threshold ordinal event-surface · 0.6harris corner detector · 0.6modulo scheduling · 0.4integer linear programming · 0.4
YearPublicationVenuePosition
2022 luvHarris: A Practical Corner Detector for Event-Cameras
abstract
There have been a number of corner detection methods proposed for event cameras in the last years, since event-driven computer vision has become more accessible. Current state-of-the-art have either unsatisfactory accuracy or real-time performance when considered for practical use, for example when a camera is randomly moved in an unconstrained environment. In this paper, we present yet another method to perform corner detection, dubbed look-up event-Harris (luvHarris), that employs the Harris algorithm for high accuracy but manages an improved event throughput. Our method has two major contributions, 1. a novel 'threshold ordinal event-surface' that removes certain tuning parameters and is well suited for Harris operations, and 2. an implementation of the Harris algorithm such that the computational load per event is minimised and computational heavy convolutions are performed only 'as-fast-as-possible', i.e., only as computational resources are available. The result is a practical, real-time, and robust corner detector that runs more than 2.6× the speed of current state-of-the-art; a necessity when using a high-resolution event-camera in real-time. We explain the considerations taken for the approach, compare the algorithm to current state-of-the-art in terms of computational performance and detection accuracy, and discuss the validity of the proposed approach for event cameras.
Arren Glover, Aiko Dinale, Leandro de Souza Rosa, Simeon Bamford, Chiara Bartolozzi
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Scaling Up Modulo Scheduling for High-Level Synthesis
abstract
High-Level Synthesis tools have been increasingly used within the hardware design community to bridge the gap between productivity and the need to design large and complex systems. When targeting heterogeneous systems, where the CPU and the FPGA fabric are both available to perform computations, a design space exploration is usually carried out for deciding which parts of the initial code should be mapped to the FPGA fabric such as the overall system’s performance is enhanced by accelerating its computation via dedicated processors. As the targeted systems become more complex and larger, leading to a large design space exploration, the fast estimative of the possible acceleration that can be obtained by mapping certain functionality into the FPGA fabric is of paramount importance. Loop pipelining, which is responsible for the majority of HLS compilation time, is a key optimization towards achieving high-performance acceleration kernels. A new modulo scheduling algorithm is proposed, which reformulates the classical modulo scheduling problem and leads to a reduced number of integer linear problems solved, resulting in large computational savings. Moreover, the proposed approach has a controlled trade-off between solution quality and computation time. Results show the scalability is improved efficiently from quadratic, for the state-of-the-art method, to linear, for the proposed approach, while the optimized loop suffers a 1% (geomean) increment in the total number of cycles.
Leandro de Souza Rosa, Christos-Savvas Bouganis, Vanderlei Bonato
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Scaling Up Loop Pipelining for High-Level Synthesis: A Non-iterative Approach
abstract
High-level synthesis is a powerful tool for increasing productivity in digital hardware design. However, as digital systems become larger and more complex, designers have to consider an increased number of optimizations and directives offered by high-level synthesis tools to control the hardware generation process, resulting in a large design space to be explored. One of the most impactful optimizations is loop pipelining due to its large improvement in the hardware throughput. Nevertheless, the modulo scheduling algorithms that are used for loop pipelining are computationally expensive, and their application to the whole design space can make its exploration inviable, leading to sub-optimum solutions. Current state-of-the-art tools for modulo scheduling follow an iterative approach, which solves O(n2) optimization problems, where n is the loop code size. To address this problem, this work proposes a novel data-flow-based approach that solves exactly 2 optimization problems, independently of the loop code size. Results show orders-of-magnitude savings in the computation time, leading to significant design space exploration time savings when compared with the state-of-the-art. As such, the proposed method produces hardware designs of higher performance than the ones produced by the current state of the art for large and complex loops, maintaining a similar resource utilization.
Leandro de Souza Rosa, Vanderlei Bonato, Christos-Savvas Bouganis
FPT1