VLDB 2026 Research / reviewers in the wild / expert
Roto Le
dblp:71/768
· DBLP profile ↗
3ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 61% Reconfigurable computing and FPGAs · 30% Interconnection networks and networks-on-chip · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
3D FPGA |
0.1 | 1 | 2009 | High-performance, cost-effective heterogeneous 3D FPGA architectures · FPGA 2009 |
Electronic design automation › physical design
design partitioning |
0.1 | 1 | 2009 | High-performance, cost-effective heterogeneous 3D FPGA architectures · FPGA 2009 |
Electronic design automation
physical design |
0.1 | 1 | 2009 | High-performance, cost-effective heterogeneous 3D FPGA architectures · FPGA 2009 |
Interconnection networks and networks-on-chip › die-to-die interconnect
3d interconnect |
0.0 | 1 | 2009 | High-performance, cost-effective heterogeneous 3D FPGA architectures · FPGA 2009 |
Methods — techniques the papers use, named apart from their topics
through-silicon via · 0.1switch box design · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | High Performance Parallel JPEG2000 Streaming Decoder Using GPGPU-CPU Heterogeneous SystemabstractThe JPEG2000 image coding standard provides many superior features compared to JPEG and other compression standards. However, the relatively slow performance of JPEG2000, especially in software implementations, is a critical drawback of the standard. Moreover, as image sizes rapidly grow in size, higher demands on performance for image coding and processing are introduced, making the slow performance of JPEG2000 even further pronounced. While much effort over the past decade has been devoted to accelerating the JPEG2000 encoder, there have been very few studies focusing on improving the performance of the JPEG2000 decoder, despite the fact that the performance of the decoder is just as critical as the encoder. This paper proposes a high-performance JPEG2000 decoder that efficiently exploits the recent improvements of modern parallel programming models and hardware architectures. Specifically, a parallel streaming decoder running on a GPGPU-CPU heterogeneous system is developed to fully exploit the flexibility of the high-performance multi-core CPUs and the massively parallel capability of GPGPUs. In addition, a new task scheduling strategy is developed that exploits the soft-heterogeneity in OpenCL and C/C++ at runtime in order to gain a significant performance boost. Running on a heterogeneous configuration of one Nvidia GTX 480 GPU and one Intel Core i7 CPU, the parallel streaming decoder gains more than 8X speedup in runtime compared to the JasPer JPEG2000 software implementation. Roto Le, Joseph L. Mundy, R. Iris Bahar |
ASAP | 1 |
| 2009 | High-performance, cost-effective heterogeneous 3D FPGA architecturesabstractIn this paper, we propose novel architectural and design techniques for three-dimensional field-programmable gate arrays (3D FPGAs) with Through-Silicon Vias (TSVs). We develop a novel design partitioning methodology that maps the heterogeneous computational resources of an FPGA into a number of die such that the total die area is minimized and the FPGA performance is maximized. Minimizing the total die area leads to direct manufacturing cost savings which is an important incentive to bring 3D technology to the fab and onto the market. An estimation framework is developed to assess the impact of silicon area utilized by 3D interconnect resources while taking into account the large area occupied by TSVs which is crucial to total die area of 3D FPGA. And in order to improve area and performance of 3D FPGA, we design a novel 3D switch box with bypass TSVs. We also analyze the impact of different partitioning strategies on die area and find the optimal number of die that gives the largest reductions in total die area while maximizing the performance. Using a well-developed simulation infrastructure, we show that our methodologies can achieve an average reduction of 27.7% in total die area with a reduced interconnect path delay of about 58%. Roto Le, Sherief Reda, R. Iris Bahar |
FPGA | 1 |
| 2009 | High-performance, cost-effective heterogeneous 3D FPGA architecturesabstractIn this paper, we propose novel architectural and design techniques for three-dimensional field-programmable gate arrays (3D FPGAs) with Through-Silicon Vias (TSVs). We develop a novel design partitioning methodology that maps the heterogeneous computational resources of an FPGA into a number of die such that the total die area is minimized and the FPGA performance is maximized. Minimizing the total die area leads to direct manufacturing cost savings which is an important incentive to bring 3D technology to the fab and onto the market. An estimation framework is developed to assess the impact of silicon area utilized by 3D interconnect resources while taking into account the large area occupied by TSVs which is crucial to total die area of 3D FPGAs. In order to improve area and performance of 3D FPGAs, we design a novel 3D switch box with bypass TSVs. We also analyze the impact of different partitioning strategies on die area and find the optimal number of die that gives the largest reductions in total die area while maximizing the performance. Using a well-developed simulation infrastructure, we show that our methodologies can achieve an average reduction of 27.7% in total die area with a reduced interconnect path delay of about 58%. Roto Le, Sherief Reda, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 1 |