VLDB 2026 Research / reviewers in the wild / expert
Juan Escobedo
dblp:195/4171
· DBLP profile ↗
8ranked-venue papers
7as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Electronic design automation · 49% Memory systems · 41% High-performance computing · 9% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
1.5 | 4 | 2020 | DOMIS: Dual-Bank Optimal Micro-Architecture for Iterative Stencils · FPGA 2020 Optimizing Order-Associative Kernel Computation with Joint Memory Banking and Data Reuse · FPGA 2019 Graph-Theoretically Optimal Memory Banking for Stencil-Based Computing Kernels · FPGA 2018 |
Memory systems › DRAM › DRAM architecture
memory bank |
1.0 | 3 | 2019 | Optimizing Order-Associative Kernel Computation with Joint Memory Banking and Data Reuse · FPGA 2019 Graph-Theoretically Optimal Memory Banking for Stencil-Based Computing Kernels · FPGA 2018 Extracting data parallelism in non-stencil kernel computing by optimally coloring folded memory conflict graph · DAC 2018 |
Memory systems › data locality
data reuse |
0.8 | 2 | 2020 | DOMIS: Dual-Bank Optimal Micro-Architecture for Iterative Stencils · FPGA 2020 Optimizing Order-Associative Kernel Computation with Joint Memory Banking and Data Reuse · FPGA 2019 |
Electronic design automation › high-level synthesis › memory synthesis
memory partitioning |
0.5 | 2 | 2020 | DOMIS: Dual-Bank Optimal Micro-Architecture for Iterative Stencils · FPGA 2020 Graph-Theoretically Optimal Memory Banking for Stencil-Based Computing Kernels · FPGA 2018 |
High-performance computing
stencil computation |
0.4 | 2 | 2019 | Graph-Theoretically Optimal Memory Banking for Stencil-Based Computing Kernels · FPGA 2018 Optimizing Order-Associative Kernel Computation with Joint Memory Banking and Data Reuse · FPGA 2019 |
Electronic design automation › high-level synthesis › memory synthesis
memory architecture synthesis |
0.3 | 1 | 2018 | Extracting data parallelism in non-stencil kernel computing by optimally coloring folded memory conflict graph · DAC 2018 |
Memory systems › memory access patterns
conflict-free access |
0.1 | 1 | 2018 | Graph-Theoretically Optimal Memory Banking for Stencil-Based Computing Kernels · FPGA 2018 |
Methods — techniques the papers use, named apart from their topics
graph coloring · 0.7dual-bank micro-architecture · 0.4joint memory banking and data reuse · 0.4graph theory · 0.3conflict graph · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | SpectralFly: Ramanujan Graphs as Flexible and Efficient Interconnection NetworksabstractIn recent years, graph theoretic considerations have become increasingly important in the design of HPC interconnection topologies. One approach is to seek optimal or near-optimal families of graphs with respect to a particular graph theoretic property, such as diameter. In this work, we consider topologies which optimize the spectral gap. We study a novel HPC topology, SpectralFly, designed around the Ramanujan graph construction of Lubotzky, Phillips, and Sarnak (LPS). We show combinatorial properties, such as diameter, bisection bandwidth, average path length, and resilience to link failure, of SpectralFly topologies are better than, or comparable to, similarly constrained DragonFly, SlimFly, and BundleFly topologies. Additionally, we simulate the performance of SpectralFly on a representative sample of micro-benchmarks using the Structure Simulation Toolkit Macroscale Element Library simulator and study cost-minimizing layouts, demonstrating considerable benefit of the SpectralFly topology. Stephen J. Young, Sinan G. Aksoy, Jesun Sahariar Firoz, Roberto Gioiosa, Tobias Hagge, Mark Kempton, Juan Escobedo, Mark Raugas |
IPDPS | 7 |
| 2020 | DOMIS: Dual-Bank Optimal Micro-Architecture for Iterative StencilsabstractHigh-Level Synthesis (HLS) can achieve significant performance improvements through effective memory partitioning and meticulous data reuse. Many modern applications, such as medical imaging and convolutional layers in a CNN, mostly contain kernels where iterations can be reordered freely without compromising its correctness. In this paper, we propose an optimal micro-architecture that can be automatically implemented for simple and iterative stencil computations that utilizes only 2 banks to achieve fully parallel conflict memory accesses from single stage stencil kernels, while only requiring reuse buffers of size proportional to the kernel size to achieve an II of 1, irrespectively of the stencil geometry. We demonstrate the effectiveness of our micro-architecture by implementing it with a Kintex 7 xc7k160tg676-1 Xilinx FPGA and testing it with several stencil-based kernels found in real-world applications. On average, when compared with the mainstream GMP and SRC architectures our approach achieves approximately 30- 70% reduction in hardware usage, while improving performance by about 15%. Moreover, the number of independent memory banks required to accomplish conflict-free data accesses have dropped by more than 30% together with some increase in power consumption due to higher clock frequencies. Juan Escobedo, Mingjie Lin |
FPGA | 1 |
| 2019 | Exploiting Irregular Memory Parallelism in Quasi-Stencils through Nonlinear TransformationabstractNon-stencil kernels with irregular memory accesses pose unique challenges to achieving high computing performance and hardware efficiency in high-level synthesis (HLS) of FPGA. We present a highly versatile and systematic approach to effectively synthesizing a special and important subset of non-stencil computing kernels, quasi-stencils, which possess the mathematical property that, if studied in a particular kind of high-dimensional space corresponding to the prime factorization space, the distance between the memory accesses during each kernel iteration becomes constant and such an irregular non-stencil can be considered as a stencil. This opens the door to exploiting a vast array of existing memory optimization algorithms, such as memory partitioning/banking and data reuse, originally designed for the standard stencil-based kernel computing, therefore offering totally new opportunity to effectively synthesizing irregular non-stencil kernels. We show the feasibility of our approach implementing our methodology in a KC705 Xilinx FPGA board and tested it with several custom code segments that meet the quasi-stencil requirement vs some of the state-of the art methods in memory partitioning. We achieve significant reduction in partition factor, and perhaps more importantly making it proportional to the number of memory accesses instead of depending on the problem size with the cost of some wasted space. Juan Escobedo, Mingjie Lin |
FCCM | 1 |
| 2019 | Optimizing Order-Associative Kernel Computation with Joint Memory Banking and Data ReuseabstractIn this paper, we develop a joint strategy of memory banking and data reuse to specifically optimize the memory performance of any given order-associative and stencil-based computing kernel i.e., its iteration order can be reordered freely without compromising its correctness. Given any shape of stencil kernel, our methodology can achieve throughput of 1 kernel per clock cycle with only two memory banks and two data reuse buffers of constant small buffer sizes provided order-associativeness is given. This is a huge leap over all existing results for general stencil-based computing, where, depending the specific data reuse method, either a number of data reuse buffers proportional to the stencil size are required or a potentially problem-dependent reuse buffer size is needed. Furthermore, the optimal memory partition factor of existing methods is typically proportional to the actual stencil size of a given kernel, whereas in our method, the number of memory banks remains to be 2 irrespective of the stencil shape and size. On average, when compared with the mainstream methods, our approach achieves approximately 30-70% reduction in hardware usage, while improving performance by about 15%. Moreover, the number of independent memory banks required to accomplish conflict-free data accesses have dropped by more than 30%. Juan Escobedo, Mingjie Lin |
FPGA | 1 |
| 2018 | Extracting data parallelism in non-stencil kernel computing by optimally coloring folded memory conflict graphabstractIrregular memory access pattern in non-stencil kernel computing renders the well-known hyperplane- [1], lattice- [2], or tessellation-based [3] HLS techniques ineffective. We develop an elegant yet effective technique that synthesizes memory-optimal architecture from high level software code in order to maximize application-specific data parallelism. Our basic idea is to exploit graph structures embedded in data access pattern and computation structure in order to perform the memory banking that maximizes parallel memory accesses while conserving both hardware and energy consumption. Specifically, we priority color a weighted conflict graph generated from folding the fundamental conflict graph to maximize memory conflict reduction. Most interestingly, our graph-based methodology enables a straightforward tradeoff between the number of memory banks and minimizing memory conflicts. Juan Escobedo, Mingjie Lin |
DAC | 1 |
| 2018 | Graph-Theoretically Optimal Memory Banking for Stencil-Based Computing KernelsabstractHigh-Level Synthesis (HLS) has advanced significantly in compiling high-level "soft»» programs into efficient register-transfer level (RTL) "hard»» specifications. However, manually rewriting C-like code is still often required in order to effectively optimize the access performance of synthesized memory subsystems. As such, extensive research has been performed on developing and implementing automated memory optimization techniques, among which memory banking has been a key technique for access performance improvement. However, several key questions remain to be answered: given a stencil-based computing kernel, what constitutes an optimal memory banking scheme that minimizes the number of memory banks required for conflict-free accesses? Furthermore, if such an optimal memory banking scheme exists, how can an FPGA designer automatically determine it? Finally, does any stencil-based kernel have the optimal banking scheme? In this paper we attempt to optimally solve memory banking problem for synthesizing stencil-based computing kernels with well-known theorems in graph theory. Our graph-based methodology not only computes the minimum memory partition factor for any given stencil, but also exploits the repeatability of coloring entire memory access conflict graph, which significantly improves hardware efficiency. Juan Escobedo, Mingjie Lin |
FPGA | 1 |
| 2017 | Tessellating memory space for parallel accessabstractModern reconfigurable computing chips, such as FPGAs, offer an unprecedented opportunity to achieving both multifunctionality and real-time responsiveness for memory-intensive embedded applications. However, how to cost-effectively synthesize application-specific hardware constructs that fully exploit memory-level parallelism remains to be a key challenge. To address this problem, we propose a new tessellation-based memory partitioning and mapping scheme that aims at maximizing parallel memory accesses while conserving both hardware and energy consumption. Comparing with the existing linear skewing and hyper-plane partitioning methodologies, our proposed technique exploits the regularity of tessellation patterns to assign memory bank and calculate intra-bank offset in a direct geometric-based manner, therefore not only quite intuitive to comprehend, but also quite straightforward to implement with hardware. To empirically validate our proposed tessellation-based methodology, we have implemented a baseline prototype with a standard Virtex 7 FPGA device and the Vivado HLS engine from Xilinx. Our experimental results have shown that on average for 5 benchmark applications from SPEC2006, compared with state-of-art methods, we have improvements in clock period of around 13%, memory overhead reduction of up to 100%, and reduction of DSP usage up to 100%. Juan Escobedo, Mingjie Lin |
ASP-DAC | 1 |
| 2016 | Tessellation-based multi-block memory mapping scheme for high-level synthesis with FPGAabstractFor many intensive computing tasks, simultaneous data access into multi-dimensional data arrays is highly restricted by its data mapping strategy and memory port constraint. As such, to increase memory accessing bandwidth, innovative memory partitioning and mapping algorithms have been proposed to simultaneously access multiple memory blocks through physically distributing data elements in the same logical array onto multiple memory blocks. Fortunately, FPGA device provides an unique opportunity of implementing application-specific memory infrastructure that maximizes memory access performance. However, even with the help of existing high-level synthesis (HLS) tools, customizing memory architecture still poses severe challenges that impede the performance of data path. In fact, existing memory partitioning and mapping schemes exploit either linear skewing or hyper-plane partitioning, therefore causing excessive run-time delay and non-optimal memory block space utilization. This work presents a hardware-efficient memory partitioning and mapping scheme with both low computing complexity and low hardware overhead for accessing multidimensional arrays. Targeting at affine memory access patterns often found in many data-intensive applications, our key idea is to leverage the geometric concept of tessellation widely known in combinatorial study and adopt a partitioning scheme based on geometric arguments instead of counting integer points in polytopes for intra-block offset generation. Aiming to assist HLS, our tessellation-based memory scheme exploits hidden memory parallelism through leveraging physically independent memory blocks in FPGAs. Using FPGA devices, our experimental results have shown that our memory partitioning algorithm saves up to 63.7% in the amount of arithmetic operations, around 15% in execution time, and 31.1% in storage overhead relative to the state-of-the-art approach on average across five widely-used circuit benchmarks. Juan Escobedo, Mingjie Lin |
FPT | 1 |