EDBT 2026 Demo / reviewers in the wild / expert
Joonseok Park
dblp:66/7043
· DBLP profile ↗
16ranked-venue papers
4as first author
0since 2021 · last 2019
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Reconfigurable computing and FPGAs · 50% Memory systems · 28% Parallel and multicore computing · 19% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs › FPGA design flow
FPGA design space exploration |
0.0 | 1 | 2004 | Performance and Area Modeling of Complete FPGA Designs in the Presence of Loop Transformations · IEEE Trans. Computers 2004 |
Parallel and multicore computing
loop transformation |
0.0 | 1 | 2004 | Performance and Area Modeling of Complete FPGA Designs in the Presence of Loop Transformations · IEEE Trans. Computers 2004 |
Memory systems
memory bandwidth |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Memory systems
processing-in-memory |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Memory systems › memory interface
processor-memory interface |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 2004 | Performance and Area Modeling of Complete FPGA Designs in the Presence of Loop Transformations · IEEE Trans. Computers 2004 |
Compilers and program optimization
dependence analysis |
0.0 | 1 | 2001 | Matching and searching analysis for parallel hardware implementation on FPGAs · FPGA 2001 |
Compilers and program optimization › loop transformation
loop unrolling |
0.0 | 1 | 2001 | Matching and searching analysis for parallel hardware implementation on FPGAs · FPGA 2001 |
Hardware accelerators and domain-specific architectures
irregular application acceleration |
0.0 | 1 | 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999 |
Methods — techniques the papers use, named apart from their topics
analytical modeling · 0.1implicit loop unrolling · 0.1array data dependence analysis · 0.1spatial query processing · 0.0soft macro design · 0.0hard macro design · 0.0parcel-based communication · 0.0PIM-to-PIM interconnect · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Efficient video quality assessment for on-demand video transcoding using intensity variation analysis
Hyoungseok Kim, Joonseok Park |
J. Supercomput. | 2 |
| 2015 | Program-Invariant Checking for Soft-Error Detection using Reconfigurable HardwareabstractThere is an increasing concern about transient errors in deep submicron processor architectures. Software-only error detection approaches that exploit program invariants for silent error detection incur large execution overheads and are unreliable as state can be corrupted after invariant checkpoints. In this article, we explore the use of configurable hardware structures for the continuous evaluation of high-level program invariants at the assembly level. We evaluate the resource requirements and performance of the proposed predicate-evaluation hardware structures when integrated with a 32-bit MIPS soft core on a contemporary reconfigurable hardware device. The results, for a small set of kernel codes, reveal that these hardware structures require a very small number of hardware resources with negligible impact on the processor core that they are integrated in. Moreover, the amount of resources is fairly insensitive to the complexity of the invariants, thus making the proposed structures an attractive alternative to software-only predicate checking. Joonseok Park, Pedro C. Diniz |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2014 | Event classification for vehicle navigation system by regional optical flow analysis
Min-Kook Choi, Joonseok Park |
Mach. Vis. Appl. | 2 |
| 2004 | Data Reuse in Configurable Architectures with RAM Blocks: Extended Abstract
Nastaran Baradaran, Joonseok Park, Pedro C. Diniz |
FPL | 2 |
| 2004 | Compiler reuse analysis for the mapping of data in FPGAs with RAM blocksabstractContemporary configurable architectures have dedicated internal functional units such as multipliers, high-capacity storage RAM, and even CAM blocks. These RAM blocks allow the implementations to cache data to be reused in the near future, thereby avoiding the latency of external memory accesses. We present a data allocation algorithm that utilizes the RAM blocks in the presence of a limited number of hardware registers. This algorithm, based on a compiler data reuse analysis, determines which data should be cached in the internal RAM blocks and when. The preliminary results, for a set of image/signal processing kernels targeting a Xilinx Virtex/spl trade/ FPGA device, reveal that despite the increase latency of accessing data in RAM blocks, designs that use them require smaller configurable resources than designs that exclusively use registers, while attaining comparable and in some cases even better performance. Nastaran Baradaran, Joonseok Park, Pedro C. Diniz |
FPT | 2 |
| 2004 | Performance and Area Modeling of Complete FPGA Designs in the Presence of Loop TransformationsabstractSelecting which program transformations to apply when mapping computations to FPGA-based computing architectures can lead to prohibitively long design space exploration cycles. An alternative is to develop fast, yet accurate, performance and area models to quickly understand the Impact and interaction of the transformations. In this paper, we present a combined analytical performance and area modeling approach for complete FPGA designs in the presence of loop transformations. Our approach takes into account the impact of input/output memory bandwidth and memory interface resources, often the limiting factor in the effective implementation of computations. Our preliminary results reveal that our modeling is very accurate, being therefore amenable to be used in a compiler tool to quickly explore very large design spaces. Joonseok Park, Pedro C. Diniz, K. R. Shesha Shayee |
IEEE Trans. Computers | 1 |
| 2003 | Data Search and Reorganization Using FPGAs: Application to Spatial Pointer-based Data StructuresabstractFPGAs (field programmable gate arrays) have appealing features such as customizable internal and external bandwidth and the ability to exploit vast amounts of fine-grain parallelism. In this paper, we explore the applicability of these features in using FPGAs as smart memory engines for search and reorganization computations over spatial pointer-based data structures. The experimental results in this paper suggests that reconfigurable logic, when combined with the data reorganization, can lead to dramatic performance improvements of up to 20x over traditional computer architectures for pointer-based computations, traditionally not viewed as a good match for reconfigurable technologies. Pedro C. Diniz, Joonseok Park |
FCCM | 2 |
| 2003 | Synthesis and Estimation of Memory Interfaces for FPGA-based Reconfigurable Computing EnginesabstractAs the densities of current FPGA continue to grow it is now possible to generate System-On-a-Chip (SoC) designs where multiple computing cores are connected to various memory modules with customized topology with application specific memory access patterns. For example, Xilinx has recently introduced devices to which a paired down version of a PowerPC core can be mapped and connected to a set of internal memories. In this paper we address the problem of synthesizing and estimating the area and speed of memory interfacing for Static RAM (SRAM) and Synchronous Dynamic RAM (SDRAM) with various latency parameters and access modes. We describe a set of synthesizable and programmable memory interfaces a compiler can use to automatically generate the appropriate designs for mapping computations to FPGA-based architectures. Our preliminary results reveal that it is possible to accurately model the area and timing requirements using a linear estimation function. We have successfully integrated the proposed memory interface designs with simple image processing kernels generated using commercially available behavioral synthesis tools. Joonseok Park, Pedro C. Diniz |
FCCM | 1 |
| 2003 | Performance and Area Modeling of Complete FPGA Designs in the presence of Loop TransformationsabstractDigital image processing algorithms are a good match for direct implementation on FPGAs as current FPGA architectures can naturally match the fine grain parallelism in these applications. Typically, these algorithms are structured as a sequence of operations, expressed in high-level programming languages as tight loop nests. The loops usually define a shifting-window region over which the algorithm applies a simple localized operator (e.g., a differential gradient, or a min/max). In this research we focus on the development of fast, yet accurate performance and area modeling of complete FPGA designs that combine analytical, empirical and behavioral estimation techniques. We model the application of a set of important program transformations for image processing algorithms, namely loop unrolling, tiling, loop interchanging, loop fission and array privatization, and explore pipelined and non-pipelined execution modes. We take into consideration the impact of various transformations, in the presence of limited I/O resources like address generators and external memory data channels, on the performance of a complete design implemented in a FPGA based architecture. K. R. Shesha Shayee, Joonseok Park, Pedro C. Diniz |
FCCM | 2 |
| 2003 | Using FPGAs for data and reorganization engines: preliminary results for spatial pointer-based data structuresabstractFPGAs have appealing features such as customizable internal and external bandwidth and the ability to exploit vast amounts of fine-grain instruction-level parallelism. In this paper we explore the applicability of these features in using FPGAs as data search and reorganization engines for performing search and reorganization computations over spatial pointer-based data structures for which traditional computing platforms perform poorly. The preliminary experiments, for a set of simple spatial queries over spatial sparse-mesh and quad-tree data structures, reveal that 3 year-old FPGA devices can deliver performance that is on par and in some instances even superior to that of today's workstations. This experience suggests that the integration in memory of FPGA-like fabrics for implementing smart memory engines should be performance-wise very advantageous. Pedro C. Diniz, Joonseok Park |
FPGA | 2 |
| 2003 | Performance and Area Modeling of Cmplete FPGA Designs in the Presence of Loop Transformations
K. R. Shesha Shayee, Joonseok Park, Pedro C. Diniz |
FPL | 2 |
| 2002 | Data reorganization engines for the next generation of system-on-a-chip FPGAsabstractField-Programmable-Core-Arrays (FPCA) will include various computing cores for a wide variety of applications ranging from DSP to general purpose computing. With the increasing gap between core computing speeds and memory access latency, managing and orchestrating the movement of data across multiple cores will become increasingly important. In this paper we propose data reorganization engines that allow a wide variety of data reorganizations intra- as well as inter-memory modules for future FPCAs. We have experimented with a suite of data reorganizations pervasive in DSP applications. Our limited set of experiments reveals that the proposed designs for these engines are flexile and use little design area in current FPGA fabrics, making them amenable to be easily integrated in future FPCAs as either soft- or hard- macros. Pedro C. Diniz, Joonseok Park |
FPGA | 2 |
| 2001 | An External Memory Interface for FPGA-Based Computing Engines
Joonseok Park, Pedro C. Diniz |
FCCM | 1 |
| 2001 | Matching and searching analysis for parallel hardware implementation on FPGAsabstractMatching and searching computations play an important role in the indexing of data. These computations are typically encoded in very tight loops with a single index variable and a simple search/ matching predicate. Their inherent sequential nature, either because of data dependences but more often because of very strong control dependences, makes it impossible to apply existing data dependence and parallelization analysis to exploit significant levels parallelism on traditional architectures. This paper describes a class of searching and matching computations and describes a mapping strategy to map these computations to hardware. We have developed a compiler analysis in SUIF using array data dependence analysis and implicit loop unrolling analysis to expose more parallelism for the parallel evaluation of these computations. Our compiler generates parallel hardware specifications in VHDL. The resulting parallel hardware yields significant performance improvements when these kernel operators are repeated over shifted portions of the input data on FPGA-based computing architectures. 1. Pablo Moisset, Pedro C. Diniz, Joonseok Park |
FPGA | 3 |
| 2000 | Automatic Synthesis of Data Storage and Control Structures for FPGA-Based Computing EnginesabstractMapping computations written in high-level programming languages to FPGA-based computing engines requires programmers to create the datapath responsible for the core of the computation as well as the control structures to generate the appropriate signals to orchestrate its execution. This paper addresses the issue of automatic generation of data storage and control structures for FPGA-based reconfigurable computing engines using existing compiler data dependence analysis techniques. We describe a set of parameterizable data storage and control structures used as the target of our prototype compiler. We present a compiler analysis algorithm to derive the parameters of the data storage structures to minimize the required memory bandwidth of the implementation. We also describe a complete compilation scheme for mapping loops that manipulate multi-dimensional array variables to hardware. We present preliminary simulation results for complete designs generated manually using the results of the compiler analysis. These preliminary results show that it is possible to successfully integrate compiler data dependence analysis with existing commercial synthesis tools. Pedro C. Diniz, Joonseok Park |
FCCM | 2 |
| 1999 | Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive ArchitectureabstractProcessing-in-memory (PIM) chips that integrate processor logic into memory devices offer a new opportunity for bridging the growing gap between processor and memory speeds, especially for applications with high memory-bandwidth requirements.The Data-IntensiVe Architecture (DIVA) system combines PIM memories with one or more external host processors and a PIM-to-PIM interconnect.DIVA increases memory bandwidth through two mechanisms: (1) performing selected computation in memory, reducing the quantity of data transferred across the processor-memory interface; and (2) providing communication mechanisms called parcels for moving both data and computation throughout memory, further bypassing the processor-memory bus.DIVA uniquely supports acceleration of important irregular applications, including sparse-matrix and pointer-based computations.In this paper, we focus on several aspects of DIVA designed to effectively support such computations at very high performance levels: (1) the memory model and parcel definitions; (2) the PIM-to-PIM interconnect; and, (3) requirements for the processor-to-memory interface.We demonstrate the potential of PIMbased architectures in accelerating the performance of three irregular computations, sparse conjugate gradient, a natural-join database operation and an object-oriented database query. Mary W. Hall, Peter M. Kogge, Jefferey G. Koller, Pedro C. Diniz, Jacqueline Chame, Jeffrey T. Draper, Jeff LaCoss, John J. Granacki, Jay B. Brockman, Apoorv Srivastava, William C. Athas, Vincent W. Freeh, Joonseok Park |
SC | 14 |