EDBT 2026 Demo / reviewers in the wild / expert
Soojung Ryu
dblp:33/9419
· DBLP profile ↗
28ranked-venue papers
0as first author
4since 2021 · last 2024
0009-0000-5228-9801ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 3 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | USDN: A Unified Sample-wise Dynamic Network with Mixed-Precision and Early-ExitabstractTo reduce computation in deep neural network inference, a promising approach is to design a network with multiple internal classifiers (ICs) and adaptively select an execution path based on the complexity of a given input. However, quantizing an input-adaptive network, a must-do task for network deployment on edge devices, is a non-trivial task due to jointly allocating its computation budget along with network layers and IC locations. In this paper, we propose Unified Sample-wise Dynamic Network (USDN) with a mixed-precision and early-exit framework that obtains both the optimal location of ICs and layer-wise bit configurations under a given computation budget. The proposed USDN comprises multiple groups of layers, with each group representing a varying degree of complexity for input samples. Experimental results demonstrate that our approach reduces computational cost of the previous work by 12.78% while achieving higher accuracy on ImageNet dataset. Ji-Ye Jeon, Xuan Truong Nguyen, Soojung Ryu |
WACV | 3 |
| 2022 | ShortcutFusion: From Tensorflow to FPGA-Based Accelerator With a Reuse-Aware Memory Allocation for Shortcut DataabstractResidual block is a very common component in recent state-of-the art CNNs such as EfficientNet or EfficientDet. Shortcut data accounts for nearly 40% of feature-maps access in ResNet152 [8]. Most of the previous DNN compilers, accelerators ignore the shortcut data optimization. This paper presents ShortcutFusion, an optimization tool for FPGA-based accelerator with a reuse-aware static memory allocation for shortcut data, to maximize on-chip data reuse given resource constraints. From TensorFlow DNN models, the proposed design generates instruction sets for a group of nodes which uses an optimized data reuse for each residual block. The accelerator design implemented on the Xilinx KCU1500 FPGA card 2.8x faster and 9.9x more power efficient than NVIDIA RTX 2080 Ti for 256x256 input size. . Compared to the result from baseline, in which the weights, inputs, and outputs are accessed from the off-chip memory exactly once per each layer, ShortcutFusion reduces the DRAM access by 47.8-84.8% for RetinaNet, Yolov3, ResNet152, and EfficientNet. Given a similar buffer size to ShortcutMining [8], which also mine the shortcut data in hardware, the proposed work reduces off-chip access for feature-maps 5.27x while accessing weight from off-chip memory exactly once. Duy Thanh Nguyen, Hyeonseung Je, Tuan Nghia Nguyen 0002, Soojung Ryu, Kyujoong Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Fast Simulation of a Many-NPU Network-on-Chip for Microarchitectural Design Space ExplorationabstractA viable solution to cope with the ever-increasing computation complexity of deep learning applications is to integrate many neural processing units (NPUs) in a chip where a network-on-chip (NoC) is used as the communication fabric. Since the design space of an NoC is huge, the network topology is first selected based on the communication patterns of applications with a high-level performance estimation method. After the network topology is selected, the microarchitectural design space exploration is performed with a cycle-level NoC simulator. However, the existing NoC simulator is so slow that design space exploration of the microarchitecture is usually conducted manually in a narrow space. Since a synthetic trace is used, the simulation accuracy is also limited. To overcome these weak-nesses, we present a simulation technique that is fast and accurate enough for microarchitectural design space of an NoC. In the proposed technique, we use the real communication trace from the many-NPU simulation without NoC consideration. To this end, we define the trace format that defines the interface between a many-NPU simulator and the NoC simulator. To accelerate simulation speed, we propose a parallelization technique at the cluster level in the simulation of the hierarchical NoC. The key technique is to manage the timestamps of events at the cluster boundary to do without time synchronization error. And, we adjust the abstraction level of simulation models to reduce the number of modules in the SystemC NoC simulation. With the proposed technique, we could achieve up to 40 times speed-up for 32 NPU system, compared with the FlexNoC simulator. Jintaek Kang, Changjae Yi, Keonjoo Lee, Seungwook Lee, Soojung Ryu, Soonhoi Ha |
DSD | 5 |
| 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient DescentabstractIn this paper, we present GradPIM, a processingin-memory architecture which accelerates parameter updates of deep neural networks training. As one of processing-in-memory techniques that could be realized in the near future, we propose an incremental, simple architectural design that does not invade the existing memory protocol. Extending DDR4 SDRAM to utilize bank-group parallelism makes our operation designs in processing-in-memory (PIM) module efficient in terms of hardware cost and performance. Our experimental results show that the proposed architecture can improve the performance of DNN training and greatly reduce memory bandwidth requirement while posing only a minimal amount of overhead to the protocol and DRAM area. Heesu Kim, Hanmin Park, Kwanheum Cho, Eojin Lee, Soojung Ryu, Kiyoung Choi, Jinho Lee 0001 |
HPCA | 6 |
| 2018 | Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore SystemsabstractThis article proposes Benzene, an energy-efficient distributed SRAM/STT-RAM hybrid cache for manycore systems running multiple applications. It is based on the observation that a naïve application of hybrid cache techniques to distributed caches in a manycore architecture suffers from limited energy reduction due to uneven utilization of scarce SRAM. We propose two-level optimization techniques: intra-bank and inter-bank. Intra-bank optimization leverages highly associative cache design, achieving more uniform distribution of writes within a bank. Inter-bank optimization evenly balances the amount of write-intensive data across the banks. Our evaluation results show that Benzene significantly reduces energy consumption of distributed hybrid caches. Namhyung Kim, Junwhan Ahn, Kiyoung Choi, Daniel Sánchez 0003, Donghoon Yoo, Soojung Ryu |
ACM Trans. Archit. Code Optim. | 6 |
| 2017 | Dirty-Block Tracking in a Direct-Mapped DRAM Cache with Self-Balancing DispatchabstractRecently, processors have begun integrating 3D stacked DRAMs with the cores on the same package, and there have been several approaches to effectively utilizing the on-package DRAMs as caches. This article presents an approach that combines the previous approaches in a synergistic way by devising a module called the dirty-block tracker to maintain the dirtiness of each block in a dirty region. The approach avoids unnecessary tag checking for a write operation if the corresponding block in the cache is not dirty. Our simulation results show that the proposed technique achieves a 10.3% performance improvement on average over the state-of-the-art DRAM cache technique. Sangheon Lee 0006, Soojung Ryu, Kiyoung Choi |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | iPAWS: Instruction-issue pattern-based adaptive warp scheduling for GPGPUsabstractThread or warp scheduling in GPGPUs has been shown to have a significant impact on overall performance. Recently proposed warp schedulers have been based on a greedy warp scheduler where some warps are prioritized over other warps. However, a single warp scheduling policy does not necessarily provide good performance across all types of workloads; in particular, we show that greedy warp schedulers are not necessarily optimal for workloads with inter-warp locality while a simple round-robin warp scheduler provides better performance. Thus, we argue that instead of single, static warp scheduling, an adaptive warp scheduler that dynamically changes the warp scheduler based on the workload characteristics should be leveraged. In this work, we propose an instruction-issue pattern-based adaptive warp scheduler (iPAWS) that dynamically adapts between a greedy warp scheduler and a fair, round-robin scheduler. We exploit the observation that workloads that favor a greedy warp scheduler will have an instruction-issue pattern that is biased towards some warps while workloads that favor a fair, round-robin warp scheduler will tend to issue instructions across all of the warps. Our evaluations show that iPAWS is able to adapt to the more optimal warp scheduler dynamically and achieve performance that is within a few percent of the statically determined, more optimal warp scheduler. We also show that iPAWS can be extended to other warp schedulers, including the cache-conscious wavefront scheduling (CCWS) and Memory Aware Scheduling and Cache Access Re-execution (MASCAR) to exploit the benefits of other warp schedulers while still providing adaptivity in warp scheduling. Minseok Lee, Gwangsun Kim, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
HPCA | 6 |
| 2016 | Dynamic clock synchronization scheme between voltage domains in multi-core architectureabstractUsing independent voltage (and frequency) domains for cores and caches allows us to achieve high energy efficiency since it enables operating the cores and caches at their own optimal voltages. However, it incurs a clock synchronization problem between the core and cache voltage domains. One of the conventional solutions is to add asynchronous FIFOs on the domain crossing boundary, but it degrades performance due to the increased latency. This paper presents a dynamic clock synchronization scheme between two different voltage domains. It uses a fast clock phase detector to monitor the clock phase difference between two voltage domains, a variable delay element consisting of multiple-stage thyristor-like delay circuits to support a wide delay range, and a small hardware module that uses the phase detector output to control the variable delay element. In this way, the scheme automatically adjusts the delay of a clock until the phase is aligned with that of the other clock. Also presented is an algorithm that speeds up the synchronization process. Experimental results show 1.4× speedup and 13% energy saving on average compared to the conventional approach. Kiyoung Choi, Sangheon Lee 0006, Soojung Ryu |
VLSI-SoC | 4 |
| 2014 | Energy-efficient scheduling for memory-intensive GPGPU workloadsabstractHigh performance for a GPGPU workload is obtained by maximizing parallelism and fully utilizing the available resources. However, this is not necessarily energy efficient, especially for memory-intensive GPGPU workloads. In this work, we propose Throttle CTA (cooperative-thread array) Scheduling (TCS) where we leverage two type of throttling - throttling the number of actives cores and throttling of warp execution in the cores - to improve energy-efficiency for memory-intensive GPGPU workloads. The algorithm requires the global CTA or thread block scheduler to reduce the number of cores with assigned thread blocks while leveraging the local warp scheduler to throttle memory requests for some of the cores to further reduce power consumption. The proposed TCS scheduling does not require off-line analysis but can be done dynamically during execution. Instead of relying on conventional metrics such as miss-per-kilo-instruction (MPKI), we leverage the memory access latency metric to determine the memory intensity of the workloads. Our evaluations show that TCS reduces energy by up to 48% (38% on average) across different memory-intensive workload while having very little impact on performance for compute-intensive workloads. Seokwoo Song, Minseok Lee, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
DATE | 6 |
| 2014 | Improving GPGPU resource utilization through alternative thread block schedulingabstractHigh performance in GPGPU workloads is obtained by maximizing parallelism and fully utilizing the available resources. The thousands of threads are assigned to each core in units of CTA (Cooperative Thread Arrays) or thread blocks - with each thread block consisting of multiple warps or wavefronts. The scheduling of the threads can have significant impact on overall performance. In this work, explore alternative thread block or CTA scheduling; in particular, we exploit the interaction between the thread block scheduler and the warp scheduler to improve performance. We explore two aspects of thread block scheduling - 1) LCS (lazy CTA scheduling) which restricts the maximum number of thread blocks allocated to each core, and 2) BCS (block CTA scheduling) where consecutive thread blocks are assigned to the same core. For LCS, we leverage a greedy warp scheduler to help determine the optimal number of thread blocks by only measuring the number of instructions issued while for BCS, we propose an alternative warp scheduler that is aware of the “block” of CTAs allocated to a core. With LCS and the observation that maximum number of CTAs does not necessary maximize performance, we also propose mixed concurrent kernel execution that enables multiple kernels to be allocated to the same core to maximize resource utilization and improve overall performance. Minseok Lee, Seokwoo Song, Joosik Moon, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
HPCA | 7 |
| 2014 | SimParallel: A high performance parallel SystemC simulator using hierarchical multi-threadingabstractAs the system complexity increases, the simulation performance becomes one of the most important issues in virtual prototyping. Parallel simulation is a fascinating technique for high-speed simulation utilizing state of the art multi-core processors on a host workstation, but the efficiency of the parallel simulation is low because of the synchronization and communication overhead and unbalanced workloads among cores in the host. This paper proposes a novel technique, hierarchical multi-threading for the efficient parallel simulation of SystemC models where the host cores are able to be maximally utilized with the same number of thread groups. We also present an efficient synchronization and dynamic load balancing scheme for the proposed parallel simulation. Experimental results show that the proposed method achieves speed-up of from 2.9 to 3.3 in quad-core host workstation. Moo-Kyoung Chung, Jun-Kyoung Kim, Soojung Ryu |
ISCAS | 3 |
| 2014 | Quantitative comparison of the power reduction techniques for samsung reconfigurable processorabstractWith significant growth in portable multimedia devices such as smartphones, application processors (AP) play a critical role for running various multimedia applications on these devices. By considering the power constraints of such devices, we often integrate reconfigurable processors (RPs) into APs. This is because RPs offer flexibility and good performance, thereby greatly improving the power efficiency for running these multimedia applications. Like many other processors, RPs also exploit the dynamic voltage/frequency scaling (DVFS) to improve their power efficiency. Owing to the platform cost constraints, however, these RPs are often integrated to low dropout (LDO) voltage regulators (VRs) instead of switching VRs. When compared with switching VRs, LDO VRs are very inexpensive; however, they suffer from considerable power loss when they are required to deliver a low output voltage. However, many previous studies focused on analyzing the power efficiency of various DVFS techniques only with regard to the processors and did not consider the negative impact of the VR power losses on the overall power efficiency of the platform. In this work, we quantitatively compare the power efficiency of a Samsung RP (SRP) adopting the race-to-halt technique with that of the SRP exploiting the DVFS supported by LDO VRs, by considering the effect of the VR power losses. Finally, we demonstrate that using the race-to-halt technique results in high power efficiency when compared with the DVFS in a commercial processor, by considering the VR power efficiency. Hoyoung Kim, Soojung Ryu, Abhishek A. Sinkar, Nam Sung Kim |
ISCAS | 2 |
| 2014 | Simulation-based memory dependence checker for CGRA-mapped code verificationabstractIn a coarse-grained reconfigurable array (CGRA) architecture, software pipelining is primarily used to improve performance by exploiting loop-level parallelism (LLP). In this technique, the loop-carried memory dependence in user code prevents high parallelism, and it is difficult to be detected. In this paper, we propose a simulation-based memory dependence checker, which is used in the verification of CGRA-mapped code. We use as a reference the memory access behavior of the sequential processor and compare it to that of the CGRA-mapped code. Although it cannot guarantee perfect verification of memory dependence violations, our approach is useful by guiding the programmer to modify the source code. When a memory dependence violation is detected, our approach provides debugging information from the sequential compiled code. Moreover, our checker is implemented in the register transfer level; it enables verification time reduction and the testing of the CGRA-mapped code with a large test input stream in FPGA or ASIC implementations. Heejun Shim, Soojung Ryu |
ISCAS | 2 |
| 2014 | Full-stream architecture for ray tracing with efficient data transmissionabstractIn this paper, we focus on the impact of a memory bandwidth limitation by analyzing the bandwidth consumption for a ray tracing system and present an energy efficient data transmission method using a dedicated interface between the processor and ray tracing hardware engine. To achieve real-time ray tracing, we propose a full-stream architecture through the use of this dedicated interface. For an evaluation of our approach, we implemented a prototype ray tracing architecture using our approach on an FPGA platform. Our experimental results, indicate that our approach shows an average reduction in system memory bandwidth of 48% and an average performance improvement of 50%. Youngsam Shin, Jae Don Lee, Won-Jong Lee, Soojung Ryu, Jeongwook Kim |
ISCAS | 4 |
| 2014 | ULP-SRP: Ultra Low-Power Samsung Reconfigurable Processor for Biomedical ApplicationsabstractThe latest biomedical applications require low energy consumption, high performance, and wide energy-performance scalability to adapt to various working environments. In this study, we present ULP-SRP, an energy-efficient reconfigurable processor for biomedical applications. ULP-SRP uses a Coarse-Grained Reconfigurable Array (CGRA) for high-performance data processing with low energy consumption. We adopted a compact-size CGRA and modified it to support dynamically switchable three performance modes with fine-grained power gating in order to further optimize the energy consumption. The energy-performance scalability is also accomplished with multiple performance modes and a Unified Memory Architecture (UMA). Experimental results show that ULP-SRP achieved 59% energy reduction compared to previous works. A technique of dynamic CGRA mode changing gives 18.9% energy reduction. ULP-SRP is a good candidate for future mobile healthcare devices. Changmoo Kim, Moo-Kyoung Chung, Yeongon Cho, Mario Konijnenburg, Soojung Ryu, Jeongwook Kim |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2013 | Reevaluating the latency claims of 3D stacked memoriesabstractIn recent years, 3D technology has been a popular area of study that has allowed researchers to explore a number of novel computer architectures. One of the more popular topics is that of integrating 3D main memory dies below the computing die and connecting them with through-silicon vias (TSVs). This is assumed to reduce off-chip main memory access latencies by roughly 45% to 60%. Our detailed circuit-level models, however, demonstrate that this latency reduction from the TSVs is significantly less. In this paper, we present these models, compare 2D and 3D main memory latencies, and show that the reduction in latency from using 3D main memory to be no more than 2.4 ns. We also show that although the wider I/O bus width enabled by using TSVs increases performance, it may do so with an increase in power consumption. Although TSVs consume less power per bit transfer than off-chip metal interconnects (11.2 times less power per bit transfer), TSVs typically use considerably more bits and may result in a net increase in power due to the large number of bits in the memory I/O bus. Our analysis shows that although a 3D memory hierarchy exploiting a wider memory bus can increase performance, this performance increase may not justify the net increase in power consumption. Daniel W. Chang, Gyungsu Byun, Hoyoung Kim, Minwook Ahn, Soojung Ryu, Nam Sung Kim, Michael J. Schulte |
ASP-DAC | 5 |
| 2013 | Adaptive compression for instruction code of Coarse Grained Reconfigurable ArchitecturesabstractCoarse Grained Reconfigurable Architecture (CGRA) achieves high performance by exploiting instruction-level parallelism with software pipeline. Large instruction memory is, however, a critical problem of CGRA, which requires large silicon area and power consumption. Code compression is a promising technique to reduce the memory area, bandwidth requirements, and power consumption. We present an adaptive code compression scheme for CGRA instructions based on dictionary-based compression, where compression mode and dictionary contents are adaptively selected for each execution kernel and compression group. In addition, it is able to design hardware decompressor efficiently with two-cycle latency and negligible silicon overhead. The proposed method achieved an average compression ratio 0.52 in a CGRA of 16-functional unit array with the experiments of well-optimized applications. Moo-Kyoung Chung, Jun-Kyoung Kim, Yeongon Cho, Soojung Ryu |
FPT | 4 |
| 2013 | Mobile GPU shader processor based on non-blocking Coarse Grained Reconfigurable Arrays architectureabstractCoarse-grained reconfigurable arrays (CGRAs) based processors provide high performance and energy-efficiency as well as programmability by means of the ability to reconfigure the datapath connecting the ALU arrays. A CGRA based processor executes loop kernels whose schedule should be fixed at compile time. This restriction hinders CGRA from being efficient particularly in accessing external memories or caches whose access time varies greatly. This makes it challenging to build a CGRA based high-performance, energy-efficient mobile GPU because GPU shader execution usually involves massive texture memory accesses which consist of accesses to texture cache and external texture memory. In this paper, we present an Non-blocking Coarse Grained Reconfigurable Arrays (NBC-GRA) architecture which can handle varying-latency operations efficiently. We also propose an improved CGRA based GPU shader processor architecture based on it. Retry buffer enables threads to re-execute later when the required memory access completes. With a non-blocking texture cache, the shader core can execute without stalls even in the case of cache misses. All of these components help to improve CGRA core throughput greatly despite of longer memory access latencies. Evaluation results show that our NBCGRA architecture based shader processor could perform efficiently despite extreme variation of texture cache access latencies and could reduce the shader execution cycles by upto 68% with minimal hardware cost overhead. Kwontaek Kwon, Sungjin Son, Jeong-Soo Park 0004, Jeongae Park, Sangoak Woo, Seokyoon Jung, Soojung Ryu |
FPT | 7 |
| 2013 | Real-time ray tracing on coarse-grained reconfigurable processorabstractRay tracing is a 3D rendering method for generating an image by simulating the path of light. It can generate high quality images, but it requires great computing power. Recent advances in ray tracing technology enable realtime ray tracing on modern desktop CPUs/GPUs. But in the current mobile environment, it is difficult because of inadequate computing power, memory bandwidth, and flexibility in mobile GPUs. In this paper, we present a mobile ray tracing system using Samsung Reconfigurable Processor (SRP). SRP architecture includes a tightly coupled very long instruction word (VLIW) engine and coarse-grained reconfigurable array (CGRA). The VLIW engine is designed for general-purpose computations, such as function invocation and branch selection, and the coarsegrained reconfigurable array is specialized for data-intensive part of the program and can be configured dynamically. We proposed iterative batch-based ray tracing algorithm for SRP, and optimized memory bandwidth with local memory and data cache. Our ray tracing system is implemented on a commercial FPGA-based prototyping system. The experimental results show that our system is suitable for the mobile ray tracing. Jae Don Lee, Youngsam Shin, Won-Jong Lee, Soojung Ryu, Jeongwook Kim |
FPT | 4 |
| 2013 | Mapping and Scheduling of Tasks and Communications on Many-Core SoC Under Local Memory ConstraintabstractThere has been extensive research on mapping and scheduling tasks on a many-core SoC. However, none considers the optimization of communication types, which can significantly affect performance, energy consumption, and local memory usage of the SoC. This paper presents an approach to automatic mapping and scheduling of tasks and communications on a many-core SoC. The key idea is to decide the type of each communication between message passing and shared memory when we do the mapping and scheduling. By assigning a proper type to each communication, we can optimize the energy consumption, performance, or energy-delay product. To solve the optimization problem, the approach adopts a probabilistic algorithm coupled with some heuristics. To enhance throughput of the system, it performs software pipelined scheduling of the tasks using a modified iterative modulo scheduling technique. Experiments show that our algorithm achieves on average 50.1% lower energy consumption, 21.0% higher throughput, and 64.9% lower energy- delay product, compared to shared memory only communication. Jinho Lee 0001, Moo-Kyoung Chung, Yeongon Cho, Soojung Ryu, Jung Ho Ahn, Kiyoung Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Implementation of a volume rendering on coarse-grained reconfigurable multiprocessorabstractIn this paper, we present reconfigurable multiprocessor architecture for volume rendering. The multiprocessor consists of sixteen reconfigurable processors to exploit data parallelism of the volume rendering. Each processor has VLIW core and reconfigurable coarse-grained array specialized for control and data-intensive part of the program, respectively. The coarse-grained array can be configured dynamically, so that it can efficiently process different kernels of the volume rendering. The multiprocessor is implemented using verilog HDL and realized onto a commercial FPGA-based prototyping system. The experimental result shows that the presented multiprocessor has comparable performance to high-end desktop GPUs. Seunghun Jin, Sangheon Lee 0006, Moo-Kyoung Chung, Yeongon Cho, Soojung Ryu |
FPT | 5 |
| 2012 | ULP-SRP: Ultra low power Samsung Reconfigurable Processor for biomedical applicationsabstractThe latest biomedical applications require low energy consumption, high performance and wide energy-performance scalability to adapt to various working environments. This paper presents ULP-SRP, an energy efficient reconfigurable processor for the biomedical applications. ULP-SRP uses a Coarse Grained Reconfigurable Array (CGRA) for high performance data processing with low energy consumption. For the scalability, we propose three performance modes and Unified Memory Architecture (UMA). Energy optimization is accomplished by run-time mode switching along with automatic power gating. Experimental results show that ULP-SRP achieved 46.1% energy reduction compared to previous works. Changmoo Kim, Moo-Kyoung Chung, Yeongon Cho, Mario Konijnenburg, Soojung Ryu, Jeongwook Kim |
FPT | 5 |
| 2012 | Design space exploration and implementation of a high performance and low area Coarse Grained Reconfigurable ProcessorabstractCoarse Grained Reconfigurable Architectures (CGRAs) have played a key role in the area of domain specific processors due to their programmability and runtime reconfigurability. The Coarse Grained Array (CGA) structure enables target designs to achieve high performance, but it is easy to fall into over-design in term of area. Moreover, the network overhead between the function units (FUs) seriously degrades its clock speed. In this paper, we propose a high performance CGRA that facilitates design space exploration (DSE) to reduce these overheads. It employs a concept of building blocks, named mini cores, to mitigate overhead involved in DSE that aims to achieve high clock speed and small area in the target design. The proposed approach reduces the design time more than 100 times compared with previous design. Experimental results show that the implemented architecture reduces logic area by 14.38% and improves clock frequency by 59.34% without performance loss. Dongkwan Suh, Kiseok Kwon, Sukjin Kim, Soojung Ryu, Jeongwook Kim |
FPT | 4 |
| 2012 | Efficient code compression for coarse grained reconfigurable architecturesabstractThough Coarse Grained Reconfigurable Architecture (CGRA) is a flexible alternative for high performance computing, it has a crucial problem on instruction code whose size is so large that the instruction memory takes a significant portion of silicon area and power consumption. This article proposes an efficient dictionary-based compression method for the CGRA instruction code, where code bit-fields are rearranged and grouped together according to locality characteristics and the most efficient compression mode is selected for each group and kernel. The proposed method can reinstall the dictionary contents adaptively for each kernel. Experimental results show that the proposed method achieved an average compression ratio 0.56 in 4×4 array of function units for well-optimized applications. Moo-Kyoung Chung, Yeongon Cho, Soojung Ryu |
ICCD | 3 |
| 2012 | Providing cost-effective on-chip network bandwidth in GPGPUsabstractNetwork-on-chip (NoC) bandwidth has a significant impact on overall performance in throughput-oriented processors such as GPG-PUs. Although it has been commonly assumed that high NoC bandwidth can be provided through abundant on-chip wires, we show that increasing NoC router frequency results in a more cost-effective NoC. However, router arbitration critical path can limit the NoC router frequency. Thus, we propose a direct all-to-all network overlaid on mesh (DA2mesh) NoC architecture that exploits the traffic characteristics of GPGPU and removes arbitration from the router pipeline. DA2mesh simplifies the router pipeline with 36% improvement of performance while reducing NoC energy by 15%. Hanjoon Kim, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
ICCD | 5 |
| 2012 | Design and evaluation of a four-port data cache for high instruction level parallelism reconfigurable processorsabstractThis paper explores high-bandwidth data cache designs for a coarse-grained reconfigurable architecture processor family capable of achieving a high degree of instruction level parallelism. To meet stringent power, area and time-to-market constraints, we take an architectural approach rather than circuit-level multi-porting approaches. We closely examine two design choices: single-level banked cache (SLC) and two-level cache (TLC). A detailed simulation study using a set of microbenchmarks and industry-strength benchmarks finds that both SLC and TLC offer a reasonably competitive performance at a small implementation cost compared with a hypothetical cache with perfect ports and a multi-bank scratchpad memory. Kiyeon Lee, Moo-Kyoung Chung, Soojung Ryu, Yeongon Cho, Sangyeun Cho |
ICCD | 3 |
| 2012 | Automatic code overlay generation and partially redundant code fetch eliminationabstractThere is an increasing interest in explicitly managed memory hierarchies, where a hierarchy of distinct memories is exposed to the programmer and managed explicitly in software. These hierarchies can be found in typical embedded systems and an emerging class of multicore architectures. To run an application that requires more code memory than the available higher-level memory, typically an overlay structure is needed. The overlay structure is generated manually by the programmer or automatically by a specialized linker. Manual code overlaying requires the programmer to deeply understand the program structure for maximum memory savings as well as minimum performance degradation. Although the linker can automatically generate the code overlay structure, its memory savings are limited and it even brings significant performance degradation because traditional techniques do not consider the program context. In this article, we propose an automatic code overlay generation technique that overcomes the limitations of traditional automatic code overlaying techniques. We are dealing with a system context that imposes two distinct constraints: (1) no hardware support for address translation and (2) a spatially and temporally coarse grained faulting mechanism at the function level. Our approach addresses those two constraints as efficiently as possible. Our technique statically computes the Worst-Case Number of Conflict misses (WCNC) between two different code segments using path expressions. Then, it constructs a static temporal relationship graph with the WCNCs and emits an overlay structure for a given higher-level memory size. We also propose an inter-procedural partial redundancy elimination technique that minimizes redundant code copying caused by the generated overlay structure. Experimental results show that our approach is promising. Choonki Jang, Jaejin Lee, Bernhard Egger 0002, Soojung Ryu |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | An instruction-scheduling-aware data partitioning technique for coarse-grained reconfigurable architecturesabstractIn this paper, we propose a data partitioning technique for the memory subsystem that consists of a multi-ported scratchpad memory (SPM) unit and a single-ported data cache in coarse-grained reconfigurable arrays (CGRA) architecture. The embedded reconfigurable processor executes programs by switching between the Non-VLIW and VLIW modes depending on the type of the code region to achieve high performance. The VLIW mode exploits code regions with high ILP that require high memory bandwidth and the Non-VLIW mode exploits those with low ILP that require low memory latency. Our data partitioning technique between the SPM and the data cache is based on data interference graph reduction and profiling information. Given an SPM size, it finds the optimal data partitions by taking the VLIW instruction schedule into consideration. We evaluate our data partitioning technique for the CGRA architecture with three representative multimedia applications. Choonki Jang, Jaejin Lee, Hee-Seok Kim, Donghoon Yoo, Sukjin Kim, Hongseok Kim, Soojung Ryu |
LCTES | 8 |