EDBT 2026 Demo / reviewers in the wild / expert
Shin-Dug Kim
dblp:46/1853
· DBLP profile ↗
71ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0002-2642-6662ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 12Graphics, computer vision, multimedia, augmented reality and games · 6Artificial intelligence and machine learning · 4Human-computer interaction and ubiquitous computing · 4Computer networks · 3Software engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Storage systems · 68% Memory systems · 29% Distributed systems · 4% | |
| Artificial intelligence
1 paper |
Robot navigation and mapping · 100% | |
| Computer graphics and multimedia
2 papers |
Virtual and augmented reality · 82% Rendering · 18% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 50% Collaborative and social computing · 50% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot navigation and mapping
SLAM |
0.4 | 1 | 2019 | Real-Time Visual-Inertial SLAM Based on Adaptive Keyframe Selection for Mobile AR Applications · IEEE Trans. Multim. 2019 |
Robotics › Robot navigation and mapping › SLAM › multi-sensor SLAM
visual-inertial SLAM |
0.4 | 1 | 2019 | Real-Time Visual-Inertial SLAM Based on Adaptive Keyframe Selection for Mobile AR Applications · IEEE Trans. Multim. 2019 |
Virtual and augmented reality › augmented reality
mobile augmented reality |
0.4 | 1 | 2019 | Real-Time Visual-Inertial SLAM Based on Adaptive Keyframe Selection for Mobile AR Applications · IEEE Trans. Multim. 2019 |
Memory systems
non-volatile memory |
0.2 | 1 | 2015 | A New Memory-Disk Integrated System with HW Optimizer · ACM Trans. Archit. Code Optim. 2015 |
Storage systems
non-volatile memory storage |
0.2 | 1 | 2015 | A New Memory-Disk Integrated System with HW Optimizer · ACM Trans. Archit. Code Optim. 2015 |
Storage systems › data reduction
write reduction |
0.2 | 1 | 2015 | A New Memory-Disk Integrated System with HW Optimizer · ACM Trans. Archit. Code Optim. 2015 |
Storage systems › i/o workload characterization
access pattern classification |
0.1 | 1 | 2012 | A Pattern Adaptive NAND Flash Memory Storage Structure · IEEE Trans. Computers 2012 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.1 | 1 | 2012 | A Pattern Adaptive NAND Flash Memory Storage Structure · IEEE Trans. Computers 2012 |
Storage systems › flash and SSD
solid-state drive |
0.1 | 1 | 2012 | A Pattern Adaptive NAND Flash Memory Storage Structure · IEEE Trans. Computers 2012 |
Storage systems › storage performance
write performance |
0.1 | 1 | 2012 | A Pattern Adaptive NAND Flash Memory Storage Structure · IEEE Trans. Computers 2012 |
Robotics › Robot navigation and mapping
visual odometry |
0.1 | 1 | 2019 | Real-Time Visual-Inertial SLAM Based on Adaptive Keyframe Selection for Mobile AR Applications · IEEE Trans. Multim. 2019 |
Ubiquitous computing and smart environments › mobile computing
mobile media sharing |
0.1 | 1 | 2009 | A Profile-Based Multimedia Sharing Scheme With Virtual Community, Based on Personal Space in a Ubiquitous Computing Environment · IEEE Trans. Multim. 2009 |
Collaborative and social computing
online communities |
0.1 | 1 | 2009 | A Profile-Based Multimedia Sharing Scheme With Virtual Community, Based on Personal Space in a Ubiquitous Computing Environment · IEEE Trans. Multim. 2009 |
Operating systems › resource management › memory management
virtual memory |
0.1 | 1 | 2015 | A New Memory-Disk Integrated System with HW Optimizer · ACM Trans. Archit. Code Optim. 2015 |
Rendering
visibility culling |
0.1 | 1 | 2006 | An Effective Visibility Culling Method Based on Cache Block · IEEE Trans. Computers 2006 |
Storage systems › flash and SSD › flash memory
flash storage |
0.0 | 1 | 2012 | A Pattern Adaptive NAND Flash Memory Storage Structure · IEEE Trans. Computers 2012 |
Memory systems
cache |
0.0 | 1 | 2003 | An Intelligent Cache System with Hardware Prefetching for High Performance · IEEE Trans. Computers 2003 |
Memory systems › cache management
cache replacement |
0.0 | 1 | 2003 | An Intelligent Cache System with Hardware Prefetching for High Performance · IEEE Trans. Computers 2003 |
Memory systems › cache › prefetching
hardware prefetching |
0.0 | 1 | 2003 | An Intelligent Cache System with Hardware Prefetching for High Performance · IEEE Trans. Computers 2003 |
Memory systems › cache
prefetching |
0.0 | 1 | 2003 | An Intelligent Cache System with Hardware Prefetching for High Performance · IEEE Trans. Computers 2003 |
Distributed systems › peer-to-peer systems
distributed hash table |
0.0 | 1 | 2009 | A Profile-Based Multimedia Sharing Scheme With Virtual Community, Based on Personal Space in a Ubiquitous Computing Environment · IEEE Trans. Multim. 2009 |
Distributed systems
peer-to-peer systems |
0.0 | 1 | 2009 | A Profile-Based Multimedia Sharing Scheme With Virtual Community, Based on Personal Space in a Ubiquitous Computing Environment · IEEE Trans. Multim. 2009 |
Rendering › graphics pipeline
rasterization pipeline |
0.0 | 1 | 2006 | An Effective Visibility Culling Method Based on Cache Block · IEEE Trans. Computers 2006 |
Methods — techniques the papers use, named apart from their topics
inertial measurement unit · 0.8adaptive keyframe selection · 0.8hardware-software co-design · 0.4buffer management · 0.4profile-based community construction · 0.2locality-based sharing · 0.2write cache · 0.1pattern adaptive structure · 0.1prefetching · 0.1spatial buffer · 0.0selective-mode · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Dynamic recognition prefetch engine for DRAM-PCM hybrid main memory
Mengzhao Zhang, Jeong-Geun Kim, Su-Kyung Yoon, Shin-Dug Kim |
J. Supercomput. | 4 |
| 2021 | History table-based linear analysis method for DRAM-PCM hybrid memory system
Jeong-Geun Kim, Yoon-Su Jo, Su-Kyung Yoon, Shin-Dug Kim |
J. Supercomput. | 4 |
| 2020 | Algorithm-Switching-Based Last-Level Cache Structure with Hybrid Main Memory ArchitectureabstractIn this research, we designed an algorithm-switching (AS)-based last-level cache (LLC) structure with DRAM-NAND Flash hybrid main memory architecture. In order to take full advantage of previous memory access patterns and achieve high performance in the upper level of memory hierarchy, an AS-based clustering engine that uses k-means, k-medoids and k-center clustering algorithms was applied to LLC. The proposed LLC consists of three major parts, namely a set-divisible cache, and victim and clustering buffers. The victim and clustering buffers efficiently managed the history of cache blocks evicted from the set-divisible cache through the AS-based engine mechanism. The experimental results that were evaluated using Redis application and YCSB benchmark show that compared with conventional LLC structure, the proposed AS-based LLC structure could reduce the total execution time by 19.50%, power consumption by 16.31%, and NAND-Flash memory write count by 8.6%. Xian-Shu Li, Su-Kyung Yoon, Jeong-Geun Kim, Bernd Burgstaller, Shin-Dug Kim |
Comput. J. | 5 |
| 2020 | Pattern analysis based data management method and memory-disk integrated system for high performance computing
Su-Kyung Yoon, Shin-Dug Kim |
Future Gener. Comput. Syst. | 2 |
| 2020 | Access pattern-based high-performance main memory system for graph processing on single machines
Jitae Yun, Su-Kyung Yoon, Jeong-Geun Kim, Shin-Dug Kim |
Future Gener. Comput. Syst. | 4 |
| 2020 | On-body wearable device localization with a fast and memory efficient SVM-kNN using GPUs
Quanzhe Li, Sae-Byuk Shin, Chung-Pyo Hong, Shin-Dug Kim |
Pattern Recognit. Lett. | 4 |
| 2020 | Effective data prediction method for in-memory database applications
Jitae Yun, Su-Kyung Yoon, Jeong-Geun Kim, Shin-Dug Kim |
J. Supercomput. | 4 |
| 2019 | Self-learnable Cluster-based Prefetching Method for DRAM-Flash Hybrid Main Memory ArchitectureabstractThis article presents a novel prefetching mechanism for memory-intensive workloads used in large-scale data centers. We design a negative-AND-flash/dynamic random-access memory (DRAM) hybrid memory architecture as a cost-effective memory architecture to resolve the scalability and power consumption problems of a DRAM-based model. A smart prefetching mechanism based on a cluster-management scheme to cope with dynamically varying and complex access patterns of any given application is designed for maximizing the performance of the DRAM. In this article, we propose a new concept for page management, called a cluster, which prefetches data in our hybrid memory architecture. The cluster management is based on a self-learning scheme on dynamically changeable access patterns by considering any correlation between missed pages. Experimental results show that the overall performance is significantly improved in relation to hit rate, execution time, and energy consumption. Namely, our proposed model can enhance the hit rate by 15% and reduce the execution time by 1.75 times. In addition, we can save energy consumption by around 48% by cutting the number of flushed pages to about an eighth of that in a conventional system. Su-Kyung Yoon, Young-Sun Youn, Bernd Burgstaller, Shin-Dug Kim |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2019 | Real-Time Visual-Inertial SLAM Based on Adaptive Keyframe Selection for Mobile AR ApplicationsabstractSimultaneous localization and mapping (SLAM) technology is used in many applications, such as augmented reality (AR)/virtual reality, robots, drones, and self-driving vehicles. In AR applications, rapid camera motion estimation, actual size, and scale are important issues. In this research, we introduce a real-time visual-inertial SLAM based on an adaptive keyframe selection for mobile AR applications. Specifically, the SLAM system is designed based on the adaptive keyframe selection visual-inertial odometry method that includes the adaptive keyframe selection method and the lightweight visual-inertial odometry method. The inertial measurement unit data are used to predict the motion state of the current frame and it is judged whether or not the current frame is a keyframe by an adaptive selection method based on learning and automatic setting. Relatively unimportant frames (not a keyframe) are processed using a lightweight visual-inertial odometry method for efficiency and real-time performance. We simulate it in a PC environment and compare it with state-of-the-art methods. The experimental results demonstrate that the mean translation root-mean-square error of the keyframe trajectory is 0.067 m without the ground-truth scale matching, and the scale error is 0.58% with the EuRoC dataset. Moreover, the experimental results of the mobile device show that the performance is improved by 34.5%-53.8% using the proposed method. Jin-Chun Piao, Shin-Dug Kim |
IEEE Trans. Multim. | 2 |
| 2018 | Self-Adaptive Filtering Algorithm with PCM-Based Memory Storage SystemabstractThis article proposes a new phase change memory– (PCM) based memory storage architecture with associated self-adaptive data filtering for various embedded devices to support energy efficiency as well as high computing power. In this approach, PCM-based memory storage can be used as working memory and mass storage layers simultaneously, and a self-adaptive data filtering module composed of small DRAM dual buffers was designed to improve unfavorable PCM features, such as asymmetric read/write access latencies and limited endurance and enhance spatial/temporal localities. In particular, the self-adaptive data filtering algorithm enhances data reusability by screening potentially high reusable data and predicting adequate lifetime of those data depending on current victim time decision value. We also propose the possibility that a small amount of DRAM buffer is embedded into mobile processors, keeping this as small as possible for cost effectiveness and energy efficiency. Experimental results show that by exploiting a small amount of DRAM space for dual buffers and using the self-adaptive filtering algorithm to manage them, the proposed system can reduce execution time by a factor of 1.9 compared to the unified conventional model with same the DRAM capacity and can be considered comparable to 1.5× DRAM capacity. Su-Kyung Yoon, Jitae Yun, Jung-Geun Kim, Shin-Dug Kim |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Adaptive correlated prefetch with large-scale hybrid memory system for stream processing
Sung Min Lee, Su-Kyung Yoon, Jeong-Geun Kim, Shin-Dug Kim |
J. Supercomput. | 4 |
| 2018 | Design of DRAM-NAND flash hybrid main memory and Q-learning-based prefetching method
Su-Kyung Yoon, Young-Sun Youn, Jeong-Geun Kim, Shin-Dug Kim |
J. Supercomput. | 4 |
| 2017 | Intelligent clustering guided adaptive prefetching and buffer management for stream processingabstractReal-time stream data processing is required to handle big data with low latency to process a large amount of incessant streaming data. We propose a clustering based inteligent prefetching scheme and its associated DRAM-PCM (phase change memory) hybrid memory management policy especially for stream processing. To alleviate stream processing's burden of requests accessing main memory, an intelligence prefetching technique is designed to reflect stream processing behavior to reduce the amount of memory access by improving buffer hit ratio. Because stream processing has to guarantee relatively small latency to the users, it is especially important for stream processsing to process the stream data in strict time. By using a hybrid memory structure, we take such advantages like high performance, low energy consumption, and memory scalability. And by using clustering based smart prefetching, we could improve buffer hit rate and because of this, overall system performance can be enhanced. Our proposed architecture and clustering based prefetching method can improve system performance by 1.15 times, compared with hybrid memory and buffer architecture without any prefetching scheme of conventional model and also energy consumption by 1.23 times, compared with DRAM only conventional model. Sung Min Lee, Young-Sun Youn, Su-Kyung Yoon, Shin-Dug Kim |
SMC | 4 |
| 2017 | Effective lazy training method for deep q-network in obstacle avoidance and path planningabstractDeep reinforcement learning technique combines reinforcement learning and neural network for various applications. This paper is to propose an effective lazy training method for deep reinforcement learning, especially for deep Q-network combining neural network with Q-learning to be used for the obstacle avoidance and path planning applications. The proposed method can reduce the overall training time by designing a lazy learning method and a method removing unnecessary repetitions in the training step. These two methods can reduce a significant portion of total execution time without losing any required accuracy. The proposed method is evaluated for the obstacle avoidance and path planning tasks, where an agent trapped in an unknown environment is trying to find out the shortest path to the destination without any collision, through its self-study. And the experiment results show that the proposed method reduces 53.38% of training time on average, compared to the traditional method with no performance loss and make the training procedure more stable. Seabyuk Shin, Cheong-Ghil Kim, Shin-Dug Kim |
SMC | 4 |
| 2017 | Hot-cold data filtering and management for PRAM based memory-storage unified systemabstractNext generation non-volatile memory technologies draw attention recently as promising candidates for replacing conventional DRAM-based memory devices. By using the favorable features such as low energy consumption, non-volatility, and good scalability, the next generation non-volatile memory devices can replace conventional main memory and secondary storage layers. This paper proposes a hot-cold data filtering adapter with it management algorithm for the memory-storage unified system which uses phase change RAM (PRAM) devices. The proposed hot-cold data filtering adapter consisting of dual decoupled adaptive buffers can improve drawbacks of PRAM, such as slower read/write access latencies and limited life time compared to DRAM, enhance spatial/temporal localities, and optimize the data reusability by analyzing data access pattern. Experimental results show that the proposed architecture and its algorithm improves the miss rate by about 69.2% and write count is decreased by 96.4% compared with a uniform buffer of the same size. Su-Kyung Yoon, Kwang-Su Jung, Xian-Shu Li, Shin-Dug Kim |
SMC | 5 |
| 2017 | Mobile Unified Memory-Storage Structure Based on Hybrid Non-Volatile MemoriesabstractIn mobile computing systems, the limited amount of main memory space leads to page swap operation overhead and data duplication in both main memory and secondary storage. Furthermore, SQLite write operations in mobile devices such as smartphones and tablet PCs tend to frequently overwrite data to storage, significantly degrading performance. Thus, this article presents a unified memory-storage structure that is optimized for mobile devices and blurs the boundary between the existing main memory layer and secondary storage layer. This structure can eliminate the conventional page-swap operations that cause significant performance degradation and support fast program execution time. The unified memory-storage structure consists of a dynamic RAM (DRAM) and phase change memory (PCM) -based dual buffering module, a hybrid unified memory-storage array consisting of DRAM and NAND Flash memory, and an associated unified storage translation layer devised for the memory address and file translation mechanism as a system software module. This hybrid array of non-volatile memories is formed as a single memory-disk integrated storage space that can be logically divided into static and dynamic spaces. Experimental results show that the overall performance of the hybrid unified memory-storage system with the buffering structure increases by around 13% and power consumption is also improved by 35%, compared to current mobile system. Su-Kyung Yoon, Young-Sun Youn, Kihyun Park, Shin-Dug Kim |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2017 | Cloud computing burst system (CCBS): for exa-scale computing system
Young-Sun Youn, Su-Kyung Yoon, Shin-Dug Kim |
J. Supercomput. | 3 |
| 2016 | Dynamic partitioning-based JPEG decompression on heterogeneous multicore architecturesabstractSummary With the emergence of social networks and improvements in computational photography, billions of JPEG images are shared and viewed on a daily basis. Desktops, tablets, and smartphones constitute the vast majority of hardware platforms used for displaying JPEG images. Despite the fact that these platforms are heterogeneous multicores, no approach exists yet that is capable of joining forces of a system's CPU and graphics processing unit (GPU) for JPEG decoding. In this paper, we introduce a novel JPEG decoding scheme for heterogeneous architectures consisting of a CPU and a general‐purpose GPU. We employ an offline profiling step to determine the performance of a system's CPU and GPU with respect to JPEG decoding. For a given JPEG image, our performance model uses: (1) the CPU and GPU performance characteristics, (2) the image entropy, and (3) the width and height of the image to balance the JPEG decoding workload on the underlying hardware. Our run‐time partitioning and scheduling scheme exploits task, data, and pipeline parallelism by scheduling the non‐parallelizable entropy‐decoding task on the CPU, whereas inverse discrete cosine transformations, color conversions, and upsampling are conducted on both the CPU and the GPU. We have implemented the proposed method in the context of the libjpeg‐turbo library, which is an industrial‐strength JPEG encoding and decoding engine. Libjpeg‐turbo's hand‐optimized SIMD routines for ARM and x86 architectures constitute a competitive yardstick for the comparison with the proposed approach. We have evaluated our approach for a total of 7194 JPEG images across four high‐end and middle‐end CPU–GPU combinations including a mobile GPU. We achieve speedups of up to 5.2× over the SIMD version of libjpeg‐turbo, and speedups of up to 10.5× over its sequential code. Taking into account the non‐parallelizable JPEG entropy‐decoding part, our approach achieves up to 97% of the theoretically attainable maximal speedup, with an average of 94%. Copyright © 2015 John Wiley & Sons, Ltd. Wasuwee Sodsong, Jingun Hong, Seongwook Chung, Yeong-Kyu Lim, Shin-Dug Kim, Bernd Burgstaller |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | A Unified Buffering Management with Set Divisible Cache for PCM Main Memory
Mei-Ying Bian, Su-Kyung Yoon, Jeong-Geun Kim, Sangjae Nam, Shin-Dug Kim |
J. Comput. Sci. Technol. | 5 |
| 2016 | Improving performance on object recognition for real-time on mobile devices
Jin-Chun Piao, Hyeon-Sub Jung, Chung-Pyo Hong, Shin-Dug Kim |
Multim. Tools Appl. | 4 |
| 2015 | Fast bootstrapping method for the memory-disk integrated memory systemabstractThis research is to design system software modules for main memory-disk integrated system. To utilize many useful features of recent non-volatile memory, a horizontal memory hierarchy should be considered as a new approach. Main objective of this study is to design an optimized bootstrap method for the integrated memory-disk system, along with compatibility with conventional bootloader and operating system. In this memory-disk integrated system, stored files and data are accessed in place reducing duplication between the main memory and storage. Thus, previous booting sequence of loading bootstrap and operating system from secondary storage to main memory should be changed to reflect the architectural change. Our proposed technique provides compatibility with conventional system software and fast booting time by utilizing prominent advantages of the integrated memory-disk system and optimizing booting sequence for this system. To implement this method, we modified BIOS (basic input/output system) to change booting from integrated memory-disk system. We used QEMU machine emulator to simulate the proposed booting method. Our experimental results show that the proposed bootstrap method can reduce the number of instructions to execute by around ten millions, resulting in fast boot performance and keeping the core of the bootloader. Sangjae Nam, Su-Kyung Yoon, Shin-Dug Kim |
ICIS | 3 |
| 2015 | Selective data buffering module for unified hybrid storage systemabstractA new design methodology is aggressively considered on conventional memory hierarchy structure because of newly emerging non-volatile memory technologies. This paper presents a unified single memory structure that merges existing main memory layer and secondary storage layer together. This structure eliminates conventional page swap operations causing significant performance degradation and supports fast program execution time. The unified single memory structure consists of a selective data buffering module and a hybrid array of PCM (phase change memory) and NAND Flash storage. This hybrid array of non-volatile memories is formed as a single memory-disk integrated storage space which can be logically divided into static and dynamic spaces. This research is to design a selective data buffering structure to reduce asymmetric read/write latencies and increase the lifetime of hybrid storage based on PCM and NAND Flash. The proposed structure is compared with the interfacing adapter structure of the memory-disk integrated system. Experimental results show that the hit rate of the selective data buffering structure increases by 90.5%, compared with the NVPO under the memory-disk integrated system [10], which shows 76.9% hit rate. The access latency is also improved by around 20.6%. Kihyun Park, Su-Kyung Yoon, Shin-Dug Kim |
ICIS | 3 |
| 2015 | Accelerating the pre-processing stages of JPEG encoder on a heterogenous system using OpenCLabstractColor space conversion and downsampling are among the major computationally intensive steps in typical image and video codec standards, and accelerating these steps will improve the performances of these applications significantly. In this paper, we describe the parallel implementation of the color space conversion and downsampling as pre-processing steps for the JPEG encoder in a heterogeneous environment using the most recent cross-platform Open Computing Language (OpenCL). This work combines a multi-core CPU and a many-core GPU in a single solution to perform the computation of the JPEG encoder pre-processing stages. In comparing with CPU-based implementation, our OpenCL parallel implementation results in an increase in the speed of the computations by factors of 8.78 on both CPU and GPU devices. Nasser Alqudami, Shin-Dug Kim |
SNPD | 2 |
| 2015 | Data Classification Management with its Interfacing Structure for Hybrid SLC/MLC PRAM Main MemoryabstractThis research aims to design a new phase-change RAM (PRAM)-based main memory structure, supporting the advantages of PRAM while providing performance similar to that of conventional DRAM main memory. To replace conventional DRAMs with non-volatile PRAMs as the main memory components, comparable memory access latency, overall cost, power dissipation, and memory cell endurance should be supported. For these goals, we propose a new main memory system consisting of a DRAM converter and an array of single-level cell (SLC)/multi-level cell (MLC) PRAMs. The DRAM converter consists of an aggressive fetching superblock buffer to assure better use of spatial locality and a selective filtering buffer for better use of temporal locality. The array of the SLC/MLC hybrid PRAM structure includes a combination of SLC and MLC PRAMs to enhance the lifetime of the MLC PRAM main memory and hide asymmetric read/write access latency. The proposed structure is evaluated by a trace-driven simulator using SPEC CPU 2006 and SPLASH-2 traces. Experimental results show that the proposed DRAM converter can reduce the miss rate by ∼37% and write count by ∼55% in comparison with the uniform buffer case. Also, the SLC/MLC PRAM with DRAM converter shows performance in terms of access latency and power consumption close to that of the conventional memory architecture. Thus, our proposed memory architecture can be used to replace the current DRAM-based main memory system. Sung-In Jang, Su-Kyung Yoon, Kihyun Park, Gi-Ho Park, Shin-Dug Kim |
Comput. J. | 5 |
| 2015 | A polymorphic service management scheme based on virtual object for ubiquitous computing environment
Chung-Pyo Hong, Cheong-Ghil Kim, Kuinam J. Kim, Shin-Dug Kim |
Multim. Tools Appl. | 4 |
| 2015 | Design of configurable I/O pin control block for improving reusability in multimedia SoC platforms
Myoung-Seo Kim, Cheong-Ghil Kim, Shin-Dug Kim, Jean-Luc Gaudiot |
Multim. Tools Appl. | 3 |
| 2015 | Advanced feature point transformation of corner points for mobile object recognition
Xiyuan Yin, Chung-Pyo Hong, Cheong-Ghil Kim, Kuinam J. Kim, Shin-Dug Kim |
Multim. Tools Appl. | 6 |
| 2015 | A New Memory-Disk Integrated System with HW OptimizerabstractCurrent high-performance computer systems utilize a memory hierarchy of on-chip cache, main memory, and secondary storage due to differences in device characteristics. Limiting the amount of main memory causes page swap operations and duplicates data between the main memory and the storage device. The characteristics of next-generation memory, such as nonvolatility, byte addressability, and scaling to greater capacity, can be used to solve these problems. Simple replacement of secondary storage with new forms of nonvolatile memory in a traditional memory hierarchy still causes typical problems, such as memory bottleneck, page swaps, and write overhead. Thus, we suggest a single architecture that merges the main memory and secondary storage into a system called a Memory-Disk Integrated System (MDIS). The MDIS architecture is composed of a virtually decoupled NVRAM and a nonvolatile memory performance optimizer combining hardware and software to support this system. The virtually decoupled NVRAM module can support conventional main memory and disk storage operations logically without data duplication and can reduce write operations to the NVRAM. To increase the lifetime and optimize the performance of this NVRAM, another hardware module called a Nonvolatile Performance Optimizer (NVPO) is used that is composed of four small buffers. The NVPO exploits spatial and temporal characteristics of static/dynamic data based on program execution characteristics. Enhanced virtual memory management and address translation modules in the operating system can support these hardware components to achieve a seamless memory-storage environment. Our experimental results show that the proposed architecture can improve execution time by about 89% over a conventional DRAM main memory/HDD storage system, and 77% over a state-of-the-art PRAM main memory/HDD disk system with DRAM buffer. Also, the lifetime of the virtually decoupled NVRAM is estimated to be 40% longer than that of a traditional hierarchy based on the same device technology. Do-Heon Lee, Su-Kyung Yoon, Jung-Geun Kim, Charles C. Weems, Shin-Dug Kim |
ACM Trans. Archit. Code Optim. | 5 |
| 2014 | Performance optimization of 3D applications by OpenGL ES library hooking in mobile devicesabstractThe mobile GPU (Graphic Processing Unit) market has grown steadily due to expansion of the mobile game industry. Despite the rapid computation capability of mobile devices, handling a large amount of high-quality graphics in real-time is difficult. Therefore, effective technologies for improving mobile GPU in smartphones are required. In this thesis, we examine the trade-off between quality and performance, and address the benefits of graphic performance improvement by degrading quality. To implement this idea, we propose performance optimization methodologies for 3D applications using an OpenGL ES library hooking method. Our methodologies do not require any source code from 3D applications, and can be applied to any Android phones that use OpenGL ES in real-time. To demonstrate the benefits of our methodology, we conducted performance verifications of five well-known benchmarks using a smartphone, and measured the quality in accordance with each methodology. In addition, we showed the optimal trade-offs between quality and performance. By using the proposed technique, the performance of mobile GPU can be significantly improved to achieve a better trade-off between quality and performance. Chang-Woo Cho, Chung-Pyo Hong, Jin-Chun Piao, Yeong-Kyu Lim, Shin-Dug Kim |
ICIS | 5 |
| 2014 | Designing virtual accessing adapter and non-volatile memory management for memory-disk integrated systemabstractThis paper presents a new memory hierarchy system using next-generation non-volatile memory devices merging a conventional main memory layer and a disk storage layer into a single memory layer, called a memory-disk integrated system (MDIS). In the MDIS, files are stored in MDIS storage and accessed in place as if they are in conventional DRAM main memory without any duplication between the main memory and secondary storage. The MDIS comprises three components: a virtual accessing adapter (VAA), an array of non-volatile hybrid memory (NVHM), and its associated virtual memory management module. The NVHM can be divided into two spaces-specifically, static and dynamic spaces-to support conventional memory management. Further, the VAA comprises four small buffers that support static and dynamic virtualization according to program execution characteristics. Experimental results show that the read/write load traffics for the proposed MDIS can be reduced by about 73% as compared with the cases of conventional main memory system. In addition, access latency of non-volatile hybrid memory is reduced by 51%. Su-Kyung Yoon, Do-Heon Lee, Shin-Dug Kim |
ICIS | 4 |
| 2012 | A Pattern Adaptive NAND Flash Memory Storage StructureabstractTo enhance performance of flash memory-based solid state disk (SSD), large logically chained blocks can be assembled by binding adjacent flash blocks across several flash memory chips. However, flash memory does not allow in-place overwriting and thus the operations that merge writes on these blocks suffer a visible decrease in performance. Furthermore, when small random writes are spread over the disk address space, performance tends to be degraded significantly. We thus present a technique to manage random writes efficiently to achieve stable SSD performance. In this paper, we propose a pattern adaptive SSD structure, which classifies access patterns as either random or sequential. The structure primarily consists of a write cache and a flash translation layer that separates groups of writes by access pattern (S-FTL). Separately managing the two types of write patterns enables greater parallelism and reduces the cost of large block management, thus enhancing the performance of the proposed SSD. Simulation experiments show that the proposed pattern adaptive structure can provide 39 percent decrease in extra flash block erase overhead on the average, and write performance can be improved by around 60 percent, compared with a basic FTL applied to existing parallel SSD structures. Seung-Ho Park, Jung-Wook Park, Shin-Dug Kim, Charles C. Weems |
IEEE Trans. Computers | 3 |
| 2010 | An instruction-systolic programmable shader architecture for multi-threaded 3D graphics processing
Jung-Wook Park, Hoon-Mo Yang, Gi-Ho Park, Shin-Dug Kim, Charles C. Weems |
J. Parallel Distributed Comput. | 4 |
| 2009 | A Profile-Based Multimedia Sharing Scheme With Virtual Community, Based on Personal Space in a Ubiquitous Computing EnvironmentabstractFor ubiquitous computing environments, an important parameter is whether all the components in the specific environment can connect with one another. Given this capability, we can share various kinds of content across mobile terminals. This paper introduces an effective scheme to manage multimedia sharing based on specially designed profiles and a virtual community. A virtual community is defined as any specific group of users connected for a common interest. Specifically the proposed scheme consists of two layers, i.e., a community construction layer and a multimedia sharing layer, based on personal spaces, which are responsible for constructing and managing the multimedia sharing community. The community construction layer, which is designed to be run on the mobile terminals, provides an effective way to find community members simultaneously, based on specially designed profiles, such as a user profile and an abstract profile. The multimedia sharing layer is responsible for sharing multimedia content, and is constructed as a specially designed scheme based on locality. The proposed scheme provides an effective multimedia sharing mechanism within a community. Simulation results show that the number of messages and the time required for community member discovery is reduced by 32% and 73%, respectively, in comparison with the conventional DHT-based scheme. The approach also reduces the time to exchange content by 50% with respect to the same baseline. Chung-Pyo Hong, Eo-Hyung Lee, Charles C. Weems, Shin-Dug Kim |
IEEE Trans. Multim. | 4 |
| 2008 | An effective vertical handoff scheme based on service management for ubiquitous computing
Chung-Pyo Hong, Charles C. Weems, Shin-Dug Kim |
Comput. Commun. | 3 |
| 2008 | A small data cache for multimedia-oriented embedded systems
Cheong-Ghil Kim, Jung-Wook Park, Shin-Dug Kim |
J. Syst. Archit. | 4 |
| 2008 | An Effective Model and Management Scheme of Personal Space for Ubiquitous Computing ApplicationsabstractIn ubiquitous computing, the computing environment for a user is no longer a fixed computer, but a space that includes multiple heterogeneous devices that can change dynamically according to the user's situation. Managing the space is an essential part of ubiquitous computing because application services in this environment need to be adaptive to the users' current situation. However, previous approaches oversimplified the model of personal space and demonstrated some limitations in developing user-centric adaptive services. In this paper, we propose an effective personal space model, defined as virtual personal world (VPW), and a sophisticated method to manage personal spaces. The VPW represents a personal space by using a set of stateful elements and their relationships, which are denoted as virtual objects, services, and neighbors. The VPW provides expressive and accurate information for a particular user, thereby helping application services adapt their operations for the user dynamically. Our conceptual model is designed as personal operating middleware software that manages the user's VPW and provides application services. Experimental results show that the prototype system based on VPW has reasonable performance in running application services and managing personal spaces. We also found that the VPW model can increase the average user satisfaction rate by up to 40% compared to other models in our simulation environment. Kyung-Lang Park, Joo-Kyoung Park, Shin-Dug Kim |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2006 | A Profile Based Vertical Handoff Scheme for Ubiquitous Computing Environment
Chung-Pyo Hong, Tae-Hoon Kang, Shin-Dug Kim |
APNOMS | 3 |
| 2006 | A Seamless Service Management with Context-Aware Handoff Scheme in Ubiquitous Computing Environment
Tae-Hoon Kang, Chung-Pyo Hong, Won-Joo Jang, Shin-Dug Kim |
APNOMS | 4 |
| 2006 | Practice and Experience of an Embedded Processor Core Modeling
Gi-Ho Park, Sung Woo Chung, Han-Jong Kim, Jung-Bin Im, Jung-Wook Park, Shin-Dug Kim, Sung-Bae Park |
HPCC | 6 |
| 2006 | VPW: An Effective Personal Space Model for Providing UbiquitousabstractUser-centrism is one of the most important things in ubiquitous computing. Previously, lots of researches have addressed issues for ubiquitous computing, but most of them oversimplified a personal space as a set of devices in a fixed location or in a communicationranged space. Thus, they have had a problem to describe sophisticated user-centric services. In this paper, we propose an effective personal space model called VPW (Virtual Personal World) for providing ubiquitous services. VPW includes various objects over several locations, tasks in progress, and proxies of other users. Thus, service provision based on VPW supports more delicate user-centric adaptation, use of multiple devices spread over locations, and harmony between multiple users and multiple services. Eventually, it increases user satisfaction in providing ubiquitous services. Experimental results show the VPW-based approach increase user satisfaction by around 20% compared to the location-based approach and by around 15% compared to the range-based approach. Kyung-Lang Park, Joo-Kyoung Park, Chang-Deok Kang, Hoon-Ki Lee, Eui-Hyun Baek, Shin-Dug Kim |
SERA | 6 |
| 2006 | An Effective Visibility Culling Method Based on Cache BlockabstractAs the complexity of 3D scenes is on the increase, the search for an effective visibility culling method has become one of the most important issues to be addressed in the design of 3D rendering processors. Here, we propose a new rasterization pipeline with visibility culling; the proposed architecture performs the visibility culling at an early stage of the rasterization pipeline (especially at the traversal stage) by retrieving data in a pixel cache without any significant hardware logics such as the hierarchical z-buffer. If the data to be retrieved does not exist in the pixel cache, the proposed architecture performs a prefetch operation in order to reduce the miss penalty of the pixel cache. That is, the cache miss penalty can be reduced as the transfer of a missed cache block from the frame memory into the pixel cache can be handled simultaneously with the rasterization pipeline executions. Simulation results show that the proposed architecture can achieve a performance gain of about 32% compared with the conventional pretexturing architecture and about 7% compared to the hierarchical z-buffer visibility scheme. Moon-Hee Choi, Woo-Chan Park, Francis Neelamkavil, Tack-Don Han, Shin-Dug Kim |
IEEE Trans. Computers | 5 |
| 2005 | A Programmable Context Interface to Build a Context Infrastructure for Worldwide Smart Applications
Kyung-Lang Park, Chang-Soon Kim, Chang-Duk Kang, Shin-Dug Kim |
EUC | 4 |
| 2005 | A Personalized and Scalable Service Broker for the Global Computing Environment
Kyung-Lang Park, Chang-Soon Kim, Oh-Young Kwon, Hyoung-Woo Park, Shin-Dug Kim |
ISPA | 5 |
| 2005 | A new NAND-type flash memory package with smart buffer system for spatial and temporal localities
Gi-Ho Park, Shin-Dug Kim |
J. Syst. Archit. | 3 |
| 2004 | Power-Aware Deterministic Block Allocation for Low-Power Way-Selective Cache StructureabstractThis paper proposes a power-aware cache block allocation algorithm for the way-selective set-associative cache on embedded systems to reduce energy consumption without additional delay or performance degradation. For this goal, way selection logic and specialized replacement policy are designed to enable only one way of set-associative cache as in the direct-mapped cache. Overall cache access time becomes almost the same as that of a conventional set associative cache with accessing additional way selection logic. Because data array can be accessed without waiting for tag comparison, multiplexer delay can be removed totally. The simulation result shows that the proposed architecture can reduce a per access power consumption by 59% over conventional set-associative caches with average 0.06% of negligible performance loss. Jung-Wook Park, Gi-Ho Park, Sung-Bae Park, Shin-Dug Kim |
ICCD | 4 |
| 2004 | A Smart Agent-Based Grid Computing Platform
Kwang-Won Koh, Hiecheol Kim, Kyung-Lang Park, Hwang-Jik Lee, Shin-Dug Kim |
ICCSA (2) | 5 |
| 2004 | A Space-Efficient On-Chip Compressed Cache Organization for High Performance Computing
Keun Soo Yim, Jang-Soo Lee, Jihong Kim 0001, Shin-Dug Kim, Kern Koh |
ISPA | 4 |
| 2003 | A selective filter-bank TLB systemabstractWe present a selective filter-bank translation lookaside buffer (TLB) system with low power consumption for embedded processors. The proposed TLB is constructed as multiple banks with a small two-bank buffer, called as a filter-bank buffer, located above its associated bank. Either a filter-bank buffer or a main bank TLB can be selectively accessed based on two bits in the filter-bank buffer. Energy savings are achieved by reducing the number of entries accessed at a time, by using filtering and bank mechanism. The overhead of the proposed TLB turns out to be negligible compared with other hierarchical structures. Simulation results show that the Energy*Delay product can be reduced by about 88% compared with a fully associative TLB, 75% with respect to a filter-TLB, and 51% relative to a banked-filter TLB. Gi-Ho Park, Sung-Bae Park, Shin-Dug Kim |
ISLPED | 4 |
| 2003 | An Intelligent Cache System with Hardware Prefetching for High PerformanceabstractWe present a high performance cache structure with a hardware prefetching mechanism that enhances exploitation of spatial and temporal locality. The proposed cache, which we call a selective-mode intelligent (SMI) cache, consists of three parts: a direct-mapped cache with a small block size, a fully associative spatial buffer with a large block size, and a hardware prefetching unit. Temporal locality is exploited by selectively moving small blocks into the direct-mapped cache after monitoring their activity in the spatial buffer for a time period. Spatial locality is enhanced by intelligently prefetching a neighboring block when a spatial buffer hit occurs. The overhead of this prefetching operation is shown to be negligible. We also show that the prefetch operation is highly accurate: Over 90 percent of all prefetches generated are for blocks that are subsequently accessed. Our results show that the system enables the cache size to be reduced by a factor of four to eight relative to a conventional direct-mapped cache while maintaining similar performance. Also, the SMI cache can reduce the miss ratio by around 20 percent and the average memory access time by 10 percent, compared with a victim-buffer cache configuration. Seh-Woong Jeong, Shin-Dug Kim, Charles C. Weems |
IEEE Trans. Computers | 3 |
| 2002 | An Advanced Filtering TLB for Low Power ConsumptionabstractThis research is to design a new two-level TLB (translation look-aside buffer) architecture that integrates a 2-way banked filter TLB with a 2-way banked main TLB. One of the main objectives is to reduce power consumption in embedded processors by distributing the accesses to the TLB entries across several banks in a balanced manner. Thus, an advanced filtering technique is devised to reduce power dissipation by adopting a sub-bank structure at the filter TLB. And also a bank-associative structure is applied to each level of the TLB hierarchy. Simulation result shows that the miss ratio and Energy*Delay product can be improved by 59.26% and 24.9%, respectively, compared with a micro TLB with 4-32 entries, and 40.81% and 12.18%, compared with a micro TLB with 16-32 entries. Jin-Hyuck Choi, Gi-Ho Park, Shin-Dug Kim |
SBAC-PAD | 4 |
| 2002 | A banked-promotion translation lookaside buffer system
Seh-Woong Jeong, Shin-Dug Kim, Charles C. Weems |
J. Syst. Archit. | 3 |
| 2002 | Application-adaptive intelligent cache memory systemabstractThis article presents the design of a simple hardware-controlled, high performance cache system. The design supports fast access time, optimal utilization of temporal and spatial localities adaptive to given applications, and a simple dynamic fetching mechanism with different fetch sizes. Support for dynamically varying the fetch size makes the cache equally effective for general-purpose as well as multimedia applications. Our cache organization and operational mechanism are especially designed to maximize temporal locality and spatial locality, selectively and adaptively. Simulation shows that the average memory access time of the proposed cache is equal to that of a conventional direct-mapped cache with eight times as much space. In addition, the simulations show that our cache achieves better performance than a 2-way or 4-way set associative cache with twice as much space. The average miss ratio, compared with the victim cache with 32-byte block size, is improved by about 41% or 60% for general applications and multimedia applications, respectively. It is also shown that power consumption of the proposed cache is around 10% to 60% lower than other cache systems that we examine. Our cache system thus offers high performance with low power consumption and low hardware cost. Shin-Dug Kim, Charles C. Weems |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2001 | A Banked-Promotion TLB for High Performance and Low PowerabstractThis research is to design a simple but high performance TLB (translation lookaside buffer) system with low power consumption. Thus, we propose a new TLB structure supporting two page sizes dynamically and selectively for high performance and low cost design without any operating system support. For high performance, a promotion-TLB is designed by supporting two page sizes. Also in order to attain low power consumption, a banked-TLB is constructed by dividing one fully associative TLB space into two sub fully associative TLBs. These two structures are integrated to form a banked-promotion TLB as a low power and high performance TLB structure for embedded processors. According to the results of comparison and analysis, a similar performance can be achieved by using fewer TLB entries and also energy dissipation can be reduced by around 50% compared with the fully associative TLB. Jang-Soo Lee, Seh-Woong Jeong, Shin-Dug Kim |
ICCD | 4 |
| 2000 | A Selective Temporal and Aggressive Spatial Cache System Based on Time IntervalabstractThis paper proposes a new cache system that can increase the effect by temporal and spatial locality by using only simple hardware control without any locality detection hardware or compiler aid. The proposed cache system consists of two caches with different associativities and different block sizes, i.e., a direct-mapped cache with small block size and a fully associative spatial buffer with large block size as a multiple of small blocks. Therefore, the spatial locality can be exploited by aggressively fetching large blocks including any missed small block into the buffer, and the temporal locality can also be exploited by selectively storing small blocks that were referenced at the spatial buffer in the past. To determine the blocks to be stored at the direct-mapped cache, the proposed cache system uses a time interval-based selection mechanism. According to the simulation results, similar performance can be achieved by using four times smaller cache size compared with the conventional direct-mapped cache. Jang-Soo Lee, Shin-Dug Kim |
ICCD | 3 |
| 2000 | A section cache system designed for VLIW architectures
Won-Kee Hong, Shin-Dug Kim |
J. Syst. Archit. | 2 |
| 2000 | Impact of the memory interface structure in the memory-processor integrated architecture for computer vision
Young-Sik Kim, Tack-Don Han, Shin-Dug Kim |
J. Syst. Archit. | 3 |
| 2000 | An on-chip cache compression technique to reduce decompression overhead and design complexity
Jang-Soo Lee, Won-Kee Hong, Shin-Dug Kim |
J. Syst. Archit. | 3 |
| 2000 | A new cache architecture based on temporal and spatial locality
Jang-Soo Lee, Shin-Dug Kim |
J. Syst. Archit. | 3 |
| 1999 | Design and Evaluation of a Selective Compressed Memory SystemabstractThis research explores any potential for an on-chip cache compression which can reduce not only cache miss ratio but also miss penalty, if main memory is also managed in compressed form. However, the decompression time causes a critical effect on the memory access time and variable-sized compressed blocks tend to increase the design complexity of the compressed cache architecture. This paper suggests several techniques to reduce the decompression overhead and to manage the compressed blocks efficiently which include selective compression, fixed space allocation for the compressed blocks, parallel decompression, the use of a decompression buffer, and so on. Moreover a simple compressed cache architecture based on the above techniques and its management method are proposed. The results from trace-driven simulation show that this approach can provide around 35% decrease in the on-chip cache miss ratio as well as a 53% decrease in the data traffic over the conventional memory systems. Also, a large amount of the decompression overhead can be reduced, and thus the average memory access time can also be reduced by maximum 20% against the conventional memory systems. Jang-Soo Lee, Won-Kee Hong, Shin-Dug Kim |
ICCD | 3 |
| 1999 | A floating point multiplier performing IEEE rounding and addition in parallel
Woo-Chan Park, Tack-Don Han, Shin-Dug Kim, Sung-Bong Yang |
J. Syst. Archit. | 3 |
| 1998 | A Dualthreaded Java Processor for Java MultithreadingabstractThe Java-Web computing paradigm has changed the Internet into a computing environment. For Java-Web computing and many Java applications, a new Java processor called simultaneous multithreaded (SMT) JavaChip, is proposed to enhance the performance of previous Java processors by hardware support of Java multithreading. SMT JavaChip is a modified architecture with the enhanced mechanism of stack cache, instruction cache, functional units, etc. It executes dual independent threads simultaneously and enhances instruction level parallelism. The performance of SMT JavaChip is evaluated through the simulation using JavaSim, a Java processor simulator. This research is focused to enhance the performance of the Java processor by considering the characteristics of the Java language and computation environment. Performance results show that SMT JavaChip can provide an execution speedup of between 1.28 and 2.00 compared with the single threaded Java processors. Chun-Mok Chung, Shin-Dug Kim |
ICPADS | 2 |
| 1998 | Design and Performance Evaluation of an Adaptive Cache Coherence ProtocolabstractIn shared-memory multiprocessor systems, the local caches which are used to tolerate the performance gap between processor and memory cause additional bus transactions to maintain the coherency of shared data. Especially, coherency misses and data traffic due to spatial locality and false sharing have a significant effect on the system performance. In this approach, an adaptive cache coherence protocol based on the sectored cache is introduced. It determines the size of a block to be migrated or invalidated dynamically, depending on the transfer mode, so that it can exploit the spatial locality and reduce useless data traffic due to false sharing at the same time. This protocol is evaluated via event-driven simulation, and its results show a 58% decrease in the data traffic and a 45% decrease in the cache miss ratio. Thus, the adaptive cache coherence protocol provides about a 56% improvement in the execution time. Won-Kee Hong, Nam-Hee Kim, Shin-Dug Kim |
ICPADS | 3 |
| 1998 | An Adaptive Parallel Computer Vision SystemabstractAn approach for designing a hybrid parallel system that can perform different levels of parallelism adaptively is presented. An adaptive parallel computer vision system (APVIS) is proposed to attain this goal. The APVIS is constructed by integrating two different types of parallel architectures, i.e. a multiprocessor based system (MBS) and a memory based processor array (MPA), tightly into a single machine. One important feature in the APVIS is that the programming interface to execute data parallel code onto the MPA is the same as the usual subroutine calling mechanism. Thus the existence of the MPA is transparent to the programmers. This research is to design an underlying base architecture that can be optimally executed for a broad range of vision tasks. A performance model is provided to show the effectiveness of the APVIS. It turns out that the proposed APVIS can provide significant performance improvement and cost effectiveness for highly parallel applications having a mixed set of parallelisms. Also an example application composed of a series of vision algorithms, from low-level and medium-level processing steps, is mapped onto the MPA. Consequently, the APVIS with a few or tens of MPA modules can perform the chosen example application in real time when multiple images are incoming successively with a few seconds inter-arrival time. Young-Sik Kim, Shin-Dug Kim, Tack-Don Han, Sung-Bong Yang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1998 | Methods to improve performance of instruction prefetching through balanced improvement of two primary performance factors
Gi-Ho Park, Oh-Young Kwon, Tack-Don Han, Shin-Dug Kim, Sung-Bong Yang |
J. Syst. Archit. | 4 |
| 1998 | Mapping of neural networks onto the memory-processor integrated architecture
Young-Sik Kim, Mi-Jung Noh, Tack-Don Han, Shin-Dug Kim |
Neural Networks | 4 |
| 1997 | An Effective Memory-Processor Integrated Architecture for Computer VisionabstractIn this paper an effective memory-processor integrated architecture, called memory based processor array (MPA), for computer vision is proposed. The MPA can be easily attached into any host system via memory interface. In order to measure the impact of the memory interface structure an analytical model is derived. The performance improvement on the proposed model for the memory interface architecture of the MPA system can be 6%/spl sim/40% for vision tasks consisting of sequential and data parallel tasks. The asymptotic time complexities of the mapping algorithms are evaluated to verify the cost-effectiveness and the efficiency of the MPA system. Young-Sik Kim, Tack-Don Han, Shin-Dug Kim, Sung-Bong Yang |
ICPP | 3 |
| 1994 | Multiple Quadratic Forms: A Case Study in the Design of Data.Parallel Algorithms
Mu-Cheng Wang, Wayne G. Nation, James B. Armstrong, Howard Jay Siegel, Shin-Dug Kim, Mark A. Nichols, Michael Gherrity |
J. Parallel Distributed Comput. | 5 |
| 1993 | Multiple Quadratic Forms: A Case Study in the Design of Scalable AlgorithmsabstractParallel implementations of the computationally intensive task of solving multiple quadratic forms (MQFs) have been examined. Coupled and uncoupled parallel methods are investigated, where coupling relates to the degree of interaction among the processors. Also, the impact of partitioning a large MQF problem into smaller non-interacting subtasks is studied. Trade-offs among the implementations fo various data-size/machine-size ratios are categorized in terms of complex arithmetic operation counts, communicatino overhead, and memory storage requirements. Mu-Cheng Wang, Wayne G. Nation, James B. Armstrong, Howard Jay Siegel, Shin-Dug Kim, Mark A. Nichols, Michael Gherrity |
ICPP (3) | 5 |
| 1991 | Impact of Temporal Juxtaposition on the Isolated Phase Optimization Approach to Mapping an Algorithm to Mixed-Mode Architectures
Thomas B. Berg, Shin-Dug Kim, Howard Jay Siegel |
ICPP (1) | 2 |
| 1991 | Limitations Imposed on Mixed-Mode Performance of Optimized Phases Due to Temporal Juxtaposition
Thomas B. Berg, Shin-Dug Kim, Howard Jay Siegel |
J. Parallel Distributed Comput. | 2 |
| 1991 | Modeling Overlapped Operation between the Control Unit and Processing Elements in an SIMD Machine
Shin-Dug Kim, Mark A. Nichols, Howard Jay Siegel |
J. Parallel Distributed Comput. | 1 |