EDBT 2026 Demo / reviewers in the wild / expert
Eun Jung Kim 0001
dblp:87/5080-1 · also Eunjung Kim 0001
· DBLP profile ↗
57ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-2590-698XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 6 first-author · 8 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compression-Aware Gradient Splitting for Collective Communications in Distributed TrainingabstractWhile distributed training is crucial for scaling deep learning models, it incurs significant overhead due to the collective communication of gradients. To alleviate the burden, compression techniques are commonly used to improve network bandwidth utilization. However, compression poses challenges for synchronized AllReduce collective communications, even more so in scalable systems. Non-uniform data sizes resulting from compression can cause bandwidth under-utilization, as faster nodes remain idle while waiting for slower nodes to complete data exchanges, increasing overall communication and consequently, training time. However, the inherent similarity in gradients across consecutive batches presents an opportunity to mitigate these inefficiencies. By leveraging the quantization of gradients and consistent distribution of zeros, the gradients can be partitioned logically to speedup communication. Splitting them into groups with and without zeros can allow different compression approaches for both. The bandwidth under-utilization due to nonuniform data size can also be solved by partitioning the gradients into variable-sized chunks, leading to more balanced compressed data sizes and reduced idle waiting time. We propose two novel strategies in Oscar, where gradient splitting is designed to improve communication and training. Oscar-SW is a novel software-based technique supporting direct AllReduce that splits gradients into probable zeros and non zeros to apply count sketch compression. Oscar-HW, a novel hardware/software codesigned gradient splitting technique is proposed with ASC (Adaptive Stepwise Coding), an encoding technique for gradient compression in distributed training. Oscar-HW dynamically splits fixed-point quantized gradients for AllReduce communications and maximizes bandwidth utilization for state-of-the-art hardware compression techniques. ASC is a variant of Adaptive Arithmetic Coding (AAC) that generates a distinct probability table for each timestep of AllReduce to adapt to its unique value ranges and avoids sending the probability table during the communication of gradients. Our experimental results show that Oscar-SW achieves$1.22 \times$speedup and 7 % better accuracy over the SOTA CountSketch algorithm. Oscar-HW achieves an average AllReduce speedup of$3.77 \times$, and an average end-to-end training speedup of$1.38 \times$. ASC achieves an average AllReduce speedup of$1.05 \times$over Atalanta and$4.66 \times$over no compression. Pranati Majhi, Sabuj Laskar, Abdullah Muzahid, Eun Jung Kim 0001 |
HPCA | 4 |
| 2025 | Enhancing Program Analysis with Deterministic Distinguishable Calling ContextabstractCalling context is crucial for improving the precision of program analyses in various use cases (clients), such as profiling, debugging, optimization, and security checking. Often the calling context is encoded using a numerical value. We have observed that many clients benefit not only from a deterministic but also globally distinguishable value across runs to simplify bookkeeping and guarantee complete uniqueness. However, existing work only guarantees determinism, not global distinguishability. Clients need to develop auxiliary helpers, which incurs considerable overhead to distinguish encoded values among all calling contexts. In this paper, we propose Deterministic Distinguishable Calling Context Encoding () that can enable both properties of calling context encoding natively. The key idea of is leveraging the static call graph and encoding each calling context as the running call path count. Thereby, a mapping is established statically and can be readily used by the clients. Our experiments with two client tools show that has a comparable overhead compared to two state-of-the-art encoding schemes, PCCE and PCC, and further avoids the expensive overheads of collision detection, up to 2.1× and 50%, for Splash-3 and SPEC CPU 2017, respectively. Sungkeun Kim, Khanh Nguyen 0001, Chia-Che Tsai, Abdullah Muzahid, Eun Jung Kim 0001 |
CC | 6 |
| 2025 | SuperMesh: Energy-Efficient Collective Communications for AcceleratorsabstractChiplet-based Deep Neural Network (DNN) accelerators are a promising approach to meet the scalability demands of modern DNN models.Such accelerators usually utilize 2D mesh topologies.However, state-of-the-art collective communication algorithms often struggle within these topologies due to limited connectivity at border nodes, leading to communication bottlenecks and performance degradation.To address this challenge, we propose two novel topologies for chiplet-based accelerators aimed at improving collective communication performance and energy efficiency by integrating additional links parallel to the existing peripheral links of mesh topologies.The first proposed topology, SuperMesh Bi adds bidirectional links parallel to all peripheral links.In contrast, the second proposed topology, SuperMesh Alter , alternately adds bidirectional links parallel to the peripheral links, offering additional paths for data traversal.Both of the topologies adhere to a core principle-augmenting the outer region of mesh topologies with extra links to retain the original structure's latency and scalability, ensuring compatibility with chiplet-based accelerator designs and maintaining energy efficiency.To fully utilize these enhanced topologies, we co-designed pipelined collective algorithms for AllReduce, ReduceScatter, and AllGather.Our proposed algorithms and topologies achieve an average AllReduce speedup of 1.18-1.33×and a 1.77-2.22×speedup in ReduceScatter and AllGather compared to conventional 2D-mesh topologies. Sabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung Kim 0001 |
MICRO | 4 |
| 2024 | Enhancing Collective Communication in MCM Accelerators for Deep Learning TrainingabstractWith the widespread adoption of Deep Learning (DL) models, the demand for DL accelerator hardware has risen. On top of that, DL models are becoming massive in size. To accommodate those models, multi-chip-module (MCM) emerges as an effective approach for implementing large-scale DL accelerators. While MCMs have shown promising results for DL inference, its potential for Deep Learning Training remains largely unexplored. Current approaches fail to fully utilize available links in a mesh interconnection network of an MCM accelerator. To address this issue, we propose two novel AllReduce algorithms for mesh-based MCM accelerators - RingBiOdd and Three Tree Overlap (TTO). RingBiOdd is a ring-based algorithm that enhances the bandwidth of AllReduce by creating two unidirectional rings using bidirectional interconnects. On the other hand, TTO is a tree-based algorithm that improves AllReduce performance by overlapping data chunks. TTO constructs three topology-aware disjoint trees and runs different steps of the AllReduce operation in parallel. We present a detailed design and implementation of the proposed approaches. Our experimental results over seven DL models indicate that RingBiOdd achieves 50% and 8% training time reduction over unidirectional Ring AllReduce and MultiTree. Furthermore, TTO demonstrates 33% and 29% training time reduction over state-ofthe-art MultiTree and Bidirectional Ring AllReduce, respectively. Sabuj Laskar, Pranati Majhi, Sungkeun Kim, Farabi Mahmud, Abdullah Muzahid, Eun Jung Kim 0001 |
HPCA | 6 |
| 2023 | Attack of the Knights: Non Uniform Cache Side Channel AttackabstractFor a distributed last-level cache (LLC) in a large multicore chip, the access time to one LLC bank can significantly differ from that to another due to the difference in physical distance. In this paper, we successfully demonstrate a new distance-based side-channel attack by timing the AES decryption operation and extracting part of an AES secret key on an Intel Knights Landing CPU. We introduce several techniques to overcome the challenges of the attack, including the use of multiple attack threads to ensure LLC hits, to detect vulnerable memory locations, and to obtain fine-grained timing of the victim operations. While operating as a covert channel, this attack can reach a bandwidth of 205 KBPS with an error rate of only 0.02%. We also observed that the side-channel attack can extract 4 bytes of an AES key with 100% accuracy with only 4000 trial rounds of encryption. Farabi Mahmud, Sungkeun Kim, Harpreet Singh Chawla, Eun Jung Kim 0001, Chia-Che Tsai, Abdullah Muzahid |
ACSAC | 4 |
| 2023 | WHISTLE: CPU Abstractions for Hardware and Software Memory Safety InvariantsabstractMemory safety invariants extracted from a program can help defend and detect against both software and hardware memory violations. For instance, by allowing only specific instructions to access certain memory locations, system can detect out-of-bound or illegal pointer dereferences that lead to correctness and security issues. In this paper, we propose CPU abstractions, called, to specify and check program invariants to provide defense mechanism against both software and hardware memory violations at runtime. ensures that the invariants must be satisfied at every memory accesses. We present a fast invariant address translation and retrieval scheme using a specialized cache. It stores and checks invariants related to global, stack and heap objects. The invariant checks can be performed synchronously or asynchronously. uses synchronous checking for high security-critical programs, while others are protected by asynchronous checking. A fast exception is proposed to alert any violations as soon as possible in order to close the gap for transient attacks. Our evaluation shows that can detect both software and hardware, spatial and temporal memory violations. incurs 53% overhead when checking synchronously, or 15% overhead when checking asynchronously. Sungkeun Kim, Farabi Mahmud, Jiayi Huang 0001, Pritam Majumder, Chia-Che Tsai, Abdullah Muzahid, Eun Jung Kim 0001 |
IEEE Trans. Computers | 7 |
| 2021 | Communication Algorithm-Architecture Co-Design for Distributed Deep LearningabstractLarge-scale distributed deep learning training has enabled developments of more complex deep neural network models to learn from larger datasets for sophisticated tasks. In particular, distributed stochastic gradient descent intensively invokes all-reduce operations for gradient update, which dominates communication time during iterative training epochs. In this work, we identify the inefficiency in widely used all-reduce algorithms, and the opportunity of algorithm-architecture co-design. We propose MultiTree all-reduce algorithm with topology and resource utilization awareness for efficient and scalable all-reduce operations, which is applicable to different interconnect topologies. Moreover, we co-design the network interface to schedule and coordinate the all-reduce messages for contention-free communications, working in synergy with the algorithm. The flow control is also simplified to exploit the bulk data transfer of big gradient exchange. We evaluate the co-design using different all-reduce data sizes for synthetic study, demonstrating its effectiveness on various interconnection network topologies, in addition to state-of-the-art deep neural networks for real workload experiments. The results show that MultiTree achieves 2.3× and 1.56× communication speedup, as well as up to 81% and 30% training time reduction compared to ring all-reduce and state-of-the-art approaches, respectively. Jiayi Huang 0001, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid, Ki Hwan Yum, Eun Jung Kim 0001 |
ISCA | 6 |
| 2021 | A Voting Approach for Adaptive Network-on-Chip Power-GatingabstractScalable Networks-on-Chip (NoCs) have become the standard interconnection mechanisms in large-scale multicore architectures. These NoCs consume a large fraction of the on-chip power budget, where the static portion is becoming dominant as technology scales down to sub-10nm node. Therefore, it is essential to reduce static power so as to achieve power- and energy-efficient computing. Power-Gating as an effective static power saving technique can be used to power off inactive routers for static power saving. However, packet deliveries in irregular power-gated networks suffer from detour or waiting time overhead to either route around or wake up power-gated routers. In this article, we proposeFly-Over (Flov), a voting approach for dynamic router power-gating in a light-weight and distributed manner, which includesFlovrouter microarchitecture, adaptive power-gating policy, and low-latency dynamic routing algorithms. We evaluateFlovusing synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. Our full-system evaluations show thatFlovreduces the power consumption of NoC by 31 and 20 percent, respectively, on average across several benchmarks, compared to the baseline and the state-of-the-art while maintaining the similar performance. Jiayi Huang 0001, Shilpa Bhosekar, Rahul Boyapati, Byul Hur, Ki Hwan Yum, Eun Jung Kim 0001 |
IEEE Trans. Computers | 7 |
| 2021 | Computing En-Route for Near-Data ProcessingabstractThe data explosion and faster data analysis demand have spawned emerging applications that operate over myriads of data and exhibit large memory footprints with low data reuse rate. Such characteristics lead to enormous data movements across the memory hierarchy and pose significant pressure on modern communication fabrics and memory subsystems. To mitigate the worsening gap between high processor computation density and deficient memory bandwidth, memory networks, and near-data processing techniques are proposed to keep improving system performance and energy efficiency. In this article, we propose Active-Routing, an in-network near-data processing architecture for data-flow execution, which enables computation en-route by exploiting patterns of aggregation over intermediate results. The proposed architecture leverages the massive memory cube- and vault-level parallelism as well as network concurrency to optimize the aggregation operations along a dynamically built Active-Routing Tree. It also introduces page granular computation offloading to amortize the offloading overhead and improve the throughput. Compared to the state-of-the-art processing-in-memory architecture, the evaluations show that the baseline Active-Routing can achieve up to 7× speedup with an average of 60 percent performance improvement, and reduce the energy-delay product by 80 percent across various benchmarks. Further optimizations with vault-level parallelism and page granular offloading can achieve an extra order of magnitude improvement. Jiayi Huang 0001, Pritam Majumder, Sungkeun Kim, Troy Fulton, Ramprakash Reddy Puli, Ki Hwan Yum, Eun Jung Kim 0001 |
IEEE Trans. Computers | 7 |
| 2021 | Remote Control: A Simple Deadlock Avoidance Scheme for Modular Systems-on-ChipabstractEver increasing performance demand and shrinking in the transistor size together result in complex and dense packing in large chips. That motivates designers to opt for many small specialized hardware modules in a chip to extract maximum performance benefits with relatively lower complexity and cost. These altogether opens up new directions for heterogeneous modular System-on-Chip (SoC) research, where a large system is built by assembling small independently designed chiplets (small chips). We focus on the communication aspect of such SoCs, especially the newly observed deadlock among chiplets. Even though deadlock is a classic problem in networks and many solutions are available, the modular SoC design demands customized solutions that preserves the design flexibility for chiplet designers. We proposeRemote Control (RC), a simple routing oblivious deadlock avoidance scheme based on selective injection-control mechanism. Along with guarantee on deadlock freedom,RCaims to provide a methodology to make each independently designed chiplet seamlessly integrate in any modular SoCs. We achieve up to 56.34% throughput and 15.49% zero load latency improvements on synthetic traffic and up to 20% speedup on real workloads taken from vast range of benchmark suites, over the state-of-the-art turn restriction based technique applied in the modular SoC domain. Pritam Majumder, Sungkeun Kim, Jiayi Huang 0001, Ki Hwan Yum, Eun Jung Kim 0001 |
IEEE Trans. Computers | 5 |
| 2019 | Active-Routing: Compute on the Way for Near-Data ProcessingabstractThe explosion of data availability and the demand for faster data analysis have led to the emergence of applications exhibiting large memory footprint and low data reuse rate. These workloads, ranging from neural networks to graph processing, expose compute kernels that operate over myriads of data. Significant data movement requirements of these kernels impose heavy stress on modern memory subsystems and communication fabrics. To mitigate the worsening gap between high CPU computation density and deficient memory bandwidth, solutions like memory networks and near-data processing designs are being architected to improve system performance substantially. In this work, we examine the idea of mapping compute kernels to the memory network so as to leverage in-network computing in data-flow style, by means of near-data processing. We propose Active-Routing, an in-network compute architecture that enables computation on the way for near-data processing by exploiting patterns of aggregation over intermediate results of arithmetic operators. The proposed architecture leverages the massive memory-level parallelism and network concurrency to optimize the aggregation operations along a dynamically built Active-Routing Tree. Our evaluations show that Active-Routing can achieve upto 7X speedup with an average of 60% performance improvement, and reduce the energy-delay product by 80% across various benchmarks compared to the state-of-the-art processing-in-memory architecture. Jiayi Huang 0001, Ramprakash Reddy Puli, Pritam Majumder, Sungkeun Kim, Rahul Boyapati, Ki Hwan Yum, Eun Jung Kim 0001 |
HPCA | 7 |
| 2019 | Dual Pattern Compression Using Data-Preprocessing for Large-Scale GPU ArchitecturesabstractGraphics Processing Units (GPUs) have been widely accepted for diverse general purpose applications due to a massive degree of parallelism. The demand for large-scale GPUs processing a large volume of data with high throughput has been rising rapidly. However, in large-scale GPUs, a bandwidth-efficient network design is challenging. Compression techniques are a practical remedy to effectively increase network bandwidth by reducing data size transferred. We propose a new simple compression mechanism, Dual Pattern Compression (DPC), that compresses only two patterns with a very low latency. The simplicity of compression/decompression is achieved through data remapping and data-type-aware data preprocessing which exploits bit-level data redundancy. The data type is detected during runtime. We demonstrate that our compression scheme effectively mitigates the network congestion in a large-scale GPU. It achieves IPC improvement by 33% on average (up to 126%) across various benchmarks with average space savings ratios of 61% in integer, 46% (up to 72%) in floating-point and 23% (up to 57%) in character type benchmarks. Kyung Hoon Kim, Priyank Devpura, Abhishek Nayyar, Andrew Doolittle, Ki Hwan Yum, Eun Jung Kim 0001 |
IPDPS | 6 |
| 2018 | Approximate Networks on ChipabstractThe trend of unsustainable power cosumption and large memory bandwidth demands in massively parallel multicore systems, with the advent of the big data era, has brought upon the onset of alternate computation paradigms utilizing heterogeneity, specialization, processor-in-memory and approximation. Approximate Computing is being touted as a viable solution for high performance computation by relaxing the accuracy constraints of applications. This trend has been accentuated by emerging data intensive applications in domains like image/video processing, machine learning and big data analytics that allow inaccurate outputs within an acceptable variance. Leveraging relaxed accuracy for high throughput in Networks-on-Chip (NoCs), which have rapidly become the accepted method for connecting a large number of on-chip components, has not yet been explored. In this talk, I will present APPROX-NoC, a hardware data approximation framework with an online data error control mechanism for high performance NoCs. APPROX-NoC facilitates approximate matching of data patterns, within a controllable value range, to compress them thereby reducing the volume of data movement across the chip. Eun Jung Kim 0001 |
NOCS | 1 |
| 2017 | Packet coalescing exploiting data redundancy in GPGPU architecturesabstractGeneral Purpose Graphics Processing Units (GPGPUs) are becoming a cost-effective hardware approach for parallel computing. Many executions on the GPGPUs place heavy stress on the memory system, creating network bottlenecks near memory controllers. We observe that data redundancy in communication traffic is common-place across a wide range of GPGPU applications. To exploit the data redundancy, we propose a packet coalescing mechanism to alleviate the network bottlenecks by directly reducing the traffic volume. The key idea is to coalesce multiple packets into one without increasing the packet size when they carry redundant cache blocks. To ensure that the coalesced packets are delivered to their respective destinations, we adopt multicast routing for the interconnection network of GPGPUs. Our coalescing approach yields 15% IPC improvement (up to 112%) in a large-scale GPGPU with 2D mesh across various GPGPU applications, by reducing average memory access time (AMAT) by 15.5% (up to 65.2%) and obtaining network bandwidth savings by 13% (up to 37%). Also, our coalescing approach achieves 7% IPC improvement in the NVIDIA Fermi architecture with the crossbar. Kyung Hoon Kim, Rahul Boyapati, Jiayi Huang 0001, Yuho Jin, Ki Hwan Yum, Eun Jung Kim 0001 |
ICS | 6 |
| 2017 | Fly-Over: A Light-Weight Distributed Power-Gating Mechanism for Energy-Efficient Networks-on-ChipabstractScalable Networks-on-Chip (NoCs) have become the de facto interconnection mechanism in large scale Chip Multiprocessors. Not only are NoCs devouring a large fraction of the on-chip power budget but static NoC power consumption is becoming the dominant component as technology scales down. Hence reducing static NoC power consumption is critical for energy-efficient computing. Previous research has proposed to power-gate routers attached to inactive cores so as to save static power, but requires centralized control and global network knowledge. In this paper, we propose Fly-Over (FLOV), a light-weight distributed mechanism for power-gating routers, which encompasses FLOV router architecture, handshake protocols, and a partition-based dynamic routing algorithm to maintain network functionalities. With simple modifications to the baseline router architecture, FLOV can facilitate FLOV links over power-gated routers. Then we present two handshake protocols for FLOV routers, restricted FLOV that can power-gate routers under restricted conditions and generalized FLOV with more power saving capability. The proposed routing algorithm provides best-effort minimal path routing without the necessity for global network information. We evaluate our schemes using synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. Our full system evaluations show that FLOV reduces the total and static energy consumption by 18% and 22% respectively, on average across several benchmarks, compared to state-of-the-art NoC power-gating mechanism while keeping the performance degradation minimal. Rahul Boyapati, Jiayi Huang 0001, Kyung Hoon Kim, Ki Hwan Yum, Eun Jung Kim 0001 |
IPDPS | 6 |
| 2017 | APPROX-NoC: A Data Approximation Framework for Network-On-Chip Architectures
Rahul Boyapati, Jiayi Huang 0001, Pritam Majumder, Ki Hwan Yum, Eun Jung Kim 0001 |
ISCA | 5 |
| 2016 | POSTER: Fly-Over: A Light-Weight Distributed Power-Gating Mechanism For Energy-Efficient Networks-on-ChipabstractReducing static NoC power consumption is becoming critical for energy-efficient computing as technology scales down since NoCs are devouring a large fraction of the on-chip power budget. We propose Fly-Over (FLOV), a light-weight distributed mechanism for power-gating routers. With simple modifications to the baseline router architecture, FLOV links are facilitated over power-gated routers. A Handshake protocol that allows seamless router power-gating in addition to a dynamic routing algorithm, that provides best-effort minimal path without the necessity for global network information, maintain normal NoC functionality. We evaluate our schemes using synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. The results show that FLOV can achieve on average 19.2% latency reduction and 15.9% total energy savings. Rahul Boyapati, Jiayi Huang 0001, Kyung Hoon Kim, Ki Hwan Yum, Eun Jung Kim 0001 |
PACT | 6 |
| 2015 | Bandwidth-efficient on-chip interconnect designs for GPGPUsabstractModern computational workloads require abundant thread level parallelism (TLP), necessitating highly-parallel, many-core accelerators such as General Purpose Graphics Processing Units (GPGPUs). GPGPUs place a heavy demand on the on-chip interconnect between the many cores and a few memory controllers (MCs). Thus, traffic is highly asymmetric, impacting on-chip resource utilization and system performance. Here, we analyze the communication demands of typical GPGPU applications, and propose efficient Network-on-Chip (NoC) designs to meet those demands. We show that the proposed schemes improve performance by up to 64.7%. Compared to the best of class prior work, our VC monopolizing and partitioning schemes improve performance by 25%. Hyunjun Jang, Jinchun Kim, Paul Gratz, Ki Hwan Yum, Eun Jung Kim 0001 |
DAC | 5 |
| 2015 | PID controlled thermal management in photonic network-on-chipabstractThe communication bandwidth and power consumption of network-on-chip (NoC) are going to meet their limits soon because of traditional metallic interconnects. Photonic NoC (PNoC) is emerging as a promising alternative to address these bottlenecks. However, PNoCs are highly susceptible to thermal fluctuations which is highly common in a manycore chip. This paper first introduces a low power, low cost mesh-based PNoC architecture and provides a quantitative analysis of it's power consumption over varying on-chip temperature. The paper then proposes a proportional-integral-derivative (PID) heater mechanism that minimizes the effect of thermal variation on PNoC's performance and power. Experimental results for a 8 × 8 PNoC shows that the proposed technique offers the maximum network-bandwidth considering the thermal effects. Compared to the recently reported results, the proposed design consumes 40% less power and has a temperature variation as low as 1 °C. Dharanidhar Dang, Rabi N. Mahapatra, Eun Jung Kim 0001 |
ICCD | 3 |
| 2013 | A Case for Handshake in Nanophotonic InterconnectsabstractNanophotonics has been proposed to design low latency and high bandwidth NOC for future Chip Multi-Processors (CMPs). Recent nanophotonic NOC designs adopt the token-based arbitration coupled with credit-based flow control, which leads to low bandwidth utilization. In this work, we propose two handshake schemes for nanophotonic interconnects in CMPs, Global Handshake (GHS) and Distributed Handshake (DHS), which get rid of the traditional credit-based flow control, reduce the average token waiting time, and finally improve the network throughput. Furthermore, we enhance the basic handshake schemes with setaside buffer and circulation techniques to overcome the Head-Of-Line (HOL) blocking. Our evaluation shows that the proposed handshake schemes improve network throughput by up to 62% under synthetic workloads. With the extracted trace traffic from real applications, the handshake schemes can reduce the communication delay by up to 59%. The basic handshake schemes add only 0.4% hardware overhead for optical components and negligible power consumption. In addition, the performance of the handshake schemes is independent of on-chip buffer space, which makes them feasible in a large scale nanophotonic interconnect design. Lei Wang 0041, Jagadish Jayabalan, Minseon Ahn, Haiyin Gu, Ki Hwan Yum, Eun Jung Kim 0001 |
IPDPS | 6 |
| 2012 | APCR: an adaptive physical channel regulator for on-chip interconnectsabstractChip Multi-Processor (CMP) architectures have become mainstream for designing processors. With a large number of cores, Network-On-Chip (NOC) provides a scalable communication method for CMP architectures, where wires become abundant resources available inside the chip. NOC must be carefully designed to meet constraints of power and area, and provide ultra low latencies. In this paper, we propose an Adaptive Physical Channel Regulator (APCR) for NOC routers to exploit huge wiring resources. The flit size in an APCR router is less than the physical channel width (phit size) to provide finer granularity flow control. An APCR router allows flits from different packets or flows to share the same physical channel in a single cycle. The three regulation schemes (Monopolizing, Fair-sharing and Channel-stealing) intelligently allocate the output channel resources considering not only the availability of physical channels but the occupancy of input buffers. In an APCR router, each Virtual Channel can forward a dynamic number of flits every cycle depending on the run-time network status. Our simulation results using a detailed cycle-accurate simulator show that an APCR router improves the network throughput by over 100% in synthetic workloads, compared with a traditional design with the same buffer size. An APCR router can outperform a traditional router even if the buffer size is halved. Lei Wang 0041, Poornachandran Kumar, Ki Hwan Yum, Eun Jung Kim 0001 |
PACT | 4 |
| 2012 | Efficient Data Packet Compression for Cache Coherent Multiprocessor SystemsabstractMultiprocessor systems have been popular for their high performance not only for server markets but also for computing environments for general users. With the increased software complexity, networking overheads in multiprocessor systems are becoming one of the most influential factors in overall system performance. In this paper, we attempt to reduce communication overheads through a data packet compression technique integrating a cache coherence protocol. Here we propose Variable Size Compression (VSC) scheme that compresses or completely eliminates data packets while harmonizing with existing cache coherence protocols. Simulation results show approximately 23% of improvement on average in terms of overall system performance when compared with the most recent compression scheme. VSC also improves performance by 20% on average in terms of cache miss latency. Baik Song An, Ki Hwan Yum, Eun Jung Kim 0001 |
DCC | 4 |
| 2012 | A Hybrid Buffer Design with STT-MRAM for On-Chip InterconnectsabstractAs the chip multiprocessor (CMP) design moves toward many-core architectures, communication delay in Network-on-Chip (NoC) has been a major bottleneck in CMP systems. Using high-density memories in input buffers helps to reduce the bottleneck through increasing throughput. Spin-Torque Transfer Magnetic RAM (STT-MRAM) can be a suitable solution due to its nature of high density and near-zero leakage power. But its long latency and high power consumption in write operations still need to be addressed. We explore the design issues in using STT-MRAM for NoC input buffers. Motivated by short intra-router latency, we use the previously proposed write latency reduction technique sacrificing retention time. Then we propose a hybrid design of input buffers using both SRAM and STT-MRAM to hide the long write latency efficiently. Considering that simple data migration in the hybrid buffer consumes more dynamic power compared to SRAM, we provide a lazy migration scheme that reduces the dynamic power consumption of the hybrid buffer. Simulation results show that the proposed scheme enhances the throughput by 21% on average. Hyunjun Jang, Baik Song An, Nikhil Kulkarni, Ki Hwan Yum, Eun Jung Kim 0001 |
NOCS | 5 |
| 2012 | Communication-Aware Globally-Coordinated On-Chip NetworksabstractWith continued Moore's law scaling, multicore-based architectures are becoming the de facto design paradigm for achieving low-cost and performance/power-efficient processing systems through effective exploitation of available parallelism in software and hardware. A crucial subsystem within multicores is the on-chip interconnection network that orchestrates high-bandwidth, low-latency, and low-power communication of data. Much previous work has focused on improving the design of on-chip networks but without more fully taking into consideration the on-chip communication behavior of application workloads that can be exploited by the network design. A significant portion of this paper analyzes and models on-chip network traffic characteristics of representative application workloads. Leveraged by this, the notion of globally coordinated on-chip networks is proposed in which application communication behavior-captured by traffic profiling-is utilized in the design and configuration of on-chip networks so as to support prevailing traffic flows well, in a globally coordinated manner. This is applied to the design of a hybrid network consisting of a mesh augmented with configurable multidrop (bus-like) spanning channels that serve as express paths for traffic flows benefiting from them, according to the characterized traffic profile. Evaluations reveal that network latency and energy consumption for a 64-core system running OpenMP benchmarks can be improved on average by 15 and 27 percent, respectively, with globally coordinated on-chip networks. Yuho Jin, Eun Jung Kim 0001, Timothy M. Pinkston |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Scalable and Efficient Bounds Checking for Large-Scale CMP EnvironmentsabstractWe attempt to provide an architectural support for fast and efficient bounds checking for multithread work-loads in chip-multiprocessor (CMP) environments. Bounds information sharing and smart tagging help to perform bounds checking more effectively utilizing the characteristics of a pointer. Also, the BCache architecture allows fast access to the bounds information. Simulation results show that the proposed scheme increasesμPC of memory operations by 29% on average compared to the previous hardware scheme. Baik Song An, Ki Hwan Yum, Eun Jung Kim 0001 |
PACT | 3 |
| 2011 | Fast Secure Communications in Shared Memory Multiprocessor SystemsabstractProtection and security are becoming essential requirements in commercial servers. To provide secure memory and cache-to-cache communications, we presented Interconnect-Independent Security Enhanced Shared Memory Multiprocessor System (I^{2}SEMS), mainly focusing on how to manage a global counter to encrypt, decrypt, and authenticate data messages with little performance overhead. However, I^{2}SEMS was vulnerable to replay attacks on data messages and integrity attacks on control and counter messages. This paper proposes three authentication schemes to remove those security vulnerabilities. First, we prevent replay attacks on data messages by inserting Request Counter (RC) into request messages. Second, we also use RC to detect integrity attacks on control messages. Third, we propose a new counter, referred to as GCC Counter (GC), to protect the global counter messages. We simulated our design with SPLASH-2 benchmarks on up to 16-processor shared memory multiprocessor systems by using Simics with Wisconsin multifacet General Execution-driven Multiprocessor Simulator (GEMS). Simulation results show that the overall performance slowdown is 4 percent on average with the highest keystream hit rate of 78 percent. Minseon Ahn, Eun Jung Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2010 | Efficient lookahead routing and header compression for multicasting in networks-on-chipabstractAs technology advanced, Chip Multi-processor (CMP) architectures have emerged as a viable solution for designing processors. Networks-on-Chip (NOCs) provide a scalable communication method for CMP architectures as the number of cores is increasing. Although there has been significant research on NOC designs for unicast traffic, the research on the multicast router design is still in infancy stage. Considering that one-to-many (multicast) and one-to-all (broadcast) traffic are more common in CMP applications, it is important to design a router providing efficient multicasting. In this paper, we propose an efficient lookahead routing with limited area overhead for a recently proposed multicast routing algorithm, Recursive Partitioning Multicast (RPM) [17]. Also, we present a novel compression scheme for a multicast packet header that becomes a big overhead in large networks. Comprehensive simulation results show that with our route computation logic design, providing lookahead routing in the multicast router only costs less than 20% area overhead and this percentage keeps decreasing with larger network sizes. Compared with the basic lookahead routing design, our design can save area by over 50%. With header compression and lookahead multicast routing, the network performance is improved by 22% in a (16 x 16) network on average. Lei Wang 0041, Poornachandran Kumar, Rahul Boyapati, Ki Hwan Yum, Eun Jung Kim 0001 |
ANCS | 5 |
| 2010 | Pseudo-Circuit: Accelerating Communication for On-Chip Interconnection NetworksabstractAs the number of cores on a single chip increases with more recent technologies, a packet-switched on-chip interconnection network has become a de facto communication paradigm for chip multiprocessors (CMPs). However, it is inevitable to suffer from high communication latency due to the increasing number of hops. In this paper, we attempt to accelerate network communication by exploiting communication temporal locality with minimal additional hardware cost in the existing state-of-the-art router architecture. We observe that packets frequently traverse through the same path chosen by previous packets due to repeated communication patterns, such as frequent pair-wise communication. Motivated by our observation, we propose a pseudo-circuit scheme. With previous communication patterns, the scheme reserves crossbar connections creating pseudo-circuits, sharable partial circuits within a single router. It reuses the previous arbitration information to bypass switch arbitration if the next flit traverses through the same pseudo-circuit. To accelerate communication performance further, we also propose two aggressive schemes, pseudo-circuit speculation and buffer bypassing. Pseudo-circuit speculation creates more pseudo-circuits using unallocated crossbar connections while buffer bypassing skips buffer writes to eliminate one pipeline stage. Minseon Ahn, Eun Jung Kim 0001 |
MICRO | 2 |
| 2010 | A session key caching and prefetching scheme for secure communication in cluster systems
Baik Song An, Eun Jung Kim 0001 |
J. Parallel Distributed Comput. | 3 |
| 2010 | Integration of admission, congestion, and peak power control in QoS-aware clusters
Ki Hwan Yum, Yuho Jin, Eun Jung Kim 0001, Chita R. Das |
J. Parallel Distributed Comput. | 3 |
| 2010 | Design and Analysis of On-Chip Networks for Large-Scale Cache SystemsabstractSwitched networks have been adopted in on-chip communication for their scalability and efficient resource sharing. However, using a general network for a specific domain may result in unnecessary high cost and low performance when the interconnects are not optimized for the domain. Designing an optimal network for the specific domain is challenging because in-depth knowledge of interconnects and the application domain is required. Recently proposed Nonuniform Cache Architectures (NUCAs) use wormhole-routed 2D mesh networks in L2 caches. We observe that in NUCAs, network resources are underutilized with the considerable area cost (41 percent of cache) and the network delay is significantly large (63 percent of cache access time). Motivated by our observations, we investigate both router architecture and network topology for communication behaviors in large-scale cache systems. We present Fast-LRU replacement, where cache replacement overlaps with data request delivery. Next, we propose a deadlock-free XYX routing algorithm in a mesh network and present a new halo network topology to reduce the required links. Finally, we introduce a single-cycle multicast router that needs small modification of the unicast router design. Simulation results show that our design improves the average IPC by 38 percent over the mesh design with Multicast Promotion replacement and uses 12 percent of the interconnection area of the mesh network. Yuho Jin, Eun Jung Kim 0001, Ki Hwan Yum |
IEEE Trans. Computers | 2 |
| 2009 | Temperature-aware scheduler based on thermal behavior grouping in multicore systemsabstractDynamic Thermal Management techniques have been widely accepted as a thermal solution for their low cost and simplicity. The techniques have been used to manage the heat dissipation and operating temperature to avoid thermal emergencies, but are not aware of application behavior in Chip Multiprocessors (CMPs). In this paper, we propose a temperature-aware scheduler based on applications' thermal behavior groups classified by a K-means clustering method in multicore systems. The application's thermal behavior group has similar thermal pattern as well as thermal parameters.With these thermal behavior groups, we provide thermal balances among cores with negligible performance overhead. We implement and evaluate our schemes in the 4-core (Intel Quad Core Q6600) and 8-core (two Quad Core Intel XEON E5310 processors) systems running several benchmarks. The experimental results show that the temperature-aware scheduler based on thermal behavior grouping reduces the peak temperature by up to 8degC and 5degC in our 4-core system and 8-core system with only 12% and 7.52% performance overhead, respectively, compared to L.inux standard scheduler. Inchoon Yeo, Eun Jung Kim 0001 |
DATE | 2 |
| 2009 | Adaptive Prefetching Scheme Using Web Log Mining in Cluster-Based Web SystemsabstractThe main memory management has been a critical issue to provide high performance in web cluster systems.To overcome the speed gap between processors and disks,many prefetch schemes have been proposed as memory management in web cluster systems. However, inefficient prefetch schemes can degrade the performance of the web cluster system. Dynamic access patterns due to the web cache mechanism in proxy servers increase mispredictions to waste the I/O bandwidth and available memory. Too aggressive prefetch schemes incur the shortage of available memory and performance degradation. Furthermore, modern web frameworks including persistent HTTP make the problem more challenging by reducing the available memory space with multiple connections from a client and web processes management in a prefork mode. Therefore, we attempt to design an adaptive web prefetch scheme by predicting memory status more accurately and dynamically. First, we design double prediction-by-partial-match scheme (DPS) that can be adapted to the modern web framework. Second, we propose adaptive rate controller(ARC) to determine the prefetch rate depending on the memory status dynamically. Finally, we suggest memory aware request distribution (MARD) that distributes requests based on the available web processes and memory.For evaluating the prefetch gain in a server node, we implement an Apache module in Linux. In addition, we build a simulator for verifying our scheme with cluster environments. Simulation results show 10% performance improvement on average in various workloads. Heung Ki Lee, Baik Song An, Eun Jung Kim 0001 |
ICWS | 3 |
| 2009 | Recursive partitioning multicast: A bandwidth-efficient routing for Networks-on-ChipabstractChip Multi-processor (CMP) architectures have become mainstream for designing processors. With a large number of cores, Networks-on-Chip (NOCs) provide a scalable communication method for CMP architectures. NOCs must be carefully designed to meet constraints of power consumption and area, and provide ultra low latencies. Existing NOCs mostly use Dimension Order Routing (DOR) to determine the route taken by a packet in unicast traffic. However, with the development of diverse applications in CMPs, one-to-many (multicast) and one-to-all (broadcast) traffic are becoming more common. Current unicast routing cannot support multicast and broadcast traffic efficiently. In this paper, we propose Recursive Partitioning Multicast (RPM) routing and a detailed multicast wormhole router design for NOCs. RPM allows routers to select intermediate replication nodes based on the global distribution of destination nodes. This provides more path diversities, thus achieves more bandwidth-efficiency and finally improves the performance of the whole network. Our simulation results using a detailed cycle-accurate simulator show that compared with the most recent multicast scheme, RPM saves 25% of crossbar and link power, and 33% of link utilization with 50% network performance improvement. Also RPM is more scalable to large networks than the recently proposed VCTM. Lei Wang 0041, Yuho Jin, Eun Jung Kim 0001 |
NOCS | 4 |
| 2008 | Predictive dynamic thermal management for multicore systemsabstractRecently, processor power density has been increasing at an alarming rate resulting in high on-chip temperature. Higher temperature increases current leakage and causes poor reliability. In this paper, we propose a Predictive Dynamic Thermal Management (PDTM) based on Application-based Thermal Model (ABTM) and Core-based Thermal Model (CBTM) in the multicore systems. ABTM predicts future temperature based on the application specific thermal behavior, while CBTM estimates core temperature pattern by steady state temperature and workload. The accuracy of our prediction model is 1.6% error in average compared to the model in HybDTM [8], which has at most 5% error. Based on predicted temperature from ABTM and CBTM, the proposed PDTM can maintain the system temperature below a desired level by moving the running application from the possible overheated core to the future coolest core (migration) and reducing the processor resources (priority scheduling) within multicore systems. PDTM enables the exploration of the tradeoff between throughput and fairness in temperature-constrained multicore systems. We implement PDTM on Intel's Quad-Core system with a specific device driver to access Digital Thermal Sensor (DTS). Compared against Linux standard scheduler, PDTM can decrease average temperature about 10%, and peak temperature by 5°C with negligible impact of performance under 1%, while running single SPEC2006 benchmark. Moreover, our PDTM outperforms HRTM [10] in reducing average temperature by about 7% and peak temperature by about 3°C with performance overhead by 0.15% when running single benchmark. Inchoon Yeo, Chih Chun Liu, Eun Jung Kim 0001 |
DAC | 3 |
| 2008 | Hybrid dynamic thermal management based on statistical characteristics of multimedia applicationsabstractRecently multimedia applications become one of the most popular applications in mobile devices such as wireless phones, PDAs, and laptops. However, typical mobile systems are not equipped with cooling components, which eventually causes critical thermal deficiencies. Although many low-power and low-temperature multimedia playback techniques have been proposed, they failed to provide QoS (Quality of Service) while controlling temperature due to the lack of proper understanding of multimedia applications. We propose Hybrid Dynamic Thermal Management (HDTM) which exploits thermal characteristics of both multimedia applications and systems. Specifically, we model application characteristics as the probability distribution of the number of cycles required to decode a frame. We also improve existing system thermal models by considering the effect of workload. This scheme finds an optimal clock frequency in order to prevent overheating with minimal performance degradation at runtime. The proposed scheme is implemented on Linux in a Pentium-M processor which provides variable clock frequencies. In order to evaluate the performance of the proposed scheme, we exploit three major codecs, namely MPEG-4, H.264/AVC and H.264/AVC streaming. Our results show that HDTM lowers the overall temperature by 15 ◦ C and the peak temperature by 20 ◦ C, while maintaining frame drop ratio under 0.2 % compared to previous thermal management schemes such as feedback control DTM [8], Frame-based DTM [5] and GOP-based DTM [15]. Inchoon Yeo, Eun Jung Kim 0001 |
ISLPED | 2 |
| 2008 | Adaptive data compression for high-performance low-power on-chip networksabstractWith the recent design shift towards increasing the number of processing elements in a chip, high-bandwidth support in on-chip interconnect is essential for low-latency communication. Much of the previous work has focused on router architectures and network topologies using wide/long channels. However, such solutions may result in a complicated router design and a high interconnect cost. In this paper, we exploit a table-based data compression technique, relying on value patterns in cache traffic. Compressing a large packet into a small one can increase the effective bandwidth of routers and links, while saving power due to reduced operations. The main challenges are providing a scalable implementation of tables and minimizing overhead of the compression latency. First, we propose a shared table scheme that needs one encoding and one decoding tables for each processing element, and a management protocol that does not require in-order delivery. Next, we present streamlined encoding that combines flit injection and encoding in a pipeline. Furthermore, data compression can be selectively applied to communication on congested paths only if compression improves performance. Simulation results in a 16-core CMP show that our compression method improves the packet latency by up to 44% with an average of 36% and reduces the network power consumption by 36% on average. Yuho Jin, Ki Hwan Yum, Eun Jung Kim 0001 |
MICRO | 3 |
| 2007 | I2SEMS: Interconnects-Independent Security Enhanced Shared Memory Multiprocessor Systems
Minseon Ahn, Eun Jung Kim 0001 |
PACT | 3 |
| 2007 | A Domain-Specific On-Chip Network Design for Large Scale Cache SystemsabstractAs circuit integration technology advances, the design of efficient interconnects has become critical. On-chip networks have been adopted to overcome scalability and the poor resource sharing problems of shared buses or dedicated wires. However, using a general on-chip network for a specific domain may cause underutilization of the network resources and huge network delays because the interconnects are not optimized for the domain. Addressing these two issues is challenging because in-depth knowledges of interconnects and the specific domain are required. Non-uniform cache architectures (NUCAs) use wormhole-routed 2D mesh networks to improve the performance of on-chip L2 caches. We observe that network resources in NUCAs are underutilized and occupy considerable chip area (52% of cache area). Also the network delay is significantly large (63% of cache access time). Motivated by our observations, we investigate how to optimize cache operations and and design the network in large scale cache systems. We propose a single-cycle router architecture that can efficiently support multicasting in on-chip caches. Next, we present fast-LRU replacement, where cache replacement overlaps with data request delivery. Finally we propose a deadlock-free XYX routing algorithm and a new halo network topology to minimize the number of links in the network. Simulation results show that our networked cache system improves the average IPC by 38% over the mesh network design with multicast promotion replacement while using only 23% of the interconnection area. Specifically, multicast fast-LRU replacement improves the average IPC by 20% compared with multicast promotion replacement. A halo topology design additionally improves the average IPC by 18% over a mesh topology Yuho Jin, Eun Jung Kim 0001, Ki Hwan Yum |
HPCA | 2 |
| 2007 | Effective Dynamic Thermal Management for MPEG-4 decodingabstractThis paper proposes dynamic thermal management (DTM) based on a dynamic voltage and frequency scaling (DVFS) technique for MPEG-4 decoding to guarantee thermal safety while maintaining a quality of service (QoS) constraint. Although many low-power and low-temperature multimedia playback techniques have been proposed, most of them are impractical in real-time and have several restricting assumptions. Multimedia data consists of several frames requiring different decoding efforts. Since both temperature and performance of a multimedia system are affected by the complexity of scenes, our main idea is to use the information on scene complexity to find an appropriate frequency. In order to predict the complexity of the current scene, we extract information from the previous group of pictures (GOP) using feedback control with a display buffer. Experimental results with twelve movies show that our DTM scheme guarantees the threshold of temperature (70degC) while maintaining 0% frame miss ratio. Also, our DTM scheme decreases the average temperature by up to 13% without any additional hardware and playback latency. Inchoon Yeo, Heung Ki Lee, Eun Jung Kim 0001, Ki Hwan Yum |
ICCD | 3 |
| 2007 | Design of an Active Set Top Box in a Wireless Network for Scalable Streaming ServicesabstractThe popularity of multimedia streaming services via wireless home networks has confronted major challenges in quality improvement for services through a set top box (STB). Even though scalable methods have been suggested to enhance the quality of multimedia streaming services, it is still challenging how to provide scalable streaming services in wireless home networks. Previous studies on scalable streaming services eliminate the corrupted stream at the multimedia client. In this paper, we propose a new method, ActiveSTB, which removes the distorted or unsuitable multimedia data early to save frugal resources. We use a network simulation tool, NS-2, to evaluate our method with various range of cross traffic and error rates. The simulation results show that ActiveSTB support 12.5% to 40% more packets than the original STB for all ranges of cross traffic and error rates. Heung Ki Lee, Varrian Hall, Ki Hwan Yum, Kyoung Ill Kim, Eun Jung Kim 0001 |
ICIP (6) | 5 |
| 2007 | Exploring IBA Design Space for Improved PerformanceabstractInfiniBand architecture (IBA) is envisioned to be the default communication fabric for future system area networks (SANs) or clusters. However, IBA design is currently in its infancy since the released specification outlines only higher level functionalities, leaving it open for exploring various design alternatives. In this paper, we investigate four corelated techniques for providing high and predictable performance in IBA. These are: 1) using the shortest path first (SPF) algorithm for deterministic packet routing, 2) developing a multipath routing mechanism for minimizing congestion, 3) developing a selective packet dropping scheme to handle deadlock and congestion, and 4) providing multicasting support for customized applications. These designs are implemented in a pipelined, IBA-style switch architecture, and are evaluated using an integrated workload consisting of MPEG-2 video streams, best- effort traffic, and control traffic on a versatile IBA simulation testbed. Simulation results with 15-node and 30-node irregular networks indicate that the SPF routing, multipath routing, packet dropping, and multicasting schemes are quite effective in delivering high and assured performance in clusters Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das, Mazin S. Yousif, José Duato |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2007 | A Comprehensive Framework for Enhancing Security in InfiniBand ArchitectureabstractThe InfiniBand™ Architecture (IBA) is a promising communication standard for building clusters and system area networks. However, the IBA specification has left out security aspects, resulting in potential security vulnerabilities which could be exploited with moderate effort. In this paper, we view these vulnerabilities from three classical security aspects: confidentiality, authentication, and availability and investigate the following security issues. First, as groundwork for secure services in IBA, we present partition-level and queue pair-level key management schemes, both of which can be easily integrated into IBA. Second, for confidentiality and authentication, we present a method to incorporate a scalable encryption and authentication algorithm into IBA with little performance overhead. Third, for better availability, we propose a stateful ingress filtering mechanism to block denial of service (DoS) attacks. Finally, to further improve the availability, we provide a scalable packet marking method tracing back DoS attacks. Simulation results of an IBA network show that the security performance overhead due to encryption/authentication on network latency ranges from 0.7% to 12.4%. Since the stateful ingress filtering is enabled only when a DoS attack is active, there is no performance overhead in a normal situation. Eun Jung Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | An Overview of Security Issues in Cluster Interconnects
Eun Jung Kim 0001, Ki Hwan Yum, Mazin S. Yousif |
CCGRID | 2 |
| 2006 | Assuring K-Coverage in the Presence of Mobility in Wireless Sensor NetworksabstractAlong with energy conservation, it has been a critical issue to maintain a desired degree of coverage in wireless sensor networks (WSNs), especially in a mobile environment. By enhancing a variant of random waypoint (RWP) model, we propose mobility resilient coverage control (MRCC) to assure if-coverage in the presence of mobility. Our basic goals are (1) to elaborate the probability of breaking if-coverage with moving-in and moving-out probabilities, and (2) to issue wake-up calls to sleeping sensors to meet user requirement of if-coverage even in the presence of mobility. Furthermore, by separating the mobility behavior into average and individual, the probability of breaking if-coverage can be precisely calculated, hence reducing the number of sensors to be awakened. Our experiments with NS2 show that MRCC with the individual probability achieves better coverage by 1.4% with 22% fewer numbers of active sensors than that of existing coverage configuration protocol (CCP). Heeyeol Yu, Jayakrishnan V. Iyer, Hogil Kim, Eun Jung Kim 0001, Ki Hwan Yum, Pyeong Soo Mah |
GLOBECOM | 4 |
| 2006 | Bandwidth Estimation in Wireless Lans for Multimedia Streaming ServicesabstractThe popularity of multimedia streaming services via wireless networks presents major challenges in the management of network bandwidth. One challenge is to quickly and precisely estimate the available bandwidth for the decision of streaming rates of layered and scalable multimedia services. Previous works based on wired networks are too burdensome to be applied to multimedia applications in wireless networks. In this paper, a new method, IdleGap, is suggested to estimate the available bandwidth of a wireless LAN based on the information from a low layer in the protocol stack. We use a network simulation tool, NS-2, to evaluate our new method with various range of cross traffic and observation times. Our simulation results show that IdleGap accurately estimates the available bandwidth for all ranges of cross traffic (100 Kbps~1 Mbps) with a very short observation time of 10 seconds Heung Ki Lee, Varrian Hall, Ki Hwan Yum, Kyoung Ill Kim, Eun Jung Kim 0001 |
ICME | 5 |
| 2006 | A PROactive Request Distribution (PRORD) Using Web Log Mining in a Cluster-Based Web ServerabstractWidely adopted distributor-based systems forward user requests to a balanced set of waiting servers in complete transparency to the users. The policy employed in forwarding requests from the front-end distributor to the backend servers dominates the overall system performance. The locality-aware request distribution (LARD) scheme improves the system response time by having the requests serviced by the Web servers that contain the data in their caches. In this paper, we propose a proactive request distribution (PRORD) that applies an intelligent proactive-distribution at the front-end and complementary pre-fetching at the back-end server nodes to obtain the data of high relation to the previous requests in their caches. The prefetching scheme fetches the Web pages in advance into the memory based on a confidence value of the Web page, which is predicted by the proactive distribution scheme. Designed to work with the prevailing Web technologies, such as HTTP 1.1, our scheme aims to provide reduced response time to the users. Simulations carried out with traces derived from the log files of real Web servers witness performance boost of 15-45 % compared to the existing distribution policies Heung Ki Lee, Gopinath Vageesan, Ki Hwan Yum, Eun Jung Kim 0001 |
ICPP | 4 |
| 2005 | Peak Power Control for a QoS Capable On-Chip NetworkabstractIn recent years integrating multiprocessors in a single chip is emerging for supporting various scientific and commercial applications, with diverse demands to the underlying on-chip networks. Communication traffic of these applications makes routers greedy to acquire more power such that the total consumed power of the network may exceed the supplied power and cause reliability problems. To ensure high performance and power constraint satisfaction, the on-chip network must have a peak power control mechanism. In this paper, we propose a credit-based peak power control scheme to assure power consumption to be under the given peak power constraint, without performance degradation. The peak power control scheme efficiently regulates each flow's injection rate at the sender to minimize performance penalty. We have two different throttling schemes for real-time traffic and best-effort traffic; a rate-based throttling and an energy-budget based throttling, respectively. The simulation results on mesh networks show that the credit-based peak power control effectively prevents performance degradation and meets the peak power constraint. Yuho Jin, Eun Jung Kim 0001, Ki Hwan Yum |
ICPP | 2 |
| 2005 | On Improving Performance and Conserving Power in Cluster-Based Web ServersabstractWith the growing use of cluster systems in Web servers, file distribution and database transactions, power conservation and efficiency have been identified as critical issues in the design of cluster systems. Widely adopted, distributor-based systems forward client requests to a balanced set of backend servers in complete transparency to the clients. In this paper, we use power and locality-based request distribution at the distributor to provide optimum power conservation, while maintaining the required QoS of the system. The distribution scheme uses a simple memory management technique using pinned memory on the backend servers and proactive distribution, with the aid of data organization of the Web site, to improve the locality of the files. A simple on-off based power management scheme is applied to conserve power. Our scheme provides reduced response time to the clients and improved power conservation at the backend server cluster without compromising performance. Simulations involving real-time Web traces and latest Web technologies witness performance boost of 15-23% and power conservation of 15-48% over the existing policies. Heung Ki Lee, Gopinath Vageesan, Eun Jung Kim 0001 |
ICWS | 3 |
| 2005 | Performance analysis of a QoS capable cluster interconnect
Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das |
Perform. Evaluation | 1 |
| 2005 | A Holistic Approach to Designing Energy-Efficient Cluster InterconnectsabstractDesigning energy-efficient clusters has recently become an important concern to make these systems economically attractive for many applications. Since the cluster interconnect is a major part of the system, the focus of this paper is to characterize and optimize the energy consumption in the entire interconnect. Using a cycle-accurate simulator of an InfiniBand Architecture (IBA) compliant interconnect fabric and actual designs of its components, we investigate the energy behavior on regular and irregular interconnects. The energy profile of the three major components (switches, network interface cards (NICs), and links) reveals that the links and switch buffers consume the major portion of the power budget. Hence, we focus on energy optimization of these two components. To minimize power in the links, first we investigate the dynamic voltage scaling (DVS) algorithm and then propose a novel dynamic link shutdown (DLS) technique. The DLS technique makes use of an appropriate adaptive routing algorithm to shut down the links intelligently. We also present an optimized buffer design for reducing leakage energy in 70nm technology. Our analysis on different networks reveals that, while DVS is an effective energy conservation technique, it incurs significant performance penalty at low to medium workload. Moreover, energy saving with DVS reduces as the buffer leakage current becomes significant with 70nm design. On the other hand, the proposed DLS technique can provide optimized performance-energy behavior (up to 40 percent energy savings with less than 5 percent performance degradation in the best case) for the cluster interconnects. Eun Jung Kim 0001, Greg M. Link, Ki Hwan Yum, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Chita R. Das |
IEEE Trans. Computers | 1 |
| 2003 | Performance Enhancement Techniques for InfiniBand? ArchitectureabstractThe InfiniBand/sup TM/ Architecture (IBA) is envisioned to be the default communication fabric for future system area networks (SAN). However, the released IBA specification outlines only higher level functionalities, leaving it open for exploring various design alternatives. In this paper we investigate four co-related techniques to provide high and predictable performance in IBA. These are: (i) using the shortest path first (SPF) algorithm for deterministic packet routing; (ii) developing a multipath routing mechanism for minimizing congestion; (iii) developing a selective packet dropping scheme to handle deadlock and congestion; and (iv) providing multicasting support for customized applications. These designs are evaluated using an integrated workload on a versatile IBA simulation testbed. Simulation results indicate that the SPF routing, multipath routing, packet dropping, and multicasting schemes are quite effective in delivering high and assured performance in clusters. One of the major contributions of this research is the IBA simulation testbed, which is an essential tool to evaluate various design tradeoffs. Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das, Mazin S. Yousif, José Duato |
HPCA | 1 |
| 2003 | Energy optimization techniques in cluster interconnectsabstractDesigning energy-efficient clusters has recently become an important concern to make these systems economically attractive for many applications. Since the links and switch buffers consume the major portion of the power budget of the cluster, the focus of this paper is to optimize the energy consumption in these two components. To minimize power in the links, we propose a novel dynamic link shutdown (DLS) technique. The DLS technique makes use of an appropriate adaptive routing algorithm to shutdown the links intelligently. We also present an optimized buffer design for reducing leakage energy. Our analysis on different networks using a complete system simulator reveals that the proposed DLS technique can provide optimized performance-energy behavior (up to 40% energy savings with less than 5% performance degradation in the best case) for the cluster interconnects. Eun Jung Kim 0001, Ki Hwan Yum, Greg M. Link, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Mazin S. Yousif, Chita R. Das |
ISLPED | 1 |
| 2002 | Integrated Admission and Congestion Control for QoS Support in ClustersabstractAdmission and congestion control mechanisms are integral parts of any Quality of Service (QoS) design for networks that support integrated traffic. In this paper we propose an-admission control algorithm and a congestion control algorithm for clusters, which are increasingly being used in a diverse set of applications that require QoS guarantees. The uniqueness of our approach is that we develop these algorithms for wormhole-switched networks. We use QoS-capable wormhole routers and QoS-capable network interface cards (NICs), referred to as Host Channel Adapters (HCAs) in InfiniBand/spl trade/ Architecture (IBA), to evaluate the effectiveness of these algorithms. The admission control is applied at the HCAs and the routers, while the congestion control is deployed only at the HCAs. Simulation results indicate that the admission and congestion control algorithms are quite effective in delivering the assured performance. The proposed credit-based congestion control algorithm is simple and practical in that it relies on hardware already available in the HCA to regulate traffic injection. Ki Hwan Yum, Eun Jung Kim 0001, Chita R. Das, Mazin S. Yousif, José Duato |
CLUSTER | 2 |
| 2002 | MediaWorm: A QoS Capable Router Architecture for ClustersabstractWith the increasing use of clusters in real-time applications, it has become essential to design high-performance networks with quality-of-service (QoS) guarantees. We explore the feasibility of providing QoS in wormhole switched routers, which are widely used in designing scalable, high-performance cluster interconnects. In particular, we are interested in supporting multimedia video streams with CBR and VBR traffic, in addition to the conventional best-effort traffic. The proposed MediaWorm router uses a rate-based bandwidth allocation mechanism, called Fine-Grained VirtualClock (FGVC), to schedule network resources for different traffic classes. Our simulation results on an 8-port router indicate that it is possible to provide jitter-free delivery to VBR/CBR traffic up to an input load of 70-80 percent of link bandwidth and the presence of best-effort traffic has no adverse effect on real-time traffic. Although the MediaWorm router shows a slightly lower performance than a pipelined circuit switched (PCS) router, commercial success of wormhole switching, coupled with simpler and cheaper design, makes it an attractive alternative. Simulation of a (2/spl times/2) fat-mesh using this router shows performance comparable to that of a single switch and suggests that clusters designed with appropriate bandwidth balance between links can provide required performance for different types of traffic. Ki Hwan Yum, Eun Jung Kim 0001, Chita R. Das, Aniruddha S. Vaidya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2001 | QoS provisioning in clusters: an investigation of Router and NIC designabstractDesign of high performance cluster networks (routers) with Quality-of-Service (QoS) guarantees is becoming increasingly important to support a variety of multimedia applications, many of which have real-time constraints. Most commercial routers, which are based on the wormhole-switching paradigm, can deliver high performance, but lack QoS provisioning. In this paper, we present a pipelined wormhole router architecture that can provide high and predictable performance for integrated traffic in clusters. We consider two different implementations—a non-preemptive model and a more aggressive preemptive model. We also present the design of a network interface card (NIC) based on the Virtual Interface Architecture (VIA) design paradigm to support QoS in the NIC. The QoS capable router and NIC designs are evaluated with a mixed workload consisting of best-effort traffic, multimedia streams, and control traffic. Ki Hwan Yum, Eun Jung Kim 0001, Chita R. Das |
ISCA | 2 |
| 2001 | Calculation of Deadline Missing Probability in a QoS Capable Cluster InterconnectabstractThe growing use of clusters in diverse applications, many of which have real-time constraints, requires Quality-of-Service (QoS) support from the underlying cluster interconnect. In this paper we propose an analytical model that captures the characteristics of a QoS capable wormhole router which is the basic building block of cluster networks. The model captures the behavior of integrated traffic in a cluster and computes the average deadline missing probability for real-time traffic. The cluster interconnect, considered here, is a hypercube network. Comparison of Deadline Missing Probability (DMP) using the proposed model with that of the simulation shows that our analytical model is accurate and useful. Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das |
NCA | 1 |