Junhyung Um

dblp:19/6432 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 8 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Electronic design automation · 52% Energy-efficient computing · 27% Hardware reliability and fault tolerance · 17%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › physical design
gate sizing
0.522016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Electronic design automation
logic synthesis
0.462016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Synthesis of arithmetic circuits considering layout effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Layout-aware synthesis of arithmetic circuits · DAC 2002
Energy-efficient computing › power management
supply voltage optimization
0.322016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Electronic design automation
physical design
0.332015
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Synthesis of arithmetic circuits considering layout effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Layout-aware synthesis of arithmetic circuits · DAC 2002
Hardware reliability and fault tolerance
aging
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Hardware reliability and fault tolerance › aging › transistor aging
negative bias temperature instability
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Energy-efficient computing
power management
0.212016
OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Energy-efficient computing
low-power design
0.212015
OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions · DAC 2015
Electronic design automation › logic synthesis
arithmetic circuit synthesis
0.142003
Synthesis of arithmetic circuits considering layout effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Layout-aware synthesis of arithmetic circuits · DAC 2002
A practical approach to the synthesis of arithmetic circuits usingcarry-save-adders · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Electronic design automation › physical design
interconnect optimization
0.012003
Synthesis of arithmetic circuits considering layout effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003
Integrated circuit design › digital circuit design › arithmetic circuit design
carry-save adder
0.022002
A practical approach to the synthesis of arithmetic circuits usingcarry-save-adders · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Layout-aware synthesis of arithmetic circuits · DAC 2002
Electronic design automation › logic synthesis
layout-aware synthesis
0.012002
Layout-aware synthesis of arithmetic circuits · DAC 2002
Electronic design automation › logic synthesis › circuit optimization
arithmetic circuit optimization
0.012001
An Optimal Allocation of Carry-Save-Adders in Arithmetic Circuits · IEEE Trans. Computers 2001
Electronic design automation
high-level synthesis
0.012001
An Optimal Allocation of Carry-Save-Adders in Arithmetic Circuits · IEEE Trans. Computers 2001
Electronic design automation › high-level synthesis
data path synthesis
0.012000
A fine-grained arithmetic optimization technique for high-performance/low-power data path synthesis · DAC 2000
Integrated circuit design
low-power circuit design
0.012000
A fine-grained arithmetic optimization technique for high-performance/low-power data path synthesis · DAC 2000
Electronic design automation › logic synthesis › performance-driven synthesis
timing-driven synthesis
0.012000
A practical approach to the synthesis of arithmetic circuits usingcarry-save-adders · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000
Integrated circuit design
digital circuit design
0.022002
Layout-aware synthesis of arithmetic circuits · DAC 2002
An Optimal Allocation of Carry-Save-Adders in Arithmetic Circuits · IEEE Trans. Computers 2001
Electronic design automation › high-level synthesis
register-transfer level synthesis
0.012001
An Optimal Allocation of Carry-Save-Adders in Arithmetic Circuits · IEEE Trans. Computers 2001

Methods — techniques the papers use, named apart from their topics

speed-up heuristic · 0.2design space exploration · 0.2threshold voltage assignment · 0.2optimization · 0.2carry-save-adder module generation · 0.0bit-level interconnect refinement · 0.0polynomial-time algorithm · 0.0multiplexer optimization · 0.0CSA transformation · 0.0
YearPublicationVenuePosition
2016 OSFA: A New Paradigm of Aging Aware Gate-Sizing for Power/Performance Optimizations Under Multiple Operating Conditions
abstract
Modern systems-on-a-chip and microprocessors, e.g., those in smart phones and laptops, typically have multiple operating conditions, such as video streaming, Web browsing, standby, and so on. They will have different performance targets and run under different supply voltages. Gate sizing (with threshold voltage assignment) is a fundamental step for power/performance optimization. However, conventional gate sizing algorithms only consider one scenario, e.g., the performance-critical operating condition, which may be over-design for other operating conditions. In addition, reliability has become a prime concern in nanometer designs, and gate sizing has been employed to mitigate aging. However: 1) previous aging-affected delay models do not take into account more than one operating condition to estimate the aging impact and 2) earlier aging aware gate sizing algorithms only consider one operating condition at a time. In this paper, we present a new paradigm of aging aware gate sizing, one-size-fits-all (OSFA), which performs power/performance optimizations across multiple operating conditions. The existing delay model for negative bias temperature instability (NBTI) is extended to take into account multiple operating conditions, and incorporated into our OSFA framework. Based on OSFA, we also adjust the supply voltage targeting overall power optimization. A speed-up heuristic is proposed to scale our OSFA design space exploration methodology for higher number of operating conditions. Experimental results on industry-strength benchmarks demonstrate that: 1) compared with conventional approach OSFA could provide an average 6.1% reduction in power without performance loss; 2) NBTI-aware OSFA framework can provide significant improvement in comparison with guard-band based traditional NBTI-aware gate sizing approach; and 3) percentage savings compared to conventional methodology increases with the number of operating conditions.
Subhendu Roy, Derong Liu 0002, Jagmohan Singh, Junhyung Um, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2015 OSFA: a new paradigm of gate-sizing for power/performance optimizations under multiple operating conditions
abstract
Modern SoCs and microprocessors, e.g., those in smart phones and laptops, typically have multiple operating conditions, such as video streaming, web browsing, standby, and so on. They will have different performance targets and run under different supply voltages. Gate sizing (with threshold voltage assignment) is a fundamental step for power/performance optimization. However, conventional gate sizing algorithms only consider one scenario, e.g., the performance-critical operating condition, which may be over-design for other operating conditions. In this paper, we present a new paradigm of gate sizing, OSFA (One-Size-Fits-All), which performs power/performance optimizations across multiple operating conditions. Based on OSFA, we also adjust the supply voltage targeting overall power optimization. Experimental results on industry-strength benchmarks demonstrate that compared with conventional approach OSFA could provide an average 6.1% reduction in power without performance loss.
Subhendu Roy, Derong Liu 0002, Junhyung Um, David Z. Pan
DAC3
2009 In-network reorder buffer to improve overall NoC performance while resolving the in-order requirement problem
abstract
Data-intensive functions on chip, e.g., codec, 3D graphics, pixel processing, etc. need to make best use of the increased bandwidth of multiple memories enabled by 3D die stacking via accessing multiple memories in parallel. Parallel memory accesses with originally in-order requirements necessitate reorder buffers to avoid deadlock. Reorder buffers are expensive in terms of area and power consumption. In addition, conventional reorder buffers suffer from a problem of low resource utilization. In our work, we present a novel idea, called in-network reorder buffer, to increase the utilization of reorder buffer resource. In our method, we move the reorder buffer resource and related functions from network entry/exit points to network routers. Thus, the in-network reorder buffers can be better utilized in two ways. First, they can be utilized by other packets without in-order requirements while there are no in-order packets. Second, even in-order packets can benefit from in-network reorder buffers by enjoying more shares of reorder buffers than before. Such an increase in reorder buffer utilization enables NoC performance improvement while supporting the original in-order requirements. Experimental results with an industrial strength DTV SoC example show that the presented idea improves the total execution cycle by 16.9%.
Woo-Cheol Kwon, Sungjoo Yoo, Junhyung Um, Seh-Woong Jeong
DATE3
2006 A systematic IP and bus subsystem modeling for platform-based system design
abstract
The topic on platform-based system modeling has received a great deal of attention today. One of the important tasks that significantly affect the effectiveness and efficiency of the system modeling is the modeling of IP components and communication between IPs. To be effective, it is generally accepted that the system modeling should be performed in two steps; In the first step, a fast but some inaccurate system modeling is considered to facilitate the simultaneous development of software and hardware. The second step then refines the models of the software and hardware blocks (i.e., IPs) to increase the simulation accuracy for the system performance analysis. Here, one critical factor required for a successful system modeling is a systematic modeling of the IP blocks and bus subsystem connecting the IPs. In this respect, this work addresses the problem of systematic modeling of the IPs and bus subsystem in different levels of refinements. In the experiments, we found that by applying our proposed IP and bus modeling methods to the MPEG-4 application, we are able to achieve 4/spl times/ performance improvement and at the same time, reduce the software development time by 35%, compared to that by conventional modeling methods.
Junhyung Um, Woo-Cheol Kwon, Sungpack Hong, Young-Taek Kim, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo, Taewhan Kim 0001
DATE1
2003 Code Placement with Selective Cache Activity Minimization for Embedded Real-time Software Design
Junhyung Um, Taewhan Kim 0001
ICCAD1
2003 Synthesis of arithmetic circuits considering layout effects
abstract
In deep submicron technology, wires are equally or more important than logic components since wire-related problems such as crosstalk noise is much critical in system-on-chip design. Recently, a method for generating a partial product reduction tree with optimal-timing using bit-level adders to implement arithmetic circuits has been proposed, which outperforms the current best designs. However, in the conventional approaches, interconnects are not primary components to be optimized in the synthesis of arithmetic circuits, mainly due to its integration complexity or unpredictable wire effects, thereby resulting in unsatisfactory layout results with long and messy wire connections. To overcome the limitation, we propose a new module generation/synthesis algorithm for arithmetic circuits utilizing carry-save-adder (CSA) modules, which not only optimizes the circuit timing but also generates a much more regular interconnect topology of the final circuits. Specifically, we propose a two-step algorithm: (Phase 1: CSA module generation) we propose an optimal-timing CSA module generation algorithm for an arithmetic expression under a general CSA timing model; then (Phase 2: Bit-level interconnect refinements), we optimally refine the interconnects between the CSA modules while retaining the global CSA-tree structure produced by Phase 1. We show that the timing of the circuits produced by our approach is equal or almost close to that in most test cases (even without including the interconnect delay), and at the same time, the interconnects in layout are short and regular.
Junhyung Um, Taewhan Kim 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2002 Layout-aware synthesis of arithmetic circuits
abstract
In deep sub-micron (DSM) technology, wires are equally or more important than logic components since wire-related problems such as crosstalk, noise are much critical in system-on-chip (SoC) design. Recently, a method [12] for generating a partial product reduction tree (PPRT) with optimal-timing using bit-level adders to implement arithmetic circuits, which outperforms the current best designs, is proposed. However, in the conventional approaches including [12], interconnects are not primary components to be optimized in the synthesis of arithmetic circuits, mainly due to its high integration complexity or unpredictable wire effects, thereby resulting in unsatisfactory layout results with long and messed wire connections. To overcome the limitation, we propose a new module generation/synthesis algorithm for arithmetic circuits utilizing carry-save-adder (CSA) modules, which not only optimizes the circuit timing but also generates a much regular interconnect topology of the final circuits. Specifically, we propose a two-step algorithm: (Phase 1: CSA module generation) we propose an optimal-timing CSA module generation algorithm for an arithmetic expression under a general CSA timing model;(Phase 2: Bit-level interconnect refinements) we optimally refine the interconnects between the CSA modules while retaining the global CSA-tree structure produced by Phase 1. It is shown that the timing of the circuits produced by our approach is equal or almost close to that by [12] in most testcases (even without including the interconnect delay), and at the same time, the interconnects in layout are significantly short and regular.
Junhyung Um, Taewhan Kim 0001
DAC1
2002 Layout-driven resource sharing in high-level synthesis
abstract
In deep submicron (DSM) technology, the interconnects are equally as or more important than the logic gates. In particular, to achieve timing closure in DSM technology, it is very necessary and critical to consider the interconnect delay at an early stage of the synthesis process. It has been known that resource sharing in high-level synthesis is one of the major synthesis tasks which greatly affect the final synthesis/layout results. In this paper, we propose a new layout-driven resource sharing approach to overcome some of the limitations of the previous works in which the effects of layout on the synthesis have never been taken into account or considered in local and limited ways, or whose computation time is excessively large. The proposed approach consists of two steps: (Step 1) We relax the integrated resource sharing and placement into an efficient linear programming (LP) formulation based on the concept of discretizing placement space; (Step 2) We derive a feasible solution from the solution obtained in Step 1. Then, we employ an iterative mechanism based on the two steps to tightly integrate resource sharing and placement tasks so that the slack time violation due to interconnect delay (determined by placement) as well as logic delay (determined by resource sharing) should be minimized. From experiments using a set of benchmark designs, it is shown that the approach is effective, and efficient, completely removing the slack time violation produced by conventional methods.
Junhyung Um, Jae-Hoon Kim 0001, Taewhan Kim 0001
ICCAD1
2001 An Optimal Allocation of Carry-Save-Adders in Arithmetic Circuits
abstract
Carry-save-adder (CSA) is one of the most widely used components for fast arithmetic in industry. This paper provides a solution to the problem of finding an optimal-timing allocation of CSAs in arithmetic circuits. Namely, we present a polynomial time algorithm which finds an optimal-timing CSA allocation for a given arithmetic expression. We then extend our result for CSA allocation to the problem of optimizing arithmetic expressions across the boundary of design hierarchy by introducing a new concept, called auxiliary ports. Our algorithm can be used to carry out the CSA allocation step optimally and automatically and this can be done within the context of a standard RTL synthesis environment.
Junhyung Um, Taewhan Kim 0001
IEEE Trans. Computers1
2000 A timing-driven synthesis of arithmetic circuits using carry-save-adders (short paper)
abstract
Article Free Access Share on A timing-driven synthesis of arithmetic circuits using carry-save-adders (short paper) Authors: Taewhan Kim Department of Computer Science and Advanced Information Technology Research Center (AITrc), Korea Advanced Institute of Science & Technology, Taejon, 305-701, Korea Department of Computer Science and Advanced Information Technology Research Center (AITrc), Korea Advanced Institute of Science & Technology, Taejon, 305-701, KoreaView Profile , Junhyung Um Department of Computer Science and Advanced Information Technology Research Center (AITrc), Korea Advanced Institute of Science & Technology, Taejon, 305-701, Korea Department of Computer Science and Advanced Information Technology Research Center (AITrc), Korea Advanced Institute of Science & Technology, Taejon, 305-701, KoreaView Profile Authors Info & Claims ASP-DAC '00: Proceedings of the 2000 Asia and South Pacific Design Automation ConferenceJanuary 2000 Pages 313–316https://doi.org/10.1145/368434.368656Published:28 January 2000Publication History 3citation201DownloadsMetricsTotal Citations3Total Downloads201Last 12 Months9Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Taewhan Kim 0001, Junhyung Um
ASP-DAC2
2000 A fine-grained arithmetic optimization technique for high-performance/low-power data path synthesis
abstract
Wallace-tree compressor style has been widely recognized as one of the most effective implementation schemes for arithmetic computation sin VLSI design. However, the scheme has been applied only in a rather restrictive way, that is, for implementing fast multipliers and for generating fixed structures without considering the characteristic of the input signals. The contributions of our work are (1) to extend the applicability of the Wallace scheme to any arithmetic circuit which consists of additions/substractions/multiplications globally (instead of applying it to each operation) to produce a globally efficient architecture of the circuit; (2) to optimize the timing of the circuit for uneven signal arrival profiles; (Specifically, we present an efficient algorithm for generating a delay-optimal (bit-level) carry-save addition structure of an arithmetic circuit.) (3) to provide a comprehensive analysis of the switching activity of a (bit-level) carry-save addition structure, and based on which we derive an effective algorithm for synthesizing low power circuits. Putting these arithmetic optimization solutions together, a circuit designer will be able to fully understand the synthesis of arithmetic circuit based on the bit-level carry-save addition.
Junhyung Um, Taewhan Kim 0001, C. L. Liu 0001
DAC1
2000 A practical approach to the synthesis of arithmetic circuits usingcarry-save-adders
abstract
Carry-save-adder (CSA) is one of the most widely used types of operation in implementing a fast computation of arithmetics. An inherent limitation of the conventional CSA applications is that the applications are confined to the sections of arithmetic circuit that can be directly translated into addition expressions. To overcome this limitation, from the analysis of the structures of arithmetic circuits found in industry, we derive a set of simple, but effective CSA transformation techniques other than the existing ones. These are 1) optimization across multiplexers, 2) optimization across design boundaries, and 3) optimization across multiplications. Based on the techniques, we develop a new timing-driven CSA transformation algorithm that is able to utilize CSA's extensively throughout all circuits. Experimental data for practical testcases are provided to show the effectiveness of our algorithm.
Taewhan Kim 0001, Junhyung Um
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Optimal allocation of carry-save-adders in arithmetic optimization
abstract
Carry-save-adder(CSA) is one of the most widely used schemes for fast arithmetic in industry. This paper provides a solution to the problem of finding an optimal-timing allocation of CSAs. Specifically, we present a polynomial time algorithm which finds an optimal-timing CSA allocation for a given arithmetic expression. In addition, we extend our result for CSA allocation to the problem of optimizing arithmetic expressions across the boundary of design hierarchy by introducing a new concept, called auxiliary ports. Our algorithm can be used to carry out the CSA allocation step optimally and automatically, and this can be done within the context of a standard HDL synthesis environment.
Junhyung Um, Taewhan Kim 0001, C. L. Liu 0001
ICCAD1