Jinn-Shyan Wang

dblp:94/2690 · DBLP profile ↗
← Back
49ranked-venue papers
13as first author
4since 2021 · last 2024
0000-0001-7638-0802ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 13 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1
YearPublicationVenuePosition
2024 A 40-nm 13.88-TOPS/W FC-DNN Engine for 16-bit Intelligent Audio Processing Featuring Weight-Sharing and Approximate Computing
abstract
This work presents a novel accelerating engine to efficiently compute the fully connected (FC) deep neural networks (DNN) based on weight sharing (WS). Beyond the highly reduced weights and weight accesses in previous designs, the proposed engine significantly reduces the multiplications by exploiting the distributive law of shared weights. All-digital in-memory computing (IMC) and approximate computing (AC) were applied to accelerate accumulations and improve energy efficiency. We have designed and implemented a test chip in 40nm CMOS for intelligent audio processing of dysarthric voice conversion (i.e., using an FCDNN composed of neurons with 16-bit inputs, 16-bit weights, and 16-bit activations). When turning AC off and on, the energy efficiency for 16b operation at 0.9V is 7.04 TOPS/W and 13.88 TOPS/W, respectively.
Tay-Jyi Lin, Yun-Cheng Chen, Chien-Tung Liu, Tien-Fu Chen, Jinn-Shyan Wang
HCS6
2024 Design of synthesizable period-jitter sensor IP with high power reduction and variation resiliency
Jinn-Shyan Wang, Yu-Hsuan Kuo
Integr.1
2024 Clock Period-Jitter Measurement With Low-Noise Runtime Calibration for Chips in FinFET CMOS
abstract
Process migration from bulk to FinFET CMOS pronounces die-to-die and within-die process variations. Runtime variations get severe as the clock frequency goes higher. Facing these two issues makes on-chip clock period-jitter measurement with a high resolution very difficult. In this work, we propose low-noise runtime resolution calibration and jitter measurement for designing a period-jitter measurement circuit, called a period-jitter sensor (PJS), in the face of higher variations imposed on the PJS. Essential design techniques include edge-triggered and symmetrical architecture and circuits for noise reduction, variation resiliency, and power saving. As a test vehicle, we have designed a PJS to measure the clock quality of an LPDDR4-4266 physical layer with the clock cycle time and the maximum period jitter specified to be 468.82ps and ±30ps, respectively. The design specifies the minimum runtime resolution to be better than 1.0 ps. Measurement results show that the 0.0192mm$^{2}~2.133$GHz PJS in 14nm FinFET CMOS achieves a sub-ps resolution in runtime and only consumes around 1mW across all PVT conditions, making multiple embedding of the PJS in a complex SoC feasible.
Jinn-Shyan Wang, Pei-Yuan Chou
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 A 40nm CMOS SoC for Real-Time Dysarthric Voice Conversion of Stroke Patients
abstract
This paper presents the first dysarthric voice conversion SoC, which can translate stroke patients' voice into more intelligible and clearer speech in real time. The SoC is composed of a RISC-V MPU and a compact DNN engine with a single 16-bit multiply-accumulator, which improves 12x performance and > 100x energy efficiency, and has been implemented in 40nm CMOS. The silicon area is 0.68×0.79mm2, and the measured power is 18.4mW for converting 3-sec dysarthric voice within 0.5 sec (at 200MHz and 0.8V) and 4.8mW for conversion < 1 sec (at 100MHz and 0.6V).
Tay-Jyi Lin, Chen-Zong Liao, You-Jia Hu, Wei-Cheng Hsu, Zheng-Xian Wu, Shao-Yu Wang, Chun-Ming Huang, Ying-Hui Lai, Chingwei Yeh, Jinn-Shyan Wang
ASP-DAC10
2020 A 0.21V 40nm NAND-ROM for IoT Sensing Systems with Long Standby Periods
abstract
IoT sensing systems usually have long standby periods to lengthen the battery lifetime. Low active and standby power consumption is an indispensable design goal for devices, such as the read-only memory (ROM) for code storage, used in these systems. Sub-threshold (sub-Vt) designs can help accomplish the goal. A 90nm NAND-ROM achieved a state-of-the-art minimum supply voltage (Vmin) of 0.25V by using a source-line control scheme to reduce the impact of leakage and noise and using a code-inversion-based flag-table (CIB-FT) and a data-aware sensing reference (DASR) scheme for performance improvement. For advanced IoT sensing systems designed in a more advanced CMOS process with a more aggressive Vmin, more considerable leakage and higher PVT variations become the main design challenges of the sub-Vt ROM. In this paper, we first illustrate the design challenges and considerations to reach the goal of a Vmin of no higher than 0.25V for the state-of-the-art ROM redesigned in 40nm CMOS. We then propose new design techniques, including flag-table-free architecture and the new bit-line load and sense amplifier for meeting the higher stringent design specifications. Comparison results according to post-layout simulations show that the proposed 40nm ROM achieve a 17%, 16%, 43%, and 91% reduction in area, Vmin, active power, and standby leakage power, respectively, compared to the redesigned NAND-ROM.
Jinn-Shyan Wang, Cheng-Xin Xue, Chien-Tung Liu, Tay-Jyi Lin
ISCAS1
2017 ULV-Turbo Cache for an Instantaneous Performance Boost on Asymmetric Architectures
abstract
An asymmetric architecture is commonly used in modern embedded systems to reduce energy consumption. The systems tend to execute more applications in the energy-efficient core, which typically employs ultralow voltage (ULV) to save energy. However, caches become a reliability and performance barrier that limits the minimum operating voltage and blocks system performance in the ULV environment. The poor performance of an ultralow-voltage core causes most workload requirements to awaken and then execute on the host core, leading to high energy consumption. In this paper, we propose a ULV-Turbo cache based on a ULV-selective-ally 8T static random access memory (SRAM) that is able to perform reliable ultralow-voltage operation and provide the speedup function of SRAM rows ally. The system is able to speed up the ULV core instantaneously and execute more applications with the ULV-Turbo cache. In our system-wide evaluation based on a real attitude and heading reference system workload on an asymmetric wearable system, the ULV-Turbo cache reduces the energy consumption of the entire system by approximately 36%.
Po-Hao Wang, Yung-Chen Chien, Shang-Jen Tsai, Xuan-Yu Lin, Rizal Tanjung, Yi-Sian Lin, Shu-Wei Syu, Tay-Jyi Lin, Jinn-Shyan Wang, Tien-Fu Chen
IEEE Trans. Very Large Scale Integr. Syst.9
2017 Process/Voltage/Temperature-Variation-Aware Design and Comparative Study of Transition-Detector-Based Error-Detecting Latches for Timing-Error-Resilient Pipelined Systems
abstract
Among timing-error-detecting registers, the transition-detector-based error-detecting latch (TD-EDL) is regarded as the most lightweight design. To realize a truly variation-resilient system, a TD-EDL should be designed to be robust across all process/voltage/temperature (PVT) corners. In particular, its detection window deviation must be carefully compensated for an accurate detection. It must also be small, as well as power and energy efficient, to prevent excessive overheads. This paper carefully redesigned several conventional TDs in 28-nm CMOS process, and postlayout simulations were performed to examine the pros and cons of the TDs. According to the characteristics of each TD, the transistor sizing methods were explored to achieve PVT-variation-aware designs. The redesigned TDs were then used to construct TD-EDLs, for which the causes of detection window deviation were examined and improved upon. The timing characteristics that affect the correct function and appropriate responses were discussed, and the area overhead, timing performance, power consumption, and voltage scalability of the TD-EDLs were evaluated and compared.
Jinn-Shyan Wang, Shih-Nung Wei
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Design of an all-digital temperature sensor in 28 nm CMOS using temperature-sensitive delay cells and adaptive-1P calibration for error reduction
abstract
We describe design techniques, calibration method, and measurement results of an all-digital temperature sensor in 28 nm CMOS. To deal with the issue of Vcc being near the zero-temperature-coefficient point, a new delay cell with much improved temperature sensitivity is proposed. Adaptive 1-point (1P) calibration is proposed to reduce the serious impact due to process variations, while without increasing the calibration cost. Measurement results show that, compared to the conventional 1P calibration, the new method achieves a 32% error reduction.
Shang-Yi Li, Pei-Yuan Chou, Jinn-Shyan Wang
ASP-DAC3
2016 Cross-matching caches: Dynamic timing calibration and bit-level timing-failure mask caches to reduce timing discrepancies with low voltage processors
Po-Hao Wang, Shang-Jen Tsai, Rizal Tanjung, Tay-Jyi Lin, Jinn-Shyan Wang, Tien-Fu Chen
Integr.5
2014 Low power fixed-latency DSP accelerator with autonomous minimum energy tracking (AMET)
abstract
Presents a slide covering the following: low power fixed-latency DSP accelerator; autonomous minimum energy tracking; adaptive voltage scaling; and intelligent voltage generation.
Chung-Hsun Huang, Wei-Jen Chen, Keng-Jui Chang, Yi-Hsuan Ting, Keng-Chang Hsu, Yu-Fu Pan, Chao-Chun Chen, Yuan-Hua Chu, Tay-Jyi Lin, Jinn-Shyan Wang
Hot Chips Symposium10
2013 An energy-efficient truly all-digital temperature sensor for SoC applications
abstract
This paper presents a truly all-digital one-point-calibrated temperature sensor for SoC applications. Several design techniques are developed to achieve small area, low power, low energy per sample, and high accuracy. First, in contrast to dual DLL's plus charge pump used in the traditional design, single DLL is used in both calibration and measurement operations to avoid locking mismatch and to increase process-variation tolerance, resulting in advantages in terms of area, power, energy, and accuracy. Second, a normalization-after-averaging design is developed to reduce computation cycles, also leading to low power and low energy. Finally, to meet different applications, clock chopping is used not only for adjustable conversion rate but also for low power when a high conversion rate is not needed. Compared to the state-of-the-art design, the proposed 90nm temperature sensor has 65% less normalized area and 26% higher accuracy with 79% less energy per sample. As the sampling rate is adjusted 1/50 lower, 90% of the power consumption is saved.
Tzu-Yuan Kuo, Keng-Jui Chang, Jen-Hsiang Lee, Zong-Wu He, Jinn-Shyan Wang
ACM Great Lakes Symposium on VLSI5
2013 Variation-aware and adaptive-latency accesses for reliable low voltage caches
abstract
Contemporary cache is known for consuming a large part of total power in microprocessors. Voltage scaling had been used to reduce the power consumption of the cache. However, due to the impact of variations, SRAM cells of the cache could potentially fail when voltage dropping. To against variations, we need to increase the supply voltage for the safety margin, thus the cache costs large energy consumption. For eliminating the voltage safety margin, some prior works for SRAM failure tolerance designs were proposed. These schemes will result in worse energy consumption and cannot deal with dynamic variations. They still have a safety margin to resist dynamic variations. With the supply voltage scaling down, we find out that the major reason of failures is that some slow cells have longer latency. We call these cell faults as “latency fault”. If each cache line can be accessed in an appropriate access time, the slower cells could be reused but not disable them. We propose a VAL-Cache adapting the access time to tolerate latency faults and which is able to scale down the voltage. And we also propose the latency-fault detector to detect latency faults at run-time so as to tolerate both static and dynamic variations. Our experimental results on Mibench and 0xbench benchmarks demonstrate that the energy consumption can be reduced 10%~18% in average at a cost of acceptable performance loss.
Po-Hao Wang, Wei-Chung Cheng, Yung-Hui Yu, Tang-Chieh Kao, Chi-Lun Tsai, Pei-Yao Chang, Tay-Jyi Lin, Jinn-Shyan Wang, Tien-Fu Chen
VLSI-SoC8
2013 A high-throughput and high-capacity IPv6 routing lookup system
Yi-Mao Hsiao, Yuan-Sun Chu, Jeng-Farn Lee, Jinn-Shyan Wang
Comput. Networks4
2013 Embedding Repeaters in Silicon IPs for Cross-IP Interconnections
abstract
During systems-on-a-chip (SoC) integration, silicon intellectual properties (IPs) are generally regarded as blockages to long interconnections that connect different IPs. With this constraint, conventional designs are forced to place those repeaters that drive long interconnections outside the IP. These designs either lead to a longer interconnection distance requiring more repeaters or result in a longer signal delay, since the interconnection wire is not appropriately segmented by the repeaters. To solve these problems, we designed the IPs such that designers can embed the repeaters in the IP for the SoC integration. In other words, it allows the cross-IP interconnections to be routed over the IP using repeaters inserted in the IP. The design concept, physical implementation, and application examples of the embedded repeaters are described in this brief. Experimental results show that the proposed design will not only make the floor plan of the SoC easier but will also improve the signal delay and the power consumption of the long interconnection circuits.
Jinn-Shyan Wang, Keng-Jui Chang, Chingwei Yeh, Shih-Chieh Chang 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2012 A Scalable High-Performance Virus Detection Processor Against a Large Pattern Set for Embedded Network Security
abstract
Contemporary network security applications generally require the ability to perform powerful pattern matching to protect against attacks such as viruses and spam. Traditional hardware solutions are intended for firewall routers. However, the solutions in the literature for firewalls are not scalable, and they do not address the difficulty of an antivirus with an ever-larger pattern set. The goal of this work is to provide a systematic virus detection hardware solution for network security for embedded systems. Instead of placing entire matching patterns on a chip, our solution is a two-phase dictionary-based antivirus processor that works by condensing as much of the important filtering information as possible onto a chip and infrequently accessing off-chip data to make the matching mechanism scalable to large pattern sets. In the first stage, the filtering engine can filter out more than 93.1% of data as safe, using a merged shift table. Only 6.9% or less of potentially unsafe data must be precisely checked in the second stage by the exact-matching engine from off-chip memory. To reduce the impact of the memory gap, we also propose three enhancement algorithms to improve performance: 1) a skipping algorithm; 2) a cache method; and 3) a prefetching mechanism.
Chieh-Jen Cheng, Chao-Ching Wang, Wei-Chun Ku, Tien-Fu Chen, Jinn-Shyan Wang
IEEE Trans. Very Large Scale Integr. Syst.5
2012 Towards Process Variation-Aware Power Gating
abstract
This paper presents a power gating design that considers process variation for proper wakeup control. First, the surge current constraint is examined and refined for a simpler and more realistic view of inter-module reliability. Following that, several circuits are proposed on top of a delay chain to adapt the timing control of power switches to process variations. Experimental results show that the proposed design is able to track process variation such that the surge current and the wakeup time are both kept to expectation in all process corners.
Chingwei Yeh, Yuan-Chang Chen, Jinn-Shyan Wang
IEEE Trans. Very Large Scale Integr. Syst.3
2011 A H.264/MPEG-2 dual mode video decoder chip supporting temporal/spatial scalable video
abstract
This paper proposes a dual mode video decoder with 4-level temporal/spatial scalability and 32/64-bit adjustable memory bus width. A design automation environment for simulation and verification is established to automatically verify the correctness and completeness of the proposed design. Using a 0.13 um CMOS technology, it comprises 439Kgates/10.9KB SRAM and consumes 2~328mW in decoding CIF~HD1080 videos at 3.75~30fps when operating at 1~150MHz, respectively.
Cheng-An Chien, Yao-Chang Yang, Hsiu-Cheng Chang, Cheng-Yen Chang, Jiun-In Guo, Jinn-Shyan Wang, Ching-Hwa Cheng
ASP-DAC7
2010 A 55nm 1GHz one-cycle-locking de-skewing circuit
abstract
This paper presents the design of a 55nm 1.0V 1GHz open-loop de-skewing circuit for clock synchronization in a SoC. All-digital architecture and circuit design techniques are adopted and developed to achieve small jitter, low power, super fast locking, and input duty-cycle independence. Even running at 1 GHz, the proposed circuit wakes up from sleep with only one-cycle locking time. Measurement results show a peak-to-peak jitter of 4 ps and a static phase error of 5.46 ps at 1 GHz. To the best of our knowledge, this is the first GHz open-loop de-skewing circuit achieving sub-μW/MHz active power. By power gating, 84% of leakage power is saved in the sleep mode.
Jinn-Shyan Wang, Chun-Yuan Cheng, Je-Ching Liu, Yu-Chia Liu
ISCAS1
2009 A dynamic quality-scalable H.264 video encoder chip
abstract
This paper proposes a dynamic quality-scalable H.264 video encoder that comprises 470Kgates and 13.3Kbytes SRAM using 1P8M 0.13μm CMOS technology. Exploiting parameterized algorithms for motion estimation and intra prediction, the proposed design can dynamically configure the encoding modes with the design trade-off between power consumption and video quality for various video encoding applications. It achieves real-time H.264 video encoding on CIF, D1, and HD720@30fps with 7mW-25mW, 27mW-162mW, and 122mW–183mW power dissipation in different quality modes.
Hsiu-Cheng Chang, Yao-Chang Yang, Ching-Lung Su, Cheng-An Chien, Jiun-In Guo, Jinn-Shyan Wang
ASP-DAC7
2009 No cache-coherence: a single-cycle ring interconnection for multi-core L1-NUCA sharing on 3D chips
abstract
Consistent with the trend towards the use of many cores in SOC and 3D Chip techniques, this paper proposes a "single-cycle ring" interconnection (SC_Ring) with ultra-low latency and minimal complexity. The proposed SC_Ring allows multiple single-cycle transactions in parallel. The main features of the circuit-switched design include a set of 3-ported circuit-switched routers (4~16) and a performance/timing effective arbiter. The arbiter, called "BTPC", features single-cycle arbitration and routing-control by means of the novel Binary-Tree paths convergence and path-prediction mechanisms, to provide a highly reduced time complexity. By combining this with the integration of 3D chips, the proposed ring-based interconnection offers several advantages for hierarchical clustering in future many-core systems, in terms of cost, latency, and power reductions. Moreover, based on the proposed SC_Ring, this work realizes a "level-1 non-uniform cache architecture" (L1-NUCA) for fast data communication without cache-coherency in facilitating multithreading/multi-core as a case study. Finally, experimental results show that our approach yields promising performance.
Shu-Hsuan Chou, Chien-Chih Chen, Chi-Neng Wen, Yi-Chao Chan, Tien-Fu Chen, Chao-Ching Wang, Jinn-Shyan Wang
DAC7
2009 A Dynamic Quality-scalable H.264 Video Encoder
abstract
This demo proposes a dynamic quality-scalable H.264 video encoder that has been published in ISSCC2007 [8]. Exploiting parameterized algorithms for motion estimation and intra prediction, the proposed design can dynamically configure the encoding modes with the design trade-off between power consumption and video quality for various video applications.
Hsiu-Cheng Chang, Yao-Chang Yang, Cheng-An Chien, Tzu-Chun Chang, Jinn-Shyan Wang, Jiun-In Guo
ISCAS6
2009 A Dynamic Quality-Adjustable H.264 Video Encoder for Power-Aware Video Applications
abstract
This paper proposes a dynamic quality-adjustable H.264 baseline profile (BP) video encoder that comprises 470 Kgates and 13.3 kB SRAM in a core size of 4.3 × 4.3 mm2using TSMC 0.13 ¿m 1P8M CMOS technology. Exploiting parameterized algorithms for motion estimation and intra prediction, the proposed design can dynamically configure the encoding modes with the design trade-off between power consumption and video quality for various video encoding applications. In addition, the proposed basic unit (BU)-based rate control hardware can maintain a constant and stable bit rate for network video transmission. It achieves real-time H.264 video encoding on CIF, D1, and HD720@30 frames/s with 7 mW to 25 mW, 27 mW to 162 mW, and 122 mW to 183 mW power dissipation in different quality modes.
Hsiu-Cheng Chang, Bing-Tsung Wu, Ching-Lung Su, Jinn-Shyan Wang, Jiun-In Guo
IEEE Trans. Circuits Syst. Video Technol.5
2009 VisoMT: A Collaborative Multithreading Multicore Processor for Multimedia Applications With a Fast Data Switching Mechanism
abstract
Multithreading and multicore processing are powerful ways to take advantage of parallelism in applications in order to boost a system's performance. However, exploring sufficient parallelism and achieving data locality with low communication overhead are still important research issues in embedded multithreading/multicore design. This paper introduces the design of a fast data switching mechanism between multilevel storage structures in a new multicore architecture. This paper makes several contributions to the development of contemporary sophisticated multimedia applications with advanced standards such as H.264. The first contribution,collaborative-multithreading, tightly unifies reduced instruction set computer and collaborative multithreading digital signal processing (DSP) in order to exploit high parallelism to provide sufficient computing power to applications. Each collaborative thread of our DSP is constructed by a heterogeneous-simultaneously multithreading single instruction, multiple data structure, and four media processing cores, which is connected by a fast switch for providing a fast data exchange mechanism among correlative streams on a thread-level basis. Our second contribution isone-stop streaming processing, which aims to keep data in the system for as long as possible until it is no longer needed, thus making data more efficient to access. Our third contribution is achunk threading programming model, including a thread management library and threading communication directives for reducing data communication and synchronization overhead. By a combination of coarse-grained and fine-grained threading, programmers can choose various threading levels based on the amount of data exchange in a program. With our proposed techniques and an appropriate programming model, we can reduce processing time by 54.9% in H.264 video encoding (common intermediate format video at 16.574 f/s) with the 1-virtual independent and streaming processing by open collaborative multithreading configuration, compared to the Texas Instruments C62 core that owns 8 function units. We realize our design as a prototype by chip implementation, and fabricate it as a chip based on the Taiwan Semiconductor Manufacturing Company Ltd. 0.13$\mu {\rm m}$process. The die size of the processor core is 16.12${\rm mm}^{2}$, including 414 k logic transistors and 34.4 kB of on-chip static random access memory. The processor runs at 180 MH0z/1.2-V and consumes 245 mW by postsimulation results.
Wei-Chun Ku, Shu-Hsuan Chou, Jui-Chin Chu, Chi-Lin Liu, Tien-Fu Chen, Jiun-In Guo, Jinn-Shyan Wang
IEEE Trans. Circuits Syst. Video Technol.7
2008 A low-voltage latch-adder based tree multiplier
abstract
For achieving low power and high performance simultaneously, we propose a new low-voltage latch-adder based Wallace-tree multiplier. By choosing the best circuit-structure of the latch-adder for low voltage, while optimizing the number and positions of latch-adders, the proposed 0.18-μm 0.9V 32×32 2’s complement multiplier can operate above 60MHz. As compared to the tree multiplier implementing the traditional latch-adder technique, the new tree multiplier with all the proposed techniques achieves a 22.3∼23.7% delay improvement with a 5.5∼3.3% power reduction. All best latch-adder based tree multipliers have a smaller power-delay-product than the tree multipliers without using the latch-adder technique.
Tzu-Yuan Kuo, Jinn-Shyan Wang
ISCAS2
2006 Low Complexity Architecture Design of H.264 Predictive Pixel Compensator for HDTV Application
abstract
In this paper, we propose a low-complexity architecture design of H.264 predictive pixel compensator (PPC) for HDTV application. In intra prediction, we propose a shared adder-based architecture style that supports all of the 17 intra prediction modes, and reduce computational complexity in the I4MB prediction mode 3~8 up to 50% computation. Besides, we have also proposed the distributed memory access to improve the HW usage. As well as, it can used to reduce the memory size for buffering the neighboring pixels. In inter prediction, we can save about 48% of external memory bandwidth by the data reused through the hybrid block size memory access. Adopting the mixed six-tap FIR filter architecture to design luma interpolation, we can efficiently reduce the hardware cost up to 27%. The implemental result shows the hardware cost of the proposed design is about 60854 gates under a TSMC 0.18 μ m CMOS technology, which achieves the real-time processing requirement for HD-1080 format video@30Hz at the working frequency of 87 MHz.
Chien-Chang Lin, Jiun-In Guo, Jinn-Shyan Wang
ICASSP (3)4
2006 A Condition-based Intra Prediction Algorithm for H.264/AVC
abstract
This paper proposes a condition-based algorithm for H.264/AVC 4times4 intra prediction. Exploiting high correlation existed in neighboring intra prediction modes, we propose the three conditions to skip the less possible candidates in doing intra4times4 block mode decision. When compared to the 9 prediction modes in the full search algorithm, the proposed algorithm can complete a 4times4 intra prediction using 4.4 prediction modes operation in average. The simulation result shows that the proposed algorithm can reduce computational complexity up to 44% at the cost of less than 0.1 dB PSNR loss in average
Chun-Hao Chang, Chien-Chang Lin, Yi-Huan Yang, Jiun-In Guo, Jinn-Shyan Wang
ICME6
2006 A performance-aware IP core design for multimode transform coding using scalable-DA algorithm
abstract
This paper proposes a performance-aware transform IP design which can be configured to appropriate hardware for different performance requirements on demand without requiring additional data bandwidth in multimode video coding (JPEG/MPEG-1/2/4/H.261/H.263/H.264). Based on the scalable-DA approach, three schemes of hardware configurations which are respectively composed of 3, 6, and 12 data-paths are illustrated. The three schemes of the proposed performance-aware DCT/IDCT can achieve CIF, 720HD, and digital cinema video formats when operated at 9.13 MHz, 41.48 MHz, and 188.75 MHz, respectively.
Kuan-Hung Chen, Jinn-Shyan Wang, Jiun-In Guo
ISCAS3
2006 Design of STR level converters for SoCs using the multi-island dual-VDD design technique
abstract
In a 0.13 /spl mu/m design environment, we design two level converters to convert signals from 0.6V and 0.8V to 1.2V, respectively, to fulfill the needs of a multi-island dual-VDD CMOS SOC. Heuristic sizing guidelines are proposed to achieve better conversion speed and lower conversion energy for both level converters.
Jinn-Shyan Wang, Yu-Juey Chang, Chingwei Yeh, Yuan-Hua Chu
ISCAS1
2006 An improved SAR controller for DLL applications
abstract
The conventional SAR controller for DLL applications is shown in this paper to have the problem of dead lock. With this problem, the DLLs adopting the SAR controller are useless because they will go unlocked forever in face of supply voltage and temperature variations. An improved SAR controller is proposed to overcome this problem and then to rescue the DLLs
Jinn-Shyan Wang, Chun-Yuan Cheng, Yu-Chia Liu
ISCAS1
2006 A high-performance direct 2-D transform coding IP design for MPEG-4AVC/H.264
abstract
This paper proposes a high-performance direct two-dimensional transform coding IP design for MPEG-4 AVC/H.264 video coding standard. Because four kinds of 4 /spl times/ 4 transforms, i.e., forward, inverse, forward-Hadamard, and inverse-Hadamard transforms are required in a H.264 encoding system, a high-performance multitransform accelerator is inevitable to compute these transforms simultaneously for fitting real-time processing requirement. Accordingly, this paper proposes a direct 2-D transform algorithm which suitably arranges the data processing sequences adopted in row and column transforms of H.264 CODEC systems to finish the data transposition on-the-fly. The induced new transform architecture greatly increases the data processing rate up to 8 pixels/cycle. In addition, an interlaced I/O schedule is disclosed to balance the data I/O rate and the data processing rate of the proposed multitransform design when integrated with H.264 systems. Using a 0.18-/spl mu/m CMOS technology, the optimum operating clock frequency of the proposed multitransform design is 100 MHz which achieves 800 Mpixels/s data throughput rate with the cost of 6482 gates. This performance can achieve the real-time multitransform processing of digital cinema video (4096 /spl times/ 4 2048@30 Hz). When the data throughput rate per unit area is adopted as the comparison index in hardware efficiency, the proposed design is at least 1.94 times more efficient than the existing designs. Moreover, the proposed multitransform design can achieve HDTV 720p, 1080i, digital cinema video processing requirements by consuming only 0.58, 2.91, and 24.18 mW when operated at 22, 50, and 100 MHz with 0.7, 1.0, and 1.8 V power supplies, respectively.
Kuan-Hung Chen, Jiun-In Guo, Jinn-Shyan Wang
IEEE Trans. Circuits Syst. Video Technol.3
2005 An efficient spurious power suppression technique (SPST) and its applications on MPEG-4 AVC/H.264 transform coding design
abstract
This paper proposes an efficient Spurious Power Suppression Technique (SPST) and its applications on an MPEG-4 AVC/H.264 transform coding design. There are three techniques addressed in this paper, which are (1) the SPST, (2) the direct 2-D algorithm, and (3) the interlaced I/O schedule to solve the design challenges induced by both the real-time processing and low-power requirements. The major novelty of this paper is implementing the SPST concept on the transform architecture for H.264, which save 31.9% power consumption at the cost of 20.9% area price. Moreover, the proposed transform design also possesses 60.05% higher hardware efficiency through the TPUA index than the existing designs
Kuan-Hung Chen, Kuo-Chuan Chao, Jinn-Shyan Wang, Yuan-Sun Chu, Jiun-In Guo
ISLPED3
2005 An Energy-Aware IP Core Design for the Variable-Length DCT/IDCT Targeting at MPEG4 Shape-Adaptive Transforms
abstract
This paper proposes a flexible hardware solution and the associated energy-aware IP core design for computing the variable-length discrete cosine transform/inverse discrete cosine transform (DCT/IDCT) required in the MPEG4 shape-adaptive DCT/IDCT (SA-DCT/IDCT). The proposed IP core has been developed based on the design concept of programmable processors to provide the flexibility in dynamically configuring the hardware. To achieve good performance both in area and speed, we optimize the proposed IP core both in the algorithmic computational complexity and hardware complexity. Furthermore, the proposed IP core possesses the feature of energy-aware design flexibility. The simulation shows that the proposed design has 44% energy reduction at the price of 0.3-dB signal quality degradation for the image compression applications. The implementation results show that the proposed IP core costs about 3100 gates along with 16 words (1 word = 16 bits) of memory, which can achieve the real-time processing of the texture coding in MPEG4 SP@L3 and ACE@L2 CODEC system for the CIF format video at 30 frames/s with 4:2:0 color format.
Kuan-Hung Chen, Jiun-In Guo, Jinn-Shyan Wang, Chingwei Yeh
IEEE Trans. Circuits Syst. Video Technol.3
2004 A reliable low-power fast skew-compensation circuit
Jinn-Shyan Wang
ASP-DAC2
2004 A power-aware SNR-progressive DCT/IDCT IP core design for multimedia transform coding
abstract
A power-aware SNR progressive DCT/IDCT IP core design for multimedia transform coding is proposed. The proposed IP core possesses the feature of power-aware design flexibility, allowing the trade-off of lower power consumption with less demand of data precision in developing its instruction library. Relationships of energy reduction and data quality degradation in the examples of both JPEG still images and MPEG4 video sequences have been analyzed. Since the proposed IP core is developed based on the concept of programmable processors, we can select a DCT/IDCT firmware library of different precisions of cosine coefficients according to the accuracy requirement of various applications. This design has been realized based on a 0.35-/spl mu/m CMOS technology and costs about 2175 gates with 8 words of RAM, which can achieve real-time processing of the texture coding in an MPEG4 SP@L3 codec system for CIF video at 30 frames per second (fps).
Kuan-Hung Chen, Jiun-In Guo, Jinn-Shyan Wang, Chingwei Yeh
ICME3
2004 Low-power fixed-width array multipliers
abstract
A fixed-width multiplier using the left-to-right algorithm for partial-product reduction is presented. The high-speed feature offered by this design is used to trade for low power. In one design, the proposed multiplier not only owns 8 % speed improvement but also gains 14 % power and 13 % area reduction. When applying the voltage scaling to balance the speed, the power reduction is increased to 29%. Categories and Subject Descriptors
Jinn-Shyan Wang, Chien-Nan Kuo, Tsung-Han Yang
ISLPED1
2003 Design theory and implementation for low-power segmented bus systems
abstract
The concept of bus segmentation has been proposed to minimize power consumption by reducing the switched capacitance on each bus [Chen et al. 1999]. This paper details the design theory and implementation issues of segmented bus systems. Based on a graph model and the Gomory-Hu cut-equivalent tree algorithm, a bus can be partitioned into several bus segments separated by pass transistors. Highly communicating devices are placed to adjacent bus segments, so most data communication can be achieved by switching a small portion of the bus segments. Thus, a significant amount of power consumption can be saved. It can be proved that the proposed bus partitioning method achieves an optimal solution. The concept of tree clustering is also proposed to merge bus segments for further power reduction. The design flow, which includes bus tree construction in the register-transfer level and bus segmentation cell placement and routing in the physical level, is discussed for design implementation. The technology has been applied to a μ-controller design, and simulation results by PowerMill show significant improvement in power consumption.
Wen-Ben Jone, Jinn-Shyan Wang, Hsueh-I Lu, I. P. Hsu, J.-Y. Chen
ACM Trans. Design Autom. Electr. Syst.2
2001 Charge-sharing alleviation and detection for CMOS domino circuits
abstract
Charge sharing, which occurs in any complementary metal-oxide-semiconductor (CMOS) domino gate, may degrade the output voltage level or may even cause an erroneous output value. In this paper, this problem is thoroughly investigated by considering circuit topology and circuit function. We describe a method to measure the sensitivity [called charge-sharing (CS) vulnerability] of the CS problem for each domino gate. A method to derive the CS vulnerability and the test vector for each domino gate is suggested. We also propose a transistor reordering method to dramatically reduce the CS vulnerabilities for all domino gates so that the CS problem can be alleviated. We also prove theoretically that a set of test vectors generated for single charge-sharing faults (SCSFs) can also detect all multiple charge-sharing faults (MCSFs). This good property significantly guarantees the test quality for the CS faults of domino circuits.
Shih-Chieh Chang 0001, Ching-Hwa Cheng, Wen-Ben Jone, Shin-De Lee, Jinn-Shyan Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2000 A new CMAC neural network architecture and its ASIC realization
abstract
-, -" " " ( " $# , " .#$ , " , ( .* , / ( , ( " " 0 (( , #$ ,( ( ( " / ! ( )*+ ( " 1 ( "-)2+ .. .* , 3 (( 4 ( $ ( ( )5+ " ( 0 ,( , ( #$ , " "3 , 670! , #$ ( , (" " (( %$#& $# , (
Yuan-Bao Hsu, Kao-Shing Hwang, Chien-Yuan Pao, Jinn-Shyan Wang
ASP-DAC4
2000 Power analysis and implementation of a low-power 300 MHz 8-b × 8-b pipelined multiplier
abstract
Article Power analysis and implementation of a low-power 300 MHz 8-b × 8-b pipelined multiplier Share on Authors: Jinn-Shyan Wang Department of Electrical Engineering, National Chung Cheng University, 160, San-Hsing, Ming-Hsiung, Chia-Yi, 621, Taiwan Department of Electrical Engineering, National Chung Cheng University, 160, San-Hsing, Ming-Hsiung, Chia-Yi, 621, TaiwanView Profile , Po-Hui Yang Department of Electrical Engineering, National Chung Cheng University, 160, San-Hsing, Ming-Hsiung, Chia-Yi, 621, Taiwan Department of Electrical Engineering, National Chung Cheng University, 160, San-Hsing, Ming-Hsiung, Chia-Yi, 621, TaiwanView Profile Authors Info & Claims ASP-DAC '00: Proceedings of the 2000 Asia and South Pacific Design Automation ConferenceJanuary 2000 Pages 225–228https://doi.org/10.1145/368434.368612Online:28 January 2000Publication History 1citation197DownloadsMetricsTotal Citations1Total Downloads197Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jinn-Shyan Wang, Po-Hui Yang
ASP-DAC1
2000 Charge sharing fault analysis and testing for CMOS domino logic circuits
abstract
Because domino logic design offers smaller area and faster delay than conventional CMOS design, it is very popular in the high-performance processor. However, domino logic suffers from several problems and one of the most notable ones is the charge sharing problem. In this paper, we describe a method to measure the sensitivity of the charge-sharing problem for each domino gate. In addition, our algorithm also generates test vectors to detect the worst case of charge-sharing fault.
Ching-Hwa Cheng, Wen-Ben Jone, Jinn-Shyan Wang, Shih-Chieh Chang 0001
Asian Test Symposium3
2000 Synthesis of CMOS Domino Circuits for Charge Sharing Alleviation
abstract
The Charge Sharing (CS) problem is one of notorious noise problems in domino circuits design and test. In this paper, this problem is thoroughly investigated by considering circuit topology and circuit function. The sensitivity of each domino gate to the CS problem is represented by the concept of CS-vulnerability. A method to derive the CS-vulnerability and the test pattern for each domino gate is suggested. We also propose a transition reordering method to dramatically reduce the CS-vulnerabilities for all domino gates, so that the CS problem can be alleviated. Simulation results demonstrate that our transistor reordering method can efficiently reduce the CS-vulnerabilities for most of domino circuits.
Ching-Hwa Cheng, Shih-Chieh Chang 0001, Shin-De Li, Wen-Ben Jone, Jinn-Shyan Wang
ICCAD5
2000 A high-speed single-phase-clocked CMOS priority encoder
abstract
The maximum operating frequency of a priority encoder is usually limited by the long propagation delay of the priority token, and the delay will increase as the number of bit of the priority encoder increases. The concept of look-ahead can be applied to improve the performance. In this paper, the design of a high-speed priority encoder is presented. The main idea of this new design is a multi-level look-ahead structure, which can be realized efficiently by the single-phase-clocked dynamic CMOS logic. A 32-bit priority encoder is implemented in a 3 V, 0.6 /spl mu/m CMOS technology to evaluate the performance of proposed techniques. The new priority encoder uses a 3-level look-ahead structure. As compared with the conventional design, the new design achieves 57% speed improvement with 5% layout area reduction.
Jinn-Shyan Wang, Chun-Shing Huang
ISCAS1
2000 A compact adaptive equalizer IC for HIPERLAN system
abstract
The design of a compact high-symbol-rate adaptive equalizer IC for the receiver of a high-speed local area network that meets the ETSI HIPERLAN standard is presented in this paper. Although the HIPERLAN defines a slowly time-varying multi-path fading-channel system, the Inter-Symbol Interference (ISI) problem is still very severe since its data rate is up to 23.5 Mbps. An Adaptive Decision Feedback Equalizer (ADFE) is selected to overcome this problem, however, the FIR filters of the ADFE require high hardware cost for complex-number computation. In this work, we select the sequential architecture to reduce the hardware cost. The penalty of adopting the sequential architecture is that several internal clocks with a much higher operating frequency (235 MHz) and the corresponding high-speed components are required. In order to generate internal clocks, we embed an All Digital Phase-Locked Loop (ADPLL) in this chip. Meanwhile, the design of high speed multipliers and adders are achieved based on the combination of PTL, CPL, and CPL-TG CMOS logic circuits. Finally, our ADFE chip is designed in a 3.3 V 0.35 /spl mu/m CMOS technology with only 3/spl times/2.8-mm/sup 2/ core area and 1 W power dissipation.
Jinn-Shyan Wang, Pei-Lung Lin, Wern-Ho Sheen, Duo Sheng
ISCAS1
2000 A 1-GHz low-power transposition memory using new pulse-clocked D flip-flops
abstract
This paper presents the design of a 1-GHz transposition memory (TRAM) that is designed in a 3.3-V 0.35-/spl mu/m CMOS technology. This high-speed TRAM is designed with the DFF-based architecture, and a new true-single-phase pulse-clocked D flip-flop (DFF) is developed to help achieve low power besides the high-speed performance. The new DFF is evolved from the true-single-phase-clocked (TSPC) split-output D latch, but the clock signal to the latch is locally processed to let the latch to behave as a DFF. The new DFF has a simpler circuit structure and less number of transistors triggered by the clock signal as compared to the previously reported high-speed semidynamic DFF (SD DFF). Therefore, when applying this new DFF to the TRAM, the power consumption of constituent DFFs and the clock driver in the TRAM can be reduced. The TRAM of this work has the same maximum operating frequency as the other TRAM designed with the SD DFFs, but 15% of the power is saved for the new design.
Po-Hui Yang, Jinn-Shyan Wang
ISCAS2
1999 Technnology Mapping for Low Power
abstract
Power consumption has become a great concern for IC and system designs. As a consequence, power-driven technology mapping has attracted several research attentions. However, the power model they used cannot properly capture the power dissipation when the output of a gate does not switch. In this paper, we propose a pattern oriented power modeling for improved technology mapping. We first perform a profitability study using the complete pattern to pattern transition data organized in tabular form. Then, we propose a probability-based, pattern oriented technology mapping method. Empirical results on benchmark circuits demonstrate the proposed method delivered an average of 13% power reduction compared to the traditional mapping method.
Chingwei Yeh, Chin-Chao Chang, Jinn-Shyan Wang
ASP-DAC3
1999 Layout Techniques Supporting the Use of Dual Supply Voltages for Cell-based Designs
abstract
Gate-level voltage scaling is an approach that allows different supply voltages for different gates in order to achieve power reduction. Previous researches focused on determining the voltage level for each gate and ascertaining the power saving capability of the approach via logic-level power estimation. In this paper, we present the layout techniques that feasiblize the approach in cell-based design environment. A new block layout style is proposed to support the voltage scaling with conventional standard cell libraries. The block layout can be automatically generated via a simulated annealing based placement algorithm. In addition, we propose a new cell layout style with built-in multiple supply rails. Using the cell layout, gate-level voltage scaling can be immediately embedded in a typical cell-based design flow. Experimental results show that proposed techniques produce very promising results.
Chingwei Yeh, Yin-Shuin Kang, Shan-Jih Shieh, Jinn-Shyan Wang
DAC4
1999 Segmented bus design for low-power systems
abstract
This paper proposes a bus-segmentation method that efficiently reduces the switched capacitance on the bus. The power consumed by the bus can, therefore, be substantially reduced. The basic idea of bus segmentation is to partition the bus into several bus segments separated by pass transistors. Highly communicating devices are located to adjacent bus segments, thus, most data communication can be achieved by switching a small portion of the bus segments. As a result, power consumption and critical path delay are both reduced. Experimental results obtained by simulating a delay model and a power model demonstrate that the proposed segmented bus system reduces bus power by about 60%-70% and improves critical bus delay by about 10%-30%.
J.-Y. Chen, Wen-Ben Jone, Jinn-Shyan Wang, Hsueh-I Lu, Tien-Fu Chen
IEEE Trans. Very Large Scale Integr. Syst.3
1998 Low-power embedded SRAM macros with current-mode read/write operations
abstract
The newly proposed SRAM performs both read and write operations in the current-mode. Due to the current-mode operations, voltage swings at bit-lines and data-lines are kept very small during read and write. The AC power dissipation of bit-lines and data-lines can thus be saved efficiently. For an embedded SRAM macro used in an 8-bit µ-controller, the SRAM using the fully current-mode technique consumes only 30% power dissipation as compared to the SRAM with only current-mode read operation. Experimental results show good agreement with the simulation results and prove the feasibility of the new technique.
Jinn-Shyan Wang, Po-Hui Yang, Wayne Tseng
ISLPED1
1995 Low-Voltage Low-Power CMOS True-Single-Phase Clocking Scheme with Locally Asynchronous Logic Circuits
Hong-Yi Huang, Jinn-Shyan Wang, Yuan-Hua Chu, Tain-Shun Wu, Kuo-Hsing Cheng, Chung-Yu Wu
ISCAS2