Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chung-Ho Chen

dblp:28/5335 · DBLP profile ↗
← Back
49ranked-venue papers
13as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 9 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Processor architecture and microarchitecture · 46% GPUs and heterogeneous computing · 28% Parallel and multicore computing · 12%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › multithreading
fine-grain multithreading
0.312018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Processor architecture and microarchitecture
multithreading
0.312018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
GPUs and heterogeneous computing › GPU microarchitecture
SIMT execution
0.312018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
GPUs and heterogeneous computing
control flow divergence
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Parallel and multicore computing
data parallelism
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
GPUs and heterogeneous computing › control flow divergence
reconvergence
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Parallel and multicore computing › data parallelism
SIMD vectorization
0.112018
Enabling SIMT Execution Model on Homogeneous Multi-Core System · ACM Trans. Archit. Code Optim. 2018
Processor architecture and microarchitecture
instruction scheduling
0.112007
Scalable Dynamic Instruction Scheduler through Wake-Up Spatial Locality · IEEE Trans. Computers 2007
Processor architecture and microarchitecture › instruction issue logic
wakeup logic
0.112007
Scalable Dynamic Instruction Scheduler through Wake-Up Spatial Locality · IEEE Trans. Computers 2007
Memory systems
cache
0.131999
Fault Containment in Cache Memories for TMR Redundant Processor Systems · IEEE Trans. Computers 1999
Microarchitecture Support for Improving the Performance of Load Target Prediction · MICRO 1997
Architecture Technique Trade-Offs Using Mean Memory Delay Time · IEEE Trans. Computers 1996
Memory systems
cache design
0.031999
Cache write generate for parallel image processing on shared memory architectures · IEEE Trans. Image Process. 1996
A Unified Architectural Tradeoff Methodology · ISCA 1994
An Easy-to-Use Approach for Practical Bus-Based System Design · IEEE Trans. Computers 1999
Processor architecture and microarchitecture › multiprocessor architecture
bus-based multiprocessor
0.011999
An Easy-to-Use Approach for Practical Bus-Based System Design · IEEE Trans. Computers 1999
Performance modeling and evaluation
queueing models
0.011999
An Easy-to-Use Approach for Practical Bus-Based System Design · IEEE Trans. Computers 1999
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.011999
An Easy-to-Use Approach for Practical Bus-Based System Design · IEEE Trans. Computers 1999
Hardware reliability and fault tolerance › redundancy › modular redundancy
triple modular redundancy
0.011999
Fault Containment in Cache Memories for TMR Redundant Processor Systems · IEEE Trans. Computers 1999
Energy-efficient computing
power management
0.012007
Scalable Dynamic Instruction Scheduler through Wake-Up Spatial Locality · IEEE Trans. Computers 2007
Energy-efficient computing › power management
processor power reduction
0.012007
Scalable Dynamic Instruction Scheduler through Wake-Up Spatial Locality · IEEE Trans. Computers 2007
Memory systems › memory access latency
load latency
0.011997
Microarchitecture Support for Improving the Performance of Load Target Prediction · MICRO 1997
Processor architecture and microarchitecture
superscalar processor
0.011997
Microarchitecture Support for Improving the Performance of Load Target Prediction · MICRO 1997
Memory systems › cache management › storage caching › caching policy
cache write policy
0.011996
Cache write generate for parallel image processing on shared memory architectures · IEEE Trans. Image Process. 1996
Performance modeling and evaluation › design trade-off analysis
performance trade-off analysis
0.011996
Architecture Technique Trade-Offs Using Mean Memory Delay Time · IEEE Trans. Computers 1996
Processor architecture and microarchitecture › pipelining
processor stall
0.011994
A Unified Architectural Tradeoff Methodology · ISCA 1994
Memory systems › cache design
cache sizing
0.011999
An Easy-to-Use Approach for Practical Bus-Based System Design · IEEE Trans. Computers 1999
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.011999
Fault Containment in Cache Memories for TMR Redundant Processor Systems · IEEE Trans. Computers 1999
Hardware reliability and fault tolerance
error recovery
0.011999
Fault Containment in Cache Memories for TMR Redundant Processor Systems · IEEE Trans. Computers 1999
Processor architecture and microarchitecture
instruction fetch
0.011997
Microarchitecture Support for Improving the Performance of Load Target Prediction · MICRO 1997
Processor architecture and microarchitecture
pipelining
0.011997
Microarchitecture Support for Improving the Performance of Load Target Prediction · MICRO 1997
Image and video processing
parallel image processing
0.011996
Cache write generate for parallel image processing on shared memory architectures · IEEE Trans. Image Process. 1996
Memory systems › memory architecture
pipelined memory architecture
0.011996
Architecture Technique Trade-Offs Using Mean Memory Delay Time · IEEE Trans. Computers 1996
Memory systems
shared memory
0.011996
Cache write generate for parallel image processing on shared memory architectures · IEEE Trans. Image Process. 1996

Methods — techniques the papers use, named apart from their topics

explicit vectorization · 0.3OpenCL · 0.3simulation-based analysis · 0.1simulation · 0.0queueing analysis · 0.0hardware voting · 0.0cache data broadcasting · 0.0load-prediction buffer · 0.0load stride prediction · 0.0bit-array design · 0.0register-level simulation · 0.0finite abelian group theory · 0.0cyclic group isomorphism · 0.0
YearPublicationVenuePosition
2025 Improve GPGPU Front-end Efficiency Via Inter-Warp Instruction Sharing
abstract
General-purpose GPUs (GPGPUs) leverage thousands of threads per core to attain high throughput. Threads executing the same program are dynamically grouped into SIMD batches known as warps, which exhibit high localities in accessing the instruction cache. Current designs allow warps to access the instruction cache independently in a time-sharing manner, leading to significant fetch redundancies. To address this issue, we introduce Inter-Warp Instruction Sharing, a lightweight yet effective technique to improve GPGPU front-end efficiency. The mechanism incorporates an improved fetch request arbitration algorithm, a fetch filter to prevent redundant fetching, and a modified instruction buffer that broadcasts instructions to all active warps. We show that our approach reduces I-Cache accesses by up to 85%, achieving a geometric mean reduction of 68% without performance degradation.
Yu-Yu Hsiao, Liang-Chou Chen, Chung-Ho Chen
ISCAS3
2020 Optimization of Stride Prefetching Mechanism and Dependent Warp Scheduling on GPGPU
abstract
In this paper, we propose a data prefetching scheme, History-Awoken Stride (HAS) prefetching, optimized with a warp scheduler, Prefetched-Then-Executed (PTE), and evaluate the performance on the platform that we developed. Our platform is a single instruction, multiple thread (SIMT) GPGPU environment, supporting OpenCL 1.2 runtime and TensorFlow framework with CUDA-on-CL technology. Enormous amount of executing threads in GPU demands critical memory performance. HAS exploits history table of related memory accesses in intra-warp and inter-warp of the same workgroup as well as among workgroups, and uses address strides and warp status to monitor the prefetching progress of the executed warp. PTE precisely issues warps according to prefetching status from HAS. The experimental results of LeNet-5 inference and 11 PolyBench test programs on CAS-GPU show that our mechanism can achieve an average IPC performance improvement of 10.4%, and 7.8% reduction in data cache miss rate. The prefetch accuracy can reach 67.7%, and the proportion of prefetch request arrived at the appropriate time reaches 48.2%.
Tsung-Han Tsou, Dun-Jie Chen, Sheng-Yang Hung, Yu-Hsiang Wang, Chung-Ho Chen
ISCAS5
2018 Kernel Aware Warp Scheduler
abstract
Observing that thread blocks of different kernels that use different functional units should be sent to the same SM (Streaming Multiprocessor) to promote the utilization of functional units, we propose a kernel-aware warp scheduler for GPGPUs. The proposed Kernel Aware Warp Scheduler uses the profiling information of the executed kernels to issue instructions from the right warp. The experimental results based on our HSAIL simulation platform show that the overall performance of the kernel execution improves by about 20% on average. The speedup comes from the increased utilization of functional units and effectiveness in hiding of the memory latency due to the proposed warp scheduling policy.
Sen-Chih Tsai, Yu-Xiang Su, Yu-Han Chin, Wei-Zhong Ceng, Chung-Ho Chen
ISCAS5
2018 Enabling SIMT Execution Model on Homogeneous Multi-Core System
abstract
Single-instruction multiple-thread (SIMT) machine emerges as a primary computing device in high-perfor-mance computing, since the SIMT execution paradigm can exploit data-level parallelism effectively. This article explores the SIMT execution potential on homogeneous multi-core processors, which generally run in multiple-instruction multiple-data (MIMD) mode when utilizing the multi-core resources. We address three architecture issues in enabling SIMT execution model on multi-core processor, including multithreading execution model, kernel thread context placement, and thread divergence. For the SIMT execution model, we propose a fine-grained multithreading mechanism on an ARM-based multi-core system. Each of the processor cores stores the kernel thread contexts in its L1 data cache for per-cycle thread-switching requirement. For divergence-intensive kernels, an Inner Conditional Statement First (ICS-First) mechanism helps early re-convergence to occur and significantly improves the performance. The experiment results show that effectiveness in data-parallel processing reduces on average 36% dynamic instructions, and boosts the SIMT executions to achieve on average 1.52× and up to 5× speedups over the MIMD counterpart for OpenCL benchmarks for single issue in-order processor cores. By using the explicit vectorization optimization on the kernels, the SIMT model gains further benefits from the SIMD extension and achieves 1.71× speedup over the MIMD approach. The SIMT model using in-order superscalar processor cores outperforms the MIMD model that uses superscalar out-of-order processor cores by 40%. The results show that, to exploit data-level parallelism, enabling the SIMT model on homogeneous multi-core processors is important.
Kuan-Chung Chen, Chung-Ho Chen
ACM Trans. Archit. Code Optim.2
2017 Processor shield for L1 data cache software-based on-line self-testing
abstract
Conventional software-based cache self-tests typically ignore system related testing issues, such as physical memory layout, virtual memory mapping, and isolating faulty effects, especially for on-line testing. We propose an architectural support for data cache software-based self-testing (SBST): Processor Shield, which can tackle difficult-to-test issues during on-line SBST. The proposed processor shield includes a software framework and design for testing (DFT) hardware, which enables SBST program to run without influencing other processes and on-bus devices even if a cache test fails. The proposed SBST process can be iteratively executed and cooperate with dynamic voltage frequency scaling (DVFS) system to calibrate the required guardbands to accommodate transistor aging effects. Finally, we present a case study that performs SBST programs under Linux kernel on an ARMv5-compatible processor system. Our method can successfully switch between the SBST process and the kernel process and achieve the expected high fault coverages for cache control logic and RAM module testing.
Ching-Wen Lin, Chung-Ho Chen
ASP-DAC2
2017 A Processor and Cache Online Self-Testing Methodology for OS-Managed Platform
abstract
Software-based self-test (SBST) is an effective method to detect operational faults of a processor system. We propose an architectural approach to support high fault-coverage online SBST: Processor Shield, which tackles the difficult-to-test issues raised due to the protection of an operating system. The processor shield, including a software framework and design for testing hardware, creates an online self-testing environment without influencing other processes and on-bus devices even if the SBST fails. We present a case study that demonstrates SBST executions under Linux kernel on an ARMv5-compatible processor system. For CPU testing, the stuck-at fault coverage is over 99% while the transition fault coverage is higher than 93%. For cache control logic testing, the stuck-at fault coverage is over 99% while the transition fault coverage is higher than 95%. For RAM module testing, the fault coverage is nearly 100%. Cache SBSTs finish in a context-switch interval of less than 4 ms while CPU SBST finishes in less than 8 ms for 1-GHz clock. The hardware overhead of the processor shield is only 0.494% of the whole processor area. We also present an SBST-dynamic voltage and frequency scaling application that calibrates the dynamic minimal guardbands and helps achieving lower power consumption and mitigating transistor-aging effect.
Ching-Wen Lin, Chung-Ho Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A testable and debuggable dual-core system with thermal-aware dynamic voltage and frequency scaling
abstract
A sophisticated SoC chip that incorporates many design modules including 2 ARM-like CPUs, a dynamic voltage and frequency scaling (DVFS) design, a master/slave temperature sensing system, and an on-chip test/debug platform is developed and implemented with TSMC 90 nm technology. Measurement results validate the functions and efficiencies of the whole chip.
Liang-Ying Lu, Ching-Yao Chang, Zhao-Hong Chen, Bo-Ting Yeh, Tai-Hua Lu, Pin-Hao Tang, Kuen-Jong Lee, Lih-Yih Chiou, Soon-Jyh Chang, Chien-Hung Tsai, Chung-Ho Chen, Jai-Ming Lin
ASP-DAC12
2016 Dynamic SIMD re-convergence with paired-path comparison
abstract
SIMD divergence is one of the critical factors that decrease the hardware utilization in contemporary GPGPUs (General Purpose Graphic Processor Unit). Both the re-convergence scheme and control flow detection have to be well considered. In the emerging HSA (Heterogeneous System Architecture) platform, we develop an effective dynamic stack-based re-convergence scheme that can be implemented without the insertion of re-convergence instructions generated by the finalizer. The stack keeps track of the minimal necessary information of the taken and non-taken paths; the additional end-of-branch instruction insertion is no longer required under our design. Using the scheme we propose, the divergent warp dynamically re-converges at opportunistic re-convergence points. The activity factor improves for 13.36% on average from opportunistic early re-convergence in the unstructured control flow. Our design has eased the development of a finalizer that no longer needs to reason about the re-convergence point after a branch divergence, especially for unstructured control flow.
Yun-Chi Huang, Kuan-Chieh Hsu, Wan-shan Hsieh, Chen-Chieh Wang, Chia-Han Lu, Chung-Ho Chen
ISCAS6
2015 A memory-efficient NoC system for OpenCL many-core platform
abstract
We present a memory-efficient NoC system for OpenCL many-core platforms. By offloading the OpenCL kernel programs into the many-core processor, we reveal that the memory contention overheads have dramatically increased with the scaling of the system, resulting to the poor performance scalability of the many-core system. We explore a memory-efficient NoC system design which includes a mesh network, a hybrid network interface for packet composition and decomposition, and a memory controller with access scheduling capability. Our experimental results show that a simple memory access scheduling and caching approach can easily boost the performance of the NoC and memory system up to 20 percent by eliminating the memory controller contention problem.
Chien-Hsuan Yen, Chung-Ho Chen, Kuan-Chung Chen
ISCAS2
2014 An OpenCL runtime system for a heterogeneous many-core virtual platform
abstract
We present a many-core full system simulation platform and its OpenCL runtime system. The OpenCL runtime system includes an on-the-fly compiler and resource manager for the ARM-based many-core platform. Using this platform, we evaluate approaches of work-item scheduling and memory management in OpenCL memory hierarchy. Our experimental results show that scheduling work-items on a many-core system using general purpose RISC CPU should avoid per work-item context switching. Data deployment and work-item coalescing are the two keys for significant speedup.
Kuan-Chung Chen, Chung-Ho Chen
ISCAS2
2014 Unambiguous I-cache testing using software-based self-testing methodology
abstract
We propose an unambiguous instruction cache software-based self-testing methodology that can generate a reliable result to precisely determine the test passed or not. We present testing cases that cause ambiguous cache testing results and propose five principles of test pattern selection to prevent these situations from occurring. To preserve the order of March sequence in testing an I-cache, we leverage cache bank and cache disable operations. In this way, we are able to implement any March algorithm without violating the sequence order. Finally, we present a case study for ARM v5 ISA processor that has an 8KB instruction cache. We use the March C- algorithm and achieve 100% of inter-word coverage and more than 97% of intra-word coverage evaluated by the RAMSES simulator.
Ching-Wen Lin, Chung-Ho Chen
ISCAS2
2014 Virtualization Technology for TCP/IP Offload Engine
abstract
Network I/O virtualization plays an important role in cloud computing. This paper addresses the system-wide virtualization issues of TCP/IP Offload Engine (TOE) and presents the architectural designs. We identify three critical factors that affect the performance of a TOE: I/O virtualization architectures, quality of service (QoS), and virtual machine monitor (VMM) scheduler. In our device emulation based TOE, the VMM manages the socket connections in the TOE directly and thus can eliminate packet copy and demultiplexing overheads as appeared in the virtualization of a layer 2 network card. To further reduce hypervisor intervention, the direct I/O access architecture provides the per VM-based physical control interface that helps removing most of the VMM interventions. The direct I/O access architecture out-performs the device emulation architecture as large as 30 percent, or achieves 80 percent of the native 10 Gbit/s TOE system. To continue serving the TOE commands for a VM, no matter the VM is idle or switched out by the VMM, we decouple the TOE I/O command dispatcher from the VMM scheduler. We found that a VMM scheduler with preemptive I/O scheduling and a programmable I/O command dispatcher with deficit weighted round robin (DWRR) policy are able to ensure service fairness and at the same time maximize the TOE utilization.
En-Hao Chang, Chen-Chieh Wang, Chien-Te Liu, Kuan-Chung Chen, Chung-Ho Chen
IEEE Trans. Cloud Comput.5
2013 Using condition flag prediction to improve the performance of out-of-order processors
abstract
If-conversion is a technique that reduces the misprediction penalties caused by conditional branches. However, executing If-converted code in out-of-order processors creates a naming problem which hinders the rename throughput. Predicting condition flag is an effective approach to resolve this problem. In this paper, we propose a scheme to predict the condition flag based on the ISA of ARM. By restoring two most recent unique condition flag values for each instruction dynamically in run time, and by using a condition flag selector when a condition flag-updating instruction reaches the renaming unit, we can predict the outcome of the condition flag-updating instruction. We show that such an approach is able to achieve the IPC performance increase of 6.62%.
Tzu-Hsuan Hsu, Ching-Wen Lin, Chung-Ho Chen
ISCAS3
2013 CASL hypervisor and its virtualization platform
abstract
In this paper, we present an ARM-based hardwareassisted hypervisor, named CASL-Hypervisor, and a full system virtualization platform developed in SystemC which enables software/hardware co-simulation of virtual machine systems. CASL-Hypervisor takes advantage of an additional processor mode, extended memory management unit, configurable hardware traps and specialized hardware devices to virtualize unmodified Linux-based guest operating systems. By utilizing hardware extensions, development effort of CASL-Hypervisor can be greatly reduced and the hypervisor has achieved relatively low virtualization overhead. Evaluation is demonstrated on an approximately-timed manner so it is able to do fast software/hardware co-simulation and evaluations. We use the ARM-v7A instruction set simulator as the host processor. The hypervisor overhead can be quantified through instruction count ratio of guest operating system to the hypervisor. The results show that CASL-Hypervisor successfully virtualizes four guest operating systems with about 9.78% overhead.
Chien-Te Liu, Kuan-Chung Chen, Chung-Ho Chen
ISCAS3
2012 Tile-based GPU optimizations through ESL full system simulation
abstract
We present a tile-based GPU design which is modeled in a full system simulation platform. The full system simulation platform includes a functional Linux-based system on which the GPU is incorporated for design explorations. To accurately estimate the execution time of the application graphics software, an execution time synchronization mechanism for the virtual platform is developed. We extend the Ericsson Texture Compression (ETC) scheme in our GPU to support alpha compression. In this way, we are able to reduce the external memory accesses to about one sixth, and speed up the rasterization engine (RE) by 35%. We also optimize the hardware-and-software data flow through the full system design platform and obtain significant improvements.
Hsu-Yao Huang, Chi-Yuan Huang, Chung-Ho Chen
ISCAS3
2012 NetVP: A system-level NETwork Virtual Platform for network accelerator development
abstract
In this paper, we propose a Network Virtual Platform (NetVP) to develop and verify network accelerator like an IPsec processor. The NetVP provides on-line verification mechanism and is suitable for ESL top-down design flow, supporting developments of un-timed as well as timed models. System development using this NetVP is efficient and flexible since it allows the designer to explore design spaces such as the network bandwidth and system architecture easily.
Chen-Chieh Wang, Sheng-Hsin Lo, Yao-Ning Liu, Chung-Ho Chen
ISCAS4
2011 Optimum profit model based on order quantity, product price, and process quality level
Chung-Ho Chen, Chih-Lun Lu
Expert Syst. Appl.1
2011 Effective Hybrid Test Program Development for Software-Based Self-Testing of Pipeline Processor Cores
abstract
This paper presents an effective hybrid test program for the software-based self-testing (SBST) of pipeline processor cores. The test program combines a deterministically developed program which explores different levels of processor core information and a block-based random program which consists of a combination of in-order instructions, random-order instructions, return instructions, as well as instruction sequences used to trigger exception/interrupt requests. Due to the complementary nature of this hybrid test program, it can achieve processor fault coverage that is comparable to the performance of the conventional scan chain method. The test response observation methods and their impacts on fault coverage are also investigated. We present the concept of micro observation versus macro observation and show that the most effective method of using SBST is through a multiple input signature register connected to the processor local bus, while conventional methods that observe only the program results in the memory lead to significantly less processor fault coverage.
Tai-Hua Lu, Chung-Ho Chen, Kuen-Jong Lee
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Energy-Efficient Trace Reuse Cache for Embedded Processors
abstract
For an embedded processor, the efficiency of instruction delivery has attracted much attention since instruction cache accesses consume a great portion of the whole processor power dissipation. In this paper, we propose a memory structure called Trace Reuse (TR) Cache to serve as an alternative source for instruction delivery. Through an effective scheme to reuse the retired instructions from the pipeline back-end of a processor, the TR cache presents improvement both in performance and power efficiency. Experimental results show that a 2048-entry TR cache is able to provide 75% energy saving for an instruction cache of 16 kB, at the same time boost the IPC up to 21%. The scalability of the TR cache is also demonstrated with the estimated area usage and energy-delay product. The results of our evaluation indicate that the TR cache outperforms the traditional filter cache under all configurations of the reduced cache sizes. The TR cache exhibits strong tolerance to the IPC degradation induced by smaller instruction caches, thus makes it an ideal design option for the cases of trading cache size for better energy and area efficiency.
Yi-Ying Tsai, Chung-Ho Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Full system simulation with QEMU: An approach to multi-view 3D GPU design
abstract
Hardware-and-software full system co-verification and co-simulation in the early stage of SoC development, i.e., before HDL code synthesis, is usually a big challenge for design engineers. In this paper, we propose a QEMU-based full system simulation framework to tackle the problem faced with the design of an embedded multi-view 3D GPU (graphic processing unit). Through the framework, we are able to extensively explore the multi-view GPU architecture, and at the same time software designers can develop and debug the associated device drivers and graphics applications by simulating the GPU design in full system operation. This approach greatly improves the ESL design process and shortens the development time for the complex multi-view GPU system.
Shye-Tzeng Shen, Shin-Ying Lee, Chung-Ho Chen
ISCAS3
2010 Optimal design of expected lifetime and warranty period for product with quality loss and inspection error
Chung-Ho Chen, Wan-Lin Chang
Expert Syst. Appl.1
2009 Full System Simulation and Verification Framework
abstract
In this paper, we propose a framework to develop high-performance system accelerator hardware and the corresponding software at system-level. This framework is designed by integrating a virtual machine, an electronic system level platform, and an enhanced QEMU-SystemC. The enhancement includes a local master interface for fast memory transfer, and an interrupt handling hardware for software/hardware communication that enables full system simulation. Finally, the PAC DSP core is used as examples to demonstrate the proposed framework for full system simulation.
Jing-Wun Lin, Chen-Chieh Wang, Chin-Yao Chang, Chung-Ho Chen, Kuen-Jong Lee, Yuan-Hua Chu, Jen-Chieh Yeh, Ying-Chuan Hsiao
IAS4
2009 The determination of optimum process mean and screening limits based on quality loss function
Chung-Ho Chen, Hui-Sung Kao
Expert Syst. Appl.1
2008 A Software-Based Test Methodology for Direct-Mapped Data Cache
abstract
We present a software-based test methodology that utilizes an on-chip processor to perform test procedures for direct-mapped data cache. The cache system under test is divided into two major groups, namely the memory modules and the logic modules. For the memory modules which include the tag memory, the data memory, and the physical address tag memory, systematic procedures to transform a widely-used March algorithm into various executable instruction sequences are developed. For the logic modules, extensive analysis on the functions as well as the structures (architecture, RTL, and gate-level) of these modules is carried out and effective test instruction sequences based on the analysis are derived. A 100% fault coverage for six conventional RAM fault models and 99.13% test efficiency for single stuck-at fault model are obtained on a real 32-bit RISC processor. These results validate the viability and effectiveness of the proposed methodology for data-cache testing.
Yi-Cheng Lin, Yi-Ying Tsai, Kuen-Jong Lee, Cheng-Wei Yen, Chung-Ho Chen
ATS5
2008 Avoiding unnecessary frame memory access and multi-frame motion estimation computation in H.264/AVC
abstract
H.264/AVC video compression standard uses multiple reference frame motion estimation (MRF-ME) to enhance the coding performance. The required frame memory accesses and the computational cost of the MRF-ME, however, are very high. This paper proposes an approach which exploits the characteristic of a constant luminance macroblock (L-MB) to avoid unnecessary frame memory accesses and MRF-ME computation. Simulation results show that the proposed scheme can reduce about 43% of the frame memory accesses and 33% of the MRF-ME computation for low motion video sequences without any PSNR degradation and bit-rate increase.
Chung-Ho Chen
ISCAS2
2008 A hybrid self-testing methodology of processor cores
abstract
Software-based self-test (SBST) is a promising new technology for at-speed testing of embedded processors in SoC systems. This paper introduces an effective and efficient new SBST methodology that uses information abstracted from the processor instruction set architecture (ISA), pipeline architecture model, RTL descriptions, and gate-level net-list for test program development of different types of the processor circuitry. This paper demonstrates the feasibility of the proposed methodology by the achieved fault coverage on a complex pipeline processor core. Comparisons with previous work are also made. Experimental results show its potential as an effective method for practical use.
Tai-Hua Lu, Chung-Ho Chen, Kuen-Jong Lee
ISCAS2
2008 Address compression for scalable load/store queue implementation
abstract
Contemporary superscalar processors employ the load/store queue for memory disambiguation. A load/store queue is typically implemented with a CAM structure to search the address for collision and consequently poses scalability challenges of energy consumption and area cost. This paper proposes an address compression technique for load/store queue to improve the power efficiency and scalability. Using the proposed approach, the LSQ can reduce the energy consumption ranging from 38% to 72% and area cost ranging from 32% to 66%, depending on the compression parameter and system configuration. The approach can provide 3.08% overall processor energy reduction and causes only 0.22% performance loss at a balanced configuration.
Yi-Ying Tsai, Chia-Jung Hsu, Chung-Ho Chen
ISCAS3
2008 Frame Buffer Access Reduction for MPEG Video Decoder
abstract
Frame buffer power consumption and bandwidth requirement are two critical design issues in MPEG video decoders due to the overwhelming amount of frame data accesses. This paper proposes a reusable macroblock detector (RMD) which exploits the characteristic of a stationary macroblock (MB) to identify the reusable MBs in the frame buffer and avoid unnecessary data transfers. The RMD on average reduces about one quarter of the frame buffer accesses for 18 video sequences.
Chung-Ho Chen
IEEE Trans. Circuits Syst. Video Technol.2
2008 Configurable VLSI Architecture for Deblocking Filter in H.264/AVC
abstract
In this paper, we study and analyze the computational complexity of the deblocking filter in H.264/AVC baseline decoder based on SimpleScalar/ARM simulator. The simulation result shows that the memory reference, content activity check operations, and filter operations are known to be very time consuming in the decoder of this new video coding standard. In order to improve overall system performance, we propose a configurable, extensible, and synthesizable window-based processing architecture which simultaneously processes the horizontal filtering of vertical edge and vertical filtering of horizontal edge. As a result, the memory performance of the proposed architecture is improved by four times when compared to previous designs. Moreover, the system performance of our window-based architecture significantly outperforms the previous designs from 7 times to 20 times.
Chung-Ming Chen, Chung-Ho Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Reduction of Frame Memory Accesses and Motion Estimation Computations in MPEG Video Encoder
abstract
This paper presents an approach that reuses data stored in the frame memory and in the motion estimation (ME) internal buffer to avoid unnecessary memory accesses and redundant ME computations for MPEG video encoders. This work employs a macroblock bitmap table, which can be easily maintained, to locate the reusable data. The experimental results show that the proposed scheme is particularly efficient in low motion video sequences, approximately saving 18% of the frame memory accesses as well as about 16% of the ME computations without any sacrifice in the image quality.
Chung-Ho Chen
ICCCN2
2007 Scalable Dynamic Instruction Scheduler through Wake-Up Spatial Locality
abstract
In a high-performance superscalar processor, the instruction scheduler often comes with poor scalability and high complexity due to the expensive wake-up operation. From detailed simulation-based analyses, we find that 95 percent of the wake-up distances between two dependent instructions are short, in the range of 16 instructions, and 99 percent are in the range of 31 instructions. We apply this wake-up spatial locality to the design of conventional CAM-based and matrix-based wakeup logic, respectively. By limiting the wake-up coverage to i + 16 instructions, where 0 les i les 15 for 16-entry segments, the proposed wake-up designs confine the wake-up operation to two matrix-based or three CAM-based 16-entry segments no matter how large the issue window size is. The experimental results show that, for an issue window of 128 entries (IW128) or 256 entries (IW256), the proposed CAM-based wake-up locality design saves 65 percent (IW128) and 76 percent (IW256) of the power consumption and reduces 44 percent (IW128) and 78 percent (IW256) in the wake-up latency compared to the conventional CAM-based design with almost no performance loss. For the matrix-based wake-up logic, applying wake-up locality to the design drastically reduces the area cost. Extensive simulation results, including comparisons with previous works, show that the wake-up spatial locality is the key element to achieving scalability for future sophisticated instruction schedulers.
Chung-Ho Chen, Kuo-Su Hsiao
IEEE Trans. Computers1
2007 Software-Based Self-Testing With Multiple-Level Abstractions for Soft Processor Cores
abstract
Software-based self-test (SBST) is a promising approach for testing a processor core embedded in a system-on-chip (SoC) system. Test routine development for SBST can be based on information of different abstraction levels. Multilevel abstraction-based SBST develops the test program for a pipeline processor using the information abstracted from its architecture model, register transfer level (RTL) descriptions, and gate-level netlist for different types of processor circuits. The proposed methodology uses gate-level and architecture information to improve coverage for structural faults. This SBST methodology uses an automatic test pattern generation tool to generate the constrained test patterns to effectively test the combinational fundamental intellectual properties used in the processor. The approach refers to the RTL code and processor architecture for the rest of the control and steering logic for test routine development. The effectiveness of this SBST methodology is demonstrated by the achieved fault coverage, test program size, and testing cycle count on a complex pipeline processor core. Comparisons with previous works are also made
Chung-Ho Chen, Chih-Kai Wei, Tai-Hua Lu, Hsun-Wei Gao
IEEE Trans. Very Large Scale Integr. Syst.1
2006 Improving Scalability and Complexity of Dynamic Scheduler through Wakeup-Based Scheduling
abstract
This paper presents a new scheduling technique to improve the speed, power, and scalability of a dynamic scheduler. In a high-performance superscalar processor, the instruction scheduler comes with poor scalability and high complexity due to the inefficient and costly instruction wakeup operation. From simulation-based analyses, we find that 98% of the wakeup activities are useless in the conventional wakeup logic. These useless activities consume a lot of power and slowdown the scheduling speed. To address this problem, the proposed technique schedules the instructions into the segmented issue window based on their wakeup addresses. During the wakeup process, the wakeup operation is only performed in the segment selected by the wakeup address of the result tag. The other segments are excluded from the wakeup operation to reduce the useless wakeup activities. The experimental results show that the proposed technique saves 50-61% of the power consumption, reduces 42-76% in the wakeup latency compared to the conventional design.
Kuo-Su Hsiao, Chung-Ho Chen
ICCD2
2006 Exploring reusable frame buffer data for MPEG-4 video decoding
abstract
This paper presents a novel approach that avoids unnecessary frame buffer accesses by exploring reusable frame buffer data for MPEG-4 simple profile video decoders. In an MPEG decoder, each decoded macroblock is stored in the frame buffer for future reference and video display. If a newly decoded macroblock overwrites the predecessor that is exactly the same, then the old data stored in the frame buffer are reusable. Frame buffer accesses of this kind become unnecessary. We propose a design that greatly reduces the number of such unnecessary memory accesses through the detection and management of the reusable data in the frame buffers. The experimental results show that the proposed scheme achieves as much as 53% (28% on the average) saving in frame buffer accesses, without any sacrifice of the image quality and with only a negligible hardware overhead.
Chung-Ho Chen
ISCAS2
2006 Design of a Giga-bit Hardware Accelerator for the iSCSI Initiator
abstract
We present the design of an iSCSI hardware accelerator for the initiator subsystem of a host bus adapter (iSCSI HBA). By analyzing the UNH-iSCSI open source code, first we evaluate the software performance and present a general methodology that transforms the software C code into the hardware HDL implementation. For the hardware module, the datapath design maximizes the concurrent accesses achievable within a clock cycle by using a dual-port descriptor memory. The synthesizable iSCSI hardware accelerator achieves 100 MHz speed and costs about 85K gates in the 0.18u technology. The design is able to meet the requirement of 1Gbps network when the average iSCSI PDU size is greater than 125 bytes
Chung-Ho Chen, Yi-Cheng Chung, Chen-Hua Wang, Han-Chiang Chen
LCN1
2006 Economic design of variable sampling intervals T2 control charts using genetic algorithms
Chao-Yu Chou, Chung-Ho Chen
Expert Syst. Appl.3
2006 Wake-Up Logic Optimizations Through Selective Match and Wakeup Range Limitation
abstract
This paper presents two effective wakeup designs that improve the speed, power, area, and scalability without instructions per cycle (IPC) loss for dynamic instruction schedulers. First, a wakeup design is proposed to aim at reducing the power consumption and wakeup latency. This design removes the read of the destination tags from the wakeup path by matching the source tags directly with the grant lines. Moreover, this design eliminates the redundant matches during the wakeup operations by matching the source tags with only the selected grant lines. Next, the second design explores a metric called wakeup locality to further reduce the area cost of the wakeup logic. By limiting the wakeup ranges for the instructions in the issue window, this design not only reduces the area requirement but also improves the scalability. The experimental results show that this range-limited-wakeup design saves 76%-94% of the power consumption and reduces 29%-77% in the wakeup latency compared to the conventional CAM-based scheme with only 5%-44% of the area cost in a traditional RAM-based scheme. The results also show that this design scales well with the increase of both the issue width and the window size
Kuo-Su Hsiao, Chung-Ho Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2003 An effective SDRAM power mode management scheme for performance and energy sensitive embedded systems
abstract
We present an effective power mode management scheme used in SDRAM memory controllers. The scheme employs a bus utilization monitoring mechanism to initiate proper operations of SDRAM chips. Our approach reduces energy consumption by actively switching memories to lowpower mode at low bus utilization. At higher bus utilization, the scheme switches memories to open page mode to reduce precharge energy as well as program execution time. This bus utilization predictor reduces memory energy consumption without the expense of increasing program execution time. It achieved the performance level of open page policy by consuming 20% less of memory energy.
Ning-Yaun Ker, Chung-Ho Chen
ASP-DAC2
2001 Parameterized MAC unit implementation
abstract
Ethernet communication devices, such as adapter, hub, bridge and switch, all follow IEEE 802.3 standard protocol. We have designed and implemented an integrated 10/100 Mbps Ethernet MAC (Medium Access Control) mechanism. The MAC unit is used to handle receive/transmit processes of network packet stream. To meet the requirement of different communication devices, we design an automatic MAC unit generator. Users can select the desired number of MAC units through parametric environment setup. To verify the application of MAC unit, we provide a 10/100 Mbps layer-2 switch simulator and an automatic test pattern generator. The FPGA demo system reveals the validity of the MAC unit.
Mingchih Chen, Ing-Jer Huang, Chung-Ho Chen
ASP-DAC3
1999 An Easy-to-Use Approach for Practical Bus-Based System Design
abstract
We present an easy-to-use model that addresses the practical issues in designing bus-based shared-memory multiprocessor systems. The model relates the shared-bus width, bus cycle time, cache memory, the features of a program execution, and the number of processors on a shared bus to a metric called request utilization. The request utilization is treated as the scaling factor for the effective average waiting processors in computing the queuing delay cycles. Simulation study shows that the model performs very well in estimating the shared bus response time. Using the model, a system designer can quickly decide the number of the processors that a shared bus is able to support effectively, the size of the cache memory a system should use, and the bus cycle time that the main memory system should provide. With the model, we show that the design favors caching the requests for a contention-based medium instead of speeding up the transfers although the same performance can be respectively achieved by the two techniques in a contention-free situation.
Chung-Ho Chen, Feng-Fu Lin
IEEE Trans. Computers1
1999 Fault Containment in Cache Memories for TMR Redundant Processor Systems
abstract
Cache data errors read by a processor may cause CPU control flow error and force the system to enter a CPU-cache reintegration process in redundant processor systems. The reintegration process degrades the system performance and reliability. To reduce the occurrences of such an event, we propose a real-time error recovery scheme that provides effective fault-containment for data errors in cache memories. The scheme is based on cache data broadcasting of a dirty line after modification. It effectively exploits the redundancy of a fault-tolerant system using hardware voting. The scheme recovers from erroneous cache data written by a processor with full coverage. This error recovery feature remedies the insufficiency of error-correcting codes that are unable to prevent such an error. In addition, more than 60 percent of cache lines are fully covered for recovery due to errors originated from the cache itself, including unrecoverable ECC errors. The protocol can also be used to speedup the CPU-cache reintegration process for a temporarily failed processor. The performance overhead of the protocol is to broadcast only 2-3 percent of the total memory references.
Chung-Ho Chen, Arun K. Somani
IEEE Trans. Computers1
1997 Microarchitecture Support for Improving the Performance of Load Target Prediction
abstract
Presents a load target prediction scheme that mitigates the impact of load latency for modern microprocessors. The scheme uses a cache-like buffer to provide the base address, offset and operand size at the instruction fetching stage of a pipeline so that a load target address can be computed earlier at the decode stage. With the dynamic use of a load stride, the scheme has achieved a prediction rate that is 15% higher than a previously proposed approach. By providing a 128-entry direct-mapped load-prediction buffer, two adders and two forwarding paths, for a 4-fetch processor the scheme provides an average speedup of 10% to 32% in performance improvement as the data cache latency increases from 2 cycles to 4 cycles. A bit-array design that supports multiple-cast writes and eliminates the associative logic commonly used in base register caching is developed for the prediction scheme.
Chung-Ho Chen, Akida Wu
MICRO1
1996 Architecture Technique Trade-Offs Using Mean Memory Delay Time
abstract
Many architecture features are available for improving the performance of a cache-based system. These hardware techniques include cache memories, processor stalling characteristics, memory cycle time, the external databus width of a processor, and pipelined memory system, etc. Each of these techniques affects the cost, design, and performance of a system. We present a powerful approach to assess the performance trade-offs of these architecture techniques based on the equivalence of mean memory delay time. For the same performance point, we demonstrate how each of these features can be traded off and report the ranking of the achievable performance of using them.
Chung-Ho Chen, Arun K. Somani
IEEE Trans. Computers1
1996 Cache write generate for parallel image processing on shared memory architectures
abstract
We investigate cache write generate, our cache mode invention. We demonstrate that for parallel image processing applications, the new mode improves main memory bandwidth, CPU efficiency, cache hits, and cache latency. We use register level simulations validated by the UW-Proteus system. Many memory, cache, and processor configurations are evaluated.
Craig M. Wittenbrink, Arun K. Somani, Chung-Ho Chen
IEEE Trans. Image Process.3
1995 Proteus: A reconfigurable computational network for computer vision
Robert M. Haralick, Arun K. Somani, Craig M. Wittenbrink, Kenneth Cooper, Linda G. Shapiro, Ihsin T. Phillips, Jenq-Neng Hwang, Yung Hsi Yao, Chung-Ho Chen, Larry Yang, Brian Daugherty, Bob Lorbeski, Kent Loving, Tom Miller, Larye Parkins, Steve Soos
Mach. Vis. Appl.11
1994 A Unified Architectural Tradeoff Methodology
abstract
Presents a unified approach to assess the tradeoff of architecture techniques that affect mean memory access time. The architectural features considered include cache hit ratio, processor stalling features, line size, memory cycle time, the external data bus width of a processor, pipelined memory system, and read bypassing write buffers. The authors demonstrate how each of these features can be traded off to achieve the desired performance. The performance of an architecture feature is quantified in terms of cache hit ratio based on the equivalence of mean memory delay time. They investigate the implication of architectural tradeoffs on the pin count, memory system design, and on-chip cache area for microprocessor systems.>
Chung-Ho Chen, Arun K. Somani
ISCA1
1992 Effects of Cache Traffic on Shared Bus Multiprocessor Systems
Chung-Ho Chen, Arun K. Somani
ICPP (1)1
1992 Proteus: a reconfigurable computational network for computer vision
abstract
The Proteus architecture is a highly parallel MIMD, multiple instruction, multiple-data machine, optimized for large granularity tasks such as machine vision and image processing. The system can achieve 20 Giga-flops (80 Giga-flops peak). It accepts data via multiple serial links at a rate of up to 640 megabytes/second. The system employs a hierarchical reconfigurable interconnection network with the highest level being a circuit switched Enhanced Hypercube serial interconnection network for internal data transfers. The system is designed to use 256 to 1024 RISC processors. The processors use one megabyte external Read/Write Allocating Caches for reduced multiprocessor contention. The system detects, locates, and replaces faulty subsystems using redundant hardware to facilitate fault tolerance.>
Robert M. Haralick, Arun K. Somani, Craig M. Wittenbrink, Kenneth Cooper, Linda G. Shapiro, Ihsin T. Phillips, Jenq-Neng Hwang, Yung Hsi Yao, Chung-Ho Chen, Larry Yang, Brian Daugherty, Bob Lorbeski, Kent Loving, Tom Miller, Larye Parkins, Steve Soos
ICPR (4)11
1982 An Algebraic Model of Arithmetic Codes
abstract
Arithmetic codes use a structured redundancy technique for binary number representation such that errors in an arithmetic operation of a digital computer can be detected or corrected. This correspondence studies the code structures by treating the set of redundant coded binary representations as a finite Abelian group. An algebraic model of arithmetic codes is developed, which shows that an arithmetic code is a pair of cyclic group isomorphisms. Two theorems are derived which describe the necessary and sufficient conditions for the existence of an arithmetic code. It is also shown that the group of redundant coded binary numbers is isomorphic to a cyclic group, or the direct sum of two cyclic groups. For a given code generator A and the information cardinality m, the two theorems may be applied to find all existing arithmetic codes up to an isomorphism. The algebraic structures of all codes published to date are covered by the mathematical model described in this correspondence.
Chung-Ho Chen
IEEE Trans. Computers1