Hung-Chi Fang

dblp:95/2354 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-authorSystems, architecture and hardware · 2Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Image and video coding · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 90% Integrated circuit design · 10%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding › image compression › wavelet-based image coding
JPEG2000
0.232007
Word-Level Parallel Architecture of JPEG 2000 Embedded Block Coding Decoder · IEEE Trans. Multim. 2007
High-Performance JPEG 2000 Encoder With Rate-Distortion Optimization · IEEE Trans. Multim. 2006
Precompression Quality-Control Algorithm for JPEG 2000 · IEEE Trans. Image Process. 2006
Image and video coding
rate-distortion optimization
0.122006
High-Performance JPEG 2000 Encoder With Rate-Distortion Optimization · IEEE Trans. Multim. 2006
Precompression Quality-Control Algorithm for JPEG 2000 · IEEE Trans. Image Process. 2006
Hardware accelerators and domain-specific architectures › video coding accelerator
video decoding accelerator
0.112007
Word-Level Parallel Architecture of JPEG 2000 Embedded Block Coding Decoder · IEEE Trans. Multim. 2007
Image and video coding › scalable coding › embedded coding
embedded block coding
0.112006
Precompression Quality-Control Algorithm for JPEG 2000 · IEEE Trans. Image Process. 2006
Image and video coding
image compression
0.112006
High-Performance JPEG 2000 Encoder With Rate-Distortion Optimization · IEEE Trans. Multim. 2006
Image and video coding
transform coding
0.112006
Precompression Quality-Control Algorithm for JPEG 2000 · IEEE Trans. Image Process. 2006
Hardware accelerators and domain-specific architectures › image processing accelerator
image compression accelerator
0.112006
High-Performance JPEG 2000 Encoder With Rate-Distortion Optimization · IEEE Trans. Multim. 2006
Integrated circuit design › low-power circuit design
low-power digital circuit design
0.012007
Word-Level Parallel Architecture of JPEG 2000 Embedded Block Coding Decoder · IEEE Trans. Multim. 2007
Image and video coding › video coding standards
MPEG-4
0.012005
Advances in Hardware Architectures for Image and Video Coding - A Survey · Proc. IEEE 2005

Methods — techniques the papers use, named apart from their topics

word-level decoding algorithm · 0.1column-switching scan order · 0.1precompression rate-distortion optimization · 0.1parallel embedded block coding · 0.1line-based discrete wavelet transform · 0.1dedicated hardware implementation · 0.1rate-distortion prediction · 0.1context modeling · 0.1
YearPublicationVenuePosition
2007 124 MSamples/s Pixel-Pipelined Motion-JPEG 2000 Codec Without Tile Memory
abstract
A 124 MSamples/s JPEG 2000 codec is implemented on a 20.1 mm2die with 0.18 mum CMOS technology dissipating 385 mW at 1.8 V and 42 MHz. This chip is capable of processing 1920times1080 HD video at 30 fps. For previous works, the tile-level pipeline scheduling is used between the discrete wavelet transform (DWT) and embedded block coding (EBC). For a tile with size 256times256, it costs 175 kB on-chip SRAM for the architectures using on-chip tile memory or costs 310 MB/s SDRAM bandwidth for the architectures using off-chip tile memory. In this design, a level-switched scheduling is developed to eliminate tile memory and the DWT and the EBC are pipelined at pixel-level. This scheduling eliminates 175 kB on-chip SRAM and 310 MB/s off-chip SDRAM bandwidth. The level-switched DWT (LS-DWT) and the code-block switched EBC (CS-EBC) are developed to enable this scheduling. The codec functions are realized on an unified hardware, and hardware sharing between encoder and decoder reduces silicon area by 40%
Chih-Chi Cheng, Chun-Chia Chen, Hung-Chi Fang, Liang-Gee Chen
IEEE Trans. Circuits Syst. Video Technol.4
2007 Word-Level Parallel Architecture of JPEG 2000 Embedded Block Coding Decoder
abstract
This paper presents a word-level decoding architecture of embedded block coding in JPEG 2000. This architecture decodes one coefficient per cycle based on the proposed word-level decoding algorithm. This algorithm eliminates state variable memories by decoding all bit-planes in parallel. The proposed column- switching scan order overcomes intra bit-plane dependency and inter bit-plane dependency to enable parallel processing. Implementation results show that the proposed architecture is capable of decoding 54 MSamples/s at 54 MHz, which can support HDTV 720p (1280 times 720, 4:2:2) decoding at 30 frames/s.
Hung-Chi Fang, Chun-Chia Chen, Chung-Jr Lian, Liang-Gee Chen
IEEE Trans. Multim.2
2006 Design and Implementation Of Word-Level Embedded Block Coding Architecture in JPEG 2000 Decoder
abstract
This paper presents a word-level decoding architecture of Embedded Block Coding (EBC) in JPEG 2000. This architecture decodes one coefficient per cycle based on the proposed word-level decoding algorithm. This algorithm eliminates state variable memories by decoding all bit-planes in parallel. The proposed column-switching scan order overcomes intra bit-plane dependency and inter bit-plane dependency to enable parallel processing. Implementation results show the proposed architecture can decode 54 MSamples/s at 54 MHz, which can support HDTV 720p (1280 × 720, 4:2:2) decoding at 30 frames/sec in real time.
Hung-Chi Fang, Chun-Chia Chen, Liang-Gee Chen
ICASSP (2)2
2006 Analysis of scalable architecture for the embedded block coding in JPEG 2000
abstract
In this paper, the scalable architecture of the embedded block coding (EBC) in JPEG 2000, the bit-plane parallel EBC, is proposed. We provided the analysis and an unified design methodology for the bit-plane parallel EBC architecture. To design the bit-plane parallel EBC, there exists three critical difficulties. To overcome the difficulties, three algorithms are proposed. By use of the proposed algorithms, the external bandwidth of the EBC is reduced by 55% averagely, and the throughput of a 4 bit-planes parallel EBC is higher than a word-level EBC with 10 bit-plans parallel by 1.5 times at the bitrate of 1 bits per pixel.
Chun-Chia Chen, Hung-Chi Fang, Liang-Gee Chen
ISCAS3
2006 Precompression Quality-Control Algorithm for JPEG 2000
abstract
In this paper, a precompression quality-control algorithm is proposed. It can greatly reduce computational power of the embedded block coding (EBC) and memory requirement to buffer bit streams. By using the propagation property and the randomness property of the EBC algorithm, rate and distortion of coding passes is approximately predicted. Thus, the truncation points are chosen before actual coding by the entropy coder. Therefore, the computational power, which is measured with the number of contexts to be processed, is greatly reduced since most of the computations are skipped. The memory requirement, which is measured with the amount required to buffer bit streams, is also reduced since the skipped contexts do not generate bit streams. Experimental results show that the proposed algorithm reduces the computational power of the EBC by 80% on average at 0.8 bpp compared with the conventional postcompression rate-distortion optimization algorithm. Moreover, the memory requirement is also reduced by 90%. The average PSNR degrades only about 0.1-0.3 dB, on average.
Hung-Chi Fang, Chih-Chi Cheng, Chun-Chia Chen, Liang-Gee Chen
IEEE Trans. Image Process.2
2006 High-Performance JPEG 2000 Encoder With Rate-Distortion Optimization
abstract
An 81 MSamples/s JPEG 2000 single-chip encoder is implemented on 5.5 mm/sup 2/ area using 0.25-/spl mu/m CMOS technology. This IC can losslessly encode HDTV 720p resolution at 30 frames/s in real time. Three techniques are adopted: line-based discrete wavelet transform, parallel embedded block coding, and precompression rate-distortion optimization. The line-based discrete wavelet transform achieves the minimum external memory access, while the internal memory is reduced by a proper memory access scheme. The parallel embedded block coding increases the throughput and reduces the memory bandwidth with similar hardware cost comparing to conventional architectures. By accurately estimating bit rates, the precompression rate-distortion optimization reduces the required computational power and processing time of the embedded block coding since the code-blocks are truncated before compression. Experimental results show that this encoder has the highest throughput with the smallest area compared with other designs in the literature.
Hung-Chi Fang, Tu-Chih Wang, Chao-Tsung Huang, Liang-Gee Chen
IEEE Trans. Multim.1
2005 PEG, MPEG-4, and H.264 Codec IP Development
abstract
The paper summarizes our design experiences of various image and video codec IPs. The design issues and methodology of custom video codecs are discussed. The design methodology can be summarized as four stages: system analysis; algorithm optimization; architecture exploration; code development. Based on these guidelines, several design cases are presented, including the proposed JPEG, MPEG-4, and H.264 architectures.
Chung-Jr Lian, Yu-Wen Huang, Hung-Chi Fang, Yung-Chi Chang, Liang-Gee Chen
DATE3
2005 Memory efficient JPEG2000 architecture with stripe pipeline scheme
abstract
The memory issue is the most critical problem for a high performance JPEG2000 architecture. The tile memory occupies more than 50% of the area in conventional JPEG2000 architectures. To solve this problem, we propose a stripe pipeline scheme. For this scheme, a level switch discrete wavelet transform (LS-DWT) and a code-block switch embedded block coding (CS-EBC) are proposed. With small additional memory, the LS-DWT and the CS-EBC can process multiple levels and code-blocks in parallel by an interleaved scheme. As a result, the overall memory requirements of the proposed architecture can be reduced to only 8.5% compared with conventional architectures.
Hung-Chi Fang, Chih-Chi Cheng, Chun-Chia Chen, Liang-Gee Chen
ICASSP (5)1
2005 Advances in Hardware Architectures for Image and Video Coding - A Survey
abstract
This paper provides a survey of state-of-the-art hardware architectures for image and video coding. Fundamental design issues are discussed with particular emphasis on efficient dedicated implementation. Hardware architectures for MPEG-4 video coding and JPEG 2000 still image coding are reviewed as design examples, and special approaches exploited to improve efficiency are identified. Further perspectives are also presented to address the challenges of hardware architecture design for advanced image and video coding in the future.
Po-Chih Tseng, Yung-Chi Chang, Yu-Wen Huang, Hung-Chi Fang, Chao-Tsung Huang, Liang-Gee Chen
Proc. IEEE4
2005 Parallel embedded block coding architecture for JPEG 2000
abstract
This paper presents a parallel architecture for the Embedded Block Coding (EBC) in JPEG 2000. The architecture is based on the proposed word-level EBC algorithm. By processing all the bit planes in parallel, the state variable memories for the context formation (CF) can be completely eliminated. The length of the FIFO (first-in first-out) between the CF and the arithmetic encoder (AE) is optimized by a reconfigurable FIFO architecture. To reduce the hardware cost of the parallel architecture, we proposed a folded AE architecture. The parallel EBC architecture can losslessly process 54 MSamples/s at 81 MHz, which can support HDTV 720p resolution at 30 frames/s.
Hung-Chi Fang, Tu-Chih Wang, Chung-Jr Lian, Liang-Gee Chen
IEEE Trans. Circuits Syst. Video Technol.1
2004 Architecture of MPEG-7 color structure description generator for realtime video applications
Jing-Kng Chang, Hung-Chi Fang, Yen-Wei Huang, Liang-Gee Chen
ICIP2
2004 Novel precompression rate-distortion optimization algorithm for JPEG 2000
abstract
In this paper, a novel pre-compression rate-distortion optimization algorithm is proposed, which can reduce computation power and memory requirement of JPEG 2000 encoder. It can reduce the wasted computational power of the entropy coder (EBCOT Tier-1) and unnecessary memory requirement for the code-stream. Distortion and rate of coding passes are calculated and estimated before coding, and therefore truncation point is selected before coding. Experimental results show that the computation time of EBCOT Tier-1 and memory requirement for the code-stream can be greatly reduced,especially at high compression ratio. The quality of the proposed algorithm is slightly lower than that of post-compression rate-distortion optimization algorithm.
Hung-Chi Fang, Chung-Jr Lian, Liang-Gee Chen
VCIP2
2003 Hardware oriented rate control algorithm and implementation for realtime video coding
abstract
In this paper, a novel rate control algorithm suitable for real-time video encoding is proposed. The proposed algorithm uses mean absolute error (MAE) results of motion estimation (ME) to achieve bit-rate control. Neither pre-analysis nor multi-pass encoding is required in our algorithm, which makes real-time hardware implementation possible. A new hardware oriented scene change detection method is also included in this rate control framework to achieve better video quality. Experiments show our rate control algorithm behaves well in all situations. The hardware architecture for this algorithm is also described. Implementation shows our proposed algorithm can be efficiently integrated into a low cost, high efficiency video encoder.
Hung-Chi Fang, Tu-Chih Wang, Liang-Gee Chen
ICASSP (2)1
2003 Performance analysis of hardware oriented algorithm modifications in H.264
abstract
H.264 is initiated by ITU-T as H.26L and will become a joint standard of ITU-T and MPEG. The coding complexity of H.264 is much higher than MPEG-4 simple profile and advance simple profile algorithms. In order to achieve real-time encoding, hardware implementation is required. The original test model of H.264 (JM) is designed to achieve high coding performance. Some algorithms of the test model require lots of operations with little coding efficiency improvement. And some algorithms create data dependencies that prevent parallel hardware accelerations. This paper presents analysis of the H.264 video coding algorithm from a hardware-oriented viewpoint. Intra prediction, Hadamard transform and motion estimation algorithms are reviewed and modified to a hardware friendly configuration. The rate distortion penalties of these modifications are simulated and shown in this paper.
Tu-Chih Wang, Yu-Wen Huang, Hung-Chi Fang, Liang-Gee Chen
ICASSP (2)3
2003 Hardware oriented rate control algorithm and implementation for realtime video coding
abstract
In this paper, a novel rate control algorithm suitable for realtime video encoding is proposed. The proposed algorithm uses mean absolute error (MAE) results of motion estimation (ME) to achieve bitrate control. Neither pre-analysis nor multi-pass encoding is required in our algorithm, which makes realtime hardware implementation possible. A new hardware oriented scene change detection method is also included in this rate control framework to achieve better video quality. Experiment shows our rate control algorithm behaves well in all situations. Hardware architecture for this algorithm is also described. Implementation shows our proposed algorithm can be efficiently integrated into low cost, high efficiency video encoder.
Hung-Chi Fang, Tu-Chih Wang, Liang-Gee Chen
ICME1
2003 Novel word-level algorithm of embedded block coding in JPEG 2000
abstract
In this work, a novel word-level algorithm of the embedded block coding in JPEG 2000 is proposed. Unlike conventional approaches, all bitplanes of coefficients are processed in parallel in the proposed algorithm. The algorithm is based on the observations and the analysis of the significant state of a coefficient and its contribution to neighbors. As a result of the analysis, there is only the coding pass of the MSB of a coefficient is the required information for the word-level processing. An algorithm for finding the coding pass of the MSB of a coefficient is also proposed. The number of scans of the entire code block is reduced to two scans and is independent of the number of bit planes.
Hung-Chi Fang, Tu-Chih Wang, Ya-Yun Shih, Liang-Gee Chen
ICME1
2003 Performance analysis of hardware oriented algorithm modification in H.264
abstract
H.264 [Committee Draft of Joint Video Specification, July 2002] is initiated by ITU-T as H.26L and will become a joint standard of ITU-T and MPEG. The coding complexity of H.264 is much higher than MPEG-4 simple profile and advance simple profile algorithms. In order to achieve real-time encoding, hardware implementation is required. The original test model of H.264 (JM) [Joint Video Team, August 2002] is designed to achieve high coding performance. Some algorithms of the test model require lots of operations with little coding efficiency improvement. And some algorithms create data dependencies that prevent parallel hardware accelerations. This paper presents analysis of H.264 video coding algorithm in a hardware-oriented viewpoint. Intra prediction, Hadamard transform and motion estimation algorithms are reviewed and modified to a hardware friendly configuration. The rate distortion penalties of these modifications are simulated and shown in this paper.
Tu-Chih Wang, Yu-Wen Huang, Hung-Chi Fang, Liang-Gee Chen
ICME3
2002 Low delay, error robust wireless video transmission architecture for video communication
abstract
In this paper, a novel video transmission architecture is proposed to meet the low-delay and error robust requirement of wireless video communications. This architecture uses FEC coding and ARQ protocol to provide efficient bandwidth access from wireless link. In order to reduce ARQ delay, a video proxy server is implemented at the base station. This video proxy not only reduces the ARQ response time but also provides error tracking functionality. We use H.263 as the experimental platform for our architecture. Experiments show that average luminance PSNR decreases only 0.35db for the "Foreman" sequence under a random error condition of 10/sup -3/ error probability.
Tu-Chih Wang, Hung-Chi Fang, Liang-Gee Chen
ICME (1)2
2002 Low-delay and error-robust wireless video transmission for video communications
abstract
Video communications over wireless networks often suffer from various errors. A novel video transmission architecture is proposed to meet the low-delay and error-robust requirement of wireless video communications. This architecture uses forward error correction coding and automatic repeat request (ARQ) protocol to provide efficient bandwidth access from wireless link. In order to reduce ARQ delay, a video proxy server is implemented at the base station. This video proxy not only reduces the ARQ response time, but also provides error-tracking functionality. The complexity of this video proxy server is analyzed. Experiment shows that about 8.9% of the total macroblocks need to be transcoded under a random-error condition of 10/sup -3/ error probability. Because H.263 is the most popular video coding standard for video communication, we use it as an experiment platform. A data-partition scheme is also used to enhance error-resilience performance. This architecture is also suitable for various motion-compensation-based standards like H.261, H.263 series, MPEG-1, MPEG-2, MPEG-4, and H.264. For "Foreman" sequence under a random-error condition of 10/sup -3/ error probability, luminance peak signal-to-noise ratio decreases only 0.35 dB, on average.
Tu-Chih Wang, Hung-Chi Fang, Liang-Gee Chen
IEEE Trans. Circuits Syst. Video Technol.2
2001 Error-Propagation Analysis and Concealment Strategy for MPEG-4 Video Bitstream with Data Partitioning
abstract
Data partitioning, one of the MPEG-4 video error resilience tools, enables better error robustness. However, it suffers from serious error propagation problem. In this paper, we use an experimental approach to model the error propagation with several commonly used error detection conditions. It is shown that errors detected in forward section of texture data may be propagated from motion data, while those in DCT coefficients mostly result from themselves. Furthermore, the motion marker is the major error source for several error conditions detected in motion part. According to these characteristics, motion marker assumption and backtracking-based concealment strategies are proposed to achieve more accurate error localization in bitstream domain. 1.
Yung-Chi Chang, Chao-Chih Huang, Hao-Chieh Chang, Hung-Chi Fang, Liang-Gee Chen
ICME4