Chao-Tsung Huang

dblp:h/ChaoTsungHuang · DBLP profile ↗
← Back
36ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0002-9173-520XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 12 first-author · 1 since 2021Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Conscious Frontier Analysis Model: Learning-Based Path Planning for PointGoal Navigation
abstract
The objective of PointGoal navigation is to guide mobile robot to their destination in the shortest distance possible without a prioir map. However, due to the environment uncertainty, finding a solution is complicated. To address this issue, we propose a conscious frontier analysis model based on a partially observable Markov decision process. Our model uses synthetic maps for training and selects frontiers (points between known and unknown areas). By leveraging a reinforcement learning framework for path planning, the model efficiently navigates without prior mapping, enhancing exploration and pathfinding capabilities to reach the PointGoal in unknown environments. We utilize a modified belief-based version of the Bellman equation to assess the cost of failing to reach the goal, enabling the selection of frontiers with the least associated cost. Additionally, we train a modified ResNet18 model to identify the cost and valuable properties of each frontier. We evaluated our model’s performance using the large-scale RGB-D indoor dataset Matterport3D on the Habitat simulator, achieving a 90.7% successful completion of the path in the environment, outperforming traditional frontier-based model by 27.9%, active neural SLAM by 14.3%, and learning augmented model by 2.8%. In certain paths, our model outperformed the Habitat simulator’s state-of-the-art baseline normalized path length by 3.6%. Also, we reckon the efficacy of the model on the real robot, and it proved to be effective. This model supports 3D Lidar, 2D Lidar, and RGB-D perception, enabling exploration in complex environments before pinpointing the goal frontier, while synthetic maps reduce sampling biases in new environments.
Shri Harish Manoharan, Wei-Yu Chiu, Chao-Tsung Huang
IEEE Trans Autom. Sci. Eng.3
2025 Temporal Fusion: Continuous-Time Light Field Video Factorization
abstract
A factored display emits full-parallax dense-view light fields for a glasses-free 3D experience without sacrificing the spatial resolution of a liquid-crystal display (LCD). For static light fields, it achieves high-quality reconstruction by applying frame-based low-rank factorization to time-multiplexed sub-frame contents of stacked LCDs. However, for light field videos such frame-based factorization could introduce reconstruction artifacts and visual flickers and further cause human discomfort. The artifacts mainly come from incomplete constraints for the emitted light fields that are actually perceived in continuous time, instead of discrete frames. In particular, the perceived light fields are related to the persistence-of-vision (POV) effect of human eyes and the refresh rates of LCD displays, which is not well explored in previous work. In this work, we introduce a light-field video factorization framework-temporal fusion (TF)-to resolve these issues. To begin with, we explicitly formulate the continuous-time POV effect into a global factorization objective functional to eliminate visual flickers and enhance image quality. We further show that this optimization problem can be solved by sequence-level iterative updates on LCD sub-frames. Then, to tackle the enormous requirement of memory access for the sequence-level processing flow, we devise an efficient cuboid-wise factorization algorithm which enables practical GPU implementation. We also devise another lightweight causal framework, TF-C, for supporting low-latency applications. Finally, extensive experiments are performed to verify the effectiveness. Compared to the plain frame-based factorization, TF/TF-C can improve temporal consistency by reducing flicker values by 85%/91% and enhance reconstruction quality by increasing PSNR values by 5.0dB/3.7dB. In addition, we present a prototype dual-layer factored display, which was built with two 240-Hz high-refresh-rate LCDs, to demonstrate the visual quality for real-life applications.
Li-De Chen, Li-Qun Weng, Hao-Chien Cheng, An-Yu Cheng, Chao-Tsung Huang
IEEE Trans. Image Process.5
2025 Falcon: A Fused-Layer Accelerator With Layer-Wise Hybrid Inference Flow for Computational Imaging CNNs
abstract
Computational imaging (CI) has advanced significantly due to the use of convolutional neural networks (CNNs). Its edge deployment relies on layer fusion to offload the monstrous external memory access (EMA) of feature maps, necessitating the handling of overlapped features either through reusing or recomputing them. Depending on how the boundary-handling strategy is organized, the induced computing complexity and EMA can be optimized. However, state-of-the-art CI accelerators primarily apply homogeneous inference flows, which employ a single overlap-handling strategy throughout the fused layers, limiting their ability to balance computation and data access. In this article, we explore layer-wise optimization in fused-layer CNNs by exploiting hybrid-strategy inference flows and devising a corresponding computing architecture. We categorize layer-wise strategies and put forward a layer-wise hybrid inference flow (LHIF) to integrate their advantages, and we propose an optimization procedure that explicitly analyzes essential figures of merit (FoMs), including throughput, EMA, and energy efficiency. Furthermore, we develop a high-throughput accelerator—Falcon—to efficiently support LHIF under massive parallelism, especially with a time-division-multiplexing (TDM) buffer interface that enables seamless access to feature maps stored in an interleaved manner. Layout results show that the accelerator, delivering 41 TOPS with 1.5 MB of feature-map buffers, supports LHIF while increasing the die area by only 1.4% and power consumption by only 0.7%. Extensive simulations are conducted to demonstrate the versatility of LHIF in working scenarios at operational, design, and system levels. Compared with using homogeneous inference flows, the proposed LHIF achieves Pareto optimality with up to$2.28\times $higher throughput and$3.5\times $lower EMA.
Yong-Tai Chen, Yen-Ting Chiu, Hao-Jiun Tu, Chao-Tsung Huang
IEEE Trans. Very Large Scale Integr. Syst.4
2024 VLSI Design of Light-Field Factorization for Dual-Layer Factored Display
abstract
This article introduces a VLSI design for light-field factorization, aimed at enhancing immersive 3-D visual experiences for computational light-field factored displays. The main design challenges are intensive memory-access demands and high computational complexity. Accordingly, we first propose half-block-based factorization (HBBF) and sparse ray sampling (SRS) to reduce DRAM bandwidth by 99% and SRAM size by 74%. Then, we devise integer hybrid quantization (INTH) to cut down computational logic by 41%, leading to improvements in die area and power efficiency. Finally, we fabricated a processor chip that incorporates 75.1 kB of SRAM and 5.9M logic gates using 40-nm CMOS technology. It can operate with three different performance modes: high quality (56.9 MPixel/s at 971 mW), balanced (62.5 MPixel/s at 442 mW), and low power (61.7 MPixel/s at 283 mW). Across these modes, its normalized energy ranges between 4.4 and 16.2 nJ/pixel. This implementation surpasses existing GPU platforms and offers an$85\times $increase in processing speed and a$311\times $reduction in power consumption. We also showcase a real-time computational 3-D display system with this chip, demonstrating its practical efficacy in computational 3-D display technology.
Li-De Chen, Li-Qun Weng, Hao-Chien Cheng, An-Yu Cheng, Kai-Ping Lin, Chao-Tsung Huang
IEEE Trans. Very Large Scale Integr. Syst.6
2022 Chaos LiDAR Based RGB-D Face Classification System With Embedded CNN Accelerator on FPGAs
abstract
Face classification is important in many applications such as surveillance, border control, and security systems. However, wide variations in environments such as insufficient light, large distances or pose angles make the task challenging. Depth sensors are added with RGB cameras for improving classification accuracy but commercial RGB-D sensors are most targeted for indoors applications. In this paper, we present and design a Chaos LiDAR depth sesnor that provides high-precision depth images through intelligent correlation processing for both indoors and outdoors applications. Our Chaos LiDAR depth sensor detects range from 2 to 40 meters with precision around 8mm at 20-meter. With the Chaos LiDAR depth as input, we design a RGB-D based face classification embedded CNN (eCNN) model for wide range applications such as dim illumination, various distances and large poses. Our Chaos LiDAR increases around 14.27% classification accuracy compared to RealSense D435i for distance from 3 to 5 meter. The eCNN face classification subsystem is implemented in Xilinx ZCU 102 and achieves 11.11 ms inference time. The eCNN engine achieves a peak throughput at 614.4 GOPS. The overall system including Chaos LiDAR, correlation and eCNN FPGA achieves face classification inference rate of 10fps.
Ching-Te Chiu, Yu-Chun Ding, Wei-Jyun Chen, Shu-Yun Wu, Chao-Tsung Huang, Chun-Yeh Lin, Chia-Yu Chang, Meng-Jui Lee, Shimazu Tatsunori, Tsung Chen, Fan-Yi Lin, Yuan-Hao Huang
IEEE Trans. Circuits Syst. I Regul. Pap.6
2021 RingCNN: Exploiting Algebraically-Sparse Ring Tensors for Energy-Efficient CNN-Based Computational Imaging
abstract
In the era of artificial intelligence, convolutional neural networks (CNNs) are emerging as a powerful technique for computational imaging. They have shown superior quality for reconstructing fine textures from badly-distorted images and have potential to bring next-generation cameras and displays to our daily life. However, CNNs demand intensive computing power for generating high-resolution videos and defy conventional sparsity techniques when rendering dense details. Therefore, finding new possibilities in regular sparsity is crucial to enable large-scale deployment of CNN-based computational imaging.In this paper, we consider a fundamental but yet well-explored approach—algebraic sparsity—for energy-efficient CNN acceleration. We propose to build CNN models based on ring algebra that defines multiplication, addition, and non-linearity for n-tuples properly. Then the essential sparsity will immediately follow, e.g. n-times reduction for the number of real-valued weights. We define and unify several variants of ring algebras into a modeling framework, RingCNN, and make comparisons in terms of image quality and hardware complexity. On top of that, we further devise a novel ring algebra which minimizes complexity with component-wise product and achieves the best quality using directional ReLU. Finally, we design an accelerator, eRingCNN, to accommodate to the proposed ring algebra, in particular with regular ring-convolution arrays for efficient inference and on-the-fly directional ReLU blocks for fixed-point computation. We implement two configurations, n = 2 and 4 (50% and 75% sparsity), with 40 nm technology to support advanced denoising and super-resolution at up to 4K UHD 30 fps. Layout results show that they can deliver equivalent 41 TOPS using 3.76 W and 2.22 W, respectively. Compared to the real-valued counterpart, our ring convolution engines for n = 2 achieve 2.00× energy efficiency and 2.08× area efficiency with similar or even better image quality. With n = 4, the efficiency gains of energy and area are further increased to 3.84× and 3.77× with only 0.11 dB drop of peak signal-to-noise ratio (PSNR). The results show that RingCNN exhibits great architectural advantages for providing near-maximum hardware efficiencies and graceful quality degradation simultaneously.
Chao-Tsung Huang
ISCA1
2021 Real-Time Block-Based Embedded CNN for Gesture Classification on an FPGA
abstract
This paper presents a block-based embedded convolutional neural network (CNN) for gesture classification on field-programmable gate array (FPGA) in real time. Gesture recognition is an important tool to spontaneous interact with human machine interface. Many CNN architectures using RGB images have been proposed for gesture classification. RGB based gesture classification may cause incorrect results under insufficient light or similar gestures. In addition, most of the CNN architectures cannot run in real time on edge devices due to their large number of parameters and DRAM data access. In this paper, a block-based CNN using RGB-D data is proposed for gesture classification. Adding depth images to RGB images boots the classification accuracy. A CNN architecture with block-based feature maps is built for embedded FPGA implementations. The total number of parameters of the proposed RGB-D embedded CNN (eCNN) model is only 0.17M and it achieves 99.96% and 99.88% accuracy with 32-bit floating point and 8-bit fixed point implementation for America Sign Language (ASL) data set. The RTL simulation of the proposed eCNN model has the average inference speed of 0.171 milliseconds at frequency of 250MHz for a single pair RGB-D image. Implemented on a FPGA integrated with Microsoft Kinect v2 achieve an inference time in 19.42 ms which achieves high accuracy and real-time performance.
Ching-Chen Wang, Yu-Chun Ding, Ching-Te Chiu, Chao-Tsung Huang, Yen-Yu Cheng, Shih-Yi Sun, Chih-Han Cheng, Hsueh-Kai Kuo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2020 Ernet Family: Hardware-Oriented Cnn Models For Computational Imaging Using Block-Based Inference
abstract
Convolutional neural networks (CNNs) demand huge DRAM bandwidth for computational imaging tasks, and block-based processing has recently been applied to greatly reduce the bandwidth. However, the induced additional computation for feature recomputing or the large SRAM for feature reusing will degrade the performance or even forbid the usage of state-of-the-art models. In this paper, we address these issues by considering the overheads and hardware constraints in advance when constructing CNNs. We investigate a novel model family-ERNet-which includes temporary layer expansion as another means for increasing model capacity. We analyze three ERNet variants in terms of hardware requirement and introduce a hardware-aware model optimization procedure. Evaluations on Full HD and 4K UHD applications will be given to show the effectiveness in terms of image quality, pixel throughput, and SRAM usage. The results also show that, for block-based inference, ERNet can outperform the state-of-the-art FFDNet and EDSR-baseline models for image denoising and super-resolution respectively.
Chao-Tsung Huang
ICASSP1
2020 FIR Filter Design and Implementation for Phase-Based Processing
abstract
Complex steerable pyramid (CSP) is widely used to decompose images into muti-scale and oriented subbands for phase-based processing, such as video magnification, frame interpolation, and view synthesis. The conventional implementation is based on frequency-domain bandpass filtering which relies on fast Fourier transform (FFT). However, FFT requires high-precision computation and complex memory access for hardware implementation. In this paper, we study computation- and memory-efficient finite impulse response filter implementation for CSP. We use Kaiser windowing for filter designs and adopt 9-tap radial and 11-tap angular filters which can achieve 38.6 dB of PSNR for frame interpolation. We then discuss about VLSI architecture designs and, in particular, propose a stripe-based computation flow for 2-D CSP to reduce the line buffer size down to 15%. For evaluation, we implement two VLSI circuits using TSMC 40nm technology. One is a 1-D CSP engine working at 4K UHD 30 fps, and it saves 67.8% of logic gates compared to a FFT-based design. The other is a 2-D CSP engine working at FHD 60 fps. It uses 32-KB SRAM and 3.5M-gate logic. We also implement a FPGA system for 2-D CSP engine, and it operates at 80 MHz and provides 1024×1024 resolution video at 16 fps.
Shih-Yao Huang, Wei-Chih Chen, Chao-Tsung Huang
ICASSP3
2020 Fast and Accurate Embedded DCNN for Rgb-D Based Sign Language Recognition
abstract
In this paper, fast and accurate two paths CNN architecture was designed in hardware-oriented manner. Our proposed network is composed of RGB and depth path for gesture recognition by fusing RGB and depth features, following the pre-defined constraints on dedicated hardware. The RTL simulation results indicate it only takes 0.171 milliseconds to infer a single pair of RGB image and depth maps at the operational frequency of 250MHz. Compared with running the same model at Intel i7 and GTX 1080, the speedups are 593.92x and 7.68x respectively. Besides, to increase the recognition accuracy under the diversity of the circumstance, a new RGB-D dataset, captured from Kinect, with complex background was built. Moreover, the number of parameters in our model is only 0.17M and it achieves 99.79% accuracy on the ASL Finger Spelling dataset. Compared with the Gao's CNN gesture recognition architecture, the number of parameter of our model is 2.9 times less and the accuracy is 6.49% higher. Demonstration video for sign language recognition is provided : https://youtu.be/DvO8mI7IZ5Q.
Ching-Chen Wang, Ching-Te Chiu, Chao-Tsung Huang, Yu-Chun Ding, Li-Wei Wang 0013
ICASSP3
2019 System and VLSI Implementation of Phase-based View Synthesis
abstract
View synthesis is one of the important techniques utilized in 3D TV devices. Traditional methods such as depth image-based rendering usually rely on accurate depth maps which require computation-intensive stereo matching. In this work, we present a hardware system of phase-based view synthesis that is able to convert stereoscopic videos to multi-view content with low-resolution depth maps. When compared to the view synthesis reference software, the phase-based method does not suffer from severe arifacts on object boundaries and provides higher quality views in our experimental result. There are two major contributions in our implementation. First, we propose a cross-band disparity correction scheme that not only enables the usage of low-resolution disparity maps but also improves the quality of novel views. Second, we propose a hardware-friendly wavelet re-projection engine to reduce the hardware complexity. We implemented a VLSI circuit for 8-view 4K Ultra-HD (UHD) 3DTV in TSMC 40nm technology. It delivers 30 frames per second (fps) for UHD display when operating at 200MHz. It uses 228-KB SRAM and 2M-gate logic . We also implemented the system on FPGA, and it can provide 4K UHD multi-view content at 12 fps.
Han-Chih Huang, Yu-Chih Wang, Wei-Chih Chen, Ping-Yen Lin, Chao-Tsung Huang
ICASSP5
2019 eCNN: A Block-Based and Highly-Parallel CNN Accelerator for Edge Inference
abstract
Convolutional neural networks (CNNs) have recently demonstrated superior quality for computational imaging applications. Therefore, they have great potential to revolutionize the image pipelines on cameras and displays. However, it is difficult for conventional CNN accelerators to support ultra-high-resolution videos at the edge due to their considerable DRAM bandwidth and power consumption. Therefore, finding a further memory- and computation-efficient microarchitecture is crucial to speed up this coming revolution.
Chao-Tsung Huang, Yu-Chun Ding, Huan-Ching Wang, Chi-Wen Weng, Kai-Ping Lin, Li-Wei Wang 0013, Li-De Chen
MICRO1
2019 Empirical Bayesian Light-Field Stereo Matching by Robust Pseudo Random Field Modeling
abstract
Light-field stereo matching problems are commonly modeled by Markov Random Fields (MRFs) for statistical inference of depth maps. Nevertheless, most previous approaches did not adapt to image statistics but instead adopted fixed model parameters. They explored explicit vision cues, such as depth consistency and occlusion, to provide local adaptability and enhance depth quality. However, such additional assumptions could end up confining their applicability, e.g. algorithms designed for dense view sampling are not suitable for sparse one. In this paper, we get back to MRF fundamentals and develop an empirical Bayesian framework-Robust Pseudo Random Field-to explore intrinsic statistical cues for broad applicability. Based on pseudo-likelihoods with hidden soft-decision priors, we apply soft expectation-maximization (EM) for good model fitting and perform hard EM for robust depth estimation. We introduce novel pixel difference models to enable such adaptability and robustness simultaneously. Accordingly, we devise a stereo matching algorithm to employ this framework on dense, sparse, and even denoised light fields. It can be applied to both true-color and grey-scale pixels. Experimental results show that it estimates scene-dependent parameters robustly and converges quickly. In terms of depth accuracy and computation speed, it also outperforms state-of-the-art algorithms constantly.
Chao-Tsung Huang
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 A 320M Pixel/S Vlsi Architecture Design of Weighted Mode Filter for 4K Ultra-Hd Depth Upsampling
abstract
High-quality and high-resolution depth maps have opened tremendous possibilities for various applications, such as ARlVR display, 3D reconstruction, image refocusing, and view synthesis. But high-resolution depth estimation requires heavy hardware resources. Depth upsampling with weighted mode filtering is an efficient way to overcome this challenge. However, its hardware implementation has two major design issues: large on-chip memory for storing high-precision depth labels and high logic cost for computing adaptive range weight. In this work, we present two techniques, histogram candidate mapping and binary range weight kernel, which can reduce on-chip memory size and logic gate count by 46.9 % and 64.3 % respectively. Furthermore, we also implement a VLSI circuit for 4K Ultra-HD depth video upsampling using TSMC 40nm technology. It has 25.5-KB SRAM and 420K-gate logic, and the core area is 1.1 ×1.1 mm2. When operating at 200 MHz and 0.9V, it delivers 320M pixel/s to support 4K Ultra-HD depth video at 40 fps, and consumes 104 mW based on post-layout simulation.
Bo-Hsiang Yang, Li-De Chen, Chao-Tsung Huang
ICASSP3
2017 Robust Pseudo Random Fields for Light-Field Stereo Matching
abstract
Markov Random Fields are widely used to model lightfield stereo matching problems. However, most previous approaches used fixed parameters and did not adapt to lightfield statistics. Instead, they explored explicit vision cues to provide local adaptability and thus enhanced depth quality. But such additional assumptions could end up confining their applicability, e.g. algorithms designed for dense light fields are not suitable for sparse ones.,,In this paper, we develop an empirical Bayesian framework-Robust Pseudo Random Field-to explore intrinsic statistical cues for broad applicability. Based on pseudo-likelihood, it applies soft expectation-maximization (EM) for good model fitting and hard EM for robust depth estimation. We introduce novel pixel difference models to enable such adaptability and robustness simultaneously. We also devise an algorithm to employ this framework on dense, sparse, and even denoised light fields. Experimental results show that it estimates scene-dependent parameters robustly and converges quickly. In terms of depth accuracy and computation speed, it also outperforms state-of-the-art algorithms constantly.
Chao-Tsung Huang
ICCV1
2017 Fast Physically Correct Refocusing for Sparse Light Fields Using Block-Based Multi-Rate View Interpolation
abstract
Digital refocusing has a tradeoff between complexity and quality when using sparsely sampled light fields for low-storage applications. In this paper, we propose a fast physically correct refocusing algorithm to address this issue in a twofold way. First, view interpolation is adopted to provide photorealistic quality at infocus-defocus hybrid boundaries. Regarding its conventional high complexity, we devised a fast line-scan method specifically for refocusing, and its 1D kernel can be 30× faster than the benchmark View Synthesis Reference Software (VSRS)-1D-Fast. Second, we propose a block-based multi-rate processing flow for accelerating purely infocused or defocused regions, and a further 3- 34× speedup can be achieved for high-resolution images. All candidate blocks of variable sizes can interpolate different numbers of rendered views and perform refocusing in different subsampled layers. To avoid visible aliasing and block artifacts, we determine these parameters and the simulated aperture filter through a localized filter response analysis using defocus blur statistics. The final quadtree block partitions are then optimized in terms of computation time. Extensive experimental results are provided to show superior refocusing quality and fast computation speed. In particular, the run time is comparable with the conventional single-image blurring, which causes serious boundary artifacts.
Chao-Tsung Huang, Yu-Wen Wang, Liren Huang, Jui Chin, Liang-Gee Chen
IEEE Trans. Image Process.1
2016 VLSI architecture design of weighted mode filter for Full-HD depth map upsampling at 30fps
abstract
High-resolution depth maps are necessary for advanced computer vision applications but difficult to generate on portable devices. In this paper, we aim to provide a realtime depth upsampling engine using weighted mode filtering to alleviate the hardware requirement in such scenarios. An appropriate filter window size is essential for both performance and complexity, and extensive experiments are conducted to choose an 8×8 window. The design bottlenecks are then mainly twofold: high SRAM bandwidth due to large-window source data access and complex histogram processing for a high depth-label count up to 128. These issues were addressed by two proposed techniques accordingly: source spreading and one-the-fly maximum finding. Based on the synthesis results using TSMC 40nm technology, the proposed architecture can provide Full-HD depth upsampling at 43 fps with 247k logic gates and 5.4 kbytes of SRAM.
Li-De Chen, Yu-Ling Hsiao, Chao-Tsung Huang
ISCAS3
2016 Fast realistic block-based refocusing for sparse light fields
abstract
View-interpolation-based refocusing achieves realistic quality for sparse light fields but requires lots of computation. In this paper, we aim to reduce the computation load while maintaining the superior refocusing quality. The idea is to interpolate only few views for infocused regions and to perform refocusing on downsampled pixels for defocused area. This is achieved b y a proposed block-based refocusing algorithm which consists of multi-level refocusing mode decision and timing-based bloc k merging. For each variable-size block, the former chooses the fastest downsampling mode given that the distortion is negligible based on a localized filter response analysis. Then, the latter determines the fastest quadtree block partition by minimizing the timing cost which is accurately estimated for each block. Experimental results show that a 10x speedup on average is achieved for realistic refocusing. Also, the computation time becomes comparable to the conventional fast depth-dependent blurring which has serious boundary artifacts.
Liren Huang, Yu-Wen Wang, Chao-Tsung Huang
ISCAS3
2016 Fast Distribution Fitting for Parameter Estimation of Range-Weighted Neighborhood Filters
abstract
The range variance of neighborhood filters is well estimated via distribution fitting of a chi scale mixtures model proposed in our previous work. However, it introduced computation overheads for deriving empirical distributions and performing iterative fitting. In this letter, we discuss how to greatly reduce the overheads for practical usage while maintaining denoising quality. For empirical distributions, a grid-subsampling strategy is adopted for acceleration. Regarding distribution fitting, two different methods are studied: equal-frequency merged distribution and L-moment fitting. The former reformulates the fitting process into entropy optimization for only few merged bins. It provides 6-13x speedup for model fitting with negligible quality loss and 9-20x speedup with ≤ 0.1 dB PSNR drop by using 20 and 10 bins respectively. The latter performs table lookup of L-moments, instead of conventional moments, for robust fitting of heavy-tailed distributions. The fitting time then becomes negligible with ≤ 0.2 dB drop in most cases, e.g. the overall run time for bilateral 9 × 9 filtering can be thus accelerated by around 6x. Experiments on bilateral and non-local means filters are also given to show the speedup, quality and robustness.
Chao-Tsung Huang
IEEE Signal Process. Lett.1
2015 Bayesian inference for neighborhood filters with application in denoising
abstract
Range-weighted neighborhood filters are useful and popular for their edge-preserving property and simplicity, but they are originally proposed as intuitive tools. Previous works needed to connect them to other tools or models for indirect property reasoning or parameter estimation. In this paper, we introduce a unified empirical Bayesian framework to do both directly. A neighborhood noise model is proposed to reason and infer the Yaroslavsky, bilateral, and modified non-local means filters. An EM+ algorithm is devised to estimate the essential parameter, range variance, via the model fitting to empirical distributions. Finally, we apply this framework to color-image denoising. Experimental results show that the proposed model fits noisy images well and the range variance is estimated successfully. The image quality can also be improved by a proposed recursive fitting and filtering scheme.
Chao-Tsung Huang
CVPR1
2015 Fast realistic refocusing for sparse light fields
abstract
Digital refocusing for sparsely sampled light fields results in aliasing effect. For realistic quality, previous works performed anti-aliasing by applying time-consuming view interpolation for a heuristic number of novel views. In this paper, we study this problem by first performing a spectral analysis to give an analytical rule for the novel view number, which saves 34% of views compared to the intuitional choice. Then we propose a fast refocusing algorithm using a view interpolation method which is about 30x faster than VSRS-1D-Fast. Experimental results show the effectiveness of our approach in terms of speed and quality. For showing realistic refocusing, a light-field wafer-level-optics array is implemented, and its refocused images are compared to pictures captured by a real camera.
Chao-Tsung Huang, Jui Chin, Hong-Hui Chen, Yu-Wen Wang, Liang-Gee Chen
ICASSP1
2015 Bayesian Inference for Neighborhood Filters With Application in Denoising
abstract
Range-weighted neighborhood filters are useful and popular for their edge-preserving property and simplicity, but they are originally proposed as intuitive tools. Previous works needed to connect them to other tools or models for indirect property reasoning or parameter estimation. In this paper, we introduce a unified empirical Bayesian framework to do both directly. A neighborhood noise model is proposed to reason and infer the Yaroslavsky, bilateral, and modified non-local means filters by joint maximum a posteriori and maximum likelihood estimation. Then, the essential parameter, range variance, can be estimated via model fitting to the empirical distribution of an observable chi scale mixture variable. An algorithm based on expectation-maximization and quasi-Newton optimization is devised to perform the model fitting efficiently. Finally, we apply this framework to the problem of color-image denoising. A recursive fitting and filtering scheme is proposed to improve the image quality. Extensive experiments are performed for a variety of configurations, including different kernel functions, filter types and support sizes, color channel numbers, and noise types. The results show that the proposed framework can fit noisy images well and the range variance can be estimated successfully and efficiently.
Chao-Tsung Huang
IEEE Trans. Image Process.1
2014 Energy and area-efficient hardware implementation of HEVC inverse transform and dequantization
abstract
High Efficiency Video Coding (HEVC) inverse transform for residual coding uses 2-D 4×4 to 32×32 transforms with higher precision as compared to H.264/AVC's 4×4 and 8×8 transforms resulting in an increased hardware complexity. In this paper, an energy and area-efficient VLSI architecture of an HEVC-compliant inverse transform and dequantization engine is presented. We implement a pipelining scheme to process all transform sizes at a minimum throughput of 2 pixel/cycle with zero-column skipping for improved throughput. We use data-gating in the 1-D Inverse Discrete Cosine Transform engine to improve energy-efficiency for smaller transform sizes. A high-density SRAM-based transpose memory is used for an area-efficient design. This design supports decoding of 4K Ultra-HD (3840×2160) video at 30 frame/sec. The inverse transform engine takes 98.1 kgate logic, 16.4 kbit SRAM and 10.82 pJ/pixel while the dequantization engine takes 27.7 kgate logic, 8.2 kbit SRAM and 1.10 pJ/pixel in 40 nm CMOS technology. Although larger transforms require more computation per coefficient, they typically contain a smaller proportion of non-zero coefficients. Due to this trade-off, larger transforms can be more energy-efficient.
Mehul Tikekar, Chao-Tsung Huang, Vivienne Sze, Anantha P. Chandrakasan
ICIP2
2014 Memory-Hierarchical and Mode-Adaptive HEVC Intra Prediction Architecture for Quad Full HD Video Decoding
abstract
This paper presents a high-throughput and areaefficient VLSI architecture for intra prediction in the emerging high efficiency video coding standard. Three design techniques are proposed to address the complexity systematically: 1) a hierarchical memory deployment that stores neighboring samples in 4.9 Kb of static RAM (SRAM) instead of 43.2-k gates of registers and increases throughput by processing reference samples in registers; 2) a mode-adaptive scheduling scheme for all prediction units, which provides at least 2 samples/cycle throughput while using low-throughput SRAM and can achieve 2.46 samples/cycle on the average based on the experimental results; and 3) resource sharing for multipliers and the readout circuits of reference sample registers, which can save 2.5-k gates. These techniques can efficiently reduce area by 40% but induce more power because of additional signal transitions. Signal-gating circuits are then applied to reduce 69% of SRAM power and 32% of logic power, which cost only 1.0-k gates. When synthesized at 200 MHz with 40-nm process, the proposed architecture needs only 27.0-k gates and 4.9 Kb of single-port SRAM. The layout core area is 0.036 mm2, and the power consumption is 2.11 mW in the postlayout simulation. The corresponding performance can support quad full high-definition (HD) (3840 × 2160) video decoding at 30 frames/s.
Chao-Tsung Huang, Mehul Tikekar, Anantha P. Chandrakasan
IEEE Trans. Very Large Scale Integr. Syst.1
2013 HEVC interpolation filter architecture for quad full HD decoding
abstract
In this paper, an area-efficient and high-throughput interpolation filter architecture is presented for the latest video coding standard, High Efficiency Video Coding. A unified filter design is first proposed for the 8-tap luma and 4-tap chroma filters to optimize area, which uses only 13 adders. And a 2D filter architecture is then devised with an adaptive scheduling which supports all symmetric prediction partitions with a throughput of at least two samples/cycle. Experimental results also show that this architecture can achieve 2.58 samples/cycle on the average. The total gate count is 45.2k when synthesized at 200MHz with 40nm process, and the corresponding performance can support at least 3840×2160 videos at 30 fps.
Chao-Tsung Huang, Chiraag Juvekar, Mehul Tikekar, Anantha P. Chandrakasan
VCIP1
2007 On-Chip Memory Optimization Scheme for VLSI Implementation of Line-Based Two-Dimentional Discrete Wavelet Transform
abstract
The on-chip line buffer dominates the total area and power of line-based 2-D discrete wavelet transform (DWT). In this paper, a memory-efficient VLSI implementation scheme for line-based 2-D DWT is proposed, which consists of two parts, the wordlength analysis methodology and the multiple-lifting scheme. The required wordlength of on-chip memory is determined firstly by use of the proposed wordlength analysis methodology, and a memory-efficient VLSI implementation scheme for line-based 2-D DWT, named multiple-lifting scheme, is then proposed. The proposed wordlength analysis methodology can guarantee to avoid overflow of coefficients, and the average difference between predicted and experimental quality level is only 0.1 dB in terms of PSNR. The proposed multiple-lifting scheme can reduce not only at least 50% on-chip memory bandwidth but also about 50% area of line buffer in 2-D DWT module.
Chih-Chi Cheng, Chao-Tsung Huang, Ching-Yeh Chen, Chung-Jr Lian, Liang-Gee Chen
IEEE Trans. Circuits Syst. Video Technol.2
2006 Line Buffer Wordlength Analysis for Line-Based 2-D DWT
abstract
The on-chip line buffer dominates the total area and power of line-based 2-D DWT. Therefore, the line buffer wordlength has to be carefully designed to maintain the quality level due to the dynamic range growing and the round-off errors. In this paper, a complete analysis methodology is proposed to derive the required wordlength of line buffer given the desired quality level of reconstructed image. The proposed methodology can guarantee to avoid overflow of coefficients, and the difference between predicted and experimental quality level is averagely 0.06 dB in terms of PSNR.
Chih-Chi Cheng, Chao-Tsung Huang, Jing-Ying Chang, Liang-Gee Chen
ICASSP (3)2
2006 Level C+ data reuse scheme for motion estimation with corresponding coding orders
abstract
The memory bandwidth reduction for motion estimation is important because of the power consumption and limited memory bandwidth in video coding systems. In this paper, we propose a Level C+ scheme which can fully reuse the overlapped searching region in the horizontal direction and partially reuse the overlapped searching region in the vertical direction to save more memory bandwidth compared to the Level C scheme. However, direct implementation of the Level C+ scheme may conflict with some important coding tools and then induces a lower hardware efficiency of video coding systems. Therefore, we propose n-stitched zigzag scan for the Level C+ scheme and discuss two types of 2-stitched zigzag scan for MPEG-4 and H.264 as examples. They can reduce memory bandwidth and solve the conflictions. When the specification is HDTV 720p, where the searching range is [-128,128), the required memory bandwidth is only 54%, and the increase of on-chip memory size is only 12% compared to those of traditional Level C data reuse scheme.
Ching-Yeh Chen, Chao-Tsung Huang, Yi-Hau Chen, Liang-Gee Chen
IEEE Trans. Circuits Syst. Video Technol.2
2006 High-Performance JPEG 2000 Encoder With Rate-Distortion Optimization
abstract
An 81 MSamples/s JPEG 2000 single-chip encoder is implemented on 5.5 mm/sup 2/ area using 0.25-/spl mu/m CMOS technology. This IC can losslessly encode HDTV 720p resolution at 30 frames/s in real time. Three techniques are adopted: line-based discrete wavelet transform, parallel embedded block coding, and precompression rate-distortion optimization. The line-based discrete wavelet transform achieves the minimum external memory access, while the internal memory is reduced by a proper memory access scheme. The parallel embedded block coding increases the throughput and reduces the memory bandwidth with similar hardware cost comparing to conventional architectures. By accurately estimating bit rates, the precompression rate-distortion optimization reduces the required computational power and processing time of the embedded block coding since the code-blocks are truncated before compression. Experimental results show that this encoder has the highest throughput with the smallest area compared with other designs in the literature.
Hung-Chi Fang, Tu-Chih Wang, Chao-Tsung Huang, Liang-Gee Chen
IEEE Trans. Multim.4
2005 Memory analysis of VLSI architecture for 5/3 and 1/3 motion-compensated temporal filtering [video coding applications]
abstract
To the best of authors' knowledge, this paper presents the first work on memory analysis of VLSI architectures for motion-compensated temporal filtering (MCTF). The open-loop MCTF prediction scheme has led the revolution for hybrid video coding methods that are mainly based on the close-loop MC prediction (MCP) scheme, and it also becomes the core technology of the coming video coding standard, MPEG-21 part 13-scalable video coding (SVC). In this paper, the macroblock (MB)-level and frame-level data reuse schemes are analyzed for the MCTF. The MB-level data reuse is especially for the motion estimation (ME), and the level C+ scheme is proposed, which can further reduce the memory bandwidth of the conventional level C scheme. Frame-level data reuse schemes for MCTF are proposed according to the open-loop prediction nature.
Chao-Tsung Huang, Ching-Yeh Chen, Yi-Hau Chen, Liang-Gee Chen
ICASSP (5)1
2005 System analysis of VLSI architecture for motion-compensated temporal filtering
abstract
The motion-compensated temporal filtering (MCTF) is an innovative prediction scheme for video coding and has become the core technology of the coming video coding standard, MPEG-21 part 13 - scalable video coding (SVC). This paper provides the system analysis of MCTF for VLSI implementation, which includes computational complexity, external memory access, external storage size, and coding delay. The one-level MCTF is analyzed first, and a modified double current frames scheme is introduced to address the external memory access penalty that results from fractional-pel motion compensation (MC). Then the analysis is extended to multi-level MCTF, in which many important system issues will be explored. Finally, a real-life test case was given to compare the system requirements of many different MCTF schemes and the prediction scheme of H.264/AVC.
Ching-Yeh Chen, Chao-Tsung Huang, Yi-Hau Chen, Chung-Jr Lian, Liang-Gee Chen
ICIP (3)2
2005 Advances in Hardware Architectures for Image and Video Coding - A Survey
abstract
This paper provides a survey of state-of-the-art hardware architectures for image and video coding. Fundamental design issues are discussed with particular emphasis on efficient dedicated implementation. Hardware architectures for MPEG-4 video coding and JPEG 2000 still image coding are reviewed as design examples, and special approaches exploited to improve efficiency are identified. Further perspectives are also presented to address the challenges of hardware architecture design for advanced image and video coding in the future.
Po-Chih Tseng, Yung-Chi Chang, Yu-Wen Huang, Hung-Chi Fang, Chao-Tsung Huang, Liang-Gee Chen
Proc. IEEE5
2005 Generic RAM-based architectures for two-dimensional discrete wavelet transform with line-based method
abstract
In this paper, three generic RAM-based architectures are proposed to efficiently construct the corresponding two-dimensional architectures by use of the line-based method for any given hardware architecture of one-dimensional (1-D) wavelet filters, including conventional convolution-based and lifting-based architectures. An exhaustive analysis of two-dimensional architectures for discrete wavelet transform in the system view is also given. The first proposed architecture is for 1-level decomposition, which is presented by introducing the categories of internal line buffers, the strategy of optimizing the line buffer size, and the method of integrating any 1-D wavelet filter. The other two proposed architectures are for multi-level decomposition. One applies the recursive pyramid algorithm directly to the proposed 1-level architecture, and the other one combines the two previously proposed architectures to increase the hardware utilization. According to the comparison results, the proposed architecture outperforms previous architectures in the aspects of line buffer size, hardware cost, hardware utilization, and flexibility.
Chao-Tsung Huang, Po-Chih Tseng, Liang-Gee Chen
IEEE Trans. Circuits Syst. Video Technol.1
2004 Memory analysis and architecture for two-dimensional discrete wavelet transform
abstract
The large amount of the frame memory access and the die area occupied by the embedded internal buffer are the most critical issues for the implementation of the two-dimensional discrete wavelet transform (2D DWT). The former may consume the most power and waste the system memory bandwidth. The latter may enlarge the chip size and also consume much power. We categorize and analyze the 2D DWT architectures by different external memory scan methods. Then the overlapped stripe-based scan method is proposed to provide an efficient and flexible implementation for 2D DWT. The implementation issues of the internal buffer are also discussed, including the lifting-based and convolution-based. Some real-life experiments are given to show that the performance of area and power for the internal buffer is highly related to memory technology and working frequency, instead of the required memory bits only.
Chao-Tsung Huang, Po-Chih Tseng, Liang-Gee Chen
ICASSP (5)1
2003 Hardware implementation of shape-adaptive discrete wavelet transform with the JPEG2000 defaulted (9, 7) filter bank
abstract
In this paper, an efficient hardware implementation of two-dimensional shape-adaptive discrete wavelet transform (2-D SA-DWT) with the JPEG2000 defaulted (9,7) filter bank is presented. Two techniques are used to minimize the critical path and the internal buffer size. One technique is the flipping structure which can shorten the critical path of the lifting scheme by flipping multiplier coefficients, rather than pipelining. The other technique is a shape-adaptive boundary handling strategy which can enhance lifting-based architectures to solve boundary extension problems of SA-DWT with little hardware overhead. A prototyping chip of this implementation will be fabricated with TSMC 0.25 /spl mu/m CMOS 1P5M process, and the estimated frequency and the core area are 50 MHz and 2.83 mm/sup 2/, respectively.
Chao-Tsung Huang, Po-Chih Tseng, Liang-Gee Chen
ICIP (2)1
2002 VLSI implementation of shape-adaptive discrete wavelet transform
Po-Chih Tseng, Chao-Tsung Huang, Liang-Gee Chen
VCIP2