EDBT 2026 Demo / reviewers in the wild / expert
Chia-Hung Yeh
dblp:37/2613
· DBLP profile ↗
58ranked-venue papers
24as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 16 first-author · 3 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graph convolutional network for fast video summarization in compressed domain
Chia-Hung Yeh, Chih-Ming Lien, Zhi-Xiang Zhan, Feng-Hsu Tsai, Mei-Juan Chen |
Neurocomputing | 1 |
| 2024 | Fine-grained video super-resolution via spatial-temporal learning and image detail enhancement
Chia-Hung Yeh, Hsin-Fu Yang, Yu-Yang Lin, Wan-Jen Huang, Feng-Hsu Tsai, Li-Wei Kang |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | SGENet: Spatial Guided Enhancement Network for Image Motion Deblurring
Yu-Chieh Wang, Chia-Hung Yeh |
BMVC | 2 |
| 2022 | Vision-oriented algorithm for fast decision in 3D video codingabstractAbstract This paper designs a novel method to reduce the coding complexity of 3D‐HEVC encoder by utilizing the properties of human visual perception. Two vision‐oriented edge detections are proposed: for colour texture detection, the authors adopt the Just‐Noticeable Distortion (JND); for depth map, the authors combine the Sample Adaptive Offset (SAO) and the Just Noticeable Depth Difference (JNDD) model. The authors also analyse the properties of colour texture and depth map to classify the coding tree unit (CTU) into various kinds of types, including complex‐edge CTU, moderate‐edge CTU and homogeneous CTU. Besides, fast mode decisions and early termination criteria are performed individually on each type of CTUs according to their characteristics. Especially for those CTUs with more edge information, the proposed projection‐based fast mode decision and residual‐based early termination preserve important colour texture while speeding up the coding at the same time. The proposed vision‐oriented algorithm reduces 31.981% of the overall average coding time with only 1.580% BD‐Bitrate increase. Experimental results show that the proposed algorithm can provide considerable time‐saving while still maintain the video quality, which outperforms the previous researches. Jie-Ru Lin, Mei-Juan Chen, Chia-Hung Yeh, Shinfeng D. Lin, Kuen-Liang Sue, Lih-Jen Kau, Yi-Sheng Ciou |
IET Image Process. | 3 |
| 2022 | A 40.96-GOPS 196.8-mW Digital Logic Accelerator Used in DNN for Underwater Object RecognitionabstractThis investigation presents a digital logic accelerator (DLA) design of a neural network hardware that utilizes output reuse. The DLA is used in the detection mechanism of underwater objects that was deployed in an underwater vehicle. A modified YoloV3-tiny network was also implemented to detect more than 20 underwater objects. The proposed DLA uses processing units that have parallel architectures of output windows, and output channels. Moreover, a new Inter-Controller is designed to control the direct memory access (DMA) together with a new Reshape module to improve the performance and power efficiency. A detailed description of the design as well as the measurements on silicon are presented. The chip is realized using a typical 180-nm CMOS process. It showed a performance result of 40.96 GOPS and the power consumption is 196.8 mW. The DLA was tested to demonstrate 19.88 frames per second and 40.96 GOPS. Chua-Chin Wang, Ralph Gerard B. Sangalang, Chien-Ping Kuo, Hsin-Che Wu, Yi Hsu, Shen-Fu Hsiao, Chia-Hung Yeh |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | Visual Perception Based Algorithm for Fast Depth Intra Coding of 3D-HEVCabstract3D-HEVC (The 3D Extension of High Efficiency Video Coding) is the newest 3D video coding standard, which enriches multimedia applications with the video format of multi-view plus depth. For the depth map coding in 3D-HEVC, the advanced coding tools enhance the coding efficiency of the depth map and the quality of the synthesized view. However, the time consumption and complexity of 3D-HEVC also increase significantly. This paper utilizes the characteristics of human visual system to propose a fast algorithm based on visual perception for the acceleration of the depth intra coding of 3D-HEVC. The depth map is segmented into different regions by Otsu's auto-thresholding. The dominate edge direction is categorized for each prediction unit. We detect the perceptual edge based on just noticeable depth difference model to extract the area that may affect the visual perception. According to depth map segmentation and edge distribution, we reduce the corresponding intra angular modes and determine whether to perform depth modelling mode. We also incorporate the boundary continuity and rate-distortion cost thresholding to propose the fast coding unit decision. The experimental results show that the proposed algorithm eliminates 53.09% of the depth coding time with only 0.15% BD-BR on average. The coding performance of the proposed algorithm outperforms the previous works significantly. Jie-Ru Lin, Mei-Juan Chen, Chia-Hung Yeh, Yong-Ci Chen, Lih-Jen Kau, Chuan-Yu Chang, Min-Hui Lin 0002 |
IEEE Trans. Multim. | 3 |
| 2022 | Lightweight Deep Neural Network for Joint Learning of Underwater Object Detection and Color ConversionabstractUnderwater image processing has been shown to exhibit significant potential for exploring underwater environments. It has been applied to a wide variety of fields, such as underwater terrain scanning and autonomous underwater vehicles (AUVs)-driven applications, such as image-based underwater object detection. However, underwater images often suffer from degeneration due to attenuation, color distortion, and noise from artificial lighting sources as well as the effects of possibly low-end optical imaging devices. Thus, object detection performance would be degraded accordingly. To tackle this problem, in this article, a lightweight deep underwater object detection network is proposed. The key is to present a deep model for jointly learning color conversion and object detection for underwater images. The image color conversion module aims at transforming color images to the corresponding grayscale images to solve the problem of underwater color absorption to enhance the object detection performance with lower computational complexity. The presented experimental results with our implementation on the Raspberry pi platform have justified the effectiveness of the proposed lightweight jointly learning model for underwater object detection compared with the state-of-the-art approaches. Chia-Hung Yeh, Chu-Han Lin, Li-Wei Kang, Chih-Hsiang Huang, Min-Hui Lin 0002, Chuan-Yu Chang, Chua-Chin Wang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | SRAM-Based Computation in Memory Architecture to Realize Single Command of Add-Multiply Operation and MultifunctionabstractThis paper presents a computation in memory (CIM) architecture and circuit design featured with single command to execute addition, signed multiplication, and multi-function to resolve poor computation throughput caused by von Neumann bottleneck. The proposed CIM takes advantage of 2T-Switch circuit which needs only 2 switches to select the required computation units such that the area on silicon is reduced. RCAM (ripple carry adder and multiply) unit realized with full swing gate diffusion input (FS-GDI) in a single-ended disturb- free 7T SRAM further reduces the power consumption and active circuit area. Auto-switching write-back circuit consisting of BL auto-switching circuit, Data switching circuit, and WL auto-switching circuit facilitates the automatic restore of addition and multiplication to designated memory addresses. The proposed CIM is realized using 40-nm CMOS process to demonstrated 12.18/28.19 fJ/bit normalized write/read energy at 100 MHz system clock rate. Chua-Chin Wang, Chia-Yi Huang, Chia-Hung Yeh |
ISCAS | 3 |
| 2021 | A 40-nm CMOS Multifunctional Computing-in-Memory (CIM) Using Single-Ended Disturb-Free 7T 1-Kb SRAMabstractThis investigation proposes a computing-in-memory (CIM) design to circumvent the von Neumann bottleneck which causes limited computation throughput for effective artificial intelligence (AI) applications. The proposed CIM performs multiple operations such as single-instruction basic Boolean operations, addition, and signed number multiplication, and multiple functions such as normal mode and retention mode for the built-in self-test (BIST). Its 2T-Switch requires only two transistors to be utilized for static random-access memory (SRAM) array; thus, the arithmetic unit can be chosen easily and the area overhead is minimized. Its ripple carry adder and multiplier (RCAM) unit based on single-ended disturb-free 7T 1-Kb SRAM was developed using the full swing-gate diffusion input (FS-GDI) technology that has full voltage swing resolution, low power consumption, and less chip area cost. Its Auto-Switching Write Back Circuit restores addition and multiplication operations automatically to assigned memory address. The CIM is implemented using the TSMC 40-nm CMOS process, where the core area is$432.81 \times 510.265\,\,\mu \text{m}^{2}$. Among the related works, the proposed CIM performs the most number of operations and functions. Chua-Chin Wang, Lean Karlo S. Tolentino, Chia-Yi Huang, Chia-Hung Yeh |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | Multi-Scale Deep Residual Learning-Based Single Image Haze Removal via Image DecompositionabstractImages/videos captured from outdoor visual devices are usually degraded by turbid media, such as haze, smoke, fog, rain, and snow. Haze is the most common one in outdoor scenes due to the atmosphere conditions. In this paper, a novel deep learning-based architecture (denoted by MSRL-DehazeNet) for single image haze removal relying on multi-scale residual learning (MSRL) and image decomposition is proposed. Instead of learning an end-to-end mapping between each pair of hazy image and its corresponding haze-free one adopted by most existing learningbased approaches, we reformulate the problem as restoration of the image base component. Based on the decomposition of a hazy image into the base and the detail components, haze removal (or dehazing) can be achieved by both of our multi-scale deep residual learning and our simplified U-Net learning only for mapping between hazy and haze-free base components, while the detail component is further enhanced via the other learned convolutional neural network (CNN). Moreover, benefited by the basic building block of our deep residual CNN architecture and our simplified UNet structure, the feature maps (produced by extracting structural and statistical features), and each previous layer can be fully preserved and fed into the next layer. Therefore, possible color distortion in the recovered image would be avoided. As a result, the final haze-removed (or dehazed) image is obtained by integrating the haze-removed base and the enhanced detail image components. Experimental results have demonstrated good effectiveness of the proposed framework, compared with state-ofthe-art approaches. Chia-Hung Yeh, Chih-Hsiang Huang, Li-Wei Kang |
IEEE Trans. Image Process. | 1 |
| 2019 | Fast prediction for quality scalability of High Efficiency Video Coding Scalable Extension
Chih-Hsuan Yeh, Jie-Ru Lin, Mei-Juan Chen, Chia-Hung Yeh, Cheng-An Lee, Kuang-Han Tai |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Efficient inter-prediction depth coding algorithm based on depth map segmentation for 3D-HEVC
Yi-Wen Liao, Mei-Juan Chen, Chia-Hung Yeh, Jie-Ru Lin, Chih-Wei Chen |
Multim. Tools Appl. | 3 |
| 2019 | Visual-Quality Guided Global Backlight Dimming for Video Display on Mobile DevicesabstractThis proposes a visual-quality guided global backlight dimming (VQG-GBD) algorithm to reduce the power consumption of liquid-crystal display on mobile devices. We build a backlight scaling ratio (BSR) prediction model via visual-quality assessment that not only considers the display contents but also the backlight intensity while measuring video quality. Also, we add visual uncertainty as an indicator to dim the backlight without being noticed by observers. The VQG-GBD includes a training stage and an online stage. For the training stage, first, we collect videos with distinct attributes of brightness and uncertainty. Then, the subjective rating obtains the relationship among the visual quality, BSR, brightness, and visual uncertainty. Finally, we use the trust-region method to build the BSR prediction model. In the online stage, the model is applied to mobile devices for real-time video display and a BSR optimization strategy is proposed to eliminate the flicker effect between frames, followed by three techniques to accelerate the process: 1) motion vector extraction; 2) pixel subsampling to reduce the computation while analyzing frame content; and 3) GPU rendering to speed up the pixel compensation. The experimental results show that VQG-GBD achieves 21% of the power demand reduction on average for displaying videos on mobile devices while preserving good visual quality. The VQG-GBD delivers more power reduction than the state-of-the-art algorithm image integrity-based gray-level error control and multi-histogram-based gray-level error control by 10% and 8%, respectively. Chia-Hung Yeh, Kyle Shih-Huang Lo, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Coding unit complexity-based predictions of coding unit depth and prediction unit mode for efficient HEVC-to-SHVC transcoding with quality scalability
Chia-Hung Yeh, Wen-Yu Tseng, Li-Wei Kang, Cheng-Wei Lee 0004, Kahlil Muchtar, Mei-Juan Chen |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Rain streak removal based on non-negative matrix factorization
Chia-Hung Yeh, Chih-Yang Lin, Kahlil Muchtar, Pin-Hsian Liu |
Multim. Tools Appl. | 1 |
| 2018 | Efficient CU and PU Decision Based on Motion Information for Interprediction of HEVCabstractHigh-efficiency video coding encoders provide great improvements in coding efficiency and can also support higher resolution and multiple coding tools. The new coding structures such as coding unit (CU) and prediction unit (PU) have helped a lot, but the computational complexity is much higher than those of previous standards. This paper proposes a fast algorithm combining with CU and PU early termination decisions to reduce computational demand. Based on the analytic results, we can set up an adaptive threshold that can be obtained for early termination. Meanwhile, we also develop an adaptive search range determination according to the motion vector (MV). Compared with HM 12.0, our proposed method achieves an approximate 57% time saving, whereas the average Bjøntegaard-Delta Bit-rate (BDBR) increase is only 0.43%. In addition, our fast algorithm outperforms the previous works in both coding speed and coding performance. Mei-Juan Chen, Yu-De Wu, Chia-Hung Yeh, Kao-Min Lin, Shinfeng D. Lin |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | Moving object detection in the encrypted domain
Chih-Yang Lin, Kahlil Muchtar, Jia-Ying Lin, Yu-Hsien Sung, Chia-Hung Yeh |
Multim. Tools Appl. | 5 |
| 2016 | Robust techniques for abandoned and removed object detection based on Markov random field
Chih-Yang Lin, Kahlil Muchtar, Chia-Hung Yeh |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Secure multicasting of images via joint privacy-preserving fingerprinting, decryption, and authentication
Chih-Yang Lin, Kahlil Muchtar, Chia-Hung Yeh, Chun-Shien Lu |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | A gender classification scheme based on multi-region feature extraction and information fusion for unconstrained images
Guo-Shiang Lin, Min-Kuan Chang, Yu-Jui Chang, Chia-Hung Yeh |
Multim. Tools Appl. | 4 |
| 2015 | Inter-embedding error-resilient mechanism in scalable video coding
Chia-Hung Yeh, Shu-Jhen Fan-Jiang, Chih-Yang Lin, Min-Kuan Chang, Mei-Juan Chen |
Inf. Sci. | 1 |
| 2015 | A new intra prediction with adaptive template matching through finite state machine
Chia-Hung Yeh, Shu-Jhen Fan-Jiang, Chih-Yang Lin, Pei-Lun Suei, Min-Kuan Chang |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Robust Laser Speckle Authentication System Through Data Mining TechniquesabstractThis paper proposes a speckle image recognition method using data mining techniques to ensure speckle identification system feasible for authentication. This is an interdisciplinary method that integrates the researches of optics, data mining, and image processing. Because objects have unique but imperfect surfaces, their laser speckle is capable of providing suitable identifiable features for authentication. In our method, matching points among speckle images acquired from one plastic card are extracted by scale-invariant feature transform (SIFT). The spatial relations among the matching points are then transformed to 9 direction lower triangular (9DLT) representations. Then, the Apriori algorithm mines frequent patterns so a useful association rule is obtained as the feature to identify the similarity between each of the speckle images for the purpose of authenticity verification. The proposed method is especially robust in the cases of card displacement and luminance change resulted from laser attenuation. Experimental results show that the proposed method has promising results and outperforms existing methods in identification accuracy. Chia-Hung Yeh, Guanling Lee, Chih-Yang Lin |
IEEE Trans. Ind. Informatics | 1 |
| 2015 | Learning-Based Joint Super-Resolution and Deblocking for a Highly Compressed ImageabstractA highly compressed image is usually not only of low resolution, but also suffers from compression artifacts (blocking artifact is treated as an example in this paper). Directly performing image super-resolution (SR) to a highly compressed image would also simultaneously magnify the blocking artifacts, resulting in an unpleasing visual experience. In this paper, we propose a novel learning-based framework to achieve joint single-image SR and deblocking for a highly-compressed image. We argue that individually performing deblocking and SR (i.e., deblocking followed by SR, or SR followed by deblocking) on a highly compressed image usually cannot achieve a satisfactory visual quality. In our method, we propose to learn image sparse representations for modeling the relationship between low- and high-resolution image patches in terms of the learned dictionaries for image patches with and without blocking artifacts, respectively . As a result, image SR and deblocking can be simultaneously achieved via sparse representation and morphological component analysis (MCA)-based image decomposition. Experimental results demonstrate the efficacy of the proposed algorithm. Li-Wei Kang, Chih-Chung Hsu, Boqi Zhuang, Chia-Wen Lin, Chia-Hung Yeh |
IEEE Trans. Multim. | 5 |
| 2015 | Predictive Texture Synthesis-Based Intra Coding Scheme for Advanced Video CodingabstractThis paper aims to improve the intra coding performance on H.264 and high efficient video coding (HEVC). A new intra prediction approach is proposed based on synthesizing two neighboring predictors. These two predictors are selected from different prediction directions and then the predicted block is generated by combining the two predictors using different weights. The weights for the predictors do not need to be saved, so the bit rate can be greatly reduced. The main contributions of this paper are proposing the following: 1) a highly efficient way and high compact compression to select the proper predictors; 2) a weight estimation method for the predictors; and 3) a reversible weight restoration method for the predictors to save the bit rate in the decoding phase. Experimental results show that the proposed method outperforms H.264/AVC and HEVC in intra prediction by 11.79% and 3.6%, respectively, in bitrate reduction. Chia-Hung Yeh, Tsung-Yih Tseng, Cheng-Wei Lee 0004, Chih-Yang Lin |
IEEE Trans. Multim. | 1 |
| 2014 | A joint content adaptive Rate-Quantization model and region of interest intra coding of H.264/AVCabstractAssigning an appropriate Quantization Parameter (QP) to the Intra-coded frames is important to the video coding. In this research, a content adaptive Rate-Quantization (R-Q) model is presented to predict the bit usage of Intra-coded frames in H.264/AVC. The relationship between the QP of a macroblock and the block complexity is derived so that a suitable QP can be decided under a target bit-rate. Since the proposed model is built on macroblocks, Region of Interest (ROI) coding can also be achieved. By adjusting the QP value at the macroblock level, more bits can be assigned to the ROI to better maintain its perceptual quality. The experimental results show the feasibility of the proposed R-Q model. Ching-Yu Wu, Po-Chyi Su, Chia-Hung Yeh, Hui-Chun Hsu |
ICME | 3 |
| 2014 | Grabcut-based abandoned object detectionabstractThis paper presents a detection-based method to subtract abandoned object from a surveillance scene. Unlike tracking-based approaches that are commonly complicated and unreliable on a crowded scene, the proposed method employs background (BG) modelling and focus only on immobile objects. The main contribution of our work is to build abandoned object detection system which is robust and can resist interference (shadow, illumination changes and occlusion). In addition, we introduce the MRF model and shadow removal to our system. MRF is a promising way to model neighbours' information when labeling the pixel that is either set to background or abandoned object. It represents the correlation and dependency in a pixel and its neighbours. By incorporating the MRF model, as shown in the experimental part, our method can efficiently reduce the false alarm. To evaluate the system's robustness, several dataset including CAVIAR datasets and outdoor test cases are both tested in our experiments. Kahlil Muchtar, Chih-Yang Lin, Chia-Hung Yeh |
MMSP | 3 |
| 2014 | Real-time background modeling based on a multi-level texture description
Chia-Hung Yeh, Chih-Yang Lin, Kahlil Muchtar, Li-Wei Kang |
Inf. Sci. | 1 |
| 2014 | Self-learning-based post-processing for image/video deblocking via sparse representation
Chia-Hung Yeh, Li-Wei Kang, Yi-Wen Chiou, Chia-Wen Lin, Shu-Jhen Fan-Jiang |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Popular music representation: chorus detection & emotion recognition
Chia-Hung Yeh, Wen-Yu Tseng, Chia-Yen Chen, Yu-Dun Lin, Yi-Ren Tsai, Hsuan-I Bi, Yu-Ching Lin, Ho-Yi Lin |
Multim. Tools Appl. | 1 |
| 2014 | Reversible joint fingerprinting and decryption based on side match vector quantization
Chih-Yang Lin, Panyaporn Prangjarote, Chia-Hung Yeh, Hui-Fuang Ng |
Signal Process. | 3 |
| 2014 | Efficient multi-view video coding using inter-view information
Xin-Xian Huang, Mei-Juan Chen, Chia-Hung Yeh, Hao-Wen Chi, Chia-Yen Chen |
Signal Process. Image Commun. | 3 |
| 2014 | Fast Mode Decision Algorithm Through Inter-View Rate-Distortion Prediction for Multiview Video Coding SystemabstractMultiview video coding (MVC) has attracted great attention from industries and research institutes. MVC is used to encode stereoscopic video streams for 3D playout systems such as 3D television, digital cinema, and IP network applications. MVC is an extended version of H.264/AVC that improves the performance of multiview videos. Yet, when compared with single-view video coding, MVC consumes much more time when encoding large amounts of data. Speed-up algorithms, therefore, are essential for realizing related applications. This paper presents a fast mode decision algorithm to avoid the high computational complexity of MVC. The proposed approach aims to reduce candidate modes and make mode decision process more efficient. The minimum and maximum values of rate-distortion cost (RD cost) in the previously encoded view are used to compute a threshold for each mode in the current view. Compared with joint multiview video coding, the experimental results demonstrate that the proposed algorithm provides an average of 79% in time savings with negligible bit rate increase and peak signal-to-noise ratio decrease. Chia-Hung Yeh, Ming-Feng Li, Mei-Juan Chen, Ming-Chieh Chi, Xin-Xian Huang, Hao-Wen Chi |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | Self-learning-based single image super-resolution of a highly compressed imageabstractLow-quality images are usually not only with low-resolution, but also suffer from compression artifacts (blocking artifact is treated as an example in this paper). Directly performing image super-resolution (SR) to a highly compressed (low-quality) image would also simultaneously magnify the blocking artifacts, resulting in unpleasing visual quality. In this paper, we propose a self-learning-based SR framework to simultaneously achieve single-image SR and compression artifact removal for a highly-compressed image. We argue that individually performing deblocking first, followed by SR to an image, would usually inevitably lose some image details induced by deblocking, which may be useful for SR, resulting in worse SR result. In our method, we propose to self-learn image sparse representation for modeling the relationship between low and high-resolution image patches in terms of the learned dictionaries, respectively, for image patches with and without blocking artifacts. As a result, image SR and deblocking can be simultaneously achieved via sparse representation and MCA (morphological component analysis)-based image decomposition. Experimental results demonstrate the efficacy of the proposed algorithm. Li-Wei Kang, Bo-Chi Chuang, Chih-Chung Hsu, Chia-Wen Lin, Chia-Hung Yeh |
MMSP | 5 |
| 2012 | Efficient image/video deblocking via sparse representationabstractBlocking artifact, characterized by visually noticeable changes in pixel values along block boundaries, is a common problem in block-based image/video compression, especially at low bitrate coding. Various post-processing techniques have been proposed to reduce blocking artifacts, but they usually introduce excessive blurring or ringing effects. This paper proposes a self-learning-based image/ video deblocking framework via properly formulating deblocking as an MCA (morphological component analysis)-based image decomposition problem via sparse representation. The proposed method first decomposes an image/video frame into the low-frequency and high-frequency parts by applying BM3D (block-matching and 3D filtering) algorithm. The high-frequency part is then decomposed into a “blocking component” and a “non-blocking component” by performing dictionary learning and sparse coding based on MCA. As a result, the blocking component can be removed from the image/video frame successfully while preserving most original image/video details. Experimental results demonstrate the efficacy of the proposed algorithm. Yi-Wen Chiou, Chia-Hung Yeh, Li-Wei Kang, Chia-Wen Lin, Shu-Jhen Fan-Jiang |
VCIP | 2 |
| 2012 | Mode decision acceleration for scalable video coding through coded block pattern
Chia-Hung Yeh, Wen-Yu Tseng, Bo-Yi Chou |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | New intra-prediction with finite state machine for H.264/AVCabstractThis paper proposes a new approach to improve the coding performance of intra block coding in H.264/AVC via finite state machine. Grounding on high correlation between neighboring blocks, finite state machine is employed both at encoder and decoder to reduce the number of bits required for encoding to enhance coding performance. Two extra intra prediction modes are created in our proposed method. Through these two modes, the number of bits required to denote the current block is greatly reduced and low bit rate can be achieved. Experimental results show that the proposed method can greatly improve coding efficiency of intra macroblock coding in H.264/AVC. Chia-Shiu Wu, Shu-Jhen Fan-Jiang, Chia-Hung Yeh |
VCIP | 3 |
| 2010 | Prediction Error Prioritizing Strategy for Fast Normalized Partial Distortion Motion Estimation AlgorithmabstractA prediction error prioritizing-based normalized partial distortion search algorithm for fast motion estimation is proposed in this letter. The distortion behavior of each pixel in a macroblock is first analyzed to point out the priority/order of sum of absolute difference calculation. Afterward, the normalized partial distortion search algorithm is applied for half-stop of the distortion calculation. In addition, a dynamic search range decision algorithm is adopted for automatically changing the size of the search range to further increase the motion estimation speed. The computational complexity can be reduced significantly through the proposed algorithm, though leaving a PSNR degradation that could be dismissed. Chao-Cing Yang, Gwo-Long Li, Ming-Chieh Chi, Mei-Juan Chen, Chia-Hung Yeh |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Fast Mode Decision Algorithm for Scalable Video Coding Using Bayesian Theorem Detection and Markov ProcessabstractThenewestvideo coding standard called scalable video coding (SVC) provides broad applications in multimedia communications. SVC encoder consumes great computational complexity when compared to previous video coding standards. This paper presents a fast mode decision algorithm that speeds up the SVC encoding process through probabilistic analysis. The mode of the enhancement layer is first predicted by statistical analysis. Afterward, Bayesian theorem is utilized to detect whether the prediction mode of the current macroblock is the best or not. The mode is further predicted and refined by the Markov process. Experimental results show that the proposed algorithm significantly reduces computational complexity with negligible peak signal-to-noise ratio degradation and bitrate increase in the enhancement layers. Chia-Hung Yeh, Kai-Jie Fan, Mei-Juan Chen, Gwo-Long Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Adaptive Video De-interlacing Algorithm by Motion Characteristic DetectionabstractIn this paper, an adaptive video de-interlacing algorithm based on motion characteristics detection jointly considering the advantages of three existing schemes is proposed. In the proposed method, the difference of the previous field pixels and the current field pixels are calculated. The fast or slow of the motion of the current frame is determined by the difference of the previous and the current field pixels. Afterwards, the appropriate scheme such as the weave, line average or edge line average de-interlacing schemes is selected and performed on the selected field. Through the proposed algorithms, the performance of the proposed method in terms of subjective and objective qualities can be significantly improved. Chia-Na Tsai, Ming-Te Wu, Chia-Hung Yeh, Mei-Juan Chen, Hsuan-Ting Chang |
ISCAS | 3 |
| 2009 | An Efficient Emotion Detection Scheme for Popular MusicabstractWith the rapid growth of multimedia information, the ability to efficiently manage data from large amount of multimedia database has become a crucial issue. In this paper, a framework for music emotion detection is proposed. First, a Thayer's 2-dimentinal model that represents the music emotion space is employed as our emotion model. Second, three features such as intensity, rhythm regularity, and tempo are extracted to describe a music clip. Then, features are trained by constructing Gaussian mixture models (GMM). Finally, the likelihood radios of test music clips to GMM are calculated for emotion identification. Experiemtal results show that the average recall and precision all are up to 80% for the database that is comprised of 145 music clips. Chia-Hung Yeh, Hung-Hsuan Lin, Hsuan-Ting Chang |
ISCAS | 1 |
| 2009 | Movie story intensity representation through audiovisual tempo analysis
Chia-Hung Yeh, Chih-Hung Kuo, Rung-Wen Liou |
Multim. Tools Appl. | 1 |
| 2009 | Robust Region-of-Interest Determination Based on User Attention Model Through Visual Rhythm AnalysisabstractRegion-of-interest (ROI) determination is very important for video processing and it is desirable to find a simple method to identify the ROI. Along this direction, this paper investigates a user attention model based on visual rhythm analysis for automatic determination of ROI in a video. The visual rhythm, which is an abstraction of a video, is a thumbnail version of a video by a 2-D image that captures the temporal information of a video sequence. Four sampling lines, including diagonal, anti-diagonal, vertical, and horizontal lines, are employed to obtain four visual rhythm maps in order to analyze the location of the ROI from video data. Via the variation on visual rhythms, object and camera motions can be efficiently distinguished. As for hardware design consideration, the proposed scheme can accurately extract ROI with very low computational complexity for real-time applications. The promising results from the experiments demonstrate that the moving object is effectively and efficiently extracted. Finally, we present a way to use flexible macroblock ordering in combination with ROI determination as a preprocessing step for H.264/AVC video coding, and experimental results show the quality of ROI regions is significantly enhanced. Ming-Chieh Chi, Chia-Hung Yeh, Mei-Juan Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Vision-based vehicle event detection through visual rhythm analysisabstractIn this paper, a simple and reliable on-road vehicle event detection algorithm is proposed to identify events for vehicle. A virtual line at the same position of a frame is employed to extract visual rhythm. The visual rhythm is a compact representation of a video that captures the temporal information of vehicle status of a coarsely spatially sampled video sequence. By analyzing statistical characteristics of the visual rhythm, the events such as safe distance, passing and lane changing can be effectively detected. The proposed techniques can prevent accidents and improve traffic safety by monitoring the alertness of drivers, augmenting vision fields to prevent collision. The proposed system is efficient both in terms of computational complexity and memory requirements. Experimental results show the efficiency and effectiveness of the proposed system for intelligent transport system. Chia-Hung Yeh, Jia-Chi Bai, Sun-Chen Wang, Po-Yi Sung, Ruey-Nan Yeh, Maverick Shih |
ICME | 1 |
| 2008 | Region-of-interest video coding based on rate and distortion variations for H.263+
Ming-Chieh Chi, Mei-Juan Chen, Chia-Hung Yeh, Jyong-An Jhu |
Signal Process. Image Commun. | 3 |
| 2007 | Robust Region-of-Interest Determination Based on User Attention Model Through Visual Rhythm AnalysisabstractThis paper investigates a user attention model based on the visual rhythm analysis for automatically determining the region-of-interest (ROI) in a video. The visual rhythm, an abstraction of a video, is a thumbnail version of a fully video by a 2D image that captures the temporal information of a video sequence. Four sampling lines, including diagonal, anti-diagonal, vertical and horizontal lines, are employed to obtain four visual rhythm maps in order to analyze the location of the ROI from video data. Via the variation on visual rhythms, object and camera motions can be efficiently distinguished. The proposed scheme can extract the ROI accurately with very low computational complexity. The promising results from the experiments demonstrate that the moving object is effectively and efficiently extracted. Ming-Chieh Chi, Chia-Hung Yeh, Mei-Juan Chen, Ching-Ting Hsu |
ICCCN | 2 |
| 2007 | A Real Time and Low Cost Hardware Architecture for Video Abstraction SystemabstractIn this paper, we propose a real time hardware architecture with low cost for histogram difference calculator to process and low-level extraction in our video abstraction system. The highlights are extracted from video data by directly combining simple visual and audio features of video data without specific domain knowledge and hence its low complexity makes it suitable to be implemented in embedded systems. Experimental results show that the condensed skimming clips comprise the interesting and informative parts and most meaningful semantic contents are extracted efficiently Li-Chuan Chang, Yen-Sung Chen, Rung-Wen Liou, Chih-Hung Kuo, Chia-Hung Yeh, Bin-Da Liu |
ISCAS | 5 |
| 2006 | Robust TV News Story Identification Via Visual Characteristics of Anchorperson Scenes
Chia-Hung Yeh, Min-Kuan Chang, Ko-Yen Lu, Maverick Shih |
PSIVT | 1 |
| 2004 | Video skimming based on story units via general tempo analysisabstractA skimming system for movie content exploration is proposed using story units extracted via general tempo analysis of audio and visual data. Quite a few schemes have been proposed to segment video data into shots with low-level features, yet the grouping of shots into meaningful units, called story units here, is important and challenging. In this work, we detect similar shots using keyframes and include these similar shots as a node. Then, an importance measure is calculated based on the total length of each node. Finally, we select sinks and shots according to this measure. Based on these semantic shots, a meaningful skim can be successfully generated. Simulation results are presented to show that the proposed video skimming scheme can preserve the essential and significant content of the original video data. Shih-Hung Lee, Chia-Hung Yeh, C.-C. Jay Kuo |
ICME | 2 |
| 2004 | Data hiding domain classification for blind image steganalysisabstractA statistical feature-based scheme is proposed to identify the data hiding domain of an embedded signal in this research. Two phenomena are observed for images before and after data hiding. First, the gradient energy increases as the continuity of gray levels between adjacent pixels is disturbed by the embedded signal. Second, the statistical variance of the coefficient distribution in macro-blocks tends to decrease after data hiding. These phenomena are analyzed mathematically. Then, statistical features in the pixel, DCT, and DWT domains are extracted and a maximum likelihood ratio test is adopted to solve the hiding domain classification problem. The proposed scheme has demonstrated good classification results. Guo-Shiang Lin, Chia-Hung Yeh, C.-C. Jay Kuo |
ICME | 2 |
| 2004 | Robust traffic event extraction via content understanding for highway surveillance systemabstractA method to extract traffic events by integrating the low-level, middle-level, and high-level feature extraction modules is developed in this research. The low-level module extracts features such as motion, size, and location. The middle-level module builds a bridge between the road surface plane in the real world and the captured image plane via geometric analysis. Finally, the high-level module identifies traffic events such as "traffic jam", "lane change", and "traffic rule violation", which require the understanding of video content in a specific knowledge domain. In the high-level module, various traffic events are related to motion characteristics obtained from the middle-level module. It is demonstrated by experimental results that the proposed system can achieve robust traffic event extraction. Akio Yoneyama, Chia-Hung Yeh, C.-C. Jay Kuo |
ICME | 2 |
| 2004 | Robust traffic event extraction from surveillance videoabstractAn approach to extract traffic events by integrating the low-level, middle-level, and high-level feature extraction modules is developed in this research. To be more specific, the low-level module extracts features such as motion, size, and location. The middle-level module builds a bridge between the road surface plane in the real world and the captured image plane by geometric analysis. Finally, the high-level module looks for traffic events such as "traffic jam", "lane change", and "traffic rule violation", which require the understanding of the video contents in a specific knowledge domain. In the high-level module, various traffic events are related to motion characteristics obtained from the middle-level module. It is demonstrated by experimental results that the proposed system can achieve robust traffic event extraction. The effectiveness of the proposed technique is analyzed. Conventional traffic event extraction methods demand the knowledge of capturing conditions for camera calibration. This requirement can be greatly relaxed in our proposed scheme. Akio Yoneyama, Chia-Hung Yeh, C.-C. Jay Kuo |
VCIP | 2 |
| 2003 | Moving Cast Shadow Elimination for Robust Vehicle Extraction Based on 2D Joint Vehicle/Shadow ModelsabstractA new algorithm to eliminate moving cast shadow for robust vehicle detection and extraction in a vision-based highway monitoring system is investigated. The proposed algorithm is based on a simplified 2D vehicle/shadow model of six types projected on to a 2D image plane. Parameters of the joint 2D vehicle/shadow models can be estimated from the input video without light source and camera calibration information. Simulations are performed to verify that the proposed technique is effective for vision-based highway surveillance systems. Akio Yoneyama, Chia-Hung Yeh, C.-C. Jay Kuo |
AVSS | 2 |
| 2003 | Iteration-free clustering algorithm for nonstationary image databaseabstractImage database systems must effectively and efficiently handle and retrieve images from a large collection of images. A serious problem faced by these systems is the requirement to deal with the nonstationary database. In an image database system, image features are typically organized into an indexing structure, and updating the indexing structure involves many computations. In this paper, this difficult problem is converted into a constrained optimization problem, and the iteration-free clustering (IFC) algorithm based on the Lagrangian function, is presented for adapting the existing indexing structure for a nonstationary database. Experimental results concerning recall and precision indicate that the proposed method provides a binary tree that is almost optimal. Simulation results further demonstrate that the proposed algorithm can maintain 94% precision in seven-dimensional feature space, even when the number of new-coming images is one-half the number of images in the original database. Finally, our IFC algorithm outperforms other methods usually applied to image databases. Chia-Hung Yeh, Chung J. Kuo |
IEEE Trans. Multim. | 1 |
| 2002 | Boundary block-searching algorithm for arbitrary shaped codingabstractArbitrarily shaped coding is an important tool to achieve MPEG-4, which is a object-based coding standard. We proposed an efficient shaped coding method which enhances the coding efficiency of conventional padding techniques in the object-based video coding where the shape information is provided. The boundary block-searching algorithm (BBS) is applied to the boundary blocks of a VOP (video object plane) which consists of both background and object pixels. Actually, boundary blocks have strong correlation even though they are located apart. For an input boundary block, we search the closest block (only object pixels) from the previously encoded data using the BBS algorithm. After searching, the boundary block is represented as a position vector. This process reduces the number of boundary blocks to be DCT and makes the bit rate small. Experimentation has been done on various video sequences under different test conditions and verifies significant coding efficiency improvement. Chia-Hung Yeh, Hsuan T. Chang, Chung J. Kuo |
ICME (1) | 1 |
| 2001 | Polynomial Motion-Vector Resampling AlgorithmabstractThere is a need to downscale the video (from image size to ) prior to transmission due to bandwidth limitation on the cannel. The conventional method for downscaling the video sequence is to decompress the video bitstream and then reencode it again at video server. This process is computationally intensive because of the motion estimation required during the reencoding process. In this paper, we proposed a polynomial motion-vector resampling (PMVR) algorithm to compute motion vectors for the downscaled video sequence in compressed domain. The proposed approach leads to significant computational saving compared with the conventional spatial domain approach. Simulation results suggest that the proposed method is effective for video downscaling. Chia-Hung Yeh, Chung J. Kuo |
ICME | 1 |
| 2000 | . Polynomial search algorithms for motion estimationabstractThis paper proposes a polynomial search (PS) algorithm and architecture to solve the motion estimation problem in video coding. Simulation results show that the proposed method is not only flexible, but also requires fewer computations to achieve the same mean absolute error results (for QCIF and sub-QCIF video) compared with the existing fast-search algorithms. Finally, a VLSI architecture is also developed to efficiently implement the PS algorithm. Chung J. Kuo, Chia-Hung Yeh, Souheil F. Odeh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Noise reduction of VQ encoded images through anti-gray codingabstractNoise reduction of VQ encoded images is achieved through the proposed anti-gray coding (AGC) and noise detection and correction scheme. In AGC, binary indices are assigned to the codevector in such a way that the 1-b neighbors of a code vector are as far apart as possible. To detect the channel errors, we first classify an image into uniform and edge regions. Then we propose a mask to detect the channel errors based on the image classification (uniform or edge region) and the characteristics of AGC. We also mathematically derive a criterion for error detection based on the image classification. Once error indices are detected, the recovered indices can be easily chosen from a "candidate set" by minimizing the gray-level transition across the block boundaries in a VQ encoded image. Simulation results show that the proposed technique provides detection results with smaller than 0.1% probability of error and more than 86.3% probability of detection at a random bit error rate of 0.1%, while the undetected errors are invisible. In addition, the proposed detection and correction techniques improve the image quality (compared with that encoded by AGC) by 3.9 dB. Chung J. Kuo, Chien H. Lin, Chia-Hung Yeh |
IEEE Trans. Image Process. | 3 |