EDBT 2026 Demo / reviewers in the wild / expert
Homer H. Chen
dblp:19/982
· DBLP profile ↗
163ranked-venue papers
15as first author
19since 2021 · last 2025
0000-0002-8795-1911ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 132 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 18 · 8 first-author · 2 since 2021Systems, architecture and hardware · 6Computer networks · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training A Phase Detection Autofocus Model Using Hybrid LabelsabstractPhase detection autofocus (PDAF) technology is essential to digital cameras in many applications including professional photography and autonomous navigation. A deep-learning-based autofocus method is advantageous over conventional AF methods in both speed and accuracy; however, the quality of training data remains a major bottleneck. In this paper, we propose a hybrid labeling strategy that leverages the complementary strengths of focus profiles derived from phase and RGB data. The latter has finer spatial resolution but is noisier than the former. By including both types of data for model training, better accuracy and reliability for PDAF can be achieved. Experiments on various scenes demonstrate that the proposed hybrid labeling strategy achieves higher accuracy than monotype labeling strategies, leading to a practical single-camera alternative to multi-camera-based or depth-based labeling solutions that are often clumsy and computationally expensive. Chen-Han Lin, Homer H. Chen |
ICIP | 2 |
| 2025 | Texturing Endoscopic 3D Stomach via Neural Radiance Field Under Uneven LightingabstractTexture generation is crucial for endoscopic 3D reconstruction, as it provides essential visual information for computer-assisted medical diagnosis and surgical procedures. Previous mapping-based methods for texturing 3D stomach depend on high-quality mesh structures for accurate camera view selection. However, it is challenging to obtain high-quality mesh structures in endoscopic 3D reconstruction. To eliminate this dependency, we propose an alternative texture generation method that extracts texture directly from a neural radiance field, removing the need for camera view selection. Furthermore, since endoscopic images often suffer from uneven lighting including local low light and overexposure, we develop a weight mechanism to guide our model in prioritizing the learning of pixels that clearly depict the stomach wall. Experimental results demonstrate that our method is more robust than previous approaches in texturing 3D stomach models and effectively mitigates lighting artifacts, thereby producing high-fidelity textures that are crucial for downstream tasks. Ze-Yan Lu, Ren-Hau Shiue, Ming-Lun Han, Kuang-Chen Yen, Homer H. Chen |
ICIP | 5 |
| 2025 | Generating Light Field From Stereo Images for AR Display With Matched Angular Sampling Structure and Minimal Retinal ErrorabstractNear-eye light field displays offer natural 3D visual experiences for AR/VR users by projecting light rays onto retina as if the light rays were emanated from a real object. Such displays normally take four-dimensional light field data as input. Given that sizeable existing 3D contents are in the form of stereo images, we propose a practical approach that generates light field data from such contents at minimal computational cost while maintaining a reasonable image quality. The perceptual quality of light field is ensured by making the baseline of light field subviews consistent with that of the micro-projectors of the light field display and by compensating for the optical artifact of the light field display through digital rectification. The effectiveness and efficiency of the proposed approach is verified through both quantitative and qualitative experiments. The results demonstrate that our light field converter works for real-world light field displays. Yi-Chou Chen, Homer H. Chen |
IEEE Trans. Image Process. | 2 |
| 2025 | Comparison of User Performance and Experience Between Light Field and Conventional AR GlassesabstractLight field AR glasses can provide better visual comfort than conventional AR glasses; however, studies on user performance comparison between them are notably scarce. In this article, we present a systematic method employing a serial visual search task without confounding factors to quantify and compare the user performance and experience between these two types of AR glasses at two different viewing distances, 30 cm and 60 cm, and in two modes, purely virtual VR mode and virtual-real integration AR mode. The results show that the light field AR glasses led to a significantly faster reaction speed and higher accuracy than the conventional AR glasses at 30 cm in the AR mode. The participant feedback also shows that the former led to better virtual-real integration. User performance and experience of the light field AR glasses remained consistent across different viewing distances. Although the conventional AR glasses had a better search efficiency than the light field AR glasses at 60 cm in both AR and VR modes, it had more negative feedback from the participants. Overall, the design of this experiment successfully allows us to quantify the effect of VAC and underscores the strength of the evaluation method. Wei-An Teng, Su-Ling Yeh, Homer H. Chen |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Super-Resolution for Near-Eye Light Field Display in Fourier SpaceabstractNear-eye light field displays are superior to conventional AR displays because they offer continuous focus and viewing experiences free from visual accommodation conflict (VAC). However, given a fixed number of pixels for representation of the spatio-angular information of a light field, the inherent tradeoff between angular and spatial resolutions presents a great challenge to the widespread adoption of light field technology. To address the challenge, we propose a hybrid super-resolution framework consisting of a digital neural network and an optical neural network and allowing an end-to-end optimization of the downsampling operation for fitting the light field data into a fixed-resolution display panel and the upsampling operation for enhancing the light field quality, all in the frequency domain. Experimental results show that the proposed hybrid framework is a promising approach to quality enhancement of near-eye light field displays. Yu-Hsiang Huang, Homer H. Chen |
ICIP | 3 |
| 2024 | Masked face recognition using domain adaptation
Yu-Chieh Huang, David Akas Bedjo Rahardjo, Ren-Hau Shiue, Homer H. Chen |
Pattern Recognit. | 4 |
| 2024 | Training With Uncertain Annotations for Semantic Segmentation of Basal Cell Carcinoma From Full-Field OCT ImagesabstractSemantic segmentation of basal cell carcinoma (BCC) from full-field optical coherence tomography (FF-OCT) images of human skin has received considerable attention in medical imaging. However, it is challenging for dermatopathologists to annotate the training data due to OCT's lack of color specificity. Very often, they are uncertain about the correctness of the annotations they made. In practice, annotations fraught with uncertainty profoundly impact the effectiveness of model training and hence the performance of BCC segmentation. To address this issue, we propose an approach to model training with uncertain annotations. The proposed approach includes a data selection strategy to mitigate the uncertainty of training data, a class expansion to consider sebaceous gland and hair follicle as additional classes to enhance the performance of BCC segmentation, and a self-supervised pre-training procedure to improve the initial weights of the segmentation model parameters. Furthermore, we develop three post-processing techniques to reduce the impact of speckle noise and image discontinuities on BCC segmentation. The mean Dice score of BCC of our model reaches 0.503±0.003, which, to the best of our knowledge, is the best performance to date for semantic segmentation of BCC from FF-OCT images. Li-Wei Fu, Chih-Hao Liu, Manu Jain, Chih-Shan Jason Chen, Yu-Hung Wu, Sheng-Lung Huang, Homer H. Chen |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Endoscopic Feature Enhancement for Stomach 3D Reconstruction without DyeingabstractIn the past decades, endoscopic 3D reconstruction has become increasingly important for medical diagnosis, computer-aided surgeries, and minimally invasive procedures. Many previous studies have been able to reconstruct the complete stomach shape using the structure-from-motion approach; however, such methods require real or virtual indigo carmine (IC) dyeing to enhance endoscopic images before the 3D reconstruction takes place. In this study, we propose an alternative approach by enhancing the appearance of gastric tissue texture and reducing the effect of noise. We show that it is possible to perform 3D reconstruction using monocular endoscopy and without the need for chromoendoscopy. In addition, we show that our method can increase the number of feature matches and maintain the feature consistency, which is critical to the performance of stomach 3D reconstruction. Ting-Yu Wei, Ming-Lun Han, Wei-Chih Liao, Kuang-Chen Yen, Shyh-Jye Chen, Homer H. Chen |
ICIP | 6 |
| 2023 | Perceptual Tolerance of Split-Up Effect for Near-Eye Light Field DisplayabstractWhen the light field of a scene is generated with a finite number of subviews, the defocused regions would appear to be split-up if a camera is used to capture the light field. Yet, the split-up effect is unnoticeable when the light field is viewed directly through a human eye. In this paper, we attribute the unobservability of the split-up effect to the decrease in visual acuity as a function of retinal eccentricity and to the low-pass filtering property of visual attention. Theoretical and experimental results are provided to support our claim. Furthermore, we set an observability criterion for the split-up effect and discuss design strategies for performance improvement of light field displays. Ting-Hsun Chi, Wen Perng, Homer H. Chen |
ISMAR | 3 |
| 2023 | Time-Division Multiplexing Light Field Display With Learned Coded ApertureabstractConventional stereoscopic displays suffer from vergence-accommodation conflict and cause visual fatigue. Integral-imaging-based displays resolve the problem by directly projecting the sub-aperture views of a light field into the eyes using a microlens array or a similar structure. However, such displays have an inherent trade-off between angular and spatial resolutions. In this paper, we propose a novel coded time-division multiplexing technique that projects encoded sub-aperture views to the eyes of a viewer with correct cues for vergence-accommodation reflex. Given sparse light field sub-aperture views, our pipeline can provide a perception of high-resolution refocused images with minimal aliasing by jointly optimizing the sub-aperture views for display and the coded aperture pattern. This is achieved via deep learning in an end-to-end fashion by simulating light transport and image formation with Fourier optics. To our knowledge, this work is among the first that optimize the light field display pipeline with deep learning. We verify our idea with objective image quality metrics (PSNR, SSIM, and LPIPS) and perform an extensive study on various customizable design variables in our display pipeline. Experimental results show that light fields displayed using the proposed technique indeed have higher quality than that of baseline display designs. Chun-Hao Chao, Chang-Le Liu, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2023 | Acquiring 360° Light Field by a Moving Dual-Fisheye CameraabstractIn this paper, we propose an efficient deep learning pipeline for light field acquisition using a back-to-back dual-fisheye camera. The proposed pipeline generates a light field from a sequence of 360° raw images captured by the dual-fisheye camera. It has three main components: a convolutional network (CNN) that enforces a spatiotemporal consistency constraint on the subviews of the 360° light field, an equirectangular matching cost that aims at increasing the accuracy of disparity estimation, and a light field resampling subnet that produces the 360° light field based on the disparity information. Ablation tests are conducted to analyze the performance of the proposed pipeline using the HCI light field datasets with five objective assessment metrics (MSE, MAE, PSNR, SSIM, and GMSD). We also use real data obtained from a commercially available dual-fisheye camera to quantitatively and qualitatively test the effectiveness, robustness, and quality of the proposed pipeline. Our contributions include: 1) a novel spatiotemporal consistency loss that enforces the subviews of the 360° light field to be consistent, 2) an equirectangular matching cost that combats severe projection distortion of fisheye images, and 3) a light field resampling subnet that retains the geometric structure of spherical subviews while enhancing the angular resolution of the light field. I-Chan Lo, Homer H. Chen |
IEEE Trans. Image Process. | 2 |
| 2022 | Face Recognition for Fisheye ImagesabstractFace recognition often suffers from severe degradation in accuracy when applied to images captured by fisheye cameras. One way to resolve the issue is to rectify the fisheye images before classification; however, it can only achieve local optimum. In this paper, we present an end-to-end model with global optimum. To tackle the challenges due to intra-class variance and diversity of fisheye transformations, we propose 1) a structural correction to guide the model learning, and 2) a spatial-transformer-networks embedded model to compensate for the non-linear distortion of fisheye lenses. We test the proposed model on the CelebA dataset and a real image dataset and achieve an average accuracy of 98.7% and 98.0%, respectively, which represent improvements of 4.05% and 5.72% over the state-of-the-art results. Yi-Cheng Lo, Chiao-Chun Huang, Yueh-Feng Tsai, I-Chan Lo, An-Yeu Wu, Homer H. Chen |
ICIP | 6 |
| 2022 | Deep face recognition for dim images
Homer H. Chen |
Pattern Recognit. | 2 |
| 2022 | Efficient and Accurate Stitching for 360° Dual-Fisheye Images and VideosabstractBack-to-back dual-fisheye cameras are the most cost-effective devices to capture 360° visual content. However, image and video stitching for such cameras often suffer from the effect of fisheye distortion, photometric inconsistency between the two views, and non-collocated optical centers. In this paper, we present algorithms for geometric calibration, photometric compensation, and seamless stitching to address these issues for back-to-back dual-fisheye cameras. Specifically, we develop a co-centric trajectory model for geometric calibration to characterize both intrinsic and extrinsic parameters of the fisheye camera to fifth-order precision, a photometric correction model for intensity and color compensation to provide efficient and accurate local color transfer, and a mesh deformation model along with an adaptive seam carving method for image stitching to reduce geometric distortion and ensure optimal spatiotemporal alignment. The stitching algorithm and the compensation algorithm can run efficiently for 1920×960 images. Quantitative evaluation of geometric distortion, color discontinuity, jitter, and ghost artifact of the resulting image and video shows that our solution outperforms the state-of-the-art techniques. I-Chan Lo, Kuang-Tsu Shih, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2021 | Robust Light Field Synthesis From Stereo Images With Left-Right Geometric ConsistencyabstractWe propose a lightweight yet effective deep learning pipeline for light field synthesis from a single stereo image pair. Our pipeline consists of a convolutional network (CNN) that enforces a left-right consistency constraint on the light fields synthesized from left and right stereo views, a stage that merges light fields synthesized from left and right stereo views with a novel alpha blending technique, and a final refinement network using a unique 3D convolution operation. Our experiments quantitatively and qualitatively confirm the effectiveness and robustness of the proposed model, which performs favorably against state-of-the-art algorithms for light field synthesis from extremely sparse (only one, two, or four) views while using much fewer parameters. Chun-Hao Chao, Chang-Le Liu, Homer H. Chen |
ICIP | 3 |
| 2021 | Deep Face Rectification for 360° Dual-Fisheye CamerasabstractRectilinear face recognition models suffer from severe performance degradation when applied to fisheye images captured by 360° back-to-back dual fisheye cameras. We propose a novel face rectification method to combat the effect of fisheye image distortion on face recognition. The method consists of a classification network and a restoration network specifically designed to handle the non-linear property of fisheye projection. The classification network classifies an input fisheye image according to its distortion level. The restoration network takes a distorted image as input and restores the rectilinear geometric structure of the face. The performance of the proposed method is tested on an end-to-end face recognition system constructed by integrating the proposed rectification method with a conventional rectilinear face recognition system. The face verification accuracy of the integrated system is 99.18% when tested on images in the synthetic Labeled Faces in the Wild (LFW) dataset and 95.70% for images in a real image dataset, resulting in an average accuracy improvement of 6.57% over the conventional face recognition system. For face identification, the average improvement over the conventional face recognition system is 4.51%. Yi-Hsin Li, I-Chan Lo, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2021 | Enhancement and Speedup of Photometric Compensation for Projectors by Reducing Inter-Pixel Coupling and Calibration PatternsabstractFor a procam to preserve the color appearance of an image projected on a color surface, the photometric distortion introduced by the color surface has to be properly compensated. The performance of such photometric compensation relies on an accurate estimation of the projector nonlinearity. In this paper, we improve the accuracy of projector nonlinearity estimation by taking inter-pixel coupling into consideration. In addition, to respond quickly to the change of projection area due to projector movement, we reduce the number of calibration patterns from six to one and use the projected image as the calibration pattern. This greatly improves the computational efficiency of re-calibration that needs to be performed on the fly during a multimedia presentation without breaking its continuity. Both objective and subjective results are provided to illustrate the effectiveness of the proposed method for color compensation. Kuang-Tsu Shih, Jen-Shuo Liu, Frank Shyu, Homer H. Chen |
IEEE Trans. Image Process. | 4 |
| 2021 | Efficient Face Detection in the Fisheye Image DomainabstractSignificant progress has been made for face detection from normal images in recent years; however, accurate and fast face detection from fisheye images remains a challenging issue because of serious fisheye distortion in the peripheral region of the image. To improve face detection accuracy, we propose a light-weight location-aware network to distinguish the peripheral region from the central region in the feature learning stage. To match the face detector, the shape and scale of the anchor (bounding box) is made location dependent. The overall face detection system performs directly in the fisheye image domain without rectification and calibration and hence is agnostic of the fisheye projection parameters. Experiments on Wider-360 and real-world fisheye images using a single CPU core indeed show that our method is superior to the state-of-the-art real-time face detector RFB Net. Cheng-Yun Yang, Homer H. Chen |
IEEE Trans. Image Process. | 2 |
| 2021 | Tag Propagation and Cost-Sensitive Learning for Music Auto-TaggingabstractThe performance of music auto-tagging depends on the quality of training data. In practice, the links between songs and tags in the manually labeled training data can be incorrect (false positive) or missing (false negative). In this paper, we propose a cost-sensitive tag propagation learning method to improve auto-tagging. Specifically, we exploit music context to determine similar songs and propagate tags between them. Both propagated tags and original tags are used to optimize the auto-tagging models, and cost-sensitivity is incorporated into the loss function to enhance the robustness by adjusting the weight of relevant (positive) links with respect to irrelevant (negative) links. The proposed method is tested on three auto-tagging models: 2D-CNN, CRNN, and SampleCNN. The Million Song Dataset is used for training, and four music contexts, artist, playlist, tag, and listener, are used for song similarity measurement. The experimental results show 1) The proposed method can successfully improve the performance of the three auto-tagging models, 2) The cost-sensitive loss function helps reduce the impact of missing tags, and 3) The artist music context is more powerful for tag propagation than the other three music contexts. Yi-Hsun Lin, Homer H. Chen |
IEEE Trans. Multim. | 2 |
| 2020 | Face Recognition Under Low Illumination Via Deep Feature Reconstruction NetworkabstractRecent major benchmarks show that deep-learning-based face recognition can achieve superb performance, even surpassing human capability. However, many state-of-the-art face recognition models suffer from severe performance degradation for images captured under low illumination. The issue can be addressed by enhancing the illumination of face images before performing face recognition. In this paper, we evaluate such enhancement methods and, based on the findings, propose a novel feature reconstruction network to make face features illumination-invariant by generating a feature image from both the raw face image and the illumination-enhanced face image. The performance of the proposed approach is tested on the Specs on Faces (SoF) dataset. The overall verification accuracy is improved by 0.5% to 2.5% and the rank-l identification accuracy is improved by 2.1%. Homer H. Chen |
ICIP | 2 |
| 2020 | Photometric Consistency For Dual Fisheye CamerasabstractBack-to-back dual fisheye cameras are the most popular and cost-effective device to capture 360° images. However, inconsistent intensity and color between the pair of fisheye images often cause visible artifacts in the final 360° image. In this paper, we present a method to create photometric consistency between the two fisheye images. Specifically, we propose a loss function for image intensity compensation and a local color transfer model for color correction. Experimental results show that our method is able to correct the photometric inconsistency between dual fisheye images for high-quality 360° imaging. I-Chan Lo, Kuang-Tsu Shih, Gwo-Hwa Ju, Homer H. Chen |
ICIP | 4 |
| 2020 | AF-Net: A Convolutional Neural Network Approach to Phase Detection AutofocusabstractIt is important for an autofocus system to accurately and quickly find the in-focus lens position so that sharp images can be captured without human intervention. Phase detectors have been embedded in image sensors to improve the performance of autofocus; however, the phase shift estimation between the left and right phase images is sensitive to noise. In this paper, we propose a robust model based on convolutional neural network to address this issue. Our model includes four convolutional layers to extract feature maps from the phase images and a fully-connected network to determine the lens movement. The final lens position error of our model is five times smaller than that of a state-of-the-art statistical PDAF method. Furthermore, our model works consistently well for all initial lens positions. All these results verify the robustness of our model. Chi-Jui Ho, Chin-Cheng Chan, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2020 | Light Field Synthesis by Training Deep Network in the Refocused Image DomainabstractLight field imaging, which captures spatial-angular information of light incident on image sensors, enables many interesting applications such as image refocusing and augmented reality. However, due to the limited sensor resolution, a trade-off exists between the spatial and angular resolutions. To increase the angular resolution, view synthesis techniques have been adopted to generate new views from existing views. However, traditional learning-based view synthesis mainly considers the image quality of each view of the light field and neglects the quality of the refocused images. In this paper, we propose a new loss function called refocused image error (RIE) to address the issue. The main idea is that the image quality of the synthesized light field should be optimized in the refocused image domain because it is where the light field is viewed. We analyze the behavior of RIE in the spectral domain and test the performance of our approach against previous approaches on both real (INRIA) and software-rendered (HCI) light field datasets using objective assessment metrics such as MSE, MAE, PSNR, SSIM, and GMSD. Experimental results show that the light field generated by our method results in better refocused images than previous methods. Chang-Le Liu, Kuang-Tsu Shih, Jiun-Woei Huang, Homer H. Chen |
IEEE Trans. Image Process. | 4 |
| 2019 | 360° Video Stitching for Dual Fisheye CamerasabstractBack-to-back dual fisheye camera configuration offers a cost-effective 360° video solution. However, the inherent parallax of this camera configuration and the unmatched scene along the seam between the two camera views present a great challenge to video stitching. To address the jittering and fragmented visual appearance issues, we present in this paper a robust method that preserves the geometric structure of the scene and enhances the stableness of the video along the temporal dimension. The proposed video stitching method entails a mesh deforming operation prior for the minimization of the geometric distortion and an adaptive seam carving operation that generates optimal spatial and temporal alignment. Experimental results show that our method can produce 360° videos without jitter and ghost artifact. I-Chan Lo, Kuang-Tsu Shih, Homer H. Chen |
ICIP | 3 |
| 2019 | Seamless Stitching Dual Fisheye Images For 360° Free ViewabstractWe demonstrate a seamless stitching technology for dual-fisheye video cameras that combines two images/frames of ultra-wide field of view into a 360° image/frame. Our technology generates high-quality 360° free view without the disjoint appearance and ghost artifact that plague most 360° cameras in the market today. Our technology is efficient, robust, and stunning. Target customers include VR device manufactures, telco operators, broadcasters, surveillance service providers, and immersive free view users. I-Chan Lo, Kuang-Tsu Shih, Po Chin Yu, Chun-Ting Hung, Minghuang Shih, Makoto Odamaki, Homer H. Chen |
ICIP | 7 |
| 2018 | Depth from GazeabstractEye trackers are found on various electronic devices. In this paper, we propose to exploit the gaze information acquired by an eye tracker for depth estimation. The data collected from the eye tracker in a fixation interval are used to estimate the depth of a gazed object. The proposed method can be used to construct a sparse depth map of an augmented reality space. The resulting depth map can be applied to, for example, controlling the visual information displayed to the viewer. A mathematical model for determining whether two depths in the augmented reality space are statistically distinguishable is also developed. Experimental results show that the proposed method can estimate and distinguish different object depths effectively. Tzu-Sheng Kuo, Kuang-Tsu Shih, Sheng-Lung Chung, Homer H. Chen |
ICIP | 4 |
| 2018 | Image Stitching for Dual Fisheye CamerasabstractPanoramic photography creates stunning immersive visual experiences for viewers. In this paper, we investigate how to seamlessly stitch a pair of images captured by two uncalibrated, back-to-back, 195-degree fisheye cameras to generate a surround view of a 3D scene. It is a challenging task because the two camera centers are displaced and because the common region is the most distorted area. To enhance the robustness of feature matching and hence the quality of stitching, we propose a novel technique that projects the image rectilinearly onto an equirectangular plane. Unlike most previous global warping methods that are sensitive to the disparity between back-to-back fisheye images, our method employs local warping for image alignment while preserving the geometric structure of the scene. Experimental results show that our method effectively produces high-quality seamless panoramic images without stitching artifacts. I-Chan Lo, Kuang-Tsu Shih, Homer H. Chen |
ICIP | 3 |
| 2018 | Cross-Cultural Music Emotion Recognition by Adversarial Discriminative Domain AdaptationabstractAnnotation of the perceived emotion of a music piece is required for an automatic music emotion recognition system. Most music emotion datasets are developed for Western pop songs. The problem is that a music emotion recognizer trained on such datasets may not work well for non-Western pop songs due to the differences in acoustic characteristics and emotion perception that are inherent to cultural background. The problem was also found in cross-cultural and cross-dataset studies; however, little has been done to learn how to adapt a model pre-trained on a source music genre to a target music genre of interest. In this paper, we propose to address the problem by an unsupervised adversarial domain adaptation method. It employs neural network models to make the target music indistinguishable from the source music in a learned feature representation space. Because emotion perception is multifaceted, three types of input feature representations related to timbre, pitch, and rhythm are considered for performance evaluation. The results show that the proposed method effectively improves the prediction of the valence of Chinese pop songs from a model trained for Western pop songs. Yi-Wei Chen, Yi-Hsuan Yang, Homer H. Chen |
ICMLA | 3 |
| 2018 | Analysis of Disparity Error for Stereo AutofocusabstractAs more and more stereo cameras are installed on electronic devices, we are motivated to investigate how to leverage disparity information for autofocus. The main challenge is that stereo images captured for disparity estimation are subject to defocus blur unless the lenses of the stereo cameras are at the in-focus position. Therefore, it is important to investigate how the presence of defocus blur would affect stereo matching and, in turn, the performance of disparity estimation. In this paper, we give an analytical treatment of this fundamental issue of disparity-based autofocus by examining the relation between image sharpness and disparity error. A statistical approach that treats the disparity estimate as a random variable is developed. Our analysis provides a theoretical backbone for the empirical observation that, regardless of the initial lens position, disparity-based autofocus can bring the lens to the hill zone of the focus profile in one movement. The insight gained from the analysis is useful for the implementation of an autofocus system. Cheng-Chieh Yang, Shao-Kang Huang, Kuang-Tsu Shih, Homer H. Chen |
IEEE Trans. Image Process. | 4 |
| 2018 | Extracting Blood Vessels From Full-Field OCT Data of Human Skin by Short-Time RPCAabstractRecent advances in optical coherence tomography (OCT) lead to the development of OCT angiography to provide additional helpful information for diagnosis of diseases like basal cell carcinoma. In this paper, we investigate how to extract blood vessels of human skin from full-field OCT (FF-OCT) data using the robust principal component analysis (RPCA) technique. Specifically, we propose a short-time RPCA method that divides the FF-OCT data into segments and decomposes each segment into a low-rank structure representing the relatively static tissues of human skin and a sparse matrix representing the blood vessels. The method mitigates the problem associated with the slow-varying background and is free of the detection error that RPCA may have when dealing with FF-OCT data. Both short-time RPCA and RPCA methods can extract blood vessels from FF-OCT data with heavy speckle noise, but the former takes only half the computation time of the latter. We evaluate the performance of the proposed method by comparing the extracted blood vessels with the ground truth vessels labeled by a dermatologist and show that the proposed method works equally well for FF-OCT volumes of different quality. The average F-measure improvements over the correlation-mapping OCT method, the modified amplitude-decorrelation OCT angiography method, and the RPCA method, respectively, are 0.1835, 0.1032, and 0.0458. Pin-Hsien Lee, Chin-Cheng Chan, Sheng-Lung Huang, Andrew Chen 0002, Homer H. Chen |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Enhancement of phase detection for autofocusabstractPhase detection autofocus (PDAF) is a technique that uses sensors on left and right pixels to determine the relative position between the object and the focal plane. When an image is out of focus, a shift between the assembled left and right pixels is resulted, which can be used to determine the lens movement during the autofocus process. However, the presence of noise and blur often affects the accuracy of shift estimation and hence the performance of autofocus. In this paper, we propose a method to enhance phase detection for such sensors. A Gaussian filter is applied to overcome the adverse effect of image noise and blur. Experiments are conducted to show that the proposed method leads to a more accurate estimate of lens movement than the traditional phase correlation. Chin-Cheng Chan, Shao-Kang Huang, Homer H. Chen |
ICIP | 3 |
| 2017 | Enhancing the perception of a hazy visual world using a see-through head-mounted deviceabstractHuman perception of the visual world in the presence of fog or haze usually suffers from loss of color vibrance and details. In this paper we propose a novel method to provide a see-through head-mounted display with image dehazing capability. The proposed method projects an auxiliary image from the head-mounted device to the user's retina. The introduction of the auxiliary image enhances the contrast and brightness of the perceived visual world. This method can be applied to tourism, military, and navigation to overcome poor visual condition. Kai-En Lin, Kuang-Tsu Shih, Homer H. Chen |
ICIP | 3 |
| 2017 | Performance analysis of reconstruction-based super-resolution for camera arraysabstractIn this paper, we look into the possibility of combining multiple images captured by a camera array into a high-resolution image by reconstruction-based super-resolution. Specifically, we analyze the effect of camera array parameters, including the number of cameras, pixel size, and sampling interval, on the quality of the super-resolved image. Our experimental results show that the pixel size of the cameras is critical to the success of resolution enhancement and that only limited resolution gain is achievable if the pixel of the cameras is not sufficiently small. Our analysis serves as a guide to the design of high-resolution camera arrays. Kuang-Tsu Shih, Homer H. Chen |
ICIP | 2 |
| 2017 | Component Tying for Mixture Model Adaptation in Personalization of Music Emotion RecognitionabstractPersonalizing a music emotion recognition model is needed because the perception of music emotion is highly subjective, but it is a time-consuming process. In this paper, we consider how to expedite the personalization process that begins with a general model trained offline using a general user base and progressively adapts the model to a music listener using the emotion annotations of the listener. Specifically, we focus on reducing the number of user annotations needed for the personalization. We investigate and evaluate four component tying methods: single group tying, quadrantwise tying, hierarchical tying, and random tying. These methods aim to exploit the available annotations by identifying related model parameters on-the-fly and updating them jointly. In the evaluation, we use the AMG1608 dataset, which contains the clip-level valence-arousal emotion ratings of 1608 30-s music clips annotated by 665 listeners. Also, we use the acoustic emotion Gaussians model as the general model that uses a mixture of Gaussian components to learn the mapping between the acoustic feature space and the emotion space. The results show that the model adaptation with component tying requires only 10-20 personal annotations to obtain the same level of prediction accuracy as the baseline model adaptation method that uses 50 personal annotations without component tying. Yu-An Chen, Ju-Chiang Wang, Yi-Hsuan Yang, Homer H. Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Radiometric Compensation of Images Projected on Non-White Surfaces by Exploiting Chromatic Adaptation and Perceptual AnchoringabstractFlat surfaces in our living environment to be used as replacements of a projection screen are not necessarily white. We propose a perceptual radiometric compensation method to counteract the effect of color projection surfaces on image appearance. It reduces color clipping while preserving the hue and brightness of images based on the anchoring property of human visual system. In addition, it considers the effect of chromatic adaptation on perceptual image quality and fixes the color distortion caused by non-white projection surfaces by properly shifting the color of the image pixels toward the complementary color of the projection surface. User ratings show that our method outperforms the existing methods in 974 out of 1020 subjective tests. Tai-Hsiang Huang, Ting-Chun Wang, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2016 | Emotion-flow guided music accompaniment generationabstractThe emotion of a music piece varies as it unrolls in time. We develop a system that takes a melody and an expected emotion flow as input and automatically generates an accompaniment. The accompaniment is composed of chord progression and accompaniment pattern. The former is generated from melody and valence data through dynamic programming, and the latter from arousal data. A mathematical model is developed to describe the relation between valence and chord progression. The performance of the system is evaluated subjectively. The cross-correlation coefficient between the expected arousals and the perceived ones is 0.84, and the cross-correlation coefficient between the expected valences and the perceived ones is 0.52. Both coefficients exceed 0.90 for musician subjects. Yi-Chan Wu, Homer H. Chen |
ICASSP | 2 |
| 2016 | Empirical reliability analysis of disparity-based autofocusabstractConventional autofocus methods based on contrast detection are often unable to reliably decide the direction of initial lens movement. In this paper, we show that even using the disparity data obtained from blurry stereo images can effectively solve the problem. This approach is developed for stereo cameras with adjustable focal distance. Such stereo cameras provide sharp images over a wide range of object distance. The disparity-based autofocus approach allows the lens to instantly come to a position near the in-focus position or inside the rising zone of the focus profile. We investigate the impact of disparity variation on the first lens movement driven by disparity-based autofocus and suggest how it can be integrated with contrast detection autofocus. Shao-Kang Huang, Cheng-Chieh Yang, Homer H. Chen |
ICIP | 3 |
| 2016 | Blood vessel extraction from OCT data by short-time RPCAabstractOptical coherence tomography (OCT) is a medical imaging technology that allows for non-invasive diagnosis of diseases in the early stage. Because blood flow anomalies provide useful information for many diseases, we develop an automatic blood vessel detection algorithm based on the robust principle component analysis (RPCA) technique. Specifically, we propose a short-time RPCA method that divides an OCT volume into segments and decomposes each segment into a low-rank structure representing relatively static tissues and a sparse matrix representing the blood vessels. It efficiently extracts blood vessel structure from OCT data by distinguishing static and dynamic components. This work serves as the foundation for further blood flow analysis. Pin-Hsien Lee, Chin-Cheng Chan, Sheng-Lung Huang, Andrew Chen 0002, Homer H. Chen |
ICIP | 5 |
| 2016 | Gaussian noise approximation for disparity-based autofocusabstractAn analytical characterization of the accuracy of disparity-based autofocus requires an effective error model of disparity estimation. In this paper, we investigate the approximation of photon shot noise of images by a Gaussian model for disparity error analysis that takes defocus blur and image noise into account. We show that, counterintuitively, defocus blur alone does not affect the disparity estimation. However, its presence makes disparity estimation more vulnerable to image noise. In addition, we show that the Gaussian model is a fair approximation of the photon shot noise for disparity-based autofocus. The Gaussian model allows defocus blur to be mathematically separable from noise and facilitates further analysis of disparity error for disparity-based autofocus. Cheng-Chieh Yang, Homer H. Chen |
ICIP | 2 |
| 2016 | Generation of Affective Accompaniment in Accordance With Emotion FlowabstractThe emotion expressed by a music piece varies as the music unrolls in time. To create such dynamic expression, we develop an algorithm that automatically generates the accompaniment for a melody according to the emotion flow specified by a user. The emotion flow is given in the form of arousal and valence curves, each as a function of time. The affective accompaniment is composed of chord progression and accompaniment pattern. The chord progression, which controls the valence of the composed music, is generated by dynamic programming using the input melody and valence data as constraints. A mathematical model is developed to describe the temporal relationship between valence and chord progression. The accompaniment pattern, which controls the arousal of the composed music, is determined according to the quantized arousal values. The performance of the system is evaluated by subjective tests. The cross-correlation coefficient between the input arousal (valence) and the perceived arousal (valence) of the composed music is 0.85 (0.52). It rises to 0.92 for arousal and 0.88 for valence if only musician subjects in the test are considered. Overall, the proposed system is capable of generating subjectively appropriate accompaniments conforming to the user specification. Yi-Chan Wu, Homer H. Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Efficient Quantization Based on Rate-Distortion Optimization for Video CodingabstractMost rate-distortion (R-D) optimized quantization methods of video coding involve an exhaustive search process to determine the optimal quantized transform coefficients of a coding block and are computationally more expensive than the conventional quantization. In this paper, we present a novel analytical method that directly solves the rate-distortion optimization problem in a closed form by employing a rate model for entropy coding. It has the appealing property of low complexity and is easy to implement. The results show that the proposed method is $4\times $ to $40\times $ faster than the previous methods and 3%-5% more efficient in bitrate than the H.264/AVC reference encoder, which uses the conventional quantization, for video coded with the IBBP structure of group of pictures (GOP) in the normal peak signal-to-noise ratio (PSNR) range (30-40 dB). Tsung-Yau Huang, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Focus Profile ModelingabstractA focus profile depicts the image sharpness (or focus value) as the lens sweeps along the optical axis of a camera. Accurate modeling of the focus profile is important to many imaging tasks. In this paper, we present an approach to focus profile modeling that makes the search of in-focus lens position a mathematically tractable problem, and hereby improves the efficiency and accuracy of image acquisition. The proposed approach entails a transformation that converts the representation of a focus profile to quadratic form. An important feature of the approach is that no prior knowledge of the focus measurement technique is required. Experimental results are provided to demonstrate the effectiveness of the approach. Dong-Chen Tsai, Homer H. Chen |
IEEE Trans. Image Process. | 2 |
| 2016 | Exploiting Perceptual Anchoring for Color Image EnhancementabstractThe preservation of image quality under various display conditions becomes more and more important in the multimedia era. A considerable amount of effort has been devoted to compensating the quality degradation caused by dim LCD backlight for mobile devices and desktop monitors. However, most previous enhancement methods for backlight-scaled images only consider the luminance component and overlook the impact of color appearance on image quality. In this paper, we propose a fast and elegant method that exploits the anchoring property of human visual system to preserve the color appearance of backlight-scaled images as much as possible. Our approach is distinguished from previous ones in many aspects. First, it has a sound theoretical basis. Second, it takes the luminance and chrominance components into account in an integral manner. Third, it has low complexity and can process 720p high-definition videos at 35 frames per second without flicker. The superior performance of the proposed method is verified through psychophysical tests. Kuang-Tsu Shih, Homer H. Chen |
IEEE Trans. Multim. | 2 |
| 2016 | Blocking harmful blue light while preserving image color appearanceabstractRecent study in vision science has shown that blue light in a certain frequency band affects human circadian rhythm and impairs our health. Although applying a light blocker to an image display can block the harmful blue light, it inevitably makes an image look like an aged photo. In this paper, we show that it is possible to reduce harmful blue light while preserving the blue appearance of an image. Moreover, we optimize the spectral transmittance profile of blue light blocker based on psychophysical data and develop a color compensation algorithm to minimize color distortion. A prototype using notch filters is built as a proof of concept. Kuang-Tsu Shih, Jen-Shuo Liu, Frank Shyu, Su-Ling Yeh, Homer H. Chen |
ACM Trans. Graph. | 5 |
| 2015 | The AMG1608 dataset for music emotion recognitionabstractAutomated recognition of musical emotion from audio signals has received considerable attention recently. To construct an accurate model for music emotion prediction, the emotion-annotated music corpus has to be of high quality. It is desirable to have a large number of songs annotated by numerous subjects to characterize the general emotional response to a song. Due to the need for personalization of the music emotion prediction model to address the subjective nature of emotion perception, it is also important to have a large number of annotations per subject for training and evaluating a personalization method. In this paper, we discuss the deficiency of existing datasets and present a new one. The new dataset, which is publically available to the research community, is composed of 1608 30-second music clips annotated by 665 subjects. Furthermore, 46 subjects annotated more than 150 songs, making this dataset the largest of its kind to date. Yu-An Chen, Yi-Hsuan Yang, Ju-Chiang Wang, Homer H. Chen |
ICASSP | 4 |
| 2015 | Compensation of spectral mismatch to enhance WRGB demosaickingabstractThe presence of spectral mismatch between the components of a WRGB color filter array severely affects the performance of demosaicking. This paper presents a novel method that compensates for the spectral mismatch and greatly enhances the accuracy of R, G, and B interpolation. The method is tested on a two-megapixel WRGB CMOS image sensor. The results show that the proposed method brings the performance of WRGB demosaicking to an unprecedented level competitive with that of the state-of-the-art Bayer demosaicking. Po-Hsun Su, Po-Chang Chen, Homer H. Chen |
ICIP | 3 |
| 2015 | Transformation of focus profiles for digital autofocusabstractA focus profile model is used to estimate the in-focus lens position so that a sharp image or video can be captured. In this paper, we propose an approach that transforms a wide variety of focus profiles generated by different focus measurement techniques into an analytic form and thereby converts the search of in-focus lens position into a mathematically tractable problem, resulting in a wide effective range and efficient image acquisition. An important feature of the approach is that it does not require any knowledge of the focus measurement technique used in a camera. Experimental results, including demo videos, are provided to demonstrate the effectiveness of the approach. Dong-Chen Tsai, Homer H. Chen |
ICIP | 2 |
| 2015 | Vector representation of emotion flow for popular musicabstractThe flow of emotion expressed by music through time is a useful feature for music information indexing and retrieval. In this paper, we propose a novel vector representation of emotion flow for popular music. It exploits the repetitive verse-chorus structure of popular music and connects a verse (represented by a point) and its corresponding chorus (another point) in the valence-arousal emotion plane. The proposed vector representation visually gives users a snapshot of the emotion flow of a popular song in an intuitive and instant manner, more effective than the point and curve representations of music emotion flow. Because many other genres also have repetitive music structure, the vector representation has a wide range of applications. Chia-Hao Chung, Homer H. Chen |
MMSP | 2 |
| 2015 | Method and experiments of subliminal cueing for real-world images
Tai-Hsiang Huang, Su-Ling Yeh, Yung-Hao Yang, Hsin-I Liao, Ya-Yeh Tsai, Pai-Ju Chang, Homer H. Chen |
Multim. Tools Appl. | 7 |
| 2015 | Guest Editorial: Visual Information Processing and Perception
Hari Kalva, Homer H. Chen, Velibor Adzic, Gerardo Fernández-Escribano |
Multim. Tools Appl. | 2 |
| 2014 | Linear regression-based adaptation of music emotion recognition models for personalizationabstractPersonalization techniques can be applied to address the subjectivity issue of music emotion recognition, which is important for music information retrieval. However, achieving satisfactory accuracy in personalized music emotion recognition for a user is difficult because it requires an impractically huge amount of annotations from the user. In this paper, we adopt a probabilistic framework for valence-arousal music emotion modeling and propose an adaptation method based on linear regression to personalize a background model in an online learning fashion. We also incorporate a component-tying strategy to enhance the model flexibility. Comprehensive experiments are conducted to test the performance of the proposed method on three datasets, including a new one created specifically in this work for personalized music emotion recognition. Our results demonstrate the effectiveness of the proposed method. Yu-An Chen, Ju-Chiang Wang, Yi-Hsuan Yang, Homer H. Chen |
ICASSP | 4 |
| 2014 | Analysis of the effect of calibration error on light field super-resolution renderingabstractLight field photography, which has recently drawn considerable attention, provides novel functionalities such as refocusing and depth estimation at the same time. However, the resolution of the rendered refocus image is incomparably lower than the number of light ray samples in the light field. In this paper, we show that super-resolution can be performed to bridge the gap and that deconvolution, which is missing in most previous methods, is essential to the success of light field super-resolution rendering. In addition, we investigate the effect of four different kinds of camera calibration error on the quality of rendered images. Given an expected level of image quality, the upperbound of the camera calibration error is analyzed. Kuang-Tsu Shih, Chen-Yu Hsu 0001, Cheng-Chieh Yang, Homer H. Chen |
ICASSP | 4 |
| 2014 | Anti-aliasing for light field renderingabstractGhosting artifact resulted from angular undersampling is a major issue of light field rendering. In this paper, we propose an anti-aliasing method that compensates the effect of undersampling by making use of depth information. It can be incorporated into most existing refocusing algorithms to render high-quality images from light field. The method yields proper anti-aliasing across all depth range. Mathematical derivation of the method and the physical intuition behind the idea are described. Results are shown to demonstrate the performance of the method compared to previous ones. An-Cheng Chang, Tzu-Pin Sung, Kuang-Tsu Shih, Homer H. Chen |
ICME | 4 |
| 2014 | PIVP 2014: First International Workshop on Perception Inspired Video ProcessingabstractThis is a MM'14 Workshop Summary Abstract for PIVP'14 - 1st International Workshop on Perception Inspired Video Processing. Workshop provided a venue for researchers involved with perceptual video processing and coding to present their current work and to discuss future directions. The workshop program consisted of a keynote talk, oral presentations of full papers and interactive poster session for the short papers. All participants took part in the discussion panel at the end of workshop, exchanging ideas and suggestions for future work. Hari Kalva, Homer H. Chen, Gerardo Fernández-Escribano, Velibor Adzic |
ACM Multimedia | 2 |
| 2014 | Music recommendation based on artist novelty and similarityabstractMost existing systems recommend songs to the user based on the popularity of songs and singers. However, the system proposed in this paper is driven by an emerging and somewhat different need in the music industry-promoting new talents. The system recommends songs based on the novelty of singers (or artists) and their similarity to the user's favorite artists. Novel artists whose popularity is on the rise have a higher priority to be recommended. Specifically, given a user's favorite artists, the system first determines the candidate artists based on their similarity with the favorite artists and then selects those who have a higher novelty score than the favorite artists. Then, the system outputs a playlist composed of the most popular songs of the selected artists. The proposed system can be integrated into most existing systems. Its performance is evaluated using the Spotify Radio Recommender as a reference and a pool of 100 subjects recruited on campus. Experimental results show that our system achieves a high novelty score and a competitive user-preference score. Ning Lin, Ping-Chia Tsai, Yu-An Chen, Homer H. Chen |
MMSP | 4 |
| 2013 | Extraction and alignment evaluation of motion beats for street danceabstractThe coordination between the dancer's movement and the accompaniment of music is an important element of dance performance. In this paper, we propose a system for coordination evaluation of street dance. Given a dance video clip as input, the system first extracts motion beats from the video and then measures how well the motion beats correlate with the music beats. The motion beats are obtained by analyzing the speed and the change of direction of the dancer's movement. Unlike most previous work, which mainly focuses on 2D motion trajectory analysis, our system provides a more efficient and accurate dance movement analysis by using the 3D joint data of the dancer acquired by Kinect. Another distinction of our system is that, for beat correlation, it considers not only the underlying steady beat of music but also the groove pattern that gives the propulsive rhythmic feel of the music. We believe the differentiation of music beats into these two categories leads to a finer dance coordination evaluation. The test video clips for performance evaluation are generated by professional dancers. The average F-score of our system is about 80%. Chieh Ho, Wei-Tze Tsai, Keng-Sheng Lin, Homer H. Chen |
ICASSP | 4 |
| 2013 | Singing voice timbre classification of Chinese popular musicabstractSinging voice plays an important role in the listening experience of music. In this paper, we propose to classify popular music by the timbre quality of the singing voice. Specifically, we adopt six singing voice timbre classes as the taxonomy and build a new data set, KKTIC, that contains the expert annotations of 387 Chinese popular songs. To build an automatic classifier, we resort to signal processing and machine learning techniques and extract a number of singing voice-related features such as vibrato and harmonic-to-noise ratio. We also propose the use of vocal segment detection and singing voice separation as preprocessing steps. Our evaluation identifies the relevant acoustic features and validates the importance of these preprocessing steps. The accuracy in timbre classification reaches 79.84% in a five-fold stratified cross validation. Cheng-Ya Sha, Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
ICASSP | 4 |
| 2013 | Radiometric compensation for procam system based on anchoring theoryabstractFor a procam system consisting of a projector and a camera, the chroma distortion of an image projected on a non-white surface can be compensated by increasing the intensity of the complementary color of the projection surface. However, this inevitably trades image brightness for chroma correctness because the required intensity often outreaches the projector's dynamic range. In this paper, we propose a solution for the compensation problem using the anchoring theory of lightness perception. A notable feature of our technique is that it takes the chromatic adaptation and viewing condition into account, which leads to a better characterization of the image quality perceived by human eyes. As a result, the compensated image appears much closer to the original image. A computational framework for the tradeoff between brightness and chroma distortion is derived, and the performance of the proposed algorithm is verified by subjective experiments. Ting-Chun Wang, Tai-Hsiang Huang, Homer H. Chen |
ICIP | 3 |
| 2013 | AUtomatic accompaniment generation to evoke specific emotionabstractA music piece consists of melody and accompaniment in most genres. In this paper, we present a system to automatically generate accompaniment that evokes specific emotions for a given melody. In particular, we propose harmonic progression and onset rate as two key features for emotion-based accompaniment generation. The former refers to the progression of chords, and the latter refers to the number of music events (such as notes and drums) in a unit time. The harmonic progression and the onset rate are altered according to the specified emotion represented by the valence and arousal parameters, respectively. The performance of the system is evaluated subjectively, and the result shows a perfect positive Spearman correlation between the specified emotion and the perceived emotion. Pei-Chun Chen, Keng-Sheng Lin, Homer H. Chen |
ICME | 3 |
| 2013 | Acceleration of rate-distortion optimized quantization for H.264/AVCabstractRate-distortion optimized quantization improves the coding performance of video compression. However, the search process involved in most existing methods is computationally expensive. In this paper, we present a novel method for accelerating the rate-distortion optimized quantization process for an H.264/AVC video encoder. The acceleration is achieved by using a rate model of entropy coding to directly solve the rate-distortion optimization problem. Compared with the H.264/AVC reference encoder, our method achieves an average 5.6% bitrate reduction for IBBP GOP coding structure, with less than 0.1% increase of total encoding time. Tsung-Yau Huang, Chieh-Kai Kao, Homer H. Chen |
ISCAS | 3 |
| 2013 | Color enhancement based on the anchoring theoryabstractProviding consistent viewing experience across different reproduction conditions is an important issue for multimedia systems in the real world. In this paper, we propose a method to improve the perceptual quality of color reproduction for backlight-scaled images displayed on a liquid crystal display. Supported by the anchoring theory of lightness perception developed in psychology, this method is able to enhance the color appearance of images even when the backlight is only 5% of the original intensity. The goal is to make the appearance of the resulting images as close as possible to the original images illuminated with full backlight. The concept behind the proposed color enhancement method is general enough for many other applications. The effectiveness of the method is verified by subjective experiments. Kuang-Tsu Shih, Homer H. Chen |
MMSP | 2 |
| 2013 | Automatic highlights extraction for drama video using music emotion and human face features
Keng-Sheng Lin, Ann Lee 0002, Yi-Hsuan Yang, Cheng-Te Lee, Homer H. Chen |
Neurocomputing | 5 |
| 2013 | Enhancement of Backlight-Scaled ImagesabstractSwitching the liquid crystal display (LCD) backlight of a portable multimedia device to a low power level saves energy but results in poor image quality especially for the low-luminance image areas. In this paper, we propose an image enhancement algorithm that overcomes such effects of dim LCD backlight by taking the human visual property into consideration. It boosts the luminance of image areas below the perceptual threshold while preserving the contrast of the other image areas. We apply the just noticeable difference theory and decompose an image into an HVS response layer and a background luminance layer. The boosting and compression processes, which enhance the visibility of the low-luminance image areas, are carried out in the background luminance layer to avoid luminance gradient reversal and over-compensation. The contrast of the processed image is further enhanced by exploiting the Craik-O'Brein-Cornsweet visual illusion. Experimental results are provided to show the performance of the proposed algorithm. Tai-Hsiang Huang, Kuang-Tsu Shih, Su-Ling Yeh, Homer H. Chen |
IEEE Trans. Image Process. | 4 |
| 2013 | Emotional Accompaniment Generation System Based on Harmonic ProgressionabstractA music piece consists of melody and accompaniment in many genres. In this paper, we present a system to automatically generate accompaniment that evokes specific emotions for a given melody. In particular, we propose harmonic progression and onset rate as two key features for emotion-based accompaniment generation. The former refers to the progression of chords, and the latter refers to the number of music events (such as notes and drums) in a unit time. The harmonic progression and the onset rate are altered according to the specified emotion represented by the valence and arousal parameters, respectively. The performance of the system is evaluated subjectively, and the result shows a perfect positive Spearman correlation between the specified emotion and the perceived emotion. Pei-Chun Chen, Keng-Sheng Lin, Homer H. Chen |
IEEE Trans. Multim. | 3 |
| 2012 | Directing visual attention by subliminal cuesabstractUnconscious attention shift has been proven to be automatic, effortless, and unresisting; however, whether it is applicable to images of complex scene remains to be verified. We propose a novel method for directing human visual attention by flashing a short-duration visual stimulus before presenting the image. This method directs viewer's visual attention to an arbitrary location in the image without engaging the viewer's awareness. Unlike previous studies using simple stimuli (e.g. simple geometric figures, image of uniform color, or artificial white-noise), we use complex color images for guiding unconscious visual attention in real-world. To our best knowledge, this is the first study demonstrating that short-duration cues presented at arbitrary locations can attract human visual attention in a complex scene. The proposed method is useful for applications that require efficient and unresisting attention shift to the target image area. Tai-Hsiang Huang, Yung-Hao Yang, Hsin-I Liao, Su-Ling Yeh, Homer H. Chen |
ICIP | 5 |
| 2012 | Smooth control of continuous autofocusabstractFor a video camera to continuously focus on a moving object in the autofocus process, accurate estimate of the in-focus lens position is required in order to avoid the so-called bouncing phenomenon that is resulted from back-and-forth lens movements around the in-focus lens position. In this paper, we present a novel continuous focus method that effectively improves the accuracy of in-focus lens position estimate for dynamic scenes by a Kalman filter and provides good user experience by a smooth control of the lens movements in the autofocus process. Experimental results are shown to demonstrate the performance of our method. Dong-Chen Tsai, Homer H. Chen |
ICIP | 2 |
| 2012 | Quality enhancement of procam system by radiometric compensationabstractA procam system consists of a projector and a camera. In this paper, a radiometric compensation scheme is proposed for improving the projection quality of a procam system that uses a nearby non-white wall as the projection screen. The compensation scheme is capable of correcting the radiometric errors such as chroma and brightness distortions caused by the projection surface and the ambient light. It is also capable of compensating the effects of nonlinear spectral response (including vignetting) of the procam system. These important functions are achieved through a computational framework that optimizes the tradeoff between chroma and brightness distortions of the projected image and accelerates the computation of the penalty in the optimization process. Experimental results are shown to demonstrate the performance of the proposed radiometric compensation scheme. Tai-Hsiang Huang, Chen-Tai Kao, Homer H. Chen |
MMSP | 3 |
| 2012 | Reciprocal Focus ProfileabstractA focus profile having a steeper peak is more resistant to image noise in the autofocus (AF) process of a digital camera. However, a focus profile of such shape normally has a flatter out-of-focus region on either side of the profile, resulting in a slow AF process due to the lack of clue about where the lens should move when the lens is in such regions. To address the problem, we provide a statistical analysis of the focus profile and show that a strictly monotonic transformation of the focus profile preserves the accuracy of the AF. On the basis of this analysis, we propose a new focus profile representation that transforms the focus profile to the reciprocal domain in which the reciprocal focus profile is modeled by a polynomial function. This transformation makes the AF mathematically tractable and boosts the search speed. Experimental results are shown to demonstrate the advantage of the proposed representation. Dong-Chen Tsai, Homer H. Chen |
IEEE Trans. Image Process. | 2 |
| 2012 | Machine Recognition of Music Emotion: A ReviewabstractThe proliferation of MP3 players and the exploding amount of digital music content call for novel ways of music organization and retrieval to meet the ever-increasing demand for easy and effective information access. As almost every music piece is created to convey emotion, music organization and retrieval by emotion is a reasonable way of accessing music information. A good deal of effort has been made in the music information retrieval community to train a machine to automatically recognize the emotion of a music signal. A central issue of machine recognition of music emotion is the conceptualization of emotion and the associated emotion taxonomy. Different viewpoints on this issue have led to the proposal of different ways of emotion annotation, model training, and result visualization. This article provides a comprehensive review of the methods that have been proposed for music emotion recognition. Moreover, as music emotion recognition is still in its infancy, there are many open issues. We review the solutions that have been proposed to address these issues and conclude with suggestions for further research. Yi-Hsuan Yang, Homer H. Chen |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Multipitch Estimation of Piano Music by Exemplar-Based Sparse RepresentationabstractPitch, together with other midlevel music features such as rhythm and timbre, holds the promise of bridging the semantic gap between low-level features and high-level semantics for music understanding. This paper investigates the pitch estimation of a piano music signal by exemplar-based sparse representation. A note exemplar is a segment of a piano note, stored in the dictionary. We first describe how to represent a segment of the piano music signal as a linear combination of a small number of note exemplars from a large note exemplar dictionary and then show how the sparse representation problem can be solved by -regularized minimization. The proposed approach incorporates tuning factor estimation, note candidate selection, and hidden-Markov-model-based smoothing into the estimation process to improve accuracy. Unlike previous approaches, the proposed approach does not require retraining for a new piano. Instead, only a dozen notes of the new piano are needed. This feature is computationally attractive and avoids intense manual labeling. The system performance is evaluated using 70 classical music recordings of two real pianos under different recording conditions. The results show that the proposed system outperforms four state-of-the-art systems. Cheng-Te Lee, Yi-Hsuan Yang, Homer H. Chen |
IEEE Trans. Multim. | 3 |
| 2012 | Single Image Realism Assessment and Recoloring by Color CompatibilityabstractIn this paper, we investigate the assessment of image realism by focusing our attention on the color compatibility between an inserted object and the background in an image composite. We propose a technique based on two color compatibility properties to achieve realistic image composition. The first property is related to the color similarity, and the second one to the consistence of color tendency between image regions. We further propose algorithms based on these two properties for image realism assessment and recoloring. These algorithms only require information from the image to be tested, making them suitable for practical applications where real images are unavailable. Effectiveness of the algorithms is demonstrated through various images and verified by ground truth. Bing-Yi Wong, Kuang-Tsu Shih, Chia-Kai Liang, Homer H. Chen |
IEEE Trans. Multim. | 4 |
| 2011 | Enhancement of LCD Images Illuminated with Dim BacklightabstractDimming the LCD backlight saves the battery power of a portable multimedia device but causes image quality degradation. In this paper, we propose an algorithm to overcome the effect of dim backlight on image quality. This algorithm enhances the backlight-scaled image by exploiting the human visual property. It performs a two-layer image decomposition based on the just noticeable difference theory and enhances the backlight-scaled image by luminance boosting and compression. Experimental results are provided to demonstrate the performance of the proposed algorithm. Tai-Hsiang Huang, Su-Ling Yeh, Homer H. Chen |
ICCCN | 3 |
| 2011 | Local Dimming of Liquid Crystal Display Using Visual Attention Prediction ModelabstractLocal dimming of the LED backlight is a popular technique for saving the power of a liquid crystal display. This paper presents a novel approach that improves the performance of conventional local dimming algorithms by incorporating a visual attention prediction model in the backlight dimming process. The approach saves the energy of the liquid crystal display and, in the mean time, maintains the perceptual image quality. This is achieved by preserving the backlight luminance of the image areas that attract human attention while reducing that of the other image areas. Experimental results are provided to show the effectiveness of the proposed approach. Chia-Hang Lee, Wen-Hsiang Shaw, Hsin-I Liao, Su-Ling Yeh, Homer H. Chen |
ICCCN | 5 |
| 2011 | Fusion of visual attention cues by machine learningabstractA new computational scheme for visual attention modeling is proposed. It adopts both low-level and high-level features to predict visual attention from a video signal and fuses the features by using machine learning. We show that such a scheme is more robust than those using purely single level features. Unlike conventional techniques, our scheme is able to avoid perceptual mismatch between the estimated saliency and the actual human fixation. We show that selecting the representative training samples according to the fixation distribution improves the efficacy of regressive training. Experimental results are shown to demonstrate the advantages of the proposed scheme. Wen-Fu Lee, Tai-Hsiang Huang, Su-Ling Yeh, Homer H. Chen |
ICIP | 4 |
| 2011 | Effective autofocus decision using reciprocal focus profileabstractIn this paper, we propose a new autofocus approach that transforms the focus profile to the reciprocal domain. A key feature of this approach is that it makes the autofocus process mathematically tractable in both in-focus and out-of- focus regions, thereby allowing a digital camera or camcorder to expedite the search of the best lens position for capturing the image that has the maximum focus value. Experimental results are provided to illustrate the advantage of the proposed approach. Dong-Chen Tsai, Homer H. Chen |
ICIP | 2 |
| 2011 | Automatic transcription of piano music by sparse representation of magnitude spectraabstractAssuming that the waveforms of piano notes are pre-stored and that the magnitude spectrum of a piano signal segment can be represented as a linear combination of the magnitude spectra of the pre-stored piano waveforms, we formulate the automatic transcription of polyphonic piano music as a sparse representation problem. First, the note candidates of the piano signal segment are found by using heuristic rules. Then, the sparse representation problem is solved by l1-regularized minimization, followed by temporal smoothing the frame-level results based on hidden Markov models. Evaluation against three state-of-the-art systems using ten classical music recordings of a real piano is performed to show the performance improvement of the proposed system. Cheng-Te Lee, Yi-Hsuan Yang, Homer H. Chen |
ICME | 3 |
| 2011 | Quality Improvement of Video Codec by Rate-Distortion Optimized QuantizationabstractConventional quantization methods consider only the distortion between original and reconstructed video as the cost of compression. Considering the time-varying nature of network bandwidth for multimedia services, we believe a video coding system can provide a better quality of experience if it takes the bit rate of the compressed bit stream into consideration as well when optimizing the quantization. In this paper we present a rate-distortion optimization approach to the quantization of video coding. This approach is able to balance between rate and distortion for quantization and enhance the overall quality of the entire coding system, with only a slight increase in computational overhead. We implement this method in H.264/AVC, and the extensive experimental data obtained under various test conditions show that the performance of the R-D optimized quantization is indeed better than the H.264 reference software. Tsung-Yau Huang, Po-Yen Su, Chieh-Kai Kao, Tao-Sheng Ou, Homer H. Chen |
ISM | 5 |
| 2011 | Automatic highlights extraction for drama video using music emotion and human face featuresabstractThe rich emotion part of a drama video is often the center of attraction to the viewer. Emotion-based highlights extraction is useful for applications such as drama video retrieval and automatic trailer generation. In this paper, we propose a system that uses music emotion and human face as features for automatic extraction of the emotion highlights of a drama video. These high-level audiovisual features are used because music invokes emotion response from the viewer and characters express emotion on their faces. To avoid the interference of speech signal and environmental noise, a novel two-stage music emotion recognition scheme is developed. We first detect the presence of incidental music in a drama video using an audio fingerprint technique, and then perform emotion recognition on the noise-free music available from the album of the incidental music. This simple but effective approach greatly improves the accuracy of music emotion recognition. Besides the conventional subjective evaluation, we propose a new metric for quantitative performance evaluation of highlights extraction. Evaluation results are provided to illustrate the performance of the system. Keng-Sheng Lin, Ann Lee 0002, Yi-Hsuan Yang, Cheng-Te Lee, Homer H. Chen |
MMSP | 5 |
| 2011 | Recent progress on perceptual video codingabstractThe purpose of this technical demonstration is to show how the compressed video quality can be improved by taking the characteristics of human visual system into account in the rate distortion optimization and rate control processes of a video encoder. Both the R-D performance of the video encoder and the picture quality can be evaluated subjectively by the viewers in this demonstration to witness that perceptual video coding is indeed a promising direction for next generation video coding. Po-Yen Su, Yi-Hsin Huang, Tao-Sheng Ou, Homer H. Chen |
VCIP | 4 |
| 2011 | Ranking-Based Emotion Recognition for Music Organization and RetrievalabstractDetermining the emotion of a song that best characterizes the affective content of the song is a challenging issue due to the difficulty of collecting reliable ground truth data and the semantic gap between human's perception and the music signal of the song. To address this issue, we represent an emotion as a point in the Cartesian space with valence and arousal as the dimensions and determine the coordinates of a song by the relative emotion of the song with respect to other songs. We also develop an RBF-ListNet algorithm to optimize the ranking-based objective function of our approach. The cognitive load of annotation, the accuracy of emotion recognition, and the subjective quality of the proposed approach are extensively evaluated. Experimental results show that this ranking-based approach simplifies emotion annotation and enhances the reliability of the ground truth. The performance of our algorithm for valence recognition reaches 0.326 in Gamma statistic. Yi-Hsuan Yang, Homer H. Chen |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Prediction of the Distribution of Perceived Music Emotions Using Discrete SamplesabstractTypically, a machine learning model of automatic music emotion recognition is trained to learn the relationship between music features and perceived emotion values. However, simply assigning an emotion value to a clip in the training phase does not work well because the perceived emotion of a clip varies from person to person. To resolve this problem, we propose a novel approach that represents the perceived emotion of a clip as a probability distribution in the emotion plane. In addition, we develop a methodology that predicts the emotion distribution of a clip by estimating the emotion mass at discrete samples of the emotion plane. We also develop model fusion algorithms to integrate different perceptual dimensions of music listening and to enhance the modeling of emotion perception. The effectiveness of the proposed approach is validated through an extensive performance study. An averageR2statistics of 0.5439 for emotion prediction is achieved. We also show how this approach can be applied to enhance our understanding of music emotion. Yi-Hsuan Yang, Homer H. Chen |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Hardware-Efficient Belief PropagationabstractLoopy belief propagation (BP) is an effective solution for assigning labels to the nodes of a graphical model such as the Markov random field (MRF), but it requires high memory, bandwidth, and computational costs. Furthermore, the iterative, pixel-wise, and sequential operations of BP make it difficult to parallelize the computation. In this paper, we propose two techniques to address these issues. The first technique is a new message passing scheme named tile-based BP that reduces the memory and bandwidth to a fraction of the ordinary BP algorithms without performance degradation by splitting the MRF into many tiles and only storing the messages across the neighboring tiles. The tile-wise processing also enables data reuse and pipeline, resulting in efficient hardware implementation. The second technique is anO(L) fast message construction algorithm that exploits the properties of robust functions for parallelization. We apply these two techniques to a very large-scale integration circuit for stereo matching that generates high-resolution disparity maps in near real-time. We also implement the proposed schemes on graphics processing unit (GPU) which is four-time faster than standard BP on GPU. Chia-Kai Liang, Chao-Chung Cheng, Yen-Chieh Lai, Liang-Gee Chen, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | SSIM-Based Perceptual Rate Control for Video CodingabstractThe quality of video is ultimately judged by human eye; however, mean squared error and the like that have been used as quality metrics are poorly correlated with human perception. Although the characteristics of human visual system have been incorporated into perceptual-based rate control, most existing schemes do not take rate-distortion optimization into consideration. In this paper, we use the structural similarity index as the quality metric for rate-distortion modeling and develop an optimum bit allocation and rate control scheme for video coding. This scheme achieves up to 25% bit-rate reduction over the JM reference software of H.264. Under the rate-distortion optimization framework, the proposed scheme can be easily integrated with the perceptual-based mode decision scheme. The overall bit-rate reduction may reach as high as 32% over the JM reference software. Tao-Sheng Ou, Yi-Hsin Huang, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | A Collaborative Transcoding Strategy for Live Broadcasting Over Peer-to-Peer IPTV NetworksabstractReal-time video transcoding that is often needed for robust video broadcasting over heterogeneous networks is not supported in most existing devices. To address this problem, we propose a collaborative strategy that leverages the peering architecture of peer-to-peer Internet protocol television networks and makes the computational resources of peers sharable. The video transcoding task is distributed among the peers and completed collaboratively. A prototype of the live video broadcasting system is evaluated over a 100-node testbed on the PlanetLab. The experimental results show that the proposed strategy works effectively even when the majority of the peers have limited computational resource and bandwidth. Jui-Chieh Wu, Polly Huang, Jason J. Yao, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Learning-Based Prediction of Visual Attention for Video SignalsabstractVisual attention, which is an important characteristic of human visual system, is a useful clue for image processing and compression applications in the real world. This paper proposes a computational scheme that adopts both low-level and high-level features to predict visual attention from video signal by machine learning. The adoption of low-level features (color, orientation, and motion) is based on the study of visual cells, and the adoption of the human face as a high-level feature is based on the study of media communications. We show that such a scheme is more robust than those using purely single low- or high-level features. Unlike conventional techniques, our scheme is able to learn the relationship between features and visual attention to avoid perceptual mismatch between the estimated salience and the actual human fixation. We also show that selecting the representative training samples according to the fixation distribution improves the efficacy of regressive training. Experimental results are shown to demonstrate the advantages of the proposed scheme. Wen-Fu Lee, Tai-Hsiang Huang, Su-Ling Yeh, Homer H. Chen |
IEEE Trans. Image Process. | 4 |
| 2011 | Light Field Analysis for Modeling Image FormationabstractImage formation is traditionally described by a number of individual models, one for each specific effect in the image formation process. However, it is difficult to aggregate the effects by concatenating such individual models. In this paper, we apply light transport analysis to derive a unified image formation model that represents the radiance along a light ray as a 4-D light field signal and physical phenomena such as lens refraction and blocking as linear transformations or modulations of the light field. This unified mathematical framework allows the entire image formation process to be elegantly described by a single equation. It also allows most geometric and photometric effects of imaging, including perspective transformation, defocus blur, and vignetting, to be represented in both 4-D primal and dual domains. The result matches that of traditional models. Generalizations and applications of this theoretic framework are discussed. Chia-Kai Liang, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2011 | Introduction to the ICME2010 Special IssueabstractThe 15 papers in this special issue are extended versions of papers presented at the 2010 IEEE International Conference on Multimedia and Expo (ICME), held in Singapore on July 19-23, 2010. These papers cover a wide range of topics in multimedia including user interface, content understanding, mobility, 3-D processing, storage, and forensics. Zicheng Liu 0001, Ming-Ting Sun, Chia-Wen Lin, Zhengyou Zhang, Zhu Liu 0001, Homer H. Chen, Yap-Peng Tan, Oscar C. Au |
IEEE Trans. Multim. | 6 |
| 2011 | Exploiting online music tags for music emotion classificationabstractThe online repository of music tags provides a rich source of semantic descriptions useful for training emotion-based music classifier. However, the imbalance of the online tags affects the performance of emotion classification. In this paper, we present a novel data-sampling method that eliminates the imbalance but still takes the prior probability of each emotion class into account. In addition, a two-layer emotion classification structure is proposed to harness the genre information available in the online repository of music tags. We show that genre-based grouping as a precursor greatly improves the performance of emotion classification. On the average, the incorporation of online genre tags improves the performance of emotion classification by a factor of 55% over the conventional single-layer system. The performance of our algorithm for classifying 183 emotion classes reaches 0.36 in example-based f-score. Yu-Ching Lin, Yi-Hsuan Yang, Homer H. Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2010 | Improving video coding quality by perceptual rate-distortion optimizationabstractThe goal of this work is to seek a feasible direction of video coding that can provide a significant quality improvement over H.264. In light of the well-known findings that the distortion metric for video quality has a profound impact on video coding performance and that traditional metrics such as mean square error are poorly correlated with human perception, we identify perceptual video coding, more specifically, perceptual-based rate-distortion optimization, as a sensible approach that has the potential to help drive the performance of video coding to a significantly higher quality level. This technology assessment is supported by experiments with various video sequences, bit-rates, and encoding profiles. The results show that the perceptual-based RDO can indeed bring significant quality improvement for H.264. Homer H. Chen, Yi-Hsin Huang, Po-Yen Su, Tao-Sheng Ou |
ICME | 1 |
| 2010 | Perceptual-based coding mode decisionabstractThe framework of rate-distortion optimization (RDO) has been widely adopted for video coding to achieve a good trade-off between bit-rate and distortion. However, objective distortion metrics such as mean square error traditionally used in this framework are poorly correlated with perceptual video quality. To address this issue, we incorporate the structural similarity index as a quality metric into the framework and develop a predictive Lagrange multiplier selection technique to resolve the chicken-and-egg dilemma of perceptual-based RDO. The resulting perceptual-based RDO is then applied to H.264 intra mode decision as an illustration of the application of the proposed technique. Given a perceptual quality level, 5%-10% bit rate reduction over the JM reference software of H.264 is achieved. Subjective evaluation further confirms that, at the same bit-rate, the proposed perceptual RDO preserves image details and prevents block artifact better than the traditional RDO. Yi-Hsin Huang, Tao-Sheng Ou, Homer H. Chen |
ISCAS | 3 |
| 2010 | Shadow removal from natural imagesabstractIn this paper, we present a system for estimating the shadow field from a single natural image. The user of our system is provided with a broad brush to roughly specify the shadow boundary. As the user finishes drawing a stroke, the system starts to estimate the shadow field around the stroke and generates pretty accurate result even if the underlying surface is highly textured. Since the strokes provided by the user may be sparse, an optimization scheme is proposed to propagate the estimated shadow field to the entire image. With this scheme, the amount of user interactions required for the system is reduced. The shadow field estimated by our system can be used to seamlessly remove the shadow from the image. It is also useful for other shadow editing tasks. Experimental results on a variety of photos are provided to show the effectiveness of the proposed system. Ya-Fan Su, Homer H. Chen |
ISCAS | 2 |
| 2010 | A perceptual-based approach to bit allocation for H.264 encoderabstractSince the ultimate receivers of encoded video are human eyes, the characteristics of human visual system should be taken into consideration in the design of bit allocation to improve the perceptual video quality. In this paper, we incorporate the structural similarity index as a distortion metric and propose a novel rate-distortion model to characterize the relationship between rate and the structural similarity index. Based on the model, we develop an optimum bit allocation and rate control scheme for H.264 encoders. Experimental results show that up to 25% bitrate reduction over the JM reference software can be achieved. Subjective evaluation further confirms that the proposed scheme preserves more structural information and improves the perceptual quality of the encoded video. Tao-Sheng Ou, Yi-Hsin Huang, Homer H. Chen |
VCIP | 3 |
| 2010 | Video Object Extraction via MRF-Based Contour TrackingabstractVideo object segmentation is a critical task in multimedia analysis and editing. Normally, the user provides some hints of foreground and background, then the target object is extracted from the video sequence. Most previous methods are either computation-expensive or labor-intensive, and approaches that assume static background have limited applications. In this letter, we propose a novel video segmentation system that integrates Markov random field-based contour tracking with graph-cut image segmentation. The contour tracking propagates the shape of the target object, whereas the graph-cut refines the shape and improves the accuracy of video segmentation. Experimental results show that our segmentation system is efficient and requires less key-frames and user interactions. Chih-Yuan Chung, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Fast Decision of Block Size, Prediction Mode, and Intra Block for H.264 Intra PredictionabstractThe spatial-domain intra prediction scheme of H.264 has high computational complexity, especially for the High Profile as it incorporates the additional intra 8 × 8 prediction mode. To address this issue, we explore the hierarchy of H.264 mode decision process in this paper and adopt an approach that is in synchrony with the mode decision hierarchy. In particular, we propose a variance-based algorithm for block size decision, an improved filter-based algorithm for prediction mode decision using contextual information, and a selection algorithm for intra block decision that exploits the relation between the rate-distortion characteristic and the best coding type. Performance comparison is provided to show the improvement of the proposed algorithms over previous methods. Yi-Hsin Huang, Tao-Sheng Ou, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Perceptual Rate-Distortion Optimization Using Structural Similarity Index as Quality MetricabstractThe rate-distortion optimization (RDO) framework for video coding achieves a tradeoff between bit-rate and quality. However, objective distortion metrics such as mean squared error traditionally used in this framework are poorly correlated with perceptual quality. We address this issue by proposing an approach that incorporates the structural similarity index as a quality metric into the framework. In particular, we develop a predictive Lagrange multiplier estimation method to resolve the chicken and egg dilemma of perceptual-based RDO and apply it to H.264 intra and inter mode decision. Given a perceptual quality level, the resulting video encoder achieves on the average 9% bit-rate reduction for intra-frame coding and 11% for inter-frame coding over the JM reference software. Subjective test further confirms that, at the same bit-rate, the proposed perceptual RDO indeed preserves image details and prevents block artifact better than traditional RDO. Yi-Hsin Huang, Tao-Sheng Ou, Po-Yen Su, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | A Three-Stage Approach to Shadow Field Estimation From Partial Boundary InformationabstractIn this paper, we present a system for estimating the shadow field from a single natural image. Unlike previous works that require extensive user assistance, our system only needs the user to roughly specify the shadow boundary with a broad brush. As the user finishes drawing a stroke, the system starts to estimate the shadow field around the stroke and generates pretty accurate result even if the underlying surface is highly textured. We also propose an optimization scheme to propagate the estimated shadow field to the entire image, achieving a further reduction of user effort required for the system. The shadow field estimated by our system can be used to seamlessly remove the shadow from the image. It is also useful for many shadow editing tasks such as pasting an object's shadow from one image to another. Experimental results on a variety of photos are provided to show the effectiveness of the proposed system. Ya-Fan Su, Homer H. Chen |
IEEE Trans. Image Process. | 2 |
| 2009 | Hardware-efficient belief propagationabstractBelief propagation (BP) is an effective algorithm for solving energy minimization problems in computer vision. However, it requires enormous memory, bandwidth, and computation because messages are iteratively passed between nodes in the Markov random field (MRF). In this paper, we propose two methods to address this problem. The first method is a message passing scheme called tile-based belief propagation. The key idea of this method is that a message can be well approximated from other faraway ones. We split the MRF into many tiles and perform BP within each one. To preserve the global optimality, we store the outgoing boundary messages of a tile and use them when performing BP in the neighboring tiles. The tile-based BP only requires 1-5% memory and 0.2-1% bandwidth of the ordinary BP. The second method is an O(L) message construction algorithm for the robust functions commonly used for describing the smoothness terms in the energy function. We find that many variables in constructing a message are repetitive; thus these variables can be calculated once and reused many times. The proposed algorithms are suitable for parallel implementations. We design a low-power VLSI circuit for disparity estimation that can construct 440 M messages per second and generate high quality disparity maps in near real-time. We also implement the proposed algorithms on a GPU, which can calculate messages 4 times faster than the sequential O(L) method. Chia-Kai Liang, Chao-Chung Cheng, Yen-Chieh Lai, Liang-Gee Chen, Homer H. Chen |
CVPR | 5 |
| 2009 | Fast belief propagation process element for high-quality stereo estimationabstractBelief propagation is a popular global optimization technique for many computer vision problems. However, it requires extensive computation due to the iterative message passing operations. In this paper, we present a new process element (PE) for efficient message construction. The efficiency is gained by exploiting the unique characteristics of the generalized Potts model (truncated linear mode) of the smoothness term in the Markov random field. For stereo estimation with L disparity values, the algorithm successfully reduces the computation from O(L2) to O(L) and retains the high throughput and low latency. Compared with the direct message construction PE, our method achieves 87.14% computation saving and a 94.38% PE area reduction. Chao-Chung Cheng, Chia-Kai Liang, Yen-Chieh Lai, Homer H. Chen, Liang-Gee Chen |
ICASSP | 4 |
| 2009 | Music emotion rankingabstractContent-based retrieval has emerged as a promising approach to information access. In this paper, we propose an approach to music emotion ranking. Specifically, we rank music in terms of arousal and valence and represent each song as a point in the 2D emotion space. Novel ranking-based methods for annotation, learning, and evaluation of music emotion recognition are developed and tested on a moderately large-scale database composed of 1240 pop songs. Results are provided to show the feasibility of the proposed approach. Yi-Hsuan Yang, Homer H. Chen |
ICASSP | 2 |
| 2009 | Perception-based high dynamic range compression in gradient domainabstractIt is often required to map the radiances of a real scene to a smaller dynamic range so that the image can be properly displayed. However, most such algorithms suffer from the halo artifact or require manual parameter tweaking that is often a tedious process for the user. We propose an automatic algorithm for high dynamic range compression based on the properties of human visual system. The algorithm is performed in the gradient domain to avoid the halo artifact. It automates the parameter adjustment process while preserving the image details. Performance comparison is provided to illustrate the advantages of the proposed algorithm. Wen-Fu Lee, Tsung-Yi Lin, Mei-Lan Chu, Tai-Hsiang Huang, Homer H. Chen |
ICIP | 5 |
| 2009 | Realism assessment of color compatibility using a single imageabstractIn this paper, we present a study of the realism of color image composites. Assessing the realism of image composites has emerged as a new field of image processing due to the advances in digital imaging and communications. However, when making image composites, users often suffer from color incompatibility between the inserted object and the background. We observe two properties that help make an image composite look realistic. The first property is related to the color similarity between different segments of the image, and the second one is related to the consistency of color deviation between the segments. These two properties only require information available from a single image. An algorithm based on these two properties is proposed for assessment of image realism. Effectiveness of the algorithm is demonstrated. Bing-Yi Wong, Chia-Kai Liang, Tai-Hsu Lin, Homer H. Chen |
ICIP | 4 |
| 2009 | Research on light field camera and music emotion recognitionabstractIn this talk, I will give a presentation of some recent progress we have made at the Multimedia Processing and Communications Lab of National Taiwan University. Specifically, I will describe the multimedia signal processing techniques we have developed for applications to digital video camera and emotion-based multimedia presentation. The demos to be shown in this talk include 1) programmable aperture photography: multiplexed light field acquisition, 2) music emotion recognition for multimedia retrieval and presentation, 3) integration of video stabilizer with digital video codec, 4) auto focusing, 5) auto white balance, 6) color transfer, and 7) Wii-like 3D pointer. Homer H. Chen |
ICME | 1 |
| 2009 | Exploiting genre for music emotion classificationabstractGenre and emotion have been applied to content-based music retrieval and organization; however, the intrinsic correlation between them has not been explored. In this paper we present a statistical association analysis to examine such intrinsic correlation and propose a two-layer scheme that exploits the correlation for emotion classification. Significant improvement of classification accuracy over the traditional single-layer scheme is obtained. Yu-Ching Lin, Yi-Hsuan Yang, Homer H. Chen, I-Bin Liao, Yeh-Chin Ho |
ICME | 3 |
| 2009 | Clustering for music search resultsabstractClustering for better representation of the diversity of text or image search results has been studied extensively. In this paper, we extend this methodology to the novel domain of music search. We conduct empirical evaluation of different clustering algorithms, audio feature representations, and the incorporation of lyrics for music clustering. Our evaluation shows the fusion of audio and text features yields the best clustering accuracy. Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
ICME | 3 |
| 2009 | Multimodal Structure Segmentation and Analysis of Music using Audio and Textual InformationabstractIn this paper, we present a multimodal approach to structure segmentation of music with applications to audio content analysis and music information retrieval. In particular, since lyrics contain rich information about the semantic structure of a song, our approach incorporates lyrics to overcome the existing difficulties associated with large acoustic variation in music. We further design a constrained clustering algorithm for music segmentation and evaluate its performance on commercial recordings. Experimental results show that our method can effectively detect the boundaries and the types of semantic structure of music segments. Heng Tze Cheng, Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
ISCAS | 4 |
| 2009 | Fast H.264 selective intra mode decision for inter-frame codingabstractDespite its remarkable coding efficiency, the computational complexity of the spatial-domain intra prediction scheme of H.264 remains a challenging issue especially for applications of the high profile, as it incorporates the additional intra 8times8 prediction mode. To address this issue, we propose in this paper a fast selective intra mode decision algorithm that exploits a newly discovered property of inter-frame coding that the difference in rate-distortion cost between the best inter mode and the intra 16times16 mode is highly related to the type of the best coding mode. The proposed algorithm has high prediction accuracy, low computational overhead, and negligible effect on PSNR and bit-rate. Performance comparison is provided to show the superiority of the algorithm. Yi-Hsin Huang, Tao-Sheng Ou, Homer H. Chen |
PCS | 3 |
| 2009 | Efficient MB and prediction mode decisions for intra prediction of H.264 High ProfileabstractThe rate-distortion optimization framework employed in H.264 for the selection of best coding mode is computationally expensive, although it achieves remarkable coding efficiency. In this paper, an integrated intra prediction algorithm for intra-frame coding of the H.264 high profile is proposed. It consists of two components: a variance-based method for macroblock mode decision and an improved filter-based method for prediction mode decision using contextual information. The two components can work independently or jointly to achieve higher efficiency. The integrated algorithm achieves an average saving of 51.3% (maximum 66.0%) in encoding time for intra coded frames, with negligible effect on PSNR and bit-rate. Tao-Sheng Ou, Yi-Hsin Huang, Homer H. Chen |
PCS | 3 |
| 2009 | Personalized music emotion recognitionabstractIn recent years, there has been a dramatic proliferation of research on information retrieval based on highly subjective concepts such as emotion, preference and aesthetic. Such retrieval methods are fascinating but challenging since it is difficult to built a general retrieval model that performs equally well to everyone. In this paper, we propose two novel methods, bag-of-users model and residual modeling, to accommodate the individual differences for emotion-based music retrieval. The proposed methods are intuitive and generally applicable to other information retrieval tasks that involve subjective perception. Evaluation result shows the effectiveness of the proposed methods. Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
SIGIR | 3 |
| 2009 | Image Enhancement for Backlight-Scaled TFT-LCD DisplaysabstractOne common way to extend the battery life of a portable device is to reduce the LCD backlight intensity. In contrast to previous approaches that minimize the power consumption by adjusting the backlight intensity frame by frame to reach a specified image quality, the proposed method optimizes the image quality for a given backlight intensity. Image is enhanced by performing brightness compensation and local contrast enhancement. For brightness compensation, global image statistics and backlight level are considered to maintain the overall brightness of the image. For contrast enhancement, the local contrast property of human visual system (HVS) is exploited to enhance the local image details. In addition, a brightness prediction scheme is proposed to speed up the algorithm for display of video sequences. Experimental results are presented to show the performance of the algorithm. Pei-Shan Tsai, Chia-Kai Liang, Tai-Hsiang Huang, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Online Reranking via Ordinal Informative Concepts for Context Fusion in Concept Detection and Video SearchabstractTo exploit the co-occurrence patterns of semantic concepts while keeping the simplicity of context fusion, a novel reranking approach is proposed in this paper. The approach, called ordinal reranking, adjusts the ranking of an initial search (or detection) list based on the co-occurrence patterns obtained by using ranking functions such as ListNet. Ranking functions are by nature more effective than classification-based reranking methods in mining ordinal relationships. In addition, the ordinal reranking is free of thead hocthresholding for noisy binary labels and requires no extra offline learning or training data. To select informative concepts for reranking, we also propose a new concept selection measurement,wc-tf-idf, which considers the underlying ordinal information of ranking lists and is thus more effective than the feature selection algorithms for classification. Being largely unsupervised, the reranking approach to context fusion can be applied equally well to concept detection and video search. While being extremely efficient, ordinal reranking outperforms existing methods by up to 40% in mean average precision (MAP) for the baseline text-based search and 12% for the baseline concept detection over TRECVID 2005 video search and concept detection benchmark. Yi-Hsuan Yang, Winston H. Hsu, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Video Streaming Over In-Home Power Line NetworksabstractThe deployment of power line communication technology for broadband video streaming remains a challenge because power lines are not originally designed for signal transmission. Scalable video is a viable approach that can cope with the bandwidth fluctuation of power line communication networks provided that the bandwidth information is available. In this paper we first investigate how the interference caused by electrical appliances or power supplies affects the power line channel bandwidth and packet transmission. Then we take the obtained characteristics of in-home power line network into account in the design of a simple but effective heuristic-based application-layer bandwidth estimation scheme, for which the cutoff rate is estimated from the packet size and the physical-layer data rates. Experimental results show that the proposed approach can effectively combat the noise interference and deliver robust video streaming over power line. Chang-Kuan Lin, Meng-Ting Lu, Shiann-Chang Yeh, Homer H. Chen |
IEEE Trans. Multim. | 4 |
| 2008 | JND-based enhancedment of perceptibility for dim imagesabstractMaintaining image quality under various lighting conditions is critical to portable multimedia devices. In this paper, an image enhancement algorithm is proposed to deal with extremely dim LCD backlight (or strong ambient light) under which the image often becomes imperceptible. The crux of our idea is to keep the detail content of the image in the visible luminance range. To do that, a two-layer image decomposition architecture based on the just noticeable difference (JND) theory is proposed for extracting the details, and a boosting scheme is applied to the details that are in the low intensity regions of the image. Experimental results are provided to show the superiority of the proposed algorithm. Tai-Hsiang Huang, Chia-Kai Liang, Su-Ling Yeh, Homer H. Chen |
ICIP | 4 |
| 2008 | Automatic chord recognition for music classification and retrievalabstractAs one of the most important mid-level features of music, chord contains rich information of harmonic structure that is useful for music information retrieval. In this paper, we present a chord recognition system based on the N-gram model. The system is time-efficient, and its accuracy is comparable to existing systems. We further propose a new method to construct chord features for music emotion classification and evaluate its performance on commercial song recordings. Experimental results demonstrate the advantage of using chord features for music classification and retrieval. Heng Tze Cheng, Yi-Hsuan Yang, Yu-Ching Lin, I-Bin Liao, Homer H. Chen |
ICME | 5 |
| 2008 | Fast H.264 Mode Decision Using Previously Coded InformationabstractThe seven prediction modes for P-slice provided in the inter-frame coding tool of H.264/AVC enables significant compression gain but at the cost of computational complexity. To address the issue, we propose in this paper a fast mode decision algorithm that is able to reduce the computational complexity of this coding tool by exploiting the coding modes of previously coded neighboring blocks and by leveraging the rate-distortion characteristics of processed modes. The target efficiency is achieved because the combination of various techniques used in the algorithm incurs little overhead. The algorithm provides an average 70% saving of encoding time for QCIF video and 73% for CIF video in comparison with the brute-force scheme, with only 0.032dB PSNR drop or, equivalently, 0.38% bit rate increase. Yi-Hsin Huang, Che-Yu Chang, Homer H. Chen |
ISM | 3 |
| 2008 | Interactive content presentation based on expressed emotion and physiological feedbackabstractIn this technical demonstration, we showcase an interactive content presentation (ICP) system that integrates media-expressed-emotion-based composition, user-perceived preference feedback, and interactive digital art creation. ICP harmonizes the browsing of multimedia contents by presenting them in the form of music videos (photos, blog articles with accompanied music) based on their expressed emotion similarity. ICP facilitates content browsing by automatically and dynamically selecting the media to be played next in real time, responding to user's preference feedback measured from physiological signals. In addition, ICP enhances the enjoyments of content browsing by incorporating interactive digital art creation. ICP achieves these goals by properly integrating recent researches on media-expressed emotion classification,cross-media composition, and physiological signal processing. Tien-Lin Wu, Hsuan-Kai Wang, Murphy Chien-Chang Ho, Yuan-Pin Lin, Ting-Ting Hu, Ming-Fang Weng, Li-Wei Chan 0001, Changhua Yang, Yi-Hsuan Yang, Yi-Ping Hung, Yung-Yu Chuang, Hsin-Hsi Chen, Homer H. Chen, Jyh-Horng Chen, Shyh-Kang Jeng |
ACM Multimedia | 13 |
| 2008 | Mr. Emo: music retrieval in the emotion planeabstractThis technical demo presents a novel emotion-based music retrieval platform, called Mr. Emo, for organizing and browsing music collections. Unlike conventional approaches which quantize emotions into classes, Mr. Emo defines emotions by two continuous variables arousal and valence and employs regression algorithms to predict them. Associated with arousal and valence values (AV values), each music sample becomes a point in the arousal-valence emotion plane, so a user can easily retrieve music samples of certain emotion(s) by specifying a point or a trajectory in the emotion plane. Being content centric and functionally powerful, such emotion-based retrieval complements traditional keyword- or artist-based retrieval. The demo shows the effectiveness and novelty of music retrieval in the emotion plane. Yi-Hsuan Yang, Yu-Ching Lin, Heng Tze Cheng, Homer H. Chen |
ACM Multimedia | 4 |
| 2008 | ContextSeer: context search and recommendation at query time for shared consumer photosabstractThe advent of media-sharing sites like Flickr has drastically increased the volume of community-contributed multimedia resources on the web. However, due to their magnitudes, these collections are increasingly difficult to understand, search and navigate. To tackle these issues, a novel search system, ContextSeer, is developed to improve search quality (by reranking) and recommend supplementary information (i.e., search-related tags and canonical images) by leveraging the rich context cues, including the visual content, high-level concept scores, time and location metadata. First, we propose an ordinal reranking algorithm to enhance the semantic coherence of text-based search result by mining contextual patterns in an unsupervised fashion. A novel feature selection method, wc-tf-idf is also developed to select informative context cues. Second, to represent the diversity of search result, we propose an efficient algorithm cannoG to select multiple canonical images without clustering. Finally, ContextSeer enhances the search experience by further recommending relevant tags. Besides being effective and unsupervised, the proposed methods are efficient and can be finished at query time, which is vital for practical online applications. To evaluate ContextSeer, we have collected 0.5 million consumer photos from Flickr and manually annotated a number of queries by pooling to form a new benchmark, Flickr550. Ordinal reranking achieves significant performance gains both in Flcikr550 and TRECVID search benchmarks. Through a subjective test, cannoG expresses its representativeness and excellence for recommending multiple canonical images. Yi-Hsuan Yang, Po Tun Wu, Ching-Wei Lee, Kuan Hung Lin, Winston H. Hsu, Homer H. Chen |
ACM Multimedia | 6 |
| 2008 | H.264/AVC-based multiple description video coding using dynamic slice groups
Che-Chun Su, Homer H. Chen, Jason J. Yao, Polly Huang |
Signal Process. Image Commun. | 2 |
| 2008 | A Regression Approach to Music Emotion RecognitionabstractContent-based retrieval has emerged in the face of content explosion as a promising approach to information access. In this paper, we focus on the challenging issue of recognizing the emotion content of music signals, or music emotion recognition (MER). Specifically, we formulate MER as a regression problem to predict the arousal and valence values (AV values) of each music sample directly. Associated with the AV values, each music sample becomes a point in the arousal-valence plane, so the users can efficiently retrieve the music sample by specifying a desired point in the emotion plane. Because no categorical taxonomy is used, the regression approach is free of the ambiguity inherent to conventional categorical approaches. To improve the performance, we apply principal component analysis to reduce the correlation between arousal and valence, and RReliefF to select important features. An extensive performance study is conducted to evaluate the accuracy of the regression approach for predicting AV values. The best performance evaluated in terms of theR2statistics reaches 58.3% for arousal and 28.1% for valence by employing support vector machine as the regressor. We also apply the regression approach to detect the emotion variation within a music selection and find the prediction accuracy superior to existing works. A group-wise MER scheme is also developed to address the subjectivity issue of emotion perception. Yi-Hsuan Yang, Yu-Ching Lin, Ya-Fan Su, Homer H. Chen |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Analysis and Compensation of Rolling Shutter EffectabstractDue to the sequential-readout structure of complementary metal-oxide semiconductor image sensor array, each scanline of the acquired image is exposed at a different time, resulting in the so-called electronic rolling shutter that induces geometric image distortion when the object or the video camera moves during image capture. In this paper, we propose an image processing technique using a planar motion model to address the problem. Unlike previous methods that involve complex 3-D feature correspondences, a simple approach to the analysis of inter- and intraframe distortions is presented. The high-resolution velocity estimates used for restoring the image are obtained by global motion estimation, BEzier curve fitting, and local motion estimation without resort to correspondence identification. Experimental results demonstrate the effectiveness of the algorithm. Chia-Kai Liang, Li-Wen Chang, Homer H. Chen |
IEEE Trans. Image Process. | 3 |
| 2008 | Programmable aperture photography: multiplexed light field acquisitionabstractIn this paper, we present a system including a novel component called programmable aperture and two associated post-processing algorithms for high-quality light field acquisition. The shape of the programmable aperture can be adjusted and used to capture light field at full sensor resolution through multiple exposures without any additional optics and without moving the camera. High acquisition efficiency is achieved by employing an optimal multiplexing scheme, and quality data is obtained by using the two post-processing algorithms designed for self calibration of photometric distortion and for multi-view depth estimation. View-dependent depth maps thus generated help boost the angular resolution of light field. Various post-exposure photographic effects are given to demonstrate the effectiveness of the system and the quality of the captured light field. Chia-Kai Liang, Tai-Hsu Lin, Bing-Yi Wong, Homer H. Chen |
ACM Trans. Graph. | 5 |
| 2007 | A Scalable Peer-to-Peer IPTV SystemabstractHotStreaming is an overlay peer-to-peer (P2P) based IPTV system. It integrates the innovations in both overlay networking and video coding for optimal user experience. The HotStreaming system is composed of three key components: partnership formation, data request scheduling and multiple description coding (MDC). In the partnership formation component, we propose two policies to reduce the time of disconnection and the number of isolated peers. In the MDC component, we adopt MDC with spatial-temporal hybrid interpolation (MDC-STHI), which makes the peers in a P2P network to adjust the streaming traffic according to their bandwidth limitation and capability of devices. The experimental results show that the HotStreaming system improves the video quality over a lossy and dynamic networking environment. Meng-Ting Lu, Hung Nien, Jui-Chieh Wu, Kuan-Jen Peng, Polly Huang, Jason J. Yao, Chih-Chun Lai, Homer H. Chen |
CCNC | 8 |
| 2007 | Depth Detection of Light FieldabstractWe propose an algorithm to detect depths in a light field. Specifically, given a 4D light field, we find all planes at which objects are located. Although the exact depth of each pixel in the space is left unknown, the partial information obtained is very useful for many applications, such as synthetic aperture photography and all-focused rendering. Our algorithm measures the degree of focus of different planes by calculating the ratio of high frequencies to the low frequencies. To handle different depth distributions, we reformulate the maximum detection problem to a maximum-cover problem that can be solved efficiently by dynamic programming. Compared with auto-focusing and per-pixel depth estimation, our algorithm is much faster yet sufficiently accurate. Yi-Hao Kao, Chia-Kai Liang, Li-Wen Chang, Homer H. Chen |
ICASSP (1) | 4 |
| 2007 | Light Field Acquisition using Programmable Aperture CameraabstractWe propose a new device, programmable aperture camera (PAC), to capture 4D light field in a camera. PAC can adjust the shape of the aperture in each exposure. This allows us to capture the angular information of the light field, which is lost in regular photography. Although multiple exposures are needed to obtain a light field, the total exposure time remains the same as that of taking a single regular photograph at the same image quality level. As opposed to previous techniques that seriously reduce the spatial resolution, PAC captures the image at full spatial resolution and allows adjustable angular resolution. Also its manufacturing cost is much lower than previous techniques. We describe the PAC prototype and demonstrate how digital refocusing is made possible by using the captured light field. Chia-Kai Liang, Gene Liu, Homer H. Chen |
ICIP (5) | 3 |
| 2007 | H.264/AVC-Based Multiple Description Coding SchemeabstractA new multiple description coding (MDC) scheme is proposed in this paper. It utilizes the advanced video coding tools and features provided in H.264/AVC to introduce redundancy into descriptions. The proposed MDC scheme produces two descriptions, each consisting of two slice groups. One of them, called main slice group (MSG), is encoded in the normal way as main information. The other one, called side slice group (SSG), is encoded with fewer bits as redundancy by using larger quantization step sizes. Spatial and temporal correlations between neighboring macroblocks in video frames are exploited to achieve efficient redundancy coding. Experimental results show that the proposed MDC scheme achieves better rate-distortion (R-D) performance than previous slice-group based MDC schemes. Che-Chun Su, Jason J. Yao, Homer H. Chen |
ICIP (4) | 3 |
| 2007 | Image Quality Enhancement for Low Backlight TFT-LCD DisplaysabstractReducing LCD backlight saves power consumption of a portable device, but it also decreases the contrast and brightness of the displayed image. Previous approaches adjust the backlight level frame by frame to reach a specified image quality level without optimizing the image quality. In contrast, the proposed method adjusts the backlight to meet the target power level while maintaining the image quality. This is achieved by performing brightness compensation and local contrast enhancement in accordance with the given backlight level. Experimental results show that the proposed algorithm outperforms previous methods. Pei-Shan Tsai, Chia-Kai Liang, Homer H. Chen |
ICIP (3) | 3 |
| 2007 | Music Emotion Classification: A Regression ApproachabstractTypical music emotion classification (MEC) approaches categorize emotions and apply pattern recognition methods to train a classifier. However, categorized emotions are too ambiguous for efficient music retrieval. In this paper, we model emotions as continuous variables composed of arousal and valence values (AV values), and formulate MEC as a regression problem. The multiple linear regression, support vector regression, and AdaBoost.RT are adopted to evaluate the prediction accuracy. Since the regression approach is inherently continuous, it is free of the ambiguity problem existing in its categorical counterparts. Yi-Hsuan Yang, Yu-Ching Lin, Ya-Fan Su, Homer H. Chen |
ICME | 4 |
| 2007 | Feature-Based Full-Frame Image StabilizationabstractDigital image stabilization usually discards boundary pixels and outputs a smaller video. In this paper, we present a new digital image stabilization algorithm that preserves the frame size of output video by pixel filling. The proposed algorithm eliminates the accumulation error by directly estimating the global motions in a transformation chain with reference to a fixed frame. A feature matching method is adopted to save the computational cost of the global motion estimation and to handle large motions. The experimental results show that the proposed algorithm produces stabilized full-frame video sequences with better frame alignment. Chih-Yuan Chung, Homer H. Chen |
ISM | 2 |
| 2007 | Edge-based automatic white balancing with linear illuminant constraintabstractAutomatic white balancing is an important function for digital cameras. It adjusts the color of an image and makes the image look as if it is taken under canonical light. White balance is usually achieved by estimating the chromaticity of the illuminant and then using the resulting estimate to compensate the image. The grey world method is the base of most automatic white balance algorithms. It generally works well but fails when the image contains a large object or background with a uniform color. The algorithm proposed in this paper solves the problem by considering only pixels along edges and by imposing an illuminant constraint that confines the possible colors of the light source to a small range during the estimation of the illuminant. By considering only edge points, we reduce the impact of the dominant color on the illuminant estimation and obtain a better estimate. By imposing the illuminant constraint, we further minimize the estimation error. The effectiveness of the proposed algorithm is tested thoroughly. Both objective and subjective evaluations show that the algorithm is superior to other methods. Homer H. Chen, Chun-Hung Shen, Pei-Shan Tsai |
VCIP | 1 |
| 2007 | Stochastic Color Interpolation for Digital CamerasabstractThis paper presents a stochastic estimation approach to adaptive interpolation of color filter array. It models an image as a 2D locally stationary Gaussian process and achieves robustness against aliasing by employing an edge-sensitive weighting policy based on the stochastic characteristics of uniformly oriented edge indicators. Experimental results show that the algorithm can effectively eliminate the occurrence of perceptible artifact. Performance comparison in terms of peak signal-to-noise ratio and mean square error is provided to demonstrate the superiority of the proposed algorithm. Hung-An Chang, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Integration of Digital Stabilizer With Video Codec for Digital Video CamerasabstractThis paper presents three novel schemes for integrating digital stabilizer with video codec of a digital video camera, each designed for a different application scenario. Scheme 1 confines the global motion estimation (ME) to within a small background region determined by clustering the motion vectors (MVs) generated by the video encoder. Scheme 2 performs the main ME task at the digital stabilizer and sends the resulting motion vectors to the video encoder, where the motion vectors are further refined to subpixel accuracy. Scheme 3, which is applied on the decoder side, acquires the motion information directly from the video decoder to perform digital stabilization. These integration schemes achieve computational efficiency without affecting the performance of video coding and digital stabilization. Homer H. Chen, Chia-Kai Liang, Yu-Chun Peng, Hung-An Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Frame-Layer Constant-Quality Rate Control of Regions of Interest for Multiple Encoders With Single Video SourceabstractIn this work, we develop a constant-quality rate control algorithm for a surveillance system which consists of one ldquobase encoderrdquo that encodes a down-sampled full-view version of the input video sequence, and one ldquoregion of interest (ROI) encoderrdquo that encodes the region of interest of the input video at the original resolution. Exploiting the inter-relationship between these two independent encoders, the algorithm allocates the bits for the ROI encoder according to the distortion obtained from the corresponding region in the base encoder. Simulation results show that the proposed algorithm can achieve significant reduction in the image quality variation. Compared to the rate control algorithm in JM 8.4, the overall quality is improved and the bit rate is saved. Ping-Hao Wu, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Design and Evaluation of a P2P IPTV System for Heterogeneous NetworksabstractNTUStreaming is an overlay P2P-based IPTV system that integrates innovations in both overlay networking and video coding for optimal user experience. The system consists of three key components: partnership formation, robust video coding, and video segment request scheduling. For partnership formation, a graph construction mechanism TYPHOON based on epidemic algorithms is developed to reduce disconnect time and isolated peers. For robust video coding, a multiple description coding (MDC) scheme with spatial-temporal hybrid interpolation (STHI) is proposed to adjust streaming traffic according to the bandwidth and device capability of each peer. For request scheduling, an optimization algorithm is developed by taking the available bandwidth and the video segment type into account. Experimental results show that NTUStreaming is able to deliver optimal video quality in lossy and dynamic networking environments. Meng-Ting Lu, Jui-Chieh Wu, Kuan-Jen Peng, Polly Huang, Jason J. Yao, Homer H. Chen |
IEEE Trans. Multim. | 6 |
| 2006 | A robust DRM system on the DVB multimedia home platformabstractIn a digital home, the copy rights of the high-quality multimedia content broadcasted from a DVB system need to be protected. However, there is no specification on how to enforce the usage rights of digital content in the DVB standards. As a result, even if the digital content is protected under the conditional access sub-system, end users can still copy and redistribute the digital content once it is descrambled. In this paper, we proposed a DRM system for set-top box which supports Multimedia Home Platform middleware. The rights of the protected digital content are described using the MPEG-21 Rights Expression Language and broadcasted with the digital content. In the proposed system, the rights are stored in the smart card. The proposed system is highly renewable and extensible. The service providers can integrate this system with their existing broadcast services without additional hardware cost. Chia-Kai Liang, Chia Chu Liu, Homer H. Chen |
CCNC | 3 |
| 2006 | Improving the coding of regions of interestabstractThis paper considers a video coding system for surveillance applications. It consists of one "base encoder" that encodes a down-sampled, full-view version of the input video sequence and one "region of interest" (ROI) encoder that encodes an ROI of the video sequence at the original image resolution. An important requirement of the video coding system is that the ROI bit stream and the base bit stream should be independently decodable. We explore the inter-relationship between the full-view video sequence and the ROI video sequence and apply it to improve the computational efficiency of the ROI encoder. The proposed algorithm achieves an average of 170% speedup. Shu-Fa Lin, Homer H. Chen, Yuh-Feng Hsu |
ISCAS | 3 |
| 2006 | Constant-Quality Rate Control Algorithm for Multiple Encoders with Single Video SourceabstractIn this paper, we develop a constant-quality rate control algorithm for a surveillance system that consists of one "base encoder" for encoding a down-sampled full-view version of the input video sequence and one "region of interest (ROI) encoder" for encoding the region of interest of the input video at the original resolution. Exploiting the inter-relationship between these two independent encoders, the bits for the ROI encoder are allocated according to the distortion obtained from the base encoder. Simulation results show that the proposed algorithm significantly reduces the image quality variation. The overall quality is improved and the bit-rate is saved as compared to the rate control algorithm in JM 8.4 Ping-Hao Wu, Homer H. Chen |
ISM | 2 |
| 2006 | Smooth Playout Control for Video Streaming over Error-Prone ChannelsabstractThe quality of media streaming over best-effort networks suffers from network delays and packet losses. The latter is more profound for wireless video. To enhance the QoS of streaming services, adaptive media playout (AMP) has been developed to adjust the playout interval. With AMP, the risk of delay and buffer underflow is reduced. However, the smoothness of playback is not guaranteed. In this paper, we propose a novel AMP control that enables smooth playout and meanwhile maintains reliable visual quality. Our AMP control adjusts the playout interval based on an estimation of channel quality, so it is more adaptive than conventional AMP controls that are based on buffer fullness. Experimental results are provided to justify our approach. Even at 20% packet loss rate, the proposed AMP control is still able to provide smooth and reliable playback Yi-Hsuan Yang, Meng-Ting Lu, Homer H. Chen |
ISM | 3 |
| 2006 | Music emotion classification: a fuzzy approachabstractDue to the subjective nature of human perception, classification of the emotion of music is a challenging problem. Simply assigning an emotion class to a song segment in a deterministic way does not work well because not all people share the same feeling for a song. In this paper, we consider a different approach to music emotion classification. For each music segment, the approach determines how likely the song segment belongs to an emotion class. Two fuzzy classifiers are adopted to provide the measurement of the emotion strength. The measurement is also found useful for tracking the variation of music emotions in a song. Results are shown to illustrate the effectiveness of the approach. Yi-Hsuan Yang, Chia Chu Liu, Homer H. Chen |
ACM Multimedia | 3 |
| 2006 | Rounding Mismatch Between Spatial-Domain and Transform-Domain Video CodecsabstractThe mismatch between a spatial-domain encoder and a transform-domain decoder due to rounding can cause serious video quality degradation after a series of inter-coded frames. The mismatch is rooted on the fact that rounding is a nonlinear operation, and hence its equivalent in the transform domain does not exist. We analyze the mathematical property of the rounding operation and propose a practical solution that provides an optimal approximation of the rounding operation in the transform-domain. The proposed solution is able to reduce the mismatch and preserve image quality. Experimental results are shown to demonstrate the effectiveness of the solution Ping-Hao Wu, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Digital image stabilization and its integration with video encoderabstractDigital image stabilizer and video encoder are two important components of a digital video camera. The digital image stabilizer compensates the image movement caused by hand jiggles and thereby improves the perceptual quality of the captured image sequence, while the video encoder compresses the huge amount of video data down to a reasonable size. Both components require the motion information of the image sequence to perform their respective tasks. Since motion estimation is a computation intensive operation, we present an integration scheme for integrating the digital image stabilizer with the video encoder. The image stabilization algorithms and the technical issues involved in the integration are discussed. Simulation results are shown to illustrate the effectiveness of the proposed digital image stabilization system. Yu-Chun Peng, Hung-An Chang, Homer H. Chen, Chang-Jung Kao |
CCNC | 3 |
| 2005 | Complexity-aware live streaming systemabstractThe number of client requests that a streaming server can handle is limited by both its computational resources and available bandwidth. While bandwidth capacity is critical for most streaming applications, computational resources for a server encoding live videos often become a critical factor as well. In order to serve more client requests or provide higher quality for high priority clients, it is desirable to allocate and adjust the computational resources on a per channel basis. In this paper, we proposed a complexity-aware live video streaming server system that manages the computational resources dynamically. In the proposed system, input videos are encoded with different quality levels based on their priorities and available computational resources. The computational resources for each encoder are adaptively allocated to match the time constraints. The seven quality levels defined in the XviD MPEG-4 encoder are used in our experiments. The results show that the new design is able to maximize the resource utilization by maintaining the highest priority channels quality while providing the other channels with best-effort quality. Meng-Ting Lu, Chang-Kuan Lin, Jason Yao, Homer H. Chen |
ICIP (1) | 4 |
| 2005 | DSP implementation of digital image stabilizerabstractA digital image stabilization system compensates the image movement caused by hand jiggle for the image sequence captured by a hand-held video camera. In this paper, a simplified stabilization algorithm based on our previous work is presented. The algorithm performs block-based motion estimation on 16 local 16/spl times/16 blocks and uses a median filter to estimate the global motion. It reduces the complexity by confining the motion estimation to a small number of blocks of the image. This greatly facilitates the implementation of the algorithm on BF561, a DSP processor of analog device. Details of the DSP implementation are described. Yu-Chun Peng, Meng-Ting Lu, Homer H. Chen |
ICME | 3 |
| 2005 | A Complexity-Aware Live Streaming System with Bit Rate AdjustmentabstractFor a live streaming server, it is highly desirable to allocate the available computational resource and bandwidth to each channel fairly and efficiently. In this paper, we propose a complexity-aware live streaming system with bit rate adjustment that can handle the allocation of bandwidth and computational resource of a live streaming server. The proposed system encodes the input videos at different quality levels based on the priority of the input videos and the available computational resource. It incorporates a bit rate adjustment mechanism to compensate for the video quality drop resulted from the quality level change of high priority encoders. The resulting system is able to handle channels more efficiently because the complexity of high priority encoders can also be dynamically adjusted with little quality drop. A new complexity adjustment method is developed that enables the system to stabilize more quickly and minimizes the variance of the time buffer. The experimental results show that the system can handle more channels while still maintaining the quality of high priority encoders. The proposed system is applicable to multimedia home gateways, surveillance, IP-based TV, and on-line sports game relays. Meng-Ting Lu, Chang-Kuan Lin, Jason Yao, Homer H. Chen |
ISM | 4 |
| 2000 | Error-resilient coding in JPEG-2000 and MPEG-4abstractThe rapid growth of mobile communications and the widespread access to information via the Internet have resulted in a strong demand for robust transmission of compressed image and video data for various multimedia applications and services. The challenge of robust transmission is to protect the compressed image/video data against hostile channel conditions while bringing little impact on bandwidth efficiency. This paper addresses this critical problem and provides an overview of the error-resilient approaches that have been evaluated and inserted into the emerging JPEG-2000 wavelet-based image coding standard. We also review the state-of-the-art techniques adopted in the MPEG-4 standard for robust transmission of video and still texture data. These techniques include resynchronization strategies, data partitioning, reversible VLCs, and header extension codes. The performance of these approaches under various channel conditions is evaluated. Iole Moccagatta, Salma Soudagar, Homer H. Chen |
IEEE J. Sel. Areas Commun. | 4 |
| 2000 | Minimum-drift digital video downconversionabstractThis paper presents a new technique for decoding a full-resolution video bitstream at low memory cost and displaying the signal at a lower resolution. Existing techniques solve the problem by storing the downconverted blocks into memory instead of the full-resolution blocks. While the memory is reduced, these techniques introduce drift errors because the decoder does not have the same pixels as the encoder in performing motion-compensated prediction. The approach proposed here alleviates the problem by tracking the drift at the decoder. It improves the video quality without any increase in decoder complexity. The effectiveness of the approach is evaluated using both objective and subjective tests. This minimum-drift approach is very simple to implement and can also be applied for memory reduction of a full-resolution HDTV decoder. Osama K. Al-Shaykh, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Minimum-drift digital video down-conversionabstractThis paper presents a new technique for decoding a full-resolution video bitstream at low memory cost and displaying the signal at a lower resolution. Existing techniques solve the problem by storing the down-converted blocks into memory instead of the full-resolution blocks. While the memory is reduced, these techniques introduce drift errors because the decoder does not have the same pixels as the encoder in performing motion-compensated prediction. The approach proposed here alleviates the problem by tracking the drift at the decoder. It improves the video quality without any increase in decoder complexity. The effectiveness of the approach is evaluated using both objective and subjective tests. This minimum-drift approach is very simple to implement and can also be applied for memory reduction of a full resolution HDTV decoder. Osama K. Al-Shaykh, Homer H. Chen |
ACM Multimedia (1) | 2 |
| 1999 | Compression of MPEG-4 facial animation parameters for transmission of talking headsabstractThe emerging MPEG-4 standard supports the transmission and composition of facial animation with natural video. The new standard will include a facial animation parameter (FAP) set that is defined based on the study of minimal facial actions and is closely related to muscle actions. The FAP set enables model-based representation of natural or synthetic talking-head sequences and allows intelligible visual reproduction of facial expressions, emotions, and speech pronunciations at the receiver. This paper addresses the data-compression issue of talking heads and presents three methods for bit-rate reduction of FAPs. Compression efficiency is achieved by way of transform coding, principal component analysis, and FAP interpolation. These methods are independent of each other in nature and thus can be applied in combination to lower the bit-rate demand of FAPs, making possible the transmission of multiple talking heads over band-limited channels. The basic methods described here have been adopted into the MPEG-4 Visual Committee Draft and are readily applicable to other articulation data such as body animation parameters. The efficacy of the methods is demonstrated by both subjective and objective results. Homer H. Chen, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | Robust image compression with packetization: the MPEG-4 still texture caseabstractIn this paper, we propose a bit-stream packetization approach to make the MPEG-4 still texture bit-stream robust to channel degradation. Our approach does not affect the spatial/quality scalability features of the still texture algorithm, and it has limited effect on coding efficiency. Also, this packetization approach requires minimal changes to the syntax and a small overhead. Finally, the proposed technique is shown to provide good error robustness over a wide range of error conditions. Iole Moccagatta, Shankar L. Regunathan, Osama K. Al-Shaykh, Homer H. Chen |
MMSP | 4 |
| 1997 | Analysis and compression of facial animation parameter set (FAPs)abstractIn this paper, a new representation of FAPs based on principal component analysis is proposed. Based on this compact representation, a FAPs compression scheme is designed. A facial expression recognition algorithm using recurrent neural network is also investigated. The inputs to the network are the most significant components of this new data representation. Experimental results show that computational complexity is reduced and expressions can be correctly recognized even with changed sampling rate. Homer H. Chen, Thomas S. Huang |
MMSP | 2 |
| 1995 | Speech recognition for image animation and codingabstractWe discuss some issues related to acoustic assisted image coding and animation. An approach of talker independent acoustic assisted image coding and animation scheme is studied. A perceptually based sliding window encoder is proposed. It utilizes the high rate (or oversampled) viseme sequence from the audio domain for image domain viseme interpolation and smoothing. The image domain visemes in our approach are dynamically constructed from a set of basic visemes. The look-ahead and look-back moving interpolations in the proposed approach provide an effective way to compensate the mismatch between auditory and visual perceptions. Wu Chou, Homer H. Chen |
ICASSP | 2 |
| 1995 | Speech-assisted lip synchronization in audio-visual communicationsabstractWe utilize speech information to improve the quality of audio-visual communications such as video telephony and videoconferencing. We show that the marriage of speech analysis and image processing can solve problems related to lip synchronization. We present a technique called speech-assisted frame-rate conversion, and apply it to coding of talking head video. Demonstration sequences are presented. Extensions and other applications are outlined. Tsuhan Chen, Hans Peter Graf, Barry G. Haskell, Eric Petajan, Yao Wang 0001, Homer H. Chen, Wu Chou |
ICIP | 6 |
| 1995 | A Region Based Motion Compensated Video Codec for Very Low Bitrate ApplicationsabstractA motion compensated video coding technique is discussed where square macro-blocks in the coding process are replaced by arbitrary shaped regions. Different components of this technique are presented in detail, such as: segmentation, region shape coding, motion estimation, and mode decisions. Simulation results show good quality images with sharper details achieved at bitrates as low as 9.6 Kb/s. Touradj Ebrahimi, Homer H. Chen, Barry G. Haskell |
ISCAS | 2 |
| 1994 | A Block Transform Coder for Arbitarily Shaped Image SegmentsabstractThis paper describes a method for coding arbitrarily shaped image segments. The method uses an iterative technique based on the theory of successive projection onto convex sets to determine the best transform coefficients. It uses block transforms with frequency domain region-zeroing and space domain region-enforcing operations for effective coding of image segments of arbitrary shape. A major strength of this method is that it can be implemented in real-time using existing codec hardware at an insignificant additional cost.> Homer H. Chen, M. Reha Civanlar, Barry G. Haskell |
ICIP (1) | 1 |
| 1992 | Comments on 'Comments on "Calibration of wrist-mounted robotic sensors by solving homogeneous transform equations of the form AX=XB" ' [with reply]abstractIn the above-named work (ibid., vol.7, p.877-8, (Dec. 1991)), H. Zhuang and Z. S. Roth point out that a particular solution can be simplified by using quaternions to represent rotations. While it is true that this approach gives rise to an algorithm more efficient than the one considered, the commenter argues that it is incorrect to claim that a unique solution for R/sub X/ exists if and only if the axes of rotation of R/sub A1/ and R/sub A2/ are nonzero (theorem 1 of the original work). A counterexample is presented to prove this point. In replying, Zhuang and Roth note the error in the original work and provide an analysis leading to the revision of their original theorem 1.> Homer H. Chen, Hanqi Zhuang, Zvi S. Roth |
IEEE Trans. Robotics Autom. | 1 |
| 1991 | A screw motion approach to uniqueness analysis of head-eye geometryabstractThe screw motion theory is used to solve a class of pose determination problems that can be characterized by a homogeneous transform equation of the form AX=XB, where A and B are known motions and X is an unknown coordinate transformation. Unlike existing methods, this method gives rise to a sound geometric interpretation that takes both rotation and translation into consideration. The author derives a screw congruence theorem and shows that the problem is to find a rigid transformation which will bring one group of lines to overlap another. He also provides a complete analysis of the conditions under which the solution can be uniquely determined.> Homer H. Chen |
CVPR | 1 |
| 1991 | Determining motion and depth from binocular orthographic views
Homer H. Chen |
CVGIP Image Underst. | 1 |
| 1991 | Pose Determination from Line-to-Plane Correspondences: Existence Condition and Closed-Form SolutionsabstractA class of pose determination problems in which the sensory data are lines and the corresponding reference data are planes is discussed. The lines considered are different from edge lines in that they are not the intersection of boundary faces of the object. The author describes a polynomial approach that does not require a priori knowledge about the object location. Closed-form solutions for orthogonal, parallel, and coplanar feature configurations of critical importance in real applications are derived. Findings concerning the necessary and sufficient conditions under which the line-to-plane pose determination problem can be solved are described.> Homer H. Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1991 | Using Motion from Orthographic Views to Verify 3-D Point MatchesabstractThe specific problem addressed by the authors is how to detect the true match of a fourth point from among candidate matches in a situation in which three points have already been matched. The two sets of points to be matched are both subject to measurement errors. The depth error is more dominant than errors in the other two coordinates; however, the exact statistical distribution of the measurement errors is not known. The authors present a new method for solving the problem. The method is based on the technique of motion analysis using orthographic views. It discards the noisy z (depth) coordinates and uses only the x and y coordinates of the points to verify the match. The effect of depth errors on the motion estimate is completely prevented. Results show that this method is substantially more effective than previous methods that use all three coordinates.> Homer H. Chen, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1990 | Pose determination from line-to-plane correspondences: existence condition and closed-form solutionsabstractConsideration is given to a specific pose determination problem in which the sensory features are lines and the matched reference features are planes. The lines discussed are different from edge lines of an object in that they are not the intersection of boundary faces of the object. The author describes a polynomial method that, unlike previous methods, does not require prior knowledge about the location of the object. Closed-form solutions for orthogonal, coplanar, and parallel feature configurations of critical importance in real applications are derived. Basic findings concerning the necessary and sufficient conditions under which the pose determination problem can be solved are presented.> Homer H. Chen |
ICCV | 1 |
| 1990 | Matching 3-D Line Segments with Applications to Multiple-Object Motion EstimationabstractA two-stage algorithm for matching line segments using three-dimensional data is presented. In the first stage, a tree-search based on the orientation of the line segments is applied to establish potential matches. the sign ambiguity of line segments is fixed by a simple congruency constraint. In the second stage, a Hough clustering technique based on the position of line segments is applied to verify potential matches. Any paired line segments of a match that cannot be brought to overlap by the translation determined by the clustering are removed from the match. Unlike previous methods, this algorithm combats noise more effectively, and ensures the global consistency of a match. While the original motivation for the algorithm is multiple-object motion estimation from stereo image sequences, the algorithm can also be applied to other domains, such as object recognition and object model construction from multiple views.> Homer H. Chen, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1988 | Motion And Depth From Binocular Orthographic ViewsabstractThe paper describes a method using orthographic views that are taken from two cameras to determine the motion and depth of a rigid object. The cameras are widely separated in space and hence they may observe different object points. By tracking the object points for each camera, we show how the motion and depth of the object can be determined in two views when the relative position and orientation between the cameras are known. Both exact and noisy cases are discussed. Homer H. Chen |
ICCV | 1 |
| 1988 | A survey of construction and manipulation of octrees
Homer H. Chen, Thomas S. Huang |
Comput. Vis. Graph. Image Process. | 1 |
| 1988 | Maximal matching of 3-D points for multiple-object motion estimation
Homer H. Chen, Thomas S. Huang |
Pattern Recognit. | 1 |