EDBT 2026 Demo / reviewers in the wild / expert
Chuohao Yeo
dblp:41/5910
· DBLP profile ↗
45ranked-venue papers
19as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 17 first-authorArtificial intelligence and machine learning · 4 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
9 papers |
Image and video coding · 29% Image and video processing · 23% Computational photography and imaging · 21% | |
| Human-computer interaction and pervasive computing
2 papers |
Collaborative and social computing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational photography and imaging
intrinsic image decomposition |
0.3 | 2 | 2013 | Intrinsic Image Decomposition Using a Sparse Representation of Reflectance · IEEE Trans. Pattern Anal. Mach. Intell. 2013 Intrinsic images decomposition using a local and global sparse representation of reflectance · CVPR 2011 |
Image and video coding
image quality assessment |
0.2 | 1 | 2013 | A Perceptually Relevant MSE-Based Image Quality Metric · IEEE Trans. Image Process. 2013 |
Image and video processing
image restoration |
0.2 | 1 | 2013 | A Perceptually Relevant MSE-Based Image Quality Metric · IEEE Trans. Image Process. 2013 |
Image and video coding › image quality assessment › perceptual quality metric
perceptual image quality metric |
0.2 | 1 | 2013 | A Perceptually Relevant MSE-Based Image Quality Metric · IEEE Trans. Image Process. 2013 |
Computational photography and imaging › intrinsic image decomposition
reflectance and shading |
0.2 | 1 | 2013 | Intrinsic Image Decomposition Using a Sparse Representation of Reflectance · IEEE Trans. Pattern Anal. Mach. Intell. 2013 |
Image and video processing
sparse representation |
0.2 | 1 | 2013 | Intrinsic Image Decomposition Using a Sparse Representation of Reflectance · IEEE Trans. Pattern Anal. Mach. Intell. 2013 |
Image and video processing › image filtering › optimal filter design
wiener filtering |
0.2 | 1 | 2013 | A Perceptually Relevant MSE-Based Image Quality Metric · IEEE Trans. Image Process. 2013 |
Multimedia analysis and retrieval
feature coding |
0.1 | 1 | 2011 | Coding of Image Feature Descriptors for Distributed Rate-efficient Visual Correspondences · Int. J. Comput. Vis. 2011 |
Image and video coding
multiview video coding |
0.1 | 1 | 2010 | Robust Distributed Multiview Video Compression for Wireless Camera Networks · IEEE Trans. Image Process. 2010 |
Audio and music processing › speaker diarization
audio-visual speaker diarization |
0.1 | 1 | 2009 | Visual speaker localization aided by acoustic models · ACM Multimedia 2009 |
Audio and music processing
speaker diarization |
0.1 | 1 | 2009 | Visual speaker localization aided by acoustic models · ACM Multimedia 2009 |
Image and video coding
distributed source coding |
0.1 | 1 | 2008 | Toward Compression of Encrypted Images and Video Sequences · IEEE Trans. Inf. Forensics Secur. 2008 |
Multimedia analysis and retrieval › multimodal communication
social signal processing |
0.1 | 1 | 2008 | Predicting the dominant clique in meetings through fusion of nonverbal cues · ACM Multimedia 2008 |
Image and video coding
video compression |
0.1 | 1 | 2008 | Toward Compression of Encrypted Images and Video Sequences · IEEE Trans. Inf. Forensics Secur. 2008 |
Storage systems › file systems
file synchronization |
0.1 | 1 | 2008 | VSYNC: a novel video file synchronization protocol · ACM Multimedia 2008 |
Multimedia analysis and retrieval › multimodal fusion
multimodal meeting analysis |
0.1 | 1 | 2007 | Using audio and video features to classify the most dominant person in a group meeting · ACM Multimedia 2007 |
Multimedia analysis and retrieval › multimedia analysis › visual content analysis
visual correspondence |
0.0 | 1 | 2011 | Coding of Image Feature Descriptors for Distributed Rate-efficient Visual Correspondences · Int. J. Comput. Vis. 2011 |
Image and video coding
error resilience |
0.0 | 1 | 2010 | Robust Distributed Multiview Video Compression for Wireless Camera Networks · IEEE Trans. Image Process. 2010 |
Internet of things and sensor networks › camera sensor networks › camera networks
wireless camera networks |
0.0 | 1 | 2010 | Robust Distributed Multiview Video Compression for Wireless Camera Networks · IEEE Trans. Image Process. 2010 |
Collaborative and social computing › computer-supported cooperative work › work practice
meeting analysis |
0.0 | 1 | 2009 | Modeling Dominance in Group Conversations Using Nonverbal Activity Cues · IEEE Trans. Speech Audio Process. 2009 |
Information theory › information measures › entropy
entropy rate |
0.0 | 1 | 2008 | Toward Compression of Encrypted Images and Video Sequences · IEEE Trans. Inf. Forensics Secur. 2008 |
Coding theory › source coding
universal coding |
0.0 | 1 | 2008 | Toward Compression of Encrypted Images and Video Sequences · IEEE Trans. Inf. Forensics Secur. 2008 |
Methods — techniques the papers use, named apart from their topics
distributed source coding · 0.5rsync · 0.2hierarchical hashing · 0.2disparity estimation · 0.2structural similarity index · 0.2second-generation wavelet representation · 0.2mean squared error · 0.2l1-norm minimization · 0.2sparse representation · 0.1rate-distortion optimization · 0.1l1-regularized least squares · 0.1edge-avoiding wavelets · 0.1view interpolation · 0.1epipolar geometry · 0.1unsupervised learning · 0.1supervised learning · 0.1audiovisual activity cues · 0.1statistical modeling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | High-throughput and low-cost hardware-oriented integer transforms for HEVCabstractTo achieve a target bit-rate reduction of 50% over H.264/AVC while maintaining equivalent perceptual video quality, High Efficiency Video Coding (HEVC) includes several new coding tools including a new set of integer transforms. Since these transforms are more complex than the H.264/AVC transforms, it is more challenging to design and develop high-performance integer transform hardware for HEVC. In this paper, we propose a series of high-throughput and low-cost hardware-oriented HEVC transform algorithms by using a butterfly structure and replacing multiplications by additions and shift operations in a way that minimizes critical computation paths. Compared to the algorithms using other methodologies like multiplierless multiple constant multiplication (MMCM) or decomposition to sparse matrices, our algorithms achieve around 20% shorter critical paths while consuming relatively small numbers of additions and shift operations. Hardware designs applying our proposed algorithms can increase their throughput by 20% while maintaining a reasonable resource consumption compared to when applying other algorithms. Trang T. T. Do 0001, Yih Han Tan, Chuohao Yeo |
ICIP | 3 |
| 2014 | Efficient Integer DCT Architectures for HEVCabstractIn this paper, we present area- and power-efficient architectures for the implementation of integer discrete cosine transform (DCT) of different lengths to be used in High Efficiency Video Coding (HEVC). We show that an efficient constant matrix-multiplication scheme can be used to derive parallel architectures for 1-D integer DCT of different lengths. We also show that the proposed structure could be reusable for DCT of lengths 4, 8, 16, and 32 with a throughput of 32 DCT coefficients per cycle irrespective of the transform size. Moreover, the proposed architecture could be pruned to reduce the complexity of implementation substantially with only a marginal affect on the coding performance. We propose power-efficient structures for folded and full-parallel implementations of 2-D DCT. From the synthesis result, it is found that the proposed architecture involves nearly 14% less area-delay product (ADP) and 19% less energy per sample (EPS) compared to the direct implementation of the reference algorithm, on average, for integer DCT of lengths 4, 8, 16, and 32. Also, an additional 19% saving in ADP and 20% saving in EPS can be achieved by the proposed pruning algorithm with nearly the same throughput rate. The proposed architecture is found to support ultrahigh definition 7680 × 4320 at 60 frames/s video, which is one of the applications of HEVC. Pramod Kumar Meher, Basant K. Mohanty, Khoon Seong Lim, Chuohao Yeo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Perceptually relevant energy function for seam carvingabstractSeam carving, an image re-targeting method, works by progressively finding and removing connected paths of low energy pixels in an image until a desired image aspect ratio is reached. In this paper, we first cast the problem of minimizing an energy function as that of minimizing a distortion cost. We then leverage on our understanding of image quality metrics/ distortion metrics in proposing a perceptually relevant energy function. Experimental results show that our proposed energy function can generate more desirable resized images in which the original structures of the images are better preserved. Hui Li Tan, Yih Han Tan, Zhengguo Li, Susanto Rahardja, Chuohao Yeo |
ICASSP | 5 |
| 2013 | Residual DPCM for lossless coding in HEVCabstractIncorporating sample-based prediction during lossless coding can significantly improve coding performance. However, its use within a codec designed for lossy coding requires a modification of the available prediction scheme. When implementing the codec, two different prediction processes will have to be implemented. This paper describes a lossless coding scheme that delays the sample-based prediction till the residue coding stage of the codec and carries out prediction in the residual domain. In this way, the prediction scheme of the lossy coder can be retained while realizing the coding gains associated with sample-based prediction. The proposed scheme improves lossless intra coding performance in HEVC Main Profile by an average of 6.5%. Yih Han Tan, Chuohao Yeo, Zhengguo Li |
ICASSP | 2 |
| 2013 | SSIM-based adaptive quantization in HEVCabstractHEVC is an emerging video coding standard that can achieve significant compression gains compared to H.264/AVC due to the inclusion of numerous new coding tools. In particular, it allows for a flexible quadtree based block partitioning of each coding tree unit (CTU) and an ability to switch quantization parameters (QP) on a sub-CTU level. In this paper, we present an approach for selecting quantization parameters for each block of pixels on the basis of optimizing the SSIM of the entire picture. Our simulation results show that when SSIM is the quality metric, the proposed approach is able to give average BD-Rate gains of 5.5% to 7.4% compared to using a constant QP per picture while having a negligible increase in encoding runtime. In addition, our proposed method also significantly outperforms the MPEG-2 TM5 adaptive quantization algorithm implemented in the HEVC reference software. Chuohao Yeo, Hui Li Tan, Yih Han Tan |
ICASSP | 1 |
| 2013 | Intrinsic Image Decomposition Using a Sparse Representation of ReflectanceabstractIntrinsic image decomposition is an important problem that targets the recovery of shading and reflectance components from a single image. While this is an ill-posed problem on its own, we propose a novel approach for intrinsic image decomposition using reflectance sparsity priors that we have developed. Our sparse representation of reflectance is based on a simple observation: Neighboring pixels with similar chromaticities usually have the same reflectance. We formalize and apply this sparsity constraint on local reflectance to construct a data-driven second-generation wavelet representation. We show that the reflectance component of natural images is sparse in this representation. We further propose and formulate a global sparse constraint on reflectance colors using the assumption that each natural image uses a small set of material colors. Using this sparse reflectance representation and the global constraint on a sparse set of reflectance colors, we formulate a constrained l₁-norm minimization problem for intrinsic image decomposition that can be solved efficiently. Our algorithm can successfully extract intrinsic images from a single image without using color models or any user interaction. Experimental results on a variety of images demonstrate the effectiveness of the proposed technique. Li Shen 0003, Chuohao Yeo, Binh-Son Hua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Intra Coding With Adaptive Partial ReconstructionabstractIntra prediction improves coding performance by reducing inter pixel redundancy. However, to accommodate the use of block transforms, not all pixels can be predicted from reconstructed pixels that are located close to themselves. This causes prediction performance to suffer as pixel values further apart are less correlated. This paper presents additional intra coding modes designed with the goal of improving prediction performance. Experimental results show an average gain of about 2% in the key technical area software when the new modes are incorporated in the current 8$\,\times\,$8 prediction modes. Since the new coding modes (8$\,\times\,$8) are designed with transform size smaller than coding block size, the modes can also be useful when the source block is larger than the maximum transform size. Yih Han Tan, Chuohao Yeo, Zhengguo Li, Susanto Rahardja |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Dynamic Range Analysis in High Efficiency Video Coding Residual Coding and ReconstructionabstractWe present a method for analyzing dynamic range along the residual coding and reconstruction pathways during video coding by deriving bounds on the maximum absolute value of intermediary data using a simple combination of triangle inequality and reformulation using Kronecker products. The proposed method is then applied toward analyzing the residual coding and reconstruction process in the emerging High Efficiency Video Coding (HEVC) standard. Our analysis shows that, for an input residual with a bitdepth of (B+1) that uses a uniform quantizer, the dynamic range of quantized levels and dequantized coefficients are no more than (B+7) bits and 17 bits, respectively. Furthermore, a 16 bits transpose buffer is sufficient, while up to 5 bits of bitdepth expansion can occur in the reconstructed residual. The analysis is validated by simulation results with both randomly generated residual and encoding/decoding test video sequences using the HEVC reference software. Chuohao Yeo, Yih Han Tan, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | On Rate Distortion Optimization Using SSIMabstractIn this paper, we present a method for performing rate-distortion optimization (RDO) using a perceptual visual quality metric, the structural similarity index (SSIM), as the target of optimization. Rate-distortion optimization is widely used in modern video codecs to make various encoder decisions to optimize the rate-distortion tradeoff. Typically, the distortion measure used is either sum-of-square error or sum-of-absolute distance, both of which are convenient when used in the RDO framework but not always reflective of a perceptual visual quality. We show that SSIM can be used as the distortion metric in the RDO framework in a simple, yet effective, manner by scaling the Lagrange multiplier used in RDO based on the local variance in that region. The experimental results on the H.264/AVC reference software show that compared to traditional RDO approaches, for the same SSIM score, the proposed approach can achieve an average rate reduction of about 9% and 14% for random access and low-delay encoding configurations. At the same time, there is no significant change in the encoding runtime. Chuohao Yeo, Hui Li Tan, Yih Han Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | A Perceptually Relevant MSE-Based Image Quality MetricabstractImage quality metrics (IQMs), such as the mean squared error (MSE) and the structural similarity index (SSIM), are quantitative measures to approximate perceived visual quality. In this paper, through analyzing the relationship between the MSE and the SSIM under an additive noise distortion model, we propose a perceptually relevant MSE-based IQM, MSE-SSIM, which is expressed in terms of the variance of the source image and the MSE between the source and distorted images. Evaluations on three publicly available databases (LIVE, CSIQ, and TID2008) show that the proposed metric, despite requiring less computation, compares favourably in performance to several existing IQMs. In addition, due to its simplicity, MSE-SSIM is amenable for the use in a wide range of image and video tasks that involve solving an optimization problem. As an example, MSE-SSIM is used as the objective function in designing a Wiener filter that aims at optimizing the perceptual visual quality of the output. Experimental results show that the images filtered with a MSE-SSIM-optimal Wiener filter have better visual quality than those filtered with a MSE-optimal Wiener filter. Hui Li Tan, Zhengguo Li, Yih Han Tan, Susanto Rahardja, Chuohao Yeo |
IEEE Trans. Image Process. | 5 |
| 2012 | A local intensity adaptive structural similarity indexabstractExisting structural similarity (SSIM) index comprises of one term on luminance comparison and the other term on contrast and structure comparison. In this paper, the SSIM index is first improved by introducing three weighting factors to the second term such that it is adaptive to local intensities of two images to be compared. The improved SSIM (iSSIM) index is further extended for two images with possibly different exposures. Experimental results show that the proposed indices are more robust to large intensity changes of two images from the same scene and more sensitive to two images from different scenes than the existing SSIM index. Zhengguo Li, Chuohao Yeo, Yih Han Tan, Susanto Rahardja |
ICASSP | 2 |
| 2012 | On fast coding tree block and mode decision for high-Efficiency Video Coding (HEVC)abstractIn the current HEVC test model (HM), a quad-tree based coding tree block (CTB) representation is used to signal mode, partition, prediction and residual information. The large number of combinations of quad-tree partitions and modes to be tested during rate-distortion optimization (RDO) results in a high encoding complexity. In this paper, we investigate and compare a variety of algorithms for fast CTB and mode decision. Experimental results from HM4-based implementations show that different strategies can provide a range of complexity-performance trade-offs. In particular, our proposed CU Depth Pruning algorithm can reduce encoding time by about 10% with only 0.1% coding loss, while a combination of our proposed Early Partition Decision and an early CU termination approach can reduce encoding time by about 40% with about 1% coding loss. Hui Li Tan, Fengjiao Liu, Yih Han Tan, Chuohao Yeo |
ICASSP | 4 |
| 2012 | On rate distortion optimization using SSIMabstractRate-distortion optimization is widely used in modern video codecs to make various encoder decisions in order to optimize the rate-distortion trade-off. Typically, the distortion measure used is either sum-of-square error (SSE) or sum-of-absolute distance (SAD), both of which are convenient when used in the RDO framework but not always reflective of perceptual quality. In this paper, we show that by expressing SSIM in terms of SSE, SSIM can be used as the distortion metric in the RDO framework in an effective and efficient manner by simply scaling the Lagrange multiplier used in RDO based on the local variance in that region without further changes to the RDO engine. Experimental results show that compared to traditional RDO approaches, for the same SSIM score, the proposed approach can achieve an average rate decrease of 8% and 11% for random access and low-delay encoding configurations, with no significant change in encoding runtime. Chuohao Yeo, Hui Li Tan, Yih Han Tan |
ICASSP | 1 |
| 2012 | Single-Pass Rate Control With Texture and Non-Texture Rate-Distortion ModelsabstractOne of the challenges in video rate control lies in determining a quantization parameter (Qp) that will be used for both the rate-distortion (R-D) optimization process and the quantization of transform coefficients. In this paper, we attempt to achieve effective rate control with a different approach. By modeling the relationships of distortion, texture bits, non-texture bits, and Qp, we derive the Qp required for both R-D optimization and quantization through Lagrangian optimization. From experiments with several video sequences, we found that our rate control scheme is capable of effective rate control with only a few model updates during encoding. The proposed rate control scheme adapts quickly to the characteristics of the source data and is particularly effective at controlling the rate of videos with high and unpredictable motion content. Yih Han Tan, Chuohao Yeo, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Mode-Dependent Transforms for Coding Directional Intra Prediction ResidualsabstractThe use of mode-dependent transforms for coding directional intra prediction residuals has been previously shown to provide coding gains, but the transform matrices have to be derived from training. In this paper, we derive a set of separable mode-dependent transforms by using a simple separable, directional, and anisotropic image correlation model. Our analysis shows that only one additional transform, the odd type-3 discrete sine transform (ODST-3), is required for the optimal implementation of mode-dependent transforms. In addition, the four-point ODST-3 also has a structure that can be exploited to reduce the operation count of the transform operation. Experimental results show that in terms of coding efficiency, our proposed approach matches or improves upon the performance of a mode-dependent transforms approach that uses transform matrices obtained through training. Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | VSYNC: Bandwidth-Efficient and Distortion-Tolerant Video File SynchronizationabstractWe introduce video-sync (VSYNC), a video file synchronization system that efficiently uses a bidirectional communications link to maintain up-to-date video sources at remote ends to a desired resolution and distortion level. By automatically detecting and transmitting only the differences between video files, VSYNC is able to avoid unnecessary re-transmission of the entire video when there are only minor differences between video copies. A hierarchical hashing scheme is designed to allow synchronization to within some user-defined distortion, white being rate-efficient and computationally tractable. Distributed video coding is used to realize further rate savings when transmitting video updates. VSYNC is bandwidth-efficient and is useful in many scenarios including video backup, video sharing, and video authentication applications. Experimental results show that rate-savings ranging from 2× to 10× can be obtained by VSYNC with about 10% of the frames being edited, compared to re- transmitting the compressed video or using a file synchronization utility such as rsync. Hao Zhang 0006, Chuohao Yeo, Kannan Ramchandran |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Intrinsic images decomposition using a local and global sparse representation of reflectanceabstractIntrinsic image decomposition is an important problem that targets the recovery of shading and reflectance components from a single image. While this is an ill-posed problem on its own, we propose a novel approach for intrinsic image decomposition using a reflectance sparsity prior that we have developed. Our method is based on a simple observation: neighboring pixels usually have the same reflectance if their chromaticities are the same or very similar. We formalize this sparsity constraint on local reflectance, and derive a sparse representation of reflectance components using data-driven edge-avoiding-wavelets. We show that the reflectance component of natural images is sparse in this representation. We also propose and formulate a novel global reflectance sparsity constraint. Using this sparsity prior and global constraints, we formulate a l1-regularized least squares minimization problem for intrinsic image decomposition that can be solved efficiently. Our algorithm can successfully extract intrinsic images from a single image, without using other reflection or color models or any user interaction. The results on challenging scenes demonstrate the power of the proposed technique. Li Shen 0003, Chuohao Yeo |
CVPR | 2 |
| 2011 | Quadratic optimization based small scale details extractionabstractIn many image processing problems, it is required to extract small scale details from an image or a set of images. In this paper, we introduce a new framework for extracting small scale details from a single input image or a set of input images. We then show how to apply the framework to address several important problems in the field of image processing including tone mapping of high dynamic range images, de-noising of a non-flash image with a pair of non flash and flash images, as well as details enhancement via multi light images and a single input image. Experimental results show that the proposed framework outperforms existing methods. Zhengguo Li, Jinghong Zheng 0001, Chuohao Yeo, Susanto Rahardja |
ICASSP | 3 |
| 2011 | Intra-prediction with adaptive sub-samplingabstractIntra-prediction improves coding performance by reducing inter-pixel redundancy. However, to accommodate the use of block transforms, not all pixels can be predicted from reconstructed pixels that are located close to themselves. This causes prediction performance to suffer as pixel values further apart are less correlated. This paper presents additional intra-prediction modes designed with the goal of improving prediction performance. Experimental results show an average gain of about 2% in KTA when the new modes are incorporated in the current 8×8 prediction modes. Since the new coding modes (8×8) are designed with transform size smaller than coding block size, the modes can also be useful when the prediction unit is larger than the maximum transform size. The use of smaller transform sizes can potentially lead to reduction of decoder complexity and implementation costs. Yih Han Tan, Chuohao Yeo, Zhengguo Li, Susanto Rahardja |
ICIP | 2 |
| 2011 | Low-complexity mode-dependent KLT for block-based intra codingabstractApplying mode-dependent separable transforms, e.g., mode- dependent directional transform (MDDT), is an effective method for improving transform coding of intra prediction residuals. However, two transform matrices typically need to be stored for each intra prediction mode. By using a simple image correlation mode, we have previously derived and proposed a simplified mode-dependent separable transforms scheme that uses a combination of two well-known trans- forms: Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST). In this paper, we propose an orthogonal 4-point integer DST that has a multiplier-less implementation consisting of only adds and bit-shifts. We also propose a simple set of mode-dependent scans for coefficient coding that can be used on top of mode-dependent transforms. Our experimental results on the current HEVC reference software show that in terms of coding efficiency, our proposed approach has comparable performance to MDDT. More importantly, compared to MDDT, our approach requires no training and has lower computational and storage costs. Chuohao Yeo, Yih Han Tan, Zhengguo Li |
ICIP | 1 |
| 2011 | Chroma intra prediction using template matching with reconstructed luma componentsabstractIntra coding in the current H.264/AVC video coding standard achieves high compression efficiency, in part due to the highly effective intra prediction process that exploits spatial directional correlation. However, intra prediction of chroma components in YUV 4:2:0 videos uses a limited set of possible predictions available for coding of luma components. Furthermore, coding of chroma components proceeds somewhat independently of luma components, and ignores any possible correlation between them. In this paper, we show a way of using reconstructed luma pixels to help with intra prediction of chroma pixels. By making use of the reconstructed co-located luma block to perform template matching in the luma plane, we are able to use as predictors the co-located chroma blocks of the matched luma blocks. Simulations results indicate that the proposed approach is able to obtain up to 33% chroma bit-rate reduction and up to 8% overall bit-rate reduction over H.264/AVC. Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja |
ICIP | 1 |
| 2011 | Mode-dependent fast separable KLT for block-based intra codingabstractIn this paper, we derive separable KLTs for coding H.264/AVC intra prediction residuals, using a simple image correlation model. Our analysis shows that for some intra prediction modes, we can in fact just use the DCT for performing either the row-wise or column-wise transform. Furthermore, we also compute the KLT that should be used based on the image correlation model, which happens to have sinosuidal terms. The 4×4 transform also has a structure that can be exploited to reduce the operation count of the transform operation. In our simplified implementation of mode-dependent directional transforms (MDDT), we only need to make use of two matrices: the DCT and the derived KLT. Our experimental results show that in terms of coding efficiency, our proposed approach has similar performance when compared with MDDT. More importantly, compared to MDDT, our approach requires no training and has lower computational and storage costs. Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja |
ISCAS | 1 |
| 2011 | On residual quad-tree coding in HEVCabstractIn the current working draft of HEVC, residual quad-tree (RQT) coding is used to encode prediction residuals in both Intra and Inter coding units (CU). However, the rationale for using RQT as a coding tool is different in the two cases. For Intra prediction units, RQT provides an efficient syntax for coding a number of sub-blocks with the same intra prediction mode. For Inter CUs, RQT adapts to the spatial-frequency variations of the CU, using as large a transform size as possible while catering to local variations in residual statistics. While providing coding gains, effective use of RQT currently requires an exhaustive search of all possible combinations of transform sizes within a block. In this paper, we exploit our insights to develop two fast RQT algorithms, each designed to meet the needs of Intra and Inter prediction residual coding. Yih Han Tan, Chuohao Yeo, Hui Li Tan, Zhengguo Li |
MMSP | 2 |
| 2011 | Coding of Image Feature Descriptors for Distributed Rate-efficient Visual Correspondences
Chuohao Yeo, Parvez Ahammad, Kannan Ramchandran |
Int. J. Comput. Vis. | 1 |
| 2010 | Robust Distributed Multiview Video Compression for Wireless Camera NetworksabstractWe present a novel framework for robustly delivering video data from distributed wireless camera networks that are characterized by packet drops. The main focus in this work is on robustness which is imminently needed in a wireless setting. We propose two alternative models to capture interview correlation among cameras with overlapping views. The view-synthesis-based correlation model requires at least two other camera views and relies on both disparity estimation and view interpolation. The disparity-based correlation model requires only one other camera view and makes use of epipolar geometry. With the proposed models, we show how interview correlation can be exploited for robustness through the use of distributed source coding. The proposed approach has low encoding complexity, is robust while satisfying tight latency constraints and requires no intercamera communication. Our experiments show that on bursty packet erasure channels, the proposed H.263+ based method outperforms baseline methods such as H.263+ with forward error correction and H.263+ with intra refresh by up to 2.5 dB. Empirical results further support the relative insensitivity of our proposed approach to the number of additional available camera views or their placement density. Chuohao Yeo, Kannan Ramchandran |
IEEE Trans. Image Process. | 1 |
| 2010 | Dialocalization: Acoustic speaker diarization and visual localization as joint optimization problemabstractThe following article presents a novel audio-visual approach for unsupervised speaker localization in both time and space and systematically analyzes its unique properties. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the-art audio-only speaker diarization system (speaker localization in time) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the speech regions and estimates “who spoke when,” then, in a second step, the visual models are used to infer the location of the speakers in the video. We call this process “dialocalization.” The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audio-only) speaker diarization, but also adds visual speaker localization at little incremental engineering and computation costs. The combined algorithm has different properties, such as increased robustness, that cannot be observed in algorithms based on single modalities. The article describes the algorithm, presents benchmarking results, explains its properties, and systematically discusses the contributions of each modality. Gerald Friedland, Chuohao Yeo, Hayley Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2009 | Multi-modal speaker diarization of real-world meetings using compressed-domain video featuresabstractSpeaker diarization is originally defined as the task of determining ldquowho spoke whenrdquo given an audio track and no other prior knowledge of any kind. The following article shows a multi-modal approach where we improve a state-of-the-art speaker diarization system by combining standard acoustic features (MFCCs) with compressed domain video features. The approach is evaluated on over 4.5 hours of the publicly available AMI meetings dataset which contains challenges such as people standing up and walking out of the room. We show a consistent improvement of about 34% relative in speaker error rate (21% DER) compared to a state-of-the-art audio-only baseline. Gerald Friedland, Hayley Hung, Chuohao Yeo |
ICASSP | 3 |
| 2009 | Rate-constrained distributed distance testing and its applicationsabstractWe investigate a practical approach to solving one instantiation of a distributed hypothesis testing problem under severe rate constraints that shows up in a wide variety of applications such as camera calibration, biometric authentication and video hashing: given two distributed continuous-valued random sources, determine if they satisfy a certain Euclidean distance criterion. We show a way to convert the problem from continuous-valued to binary-valued using binarized random projections and obtain rate savings by applying a linear syndrome code. In finding visual correspondences, our approach uses just 49% of the rate of scalar quantization to achieve the same level of retrieval performance. To perform video hashing, our approach requires only a hash rate of 0.0142 bpp to identify corresponding groups of pictures correctly. Chuohao Yeo, Parvez Ahammad, Hao Zhang 0006, Kannan Ramchandran |
ICASSP | 1 |
| 2009 | Receiver error concealment using acknowledge preview (RECAP) - An approach to resilient video streamingabstractHigh-quality and low-latency video streaming is essential to providing a natural user experience in video conferencing. This is challenging over lossy networks since compressed video is highly fragile while the low-latency requirement limits the effectiveness of traditional error control approaches such as retransmission and forward error correction. In this paper, we advocate a practical solution for low-latency video communications over best-effort networks that employs an additional low-quality, low-resolution but robustly coded copy of the video. This approach, called RECAP, incurs minimal rate overhead, and can be combined with previously decoded frames to achieve effective concealment of isolated and burst losses even under tight delay constraints. RECAP achieves PSNR gains of 2-6 dB against complete frame loss. Chuohao Yeo, Wai-tian Tan, Debargha Mukherjee |
ICASSP | 1 |
| 2009 | Rate efficient remote video file synchronizationabstractVideo file synchronization between remote users is an important task in many applications. Re-transmission of a video that has been only slightly modified is expensive, wasteful and avoidable. We propose a scheme that automatically detects and sends only modified content according to some user defined distortion constraint to enable rate savings under a wide range of video edits. Through the use of a low-rate hierarchical hashing scheme, we can detect modifications with some spatial granularity. We also apply distributed source coding techniques to exploit correlation between remote copies for a further rate rebate. Experimental results show that the proposed approach achieves up to 7times rate reduction when compared to re-transmitting. Hao Zhang 0006, Chuohao Yeo, Kannan Ramchandran |
ICASSP | 2 |
| 2009 | On low-complexity video encoding through feedbackabstractWe consider the design of a low-complexity video encoder in a scenario where an unconstrained receiver can help the encoder reduce complexity by exploiting feedback. To aid us in making design choices, we propose a simple model for video data that captures key properties of the temporal correlation between video frames and the predictability of motion vectors. Our analysis identifies the strengths and weaknesses of previous approaches and motivates a hybrid approach in which the decoder sends motion information to reduce motion estimation workload, the encoder uses incremental distributed source coding on the blocks with poor motion estimates, and the decoder performs decoder motion search and requests additional bits when necessary. Under the model, our analysis shows that such a hybrid scheme outperforms prior approaches. Experimental results of comparable implementations indicate that the hybrid scheme, while not fully implemented, is very promising in practice. Chuohao Yeo, Kannan Ramchandran |
ICIP | 1 |
| 2009 | Visual speaker localization aided by acoustic modelsabstractThe following paper presents a novel audio-visual approach for unsupervised speaker locationing. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the art audio-only speaker localization system (traditionally called speaker diarization) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the number of speakers and estimates "who spoke when", then, in a second step, the visual models are used to infer the location of the speakers in the video. The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audio-only) speaker diarization, but also adds visual speaker locationing at little incremental engineering and computation costs. Gerald Friedland, Chuohao Yeo, Hayley Hung |
ACM Multimedia | 2 |
| 2009 | Modeling Dominance in Group Conversations Using Nonverbal Activity CuesabstractDominance - a behavioral expression of power - is a fundamental mechanism of social interaction, expressed and perceived in conversations through spoken words and audiovisual nonverbal cues. The automatic modeling of dominance patterns from sensor data represents a relevant problem in social computing. In this paper, we present a systematic study on dominance modeling in group meetings from fully automatic nonverbal activity cues, in a multi-camera, multi-microphone setting. We investigate efficient audio and visual activity cues for the characterization of dominant behavior, analyzing single and joint modalities. Unsupervised and supervised approaches for dominance modeling are also investigated. Activity cues and models are objectively evaluated on a set of dominance-related classification tasks, derived from an analysis of the variability of human judgment of perceived dominance in group discussions. Our investigation highlights the power of relatively simple yet efficient approaches and the challenges of audiovisual integration. This constitutes the most detailed study on automatic dominance modeling in meetings to date. Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Rate-efficient visual correspondences using random projectionsabstractWe consider the problem of establishing visual correspondences in a distributed and rate-efficient fashion by broadcasting compact descriptors. Establishing visual correspondences is a critical task before other vision tasks can be performed in a wireless camera network. We propose the use of coarsely quantized random projections of descriptors to build binary hashes, and use the Hamming distance between binary hashes as the matching criterion. In this work, we derive the analytic relationship of Hamming distance between the binary hashes to Euclidean distance between the original descriptors. We present experimental verification of our result, and show that for the task of finding visual correspondences, sending binary hashes is more rate-efficient than prior approaches. Chuohao Yeo, Parvez Ahammad, Kannan Ramchandran |
ICIP | 1 |
| 2008 | Predicting the dominant clique in meetings through fusion of nonverbal cuesabstractThis paper addresses the problem of automatically predicting the dominant clique (i.e., the set of K-dominant people) in face-to-face small group meetings recorded by multiple audio and video sensors. For this goal, we present a framework that integrates automatically extracted nonverbal cues and dominance prediction models. Easily computable audio and visual activity cues are automatically extracted from cameras and microphones. Such nonverbal cues, correlated to human display and perception of dominance, are well documented in the social psychology literature. The effectiveness of the cues were systematically investigated as single cues as well as in unimodal and multimodal combinations using unsupervised and supervised learning approaches for dominant clique estimation. Our framework was evaluated on a five-hour public corpus of teamwork meetings with third-party manual annotation of perceived dominance. Our best approaches can exactly predict the dominant clique with 80.8% accuracy in four-person meetings in which multiple human annotators agree on their judgments of perceived dominance. Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2008 | VSYNC: a novel video file synchronization protocolabstractVSYNC is a novel incremental video file synchronization system that efficiently synchronizes two video files at remote ends through a bi-directional communications link. Retransmission of a video file that has been modified only slightly, for the purpose of synchronization with a remote-end copy, is extremely expensive but avoidable. VSYNC is a bi-directional algorithm designed to automatically detect and transmit changes in the modified video file without the knowledge of what was changed. Another feature of VSYNC is that it allows synchronization to within some user defined distortion constraint. A hierarchical hashing scheme is designed to compare video chunks, converting the high-level content information to a low-level hash stream that is more amenable to the tools of coding theory. Our approach shows impressive gains in transmission rate-savings. In a typical example of two 12 sec video files with about 10% of the frames being edited, transmission savings of 44% to 87% can be obtained compared to directly sending the updated video files using H.264 and rsync [1]. Hao Zhang 0006, Chuohao Yeo, Kannan Ramchandran |
ACM Multimedia | 2 |
| 2008 | High-Speed Action Recognition and Localization in Compressed Domain VideosabstractWe present a compressed domain scheme that is able to recognize and localize actions at high speeds. The recognition problem is posed as performing an action video query on a test video sequence. Our method is based on computing motion similarity using compressed domain features which can be extracted with low complexity. We introduce a novel motion correlation measure that takes into account differences in motion directions and magnitudes. Our method is appearance-invariant, requires no prior segmentation, alignment or stabilization, and is able to localize actions in both space and time. We evaluated our method on a benchmark action video database consisting of six actions performed by 25 people under three different scenarios. Our proposed method achieved a classification accuracy of 90%, comparing favorably with existing methods in action classification accuracy, and is able to localize a template video of 80 x 64 pixels with 23 frames in a test video of 368 x 184 pixels with 835 frames in just 11 s, easily outperforming other methods in localization speed. We also perform a systematic investigation of the effects of various encoding options on our proposed approach. In particular, we present results on the compression-classification tradeoff, which would provide valuable insight into jointly designing a system that performs video encoding at the camera front-end and action classification at the processing back-end. Chuohao Yeo, Parvez Ahammad, Kannan Ramchandran, S. Shankar Sastry |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Toward Compression of Encrypted Images and Video SequencesabstractWe present a framework for compressing encrypted media, such as images and videos. Encryption masks the source, rendering traditional compression algorithms ineffective. By conceiving of the problem as one of distributed source coding, it has been shown in prior work that encrypted data are as compressible as unencrypted data. However, there are two major challenges to realize these theoretical results. The first is the development of models that capture the underlying statistical structure and are compatible with our framework. The second is that since the source is masked by encryption, the compressor does not know what rate to target. We tackle these issues in this paper. We first develop statistical models for images before extending it to videos, where our techniques really gain traction. As an illustration, we compare our results to a state-of-the-art motion-compensated lossless video encoder that requires unencrypted video input. The latter compresses each unencrypted frame of the ldquoForemanrdquo test sequence by 59% on average. In comparison, our proof-of-concept implementation, working on encrypted data, compresses the same sequence by 33%. Next, we develop and present an adaptive protocol for universal compression and show that it converges to the entropy rate. Finally, we demonstrate a complete implementation for encrypted video. Daniel Schonberg, Stark C. Draper, Chuohao Yeo, Kannan Ramchandran |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2007 | On Compression of Encrypted VideoabstractWe consider video sequences that have been encrypted uncompressed. Since encryption masks the source, traditional data compression algorithms are rendered ineffective. However, it has been shown that through the use of distributed source-coding techniques, the compression of encrypted data is in fact possible. This means that it is possible to reduce data size without requiring that the data be compressed prior to encryption. Indeed, under some reasonable conditions, neither security nor compression efficiency need be sacrificed when compression is performed on the encrypted data (Johnson et al., 2004). In this paper we develop an algorithm for the practical lossless compression of encrypted gray scale video. Our method is based on considering the temporal correlations in the video. This move to temporal dependence builds on our previous work on memoryless sources, and one- and two-dimensional Markov sources. For comparison, a motion-compensated lossless video encoder can compress each unencrypted frame of the standard "Foreman" test video sequence by about 57%. Our algorithm can compress the same frames, after encryption, by about 33% Daniel Schonberg, Chuohao Yeo, Stark C. Draper, Kannan Ramchandran |
DCC | 2 |
| 2007 | View Synthesis for Robust Distributed Video Compression in Wireless Camera NetworksabstractWe propose a method for delivering error-resilient video from wireless camera networks in a distributed fashion over lossy channels. Our scheme is based on distributed source coding that exploits inter-view correlation among cameras with overlapping views. The main focus in this work is on robustness which is imminently needed in a wireless setting. The proposed approach has low encoding complexity, is robust while satisfying tight latency constraints, and requires no inter-camera communication. Our system is built on and is a multi-camera extension of PRISM[1], an earlier proposed single-camera distributed video compression system. Decoder motion search, a key attribute of single-camera PRISM, is extended to the multi-view setting by using estimated scene depth information when it is available. In particular, dense stereo correspondence and view synthesis are utilized to generate side-information. When combined with decoder motion search, our proposed method can be made insensitive to small errors in camera calibration, disparity estimation and view synthesis. In experiments over a simulated wireless channel, the proposed approach achieves up to 2.1 dB gain in PSNR over a system using H.263+ with forward error correction. Chuohao Yeo, Kannan Ramchandran |
ICIP (3) | 1 |
| 2007 | Using audio and video features to classify the most dominant person in a group meetingabstractThe automated extraction of semantically meaningful information from multi-modal data is becoming increasingly necessary due to the escalation of captured data for archival. A novel area of multi-modal data labelling, which has received relatively little attention, is the automatic estimation of the most dominant person in a group meeting. In this paper, we provide a framework for detecting dominance in group meetings using different audio and video cues. We show that by using a simple model for dominance estimation we can obtain promising results. Hayley Hung, Dinesh Babu Jayagopi, Chuohao Yeo, Gerald Friedland, Sileye O. Ba, Jean-Marc Odobez, Kannan Ramchandran, Nikki Mirghafori, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2007 | Unsupervised Discovery of Action Hierarchies in Large Collections of Activity VideosabstractGiven a large collection of videos containing activities, we investigate the problem of organizing it in an unsupervised fashion into a hierarchy based on the similarity of actions embedded in the videos. We use spatio-temporal volumes of filtered motion vectors to compute appearance-invariant action similarity measures efficiently - and use these similarity measures in hierarchical agglomerative clustering to organize videos into a hierarchy such that neighboring nodes contain similar actions. This naturally leads to a simple automatic scheme for selecting videos of representative actions (exemplars) from the database and for efficiently indexing the whole database. We compute a performance metric on the hierarchical structure to evaluate goodness of the estimated hierarchy, and show that this metric has potential for predicting the clustering performance of various joining criteria used in building hierarchies. Our results show that perceptually meaningful hierarchies can be constructed based on action similarities with minimal user supervision, while providing favorable clustering performance and retrieval performance. Parvez Ahammad, Chuohao Yeo, Kannan Ramchandran, S. Shankar Sastry |
MMSP | 2 |
| 2007 | Robust distributed multi-view video compression for wireless camera networksabstractWe propose a novel method of exploiting inter-view correlation among cameras that have overlapping views in order to deliver error-resilient video in a distributed multi-camera system. The main focus in this work is on robustness which is imminently needed in a wireless setting. Our system has low encoding complexity, is robust while satisfying tight latency constraints, and requires no inter-sensor communication. In this work, we build on and generalize PRISM [Puri2002], an earlier proposed single-camera distributed video compression system. Specifically, decoder motion search, a key attribute of single-camera PRISM, is extended to the multi-view setting to include decoder disparity search based on two-view camera geometry. Our proposed system, dubbed PRISM-MC (PRISM multi-camera), achieved PSNR gains of up to 1.7 dB over a PRISM based simulcast solution in experiments over a wireless channel simulator. Chuohao Yeo, Kannan Ramchandran |
VCIP | 1 |
| 2006 | Compressed Domain Real-time Action RecognitionabstractWe present a compressed domain scheme that is able to recognize and localize actions in real-time. The recognition problem is posed as performing a video query on a test video sequence. Our method is based on computing motion similarity using compressed domain features which can be extracted with low complexity. We introduce a novel motion correlation measure that takes into account differences in motion magnitudes. Our method is appearance invariant, requires no prior segmentation, alignment or stabilization, and is able to localize actions in both space and time. We evaluated our method on a large action video database consisting of 6 actions performed by 25 people under 3 different scenarios. Our classification results compare favorably with existing methods at only a fraction of their computational cost Chuohao Yeo, Parvez Ahammad, Kannan Ramchandran, S. Shankar Sastry |
MMSP | 1 |
| 2005 | A Framework for Sub-Window Shot DetectionabstractBrowsing a digital video library can be very tedious especially with an ever expanding collection of multimedia material. We present a novel framework for extracting sub-window shots from MPEG encoded news video with the expectation that this will be another tool that can be used by retrieval systems. Sub-windows shots are also useful for tying in relevant material from multiple video sources. The system makes use of Macroblock parameters to extract visual features, which are then combined to identify possible sub-windows in individual frames. The identified sub-widows are then filtered by a non-linear Spatial-Temporal filter to produce sub-window shots. By working only on compressed domain information, this system avoids full frame decoding of MPEG sequences and hence achieves high speeds of up to 11 times real time. Chuohao Yeo, Yongwei Zhu, Qibin Sun, Shih-Fu Chang |
MMM | 1 |