EDBT 2026 Demo / reviewers in the wild / expert
Yo-Sung Ho
dblp:83/2498
· DBLP profile ↗
133ranked-venue papers
4as first author
33since 2021 · last 2025
0000-0002-7220-1034ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 119 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 3Computer networks · 2 · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Geometry-Guided Latent Diffusion Model for Static Point Cloud Color Attribute Denoising
Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Signal Process. Lett. | 5 |
| 2024 | Hallucinated-PQA: No reference point cloud quality assessment via injecting pseudo-reference features
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
Expert Syst. Appl. | 6 |
| 2024 | Visual Prompt Multibranch Fusion Network for RGB-Thermal Crowd CountingabstractAs population growth and urbanization continue, accurate crowd counting is increasingly important for public safety management and the Internet of Video Things (IOVT). However, RGB and thermal infrared (RGB-T) crowd counting still faces challenges in improving feature extraction capability for RGB streams and reducing multimodality differences. For this, we propose a visual prompt multibranch fusion network (VPMFNet) to tackle the above challenges. Specifically, to improve the ability of crowd analysis of the RGB stream in RGB-T crowd counting, through designing the prompt enhancement module, we take the prior features of head perception in the crowd as visual prompt cues to embed into the RGB stream. In terms of RGB and thermal image feature fusion, we fully reduce the modality differences from the perspectives of local fusion, global fusion, and multireceptive field fusion to accurately estimate the pedestrian number. Various experiments on two RGB-T crowd counting data sets demonstrate that our VPMFNet achieves a smaller estimation error in the number of pedestrians. Besides, our VPMFNet outperforms existing methods (i.e., multicolumn convolutional neural network, BL, SANet, UCNet, HDFNet, BBSNet, BL+IDAM, BL+CSCA, dual-branch enhanced feature fusion network, and GETANet) on the RGB-D data set. Our code will be released athttps://github.com/QSBAOYANGMU/VPMFNet. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Internet Things J. | 6 |
| 2024 | HDR light field imaging of dynamic scenes: A learning-based method and a benchmark dataset
Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Chongchong Jin, Haiyong Xu, Yo-Sung Ho |
Pattern Recognit. | 6 |
| 2024 | Plain-PCQA: No-Reference Point Cloud Quality Assessment by Analysis of Plain Visual and Geometrical ComponentsabstractIn reviewing the research progress in Point Cloud Quality Assessment (PCQA), two main pathways have emerged, i.e., 2D projections and 3D point descriptors. The former primarily focuses on visual information, while the latter concentrates on crucial geometrical information in three-dimensional space. However, the current studies lack a thorough investigation of the impact of visual components and seldom pay special attention to plane-point fusion strategies. To comprehensively represent features and effectively tackle various types of impairments, we propose an end-to-end learning paradigm, only considering plain visual and geometrical factors called Plain-PCQA, for quantitatively evaluating objective metrics of 3D dense point clouds associated with human perception. Firstly, we explore a sophisticated preprocessing technique. The entire point clouds are packaged into six projections by moving virtual cameras, which can conveniently increase the visual samples during the training stage. Given the high resolution of the projected image, we have opted for a relatively lightweight network, namely ResNet-18, as the backbone to enable higher resolution input data. Five cropped patches from the projected image are collectively fed into this network. In light of the presence of some invalid information in the projections, a mask weight is devised to calculate the significance of each patch based on its effective informational content. Secondly, dual neural networks, comprising of a No-Reference (NR) branch and a Degraded-Reference (DR) branch, are designed with fundamental visual components to provide quantitative quality metrics. Specifically, the NR branch utilizes the feature output of each block in the Vision Transformer (ViT) model to obtain long-range low-level and high-level visual NR quality. The DR branch employs KLT (Karhunen-Loève Transform) to acquire the principal component information of an image as the macro-structural image, and then feeds the difference between input images and macro-structural images into a network for DR quality extraction. Thirdly, a Plane-Point Interaction Transformer (P2IT) is presented by incorporating texture and semantic features in 2D projections and geometrical features in 3D spaces to characterize the complete features with a connected 2D-3D feature representation. With these elaborately designed deep features, the proposed model can achieve competitive performances relying solely on plain visual and geometrical components. The experimental results demonstrate the potential of the proposed approach in multiple representative databases, which surpasses existing state-of-the-art methods significantly. Xiongli Chai, Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Blind Quality Evaluator of Light Field Images by Group-Based Representations and Multiple Plane-Oriented Perceptual CharacteristicsabstractDue to the emergency of multi-view cameras and commercial Light Field (LF) cameras, the demand of high-performance LF quality evaluator is of great significance for guiding LF acquisition, processing and application and further promoting the visual perceived quality of LF visualizations. However, LF Images (LFIs), as high-dimensional data, suffer from various quality degradations not only in the spatial domain but also in the angular domain. Therefore, it is of great challenge to predict LF quality accurately. An effective LF evaluator should be able to represent these heterogeneous artifacts. In this paper, we provide a novel No-Reference LF Quality Assessment Evaluator (NR LF-QAE) to tackle this problem. Firstly, to measure angular consistency among viewports, we utilize group-based representations to character information similarity of aligned view stacks. Secondly, to better describe the texture information of LFIs, unifying spatial-angular texture statistic measurement is performed via Local Binary Patterns from Three Orthogonal Planes (LBP-TOP). Thirdly, we design 3D Log-Gabor filters to extract LF global structure information in Sub-Aperture Images (SAIs) as spatial feature characterizations and 2D Log-Gabor filters are adopted to characterize ray direction/depth information in Epipolar Plane Images (EPIs) as angular feature characterizations. By comprehensive LF information analyses in angular consistency and spatial-angular feature extraction with texture and structure descriptors, experimental results demonstrate the superiority of the proposed NR LF-QAE over the state-of-the-art comparative models in predicting the quality of LFIs on three available benchmark databases. The code will be released athttps://github.com/zerosola/NR-LF-QAE. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xuejin Wang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 6 |
| 2024 | Stitched Wide Field of View Light Field Image Quality Assessment: Benchmark Database and Objective MetricabstractDue to the limitation of commercial light field camera hardware devices, the imaging field of view is quite narrow. Numerous Light Field Image (LFI) stitching algorithms have been developed to expand the field of view. However, it is highly challenging to compare the performance of LFI stitching algorithms in a fair manner. Currently, due to the absence of a comprehensive benchmark database for subjective rating and a reliable objective quality metric, it is fairly difficult to comprehensively and accurately compare the actual performance of existing LFI stitching algorithms. In this study, we dedicate our efforts to the development of quality metrics for stitched Wide field of view LFI (WLFI) from subjective and objective assessment aspects. Specifically, we build the first stitched WLFI database, which provides the stitched WLFIs generated by eight representative LFI stitching algorithms, along with their corresponding subjective rating scores. Secondly, an effective blind stitched WLFI quality metric is developed to accurately assess the visual quality degradation. Extensive experiments conducted over our established WLFI database demonstrate that the proposed metric achieves higher consistency with subjective ratings than the competing quality metrics. Yueli Cui, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2024 | SCFANet: Semantics and Context Feature Aggregation Network for 360° Salient Object DetectionabstractHow to solve the problem of geometric distortion is the key for salient object detection (SOD) in 360° omnidirectional images. Most of the current methods integrate global and local visual cues through the fusion of the 360° equirectangular images and corresponding 360° cube-map images. The fusion in a single level cannot effectively utilize the information between the 360° equirectangular images and corresponding 360° cube-map images. In this work, we innovatively propose a semantics and context feature aggregation network (SCFANet) by fully exploring the interactivity between the two projection data. Specifically, we use Vision Transformer (ViT) to capture global visual cues for 360° equirectangular images and Convolutional Neural Network (CNN) to capture local visual cues for 360° cube-map images. To achieve effective fusion of the two projection data, we design a semantic guidance module (SGM), in which semantic features are used to guide the information fusion of the 360° equirectangular images and corresponding 360° cube-map images at each level. Then, a context fusion module (CFM) containing one local input and two context inputs is designed to integrate multi-scale features, where the local input extracts its own multi-scale information, and the context inputs complements their fine details and location information. Finally, we use feature aggregation and refinement module (FARM) to aggregate semantics and context feature and adopt a deep supervision strategy for training. Extensive experiments on two public 360° datasets show that our SCFANet exhibits competitive performance compared to other state-of-the-art (SOTA) 360° salient object detection models. Feng Shao 0001, Xiongli Chai, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2024 | Progressive Bidirectional Feature Extraction and Enhancement Network for Quality Evaluation of Night-Time ImagesabstractBlind image quality assessment (BIQA) has received increasing attention in the past decades. However, it still remains inadequately researched on BIQA for night-time images suffering from the diverse authentic degradations. Since the intrinsic content degradations of night-time images are highly related to the illumination, how to use the connection between content and illumination to enhance the feature representation ability is the key issue in designing BIQA methods for night-time images. In this article, we first construct an ultra-high-definition night-time image dataset (UHD-NID) with high image resolution and abundant parameter settings. UHD-NID contains 1600 images with a high resolution of 5616 × 3744, and each group of images contains ten exposure levels. Then, we conduct subjective assessment and analyze the subjective data to obtain a mean opinion score to each image in UHD-NID. To enhance the feature representation ability in content and illumination, we propose a Progressive Bidirectional Feature Extraction and Enhancement Network (PBFEE-Net). In addition, we use a decomposition network to decompose the input image into the reflectance and illumination, which can facilitate the ability of feature extraction to some extent. The experimental results show that our proposed method achieves superior performance in evaluating the quality of night-time images. Jiangli Shi, Feng Shao 0001, Chongzhen Tian, Hangwei Chen, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 6 |
| 2024 | Collaborative Learning and Style-Adaptive Pooling Network for Perceptual Evaluation of Arbitrary Style TransferabstractAlthough the research of arbitrary style transfer (AST) has achieved great progress in recent years, few studies pay special attention to the perceptual evaluation of AST images that are usually influenced by complicated factors, such as structure-preserving, style similarity, and overall vision (OV). Existing methods rely on elaborately designed hand-crafted features to obtain quality factors and apply a rough pooling strategy to evaluate the final quality. However, the importance weights between the factors and the final quality will lead to unsatisfactory performances by simple quality pooling. In this article, we propose a learnable network, named collaborative learning and style-adaptive pooling network (CLSAP-Net) to better address this issue. The CLSAP-Net contains three parts, i.e., content preservation estimation network (CPE-Net), style resemblance estimation network (SRE-Net), and OV target network (OVT-Net). Specifically, CPE-Net and SRE-Net use the self-attention mechanism and a joint regression strategy to generate reliable quality factors for fusion and weighting vectors for manipulating the importance weights. Then, grounded on the observation that style type can influence human judgment of the importance of different factors, our OVT-Net utilizes a novel style-adaptive pooling strategy guiding the importance weights of factors to collaboratively learn the final quality based on the trained CPE-Net and SRE-Net parameters. In our model, the quality pooling process can be conducted in a self-adaptive manner because the weights are generated after understanding the style type. The effectiveness and robustness of the proposed CLSAP-Net are well validated by extensive experiments on the existing AST image quality assessment (IQA) databases. Our code will be released at https://github.com/Hangwei-Chen/CLSAP-Net. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Modality-Induced Transfer-Fusion Network for RGB-D and RGB-T Salient Object DetectionabstractThe ability of capturing the complementary information of multi-modality data is critical to the development of multi-modality salient object detection (SOD). Most of existing studies attempt to integrate multi-modality information through various fusion strategies. However, most of these methods ignore the inherent differences in multi-modality data, resulting in poor performance when dealing with some challenging scenarios. In this paper, we propose a novel Modality-Induced Transfer-Fusion Network (MITF-Net) for RGB-D and RGB-T SOD by fully exploring the complementarity in multi-modality data. Specifically, we first deploy a modality transfer fusion (MTF) module to bridge the semantic gap between single and multi-modality data, and then mine the cross-modality complementarity based on point-to-point structural similarity information. Then, we design a cycle-separated attention (CSA) module to optimize the cross-layer information recurrently, and measure the effectiveness of cross-layer features through point-wise convolution-based multi-scale channel attention. Furthermore, we refine the boundaries in the decoding stage to obtain high-quality saliency maps with sharp boundaries. Extensive experiments on 13 RGB-D and RGB-T SOD datasets show that the proposed MITF-Net achieves a competitive and excellent performance. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Quality Evaluation of Arbitrary Style Transfer: Subjective Study and Objective MetricabstractArbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the task of arbitrary style transfer (AST) for improving the stylization quality. However, there are very few explorations about the quality evaluation of AST images, even it can potentially guide the design of different algorithms. In this paper, we first construct a new AST images quality assessment database (AST-IQAD), which consists 150 content-style image pairs and the corresponding 1200 stylized images produced by eight typical AST algorithms. Then, a subjective study is conducted on our AST-IQAD database, which obtains the subjective rating scores of all stylized images on the three subjective evaluations, i.e., content preservation (CP), style resemblance (SR), and overall vision (OV). To quantitatively measure the quality of AST image, we propose a new sparse representation-based method, which computes the quality according to the sparse feature similarity. Experimental results on our AST-IQAD have demonstrated the superiority of the proposed method. The dataset and source code will be released athttps://github.com/Hangwei-Chen/AST-IQAD-SRQE Hangwei Chen, Feng Shao 0001, Xiongli Chai, Yuese Gu, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Viewport-Sphere-Branch Network for Blind Quality Assessment of Stitched 360° Omnidirectional ImagesabstractCompared with conventional images/videos, omnidirectional data records rich information with higher resolution and wider Field-of-View. Moreover, the stitching distortions introduced in the panoramic content generation process make the quality assessment task more challenging. Targeting at designing an accurate and fast stitched 360° omnidirectional image quality evaluator, we propose a Viewport-Sphere-Branch Network (VSBNet) via dual-branch quality estimation. Specifically, for the viewport quality estimation, we extract distorted viewports around the stitching seams and conduct distortion rectification through a progressively complementary network to obtain pseudo-reference viewports. The qualitative and quantitative experiments validate that pseudo-reference viewports are reliable. Then, the differences between distorted and pseudo-reference viewports are quantified through transformer architecture to obtain quality scores of viewports. The introduction of pseudo-reference viewports can effectively improve the performance of the viewport quality prediction branch. To establish general scenario awareness and accurately evaluate the immersive experience, we extract feature representation through deformable convolutions to eliminate 2D-to-Sphere intrinsic sampling distortions and use multilayer perceptron to predict score of the whole sphere. The final prediction score is obtained by aggregating the quality scores from viewport and sphere branches. We evaluate the proposed VSBNet on two benchmark databases and results demonstrate that the combination of two branches can obtain more accurate results. Overall, our method is superior to existing full reference and no reference models designed for conventional images and 360° omnidirectional images. Chongzhen Tian, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Cross-Modality Double Bidirectional Interaction and Fusion Network for RGB-T Salient Object DetectionabstractRGB-T salient object detection (SOD) aims to detect and segment saliency regions on RGB images and the corresponding thermal maps. The ability of alleviating the modality difference between RGB and thermal modality plays a vital role in the development of RGB-T SOD. However, most of the existing methods try to integrate multi-modal information through various fusion strategies, or reduce the modality difference via unidirectional or undifferentiated bidirectional interaction, but failing in some challenging scenes. To deal with the above question, a novel Cross-Modality Double Bidirectional Interaction and Fusion Network (CMDBIF-Net) for RGB-T SOD is proposed. Specifically, we construct an interactive branch to indirectly bridge the RGB and thermal modalities. In addition, we propose a double bidirectional interaction (DBI) module composed of a forward interaction block (FIB) and a backward interaction block (BIB) to reduce the cross-modality differences. Moreover, a multi-scale feature enhancement and fusion (MSFEF) module is introduced to integrate the multi-modal features with considering the internal gap of different modality. Finally, we use a cascaded decoder and a cross-level feature enhancement (CLFE) module to generate high-quality saliency map. Extensive experiments are conducted on three publicly available RGB-T SOD datasets shows that the proposed CMDBIF-Net achieves outstanding performance against the state-of-the-art (SOTA) RGB-T SOD methods. Zhengxuan Xie, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Perceptual Quality Assessment of Cartoon ImagesabstractIn the animation industry, automatically predicting the quality of cartoon images based on the inputs of general distortions and color change is an urgent task, while the existing no-reference (NR) methods usually measure the perceptual quality of the natural images. In this paper, based on the observation that structure and color are the main factors affecting cartoon images quality, we proposed a new NR quality prediction metric for cartoon images, which fully takes gradient and color information into account. The experimental results on our newly constructed NBU-CIQAD dataset with color change and other existing cartoon image dataset demonstrate that the proposed method significantly outperforms existing no-references methods for the task of cartoon image quality assessment. The database and code will be released athttps://github.com/1010075746/NBU-CIQAD. Hangwei Chen, Xiongli Chai, Feng Shao 0001, Xuejin Wang, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 7 |
| 2023 | Composition-Guided Neural Network for Image Cropping Aesthetic AssessmentabstractHow to explore the interaction between image aesthetic rules and crops is the key to finding views with good composition. Besides, it is subjective to evaluate candidate crops, which mainly depends on aesthetic knowledge, but it is not an easy task for people without extensive photography experience. However, existing methods mostly find good views by extracting general aesthetic features of crops without fully exploring the aesthetic rules. Motivated by this, we innovatively propose a composition-guided image cropping aesthetic assessment network (CGICAANet) for efficiently finding good crops and optimizing the cropping operation. Specifically, we adopt a direct and comprehensive composition pattern module, which adaptively mines suitable compositions for the images and emphasizes the dominant position of visual elements to contribute to optimizing the best crops in an interpretable way. Moreover, we designed a multi-task loss function to train the model. Particularly, to explore the commonality between predicted crops and labels, the complete intersection-over-union loss is adopted thoroughly considering the overlap area, central point distance and the consistency of aspect ratios for crops concurrently. Therefore, the predicted best crop can preserve the visual elements and have better composition. Experimental results with lightweight MobileNetV2 and ShuffleNetV2 as backbone networks demonstrate that our method can obtain comparable or better performance in terms of efficiency and accuracy. Shijia Ni, Feng Shao 0001, Xiongli Chai, Hangwei Chen, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2023 | No-Reference Light Field Image Quality Assessment Using Four-Dimensional Sparse TransformabstractLight field imaging can simultaneously capture the intensity and direction information of light rays in the real world. Light field image (LFI) with four-dimensional (4D) data suffers from quality degradation in the process of compression, reconstruction and processing. How to evaluate the visual quality of LFI is thought-provoking. This paper proposes a no-reference LFI quality assessment metric based on high-dimensional sparse transform. Firstly, LFI's sub-aperture gradient image array (SAGIA), which is still a 4D signal, is generated by high-pass filtering between adjacent SAIs. Then, SAGIA is transformed with 4D discrete cosine transform (4D-DCT). 4D-DCT coefficients of SAGIA can characterize the angular and spatial information of LFI. And the logarithmic amplitudes of the coefficients at the same position of SAGIA?s transformed 4D blocks are averaged as the coefficient energy. Subsequently, the 4D-DCT coefficients of SAGIA are divided into the spatial-angular frequency bands and spatial-angular orientation bands, and the corresponding energy features are extracted by converging the coefficient energy of the same band. In addition, the coefficients' amplitudes at the same position of blocks are fitted by the Weibull distribution. Then, the fitted parameters of each position are concatenated, and cropped with principal component analysis to obtain the compact features. Finally, the extracted features are pooled to predict the visual quality of the distorted LFIs. The experimental results demonstrate that the proposed method is more consistent with the subjective evaluation on three LFI databases, compared with the state-of-the-art image quality assessment methods and LFI quality assessment methods. Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2023 | Deep Light Field Spatial Super-Resolution Using Heterogeneous ImagingabstractLight field (LF) imaging expands traditional imaging techniques by simultaneously capturing the intensity and direction information of light rays, and promotes many visual applications. However, owing to the inherent trade-off between the spatial and angular dimensions, LF images acquired by LF cameras usually suffer from low spatial resolution. Many current approaches increase the spatial resolution by exploring the four-dimensional (4D) structure of the LF images, but they have difficulties in recovering fine textures at a large upscaling factor. To address this challenge, this paper proposes a new deep learning-based LF spatial super-resolution method using heterogeneous imaging (LFSSR-HI). The designed heterogeneous imaging system uses an extra high-resolution (HR) traditional camera to capture the abundant spatial information in addition to the LF camera imaging, where the auxiliary information from the HR camera is utilized to super-resolve the LF image. Specifically, an LF feature alignment module is constructed to learn the correspondence between the 4D LF image and the 2D HR image to realize information alignment. Subsequently, a multi-level spatial-angular feature enhancement module is designed to gradually embed the aligned HR information into the rough LF features. Finally, the enhanced LF features are reconstructed into a super-resolved LF image using a simple feature decoder. To improve the flexibility of the proposed method, a pyramid reconstruction strategy is leveraged to generate multi-scale super-resolution results in one forward inference. The experimental results show that the proposed LFSSR-HI method achieves significant advantages over the state-of-the-art methods in both qualitative and quantitative comparisons. Furthermore, the proposed method preserves more accurate angular consistency. Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Yo-Sung Ho |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Monocular and Binocular Interactions Oriented Deformable Convolutional Networks for Blind Quality Assessment of Stereoscopic Omnidirectional ImagesabstractStereoscopic omnidirectional content, as a novel visual media, has drawn wide attention in recent years due to its ability in providing strong immersive experience. Since Stereoscopic Omnidirectional Images (SOIs) involve the properties from panoramic and stereoscopic visual perception, it is very challenging to establish an efficient and effective visual quality evaluation model for SOIs. To better measure the user’s experience in virtual reality, we put forward a novel deep learning framework to assess the quality of SOIs in this paper. Firstly, the deformable convolutions instead of standard convolutions are adopted to ensure the invariant receptive fields of convolutional kernels on Equi-Rectangular Projection (ERP). Secondly, according to the stereoscopic property, we use binocular-difference information and a coarse-to-fine mechanism to construct the binocular feature extraction network. Thirdly, a three-channel network involving left-view, right-view and binocular-difference channels is presented to simulate the process of monocular and binocular interactions, in which independent quality labels are provided for each channel to reflect the individual effect of monocular and binocular visions on the whole visual quality. Finally, experimental results on two available benchmark databases demonstrate the superiority of the proposed metric over the state-of-the-art blind quality assessment models in predicting the quality of SOIs. Moreover, our model is efficient in computational cost as the feature extraction is directly applied on ERP images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | CGMDRNet: Cross-Guided Modality Difference Reduction Network for RGB-T Salient Object DetectionabstractHow to explore the interaction between the RGB and thermal modalities is the key success of the RGB-T saliency object detection (SOD). Most of the existing methods integrate multi-modality information by designing various fusion strategies. However, the modality gap between the RGB and thermal features will lead to unsatisfactory performances by simple feature concatenation. To solve this problem, we innovatively propose a cross-guided modality difference reduction network (CGMDRNet) to achieve intrinsic consistency feature fusion via reducing the modality differences. Specifically, we design a modality difference reduction (MDR) module, which is embedded in each layer of the backbone network. The module uses a cross-guided strategy to reduce the modality difference between the RGB and thermal features. Then, a cross-attention fusion (CAF) module is designed to fuse cross-modality features with small modality differences. In addition, we use a transformer-based feature enhancement (TFE) module to enhance the high-level feature representation that contributes more to performance. Finally, the high-level features guide the fusion of low-level features to obtain a saliency map with clear boundaries. Extensive experiments on three public RGB-T datasets show that the proposed CGMDRNet achieves competitive performance compared with state-of-the-art (SOTA) RGB-T SOD models. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Texture-Aware Spherical Rotation for High Efficiency Omnidirectional Intra Video CodingabstractTo adapt to the existing video coding standards, omnidirectional videos are usually projected from Three-Dimensional (3D) sphere to Two-Dimensional (2D) plane. However, this projection will cause geometrical stretching distortion and boundary discontinuity, which may degrade coding efficiency. In this paper, we present a Spherical Rotation based Omnidirectional Video Coding (SROVC) method, which exploits the textural properties of omnidirectional videos with spherical rotation. Firstly, SROVC framework is presented and Full-traversal Spherical Rotation (FSR) is developed to derive the optimal rotation angle with frame-level Rate Distortion Optimization (RDO). Secondly, to achieve comparable coding gains and lower computational complexity when compared with FSR, a Texture-aware Spherical Rotation (TSR) method is proposed to predict the rotation angle. Finally, to further reduce complexity and maintain coding efficiency, a Group-oriented TSR (G-TSR) approach is presented, in which the group length is statistically determined. Extensive experiments demonstrate that the proposed TSR and G-TSR schemes can achieve bit rate reductions up to 4.38%, 0.94% and 0.89% on average for CubeMap Projection (CMP) based high efficiency omnidirectional video coding. Additionally, the TSR scheme achieves bit rate saving from 0.91% to 1.19% on average under three more CMP-based projection formats, and 1.89% for joint rotation of X, Y, and Z axes. Jinyong Pi, Yun Zhang 0002, Linwei Zhu, Jinzhi Lin, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality EvaluatorabstractWith the rapid development of virtual reality (VR), 360° omnidirectional images and videos have drawn wide attention. However, the quality assessment of 360° omnidirectional images is a challenging task, especially when the panoramic image contains multiple stitching distortions. We propose a viewport-based stitched 360° omnidirectional image quality evaluator (VSOIQE), by first extracting the features of salient and stitching viewports, and then inferring the overall perceptual quality via multiple linear regression (MLR). Comprehensive image attributes including edge, color, shape and information entropy are considered in the framework. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art quality models designed for 2D images and the quality models developed for 360° omnidirectional images. Chongzhen Tian, Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2022 | Deep Light Field Super-Resolution Using Frequency Domain Analysis and Semantic PriorabstractLight field (LF) camera can simultaneously capture the intensity and direction information of light rays, which has been widely concerned. However, limited by the size of the imaging sensor, the captured LF image (LFI) has a trade-off between spatial and angular resolutions. To this end, this paper proposes a new LF super-resolution method using frequency domain analysis and semantic prior, which designs a two-stage learning framework to enhance the spatial and angular resolutions of LFI. Specifically, the proposed method first decomposes the spatial and angular information to explore the 4D structure of LFI by using frequency domain transformation, and formulates the LF super-resolution as a frequency restoration process. Then, the decomposed frequency components are recovered in a progressive restoration manner, with new cascaded 2D and 3D convolutional neural networks. To further improve the quality of the reconstructed LFI, especially at the object boundary, the semantic prior is incorporated into the designed network to enhance its representation ability. Finally, the super-resolved LFI is reconstructed by inverse frequency domain transformation. Experimental results show that the proposed method can effectively generate high-resolution LFI, and outperforms other state-of-the-art methods in terms of both subjective visual perception and objective quality evaluation. Moreover, the proposed method can enhance the performance of LF applications such as depth estimation. Yeyao Chen, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2022 | List-Wise Rank Learning for Stereoscopic Image Retargeting Quality AssessmentabstractStereoscopic imageretargeting (SIR) techniques attempt to display stereoscopic images on stereoscopic devices of various resolutions and aspect ratios to provide the users with better viewing experience. However, new quality perceptual problems emerge in the retargeted stereoscopic images generated by current SIR operators are quite different from those in the retargeted 2D images. In this paper, we dedicate to exploring the perceptual quality-related factors (e.g., shape preservation, object preservation and visual comfort.) of retargeted stereoscopic images, and propose a novel quality evaluation metric for SIR to achieve a more consistent evaluation with 3D perception and image degradation mechanism in the SIR process. Moreover, image quality features and 3D perceptual features are integrated into one representation for an overall perceptual quality prediction using a list-wise ranking approach, which gives priority to the ranking among the SIR results generated from the same stereoscopic source. Experimental results demonstrate that the proposed method outperforms most quality models developed for retargeted 2D/stereoscopic images. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiongli Chai, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 6 |
| 2022 | Combining Retargeting Quality and Depth Perception Measures for Quality Evaluation of Retargeted StereopairsabstractStereoscopic Image Retargeting (SIR) aims to adapt stereoscopic images and videos to 3D display devices with various aspect ratios by emphasizing the important content while retaining surrounding context with minimal visual distortion. To address the issue of SIR evaluation, this paper presents a new objective quality assessment method for retargeted stereopairs by combining image quality and depth perception measures. Specifically, the image quality measure is conducted between the source and retargeted intermediate views generated by the view synthesis method to characterize the geometric distortion and content loss of the retargeted stereopair, while several depth-aware features are extracted to measure the visual comfort/discomfort and depth sensation when human views a 3D scene. Then, the extracted features are integrated into an overall perceptual quality prediction. Experiment results on NBU SIRQA and SIRD databases verify the superiority of our method. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Zhenqi Fu, Xiangchao Meng, Ke Gu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 7 |
| 2021 | Roundness-Preserving Warping for Aesthetic Enhancement-Based Stereoscopic Image EditingabstractImage editing is an effective solution to adapt contents for different applications. In this paper, we present a roundness-preserving warping model for stereoscopic image editing, in which energy constraints from image quality energy, aesthetics energy and depth adaptation energy are involved in the framework to solve the optimization. Specifically, to preserve object roundness during warping, the relationship between object's shape and disparity is established and is applied for depth adaptation. Different from the existing stereoscopic image editing methods, the main innovations of our method are to achieve a tradeoff in balancing information loss and reducing semantic distortion while providing a novel death adaptation model for recomposition and retargeting applications. Experimental results demonstrate the effectiveness of our method in enhancing the aesthetics of stereoscopic images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Viewport Perception Based Blind Stereoscopic Omnidirectional Image Quality AssessmentabstractCompared with traditional 2D images, stereoscopic omnidirectional images (SOIs) usually have more complex perceptual factors due to the particularities of imaging and display, making the objective quality assessment of SOIs challenging. In this paper, we construct a large and diverse subjective SOIs database named as NBU-SOID for further research demand. And then, we propose a viewport perception based blind SOIs quality assessment (VP-BSOIQA) method by considering the impacts of viewport, user behavior and stereoscopic perception on human visual system, which is mainly composed of binocular perception model (BPM) and omnidirectional perception model (OPM). In the BPM, a binocular combination perception map is generated by the dimension reduction of stereopair and the weighting of binocular energy to reflect the binocular masking effect. In the OPM, several viewports are first created to ensure the consistency of evaluation objects. Then, the intra-viewport and inter-viewport weighting factors are designed with the common influences of visual attention and peripheral vision sensitivity to aggregate the novel multi-orientation structural features extracted from all potential viewports. Experimental results on the NBU-SOID and SOLID databases demonstrate that BPM and OPM can be robustly combined with the existing 2D image quality assessment (IQA) methods, thus averagely achieving 10.2% and 12.2% performance gain in terms of SRCC, respectively. In addition, the proposed VP-BSOIQA method outperforms the state-of-the-art blind IQA methods in predicting the quality of SOIs. Yubin Qi, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Pseudo Video and Refocused Images-Based Blind Light Field Image Quality AssessmentabstractThe commercial light field camera is able to capture four-dimensional Light Field Image (LFI), which can be visualized to LFI contents on 2D displays by means of the Pseudo Video (PV) or the Refocused Images (RIs) generated with the refocusing function of LFI. However, the quality degradation of LFI will affect user’s visual experience of LFI contents. Hence, it is crucial to develop an effective LFI quality assessment method to monitor the LFI quality. Most existing subjective databases of LFI use PV and RIs visualization techniques to assess the quality of LFI. Therefore, as the way of presenting LFI on 2D display, PV and RIs are closely related to the subjective perception of LFI by human eyes. Based on these two visualization techniques, this article proposes a novel PV and RIs based blind LFI quality assessment method, in which the feature extraction is divided into two parts. In the first part, the PV’s structure, motion and disparity information are extracted with multi-scale and multi-directional Shearlet transform. In the other part, the spatial structure, depth and semantic information of the RIs are obtained. Finally, support vector regression is used to nonlinear map the perceptual features to quality score of LFI. The experimental results on four LFI databases show that the proposed method has better correlation with human visual perception, compared with the classical 2D image quality assessment methods as well as the state-of-the-art LFI quality assessment methods. Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Online Learning-Based Multi-Stage Complexity Control for Live Video CodingabstractHigh Efficiency Video Coding (HEVC) can significantly improve the compression efficiency in comparison with the preceding H.264/Advanced Video Coding (AVC) but at the cost of extremely high computational complexity. Hence, it is challenging to realize live video applications on low-delay and power-constrained devices, such as the smart mobile devices. In this article, we propose an online learning-based multi-stage complexity control method for live video coding. The proposed method consists of three stages: multi-accuracy Coding Unit (CU) decision, multi-stage complexity allocation, and Coding Tree Unit (CTU) level complexity control. Consequently, the encoding complexity can be accurately controlled to correspond with the computing capability of the video-capable device by replacing the traditional brute-force search with the proposed algorithm, which properly determines the optimal CU size. Specifically, the multi-accuracy CU decision model is obtained by an online learning approach to accommodate the different characteristics of input videos. In addition, multi-stage complexity allocation is implemented to reasonably allocate the complexity budgets to each coding level. In order to achieve a good trade-off between complexity control and rate distortion (RD) performance, the CTU-level complexity control is proposed to select the optimal accuracy of the CU decision model. The experimental results show that the proposed algorithm can accurately control the coding complexity from 100% to 40%. Furthermore, the proposed algorithm outperforms the state-of-the-art algorithms in terms of both accuracy of complexity control and RD performance. Chao Huang 0008, Zongju Peng, Yong Xu 0001, Qiuping Jiang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 8 |
| 2021 | Highly Efficient Multiview Depth Coding Based on Histogram Projection and Allowable Depth DistortionabstractMismatches between the precisions of representing the disparity, depth value and rendering position in 3D video systems cause redundancies in depth map representations. In this paper, we propose a highly efficient multiview depth coding scheme based on Depth Histogram Projection (DHP) and Allowable Depth Distortion (ADD) in view synthesis. Firstly, DHP exploits the sparse representation of depth maps generated from stereo matching to reduce the residual error from INTER and INTRA predictions in depth coding. We provide a mathematical foundation for DHP-based lossless depth coding by theoretically analyzing its rate-distortion cost. Then, due to the mismatch between depth value and rendering position, there is a many-to-one mapping relationship between them in view synthesis, which induces the ADD model. Based on this ADD model and DHP, depth coding with lossless view synthesis quality is proposed to further improve the compression performance of depth coding while maintaining the same synthesized video quality. Experimental results reveal that the proposed DHP based depth coding can achieve an average bit rate saving of 20.66% to 19.52% for lossless coding on Multiview High Efficiency Video Coding (MV-HEVC) with different groups of pictures. In addition, our depth coding based on DHP and ADD achieves an average depth bit rate reduction of 46.69%, 34.12% and 28.68% for lossless view synthesis quality when the rendering precision varies from integer, half to quarter pixels, respectively. We obtain similar gains for lossless depth coding on the 3D-HEVC, HEVC Intra coding and JPEG2000 platforms. Yun Zhang 0002, Linwei Zhu, Raouf Hamzaoui, Sam Kwong, Yo-Sung Ho |
IEEE Trans. Image Process. | 5 |
| 2021 | Subjective and Objective Quality Assessment for Stereoscopic Image RetargetingabstractBinocular stereoscopic image retargeting (SIR) aims to adjust 3D images into target aspect ratios. In recent years, various SIR methods have been proposed, but there are few researches on visual quality assessment. As a consequence, we construct a benchmark stereoscopic image retargeting quality assessment database (NBU-SIRQA), which contains 720 stereoscopic retargeted images generated by eight representative SIR operators. Subjective test is conducted to obtain the mean opinion score (MOS) for each stereoscopic retargeted image. Additionally, we propose an objective SIRQA metric based on grid deformation and information loss (GDIL). The main idea of GDIL is to decompose the SIR operator into two transformations: monocular image retargeting transformation and viewpoint transformation. In each transformation, grid deformation and information loss are extracted simultaneously to represent image quality and 3D perception quality. Experimental results validated on our established NBU-SIRQA database show the superiority of our metric in measuring the quality of stereoscopic retargeted images over the existing approaches. Zhenqi Fu, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2021 | Measuring Coarse-to-Fine Texture and Geometric Distortions for Quality Assessment of DIBR-Synthesized ImagesabstractA synthesized view can be generated via Depth-Image-Based Rendering (DIBR) technique using one (or more) color images and the associated depth maps. However, several artifacts may occur in the synthesized views due to the imperfect color images, depth maps or texture inpainting techniques, which cannot be effectively estimated by the conventional quality metrics designed for natural images. In this paper, a new quality metric is proposed to evaluate DIBR-synthesized images by measuring texture and geometric distortions. The artifacts are first analyzed on different phases of the synthesis process, and the associated features are extracted to estimate the degree of texture and geometric distortions from both coarse and fine scales. Finally, individual quality scores are aggregated into an overall quality via regression. Experimental results on three publicly available DIBR datasets demonstrate the superiority of the proposed method over the state-of-the-art quality models. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2021 | Transformation-Aware Similarity Measurement for Image Retargeting Quality Assessment via Bidirectional RewarpingabstractImage retargeting is an effective way to adapt images for target displays with different aspect ratios and sizes. Meanwhile, effective image retargeting quality assessment (IRQA) is important for optimizing the image retargeting operations. In this paper, we propose a transform-aware similarity (TRASIM) measurement metric for IRQA, including bidirectional geometric distortion measurement, bidirectional information loss measurement, and global salient structure distortion measurement. The main innovation of the TRASIM is to build a universal framework to establish the similarity transformation via bidirectional rewarping to simulate different types of retargeting operators. Based on the similarity transformation, geometric distortion and content loss are measured to determine the retargeting quality. Experimental results on two widely used databases (CUHK and RetargetMe) indicate that the proposed TRASIM has higher consistency with subjective ranks, compared with the state-of-the-art IRQA metrics. Feng Shao 0001, Zhenqi Fu, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | A large-scale remote sensing database for subjective and objective quality assessment of pansharpened images
Yiming Xiong, Feng Shao 0001, Xiangchao Meng, Qiuping Jiang, Weiwei Sun 0005, Randi Fu, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 7 |
| 2020 | Sparse Representation-Based Video Quality Assessment for Synthesized 3D VideosabstractThe temporal flicker distortion is one of the most annoying noises in synthesized virtual view videos when they are rendered by compressed multi-view video plus depth in Three Dimensional (3D) video system. To assess the synthesized view video quality and further optimize the compression techniques in 3D video system, objective video quality assessment which can accurately measure the flicker distortion is highly needed. In this paper, we propose a full reference sparse representation based video quality assessment method towards synthesized 3D videos. Firstly, a synthesized video, treated as a 3D volume data with spatial (X-Y) and temporal (T) domains, is reformed and decomposed as a number of spatially neighboring temporal layers, i.e., X-T or Y-T planes. Gradient features in temporal layers of the synthesized video and strong edges of depth maps are used as key features in detecting the location of flicker distortions. Secondly, dictionary learning and sparse representation for the temporal layers are then derived and applied to effectively represent the temporal flicker distortion. Thirdly, a rank pooling method is used to pool all the temporal layer scores and obtain the score for the flicker distortion. Finally, the temporal flicker distortion measurement is combined with the conventional spatial distortion measurement to assess the quality of synthesized 3D videos. Experimental results on synthesized video quality database demonstrate our proposed method is significantly superior to other state-of-the-art methods, especially on the view synthesis distortions induced from depth videos. Yun Zhang 0002, Huan Zhang 0008, Mei Yu 0001, Sam Kwong, Yo-Sung Ho |
IEEE Trans. Image Process. | 5 |
| 2020 | MSTGAR: Multioperator-Based Stereoscopic Thumbnail Generation With Arbitrary ResolutionabstractAt present, thumbnail generation for 2D images has been extensively studied, but the research in thumbnail generation for stereoscopic images is still relatively lacking. This paper presents a novel thumbnail generation technology for stereoscopic images based on multioperator with the following innovations: 1) The warping technique is used to retarget a stereopair into six-scale resolutions with different contexts, and the disparity is uniformly adjusted to a certain value based on just noticeable depth difference (JNDD) model, which overcomes the issues that 3D perception in stereoscopic thumbnail is uncontrollable and the sense of depth disappears in low-resolution stereoscopic images. 2) The six-scale images are cropped via cropping network, and are optimized to a target resolution based on the designed image visual representation energy. As a result, our method has better visual effect than state-of-the-art methods in generating thumbnail for stereoscopic display. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 4 |
| 2019 | Event-Based High Dynamic Range Image and Very High Frame Rate Video Generation Using Conditional Generative Adversarial NetworksabstractEvent cameras have a lot of advantages over traditional cameras, such as low latency, high temporal resolution, and high dynamic range. However, since the outputs of event cameras are the sequences of asynchronous events over time rather than actual intensity images, existing algorithms could not be directly applied. Therefore, it is demanding to generate intensity images from events for other tasks. In this paper, we unlock the potential of event camera-based conditional generative adversarial networks to create images/videos from an adjustable portion of the event data stream. The stacks of space-time coordinates of events are used as inputs and the network is trained to reproduce images based on the spatio-temporal intensity changes. The usefulness of event cameras to generate high dynamic range (HDR) images even in extreme illumination conditions and also non blurred images under rapid motion is also shown. In addition, the possibility of generating very high frame rate videos is demonstrated, theoretically up to 1 million frames per second(FPS) since the temporal resolution of event cameras is about 1 microsecond. Proposed methods are evaluated by comparing the results with the intensity images captured on the same pixel grid-line of events using online available real datasets and synthetic datasets produced by the event camera simulator. Lin Wang 0025, Sayed Mohammad Mostafavi Isfahani, Yo-Sung Ho, Kuk-Jin Yoon |
CVPR | 3 |
| 2019 | Simultaneous object size and depth adjustment for stereoscopic 3D images
Feng Shao 0001, Yanjia Fei, Randi Fu, Gangyi Jiang, Yo-Sung Ho |
Inf. Sci. | 5 |
| 2019 | A Risk-Aware Pairwise Rank Learning Approach for Visual Discomfort Prediction of Stereoscopic 3DabstractFor visual discomfort prediction (VDP) of stereoscopic 3D images, a common two-stage framework is to first extract features that are predictive of the experienced visual discomfort level when viewing stereoscopic images and then use typical regression tools to learn the mapping from the extracted features to visual discomfort scores. Most existing approaches for stereoscopic 3D VDP focus on the former stage, i.e., feature extraction, while limited efforts have dedicated to exploiting more powerful and robust learning algorithms in this field. In this letter, inspired by the pairwise comparison-based subjective evaluation methodology, we propose a novel Risk-Aware Pairwise Rank Learning (RAPRL) approach to further improve the prediction accuracy. Unlike the traditional VDP approaches using different regression tools for feature-score mapping, our proposed RARL method addresses this problem based on a completely different pairwise rank learning framework with a risk-aware constraint. Experiments have verified the effectiveness and robustness of our proposed VDP model using RAPRL as the learning algorithm. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Yo-Sung Ho |
IEEE Signal Process. Lett. | 5 |
| 2019 | Unified No-Reference Quality Assessment of Singly and Multiply Distorted Stereoscopic ImagesabstractA challenging problem in the no-reference quality assessment of multiply distorted stereoscopic images (MDSIs) is to simulate the monocular and binocular visual properties under a mixed type of distortions. Due to the joint effects of multiple distortions in MDSIs, the underlying monocular and binocular visual mechanisms have different manifestations with those of singly distorted stereoscopic images (SDSIs). This paper presents a unified no-reference quality evaluator for SDSIs and MDSIs by learning monocular and binocular local visual primitives (MB-LVPs). The main idea is to learn MB-LVPs to characterize the local receptive field properties of the visual cortex in response to SDSIs and MDSIs. Furthermore, we also consider that the learning of primitives should be performed in a task-driven manner. For this, two penalty terms including reconstruction error and quality inconsistency are jointly minimized within a supervised dictionary learning framework, generating a set of quality-oriented MB-LVPs for each single and multiple distortion modality. Given an input stereoscopic image, feature encoding is performed using the learned MB-LVPs as codebooks, resulting in the corresponding monocular and binocular responses. Finally, responses across all the modalities are fused with probabilistic weights which are determined by the modality-specific sparse reconstruction errors, yielding the final monocular and binocular features for quality regression. The superiority of our method has been verified on several SDSI and MDSI databases. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 6 |
| 2019 | WLDISR: Weighted Local Sparse Representation-Based Depth Image Super-Resolution for 3D Video SystemabstractIn this paper, we propose a Weighted Local sparse representation based Depth Image Super-Resolution (WLDISR) schemes aiming at improving the Virtual View Image (VVI) quality of 3D video system. Different from color images, depth images are mainly used to provide geometrical information in synthesizing VVI. Due to the view synthesis characteristics difference between textural structures and smooth regions of depth images, we divide the depth images into edge and smooth patches and learn two local dictionaries, respectively. Meanwhile, the weight term is derived and incorporated explicitly in the cost function to denote different importance of edge structures and smooth regions to the VVI quality. Then, local sparse representation and weighted sparse representation are jointly used in both dictionary learning and reconstruction phases in depth image super-resolution. Based on different optimizations on learning and reconstruction modules, three WLDISR schemes, WLDISR-D, WLDISR-R, and WLDISR-ALL, are proposed. Experimental results on 3D sequences demonstrate that the proposed WLDISR-D, WLDISR-R, and WLDISR-ALL schemes can achieve more than 1.9-, 2.03-, and 2.16-dB gains on average, respectively, in terms of the VVIs' quality, as compared with the state-of-the-art schemes. In addition, the visual quality of VVIs is also improved. Huan Zhang 0008, Yun Zhang 0002, Hanli Wang, Yo-Sung Ho, Shengzhong Feng |
IEEE Trans. Image Process. | 4 |
| 2018 | High-Precision 3D Coarse Registration Using RANSAC and Randomly-Picked Rejections
Jong-Hee Back, Yo-Sung Ho |
MMM (1) | 3 |
| 2018 | Place Recognition for SLAM Using Texture-independent Features and DCGANabstractIn order to make an accurate 3D map in visual SLAM, it is important to check whether the current place is already visited or not because we should close the loop at the correct position. If it is not visited before, it is very hard to register the 3D map. In this paper, we propose a method for revisited place detection based on unsupervised learning, and a method for high-quality 3D map rendering based on root finding optimization, both of which are faster than previous works. In order to detect revisited places, we generate local patches based on the superpixel and learn them using unsupervised learning based neural networks. Experimental results show that the superpixel-based feature extraction method detects image features that cover both textured and texture-less regions, unlike previous feature extraction methods. Thus, we can utilize the proposed method not only for outside scenes, but also for indoor scenes. Moreover, when the place recognition process is performed using the unsupervised learning and extracted image features, we can improve the average precision over previous methods. Yo-Sung Ho |
TENCON | 2 |
| 2018 | Edge preserving suppression for depth estimation via comparative variationabstractMost applications in computer vision manage to suppress textures and noise while maintaining meaningful structure based on colour intensity variation, but it is intractable due to texture patterns or error. This study presents an edge‐preserving suppression method for depth estimation. The authors formulate a functional energy function based on the relative total intensity and space variation, and they minimise the energy function via iteratively reweighted least squares. Assuming that textural edges most likely correspond to depth discontinuities, they exploit the comparative variations of the colour image to produce a more accurate depth map. The experimental results demonstrate the usefulness of the proposed approach, and show that texture patterns are suppressed while meaningful edges are preserved. According to the results of the depth acquisition methods, the proposed depth estimation methods generate the accurate and robust results. Eu-Tteum Baek, Yo-Sung Ho |
IET Image Process. | 2 |
| 2018 | 3D Scene Reconstruction Using Colorimetric and Geometric Constraints on Iterative Closest Point Method
Dong-Won Shin, Yo-Sung Ho |
Multim. Tools Appl. | 2 |
| 2018 | Multistage Pooling for Blind Quality Prediction of Asymmetric Multiply-Distorted Stereoscopic ImagesabstractQuality prediction for asymmetric multiply-distorted stereoscopic images (MDSIs) confronts more challenges than previous stereoscopic image quality assessment (SIQA) issues, whereas the existing no-reference SIQA methods have been limited to understand the asymmetric distortions and multiple distortions simultaneously for general-purpose blind quality prediction. In this paper, we propose a multistage pooling (MUSP) model for quality prediction of asymmetric MDSIs. In the training stage, we establish multimodal sparse representation framework for phase and amplitude components, respectively. In the testing stage, we use an MUSP strategy to simulate the pooling procedure undergoing multimodal quality pooling, feature pooling, binocular pooling, and phase-amplitude quality pooling in order. Experimental results on our new established database (NBU-MDSID Phase-II) demonstrate the effectiveness of our blind metric. Feng Shao 0001, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2017 | High throughput entropy decoder design for H.265/HEVC
Jung-Ah Choi, Yo-Sung Ho |
Multim. Tools Appl. | 2 |
| 2016 | Tutorials: 3D video processing techniques for immersive contents generationabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. With the emerging market of 3D imaging products, 3D video has become an active area of research and development in recent years. 3D video is the key to provide more realistic and immersive perceptual experiences than the existing 2D counterpart. There are many applications of 3D video, such as 3D movie and 3DTV, which are considered the main drive of the next‐generation technical revolution. Stereoscopic display is the current mainstream technology for 3DTV, while auto-stereoscopic display is a more promising solution that requires more research endeavors to resolve the associated technical difficulties. In this tutorial lecture, we are going to cover the current state-of-the-art technologies of 3D video processing. After defining the basic requirements for 3D realistic multimedia services, we will cover various multi‐modal immersive media processing techniques. We also address the depth estimation problem for natural 3D scenes and discuss several challenging issues of 3D video processing, such as camera calibration, image rectification, illumination compensation and color correction. In addition, we are going to discuss the MPEG activities for 3D video coding, including depth map estimation, prediction structure for multi‐view video coding, multi-view video-plus-depth coding, and intermediate view synthesis for multi-view video display applications. Yo-Sung Ho |
VCIP | 1 |
| 2016 | Disparity map enhancement in pixel based stereo matching method using distance transform
Yong-Jun Chang, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Depth video spatial and temporal correlation enhancement algorithm based on just noticeable rendering distortion model
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 6 |
| 2014 | Discontinuity preserving disparity estimation with occlusion handling
Woo-seok Jang, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Geometric and colorimetric error compensation for multi-view images
Jae-Il Jung, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Guest editorial: Advances in 3D video processing
Shang-Hong Lai, Gene Cheung, Dinei A. F. Florêncio, Peter Eisert, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 5 |
| 2013 | Low-bit depth-high-dynamic range image generation by blending differently exposed imagesabstractRecently, high‐dynamic range (HDR) imaging has taken the centre stage because of the drawbacks of low‐dynamic range imaging, namely detail losses in under‐ and over‐exposed areas. In this study, the authors propose an algorithm for HDR image generation of a low‐bit depth from two differently exposed images. For compatibility with conventional devices, HDR image generations of a large bit depth and bit depth compression are skipped. By using posterior probability‐based labelling, luminance adjusting and adaptive blending, the authors directly blend two input images into one while preserving the global intensity order as well as enhancing its dynamic range. From the experiments on various test images, results confirm that the proposed method generates more natural HDR images than other state‐of‐the‐art algorithms regardless of image properties. Jae-Il Jung, Yo-Sung Ho |
IET Image Process. | 2 |
| 2012 | Improved near-lossless HEVC codec for depth map based on statistical analysis of residual dataabstractA depth map represents three-dimensional (3D) data and is used for depth image-based rendering (DIBR) to synthesize virtual views. Since quality of the rendering view depends on that of the depth map, we should encode the depth map by maintaining the original quality. However, lossless depth coding requires a large amount of bits to carry the encoded depth map over the network. Thus, in this paper, we encode the depth map using high efficiency video coding (HEVC) near-lossless coding considering trade-off between the high quality and the bit saving. In addition, in order to improve the compression efficiency of HEVC near-lossless coding, we use large coefficients clipping and limited codeword lengh Golomb-Rice (GR) code, considering the statistical analysis of residual data. Experimental results show that the proposed method provides approximately 16.59% and 0.54% bit savings compared to HEVC lossless depth coding and HEVC near-lossless depth map coding, without the significant degradation of the synthesized view quality. Jung-Ah Choi, Yo-Sung Ho |
ISCAS | 2 |
| 2012 | Disparity map acquisition with occlusion handling using warping constraintabstractIn this paper, we propose a stereo matching algorithm with occlusion handling. In order to detect occlusion, we obtain an initial disparity map via optimization based on modified constant-space belief propagation (CSBP). Such a method is advantageous due to its low complexity. The initial disparity maps provide clue for occlusion detection. From such clue, an energy function for occlusion detection is defined and optimized by energy minimization framework. We classify occlusion into two types from the obtained occlusion map and apply suitable occlusion handling process, respectively. The proposed occlusion handling method based on the potential energy function extends disparity values of visible pixels to occluded pixels. Experimental results show that generated disparity map of the proposed method has satisfactory quality. Woo-seok Jang, Yo-Sung Ho |
ISCAS | 2 |
| 2012 | A framework of 3D video coding using view synthesis predictionabstractThe advanced 3D video system employs the multi-view video plus depth (MVD) format to support free-viewpoint navigation and comfortable 3D video. Therefore, the prediction structure of the multi-view video coding (MVC) can be used for 3D video coding. The view synthesis prediction method is designed to exploit inter-view correlation using the virtual view generation; hence it is suitable for 3D video coding. In this paper, we propose an efficient framework for 3D video coding using view synthesis prediction to compress multi-view color and depth data simultaneously. We designed the coding procedure of MVD data with four types of view synthesis methods according to the view position. The experimental results showed that the proposed framework improved the coding performance at most 0.9 dB for the multi-view color videos. Cheon Lee, Yo-Sung Ho |
PCS | 2 |
| 2012 | Adaptive depth boundary sharpening for effective view synthesisabstractThis paper focuses on sharpening boundaries in depth maps for multi-view video coding. Artifacts around boundaries degrade the quality of synthesized images. In order to encounter this problem, after applying the deblocking filter for each frame, we create a binary edge map and find the location of blocks that need to be altered. Subsequently, we apply a boundary sharpening filter which uses pixel frequency, similarity, and closeness as sub-costs. This filter is only applied to blocks which are near edges. Experimental results exhibit much more visual comfort level in synthesized images compared to when JMVC was used. Yunseok Song, Cheon Lee, Yo-Sung Ho, HoCheon Wey, Jaejoon Lee |
PCS | 3 |
| 2012 | Video processing techniques for 3D televisionabstractIn recent years, various multimedia services have become available and the demand for three-dimensional television (3DTV) is growing rapidly. Since 3DTV is considered as the next generation broadcasting service that can deliver real and immersive experiences by supporting user-friendly interactions, a number of advanced three-dimensional video technologies have been studied. Among them, multi-view video coding is the key technology for various applications including free-viewpoint video, free-viewpoint television, 3DTV, immersive teleconference, and surveillance systems. In this tutorial lecture, we are going to cover the current state-of-the-art technologies for 3D video: representation of 3D scenes, acquisition of 3D video contents, illumination compensation and color correction, camera calibration and image rectification, depth map modeling and enhancement, 3-D warping and depth map refinement, coding of multi-view video and depth map, hole filling for occluded objects, and view synthesis using homography. After defining the basic requirements for realistic 3D broadcasting services, we will cover various multi-modal immersive media processing technologies. Yo-Sung Ho |
VCIP | 1 |
| 2012 | Residual coding of depth map with transform skippingabstractSince advanced 3D video systems employ depth information to support free-viewpoint navigation and comfortable 3D video viewing, efficient depth map coding is necessary for future 3D video systems. Most residual data in depth map coding are generated along abrupt depth discontinuities, represented by near-zero and high-magnitude values. In this paper, we model the residual data with two representative values calculated by the K-means clustering method and send them to the decoder by skipping transformation. After best mode decision, we applied the proposed method to a block containing residual data, and then we send the quantized representative values to decoder if its coding rate is less than the conventional best mode. By conducting INTRA only coding, -20.32% bit saving was achieved. Cheon Lee, HoCheon Wey, Jaejoon Lee, Yo-Sung Ho |
VCIP | 4 |
| 2012 | Depth Video Coding Using Adaptive Geometry Based Intra Prediction for 3-D Video SystemsabstractDepth video coding is an essential part of 3-D video processing systems. Specifically, object boundary regions are important in depth video coding since these regions significantly affect the visual quality of a synthesized view. In this paper, we propose an efficient depth video coding method to determine precise intra prediction modes and thereby reduce the loss of boundary information. To achieve this objective, we analyze and exploit statistical and geometric characteristics of the depth video. Experimental results subsequently show that the proposed method performs better than the original intra prediction of H.264/AVC in terms of bit savings and rendering quality. Min-Koo Kang, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2012 | Asymmetric Coding of Multi-View Video Plus Depth Based 3-D Video for View RenderingabstractThe recent years have witnessed three-dimensional (3-D) video technology to become increasingly popular, as it can provide high-quality and immersive experience to end users, where view rendering with depth-image-based rendering (DIBR) technique is employed to generate the virtual views. Distortions in depth map may induce geometry changes in the virtual views, and distortions in texture video may be propagated to the virtual views. Thus, effective compression of both texture videos and depth maps is important for 3-D video system. From the perspective of bit allocation, asymmetric coding of the texture videos and depth maps is an effective way to get the optimal solution of 3-D video compression and view rendering problems. In this paper, a novel asymmetric coding method of multi-view video plus depth (MVD) based 3-D video is proposed on purpose of providing high-quality view rendering. In the proposed method, two models are proposed to characterize view rendering distortion and binocular suppression in 3-D video. Then, an asymmetric coding method of MVD-based 3-D video is proposed by combining two models in encoding framework. Finally, a chrominance reconstruction algorithm is presented to achieve accurate reconstruction. Experimental results show that compared with other methods, the proposed method can obtain higher performance of view rendering under the total bitrate constraint. Moreover, the perceptual visual quality of 3-D video is almost unaffected with the proposed method. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2011 | H.264/AVC based near lossless intra codec using line-based prediction and modified CABACabstractIn this paper, we propose a new H.264/AVC based intra codec for near lossless coding. The proposed algorithm is composed of two parts: line-based intra prediction and modified context-based adaptive binary arithmetic coding (CABAC). Experimental results show that the proposed method provides about 8.95% bit savings, compared to the current H.264/AVC FRExt high profile. Jung-Ah Choi, Jin Heo, Yo-Sung Ho |
ICME | 3 |
| 2011 | Improved CABAC design in H.264/AVC for lossless depth map codingabstractThe depth map represents three-dimensional (3D) data and is used for depth image-based rendering (DIBR) to synthesize virtual views. Since the quality of synthesized virtual views highly depends on the quality of the depth map, we encode the depth map by lossless coding mode. However, context-based adaptive binary arithmetic coding (CABAC) for the H.264/AVC standard does not guarantee the best coding performance for lossless depth map coding because CABAC was originally designed for lossy coding. In this paper, we propose an improved coding method of CABAC for lossless depth map coding considering the statistical properties of residual data from lossless depth map coding. Experimental results show that the proposed CABAC method provides approximately 4.3% bit saving, compared to the original CABAC in H.264/AVC. Jin Heo, Yo-Sung Ho |
ICME | 2 |
| 2011 | Improved Entropy Coder in H.264/AVC for Lossless Residual Coding in the Spatial Domain
Jin Heo, Yo-Sung Ho |
PSIVT (1) | 2 |
| 2011 | Efficient Stereo Image Rectification Method Using Horizontal Baseline
Yun-Suk Kang, Yo-Sung Ho |
PSIVT (1) | 2 |
| 2011 | Depth Map Up-Sampling Using Random Walk
Gyo-Yoon Lee, Yo-Sung Ho |
PSIVT (1) | 2 |
| 2011 | Virtual Viewpoint Disparity Estimation and Convergence Check for Real-Time View Synthesis
In-Yong Shin, Yo-Sung Ho |
PSIVT (1) | 2 |
| 2011 | Fast Mode Decision Algorithm for Depth Coding in 3D Video Systems Using H.264/AVC
Da-Hyun Yoon, Yo-Sung Ho |
PSIVT (2) | 2 |
| 2011 | Fast depth video coding method using adaptive edge classificationabstractIn this paper, we propose a fast mode decision algorithm for both the intra prediction and inter prediction in the depth video sequence. The proposed algorithm reduces the complexity of the depth video coding. According to the depth variation, depth video can be classified into depth-continuity and depth-discontinuity regions. From experiments, we determine a threshold value for classifying these regions. Since the depth-continuity region has an imbalance in the mode distribution, we limit the mode candidates to reduce the complexity of the mode decision process. Experimental results show that our proposed algorithm reduces encoding time up to 31% and 98% for the intra and inter frames, respectively, compared to the H.264/AVC standard with negligible PSNR loss and bit rate increase. Da-Hyun Yoon, Yo-Sung Ho |
VCIP | 2 |
| 2011 | Generation of high-quality depth maps using hybrid camera system for 3-D video
Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | High-quality non-blind image deconvolution with adaptive regularization
Jongho Lee 0004, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Depth Coding Using a Boundary Reconstruction Filter for 3-D Video SystemsabstractA depth image is 3-D information used for virtual view synthesis in 3-D video system. In depth coding, the object boundaries are hard to compress and severely affect the rendering quality since they are sensitive to coding errors. In this paper, we propose a depth boundary reconstruction filter and utilize it as an in-loop filter to code the depth video. The proposed depth boundary reconstruction filter is designed considering occurrence frequency, similarity, and closeness of pixels. Experimental results demonstrate that the proposed depth boundary reconstruction filter is useful for efficient depth coding as well as high-quality 3-D rendering. Kwan-Jung Oh, Anthony Vetro, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Efficient intra coding structure for high resolution videos using line-by-line prediction and adaptive transform selectionabstractIn this paper, an efficient intra coding structure for high resolution videos is proposed. First, we use a line-by-line intra 16×16 prediction to improve the prediction accuracy. Second, we modify the intra coding structure for the line-by-line prediction mode. Experimental results show that the proposed method provides approximately 10.415% bit saving, compared to the current H.264/AVC FPExt high profile over several high-definition (HD) test sequences. Jung-Ah Choi, Yo-Sung Ho |
ICIP | 2 |
| 2010 | Line-by-line intra 16×16 prediction for high-quality video codingabstractIn this paper, we proposed an intra 16×16 line-by-line (LbL) prediction. First, we find suitable quantization parameters (QP) for the high resolution videos without the significant subjective quality degradation. Then, we show that closer reference pixels give better prediction through the experiment. Using this observation, we propose an improved intra prediction method. In the proposed method, the line-based prediction is performed instead of the block-based prediction used in the current H.264/AVC. Experimental results show that the proposed algorithm reduces approximately 6.07% bit-rate compared to the H.264/AVC. Jung-Ah Choi, Yo-Sung Ho |
ICME | 2 |
| 2010 | High-quality multi-view depth generation using multiple color and depth camerasabstractIn this paper, we propose a high-quality multi-view depth generation method using multiple color and depth cameras. After we capture low-resolution depth maps by three TOF cameras, the depth information is warped into color image positions and used as the initial disparity value. By applying the stereo matching using belief propagation with the initial disparity information, we have obtained more accurate and stable multi-view disparity maps, compared to those results without the initial disparity information. Yun-Suk Kang, Yo-Sung Ho |
ICME | 2 |
| 2010 | Adaptive geometry-based intra prediction for depth video codingabstractIn this paper, we propose an efficient depth video coding method for the 3D video (3DV) system. Since the boundary information of depth video significantly affects the rendering quality in the 3DV system, the proposed method reduces the loss of the boundary information by producing precise intra prediction modes. The characteristics of depth video are analyzed and considered to modify the previous geometry-adaptive block partitioning in the proposed method. The proposed method guarantees a better performance of intra prediction than the method of the H.264/AVC. Experimental results have shown that 0.33 dB coding gain in terms of PSNR and subjective quality improvement of synthesized views are achieved by the proposed method. Min-Koo Kang, Cheon Lee, Yo-Sung Ho |
ICME | 4 |
| 2010 | Fast color correction for multi-view video by modeling spatio-temporal variation
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 4 |
| 2010 | Efficient entropy coding scheme for H.264/AVC lossless video coding
Seung-Hwan Kim 0001, Jin Heo, Yo-Sung Ho |
Signal Process. Image Commun. | 3 |
| 2010 | Efficient Level and Zero Coding Methods for H.264/AVC Lossless Intra CodingabstractSince H.264/AVC was designed mainly for lossy video coding, the entropy coding methods in H.264/AVC are not appropriate for lossless video coding. Based on statistical differences of residual data in lossy and lossless coding, we develop efficient level and zero coding methods. Therefore, we design an improved context-based adaptive variable length coding (CAVLC) scheme for lossless intra coding by modifying the relative entropy coding parts in H.264/AVC. Experimental results show that the proposed method provides approximately 6.8% bit saving, compared with the H.264/AVC FRExt high profile. Jin Heo, Yo-Sung Ho |
IEEE Signal Process. Lett. | 2 |
| 2010 | Improved Context-Based Adaptive Binary Arithmetic Coding over H.264/AVC for Lossless Depth Map CodingabstractThe depth map, which represents three-dimensional (3D) information, is used to synthesize virtual views in the depth image-based rendering (DIBR) method. Since the quality of synthesized virtual views highly depends on the quality of depth map, we encode the depth map under the lossless coding mode. The original context-based adaptive binary arithmetic coding (CABAC) that was originally designed for lossy texture coding cannot provide the best coding performance for lossless depth map coding due to the statistical differences of residual data in lossy and lossless depth map coding. In this letter, we propose an enhanced CABAC coding mechanism for lossless depth map coding based on the statistics of residual data. Experimental results show that the proposed CABAC method provides approximately 4% bit saving compared to the original CABAC in H.264/AVC. Jin Heo, Yo-Sung Ho |
IEEE Signal Process. Lett. | 2 |
| 2010 | Improved CAVLC for H.264/AVC Lossless Intra-CodingabstractContext-based adaptive variable length coding (CAVLC) for the H.264/advanced video coding (AVC) standard was originally designed for lossy video coding, and as such does not yield adequate performance for lossless video coding. In this paper, we propose an improved CAVLC for lossless intra-coding by considering the statistical differences in residual data between lossy and lossless coding. From experimental results, we confirm that the proposed method provides approximately 9% bit saving in terms of a compression ratio compared with the current H.264/AVC fidelity range extensions high profile. Jin Heo, Seung-Hwan Kim 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | New CAVLC design for lossless intra codingabstractThe context-based adaptive variable length coder (CAVLC) in H.264/AVC is not appropriate for lossless video coding because it was designed for lossy video coding. Since statistical characteristics of residual data in lossy and lossless coding are quite different, we design a new VLC table for the number of non-zero coefficients and an adaptive scheme for VLC table selection in level coding for lossless intra coding. Experimental results show that the proposed CAVLC scheme provides approximately 10% bit saving, compared to the original CAVLC scheme in H.264/AVC. Jin Heo, Seung-Hwan Kim 0001, Yo-Sung Ho |
ICIP | 3 |
| 2009 | New CAVLC encoding algorithm for lossless intra coding in H.264/AVCabstractContext-based adaptive variable length coding (CAVLC) of H.264/AVC was originally designed for quantized transform coefficients in lossy video coding, and as such does not yield adequate performance in lossless video coding. In this paper, we propose an improved CAVLC for lossless intra coding. Considering statistical differences of residual data in lossy and lossless coding, we design a new CAVLC encoding algorithm. Experimental results show that the proposed method provides approximately 7.6% bit savings, compared to the current H.264/AVC FRExt high profile. Jin Heo, Seung-Hwan Kim 0001, Yo-Sung Ho |
PCS | 3 |
| 2009 | Hole filling method using depth based in-painting for view synthesis in free viewpoint television and 3-D videoabstractDepth image-based rendering (DIBR) is generally used to synthesize virtual view images in free viewpoint television (FTV) and three-dimensional (3-D) video. One of the main problems in DIBR is how to fill the holes caused by disocclusion regions and inaccurate depth values. In this paper, we propose a new hole filling method using a depth based in-painting technique. Experimental results show that the proposed hole filling method provides improved rendering quality both objectively and subjectively. Kwan-Jung Oh, Sehoon Yea, Yo-Sung Ho |
PCS | 3 |
| 2009 | Depth Reconstruction Filter and Down/Up Sampling for Depth Coding in 3-D VideoabstractA depth image represents three-dimensional (3-D) scene information and is commonly used for depth image-based rendering (DIBR) to support 3-D video and free-viewpoint video applications. The virtual view is generally rendered by the DIBR technique and its quality depends highly on the quality of depth image. Thus, efficient depth coding is crucial to realize the 3-D video system. In this letter, we propose a depth reconstruction filter and depth down/up sampling techniques to improve depth coding performance. Experimental results demonstrate that the proposed methods reduce the bit-rate for depth coding and achieve better rendering quality. Kwan-Jung Oh, Sehoon Yea, Anthony Vetro, Yo-Sung Ho |
IEEE Signal Process. Lett. | 4 |
| 2008 | Joint coding of multi-view video and corresponding depth mapabstractIn this paper, we propose a joint coding scheme for both multi-view video and its corresponding depth map. After we synthesize a virtual image for the target view using adjacent view images and their depth information, we apply a view interpolation prediction (VIP) method for both multi-view video coding and its depth data coding. In order to improve the synthesized virtual view, we also propose a hole filling method that can compensate for empty regions caused by the 3D warping operation. With the proposed algorithm, we have obtained approximately 0.65 dB of the PSNR gain on average for the multi-view depth data, and 0.17 dB of the PSNR gain for the multi-view video data, compared to JMVM 1.0. Sang-Tae Na, Kwan-Jung Oh, Yo-Sung Ho |
ICIP | 3 |
| 2008 | Multi-view depth video coding using depth view synthesisabstractDepth information indicates the distance of an object in the three dimensional (3D) scene from the camera view-point, typically represented by eight bits. Since the depth map is useful in various multimedia applications, such as three dimensional television (3DTV) and free-viewpoint television (FTV), we need to acquire a single or multi-view depth maps and process them effectively. In this paper, we propose a new coding scheme for multi-view depth video data using depth view synthesis. We first apply a 3D warping method to synthesize a virtual depth image for the current view using the multi-view depth information. We also propose a hole filling method to compensate for the holes generated during the depth map synthesis process. Finally, we utilize the synthesized depth map for the current view as an additional reference frame in encoding the current depth map. Experimental results show that the proposed algorithm achieves approximately 0.69 dB of PSNR gain on average, compared to JMVM 1.0. Sang-Tae Na, Kwan-Jung Oh, Cheon Lee, Yo-Sung Ho |
ISCAS | 4 |
| 2007 | Mesh-Based Depth Coding for 3D Video using Hierarchical Decomposition of Depth MapsabstractIn this paper, we present a new coding scheme for depth maps using a hierarchical decomposition. After we decompose a depth map into three disjoint images and a layer descriptor according to the region of edges, we merge the disjoint images of each depth map into an image. Then, the merged images and the layer descriptor are coded by H.264/AVC. Unlike previous mesh-based depth coding methods, we compress the irregular depth information using a conventional 2D video coder. Experimental results show that our scheme improves compression efficiency of mesh-based depth coding. Sung-Yeol Kim, Yo-Sung Ho |
ICIP (5) | 2 |
| 2007 | Lossless Data Hiding for Medical Images with Patient InformationabstractThis paper presents a lossless data hiding algorithm for medical images, where we embed the patient information into the segmented liver region of the CT image. This algorithm utilizes the characteristics of difference images and modifies pixel values slightly to embed a large amount of data while keeping high visual quality. Sang-Kwang Lee, Seong-Jae Lim, Young-Ho Suh, Yo-Sung Ho |
ICIP (3) | 4 |
| 2007 | A fast inter mode decision algorithm in H.264/AVC for IPTV broadcasting servicesabstractThe new video coding standard H.264/AVC employs the rate-distortion optimization (RDO) method for choosing the best coding mode. However, since it increases the encoder complexity tremendously, it is not suitable for real-time applications, such as IPTV broadcasting services. Therefore we need a fast mode decision algorithm to reduce its encoding time. In this paper, we propose a fast mode decision algorithm considering quantization parameter (QP) because we have noticed that the frequency of best modes depends on QP. In order to consider these characteristics, we use the coded block pattern (CBP) that has "0" value when all quantized discrete cosine transform (DCT) coefficients are zero. We also use both the early SKIP mode and early 16x16 mode decisions. Experimental results show that the proposed algorithm reduces the encoding time by 74.6% for the baseline profile and 72.8% for the main profile, compared to the H.264/AVC reference software. Geun-Yong Kim, Bin-Yeong Yoon, Yo-Sung Ho |
VCIP | 3 |
| 2007 | Fine Granular Scalable Video Coding Using Context-Based Binary Arithmetic Coding for Bit-Plane CodingabstractIn this paper, we propose a new efficient entropy coding scheme, namely context-based binary arithmetic coding for bit-plane coding (CBACBP), for the enhancement layer of fine granular scalable (FGS) video coding. The basic structure of proposed CBACBP is based on the traditional bit-plane coding of MPEG-4 FGS. However, in order to enhance the coding efficiency of bit-plane coding, a newly designed context-based binary arithmetic coding scheme is used. In CBACBP, we apply three types of probability estimation techniques. The first type relies on local information of a given symbol to code, such as bit-plane level, data level and frequency component. In the second type, the preceding symbols in the same and/or higher bit-planes are used. In the third type, we consider complexity of a block including the symbol based on the number of coded nonzero coefficients in the block. Experimental results show that the proposed CBACBP improves the PSNR up to 1.2 and 0.5 dB compared with MPEG-4 FGS and JSVM, respectively. Seung-Hwan Kim 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Multiple Color and Depth Video Coding Using a Hierarchical RepresentationabstractThis paper presents coding schemes for multiple color and depth video using a hierarchical representation. We use the concept of layered depth image (LDI) to represent and process multiview video with depth. After converting those data to the proposed representation, we encode color, depth, and auxiliary data representing the hierarchical structure, respectively. Two kinds of preprocessing approaches are proposed for multiple color and depth components. In order to compress auxiliary data, we have employed a near lossless coding method. Finally, we have reconstructed the original viewpoints successfully from the decoded LDI frames. From our experiments, we realize that the proposed approach is useful for dealing with multiple color and depth data simultaneously. Seung-Uk Yoon, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | A Hierarchical Approach to Rotation-Invariant Texture Feature Extraction Based on Radon Transform ParametersabstractIn this paper, we propose an efficient hierarchical method for extracting invariant texture features using the Gabor wavelets and Radon transform parameters. The proposed method applies the Radon transform to estimate the directional information in the high-band texture image extracted by Gabor wavelets. The directional information is then used to make the texture feature invariant to rotation. To show the efficiency of our scheme, we developed a texture-based image retrieval system based on the proposed method and evaluated it on a set of images from the Brodatz album. Experimental results show that the proposed system outperforms previous rotation-invariant systems significantly. Mahmoud R. Hejazi, Yo-Sung Ho |
ICIP | 2 |
| 2006 | Public Key Watermarking for Reversible Image AuthenticationabstractIn this paper, we propose a new public key watermarking scheme for reversible image authentication where if the image is authentic, the distortion due to embedding can be completely removed from the watermarked image after the hidden data has been extracted. This technique utilizes histogram characteristics of the difference image and modifies pixel values slightly to embed more data than other lossless data hiding algorithm. We show that the lower bound of the PSNR (peak-signal-to-noise-ratio) values of watermarked images are 51.14 dB. Moreover, the proposed scheme is quite simple and the execution time is rather short. Experimental results demonstrate that the proposed scheme can detect any modifications of the watermarked image. Sang-Kwang Lee, Young-Ho Suh, Yo-Sung Ho |
ICIP | 3 |
| 2006 | Reversiblee Image Authentication Based on WatermarkingabstractIn this paper, we propose a new reversible image authentication technique based on watermarking where if the image is authentic, the distortion due to embedding can be completely removed from the watermarked image after the hidden data has been extracted. This technique utilizes histogram characteristics of the difference image and modifies pixel values slightly to embed more data than other lossless data hiding algorithm. We show that the lower bound of the PSNR (peak-signal-to-noise-ratio) values of watermarked images are 51.14 dB. Moreover, the proposed scheme is quite simple and the execution time is rather short. Experimental results demonstrate that the proposed scheme can detect any modifications of the watermarked image Sang-Kwang Lee, Young-Ho Suh, Yo-Sung Ho |
ICME | 3 |
| 2006 | Color Data Coding for Three-Dimensional Mesh Models Considering Connectivity and Geometry InformationabstractIn this paper, we propose a new predictive coding scheme for color data of three-dimensional (3-D) mesh models. We exploit connectivity and geometry information to improve coding efficiency. After ordering all vertices in a 3-D mesh model with a connectivity coding technique, we propose a geometry predictor to compress the color data efficiently. The predicted color can be obtained by a weighted sum of reconstructed colors for adjacent vertices using both angles and distances between the current vertex and adjacent vertices. Simulation results show that the proposed scheme provides enhanced coding efficiency over previous works for various 3-D mesh models Young-Suk Yoon, Sung-Yeol Kim, Yo-Sung Ho |
ICME | 3 |
| 2006 | Modified Discrete Radon Transforms and Their Application to Rotation-Invariant Image AnalysisabstractThis paper presents two novel transforms based on the discrete Radon transform. The proposed transforms smartly solve two inherent problems of the Radon transform in rotation estimation in digital images, i.e., direction-dependency and nonhomogeneity, that come from the different numbers of pixels projected on a line for different directions and/or coordinates of a direction. While the first transform considers the sample mean operator on the same sets of pixels for a direction instead of summation in the discrete Radon transform, the second transform uses the mean operator on sets of pixels with the equal number of elements. In order to show the efficiency of the proposed transforms, we apply them on image collections from the Brodatz album for estimating the directional information. Experimental results show a significant increase in correct estimation as well as in the processing time compared to the conventional Radon transform Mahmoud R. Hejazi, Georgy L. Shevlyakov, Yo-Sung Ho |
MMSP | 3 |
| 2006 | H.264-Based Depth Map Sequence Coding Using Motion Information of Corresponding Texture Video
Han Oh, Yo-Sung Ho |
PSIVT | 2 |
| 2006 | Automatic liver segmentation for volume measurement in CT Images
Seong-Jae Lim, Yong-Yeon Jeong, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 3 |
| 2006 | Predictive compression of geometry, color and normal data of 3-D mesh modelsabstractPredictive compression algorithms for geometry, color and normal data of three-dimensional (3-D) mesh models are proposed in this work. In order to eliminate redundancies in geometry data, we predict each vertex position by exploiting the position and angle information in neighboring triangles. To compress color data, we propose a mapping table scheme that compresses frequently recurring colors efficiently. For normal data, we propose an average predictor and a 6-4 subdivision quantizer to improve coding gain. Simulation results demonstrate that the proposed algorithm provides better performance than the MPEG-4 standard for 3-D mesh model coding (3-DMC). Jeong-Hwan Ahn, Chang-Su Kim 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Desing of reversible variable-length codes using propertfes of the huffman code and average length functionabstractVariable-length codes (VLCs) are generally employed lo improve compression efficiency using data statistics. However, VLCs are very sensitive to bit errors in noisy transmission environments, such as mobile channels. Recently, several reversible variable-length codes (RVLCs) have been introduced due to recovering information from corrupted compressed bitstreams and enhancing robustness of VLCs to bit errors. However, existing RVLCs have some rooms for improvement in coding efficiency. In this paper, we propose a new design algorithm for efficient symmetrical and asymmetrical RVLCs by employing essential information from the Huffman code and the property of the average length function. The proposed algorithm has demonstrated improved coding efficiency over existing RVLC algorithms. Wook-Hyun Jeong, Young-Suk Yoon, Yo-Sung Ho |
ICIP | 3 |
| 2004 | Design of robust reversible variable-length codes using the property of free distanceabstractIn recent years, reversible variable-length codes (RVLC) have been proposed to recover correct information from erroneous bitstreams. The free distance of a binary code indicates the strength of the code to transmission errors. Since the minimum free distance of RVLC is one, we cannot ensure that encoded data are little exposed. We propose a new code design algorithm for robust RVLC by enhancing the free distance. In the proposed algorithm, we exploit properties of the Huffman code and the average length function satisfying the distance condition in order to improve the code performance. Wook-Hyun Jeong, Young-Suk Yoon, Yo-Sung Ho |
ICME | 3 |
| 2004 | Rate control algorithm for H.264/AVC video coding standard based on rate-quantization modelabstractRate control is an essential part of video coding algorithms, to maintain uniform picture quality for the given coding constraints. We propose a new rate control algorithm for the H.264 video coding standard. We derive a rate-quantization model from the rate-distortion function, based on the distribution of source data to be quantized. We also classify macro-blocks in each frame into three groups according to their characteristics, and decide the quantization parameter for each macroblock. Experimental results show that the proposed scheme generates coding bits very close to target bits and provides improved coding efficiency at low bit rates. Seonki Kim, Yo-Sung Ho |
ICME | 2 |
| 2003 | Multiple Description Coding for Image Data Hiding Jointly in the Spatial and DCT Domains
Mohsen Ashourian, Yo-Sung Ho |
ICICS | 2 |
| 2003 | An efficient coding algorithm for color and normal data of three-dimensional mesh modelsabstractThree-dimensional (3D) mesh models have attribute data, such as colors, normal vectors, and texture coordinates to render or shade the surface of the mesh. Although several coding schemes have been developed to represent the topology and geometry information of the 3D mesh, coding of the attribute data has received less attention. In this paper, we propose a new predictive coding scheme for colors and normal vectors of the 3D mesh model, where we predict colors and normals based on several ancestors along the vertex ordering. In order to encode the color information, we define a mapping table that specifies how colors are mapped into other vertices. The mapping table can represent frequently occurring color patterns efficiently. For normal vectors, we also propose an average predictor and the 6-4 subdivision quantizer in the spherical coordinate system. The proposed scheme has demonstrated good coding efficiency for various VRML test data. Jeong-Hwan Ahn, Chang-Su Kim 0001, Yo-Sung Ho |
ICIP (1) | 3 |
| 2003 | Video segmentation using vector-valued diffusion and clusteringabstractIn this paper, we propose a user-assisted video segmentation algorithm based on color information to alleviate oversegmentation problems. We perform intra-frame segmentation by image simplification, region labeling, and color clustering. In this paper, we also present a discrete three dimensional diffusion model for easy implementation. The statistical property of each labeled region is used to estimate the number of total clusters, and agglomerative hierarchical clustering is performed with the estimated number of clusters. Since the proposed clustering algorithm counts each region as a unit, it does not generate oversegmentation problems along region boundaries. For inter-frame segmentation, we employ a look-up table for foreground color clusters, track the foreground regions, and utilize those information to extract moving objects. Daehee Kim 0005, Chung-Hyun Ahn, Yo-Sung Ho |
ICIP (1) | 3 |
| 2003 | A Three-Dimensional Watermarking Algorithm Using the DCT Transform of Triangle Strips
Jeonghee Jeon, Sang-Kwang Lee, Yo-Sung Ho |
IWDW | 3 |
| 2003 | Three-dimensional mesh simplification using normal variation error metric and modified subdivided edge classification
Eun-Young Chang, Chung-Hyun Ahn, Yo-Sung Ho |
VCIP | 3 |
| 2003 | Active camera tracking system using affine motion compensation
Young-Kee Jung, Yo-Sung Ho |
VCIP | 2 |
| 2003 | View-dependent transmission of 3D mesh models using hierarchical partitioning
Sung-Yeol Kim, Jeong-Hwan Ahn, Yo-Sung Ho |
VCIP | 3 |
| 2003 | Bit allocation for MPEG-4 video coding with spatio-temporal tradeoffsabstractThis paper describes rate-control algorithms that consider the tradeoff between coded quality and temporal rate. We target improved coding efficiency for both frame-based and object-based video coding. We propose models that estimate the rate-distortion characteristics for coded frames and objects, as well as skipped frames and objects. Based on the proposed models, we propose three types of rate-control algorithms. The first is for frame-based coding, in which the distortion of coded frames is balanced with the distortion incurred by frame skipping. The second algorithm applies to object-based coding, where the temporal rate of all objects is constrained to be the same, but the bit allocation is performed at the object level. The third algorithm also targets object-based coding, but in contrast to the second algorithm, the temporal rates of each object may vary. The algorithm also takes into account the composition problem, which may cause holes in the reconstructed frame when objects are encoded at different temporal rates. We propose a solution to this problem that is based on first detecting changes in the shape boundaries over time at the encoder, then employing a hole detection and recovery algorithm at the decoder. Overall, the proposed algorithms are able to achieve the target bit rate, effectively code frames and objects with different temporal rates, and maintain a stable buffer level. Anthony Vetro, Yao Wang 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2002 | Geometry Compression of 3-D Mesh Models Using a Joint PredictionabstractSummary form only given. We address geometry coding of 3D models. Conventional geometry coding schemes quantize vertex positions using a bounding box before coding it by an entropy coder. However, in the proposed scheme, we first predict vertex positions and then quantize them differentially. Our geometry encoder consists of four stages: preprocessing, prediction, quantization, and entropy coding. Using a joint prediction, the encoder predicts vertex positions in a layer traversal order. Although the joint prediction is based on the parallelogram prediction, we use every possible preceding triangle to approximate the predicted vertex. Jeong-Hwan Ahn, Yo-Sung Ho |
DCC | 2 |
| 2002 | MPEG-4 video object-based rate allocation with variable temporal ratesabstractThis paper describes a bit allocation algorithm to achieve a constant bit rate when coding multiple video objects (MVOs), while improving the rate-distortion (R-D) performance over the reference method for MPEG-4 object-based rate control. In object-based coding, bit allocation is performed at the object level and temporal rates of different objects may vary. In this paper, we deal with these two issues. We pay particular attention to maintenance of buffer occupancy levels and propose a new method for spatiotemporal trade-offs for object-based coding. The proposed algorithm is able to successfully achieve the target bit rate, effectively code arbitrarily-shaped MVOs with different temporal rates, and maintain a stable buffer level. Anthony Vetro, Yao Wang 0001, Yo-Sung Ho |
ICIP (3) | 4 |
| 2002 | Motion-compensated coding of 3D animation models
Jeong-Hwan Ahn, Chang-Su Kim 0001, C.-C. Jay Kuo, Yo-Sung Ho |
VCIP | 4 |
| 2002 | Multiresolution motion compensation in the wavelet domain for scalable video coding
Jong-Tae Kim, Chang-Mo Yang, Dong-Keun Lim, Yo-Sung Ho |
VCIP | 4 |
| 2002 | Object-based rate allocation with spatio-temporal trade-offs
Anthony Vetro, Yao Wang 0001, Yo-Sung Ho |
VCIP | 4 |
| 2002 | Semi-automatic video segmentation algorithm using virtual blue screens
Jong-Han Shin, Daehee Kim 0005, Yo-Sung Ho |
VCIP | 3 |
| 2002 | Embedded image-coding algorithm using set partitioning in block trees of wavelet coefficients
Chang-Mo Yang, Yo-Sung Ho |
VCIP | 2 |
| 2001 | Shape-preserving progressive coding of 3D modelsabstractWe propose a shape-preserving progressive coding scheme for 3D models, focusing on the visual quality of 3D meshes at low resolutions. In order to obtain good reconstruction quality of 3D meshes, we improve a mesh simplification method based on the quadric error metric (QEM) by assigning different weighting factors according to topological constraints. We also develop an efficient geometric prediction algorithm that exploits neighboring information in the 3D mesh model to achieve high compression ratios. Jeong-Hwan Ahn, Eun-Young Chang, Yo-Sung Ho |
ICIP (3) | 3 |
| 2001 | Image retrieval using multi-scale color clusteringabstractA fundamental issue in content-based image retrieval is how to select image features that can represent image contents appropriately. A multi-scale color clustering algorithm based on human perceptual properties of color images is proposed for image retrieval. The multi-scale clustering algorithm is an unsupervised clustering method that utilizes the perceptual uniformity property in the (p,q) color space. The proposed color clustering algorithm produces a small set of representative color vectors for each image that capture color properties of the image, and a set of correlogram values that contain the spatial information of the image. Sehwan Kim, Woontack Woo, Yo-Sung Ho |
ICIP (1) | 3 |
| 2001 | Segmentation and compression techniques for 3D animation models based on motion trajectory in the spherical coordinate system
Jeong-Hwan Ahn, Yo-Sung Ho |
VCIP | 2 |
| 2001 | User-assisted segmentation algorithm using B-spline curves
Daehee Kim 0005, Yo-Sung Ho |
VCIP | 2 |
| 2001 | Fast-block-matching motion estimation algorithm using optimal search patterns
Dong-Keun Lim, Yo-Sung Ho |
VCIP | 2 |
| 2001 | Content-based event retrieval using semantic scene interpretation for automated traffic surveillanceabstractThis paper proposes an object segmentation and tracking algorithm for visual surveillance applications. In order to detect moving objects from a dynamic background scene which may have temporal clutters such as swaying plants, we devised an adaptive background update method and a motion classification rule. A two-dimensional token-based tracking system using a Kalman filter is designed to track individual objects under occlusion conditions. We propose a new occlusion reasoning approach where we consider two different types of occlusion: explicit occlusion and implicit occlusion. By tracking individual objects with segmented data, we can generate motion trajectories and set a motion model using polynomial curve fitting. The trajectory model is used as an indexing key for accessing the individual object in the semantic level. We also propose an efficient way of indexing and searching based on object-specific features at different semantic levels. The proposed searching scheme supports various queries including query by example, query by sketch, and query on weighting parameters for event-based retrieval. When retrieving an interested video clip, the system returns the best matching event in the similarity order. In addition, we implement a temporal event graph for direct accessing and browsing of a specific event in the video sequence. Young-Kee Jung, Kyu-Won Lee, Yo-Sung Ho |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2000 | Multiple-object tracking under occlusion conditions
Young-Kee Jung, Yo-Sung Ho |
VCIP | 2 |
| 1999 | An Efficient Geometry Compression Method for 3D Objects in the Spherical Coordinate SystemabstractIn this paper, we propose a new geometry coding scheme for 3D objects. The prediction error of each vertex position along a vertex spanning tree is obtained in the Cartesian coordinate system (x,y,z), and transformed into the spherical coordinate system (r,θ,φ). The magnitude r of the prediction error is quantized by an optimal uniform quantizer, and the corresponding pair of (θ,φ) is encoded according to the magnitude of the quantized value of the radius. We partition the sphere of the radius r into n2regions to allocate n2quantization points over the sphere. Experimental results of the proposed method demonstrate improved coding efficiency. Jeong-Hwan Ahn, Yo-Sung Ho |
ICIP (2) | 2 |
| 1999 | Image Segmentation Using Hierachical MeshesabstractThe object boundary of an image plays an important role in image interpretation. In this paper, we introduce a concept of hierarchical mesh-based image segmentation for finding object boundaries. In each hierarchical layer, we employ neighborhood searching and boundary tracking methods to refine the initial boundary estimate. We also apply a local region growing method to define closed contours. Experimental results indicate that reliable segmentation of objects can be accomplished by the proposed low complexity technique. Dong-Keun Lim, Yo-Sung Ho |
ICIP (1) | 2 |
| 1999 | A VOP generation tool: automatic segmentation of moving objects in image sequences based on spatio-temporal informationabstractThe new MPEG-4 video coding standard enables content-based functionalities. In order to support the philosophy of the MPEG-4 visual standard, each frame of video sequences should be represented in terms of video object planes (VOPs). In other words, video objects to be encoded in still pictures or video sequences should be prepared before the encoding process starts. Therefore, it requires a prior decomposition of sequences into VOPs so that each VOP represents a moving object. This paper addresses an image segmentation method for separating moving objects from the background in image sequences. The proposed method utilizes the following spatio-temporal information. (1) For localization of moving objects in the image sequence, two consecutive image frames in the temporal direction are examined and a hypothesis testing is performed by comparing two variance estimates from two consecutive difference images, which results in an F-test. (2) Spatial segmentation is performed to divide each image into semantic regions and to find precise object boundaries of the moving objects. The temporal segmentation yields a change detection mask that indicates moving areas (foreground) and nonmoving areas (background), and spatial segmentation produces spatial segmentation masks. A combination of the spatial and temporal segmentation masks produces VOPs faithfully. This paper presents various experimental results. Munchurl Kim, Jae-Gark Choi, Daehee Kim 0005, Hyung Lee, Myoung Ho Lee, Chieteuk Ahn, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 1998 | Image Warping using Adaptive Partial MatchingabstractIn this paper, we propose an adaptive partial matching method for motion estimation to reduce the computational complexity, while maintaining the image quality comparable to the hexagonal matching method. The proposed motion compensation method combines the fast affine transformation using a vector relationship. We simulate our proposed motion estimation method in a DCT-based coder by encoding CIF (common intermediate format) images at the bitrate of below 64 kb/s. The quality of reconstructed image with our method is substantially improved compared with the block matching algorithms (BMA), and is comparable to the hexagonal matching method. Computational complexity and coding bits are also reduced significantly relative to the BMA and the conventional image warping methods. Dong-Keun Lim, Yo-Sung Ho |
ICIP (2) | 2 |
| 1993 | Advanced Digital HDTV Transmission System for Terrestrial Video SimulcastingabstractTransmission aspects of the advanced digital high definition television (AD-HDTV) system, for terrestrial simulcast delivery of HDTV are described. In AD-HDTV, two quadrature-amplitude-modulated (QAM) carriers, with different power spectral densities, are employed in a frequency division multiplex (FDM) mode within the standard 6-MHz channel. The resulting spectral shaping allows a larger power to be transmitted, compared to that for a single QAM carrier, for the same level of perceptual interference into cochannel NTSC. The coded video data are split into high-priority (HP) data and standard-priority (SP) data, and the vital information is sent on the appropriate QAM carrier, resulting in a robust audio and video transmission system. The availability is higher in scenarios where the carrier-to-noise ratio (CNR) is above the threshold for HP reception but below the threshold for SP reception; this is important in fringe areas. The NTSC planning factors, suitably modified for HDTV delivery, are used to estimate the coverage area for AD-HDTV. The calculated AD-HDTV coverage area of 54.5 miles is comparable to that for NTSC transmission.> Samir N. Hulyalkar, Yo-Sung Ho, Kiran S. Challapali, David A. Bryan, Carlo Basile, Hugh White, Newman D. Wilson, Bhavesh Bhatt |
IEEE J. Sel. Areas Commun. | 2 |
| 1989 | Classified transform coding of images using vector quantizationabstractAn adaptive image coding scheme, called classified transform vector quantization, is proposed. It efficiently exploits correlation in large image blocks by taking advantage of transform coding (TC) and vector quantization (VQ), while overcoming the suboptimalities of TC and avoiding the complexity obstacle of VQ. After local mean luminance values are removed in the spatial domain using two-stage interpolative VQ, the residual errors are encoded in the transform domain by means of perceptual block classification and adaptive subvector construction. This scheme avoids the use of scalar quantization of DC coefficients in the transform domain and yet substantially reduces the blocking effect that tends to arise at low bit rates. Good reconstructed images have been obtained at rates between 0.3 and 0.4 bits/pixel, depending on the nature of the test images. The technique also permits progressive image transmission and reproduces errorless images with compression of about 5.0 bits/pixel. Experimental results indicate that an efficient bit allocation in the coding process produces a substantial improvement in performance.> Yo-Sung Ho, Allen Gersho |
ICASSP | 1 |
| 1988 | Variable-rate multi-stage vector quantization for image codingabstractA hierarchical approach to encoding of images using vector quantization (VQ) is described which allows VQ of large blocks with a tolerable computational complexity and permits progressive image reconstruction at bit rates as low as 0.3 bit/pixel. The technique used, multi-stage hierarchical VQ (MSHVQ), is a successive approximation approach for quantization that is based on recursive decomposition of larger image blocks into smaller subblocks. The encoder of MSHVQ consists of several stages, each of which encodes the quantization error generated from the previous stage. Digital decimation and interpolation techniques are used to convert between different vector dimensions in the higher level stages and to reduce the otherwise highly noticeable blocking effect in the reconstructed image. Variable-rate encoding increases the compression efficiency by allocating greater resolution to high-detail regions and provides a perceptually more consistent image quality throughout all areas of the decoded image. Experimental results show that MSHVQ yields a significantly improved image quality over conventional VQ of equivalent complexity.> Yo-Sung Ho, Allen Gersho |
ICASSP | 1 |