VLDB 2026 Research / reviewers in the wild / expert
Xuyuan Xu
dblp:08/9875
· DBLP profile ↗
34ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0000-6609-1046ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Systems, architecture and hardware · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSBoRA: A continual learning method for large language models with true orthogonality and reduced forgetting
Lai-Man Po, Farrell Hung, Zhuohan Wang, Haoxuan Wu, Kun Li 0015, Xuyuan Xu, Kwok-Wai Cheung 0002 |
Pattern Recognit. | 8 |
| 2025 | AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic AssessmentabstractMultimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations.However, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education.In this work, we propose AesBiasBench, a benchmark designed to evaluate MLLMs along two complementary dimensions: (1) stereotype bias, quantified by measuring variations in aesthetic evaluations across demographic groups; and (2) alignment between model outputs and genuine human aesthetic preferences.Our benchmark covers three subtasks (Aesthetic Perception, Assessment, Empathy) and introduces structured metrics (IFD, NRD, AAS) to assess both bias and alignment.We evaluate 19 MLLMs, including proprietary models (e.g., GPT-4o, Claude-3.5-Sonnet)and open-source models (e.g., InternVL-2.5, Qwen2.5-VL).Results indicate that smaller models exhibit stronger stereotype biases, whereas larger models align more closely with human preferences.Incorporating identity information often exacerbates bias, particularly in emotional judgments.These findings underscore the importance of identity-aware evaluation frameworks in subjective vision-language tasks. Kun Li 0015, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao |
EMNLP | 4 |
| 2025 | Comprehensive regional guidance for attention map semantics in text-to-image diffusion models
Haoxuan Wu, Lai-Man Po, Xuyuan Xu, Kun Li 0015 |
Comput. Vis. Image Underst. | 3 |
| 2025 | RMP-adapter: A region-based Multiple Prompt Adapter for multi-concept customization in text-to-image diffusion modelabstractThis paper introduces a novel framework for multi-concept customization in text-to-image diffusion models . At its core is a Multiple Prompt Adapter (MP-Adapter) capable of processing multiple image prompts in parallel, extracting features from target concepts and projecting them into the same latent space as the text prompt. This enables simultaneous handling of multiple concepts using just one reference image per concept. To address challenges in fusing multiple concepts with complex interactions, we propose a Region-based Denoising Framework (RDF) that dynamically generates concept-specific regions of interest during inference, allowing spatially decoupled injection of concept features. By integrating the MP-Adapter and RDF, our end-to-end pipeline enables multi-concept customization with intricate occlusions and interactions while preserving concept identities. This approach surpasses current methods by resolving concept conflicts, identity degradation, and occlusion issues, allowing flexible customization without concept-specific retraining. Both qualitative and quantitative evaluations demonstrate that our framework outperforms state-of-the-art approaches in multi-concept customization tasks, while ablation studies validate the effectiveness of each proposed component. This work significantly advances text-to-image generation capabilities for complex, user-defined concept combinations. Code and models will be released at https://github.com/baojudezeze/RMP-Adapter . Lai-Man Po, Xuyuan Xu, Yexin Wang, Haoxuan Wu, Kun Li 0015 |
Expert Syst. Appl. | 3 |
| 2024 | Self-Calibration Flow Guided Denoising Diffusion Model for Human Pose TransferabstractThe human pose transfer task aims to generate synthetic person images that preserve the style of reference images while accurately aligning them with the desired target pose. However, existing methods based on generative adversarial networks (GANs) struggle to produce realistic details and often face spatial misalignment issues. On the other hand, methods relying on denoising diffusion models require a large number of model parameters, resulting in slower convergence rates. To address these challenges, we propose a self-calibration flow-guided module (SCFM) to establish precise spatial correspondence between reference images and target poses. This module facilitates the denoising diffusion model in predicting the noise at each denoising step more effectively. Additionally, we introduce a multi-scale feature fusing module (MSFF) that enhances the denoising U-Net architecture through a cross-attention mechanism, achieving better performance with a reduced parameter count. Our proposed model outperforms state-of-the-art methods on the DeepFashion and Market-1501 datasets in terms of both the quantity and quality of the synthesized images. Our code is publicly available at https://github.com/zylwithxy/SCFM-guided-DDPM. Lai-Man Po, Wing Yin Yu, Haoxuan Wu, Xuyuan Xu, Kun Li 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Darwinian Model Upgrades: Model Evolving with Selective CompatibilityabstractThe traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensive and time-consuming considering billions of instances in industrial applications. BCT presents the first step towards backward-compatible model upgrades to get rid of backfilling. It is workable but leaves the new model in a dilemma between new feature discriminativeness and new-to-old compatibility due to the undifferentiated compatibility constraints. In this work, we propose Darwinian Model Upgrades (DMU), which disentangle the inheritance and variation in the model evolving with selective backward compatibility and forward adaptation, respectively. The old-to-new heritable knowledge is measured by old feature discriminativeness, and the gallery features, especially those of poor quality, are evolved in a lightweight manner to become more adaptive in the new latent space. We demonstrate the superiority of DMU through comprehensive experiments on large-scale landmark retrieval and face recognition benchmarks. DMU effectively alleviates the new-to-new degradation at the same time improving new-to-old compatibility, rendering a more proper model upgrading paradigm in large-scale retrieval systems.Code: https://github.com/TencentARC/OpenCompatible. Binjie Zhang, Shupeng Su, Yixiao Ge, Xuyuan Xu, Yexin Wang, Chun Yuan 0003, Zheng Shou 0001, Ying Shan |
AAAI | 4 |
| 2023 | Binary Embedding-based Retrieval at TencentabstractLarge-scale embedding-based retrieval (EBR) is the cornerstone of search-related industrial applications. Given a user query, the system of EBR aims to identify relevant information from a large corpus of documents that may be tens or hundreds of billions in size. The storage and computation turn out to be expensive and inefficient with massive documents and high concurrent queries, making it difficult to further scale up. Yukang Gan, Yixiao Ge, Chang Zhou 0008, Shupeng Su, Zhouchuan Xu, Xuyuan Xu, Quanchao Hui, Yexin Wang, Ying Shan |
KDD | 6 |
| 2022 | Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video RepresentationabstractSpatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignoring the intermediate state of the learned representations, which limits the overall performance. In this work, taking into account the degree of similarity of sampled instances as the intermediate state, we propose a novel pretext task - spatio-temporal overlap rate (STOR) prediction. It stems from the observation that humans are capable of discriminating the overlap rates of videos in space and time. This task encourages the model to discriminate the STOR of two generated samples to learn the representations. Moreover, we employ a joint optimization combining pretext tasks with contrastive learning to further enhance the spatio-temporal representation learning. We also study the mutual influence of each component in the proposed scheme. Extensive experiments demonstrate that our proposed STOR task can favor both contrastive learning and pretext tasks and the joint optimization scheme can significantly improve the spatio-temporal representation in video understanding. The code is available at https://github.com/Katou2/CSTP. Yujia Zhang 0002, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing Yin Yu |
AAAI | 3 |
| 2022 | Dynamic Token Normalization improves Vision Transformers
Wenqi Shao, Yixiao Ge, Zhaoyang Zhang 0004, Xuyuan Xu, Xiaogang Wang 0001, Ying Shan, Ping Luo 0002 |
ICLR | 4 |
| 2022 | Hot-Refresh Model Upgrades with Regression-Free Compatible Training in Image Retrieval
Binjie Zhang, Yixiao Ge, Yantao Shen 0003, Yu Li 0003, Chun Yuan 0003, Xuyuan Xu, Yexin Wang, Ying Shan |
ICLR | 6 |
| 2022 | Towards Universal Backward-Compatible Representation LearningabstractConventional model upgrades for visual search systems require offline refresh of gallery features by feeding gallery images into new models (dubbed as “backfill”), which is time-consuming and expensive, especially in large-scale applications. The task of backward-compatible representation learning is therefore introduced to support backfill-free model upgrades, where the new query features are interoperable with the old gallery features. Despite the success, previous works only investigated a close-set training scenario (i.e., the new training set shares the same classes as the old one), and are limited by more realistic and challenging open-set scenarios. To this end, we first introduce a new problem of universal backward-compatible representation learning, covering all possible data split in model upgrades. We further propose a simple yet effective method, dubbed as Universal Backward-Compatible Training (UniBCT) with a novel structural prototype refinement algorithm, to learn compatible representations in all kinds of model upgrading benchmarks in a unified manner. Comprehensive experiments on the large-scale face recognition datasets MS1Mv3 and IJB-C fully demonstrate the effectiveness of our method. Source code is available at https://github.com/TencentARC/OpenCompatible. Binjie Zhang, Yixiao Ge, Yantao Shen 0003, Shupeng Su, Fanzi Wu, Chun Yuan 0003, Xuyuan Xu, Yexin Wang, Ying Shan |
IJCAI | 7 |
| 2019 | Video copy detection by conducting fast searching of inverted files
Mengyang Liu, Lai-Man Po, Yasar Abbas Ur Rehman, Xuyuan Xu, Litong Feng |
Multim. Tools Appl. | 4 |
| 2019 | A Novel Patch Variance Biased Convolutional Neural Network for No-Reference Image Quality AssessmentabstractDeep convolutional neural networks (CNNs) have been successfully applied on no-reference image quality assessment (NR-IQA) with respect to human perception. Most of these methods deal with small image patches and use the average score of the test patches for predicting the whole image quality. We discovered that image patches from homogenous regions are unreliable for both neural network training and final image quality score estimation. In addition, image patches with complex structures have much higher chances of achieving better image quality prediction. Based on these findings, we enhanced the conventional CNN-based NR-IQA algorithm to avoid homogenous patches for the network training and quality score estimation. Moreover, we also use a variance-based weighting average to bias the final image quality score to the patches with complex structure. The experimental results show that this simple approach can achieve state-of-the-art performance compared with well-known NR-IQA algorithms. Lai-Man Po, Mengyang Liu, Wilson Y. F. Yuen, Xuyuan Xu, Chang Zhou 0008, Peter Hon-Wah Wong, Kin Wai Lau, Hon-Tung Luk |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Block-based adaptive ROI for remote photoplethysmography
Lai-Man Po, Litong Feng, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
Multim. Tools Appl. | 4 |
| 2016 | Face liveness detection and recognition using shearlet based feature descriptorsabstractFace recognition is a widely used biometric technology due to its convenience but it is vulnerable to spoofing attacks made by non-real faces such as a photograph or video of valid user. Face liveness detection is a core technology to make sure that the input face is a live person. However, this is still very challenging using conventional liveness detection approaches of texture analysis and motion detection. The aim of this paper is to develop a multifunctional feature descriptor and an efficient framework which can be used to deal with both face liveness detection and recognition. In this framework, new feature descriptors are defined using a multiscale directional transform (shearlet transform). Then, stacked autoencoders and softmax classifier are concatenated to detect face liveness and identify person. We evaluated this approach using CASIA Face Anti-Spoofing Database and the results show that our approach performs better than state-of-the-art techniques following the provided evaluation protocols of this database, and is possible to significantly enhance the security of face recognition biometric system. Lai-Man Po, Xuyuan Xu, Litong Feng |
ICASSP | 3 |
| 2016 | Integration of image quality and motion cues for face anti-spoofing: A neural network approach
Litong Feng, Lai-Man Po, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | No-Reference Video Quality Assessment With 3D Shearlet Transform and Convolutional Neural NetworksabstractIn this paper, we propose an efficient general-purpose no-reference (NR) video quality assessment (VQA) framework that is based on 3D shearlet transform and convolutional neural network (CNN). Taking video blocks as input, simple and efficient primary spatiotemporal features are extracted by 3D shearlet transform, which are capable of capturing natural scene statistics properties. Then, CNN and logistic regression are concatenated to exaggerate the discriminative parts of the primary features and predict a perceptual quality score. The resulting algorithm, which we name shearlet- and CNN-based NR VQA (SACONVA), is tested on well-known VQA databases of Laboratory for Image & Video Engineering, Image & Video Processing Laboratory, and CSIQ. The testing results have demonstrated that SACONVA performs well in predicting video quality and is competitive with current state-of-the-art full-reference VQA methods and general-purpose NR-VQA algorithms. Besides, SACONVA is extended to classify different video distortion types in these three databases and achieves excellent classification accuracy. In addition, we also demonstrate that SACONVA can be directly applied in real applications such as blind video denoising. Lai-Man Po, Chun-Ho Cheung, Xuyuan Xu, Litong Feng, Kwok-Wai Cheung 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Dynamic ROI based on K-means for remote photoplethysmographyabstractRemote imaging photoplethysmography (RIPPG) can achieve contactless human vital signs monitoring. Though the remote operation mode brings a great convenience for RIPPG applications, the RIPPG signal quality is limited by the remote nature. Improving the RIPPG signal quality becomes an essential task in the clinical application of RIPPG. Since the region of interest (ROI) of the RIPPG transforms from a point to an area, there is a new approach to improving the RIPPG signal quality through refining the ROI. In this paper, we propose a dynamic ROI for RIPPG, which can automatically select the skin regions corresponding to good quality RIPPG signals. First, a fixed ROI is divided into non-overlapped blocks. Then two features are proposed to perform no-reference quality assessment for RIPPG signals from different blocks. After that, K-means clustering operates in a two dimensional feature space. A dynamic ROI can be selected for a video segment based on the clustering result, updated every two seconds. Nineteen healthy subjects were enrolled to test the proposed ROI selection method on both the facial region and the palmar region. Experimental results of heart rate measurement show that the proposed dynamic ROI method for RIPPG can effectively improve the RIPPG signal quality, compared with the state-of-the-art ROI methods for RIPPG. Litong Feng, Lai-Man Po, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
ICASSP | 3 |
| 2015 | No-reference image quality assessment using shearlet transform and stacked autoencodersabstractIn this work, we describe an efficient generalpurpose no-reference (NR) image quality assessment (IQA) algorithm that is based on a new multiscale directional transform (shearlet transform) with a strong ability to localize distributed discontinuities. The algorithm relies on utilizing the sum of subband coefficient amplitudes (SSCA) as primary features to describe the behavior of natural images and distorted images. Then, stacked autoencoders are applied to exaggerate the discriminative parts of the primary features. Finally, by translating the NR-IQA problem into classification problem, the differences of evolved features are identified by softmax classifier. The resulting algorithm, which we name SESANIA (ShEarlet and Stacked Autoencoders based No-reference Image quality Assessment), is tested on several databases (LIVE, Multiply Distorted LIVE and TID2008) and shown to be suitable to many common distortions, consistent with subjective assessment and comparable to full-reference IQA methods and state-of-the-art general purpose NR-IQA algorithms. Lai-Man Po, Xuyuan Xu, Litong Feng, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
ISCAS | 3 |
| 2015 | Frame adaptive ROI for photoplethysmography signal extraction from fingertip video captured by smartphoneabstractPhotoplethysmography (PPG) has been widely used in clinical applications for monitoring vital signs especially heart rate by pulse oximeter. Recent researches have demonstrated the possibility of using fingertip video based PPG approach to estimate heart rate by smartphones. However, due to the variation of camera sensor characteristics in difference smartphones, the conventional fixed region-of-interest (ROI) for PPG signal extraction technique is not reliable. In this paper, a novel frame adaptive ROI method is proposed to detour the color saturation or cut-off distortion in the fingertip video capturing process for improving the reliability due to variation and limited dynamic range of the camera sensors in different smartphone models. Experimental results demonstrate that the proposed method can produce good pulsatile waveform and achieve high heart rate estimation accuracy using different smartphone models as compared with a FDA (U.S. Food and Drug Administration) approved commercial pulse oximeter. Lai-Man Po, Xuyuan Xu, Litong Feng, Kwok-Wai Cheung 0002, Chun-Ho Cheung |
ISCAS | 2 |
| 2015 | No-reference image quality assessment with shearlet transform and deep neural networks
Lai-Man Po, Xuyuan Xu, Litong Feng, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
Neurocomputing | 3 |
| 2015 | Motion-Resistant Remote Imaging Photoplethysmography Based on the Optical Properties of SkinabstractRemote imaging photoplethysmography (RIPPG) can achieve contactless monitoring of human vital signs. However, the robustness to a subject's motion is a challenging problem for RIPPG, especially in facial video-based RIPPG. The RIPPG signal originates from the radiant intensity variation of human skin with pulses of blood and motions can modulate the radiant intensity of the skin. Based on the optical properties of human skin, we build an optical RIPPG signal model in which the origins of the RIPPG signal and motion artifacts can be clearly described. The region of interest (ROI) of the skin is regarded as a Lambertian radiator and the effect of ROI tracking is analyzed from the perspective of radiometry. By considering a digital color camera as a simple spectrometer, we propose an adaptive color difference operation between the green and red channels to reduce motion artifacts. Based on the spectral characteristics of photoplethysmography signals, we propose an adaptive bandpass filter to remove residual motion artifacts of RIPPG. We also combine ROI selection on the subject's cheeks with speeded-up robust features points tracking to improve the RIPPG signal quality. Experimental results show that the proposed RIPPG can obtain greatly improved performance in accessing heart rates in moving subjects, compared with the state-of-the-art facial video-based RIPPG methods. Litong Feng, Lai-Man Po, Xuyuan Xu, Ruiyi Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Adaptive block truncation filter for MVC depth image enhancementabstractIn Multiview Video plus Depth (MVD) format, virtual views are generated from decoded texture videos with decoded depth images through Depth Image based Rendering (DIBR). 3DV-ATM is a reference model for H.264/AVC based Multiview Video Coding (MVC) and aims at achieving high coding efficiency for 3D video in MVD format. Depth images are first downsampled then coded by 3DV-ATM. However, sharp object boundary characteristic of depth images does not well match with the transform coding of 3DV-ATM. Depth boundaries are often blurred with ringing artifacts in the decoded depth images that result in noticeable artifacts in synthesized views. This paper presents a low complexity adaptive block truncation filter to recover the sharp object boundaries of depth images using adaptive block repositioning and expansion for increasing the depth values refinement accuracy. This new approach is very efficient and can avoid false depth boundary refinement when block boundaries lie around the depth edge regions and ensure sufficient information within the processing block for depth layers classification. Experimental results show that sharp depth edges can be recovered using the proposed filter and boundary artifacts in the synthesized views can be removed. The proposed method can provide improvement up to 3.25dB in the depth map enhancement and bitrate reduction of 3.06% in the synthesized views. Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Litong Feng, Kwok-Wai Cheung 0002, Chi-Wang Ting, Ka-Ho Ng |
ICASSP | 1 |
| 2014 | No-reference image quality assessment using statistical characterization in the shearlet domain
Lai-Man Po, Xuyuan Xu, Litong Feng |
Signal Process. Image Commun. | 3 |
| 2014 | Adaptive depth truncation filter for MVC based compressed depth image
Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Kwok-Wai Cheung 0002, Litong Feng, Chi-Wang Ting, Ka-Ho Ng |
Signal Process. Image Commun. | 1 |
| 2013 | Watershed based depth map misalignment correction and foreground biased dilation for DIBR view synthesisabstractThe quality of the synthesized views by Depth Image Based Rendering (DIBR) highly depends on the accuracy of the depth map, especially the alignment of object boundaries of texture image. In practice, the misalignment of sharp depth map edges is the major cause of the annoying artifacts at the disoccluded regions of the synthesized views. In this paper, a new depth map preprocessing method using Watershed misalignment correction and dilation filter is proposed to align the foreground depth edges to cover the whole transitional color edge regions. This approach can handle the sharp depth map edges lying inside or outside the object boundaries in 2D sense. The quality of the disoccluded regions of the synthesized views can be significantly improved. Experimental results show that the proposed method achieves superior performance for view synthesis by DIBR especially for generating large baseline virtual views. Xuyuan Xu, Lai-Man Po, Kwok-Wai Cheung 0002, Litong Feng, Chun-Ho Cheung |
ICIP | 1 |
| 2013 | An adaptive background biased depth map hole-filling method for KinectabstractThe launch of Kinect provides a convenient way to access the depth information in real time. However the depth map quality still needs to be enhanced for 3D visual applications. In this paper, an adaptive background biased depth map hole-filling method is proposed. First, depth holes caused by abnormal reflection are filled by color similarity in-painting, and a soft decision for color similarity checking is performed by the use of probabilities in random walks color segmentation. Afterwards it is assumed that the lost information in the rest of depth holes belongs to the background. The background depth information is extracted by automatic thresholding in the neighborhood of each hole. Depth holes are in-painted with the background information in their local neighborhood. Combination of color similarity in-painting and background biased in-painting is able to perform depth map hole-filling adaptively for different kinds of depth holes for Kinect. The hole-filling results and virtual view synthesis results show that the Kinect depth map quality can be improved significantly by the proposed method. Litong Feng, Lai-Man Po, Xuyuan Xu, Ka-Ho Ng, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
IECON | 3 |
| 2013 | Depth-aided exemplar-based hole filling for DIBR view synthesisabstractQuality of synthesized view by Depth-Image-Based Rendering (DIBR) highly depends on hole filling, especially for synthesized view with large disocclusion. Many hole filling methods are proposed to improve the synthesized view quality and inpainting is the most popular approach to recover the disocclusions. However, the conventional inpainting either makes the hole regions blurred via diffusion or propagates the foreground information to the disoclusion regions. Annoying artifacts are created in the synthesized virtual views. This paper proposes a depth-aided exemplar-based inpainting method for recovering large disoclusion. It consists of two processes, warped depth map filling and warped color image filling. Since depth map can be considered as a grey-scale image without texture, it is much easier to be filled. Disoccluded regions of color image are predicted based on its associated filled depth map information. Regions with texture lying around the background have higher priority to be filled than other regions and disoccluded regions are filled by propagating the background texture through the exemplar-based inpainting. Thus artifacts created by diffusion or using foreground information for prediction can be eliminated. Experimental results show texture can be recovered in large disocclusions and the proposed method has better visual quality compared to existing methods. Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Litong Feng, Ka-Ho Ng, Kwok-Wai Cheung 0002 |
ISCAS | 1 |
| 2013 | Depth map misalignment correction and dilation for DIBR view synthesis
Xuyuan Xu, Lai-Man Po, Ka-Ho Ng, Litong Feng, Kwok-Wai Cheung 0002, Chun-Ho Cheung, Chi-Wang Ting |
Signal Process. Image Commun. | 1 |
| 2012 | A new motion compensation method using superimposed inter-frame signalsabstractA new MCP method called Neighbor Predicted Superimposed Search (NPSS) algorithm that uses superimposed inter-frame signals to achieve higher prediction accuracy is proposed in this paper. It outperforms other Multi-Hypothesis MCP (MHMCP) methods as it does not require the transmission of multiple motion vectors. The proposed method has better prediction quality and yet having comparable computational complexity as conventional block-based MCP with no extra side-information overhead. Ka-Ho Ng, Lai-Man Po, Kwok-Wai Cheung 0002, Xuyuan Xu, Ka-Man Wong |
ICASSP | 4 |
| 2012 | A foreground biased depth map refinement method for DIBR view synthesisabstractThe performance of view synthesis using depth image based rendering (DIBR) highly depends on the accuracy of depth map. Inaccurate boundary alignment between texture image and depth map especially for large depth discontinuities always cause annoying artifacts in disocclusion regions of the synthesized view. Pre-filtering approach and reliability-based approach have been proposed to tackle this problem. However, pre-filtering approach blurs the depth map with drawback of degradation of the depth map and may also cause distortion in non-hole region. Reliability-based approach uses reliable warping information from other views to fill up holes and is not suitable for the view synthesis with single texture video such as video-plus-depth based DIBR applications. This paper presents a simple and efficient depth map preprocessing method with use of texture edge information to refine depth pixels around the large depth discontinuities. The refined depth map can make the whole texture edge pixels assigned with foreground depth values. It can significantly improve the quality of the synthesized view by avoiding incorrect use of foreground texture information in hole filling. The experimental results show the proposed method achieves superior performance for view synthesis by DIBR especially for large baseline. Xuyuan Xu, Lai-Man Po, Kwok-Wai Cheung 0002, Ka-Ho Ng, Ka-Man Wong, Chi-Wang Ting |
ICASSP | 1 |
| 2012 | Horizontal Scaling and Shearing-Based Disparity-Compensated Prediction for Stereo Video CodingabstractIn multiview video coding (MVC), disparity-compensated prediction (DCP) exploits the correlation among different views. A common approach is to use block-based motion-compensated prediction (MCP) tools to predict the disparity effect among different views. However, some regions in different views may have various deformations due to nonconstant depth. Thus, performance of DCP is not satisfactory with the simple translational model assumed in conventional block-based MCP tools. Previous attempts to achieve better disparity prediction were usually too complex for practical use. In this paper, horizontal scaling and shearing (HSS) effects are investigated to increase interview prediction accuracy for stereo video. HSS deformations are common among images of horizontally aligned views, due to horizontal and vertical flat surfaces that are not parallel with projection image planes. To achieve HSS-based DCP with minimal complexity, an efficient subsampled block-matching technique is adopted and integrated into MVC extension of H.264/AVC in stereo profile. Affine parameters estimation and additional frame buffers are not required and the overall increase of computational complexity and memory requirements are moderate. Experimental results show that the new technique can achieve up to 5.25% bitrate reduction in interview prediction using JM17.0 reference software implementation. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002, Chi-Wang Ting, Ka-Ho Ng, Xuyuan Xu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2011 | Stretching, compression and shearing disparity compensated prediction techniques for stereo and multiview video codingabstractIn multiview video coding, disparity compensated prediction exploits the correlation among different views. A common approach is to use the conventional motion compensated prediction to predict disparity effect among different views. However, the same object in different views usually has deformation of different extents and, thus, accurate disparity prediction cannot be achieved with such simple translational motion model. Previous attempts to achieve more accurate disparity prediction are usually too complex for practical implementation. In this paper, stretching, compression and shearing (SCSH) effects are investigated to better model the disparity effect in disparity compensated prediction. To achieve SCSH effects with minimal computation, an efficient disparity compensated prediction using subsampled block-matching technique is proposed. No affine parameters estimation or additional frame buffers is required and the overall increase in memory requirement and computational complexity is moderate. Experimental results show that the new technique can achieve up to 4.84% bitrate reduction in inter-view prediction using JM17.0 reference software implementation. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002, Ka-Ho Ng, Xuyuan Xu |
ICASSP | 5 |
| 2011 | A new multidirectional extrapolation hole-filling method for Depth-Image-Based RenderingabstractDepth-Image-Based Rendering (DIBR) is widely used to generate virtual view of a scene from a known view with associated depth map in 3D video applications. However, disocclusion arises in image warping of DIBR. Many hole-filling methods have been proposed such as constant color, horizontal interpolation, horizontal extrapolation, and variational inpainting, but they cause different types of annoying artifact for large holes with complex texture background. In this paper, a novel multidirectional extrapolation hole-filling method is proposed to enhance visual quality for large hole-filling with complex texture background. The proposed method uses neighbor pixels' texture features to estimate hole-filling direction in a pixel-by-pixel manner. Experimental results demonstrated that the proposed method could provide better visual quality compared with conventional methods for virtual views synthesis with high-quality depth map. Lai-Man Po, Shihang Zhang, Xuyuan Xu, Yuesheng Zhu |
ICIP | 3 |