EDBT 2026 Demo / reviewers in the wild / expert
Qifei Wang
dblp:73/8011
· DBLP profile ↗
29ranked-venue papers
12as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table UnderstandingabstractYuhang Zhou, Mingrui Zhang, Ke Li, Mingyi Wang, Qiao Liu, Qifei Wang, Jiayi Liu, Fei Liu, Serena Li, Weiwei LI, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Zhuokai Zhao, Lizhu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qifei Wang, Serena Li, Weiwei Li 0006, Xiangjun Fan, Zhuokai Zhao, Lizhu Zhang |
ACL (1) | 6 |
| 2026 | Knowledge Distillation-Driven Semantic NOMA for Image Transmission With Diffusion ModelabstractAs a promising 6G enabler beyond conventional bit-level transmission, semantic communication can considerably reduce required bandwidth resources, while its combination with multiple access requires further exploration. This paper proposes a knowledge distillation-driven and diffusion-enhanced (KDD) semantic non-orthogonal multiple access (NOMA), named KDD-SemNOMA, for multi-user uplink wireless image transmission. Specifically, to ensure robust feature transmission across diverse transmission conditions, we firstly develop a ConvNeXt-based deep joint source and channel coding architecture with enhanced adaptive feature module. This module incorporates signal-to-noise ratio and channel state information to dynamically adapt to additive white Gaussian noise and Rayleigh fading channels. Furthermore, to improve image restoration quality without inference overhead, we introduce a two-stage knowledge distillation strategy, i.e., a teacher model, trained on interference-free orthogonal transmission, guides a student model via feature affinity distillation and cross-head prediction distillation. Moreover, a diffusion model-based refinement stage leverages generative priors to transform initial SemNOMA outputs into high-fidelity images with enhanced perceptual quality. Extensive experiments on CIFAR-10 and FFHQ-256 datasets demonstrate superior performance over state-of-the-art methods, delivering satisfactory reconstruction performance even at extremely poor channel conditions. These results highlight the advantages in both pixel-level accuracy and perceptual metrics, effectively mitigating interference and enabling high-quality image recovery. Qifei Wang, Zhen Gao 0001, Shuo Sun 0001, Zhijin Qin, Xiaodong Xu 0001, Meixia Tao |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | CrossMed-SAM: Cross-Modal Medical Image Segmentation via Frequency-Seale-Semantic AwarenessabstractAlthough vision foundation models such as SAM excel at natural image segmentation, their transfer to cross-modal medical segmentation remains difficult due to frequency mismatches, extreme scale variability, and modality-specific se-mantics. Existing adaptations typically require heavy fine-tuning and bespoke modules, which undermine generalization across diverse anatomies and imaging protocols. To address these chal-lenges, we propose CrossMed-SAM, which integrates three syner-gistic modules. The Cross-Frequency Attention Module (CFAM) leverages discrete wavelet transform to handle frequency domain variations across imaging modalities, while the Multi-Scale At-tention Module (MSAM) employs parallel dilated convolutions to capture extreme scale variations from microscopic lesions to large anatomical structures. The Adaptive Semantic Fusion Mod-ule (ASFM) integrates CLIP-based medical knowledge through dynamic gating mechanisms to provide semantic guidance when visual features are ambiguous. Extensive experiments on twelve datasets across five diverse medical imaging modalities demon-strate that CrossMed-SAM significantly outperforms existing state-of-the-art methods, achieving Dice coefficient improvements of 1.7% to 5.3% over the strong baseline MedSAM, with superior generalization capabilities across unseen datasets. Qifei Wang, Yuefeng Zhao, Nai Zhou, Nannan Hu |
BIBM | 1 |
| 2025 | Calibrated Multi-Preference Optimization for Aligning Diffusion ModelsabstractAligning text-to-image (T2I) diffusion models with preference optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting the rich information, as they only consider pairwise preference distribution. Furthermore, they lack generalization to multi-preference scenarios and struggle to handle inconsistencies between rewards. To address this, we present Calibrated Preference Optimization (CaPO), a novel method to align T2I diffusion models by incorporating the general preference from multiple reward models without human annotated data. The core of our approach involves a reward calibration method to approximate the general preference by computing the expected win-rate against the samples generated by the pretrained models. Additionally, we propose a frontier-based pair selection method that effectively manages the multi-preference distribution by selecting pairs from Pareto frontiers. Finally, we use regression loss to fine-tune diffusion models to match the difference between calibrated rewards of a selected pair. Experimental results show that CaPO consistently outperforms prior methods, such as Direct Preference Optimization (DPO), in both single and multi-reward settings validated by evaluation on T2I benchmarks, including GenEval and T2I-Compbench. Kyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He, Junjie Ke, Ming-Hsuan Yang 0001, Irfan A. Essa, Jinwoo Shin, Feng Yang 0008, Yinxiao Li |
CVPR | 3 |
| 2025 | View-aware Decomposition and Unification for Fast Ground-to-Aerial Person SearchabstractGround-to-aerial person search leverages cooperative efforts between unmanned aerial vehicles (UAV) and ground surveillance cameras to locate person individuals. Despite the progress made by recent works, the impact of the discrepancy between the two views is underestimated. This limits the overall person search performance when training the model in a view-agnostic way. To address this, we propose a view-aware decomposition and unification (VADU) framework for ground-to-aerial person search. Specifically, we decompose the person search model to learn view-oriented modules for image feature encoding and person proposal generation. The data sampling and retrieval feature learning are also composed to cope with the decomposed model. This decomposition improves both person detection and discriminative feature learning within each view. On top of the decomposition, we propose view-aware unification to produce unified cross-view person features. Cross-view prototypical contrastive learning is introduced to enhance the unification between different views, enhancing model robustness to retrieve a target person in cameras of a different view. As the decomposed parts of the model are deployed on different devices for inference, this overall framework adds no extra computation cost in real-world applications. Extensive experiments demonstrate that the proposed method achieves superior person search performance and guarantees the efficiency of inference. The source code is available at https://github.com/QFWang-11/vadu. Qifei Wang, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Yongsheng Gao 0001 |
IROS | 1 |
| 2025 | Adaptive Mixture of Experts for Cross-Domain Medical Image Segmentation with Vision Foundation Models
Qifei Wang, Yuefeng Zhao, Nai Zhou, Qianqian Tao, Nannan Hu |
PRCV (18) | 1 |
| 2025 | EHPR: Learning evolutionary hierarchy perception representation based on quaternion for temporal knowledge graph completion
Jiujiang Guo, Mankun Zhao, Jian Yu 0003, Jianhang Song, Qifei Wang, Linying Xu, Mei Yu 0004 |
Inf. Sci. | 6 |
| 2024 | PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion ModelsabstractReward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize rewards that reflect human preference. However, in the vision domain, existing RL-based reward finetuning methods are limited by their instability in large-scale training, rendering them incapable of generalizing to complex, unseen prompts. In this paper, we propose Proximal Reward Difference Prediction (PRDP), enabling stable black-box reward finetuning for diffusion models for the first time on large-scale prompt datasets with over 100K prompts. Our key innovation is the Reward Difference Prediction (RDP) objective that has the same optimal solution as the RL objective while enjoying better training stability. Specifically, the RDP objective is a supervised regression objective that tasks the diffusion model with predicting the reward difference of generated image pairs from their denoising trajectories. We theoretically prove that the diffusion model that obtains perfect reward difference prediction is exactly the maximizer of the RL objective. We further develop an online algorithm with proximal updates to stably optimize the RDP objective. In experiments, we demonstrate that PRDP can match the reward maximization ability of well-established RL-based methods in small-scale training. Furthermore, through large-scale training on text prompts from the Human Preference Dataset v2 and the Pick-a-Pic v1 dataset, PRDP achieves superior generation quality on a diverse set of complex, unseen prompts whereas RL-based methods completely fail. Fei Deng 0001, Qifei Wang, Tingbo Hou, Matthias Grundmann 0002 |
CVPR | 2 |
| 2024 | Parrot: Pareto-Optimal Multi-reward Reinforcement Learning Framework for Text-to-Image Generation
Yinxiao Li, Junjie Ke, Innfarn Yoo, Han Zhang 0010, Qifei Wang, Fei Deng 0001, Glenn Entis, Junfeng He, Gang Li 0021, Sangpil Kim, Irfan A. Essa, Feng Yang 0008 |
ECCV (38) | 7 |
| 2024 | Optical Diffusion Models for Image GenerationabstractDiffusion models generate new samples by progressively decreasing the noise from the initially provided random distribution. This inference procedure generally utilizes a trained neural network numerous times to obtain the final output, creating significant latency and energy consumption on digital electronic hardware such as GPUs. In this study, we demonstrate that the propagation of a light beam through a transparent medium can be programmed to implement a denoising diffusion model on image samples. This framework projects noisy image patterns through passive diffractive optical layers, which collectively only transmit the predicted noise term in the image. The optical transparent layers, which are trained with an online training approach, backpropagating the error to the analytical model of the system, are passive and kept the same across different steps of denoising. Hence this method enables high-speed image generation with minimal power consumption, benefiting from the bandwidth and energy efficiency of optical information processing. Ilker Oguz, Niyazi Ulas Dinç, Mustafa Yildirim, Junjie Ke, Innfarn Yoo, Qifei Wang, Christophe Moser, Demetri Psaltis |
NeurIPS | 6 |
| 2024 | Chinese image captioning with fusion encoder and visual keyword searchabstractAbstract Automatic generation of image captions is essentially a cross‐modal conversion from image to text. Owing to the differences in linguistic characteristics between Chinese and English, quite a few Chinese image captioning methods have recently been proposed. Nevertheless, the existing Chinese image captioning models usually lack attention to local details of images or tend to produce general descriptions. To address these challenges, a Chinese image captioning method is proposed that incorporates fusion encoder, visual keyword search, and reinforcement learning. The fusion encoder can simultaneously extract local and global features of the input image to enrich the semantic information in the decoding stage, visual keyword search can pursue potential visual words associated with the image content, and the reinforcement learning mechanism can optimize the evaluation metric CIDEr at sentence level to promote the lexical diversity of image description. The results of extensive experiments demonstrate that the proposed model outperforms the state‐of‐the‐art models and delivers expressive and informative Chinese image captions. Yang Zou 0001, Shiyu Liao, Qifei Wang |
IET Image Process. | 3 |
| 2023 | Distant Supervision Relation Extraction with Improved PCNN and Multi-level Attention
Yang Zou 0001, Qifei Wang, Xiaoqin Zeng |
KSEM (1) | 2 |
| 2021 | Adversarially Adaptive Normalization for Single Domain GeneralizationabstractSingle domain generalization aims to learn a model that performs well on many unseen domains with only one domain data for training. Existing works focus on studying the adversarial domain augmentation (ADA) to improve the model’s generalization capability. The impact on domain generalization of the statistics of normalization layers is still underinvestigated. In this paper, we propose a generic normalization approach, adaptive standardization and rescaling normalization (ASR-Norm), to complement the missing part in previous works. ASR-Norm learns both the standardization and rescaling statistics via neural networks. This new form of normalization can be viewed as a generic form of the traditional normalizations. When trained with ADA, the statistics in ASR-Norm are learned to be adaptive to the data coming from different domains, and hence improves the model generalization performance across domains, especially on the target domain with large discrepancy from the source domain. The experimental results show that ASR-Norm can bring consistent improvement to the state-of-the-art ADA approaches by 1.6%, 2.7%, and 6.3% averagely on the Digits, CIFAR-10-C, and PACS benchmarks, respectively. As a generic tool, the improvement introduced by ASR-Norm is agnostic to the choice of ADA methods. Xinjie Fan, Qifei Wang, Junjie Ke, Feng Yang 0008, Boqing Gong, Mingyuan Zhou |
CVPR | 2 |
| 2021 | MUSIQ: Multi-scale Image Quality TransformerabstractImage quality assessment (IQA) is an important research topic for understanding and improving visual experience. The current state-of-the-art IQA methods are based on convolutional neural networks (CNNs). The performance of CNN-based models is often compromised by the fixed shape constraint in batch training. To accommodate this, the input images are usually resized and cropped to a fixed shape, causing image quality degradation. To address this, we design a multi-scale image quality Transformer (MUSIQ) to process native resolution images with varying sizes and aspect ratios. With a multi-scale image representation, our proposed method can capture image quality at different granularities. Furthermore, a novel hash-based 2D spatial embedding and a scale embedding is proposed to support the positional embedding in the multi-scale representation. Experimental results verify that our method can achieve state-of-the-art performance on multiple large scale IQA datasets such as PaQ-2-PiQ [41], SPAQ [11], and KonIQ-10k [16].1 Junjie Ke, Qifei Wang, Yilin Wang 0001, Peyman Milanfar, Feng Yang 0008 |
ICCV | 2 |
| 2021 | Multi-path Neural Networks for On-device Multi-domain Visual ClassificationabstractLearning multiple domains/tasks with a single model is important for improving data efficiency and lowering inference cost for numerous vision tasks, especially on resource-constrained mobile devices. However, hand-crafting a multi-domain/task model can be both tedious and challenging. This paper proposes a novel approach to automatically learn a multi-path network for multi-domain visual classification on mobile devices. The proposed multi-path network is learned from neural architecture search by applying one reinforcement learning controller for each domain to select the best path in the super-network created from a MobileNetV3-like search space. An adaptive balanced domain prioritization algorithm is proposed to balance optimizing the joint model on multiple domains simultaneously. The determined multi-path model selectively shares parameters across domains in shared nodes while keeping domain-specific parameters within non-shared nodes in individual domain paths. This approach effectively reduces the total number of parameters and FLOPS, encouraging positive knowledge transfer while mitigating negative interference across domains. Extensive evaluations on the Visual Decathlon dataset demonstrate that the proposed multi-path model achieves state-of-the-art performance in terms of accuracy, model size, and FLOPS against other approaches using MobileNetV3-like architectures. Furthermore, the proposed method improves average accuracy over learning single-domain models individually, and reduces the total number of parameters and FLOPS by 78% and 32% respectively, compared to the approach that simply bundles single-domain models for multi-domain learning. Qifei Wang, Junjie Ke, Joshua Greaves, Grace Chu, Gabriel Bender, Luciano Sbaiz, Alec Go, Andrew G. Howard, Ming-Hsuan Yang 0001, Jeff Gilbert, Peyman Milanfar, Feng Yang 0008 |
WACV | 1 |
| 2020 | Learnable Cost Volume Using the Cayley Representation
Taihong Xiao, Jinwei Yuan, Deqing Sun, Qifei Wang, Xinyu Zhang 0023, Ming-Hsuan Yang 0001 |
ECCV (9) | 4 |
| 2020 | Multi-data UAV Images for Large Scale Reconstruction of Buildings
Menghan Zhang, Yunbo Rao, Jiansu Pu, Xun Luo, Qifei Wang |
MMM (2) | 5 |
| 2018 | Roads Detection of Aerial Image with FCN-CRF ModelabstractThis paper describes a deep learning based model for roads detection in Aerial image. In general, standard CNN networks would have less ability for tiny objects detection in remote sensing image. With this regard, we propose a novel fully convolutional network, which utilizes deconvolution layers and feature map fussing to take as input intensity and pixel-wise labeling. Moreover, the class prediction are used as the input to Condition Random Field (CRF) for the final pixel prediction. The Batch Normalization (BN) algorithm and two stages training strategy were used in our model to reduce the time cost of model training. Several experimental results conducted in Massachuseets. Road dataset demonstrate the superiority of our model with respect to accuracy and time cost. Yunbo Rao, Wei Liu 0073, Jiansu Pu, Qifei Wang |
VCIP | 5 |
| 2018 | Visual Analysis of Human Motion: A Survey on Recent Advances and ApplicationsabstractThis paper summarizes the recent progress in human motion analysis and its applications. The first part of this paper reviews the motion capture systems and the representations of human's motion data. Next, the paper sketches the advanced human motion data processing technologies, including motion data filtering, temporal alignment, and segmentation. The following parts overview the state-of-the-art approaches of action recognition and dynamics measuring. The last part discusses the emerging applications of human motion analysis in healthcare and human robot interaction. The promising research topics of human motion analysis in the future are also summarized in the last part. Qifei Wang, Yunbo Rao |
VCIP | 1 |
| 2017 | Anterior cruciate ligament reconstruction model based on anatomical position locating
Yunbo Rao, Xianshu Ding, Jianping Gou, Qifei Wang |
Multim. Tools Appl. | 5 |
| 2016 | Online distribution and interaction of video data in social multimedia network
Xiangyang Ji, Qifei Wang, Bo-Wei Chen, Seungmin Rho, C.-C. Jay Kuo, Qionghai Dai |
Multim. Tools Appl. | 2 |
| 2015 | Unsupervised Temporal Segmentation of Repetitive Human Actions Based on Kinematic Modeling and Frequency AnalysisabstractIn this paper, we propose a method for temporal segmentation of human repetitive actions based on frequency analysis of kinematic parameters, zero-velocity crossing detection, and adaptive k-means clustering. Since the human motion data may be captured with different modalities which have different temporal sampling rate and accuracy (e.g., Optical motion capture systems vs. Microsoft Kinect), we first apply a generic full-body kinematic model with an unscented Kalman filter to convert the motion data into a unified representation that is robust to noise. Furthermore, we extract the most representative kinematic parameters via the primary frequency analysis. The sequences are segmented based on zero-velocity crossing of the selected parameters followed by an adaptive k-means clustering to identify the repetition segments. Experimental results demonstrate that for the motion data captured by both the motion capture system and the Microsoft Kinect, our proposed algorithm obtains robust segmentation of repetitive action sequences. Qifei Wang, Gregorij Kurillo, Ferda Ofli, Ruzena Bajcsy |
3DV | 1 |
| 2014 | Content adaptive screen image scalingabstractThis paper proposes an efficient content adaptive screen image scaling scheme for the real-time screen applications like remote desktop and screen sharing. In the proposed screen scaling scheme, a screen content classification step is first introduced to classify the screen image into text and pictorial regions. Afterward, we propose an adaptive shift linear interpolation algorithm to predict the new pixel values with the shift offset adapted to the content type of each pixel. The shift offset for each screen content type is offline optimized by minimizing the theoretical interpolation error based on the training samples respectively. The proposed content adaptive screen image scaling scheme can achieve good visual quality and also keep the low complexity for realtime applications. Yao Zhai, Qifei Wang, Yan Lu 0001, Shipeng Li 0001 |
ICIP | 2 |
| 2013 | Complexity Reduction and Performance Improvement for Geometry Partitioning in Video CodingabstractGeometry partitioning for video coding involves establishing a partition line boundary within each block-shaped region and applying motion-compensated prediction to the two sub-regions created by the partition line. This paper presents techniques for enhancing the effectiveness and reducing the complexity of geometry partitioning schemes. A texture-difference-based approach is described to simplify the process of selecting the partition lines. Applying this approach together with a described skipping strategy for blocks with uniform texture can achieve a 94% reduction of encoding time while retaining a similar rate-distortion (R-D) performance to the full-search partitioning approach, when implemented for wedge-based geometry partitioning (WGP) in the context of H.264/MPEG-4 AVC JM 16.2. A bit rate improvement of approximately 6% is shown relative to not using geometry partitioning. For further R-D improvement, we describe a background-compensated prediction scheme to reduce the number of overhead bits used for motion vectors. Additionally, for systems in which high-quality depth maps are available, we incorporate depth map usage into the described approaches to generate a more accurate partitioning. Using these approaches with object-boundary-based geometry partitioning can achieve about 9% bit rate savings relative to using WGP, while keeping a similar computational complexity to the described complexity-reduced WGP. Qifei Wang, Xiangyang Ji, Ming-Ting Sun, Gary J. Sullivan, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | 3D spatial reconstruction and communication from vision fieldabstractVision field describes the real world visual information by summarizing the seven-dimensional plenoptic function into three domains: view, light, time, from which a better understanding of previous 3D capture and reconstruction systems can be provided. In this paper, we first show how to reconstruct 3D spatial information from all the three attributes of the vision field, namely full-space vision field reconstruction. Then, based on Laplacian iterative geometry prediction, a 3D mesh coding algorithm with cascaded quantization is presented to facilitate the communication of the reconstructed 3D models from vision field. At last, experimental results of both the 3D spatial reconstruction and the 3D mesh coding are demonstrated. Xun Cao, Qifei Wang, Xiangyang Ji, Qionghai Dai |
ICASSP | 2 |
| 2012 | Complexity-reduced geometry partition search and high efficiency prediction for video codingabstractTo reduce the complexity of searching for wedge-based geometry partitions in video coding, we propose a texture-difference based partition line selection approach with a skipping strategy. Applying this approach can reduce the encoding time by 90% while retaining the similar rate-distortion performance to that of exhaustive searching. We also propose a background-compensated prediction scheme to improve the rate-distortion performance of the geometry partitioning prediction by reducing its motion vector overhead. Incorporating our proposed approach into object-boundary-based geometry partitioning can achieve about 10% bit-rate savings relative to the full-search approach while keeping the complexity at about the same level as our proposed complexity-reduced wedge-based geometry partitioning. Qifei Wang, Ming-Ting Sun, Gary J. Sullivan |
ISCAS | 1 |
| 2012 | Free Viewpoint Video Coding With Rate-Distortion AnalysisabstractTo improve free viewpoint video (FVV) coding efficiency and optimize the quality of the synthesized virtual view video, this paper proposes a depth-assisted FVV coding framework and analyzes the rate-distortion (R-D) property of the synthesized virtual view video in FVV coding. In the depth-assisted FVV coding framework, the depth assigned disparity compensated prediction is introduced to exploit the correlation between multiview video (MVV) and depth. To model the R-D property of the synthesized virtual view video, a region-based view synthesis distortion estimation approach is investigated with respect to the distortion of MVV and depth. Subsequently, the general R-D property estimation models of MVV and depth are analyzed. Finally, a rate-allocation scheme is designed to optimize the quantization parameter pair of MVV and depth in FVV coding. The simulation results demonstrate that the proposed depth-assisted FVV coding framework can improve the FVV coding efficiency. The region-based view synthesis distortion estimation approach and the general R-D model are able to precisely approximate the R-D property of synthesized virtual view video in the multiview video plus depth based FVV coding frameworks. The proposed rate-allocation scheme can optimize the overall FVV coding efficiency to achieve a high-quality reconstructed video at the desired viewpoint with a given rate constraint. Qifei Wang, Xiangyang Ji, Qionghai Dai, Naiyao Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Reduced-complexity search for video coding geometry partitions using texture and depth dataabstractIn this paper, a texture-space geometry partitioning approach is proposed to reduce the computational complexity of searching for geometry partitions for video coding. Additionally, for systems that capture both video and depth data, an enhanced geometry partitioning approach using both texture and depth information is proposed to further improve the partitioning accuracy and reduce the search complexity. Compared with a full-search approach, the proposed geometry partition search approaches achieve about 94% reduction of the encoding time while retaining similar rate-distortion performance. Qifei Wang, Gary J. Sullivan, Ming-Ting Sun |
VCIP | 1 |
| 2010 | Region Based Rate-Distortion Analysis for 3D Video CodingabstractSummary form only given. In 3D video (3DV), the virtual view images are commonly synthesized by the color and depth images of the reference views with image based rendering (IBR). Thus, in 3DV coding, to provide the high-quality interactive viewpoint video to audience, it is necessary to jointly optimize coding efficiency of color and depth images at a given bit-rate by rate-distortion (R-D) property analysis of 3DV coding.to calculate Edw, a region based distortion model is proposed firstly. In IBR, depth quantization error will cause the disparities between the pixels of the virtual view and the correspondent pixels of the reference views changed. Qifei Wang, Xiangyang Ji, Qionghai Dai, Naiyao Zhang |
DCC | 1 |