EDBT 2026 Demo / reviewers in the wild / expert
Baoliang Chen
dblp:214/9545
· DBLP profile ↗
46ranked-venue papers
14as first author
42since 2021 · last 2026
0000-0003-4884-6956ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 12 first-author · 34 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality AssessmentabstractDespite the impressive performance of large multimodal models (LMMs) in high-level visual tasks, their capacity for image quality assessment (IQA) remains limited. One main reason is that LMMs are primarily trained for high-level tasks (e.g., image captioning), emphasizing unified image semantics extraction under varied quality. Such semantic-aware yet quality-insensitive perception bias inevitably leads to a heavy reliance on image semantics when those LMMs are forced for quality rating. In this paper, instead of retraining or tuning an LMM costly, we propose a training-free debiasing framework, in which the image quality prediction is rectified by mitigating the bias caused by image semantics. Specifically, we first explore several semantic-preserving distortions that can significantly degrade image quality while maintaining identifiable semantics. By applying these specific distortions to the query/test images, we ensure that the degraded images are recognized as poor quality while their semantics remain. During quality inference, both a query image and its corresponding degraded version are fed to the LMM along with a prompt indicating that the query image quality should be inferred under the condition that the degraded one is deemed poor quality. This prior condition effectively aligns the LMM’s quality perception, as all degraded images are consistently rated as poor quality, regardless of their semantic difference. Finally, the quality scores of the query image inferred under different prior conditions (degraded versions) are aggregated using a conditional probability model. Extensive experiments on various IQA datasets show that our debiasing framework could consistently enhance the LMM performance and the code will be publicly available. Baoliang Chen, Siyi Pan, Dongxu Wu, Liang Xie 0013, Xiangjie Sui, Lingyu Zhu 0006, Hanwei Zhu |
AAAI | 1 |
| 2026 | Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality AssessmentabstractRecent efforts have repurposed the Contrastive Language-Image Pre-training (CLIP) model for No-Reference Image Quality Assessment (NR-IQA) by measuring the cosine similarity between the image embedding and textual prompts such as "a good photo" or "a bad photo." However, this semantic similarity overlooks a critical yet underexplored cue: the magnitude of the CLIP image features, which we empirically find to exhibit a strong correlation with perceptual quality. In this work, we introduce a novel adaptive fusion framework that complements cosine similarity with a magnitude-aware quality cue. Specifically, we first extract the absolute CLIP image features and apply a Box-Cox transformation to statistically normalize the feature distribution and mitigate semantic sensitivity. The resulting scalar summary serves as a semantically-normalized auxiliary cue that complements cosine-based prompt matching. To integrate both cues effectively, we further design a confidence-guided fusion scheme that adaptively weighs each term according to its relative strength. Extensive experiments on multiple benchmark IQA datasets demonstrate that our method consistently outperforms standard CLIP-based IQA and state-of-the-art baselines, without any task-specific training. Zhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai, Hanwei Zhu, Lingyu Zhu 0006, Yuncheng Jiang 0004, Baoliang Chen |
AAAI | 8 |
| 2026 | Malice Hides in Equivalence: Attacking Text-Image Alignment Assessment with Subtle Text Variations
Kang Xiao, Xuelin Shen, Baoliang Chen, Wenhan Yang, Meng Wang 0001 |
ISCAS | 3 |
| 2026 | Temporal Quality Aggregation for VQA: Benchmark and Psychology-Inspired Model
Baoliang Chen, Changsheng Gao, Lingyu Zhu 0006, Liang Xie 0013, Hanwei Zhu, Zhijian Hao |
QoMEX | 1 |
| 2025 | AI-generated Image Quality Assessment in Visual CommunicationabstractAssessing the quality of artificial intelligence-generated images (AIGIs) plays a crucial role in their application in real-world scenarios. However, traditional image quality assessment (IQA) algorithms primarily focus on low-level visual perception, while existing IQA works on AIGIs overemphasize the generated content itself, neglecting its effectiveness in real-world applications. To bridge this gap, we propose AIGI-VC, a quality assessment database for AI-Generated Images in Visual Communication, which studies the communicability of AIGIs in the advertising field from the perspectives of information clarity and emotional interaction. The dataset consists of 2,500 images spanning 14 advertisement topics and 8 emotion types. It provides coarse-grained human preference annotations and fine-grained preference descriptions, benchmarking the abilities of IQA methods in preference prediction, interpretation, and reasoning. We conduct an empirical study of existing representative IQA methods and large multi-modal models on the AIGI-VC dataset, uncovering their strengths and weaknesses. Yu Tian 0010, Baoliang Chen, Hanwei Zhu, Shiqi Wang 0001, Sam Kwong |
AAAI | 3 |
| 2025 | The Loop Game: Quality Assessment and Optimization for Low-Light Image Enhancement
Danni Huang, Lingyu Zhu 0006, Hanwei Zhu, Shiqi Wang 0001, Baoliang Chen |
ICIC (3) | 6 |
| 2025 | Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
Jiaxi Huang, Dongxu Wu, Hanwei Zhu, Lingyu Zhu 0006, Jun Xing, Xu Wang 0006, Baoliang Chen |
PRCV (8) | 7 |
| 2025 | Simple Lines, Big Ideas: Towards Interpretable Assessment of Human Creativity from Drawings
Zhenshan Shi, Sasa Zhao, Hanwei Zhu, Lingyu Zhu 0006, Baoliang Chen, Lei Mo |
PRCV (9) | 6 |
| 2025 | Monotonic and Invertible Network: A General Framework for Learning IQA Model from Mixed Datasets
Baoliang Chen, Kang Xiao, Xuelin Shen, Shiqi Wang 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Seeking the optimal accuracy-rate equilibrium in face recognition
Yu Tian 0010, Fu-Zhao Ou, Shiqi Wang 0001, Baoliang Chen, Sam Kwong |
Neurocomputing | 4 |
| 2025 | Learning-Based Compression for Noisy Images in the WildabstractDigital images in real world applications typically undergo a wide variety of quality degradations before compression or re-compression. Existing learning based codecs are typically data-driven, relying on the predefined compression pipeline with pristine or high quality images as the input. However, the images in the wild may exhibit the substantially different characteristics compared to the high quality images, casting major challenges to the learning based image coding. In this paper, we propose a robust noisy image compression framework with the blind assumption on the specific noise type and level. The specifically designed encoder decomposes the representation of visual content into two types of features, including the Features that represent the Intrinsic Content (FIC) and the Features that account for Additive Degradation (FAD). As such, beyond the philosophy of faithfully reconstructing the given image with high fidelity, only FIC needs to be compactly represented and conveyed. The principled disentanglement strategy facilitates the removal of the redundancy from multiple perspectives (e.g., spatial, channel and content), ensuring the handling of a wide variety of noisy images in the wild. Extensive experimental results show that our model can achieve superior performance in terms of the ultimate quality and exhibit the strong generalizability across images degraded by a variety of means. The proposed scheme also points out a new research avenue on learning based compression for images in the wild, which is technically challenging but desirable in practice. Meng Wang 0017, Baoliang Chen, Rongqun Lin, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | DeepDC: Deep Distance Correlation as a Perceptual Image Quality EvaluatorabstractDeep neural networks pre-trained on ImageNet have demonstrated remarkable transferability for developing effective full-reference image quality assessment (FR-IQA) models. However, existing approaches typically demand pixel-level alignment between reference and distorted images-a requirement that poses significant challenges in practical scenarios involving natural photography and texture similarity evaluation. To address this limitation, we propose a novel FR-IQA model leveraging deep statistical similarity derived from pre-trained features without relying on spatial co-location of these features or requiring fine-tuning with mean opinion scores. Specifically, we employ distance correlation, a potent yet relatively underexplored statistical measure, to quantify similarity between reference and distorted images within a deep feature space. The distance correlation is computed via the ratio of the distance covariance to the product of their respective distance standard deviations, for which we derive a closed-form solution using the inner product of deep double-centered distance matrices. Extensive experimental evaluations across diverse IQA benchmarks demonstrate the superiority and robustness of the proposed model. Furthermore, we demonstrate the utility of our model for optimizing texture synthesis and neural style transfer tasks, achieving state-of-the-art performance in both quantitative measures and qualitative assessments. The implementation is publicly available at https://github.com/h4nwei/DeepDC. Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Shiqi Wang 0001, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2025 | Anomaly-Led Prompting Learning Caption Generating Model and BenchmarkabstractVideo anomaly detection (VAD) is an important intelligent system application, but most current research views it as a coarse binary classification task that lacks a fine-grained understanding of abnormal video sequences. We explore a new task for video anomaly analysis called Comprehensive Video Anomaly Caption (CVAC), which aims to generate comprehensive textual captions (containing scene information such as time, location, anomalous subject, anomalous behavior, etc.) for surveillance videos. CVAC is more consistent with human understanding than VAD, but it has not been well explored. We constructed a large-scale benchmark CVACBench to lead this research. For each video clip, we provide 6 fine-grained annotations, including scene information and abnormal keywords. A new evaluation metric Abnormal-F1 (A-F1) is also proposed to more accurately evaluate the caption generation performance of the model. We also designed a method called Anomaly-Led Generating Prompting Transformer (AGPFormer) as a baseline. In AGPFormer, we introduce an anomaly-led language modeling mechanism (Anomaly-Led MLM, AMLM) to focus on anomalous events in videos. To achieve more efficient cross-modal semantic understanding, we design the Interactive Generating Prompting (IGP) module and Scene Alignment Prompting (SAP) module to explore the divide between video and text modalities from multiple perspectives, and to improve the model's performance in understanding and reasoning about the complex semantics of videos. We conducted experiments on CVACBench by using traditional caption metrics and the proposed metrics, and the experimental results demonstrate the effectiveness of AGPFormer in the field of anomaly caption. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Baoliang Chen |
IEEE Trans. Multim. | 9 |
| 2025 | Debiased Mapping for Full-Reference Image Quality AssessmentabstractAn ideal full-reference image quality (FR-IQA) model should exhibit both high separability for images with different quality and compactness for images with the same or indistinguishable quality. However, existing learning-based FR-IQA models that directly compare images in deep-feature space, usually overly emphasize the quality separability, neglecting to maintain the compactness when images are of similar quality. In our work, we identify that the perception bias mainly stems from an inappropriate subspace where images are projected and compared. For this issue, we propose a Debiased Mapping based quality Measure (DMM), leveraging orthonormal bases formed by singular value decomposition (SVD) in the deep features domain. The SVD effectively decomposes the quality variations into singular values and mapping bases, enabling quality inference with more reliable feature difference measures. Extensive experimental results reveal that our proposed measure could mitigate the perception bias effectively and demonstrates excellent quality prediction performance on various IQA datasets. Baoliang Chen, Hanwei Zhu, Lingyu Zhu 0006, Shanshe Wang, Jingshan Pan, Shiqi Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | HFGlobalFormer: When High-Frequency Recovery Meets Global Context Modeling for Compressed Image DeraindropabstractWhen transmission medium and compression degradation are intertwined, new challenges emerge. This study addresses the problem of raindrop removal from compressed images, where raindrops obscure large areas of the background and compression leads to the loss of high-frequency (HF) information. The restoration of the former requires global contextual information, while the latter necessitates guidance for high-frequency details, resulting in a conflict in utilizing these two types of information when designing existing methods. To address this issue, we propose a novel transformer architecture that leverages the advantages of attention mechanism and HF-friendly design to effectively restore the compressed raindrop images at the framework, component, and module levels. Specifically, at the framework level, we integrate relative position multi-head self-attention and convolutional layers into the proposed low-high-frequency transformer (LHFT), where the former captures global contextual information and the latter focuses on high-frequency information. Their combination effectively resolves the issue of mixed degradation. At the component level, we utilize high-frequency depth-wise convolution (HFDC) with zero-mean kernels to improve the capability to extract high-frequency features, drawing inspiration from typical high-frequency filters like Prewitt and Sobel operators. Finally, at the module level, we introduce a low-high-attention module (LHAM) to adaptively allocate the importance of low and high frequencies along channels for effective fusion. We establish the JPEG-compressed raindrop image dataset and conduct extensive experiments on different compression rates. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods without increasing computational costs. Rongqun Lin, Wenhan Yang, Baoliang Chen, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2025 | RCNet: Deep Recurrent Collaborative Network for Multi-View Low-Light Image EnhancementabstractScene observation from multiple perspectives would bring a more comprehensive visual experience. However, in the context of acquiring multiple views in the dark, the highly correlated views are seriously alienated, making it challenging to improve scene understanding with auxiliary views. Recent single image-based enhancement methods may not be able to provide consistently desirable restoration performance for all views due to the ignorance of potential feature correspondence among different views. To alleviate this issue, we make the first attempt to investigate multi-view low-light image enhancement. First, we construct a new dataset called Multi-View Low-light Triplets (MVLT), including 1,860 pairs of triple images with large illumination ranges and wide noise distribution. Each triplet is equipped with three different viewpoints towards the same scene. Second, we propose a deep multi-view enhancement framework based on the Recurrent Collaborative Network (RCNet). Specifically, in order to benefit from similar texture correspondence across different views, we design the recurrent feature enhancement, alignment and fusion (ReEAF) module, in which intra-view feature enhancement (Intra-view EN) followed by inter-view feature alignment and fusion (Inter-view AF) is performed to model the intra-view and inter-view feature propagation sequentially via multi-view collaboration. In addition, two different modules from enhancement to alignment (E2A) and from alignment to enhancement (A2E) are developed to enable the interactions between Intra-view EN and Inter-view AF, which explicitly utilize attentive feature weighting and sampling for enhancement and alignment, respectively. Experimental results demonstrate that our RCNet significantly outperforms other state-of-the-art methods. All of our dataset, code, and model will be available athttps://github.com/hluo29/RCNet. Baoliang Chen, Lingyu Zhu 0006, Peilin Chen 0001, Shiqi Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Zhangkai Ni, Qi Mao 0002, Shiqi Wang 0001 |
ECCV (23) | 3 |
| 2024 | Sliced Maximal Information Coefficient: A Training-Free Approach for Image Quality Assessment EnhancementabstractFull-reference image quality assessment (FR-IQA) models generally operate by measuring the visual differences between a degraded image and its reference. However, existing FR-IQA models including both the classical ones (e.g., PSNR and SSIM) and deep-learning based measures (e.g., LPIPS and DISTS) still exhibit limitations in capturing the full perception characteristics of the human visual system (HVS). In this paper, instead of designing a new FR-IQA measure, we aim to explore a generalized human visual attention estimation strategy to mimic the process of human quality rating and enhance existing IQA models. In particular, we model human attention generation by measuring the statistical dependency between the degraded image and the reference image. The dependency is captured in a training-free manner by our proposed sliced maximal information coefficient and exhibits surprising generalization in different IQA measures. Experimental results verify the performance of existing IQA models can be consistently improved when our attention module is incorporated. The source code is available at https://github.com/KANGX99/SMIC. Kang Xiao, Xu Wang 0006, Yu-Lin He, Baoliang Chen, Xuelin Shen |
ICME | 4 |
| 2024 | Diffusion-Based Bit-Depth ExpansionabstractDiffusion-based generative models have achieved remarkable success across a variety of applications. However, the potential application for bit-depth expansion has not been extensively studied. This paper introduces a wavelet-based diffusion model for the bit-depth expansion task. In this method, the image is first decomposed into low and high-frequency components via wavelet transformation. This decomposition allows for targeted processing by specialized modules and reduces computational complexity by lowering the image resolution. The low-frequency component is processed in both the forward diffusion and reverse denoising stages. Meanwhile, the high-frequency components are filtered by the High Frequency Denoising Filter (HFDF) to eliminate noise and artifacts. Finally, the low and high-frequency components are recombined into a predicted high-bit-depth image through inverse wavelet transformation. Experimental results demonstrate the superiority of the proposed method in producing perceptually compelling outputs that outperform previous methods. Riyu Lu, Lingyu Zhu 0006, Baoliang Chen, Xiaopeng Fan 0001, Shiqi Wang 0001 |
MMSP | 3 |
| 2024 | Adaptive Image Quality Assessment via Teaching Large Multimodal Model to CompareabstractWhile recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplored. To address this gap, we introduce an all-around LMM-based NR-IQA model, which is capable of producing qualitatively comparative responses and effectively translating these discrete comparison outcomes into a continuous quality score. Specifically, during training, we present to generate scaled-up comparative instructions by comparing images from the same IQA dataset, allowing for more flexible integration of diverse IQA datasets. Utilizing the established large-scale training corpus, we develop a human-like visual quality comparator. During inference, moving beyond binary choices, we propose a soft comparison method that calculates the likelihood of the test image being preferred over multiple predefined anchor images. The quality score is further optimized by maximum a posteriori estimation with the resulting probability matrix. Extensive experiments on nine IQA datasets validate that the Compare2Score effectively bridges text-defined comparative levels during training with converted single image quality scores for inference, surpassing state-of-the-art IQA models across diverse scenarios. Moreover, we verify that the probability-matrix-based inference conversion not only improves the rating accuracy of Compare2Score but also zero-shot general-purpose LMMs, suggesting its intrinsic effectiveness. Hanwei Zhu, Haoning Wu 0001, Baoliang Chen, Lingyu Zhu 0006, Yuming Fang 0001, Guangtao Zhai, Weisi Lin, Shiqi Wang 0001 |
NeurIPS | 5 |
| 2024 | Temporally Consistent Enhancement of Low-Light Videos via Spatial-Temporal Compatible LearningabstractAbstract Temporal inconsistency is the annoying artifact that has been commonly introduced in low-light video enhancement, but current methods tend to overlook the significance of utilizing both data-centric clues and model-centric design to tackle this problem. In this context, our work makes a comprehensive exploration from the following three aspects. First, to enrich the scene diversity and motion flexibility, we construct a synthetic diverse low/normal-light paired video dataset with a carefully designed low-light simulation strategy, which can effectively complement existing real captured datasets. Second, for better temporal dependency utilization, we develop a Temporally Consistent Enhancer Network (TCE-Net) that consists of stacked 3D convolutions and 2D convolutions to exploit spatial-temporal clues in videos. Last, the temporal dynamic feature dependencies are exploited to obtain consistency constraints for different frame indexes. All these efforts are powered by a Spatial-Temporal Compatible Learning (STCL) optimization technique, which dynamically constructs specific training loss functions adaptively on different datasets. As such, multiple-frame information can be effectively utilized and different levels of information from the network can be feasibly integrated, thus expanding the synergies on different kinds of data and offering visually better results in terms of illumination distribution, color consistency, texture details, and temporal coherence. Extensive experimental results on various real-world low-light video datasets clearly demonstrate the proposed method achieves superior performance to state-of-the-art methods. Our code and synthesized low-light video database will be publicly available at https://github.com/lingyzhu0101/low-light-video-enhancement.git . Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Xiandong Meng, Shiqi Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Quality Harmonization for Virtual Composition in Online Video CommunicationsabstractRecent years have witnessed strong demands for video composition in online video communications, enabling a series of new functionalities for video conferencing including virtual conference rooms, virtual reunions, and virtual backgrounds. In video composition, typically the foreground videos including the human bodies and faces are subject to compression due to the constrained bandwidth, whereas the virtual background is uncompressed and in pristine quality. The disharmony caused by the incoherent quality of foreground and background, which may worsen the quality of experience, has not been extensively studied. In this paper, we focus on this particular problem and present an image quality harmonization framework. Our principle is to align the quality of the background with that of the foreground such that they share similar levels of distortion. This is achieved by inferring the quantization parameter for background compression based on the foreground information. In particular, we aim to learn the quality and compression parameters in a self-supervised manner without laborious human annotation. Furthermore, a large dataset is constructed to provide sufficient training samples and testing scenarios for validation. The composite videos show superior harmonized quality in both quantitative and qualitative comparisons, demonstrating the effectiveness of the proposed framework. Binzhe Li, Zhao Wang 0004, Baoliang Chen, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Causal Representation Learning for GAN-Generated Face Image Quality AssessmentabstractRecent years have witnessed significant advancements in face image generation using generative adversarial networks (GANs), leading to a high demand for GAN-generated face image quality assessment (GFIQA). However, the intrinsic distortion caused by the generation brings a significant challenge for existing image quality assessment (IQA) models which are typically designed for natural images. In addition, the image distortion usually varies depending on different GAN models, resulting in a high generalization capability that a GFIQA model should possess. To account for this, we first establish a large GFIQA database by collecting various GFIs from existing popular GAN models. Subsequently, we further propose a causal representation learning (CRL) scheme for the generalized GFIQA model (CRL-GFIQA) with the assumption that the causal knowledge of human quality assessment is shareable in different scenarios. In particular, we disentangle the learned features into casual and non-causal components by an invertible neural network, facilitating the proposed CRL-GFIQA model with a high generalization on unseen domains. Extensive experimental results demonstrate the effectiveness of our CRL-GFIQA model. The codes and the constructed dataset will be publicly available. Yu Tian 0010, Shiqi Wang 0001, Baoliang Chen, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Video Quality Assessment for Spatio-Temporal Resolution Adaptive CodingabstractSpatio-temporal resolution adaptive (STRA) coding has been repeatedly proven to be a promising way to improve coding efficiency and reduce coding complexity. The wide consensus is that the optimal subsampled resolution and frame rate should be governed by so- called generalized rate-distortion performance based on the ultimately perceived distortion. However, it is non-trivial to accurately predict the quality of reconstructed videos due to the fact that the distortion originates from both subsampling and compression. To address this issue, we propose a novel video quality assessment model that is fully aware of the information available in downsampled videos for compression, such as resolution and frame rate. More specifically, the proposed model relies on quality-aware spatial features that are extracted by an image quality fine-tuned backbone. Subsequently, the spatio-temporal quality is modeled based on the transformer encoder, which is adaptive to the downsampling spatial and temporal resolutions. This enables the transformer encoder to produce discriminative features that capture long-range temporal dependencies related to the current context. The quality score, which is the output of the transformer encoder, thus reflects both the influence of the subsampling and compression. We conduct extensive experiments that demonstrate the superiority of the proposed model over state-of-the-art methods on four subsampling and compression video quality datasets. Furthermore, we apply the proposed model to bitrate ladder optimization, leading to a perceptual-aware spatial and temporal downsampling strategy that yields promising bitrate savings. The source codes of the proposed model will be publicly available athttps://github.com/h4nwei/STRA-VQA. Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Peilin Chen 0001, Linqi Song, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | 2AFC Prompting of Large Multimodal Models for Image Quality AssessmentabstractWhile abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models (LMMs), their image quality assessment (IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice (2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posteriori estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available athttps://github.com/h4nwei/2AFC-LMMs. Hanwei Zhu, Xiangjie Sui, Baoliang Chen, Xuelin Liu, Peilin Chen 0001, Yuming Fang 0001, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Deep Feature Statistics Mapping for Generalized Screen Content Image Quality AssessmentabstractThe statistical regularities of natural images, referred to as natural scene statistics, play an important role in no-reference image quality assessment. However, it has been widely acknowledged that screen content images (SCIs), which are typically computer generated, do not hold such statistics. Here we make the first attempt to learn the statistics of SCIs, based upon which the quality of SCIs can be effectively determined. The underlying mechanism of the proposed approach is based upon the mild assumption that the SCIs, which are not physically acquired, still obey certain statistics that could be understood in a learning fashion. We empirically show that the statistics deviation could be effectively leveraged in quality assessment, and the proposed method is superior when evaluated in different settings. Extensive experimental results demonstrate the Deep Feature Statistics based SCI Quality Assessment (DFSS-IQA) model delivers promising performance compared with existing NR-IQA models and shows a high generalization capability in the cross-dataset settings. The implementation of our method is publicly available at https://github.com/Baoliang93/DFSS-IQA. Baoliang Chen, Hanwei Zhu, Lingyu Zhu 0006, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2024 | Gap-Closing Matters: Perceptual Quality Evaluation and Optimization of Low-Light Image EnhancementabstractThere is a growing consensus in the research community that the optimization of low-light image enhancement approaches should be guided by the visual quality perceived by end users. Despite the substantial efforts invested in the design of low-light enhancement algorithms, there has been comparatively limited focus on assessing subjective and objective quality systematically. To mitigate this gap and provide a clear path towards optimizing low-light image enhancement for better visual quality, we propose a gap-closing framework. In particular, our gap-closing framework starts with the creation of a large-scale dataset for Subjective QUality Assessment of REconstructed LOw-Light Images (SQUARE-LOL). This database serves as the foundation for studying the quality of enhanced images and conducting a comprehensive subjective user study. Subsequently, we propose an objective quality assessment measure that plays a critical role in bridging the gap between visual quality and enhancement. Finally, we demonstrate that our proposed objective quality measure can be incorporated into the process of optimizing the learning of the enhancement model toward perceptual optimality. We validate the effectiveness of our proposed framework through both the accuracy of quality prediction and the perceptual quality of image enhancement. Baoliang Chen, Lingyu Zhu 0006, Hanwei Zhu, Wenhan Yang, Linqi Song, Shiqi Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Perceptual Quality Assessment of Face Video Compression: A Benchmark and An Effective MethodabstractRecent years have witnessed an exponential increase in the demand for face video compression, and the success of artificial intelligence has expanded the boundaries beyond traditional hybrid video coding. Generative coding approaches have been identified as promising alternatives with reasonable perceptual rate-distortion trade-offs, leveraging the statistical priors of face videos. However, the great diversity of distortion types in spatial and temporal domains, ranging from the traditional hybrid coding frameworks to generative models, present grand challenges in compressed face video quality assessment (VQA) that plays a crucial role in the whole delivery chain for quality monitoring and optimization. In this paper, we introduce the large-scale Compressed Face Video Quality Assessment (CFVQA) database, which is the first attempt to systematically understand the perceptual quality and diversified compression distortions in face videos. The database contains 3,240 compressed face video clips in multiple compression levels, which are derived from 135 source videos with diversified content using six representative video codecs, including two traditional methods based on hybrid coding frameworks, two end-to-end methods, and two generative methods. The unique characteristics of CFVQA, including large-scale, fine-grained, great content diversity, and cross-compression distortion types, make the benchmarking for existing image quality assessment (IQA) and VQA feasible and practical. The results reveal the weakness of existing IQA and VQA models, which challenge real-world face video applications. In addition, a FAce VideO IntegeRity (FAVOR) index for face video compression was developed to measure the perceptual quality, considering the distinct content characteristics and temporal priors of the face videos. Experimental results exhibit its superior performance on the proposed CFVQA dataset. The benchmark is now made publicly available at:https://github.com/Yixuan423/Compressed-Face-Videos-Quality-Assessment. Baoliang Chen, Meng Wang 0017, Shiqi Wang 0001, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2024 | Towards Thousands to One Reference: Can We Trust the Reference Image for Quality Assessment?abstractTraditional full-reference image quality assessment (FR-IQA) methods predict the perceptual quality of a distorted image with a given pristine-quality image as the reference. However, the near-threshold visual perception suggests that there could be numerous pristine-quality representations that are indistinguishable in a scene, and the so-called pristine image used in FR-IQA for reference is just one of them. With numerous approaches proposed for FR-IQA by evaluating the perceptual similarity, much less work has been dedicated to locating the best reference for the deterministic perceptual This paper aims to answer the question that whether enabling the freedom in reference image selection could lead to better performance by designing a new FR-IQA paradigm FLexible REference (FLRE). The FLRE paradigm is developed in the feature space by attempting to obtain the feature-level reference of the distorted image via the selection of its corresponding best explanation within an equal-quality space. To this end, we devise the Perceptually Near-Threshold Estimation (PNTE) and the Pseudo-Reference Search (PRS) strategies. In particular, the PNTE module predicts the equal-quality map of a given pristine-quality feature, forming an equal-quality space. Subsequently, the PRS strategy is employed to locate the reference of the distorted feature within the equal-quality space in an element-wise minimum distance search manner. Due to the lack of the ground-truth reference (i.e., best explanation) of each distorted image, we optimize the pseudo-reference feature learning under three constraints, i.e., the quality regression loss, the disturbance maximization loss, and the content loss. We implement the FLRE as a plug-in module before the deterministic FR-IQA process, and experimental results have demonstrated that combining FLRE with the existing deep feature-based FR-IQA models can significantly improve the quality prediction performance, largely surpassing the state-of-the-art methods. The implementation of our method is publicly available onhttps://github.com/ytian73/FLRE. Yu Tian 0010, Baoliang Chen, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2024 | CodedBGT: Code Bank-Guided Transformer for Low-Light Image EnhancementabstractLow-light images commonly exhibit issues such as reduced contrast, heightened noise, faded colors, and the absence of critical details. Enhancing these images is challenging due to the complex interplay of various factors. Existing methods primarily focus on learning the intricate mapping between low-light input and normal-light output through well-designed deep neural networks, potentially overlooking the valuable priors inherent in normal-light images. In this paper, we introduce a Code Bank-Guided Transformer (CodedBGT) for low-light image enhancement. Initially, we pre-train a VQGAN on an extensive collection of high-quality normal-light images to capture a high-quality prior. This prior is stored in a discrete codebook along with its corresponding decoded feature space, forming the code bank that guides the enhancement process. To effectively align low-light features with undistorted normal-light code bank features, we design a Code Bank-Guided Block (CBGB) within our enhancement network. The CBGB is integrated into the transformer to aggregate prior information into the enhancement network. Benefiting from the high-quality code bank, our method produces results with more satisfying visual quality. In comparison with the state-of-the-art methods, higher quantitative and qualitative experimental results on the paired dataset and unpaired datasets with various evaluation metrics show the superiority of our method. Dongjie Ye, Baoliang Chen, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2023 | Troubleshooting Ethnic Quality Bias with Curriculum Domain Adaptation for Face Image Quality AssessmentabstractFace Image Quality Assessment (FIQA) lays the foundation for ensuring the stability and accuracy of face recognition systems. However, existing FIQA methods mainly formulate quality relationships within the training set to yield quality scores, ignoring the generalization problem caused by ethnic quality bias between the training and test sets. Domain adaptation presents a potential solution to mitigate the bias, but if FIQA is treated essentially as a regression task, it will be limited by the challenge of feature scaling in transfer learning. Additionally, how to guarantee source risk is also an issue due to the lack of ground-truth labels of the source domain for FIQA. This paper presents the first attempt in the field of FIQA to address these challenges with a novel Ethnic-Quality-Bias Mitigating (EQBM) framework. Specifically, to eliminate the restriction of scalar regression, we first compute the Likert-scale quality probability distributions as source domain annotations. Furthermore, we design an easy-to-hard training scheduler based on the inter-domain uncertainty and intra-domain quality margin as well as the ranking-based domain adversarial network to enhance the effectiveness of transfer learning and further reduce the source risk in domain adaptation. Extensive experiments demonstrate that the EQBM significantly mitigates the quality bias and improves the generalization capability of FIQA across races on different datasets. Fu-Zhao Ou, Baoliang Chen, Chongyi Li, Shiqi Wang 0001, Sam Kwong |
ICCV | 2 |
| 2023 | Learning Spatiotemporal Interactions for User-Generated Video Quality AssessmentabstractDistortions from spatial and temporal domains have been identified as the dominant factors that govern the visual quality. Though both have been studied independently in deep learning-based user-generated content (UGC) video quality assessment (VQA) by frame-wise distortion estimation and temporal quality aggregation, much less work has been dedicated to the integration of them with deep representations. In this paper, we propose a SpatioTemporal Interactive VQA (STI-VQA) model based upon the philosophy that video distortion can be inferred from the integration of both spatial characteristics and temporal motion, along with the flow of time. In particular, for each timestamp, both the spatial distortion explored by the feature statistics and local motion captured by feature difference are extracted and fed to a transformer network for the motion aware interaction learning. Meanwhile, the information flow of spatial distortion from the shallow layer to the deep layer is constructed adaptively during the temporal aggregation. The transformer network enjoys an advanced advantage for long-range dependencies modeling, leading to superior performance on UGC videos. Experimental results on five UGC video benchmarks demonstrate the effectiveness and efficiency of our STI-VQA model, and the source code will be available online athttps://github.com/h4nwei/STI-VQA. Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | DeepWSD: Projecting Degradations in Perceptual Space to Wasserstein Distance in Deep Feature SpaceabstractExisting deep learning-based full-reference IQA (FR-IQA) models usually predict the image quality in a deterministic way by explicitly comparing the features, gauging how severely distorted an image is by how far the corresponding feature lies from the space of the reference images. Herein, we look at this problem from a different viewpoint and propose to model the quality degradation in perceptual space from a statistical distribution perspective. As such, the quality is measured based upon the Wasserstein distance in the deep feature domain. More specifically, the 1D Wasserstein distance at each stage of the pre-trained VGG network is measured, based on which the final quality score is performed. The deep Wasserstein distance (DeepWSD) performed on features from neural networks enjoys better interpretability of the quality contamination caused by various types of distortions and presents an advanced quality prediction capability. Extensive experiments and theoretical analysis show the superiority of the proposed DeepWSD in terms of both quality prediction and optimization. The implementation of our method is publicly available at https://github.com/Buka-Xing/DeepWSD. Xingran Liao, Baoliang Chen, Hanwei Zhu, Shiqi Wang 0001, Mingliang Zhou 0001, Sam Kwong |
ACM Multimedia | 2 |
| 2022 | No-reference Image Quality Assessment via Non-local Dependency ModelingabstractIn this paper, we propose a no-reference image quality assessment method based on non-local features learned by a graph neural network (GNN). The proposed quality assessment framework is rooted in the view that the human visual system perceives image quality with long-dependency constructed among different regions, inspiring us to explore the non-local interactions in quality prediction. Instead of relying on convolutional neural network (CNN) based quality assessment methods that primarily focus on local field features, the GNN aiming for non-local quality perception facilitates modeling such long-dependency. In particular, we first adopt superpixel segmentation for the graph nodes construction. Subsequently, a spatial attention module is proposed to integrate the long- and short-range dependencies among the nodes of the whole image. The learned non-local features are finally combined with the local features extracted by the pre-trained CNN, achieving superior performance to the features utilized individually. Experimental results on intra-dataset and cross-dataset settings verify our proposed method's effectiveness and advanced generalization capability. Source codes are publicly accessible at https://github.com/SuperBruceJia/NLNet-IQA for scientific reproducible research. Shuyue Jia, Baoliang Chen, Dingquan Li, Shiqi Wang 0001 |
MMSP | 2 |
| 2022 | Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality AssessmentabstractIn this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method. The codes are released athttps://github.com/Baoliang93/GSTVQA Baoliang Chen, Lingyu Zhu 0006, Fangbo Lu, Hongfei Fan, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Appearance Matters, So Does Audio: Revealing the Hidden Face via Cross-Modality TransferabstractRecently, there has been an exponential increase in the security concerns raised by faking face (e.g., deepfake), which automatically changes the identity with a specifically learned deep generative model. With numerous approaches proposed to identify the fake content, much less work has been dedicated to automatically revealing the authentic one that is originally acquired. Here, we propose a new paradigm that seeks to reveal the authentic face hidden behind the fake one by leveraging the joint information of face and audio. More specifically, given the fake face as well as the audio segment, the cross-modality transferable capability is exploited by learning to generate the feature of the authentic face, based on the underlying clues from the audio as well as the fake face appearance. The effectiveness of the proposed scheme is validated through a series of evaluations, and experimental results show that the proposed model achieves promising face reconstruction performance in revealing the hidden faces, in terms of reconstruction quality, as well as identity and face attribute inference accuracy. Chenqi Kong, Baoliang Chen, Wenhan Yang, Haoliang Li, Peilin Chen 0001, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Enlightening Low-Light Images With Dynamic Guidance for Context EnrichmentabstractImages acquired in low-light conditions suffer from a series of visual quality degradations,e.g., low visibility, degraded contrast, and intensive noise. These complicated degradations based on various contexts (e.g., noise in smooth regions, over-exposure in well-exposed regions and low contrast around edges) cast major challenges to the low-light image enhancement. Herein, we propose a new methodology by imposing a learnable guidance map from the signal and deep priors, making the deep neural network adaptively enhance low-light images in a region-dependent manner. The enhancement capability of the learnable guidance map is further exploited with the multi-scale dilated context collaboration, leading to contextually enriched feature representations extracted by the model with various receptive fields. Through assimilating the intrinsic perceptual information from the learned guidance map, richer and more realistic textures are generated. Extensive experiments on real low-light images demonstrate the effectiveness of our method, which delivers superior results quantitatively and qualitatively. The code is available athttps://github.com/lingyzhu0101/GEMSCto facilitate future research. Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Fangbo Lu, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Detect and Locate: Exposing Face Manipulation by Semantic- and Noise-Level TelltalesabstractThe technological advancements of deep learning have enabled sophisticated face manipulation schemes, raising severe trust issues and security concerns in modern society. Generally speaking, detecting manipulated faces and locating the potentially altered regions are challenging tasks. Herein, we propose a conceptually simple but effective method to efficiently detect forged faces in an image while simultaneously locating the manipulated regions. The proposed scheme relies on a segmentation map that delivers meaningful high-level semantic information clues about the image. Furthermore, a noise map is estimated, playing a complementary role in capturing low-level clues and subsequently empowering decision-making. Finally, the features from these two modules are combined to distinguish fake faces. Extensive experiments show that the proposed model achieves state-of-the-art detection accuracy and remarkable localization performance. Chenqi Kong, Baoliang Chen, Haoliang Li, Shiqi Wang 0001, Anderson Rocha 0001, Sam Kwong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | No-Reference Image Quality Assessment by Hallucinating Pristine FeaturesabstractIn this paper, we propose a no-reference (NR) image quality assessment (IQA) method via feature level pseudo-reference (PR) hallucination. The proposed quality assessment framework is rooted in the view that the perceptually meaningful features could be well exploited to characterize the visual quality, and the natural image statistical behaviors are exploited in an effort to deliver the accurate predictions. Herein, the PR features from the distorted images are learned by a mutual learning scheme with the pristine reference as the supervision, and the discriminative characteristics of PR features are further ensured with the triplet constraints. Given a distorted image for quality inference, the feature level disentanglement is performed with an invertible neural layer for final quality prediction, leading to the PR and the corresponding distortion features for comparison. The effectiveness of our proposed method is demonstrated on four popular IQA databases, and superior performance on cross-database evaluation also reveals the high generalization capability of our method. The implementation of our method is publicly available on https://github.com/Baoliang93/FPR. Baoliang Chen, Lingyu Zhu 0006, Chenqi Kong, Hanwei Zhu, Shiqi Wang 0001, Zhu Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | PUGCQ: A Large Scale Dataset for Quality Assessment of Professional User-Generated ContentabstractRecent years have witnessed a surge of professional user-generated content (PUGC) based video services, coinciding with the accelerated proliferation of video acquisition devices such as mobile phones, wearable cameras, and unmanned aerial vehicles. Different from traditional UGC videos by impromptu shooting, PUGC videos produced by professional users tend to be carefully designed and edited, receiving high popularity with a relatively satisfactory playing count. In this paper, we systematically conduct the comprehensive study on the perceptual quality of PUGC videos and introduce a database consisting of 10,000 PUGC videos with subjective ratings. In particular, during the subjective testing, we collect the human opinions based upon not only the MOS, but also the attributes that could potentially influence the visual quality including face, noise, blur, brightness, and color. We make the attempt to analyze the large-scale PUGC database with a series of video quality assessment (VQA) algorithms and a dedicated baseline model based on pretrained deep neural network is further presented. The cross-dataset experiments reveal a large domain gap between the PUGC and the traditional user-generated videos, which are critical in learning based VQA. These results shed light on developing next-generation PUGC quality assessment algorithms with desired properties including promising generalization capability, high accuracy, and effectiveness in perceptual optimization. The dataset and the codes are released at https://github.com/wlkdb/pugcq_create. Baoliang Chen, Lingyu Zhu 0006, Qingwen He, Hongfei Fan, Shiqi Wang 0001 |
ACM Multimedia | 2 |
| 2021 | Camera Invariant Feature Learning for Generalized Face Anti-SpoofingabstractThere has been an increasing consensus in learning based face anti-spoofing that the divergence in terms of camera models is causing a large domain gap in real application scenarios. We describe a framework that eliminates the influence of inherent variance from acquisition cameras at the feature level, leading to the generalized face spoofing detection model that could be highly adaptive to different acquisition devices. In particular, the framework is composed of two branches. The first branch aims to learn the camera invariant spoofing features via feature level decomposition in the high frequency domain. Motivated by the fact that the spoofing features exist not only in the high frequency domain, in the second branch the discrimination capability of extracted spoofing features is further boosted from the enhanced image based on the recomposition of the high-frequency and low-frequency information. Finally, the classification results of the two branches are fused together by a weighting strategy. Experiments show that the proposed method can achieve better performance in both intra-dataset and cross-dataset settings, demonstrating the high generalization capability in various application scenarios. Baoliang Chen, Wenhan Yang, Haoliang Li, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | No-Reference Screen Content Image Quality Assessment With Unsupervised Domain AdaptationabstractIn this paper, we quest the capability of transferring the quality of natural scene images to the images that are not acquired by optical cameras (e.g., screen content images, SCIs), rooted in the widely accepted view that the human visual system has adapted and evolved through the perception of natural environment. Here, we develop the first unsupervised domain adaptation based no reference quality assessment method for SCIs, leveraging rich subjective ratings of the natural images (NIs). In general, it is a non-trivial task to directly transfer the quality prediction model from NIs to a new type of content (i.e., SCIs) that holds dramatically different statistical characteristics. Inspired by the transferability of pair-wise relationship, the proposed quality measure operates based on the philosophy of improving the transferability and discriminability simultaneously. In particular, we introduce three types of losses which complementarily and explicitly regularize the feature space of ranking in a progressive manner. Regarding feature discriminatory capability enhancement, we propose a center based loss to rectify the classifier and improve its prediction capability not only for source domain (NI) but also the target domain (SCI). For feature discrepancy minimization, the maximum mean discrepancy (MMD) is imposed on the extracted ranking features of NIs and SCIs. Furthermore, to further enhance the feature diversity, we introduce the correlation penalization between different feature dimensions, leading to the features with lower rank and higher diversity. Experiments show that our method can achieve higher performance on different source-target settings based on a light-weight convolution neural network. The proposed method also sheds light on learning quality assessment measures for unseen application-specific content without the cumbersome and costing subjective evaluations. Baoliang Chen, Haoliang Li, Hongfei Fan, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Single Depth Image Super-Resolution Using Convolutional Neural NetworksabstractIn this paper, we propose single depth image super-resolution using convolutional neural networks (CNN). We adopt CN-N to acquire a high-quality edge map from the input low-resolution (LR) depth image. We use the high-quality edge map as the weight of the regularization term in a total variation (TV) model for super-resolution. First, we interpolate the LR depth image using bicubic interpolation and extract its low-quality edge map. Then, we get the high-quality edge map from the low-quality one using CNN. Since the CNN output often contains broken edges and holes, we refine it using the low-quality edge map. Guided by the high-quality edge map, we upsample the input LR depth image in the TV model. The edge-based guidance in TV effectively removes noise in depth while minimizing jagged artifacts and preserving sharp edges. Various experiments on the Middle-bury stereo dataset and Laser Scan dataset demonstrate the superiority of the proposed method over state-of-the-arts in both qualitative and quantitative measurements. Baoliang Chen, Cheolkon Jung |
ICASSP | 1 |
| 2018 | Patch-Based Stereo Matching Using 3D Convolutional Neural NetworksabstractIn this paper, we propose patch-based stereo matching using 3D convolutional neural networks (CNN). We extract spatial color and disparity features simultaneously through 3D CNN. We treat stereo matching as multi-class classification that the classes are all possible disparity values. We first generate a large set of patches from stereo images for 3D CNN. Then, we get an initial disparity map through 3D CNN and refine it using color image guided filtering. The color image guided filtering minimizes outliers and refines edges in disparity without texture copying artifacts. Experimental results show that the proposed method successfully estimates disparity in smooth and discontinuity regions while preserving edges as well as outperforms state-of-the-arts in terms of average errors. Baoliang Chen, Cheolkon Jung |
ICIP | 1 |
| 2018 | Variational Fusion of Time-of-Flight and Stereo Data for Depth Estimation Using Edge-Selective Joint FilteringabstractIn this paper, we propose variational fusion of time-of-flight (TOF) and stereo data for depth estimation using edge-selective joint filtering (ESJF). ESJF is able to adaptively select edges for depth upsampling from the TOF depth map, stereo matching-based disparity map, and stereo images. We adopt ESJF to produce high-resolution (HR) depth maps with accurate edge information from low-resolution ones captured by the TOF camera. First, we measure confidences of TOF and stereo data based on a Gaussian function to be used as fusion weights. Then, we upsample the TOF depth map using ESJF and extract vertical and horizontal discontinuity maps from it. Finally, we perform variational fusion of TOF and stereo depth data guided by the discontinuity maps. Experimental results show that the proposed method successfully produces HR depth maps and outperforms the state of the art in preserving edges and removing noise. Baoliang Chen, Cheolkon Jung, Zhendong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Variational fusion of time-of-flight and stereo data using edge selective joint filteringabstractIn this paper, we propose variational fusion of time-of-flight (TOF) and stereo data using edge selective joint filtering (ESJF). We utilize ESJF to up-sample low-resolution (LR) depth captured by TOF camera and produce high-resolution (HR) depth maps with accurate edge information. First, we measure confidence of two sensor with different reliability to fuse them. Then, we up-sample TOF depth map using ESJF to generate discontinuity maps and protect edges in depth. Finally, we perform variational fusion of TOF and stereo depth data based on total variation (TV) guided by discontinuity maps. Experimental results show that the proposed method successfully produces HR depth maps and outperforms the-state-of-the-art ones in preserving edges and removing noise. Baoliang Chen, Cheolkon Jung, Zhendong Zhang 0001 |
ICIP | 1 |