VLDB 2026 Research / reviewers in the wild / expert
Kede Ma
dblp:127/1809
· DBLP profile ↗
86ranked-venue papers
18as first author
49since 2021 · last 2026
0000-0001-8608-1128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 16 first-author · 30 since 2021Artificial intelligence and machine learning · 38 · 2 first-author · 34 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised AI-Generated Image Detection: A Camera Metadata PerspectiveabstractThe proliferation of AI-generated imagery poses escalating challenges for multimedia forensics, yet many existing detectors depend on assumptions about the internals of specific generative models, limiting their cross-model applicability. We introduce a self-supervised approach for detecting AI-generated images that leverages camera metadata-specifically exchangeable image file format (EXIF) tags-to learn features intrinsic to digital photography. Our pretext task trains a feature extractor solely on camera-captured photographs by classifying categorical EXIF tags (e.g., camera model and scene type) and pairwise-ranking ordinal and continuous EXIF tags (e.g., focal length and aperture value). Using these EXIF-induced features, we first perform one-class detection by modeling the distribution of photographic images with a Gaussian mixture model and flagging low-likelihood samples as AI-generated. We then extend to binary detection that treats the learned extractor as a strong regularizer for a classifier of the same architecture, operating on high-frequency residuals from spatially scrambled patches. Extensive experiments across various generative models demonstrate that our EXIF-induced detectors substantially advance the state of the art, delivering strong generalization to in-the-wild samples and robustness to common benign image perturbations. Nan Zhong, Mian Zou, Zhenxing Qian, Xinpeng Zhang 0001, Baoyuan Wu, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Pseudocylindrical Convolutions for Learned Omnidirectional Image CompressionabstractEquirectangular projection (ERP) is a convenient form to store omnidirectional images, but it is neither equal-area nor conformal, creating challenges for subsequent visual communication. When used for image compression, ERP amplifies sampling density and deforms objects near the poles, hindering perceptually optimal bit allocation. Here, we present one of the earliest endeavors to apply deep neural networks to omnidirectional image compression. We first propose parametric pseudocylindrical representations that generalize common pseudocylindrical map projections. A tractable greedy algorithm is introduced to identify (sub-)optimal representation configurations, guided by a proxy objective for rate-distortion performance. We then develop pseudocylindrical convolutions, which can be efficiently implemented by standard convolutions with “pseudocylindrical padding.” To demonstrate the utility of the proposed pseudocylindrical representations and convolutions, we implement an end-to-end omnidirectional image compression method, consisting of an analysis transform, a uniform quantizer, a synthesis transform, and an entropy model. Experiments show that our optimized method achieves consistently better rate-distortion performance compared to the state-of-the-art. Mu Li 0005, Kede Ma, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionabstractKehua Feng, Keyan Ding, Tan Hongzhi, Kede Ma, Zhihua Wang, Shuangquan Guo, Cheng Yuzhou, Ge Sun, Guozhou Zheng, Qiang Zhang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma, Zhihua Wang 0002, Shuangquan Guo, Yuzhou Cheng, Guozhou Zheng, Qiang Zhang 0026, Huajun Chen |
ACL (1) | 4 |
| 2025 | Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality AssumptionabstractFull-Reference image quality assessment (FR-IQA) generally assumes that reference images are of perfect quality. However, this assumption is flawed due to the sensor and optical limitations of modern imaging systems. Moreover, recent generative enhancement methods are capable of producing images of higher quality than their original. All of these challenge the effectiveness and applicability of current FR-IQA models. To relax the assumption of perfect reference image quality, we build a large-scale IQA database, namely DiffIQA, containing approximately 180,000 images generated by a diffusion-based image enhancer with adjustable hyper-parameters. Each image is annotated by human subjects as either worse, similar, or better quality compared to its reference. Building on this, we present a generalized FR-IQA model, namely Adaptive Fidelity-Naturalness Evaluator (A-FINE), to accurately assess and adaptively combine the fidelity and naturalness of a test image. A-FINE aligns well with standard FR-IQA when the reference image is much more natural than the test image. We demonstrate by extensive experiments that A-FINE surpasses standard FR-IQA models on well-established IQA datasets and our newly created DiffIQA. To further validate A-FINE, we additionally construct a super-resolution IQA benchmark (SRIQA-Bench), encompassing test images derived from ten state-of-the-art SR methods with reliable human quality annotations. Tests on SRIQA-Bench re-affirm the advantages of A-FINE. The code and dataset are available at https://tianhewu.github.io/A-FINEpage.github.io/. Tianhe Wu, Kede Ma, Lei Zhang 0006 |
CVPR | 3 |
| 2025 | Hiding Images in Diffusion Models by Editing Learned Score FunctionsabstractHiding data using neural networks (i.e., neural steganography) has achieved remarkable success across both discriminative classifiers and generative adversarial networks. However, the potential of data hiding in diffusion models remains relatively unexplored. Current methods exhibit limitations in achieving high extraction accuracy, model fidelity, and hiding efficiency due primarily to the entanglement of the hiding and extraction processes with multiple denoising diffusion steps. To address these, we describe a simple yet effective approach that embeds images at specific timesteps in the reverse diffusion process by editing the learned score functions. Additionally, we introduce a parameter-efficient fine-tuning method that combines gradient-based parameter selection with low-rank adaptation to enhance model fidelity and hiding efficiency. Comprehensive experiments demonstrate that our method extracts high-quality images at human-indistinguishable levels, replicates the original model behaviors at both sample and population levels, and embeds images orders of magnitude faster than prior methods. Besides, our method naturally supports multi-recipient scenarios through independent extraction channels. Yunqiao Yang 0001, Nan Zhong, Kede Ma |
CVPR | 4 |
| 2025 | Self-Supervised Learning for Detecting AI-Generated Faces as AnomaliesabstractThe detection of AI-generated faces is commonly approached as a binary classification task. Nevertheless, the resulting detectors frequently struggle to adapt to novel AI face generators, which evolve rapidly. In this paper, we describe an anomaly detection method for AI-generated faces by leveraging self-supervised learning of camera-intrinsic and face-specific features purely from photographic face images. The success of our method lies in designing a pretext task that trains a feature extractor to rank four ordinal exchangeable image file format (EXIF) tags and classify artificially manipulated face images. Subsequently, we model the learned feature distribution of photographic face images using a Gaussian mixture model. Faces with low likelihoods are flagged as AI-generated. Both quantitative and qualitative experiments validate the effectiveness of our method. Our code is available at https://github.com/MZMMSEC/AIGFD_EXIF.git. Mian Zou, Baosheng Yu, Yibing Zhan, Kede Ma |
ICASSP | 4 |
| 2025 | Dataset Distillation as Data Compression: A Rate-Utility PerspectiveabstractDriven by the ``scale-is-everything'' paradigm, modern machine learning increasingly demands ever-larger datasets and models, yielding prohibitive computational and storage requirements. Dataset distillation mitigates this by compressing an original dataset into a small set of synthetic samples, while preserving its full utility. Yet, existing methods either maximize performance under fixed storage budgets or pursue suitable synthetic data representations for redundancy removal, without jointly optimizing both objectives. In this work, we propose a joint rate-utility optimization method for dataset distillation. We parameterize synthetic samples as optimizable latent codes decoded by extremely lightweight networks. We estimate the Shannon entropy of quantized latents as the rate measure and plug any existing distillation loss as the utility measure, trading them off via a Lagrange multiplier. To enable fair, cross-method comparisons, we introduce bits per class (bpc), a precise storage metric that accounts for sample, label, and decoder parameter costs. On CIFAR-10, CIFAR-100, and ImageNet-128, our method achieves up to $170\times$ greater compression than standard distillation at comparable accuracy. Across diverse bpc budgets, distillation losses, and backbone architectures, our approach consistently establishes better rate-utility trade-offs. Youneng Bao, Yongsheng Liang 0001, Mu Li 0005, Kede Ma |
ICCV | 6 |
| 2025 | Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
Mian Zou, Nan Zhong, Baosheng Yu, Yibing Zhan, Kede Ma |
ICCV | 5 |
| 2025 | CLDyB: Towards Dynamic Benchmarking for Continual Learning with Pre-trained ModelsabstractThe emergence of the foundation model era has sparked immense research interest in utilizing pre-trained representations for continual learning~(CL), yielding a series of strong CL methods with outstanding performance on standard evaluation benchmarks. Nonetheless, there are growing concerns regarding potential data contamination within the massive pre-training datasets. Furthermore, the static nature of standard evaluation benchmarks tends to oversimplify the complexities encountered in real-world CL scenarios, putting CL methods at risk of overfitting to these benchmarks while still lacking robustness needed for more demanding real-world applications. To solve these problems, this paper proposes a general framework to evaluate methods for Continual Learning on Dynamic Benchmarks (CLDyB). CLDyB continuously identifies inherently challenging tasks for the specified CL methods and evolving backbones, and dynamically determines the sequential order of tasks at each time step in CL using a tree-search algorithm, guided by an overarching goal to generate highly challenging task sequences for evaluation. To highlight the significance of dynamic evaluation on the CLDyB, we first simultaneously evaluate multiple state-of-the-art CL methods under CLDyB, resulting in a set of commonly challenging task sequences where existing CL methods tend to underperform. We intend to publicly release these task sequences for the CL community to facilitate the training and evaluation of more robust CL algorithms. Additionally, we perform individual evaluations of the CL methods under CLDyB, yielding informative evaluation results that reveal the specific strengths and weaknesses of each method. Shengzhuang Chen, Yikai Liao, Kede Ma, Ying Wei 0001 |
ICLR | 4 |
| 2025 | SD-LoRA: Scalable Decoupled Low-Rank Adaptation for Class Incremental LearningabstractContinual Learning (CL) with foundation models has recently emerged as a promising paradigm to exploit abundant knowledge acquired during pre-training for tackling sequential tasks. However, existing prompt-based and Low-Rank Adaptation-based (LoRA-based) methods often require expanding a prompt/LoRA pool or retaining samples of previous tasks, which poses significant scalability challenges as the number of tasks grows.
To address these limitations, we propose Scalable Decoupled LoRA (SD-LoRA) for class incremental learning, which continually separates the learning of the magnitude and direction of LoRA components without rehearsal. Our empirical and theoretical analysis reveals that SD-LoRA tends to follow a low-loss trajectory and converges to an overlapping low-loss region for all learned tasks, resulting in an excellent stability-plasticity trade-off. Building upon these insights, we introduce two variants of SD-LoRA with further improved parameter efficiency. All parameters of SD-LoRAs can be end-to-end optimized for CL objectives. Meanwhile, they support efficient inference by allowing direct evaluation with the finally trained model, obviating the need for component selection. Extensive experiments across multiple CL benchmarks and foundation models consistently validate the effectiveness of SD-LoRA. The code is available at https://github.com/WuYichen-97/SD-Lora-CL. Hongming Piao, Long-Kai Huang, Renzhen Wang, Wanhua Li 0001, Hanspeter Pfister, Deyu Meng, Kede Ma, Ying Wei 0001 |
ICLR | 8 |
| 2025 | Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank AdaptationabstractLow-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters’ activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter’s marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto. Zhan Zhuang, Xiequn Wang, Yulong Zhang 0005, Qiushi Huang, Shuhao Chen, Xuehao Wang, Yanbin Wei, Yuhe Nie, Kede Ma, Yu Zhang 0006, Ying Wei 0001 |
ICML | 10 |
| 2025 | VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankabstractDeepSeek-R1 has demonstrated remarkable effectiveness in incentivizing reasoning and generalization capabilities of large language models (LLMs) through reinforcement learning. Nevertheless, the potential of reasoning-induced computation has not been thoroughly explored in the context of image quality assessment (IQA), a task depending critically on visual reasoning. In this paper, we introduce VisualQuality-R1, a reasoning-induced no-reference IQA (NR-IQA) model, and we train it with reinforcement learning to rank, a learning algorithm tailored to the intrinsically relative nature of visual quality. Specifically, for a pair of images, we employ group relative policy optimization to generate multiple quality scores for each image. These estimates are used to compute comparative probabilities of one image having higher quality than the other under the Thurstone model. Rewards for each quality estimate are defined using continuous fidelity measures rather than discretized binary labels. Extensive experiments show that the proposed VisualQuality-R1 consistently outperforms discriminative deep learning-based NR-IQA models as well as a recent reasoning-induced quality regression method. Moreover, VisualQuality-R1 is capable of generating contextually rich, human-aligned quality descriptions, and supports multi-dataset training without requiring perceptual scale realignment. These features make VisualQuality-R1 especially well-suited for reliably measuring progress in a wide range of image processing tasks like super-resolution and image generation. Tianhe Wu, Jie Liang 0007, Lei Zhang 0006, Kede Ma |
NeurIPS | 5 |
| 2025 | AniClipart: Clipart Animation with Text-to-Video PriorsabstractAbstract Clipart, a pre-made graphic art form, offers a convenient and efficient way of illustrating visual content. Traditional workflows to convert static clipart images into motion sequences are laborious and time-consuming, involving numerous intricate steps like rigging, key animation and in-betweening. Recent advancements in text-to-video generation hold great potential in resolving this problem. Nevertheless, direct application of text-to-video generation models often struggles to retain the visual identity of clipart images or generate cartoon-style motions, resulting in unsatisfactory animation outcomes. In this paper, we introduce AniClipart, a system that transforms static clipart images into high-quality motion sequences guided by text-to-video priors. To generate cartoon-style and smooth motion, we first define Bézier curves over keypoints of the clipart image as a form of motion regularization. We then align the motion trajectories of the keypoints with the provided text prompt by optimizing the Video Score Distillation Sampling (VSDS) loss, which encodes adequate knowledge of natural motion within a pretrained text-to-video diffusion model. With a differentiable As-Rigid-As-Possible shape deformation algorithm, our method can be end-to-end optimized while maintaining deformation rigidity. Experimental results show that the proposed AniClipart consistently outperforms existing image-to-video generation models, in terms of text-video alignment, visual identity preservation, and motion consistency. Furthermore, we showcase the versatility of AniClipart by adapting it to generate a broader array of animation formats, such as layered animation, which allows topological changes. Ronghuan Wu, Wanchao Su, Kede Ma, Jing Liao 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding ApproachabstractIn recent years, the multimedia forensics and security community has seen remarkable progress in multitask learning for DeepFake (i.e., face forgery) detection. The prevailing approach has been to frame DeepFake detection as a binary classification problem augmented by manipulation-oriented auxiliary tasks. This scheme focuses on learning features specific to face manipulations with limited generalizability. In this paper, we delve deeper into semantics-oriented multitask learning for DeepFake detection, capturing the relationships among face semantics via joint embedding. We first propose an automated dataset expansion technique that broadens current face forgery datasets to support semantics-oriented DeepFake detection tasks at both the global face attribute and local face region levels. Furthermore, we resort to the joint embedding of face images and labels (depicted by text descriptions) for prediction. This approach eliminates the need for manually setting task-agnostic and task-specific parameters, which is typically required when predicting multiple labels directly from images. In addition, we employ bi-level optimization to dynamically balance the fidelity loss weightings of various tasks, making the training process fully automated. Extensive experiments on six DeepFake datasets show that our method improves the generalizability of DeepFake detection and renders some degree of model interpretation by providing human-understandable explanations. Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Semantic Contextualization of Face Forgery: A New Definition, Dataset, and Detection MethodabstractIn recent years, deep learning has greatly streamlined the process of manipulating photographic face images. Aware of the potential dangers, researchers have developed various tools to spot these counterfeits. Yet, none asks the fundamental question:What digital manipulations make a real photographic face image fake, while others do not? In this paper, we put face forgery in a semantic context and define thatcomputational methods that alter semantic face attributes to exceed human discrimination thresholds are sources of face forgery. Following our definition, we construct a large face forgery image dataset, where each image is associated with a set of labels organized in a hierarchical graph. Our dataset enables two new testing protocols to probe the generalizability of face forgery detectors. Moreover, we propose a semantics-oriented face forgery detection method that captures label relations and prioritizes the primary task (i.e., real or fake face detection). We show that the proposed dataset successfully exposes the weaknesses of current detectors as the test set and consistently improves their generalizability as the training set. Additionally, we demonstrate the superiority of our semantics-oriented method over traditional binary and multi-class classification-based detectors. Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Perceptually Optimized and Self-Calibrated Tone Mapping OperatorabstractWith the increasing popularity and accessibility of high dynamic range (HDR) photography, tone mapping operators (TMOs) for dynamic range compression are practically demanding. In this paper, we develop a two-stage neural network-based TMO that is self-calibrated and perceptually optimized. In Stage one, motivated by the physiology of the early stages of the human visual system, we first decompose an HDR image into a normalized Laplacian pyramid. We then use two lightweight deep neural networks, taking the normalized representation as input and estimating the Laplacian pyramid of the corresponding LDR image. We optimize the tone mapping network by minimizing the normalized Laplacian pyramid distance, a perceptual metric aligning with human judgments of tone-mapped image quality. In Stage two, we input the same HDR image-self-calibrated to different maximum luminance levels-into the learned tone mapping network, and generate a pseudo-multi-exposure image stack with varying detail visibility and color saturation. We then train another fusion network to merge the LDR image stack into a desired LDR image by maximizing a variant of the structural similarity index for multi-exposure image fusion, proven perceptually relevant to fused image quality. Extensive experiments show that our method produces images with consistently better visual quality while ranking among the fastest local TMOs. Peibei Cao, Chenyang Le, Yuming Fang 0001, Kede Ma |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Perceptual Assessment and Optimization of HDR Image RenderingabstractHigh dynamic range (HDR) rendering has the ability to faithfully reproduce the wide luminance ranges in natural scenes, but how to accurately assess the rendering quality is relatively underexplored. Existing quality models are mostly designed for low dynamic range (LDR) images, and do not align well with human perception of HDR image quality. To fill this gap, we propose a family of HDR quality metrics, in which the key step is employing a simple inverse display model to decompose an HDR image into a stack of LDR images with varying exposures. Subsequently, these decomposed images are assessed through well-established LDR quality metrics. Our HDR quality models present three distinct benefits. First, they directly inherit the recent advancements of LDR quality metrics. Second, they do not rely on human perceptual data of HDR image quality for re-calibration. Third, they facilitate the alignment and prioritization of specific luminance ranges for more accurate and detailed quality assessment. Experimental results show that our HDR quality metrics consistently outperform existing models in terms of quality assessment on four HDR image quality datasets and perceptual optimization of HDR novel view synthesis. Peibei Cao, Rafal Mantiuk, Kede Ma |
CVPR | 3 |
| 2024 | Learned Scanpaths Aid Blind Panoramic Video Quality AssessmentabstractPanoramic videos have the advantage of providing an immersive and interactive viewing experience. Nevertheless, their spherical nature gives rise to various and uncertain user viewing behaviors, which poses significant challenges for panoramic video quality assessment (PVQA). In this work, we propose an end-to-end optimized, blind PVQA method with explicit modeling of user viewing patterns through visual scanpaths. Our method consists of two modules: a scanpath generator and a quality assessor. The scanpath generator is initially trained to predict future scanpaths by minimizing their expected code length and then jointly optimized with the quality assessor for quality prediction. Our blind PVQA method enables direct quality assessment of panoramic images by treating them as videos composed of identical frames. Experiments on three public panoramic image and video quality datasets, encompassing both synthetic and authentic distortions, validate the superiority of our blind PVQA model over existing methods. Kanglong Fan, Wen Wen 0007, Mu Li 0005, Yifan Peng 0001, Kede Ma |
CVPR | 5 |
| 2024 | Modular Blind Video Quality AssessmentabstractBlind video quality assessment (BVQA) plays a pivotal role in evaluating and improving the viewing experience of end-users across a wide range of video-based platforms and services. Contemporary deep learning-based models primarily analyze video content in its aggressively subsampled format, while being blind to the impact of the actual spatial resolution and frame rate on video quality. In this paper, we propose a modular BVQA model and a method of training it to improve its modularity. Our model comprises a base quality predictor, a spatial rectifier, and a temporal rectifier, responding to the visual content and distortion, spatial resolution, and frame rate changes on video quality, respectively. During training, spatial and temporal rectifiers are dropped out with some probabilities to render the base quality predictor a standalone BVQA model, which should work better with the rectifiers. Extensive experiments on both professionally-generated content and user-generated content video databases show that our quality model achieves superior or comparable performance to current methods. Additionally, the modularity of our model offers an opportunity to analyze existing video quality databases in terms of their spatial and temporal complexity. Wen Wen 0007, Mu Li 0005, Yabin Zhang 0002, Yiting Liao, Kede Ma |
CVPR | 7 |
| 2024 | Learned HDR Image Compression for Perceptually Optimal Storage and Display
Peibei Cao, Jingzhe Ma, Yu-Chieh Yuan, Zhiyong Xie, Haiqing Bai, Kede Ma |
ECCV (49) | 8 |
| 2024 | Multiscale Sliced Wasserstein Distances as Perceptual Color Difference Measures
Zhihua Wang 0002, Leon Wang, Tsein-I Liu, Yuming Fang 0001, Qilin Sun 0001, Kede Ma |
ECCV (53) | 7 |
| 2024 | Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors
Wei Shang 0001, Dongwei Ren, Yuming Fang 0001, Wangmeng Zuo, Kede Ma |
ECCV (57) | 6 |
| 2024 | A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
Tianhe Wu, Kede Ma, Jie Liang 0007, Yujiu Yang 0001, Lei Zhang 0006 |
ECCV (74) | 2 |
| 2024 | Learning Where to Edit Vision TransformersabstractModel editing aims to data-efficiently correct predictive errors of large pre-trained models while ensuring generalization to neighboring failures and locality to minimize unintended effects on unrelated examples. While significant progress has been made in editing Transformer-based large language models, effective strategies for editing vision Transformers (ViTs) in computer vision remain largely untapped. In this paper, we take initial steps towards correcting predictive errors of ViTs, particularly those arising from subpopulation shifts. Taking a locate-then-edit approach, we first address the ``where-to-edit`` challenge by meta-learning a hypernetwork on CutMix-augmented data generated for editing reliability. This trained hypernetwork produces generalizable binary masks that identify a sparse subset of structured model parameters, responsive to real-world failure samples. Afterward, we solve the ``how-to-edit`` problem by simply fine-tuning the identified parameters using a variant of gradient descent to achieve successful edits. To validate our method, we construct an editing benchmark that introduces subpopulation shifts towards natural underrepresented images and AI-generated images, thereby revealing the limitations of pre-trained ViTs for object recognition. Our approach not only achieves superior performance on the proposed benchmark but also allows for adjustable trade-offs between generalization and locality. Our code is available at https://github.com/hustyyq/Where-to-Edit. Yunqiao Yang 0001, Long-Kai Huang, Shengzhuang Chen, Kede Ma, Ying Wei 0001 |
NeurIPS | 4 |
| 2024 | Steerable Pyramid Transform Enables Robust Left Ventricle Quantification
Kede Ma, Wufeng Xue |
PRCV (14) | 2 |
| 2024 | Analysis of Video Quality Datasets via Design of Minimalistic Video Quality ModelsabstractBlind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA models have been measured primarily on a few human-rated VQA datasets. Thus, it is crucial to gain a better understanding of existing VQA datasets in order to properly evaluate the current progress in BVQA. Towards this goal, we conduct a first-of-its-kind computational analysis of VQA datasets via designing minimalistic BVQA models. By minimalistic, we restrict our family of BVQA models to build only upon basic blocks: a video preprocessor (for aggressive spatiotemporal downsampling), a spatial quality analyzer, an optional temporal quality analyzer, and a quality regressor, all with the simplest possible instantiations. By comparing the quality prediction performance of different model variants on eight VQA datasets with realistic distortions, we find that nearly all datasets suffer from the easy dataset problem of varying severity, some of which even admit blind image quality assessment (BIQA) solutions. We additionally justify our claims by comparing our model generalization capabilities on these VQA datasets, and by ablating a dizzying set of BVQA design choices related to the basic building blocks. Our results cast doubt on the current progress in BVQA, and meanwhile shed light on good practices of constructing next-generation VQA datasets and models. Wei Sun 0029, Wen Wen 0007, Xiongkuo Min, Long Lan, Guangtao Zhai, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Perceptual Quality Assessment of Virtual Reality Videos in the WildabstractInvestigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due to complexauthenticdistortionslocalizedinspaceandtime.Existingpanoramic video databases only consider synthetic distortions, assume fixed viewing conditions, and are limited in size. To overcome these shortcomings, we construct the VR Video Quality in the Wild (VRVQW) database, containing 502 user-generated videos with diverse content and distortion characteristics. Based on VRVQW, we conduct a formal psychophysical experiment to record the scanpaths and perceived quality scores from 139 participants under two different viewing conditions. We provide a thorough statistical analysis of the recordeddata, observing significantimpact of viewing conditions on both human scanpaths and perceived quality. Moreover, we develop an objective quality assessment model for VR videos based on pseudocylindrical representation and convolution. Results on the proposed VRVQW show that our method is superior to existing video quality assessment models.We have made the database and code available at https://github.com/ limuhit/VR-Video-Quality-in-the-Wild. Wen Wen 0007, Mu Li 0005, Yiru Yao, Xiangjie Sui, Yabin Zhang 0002, Long Lan, Yuming Fang 0001, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Hierarchical Prior-Based Super Resolution for Point Cloud Geometry CompressionabstractThe Geometry-based Point Cloud Compression (G-PCC) has been developed by the Moving Picture Experts Group to compress point clouds efficiently. Nevertheless, in its lossy mode, the reconstructed point cloud by G-PCC often suffers from noticeable distortions due to naïve geometry quantization (i.e., grid downsampling). This paper proposes a hierarchical prior-based super resolution method for point cloud geometry compression. The content-dependent hierarchical prior is constructed at the encoder side, which enables coarse-to-fine super resolution of the point cloud geometry at the decoder side. A more accurate prior generally yields improved reconstruction performance, albeit at the cost of increased bits required to encode this piece of side information. Our experiments on the MPEG Cat1A dataset demonstrate substantial Bjøntegaard-delta bitrate savings, surpassing the performance of the octree-based and trisoup-based G-PCC v14. We provide our implementations for reproducible research at https://github.com/lidq92/mpeg-pcc-tmc13. Dingquan Li, Kede Ma, Jing Wang 0115, Ge Li 0002 |
IEEE Trans. Image Process. | 2 |
| 2024 | Task-Specific Normalization for Continual Learning of Blind Image Quality ModelsabstractIn this paper, we present a simple yet effective continual learning method for blind image quality assessment (BIQA) with improved quality prediction accuracy, plasticity-stability trade-off, and task-order/-length robustness. The key step in our approach is to freeze all convolution filters of a pre-trained deep neural network (DNN) for an explicit promise of stability, and learn task-specific normalization parameters for plasticity. We assign each new IQA dataset (i.e., task) a prediction head, and load the corresponding normalization parameters to produce a quality score. The final quality estimate is computed by a weighted summation of predictions from all heads with a lightweight K -means gating mechanism. Extensive experiments on six IQA datasets demonstrate the advantages of the proposed method in comparison to previous training techniques for BIQA. Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Learning a Deep Color Difference Metric for Photographic ImagesabstractMost well-established and widely used color difference (CD) metrics are handcrafted and subjectcalibrated against uniformly colored patches, which do not generalize well to photographic images characterized by natural scene complexities. Constructing CD formulae for photo-graphic images is still an active research topic in imaging/illumination, vision science, and color science communities. In this paper, we aim to learn a deep CD metric for photographic images with four desirable properties. First, it well aligns with the observations in vision science that color and form are linked inextricably in visual cortical processing. Second, it is a proper metric in the mathematical sense. Third, it computes accurate CDs between photographic images, differing mainly in color appearances. Fourth, it is robust to mild geometric distortions (e.g., translation or due to parallax), which are often present in photographic images of the same scene captured by different digital cameras. We show that all four properties can be satisfied at once by learning a multi-scale autoregressive normalizing flow for feature transform, followed by the Euclidean distance which is linearly proportional to the human perceptual CD. Quantitative and qualitative experiments on the large-scale SPCD dataset demonstrate the promise of the learned CD metric. Source code is available at https://github.com/haoychen3/CD-Flow. Zhihua Wang 0002, Yang Yang 0003, Qilin Sun 0001, Kede Ma |
CVPR | 5 |
| 2023 | Joint Video Multi-Frame Interpolation and Deblurring under Unknown Exposure TimeabstractNatural videos captured by consumer cameras often suffer from low framerate and motion blur due to the combination of dynamic scene complexity, lens and sensor imperfection, and less than ideal exposure setting. As a result, computational methods that jointly perform video frame interpolation and deblurring begin to emerge with the unrealistic assumption that the exposure time is known and fixed. In this work, we aim ambitiously for a more realistic and challenging task - joint video multi-frame interpolation and deblurring under unknown exposure time. Toward this goal, we first adopt a variant of supervised contrastive learning to construct an exposure-aware representation from input blurred frames. We then train two U-Nets for intramotion and inter-motion analysis, respectively, adapting to the learned exposure representation via gain tuning. We finally build our video reconstruction network upon the exposure and motion representation by progressive exposureadaptive convolution and motion refinement. Extensive experiments on both simulated and real-world datasets show that our optimized method achieves notable performance gains over the state-of-the-art on the joint video ×8 interpolation and deblurring task. Moreover, on the seemingly implausible ×16 interpolation task, our method outperforms existing methods by more than 1.5 dB in terms of PSNR. Wei Shang 0001, Dongwei Ren, Yi Yang 0001, Kede Ma, Wangmeng Zuo |
CVPR | 5 |
| 2023 | Blind Image Quality Assessment via Vision-Language Correspondence: A Multitask Learning PerspectiveabstractWe aim at advancing blind image quality assessment (BIQA), which predicts the human perception of image quality without any reference information. We develop a general and automated multitask learning scheme for BIQA to exploit auxiliary knowledge from other tasks, in a way that the model parameter sharing and the loss weighting are determined automatically. Specifically, we first describe all candidate label combinations (from multiple tasks) using a textual template, and compute the joint probability from the cosine similarities of the visual-textual embeddings. Predictions of each task can be inferred from the joint distribution, and optimized by carefully designed loss functions. Through comprehensive experiments on learning three tasks - BIQA, scene classification, and distortion type identification, we verify that the proposed BIQA method 1) benefits from the scene classification and distortion type identification tasks and outperforms the state-of-the-art on multiple IQA datasets, 2) is more robust in the group maximum differentiation competition, and 3) realigns the quality annotations from different IQA datasets more effectively. The source code is available at https://github.com/zwx8981/LIQE. Weixia Zhang, Guangtao Zhai, Ying Wei 0001, Xiaokang Yang 0001, Kede Ma |
CVPR | 5 |
| 2023 | Troubleshooting image segmentation models with human-in-the-loop
Haotao Wang, Tianlong Chen 0001, Zhangyang Wang, Kede Ma |
Mach. Learn. | 4 |
| 2023 | Measuring Perceptual Color Differences of Smartphone PhotographsabstractMeasuring perceptual color differences (CDs) is of great importance in modern smartphone photography. Despite the long history, most CD measures have been constrained by psychophysical data of homogeneous color patches or a limited number of simplistic natural photographic images. It is thus questionable whether existing CD measures generalize in the age of smartphone photography characterized by greater content complexities and learning-based image signal processors. In this article, we put together so far the largest image dataset for perceptual CD assessment, in which the photographic images are 1) captured by six flagship smartphones, 2) altered by Photoshop, 3) post-processed by built-in filters of the smartphones, and 4) reproduced with incorrect color profiles. We then conduct a large-scale psychophysical experiment to gather perceptual CDs of 30,000 image pairs in a carefully controlled laboratory environment. Based on the newly established dataset, we make one of the first attempts to construct an end-to-end learnable CD formula based on a lightweight neural network, as a generalization of several previous metrics. Extensive experiments demonstrate that the optimized formula outperforms 33 existing CD measures by a large margin, offers reasonable local CD maps without the use of dense supervision, generalizes well to homogeneous color patch data, and empirically behaves as a proper metric in the mathematical sense. Our dataset and code are publicly available at https://github.com/hellooks/CDNet. Zhihua Wang 0002, Keshuo Xu, Yang Yang 0201, Jianlei Dong, Shuhang Gu, Lihao Xu, Yuming Fang 0001, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Continual Learning for Blind Image Quality AssessmentabstractThe explosive growth of image data facilitates the fast development of image processing and computer vision methods for emerging visual applications, meanwhile introducing novel distortions to processed images. This poses a grand challenge to existing blind image quality assessment (BIQA) models, which are weak at adapting to subpopulation shift. Recent work suggests training BIQA methods on the combination of all available human-rated IQA datasets. However, this type of approach is not scalable to a large number of datasets and is cumbersome to incorporate a newly created dataset as well. In this paper, we formulate continual learning for BIQA, where a model learns continually from a stream of IQA datasets, building on what was learned from previously seen data. We first identify five desiderata in the continual setting with three criteria to quantify the prediction accuracy, plasticity, and stability, respectively. We then propose a simple yet effective continual learning method for BIQA. Specifically, based on a shared backbone network, we add a prediction head for a new dataset and enforce a regularizer to allow all prediction heads to evolve with new data while being resistant to catastrophic forgetting of old data. We compute the overall quality score by a weighted summation of predictions from all heads. Extensive experiments demonstrate the promise of the proposed continual learning method in comparison to standard training techniques for BIQA, with and without experience replay. We made the code publicly available at https://github.com/zwx8981/BIQA_CL. Weixia Zhang, Dingquan Li, Chao Ma 0004, Guangtao Zhai, Xiaokang Yang 0001, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | IconShop: Text-Guided Vector Icon Synthesis with Autoregressive TransformersabstractScalable Vector Graphics (SVG) is a popular vector image format that offers good support for interactivity and animation. Despite its appealing characteristics, creating custom SVG content can be challenging for users due to the steep learning curve required to understand SVG grammars or get familiar with professional editing software. Recent advancements in text-to-image generation have inspired researchers to explore vector graphics synthesis using either image-based methods (i.e., text → raster image → vector graphics) combining text-to-image generation models with image vectorization, or language-based methods (i.e., text → vector graphics script) through pretrained large language models. Nevertheless, these methods suffer from limitations in terms of generation quality, diversity, and flexibility. In this paper, we introduce IconShop, a text-guided vector icon synthesis method using autoregressive transformers. The key to success of our approach is to sequentialize and tokenize SVG paths (and textual descriptions as guidance) into a uniquely decodable token sequence. With that, we are able to exploit the sequence learning power of autoregressive transformers, while enabling both unconditional and text-conditioned icon synthesis. Through standard training to predict the next token on a large-scale vector icon dataset accompanied by textural descriptions, the proposed IconShop consistently exhibits better icon synthesis capability than existing image-based and language-based methods both quantitatively (using the FID and CLIP scores) and qualitatively (through formal subjective user studies). Meanwhile, we observe a dramatic improvement in generation diversity, which is validated by the objective Uniqueness and Novelty measures. More importantly, we demonstrate the flexibility of IconShop with multiple novel icon synthesis tasks, including icon editing, icon interpolation, icon semantic combination, and icon design auto-suggestion. Ronghuan Wu, Wanchao Su, Kede Ma, Jing Liao 0001 |
ACM Trans. Graph. | 3 |
| 2022 | A Database of Visual Color Differences of Modern Smartphone PhotographyabstractMeasures for visual color differences (CDs) are pivotal in hardware and software upgrading of modern smartphone photography. Towards this goal, we construct currently the largest database for visual CDs of smartphone photography. Our database consists of 15, 335 natural images 1) captured by six latest flagship smartphones, 2) altered by Photoshop®, 3) post-processed by built-in filters of smartphones, and 4) reproduced with incorrect color profiles. Moreover, we conduct a large-scale psychophysical experiment to gather visual CDs of 30, 000 image pairs from 20 human subjects in a well-designed laboratory environment. Last, we apply our human-rated database to compare a total of 27 classical and recent CD metrics. We show that existing metrics are limited in assessing CDs of smartphone photography, and point out promising future directions of learning-based CD metrics. Keshuo Xu, Zhihua Wang 0002, Yang Yang 0201, Jianlei Dong, Lihao Xu, Yuming Fang 0001, Kede Ma |
ICIP | 7 |
| 2022 | Hiding Images in Deep Probabilistic ModelsabstractData hiding with deep neural networks (DNNs) has experienced impressive successes in recent years. A prevailing scheme is to train an autoencoder, consisting of an encoding network to embed (or transform) secret messages in (or into) a carrier, and a decoding network to extract the hidden messages. This scheme may suffer from several limitations regarding practicability, security, and embedding capacity. In this work, we describe a different computational framework to hide images in deep probabilistic models. Specifically, we use a DNN to model the probability density of cover images, and hide a secret image in one particular location of the learned distribution. As an instantiation, we adopt a SinGAN, a pyramid of generative adversarial networks (GANs), to learn the patch distribution of one cover image. We hide the secret image by fitting a deterministic mapping from a fixed set of noise maps (generated by an embedding key) to the secret image during patch distribution learning. The stego SinGAN, behaving as the original SinGAN, is publicly communicated; only the receiver with the embedding key is able to extract the secret image. We demonstrate the feasibility of our SinGAN approach in terms of extraction accuracy and model security. Moreover, we show the flexibility of the proposed method in terms of hiding multiple images for different receivers and obfuscating the secret image. Linqi Song, Zhenxing Qian, Xinpeng Zhang 0001, Kede Ma |
NeurIPS | 5 |
| 2022 | Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-LoopabstractNo-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimization of man-made vision systems. Here we make one of the first attempts to examine the perceptual robustness of NR-IQA models. Under a Lagrangian formulation, we identify insightful connections of the proposed perceptual attack to previous beautiful ideas in computer vision and machine learning. We test one knowledge-driven and three data-driven NR-IQA methods under four full-reference IQA models (as approximations to human perception of just-noticeable differences). Through carefully designed psychophysical experiments, we find that all four NR-IQA models are vulnerable to the proposed perceptual attack. More interestingly, we observe that the generated counterexamples are not transferable, manifesting themselves as distinct design flows of respective NR-IQA methods. Source code are available at https://github.com/zwx8981/PerceptualAttack_BIQA. Weixia Zhang, Dingquan Li, Xiongkuo Min, Guangtao Zhai, Guodong Guo, Xiaokang Yang 0001, Kede Ma |
NeurIPS | 7 |
| 2022 | Image Quality Assessment: Unifying Structure and Texture SimilarityabstractObjective measures of image quality generally operate by comparing pixels of a "degraded" image to those of the original. Relative to human observers, these measures are overly sensitive to resampling of texture regions (e.g., replacing one patch of grass with another). Here, we develop the first full-reference image quality model with explicit tolerance to texture resampling. Using a convolutional neural network, we construct an injective and differentiable function that transforms images to multi-scale overcomplete representations. We demonstrate empirically that the spatial averages of the feature maps in this representation capture texture appearance, in that they provide a set of sufficient statistical constraints to synthesize a wide variety of texture patterns. We then describe an image quality method that combines correlations of these spatial averages ("texture similarity") with correlations of the feature maps ("structure similarity"). The parameters of the proposed measure are jointly optimized to match human ratings of image quality, while minimizing the reported distances between subimages cropped from the same texture images. Experiments show that the optimized method explains human perceptual scores, both on conventional image quality databases, as well as on texture databases. The measure also offers competitive performance on related tasks such as texture classification and retrieval. Finally, we show that our method is relatively insensitive to geometric transformations (e.g., translation and dilation), without use of any specialized training or data augmentation. Code is available at https://github.com/dingkeyan93/DISTS. Keyan Ding, Kede Ma, Shiqi Wang 0001, Eero P. Simoncelli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Active Fine-Tuning From gMAD Examples Improves Blind Image Quality AssessmentabstractThe research in image quality assessment (IQA) has a long history, and significant progress has been made by leveraging recent advances in deep neural networks (DNNs). Despite high correlation numbers on existing IQA datasets, DNN-based models may be easily falsified in the group maximum differentiation (gMAD) competition. Here we show that gMAD examples can be used to improve blind IQA (BIQA) methods. Specifically, we first pre-train a DNN-based BIQA model using multiple noisy annotators, and fine-tune it on multiple synthetically distorted images, resulting in a "top-performing" baseline model. We then seek pairs of images by comparing the baseline model with a set of full-reference IQA methods in gMAD. The spotted gMAD examples are most likely to reveal the weaknesses of the baseline, and suggest potential ways for refinement. We query human quality annotations for the selected images in a well-controlled laboratory environment, and further fine-tune the baseline on the combination of human-rated images from gMAD and existing databases. This process may be iterated, enabling active fine-tuning from gMAD examples for BIQA. We demonstrate the feasibility of our active learning scheme on a large-scale unlabeled image set, and show that the fine-tuned quality model achieves improved generalizability in gMAD, without destroying performance on previously seen databases. Zhihua Wang 0002, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Perceptual Quality Assessment of Omnidirectional Images as Moving Camera VideosabstractOmnidirectional images (also referred to as static 360$^{\circ }$panoramas) impose viewing conditions much different from those of regular 2D images. How do humans perceive image distortions in immersive virtual reality (VR) environments is an important problem which receives less attention. We argue that, apart from the distorted panorama itself, two types of VR viewing conditions are crucial in determining the viewing behaviors of users and the perceived quality of the panorama: the starting point and the exploration time. We first carry out a psychophysical experiment to investigate the interplay among the VR viewing conditions, the user viewing behaviors, and the perceived quality of 360$^{\circ }$images. Then, we provide a thorough analysis of the collected human data, leading to several interesting findings. Moreover, we propose a computational framework for objective quality assessment of 360$^{\circ }$images, embodying viewing conditions and behaviors in a delightful way. Specifically, we first transform an omnidirectional image to several video representations using different user viewing behaviors under different viewing conditions. We then leverage advanced 2D full-reference video quality models to compute the perceived quality. We construct a set of specific quality measures within the proposed framework, and demonstrate their promises on three VR quality databases. Xiangjie Sui, Kede Ma, Yiru Yao, Yuming Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Debiased Subjective Assessment of Real-World Image EnhancementabstractIn real-world image enhancement, it is often challenging (if not impossible) to acquire ground-truth data, preventing the adoption of distance metrics for objective quality assessment. As a result, one often resorts to subjective quality assessment, the most straightforward and reliable means of evaluating image enhancement. Conventional subjective testing requires manually pre-selecting a small set of visual examples, which may suffer from three sources of biases: 1) sampling bias due to the extremely sparse distribution of the selected samples in the image space; 2) algorithmic bias due to potential overfitting the selected samples; 3) subjective bias due to further potential cherry-picking test results. This eventually makes the field of real-world image enhancement more of an art than a science. Here we take steps towards debiasing conventional subjective assessment by automatically sampling a set of adaptive and diverse images for subsequent testing. This is achieved by casting sample selection into a joint maximization of the discrepancy between the enhancers and the diversity among the selected input images. Careful visual inspection on the resulting enhanced images provides a debiased ranking of the enhancement algorithms. We demonstrate our subjective assessment method using three popular and practically demanding image enhancement tasks: dehazing, super-resolution, and low-light enhancement. Peibei Cao, Zhangyang Wang, Kede Ma |
CVPR | 3 |
| 2021 | Troubleshooting Blind Image Quality Models in the WildabstractRecently, the group maximum differentiation competition (gMAD) has been used to improve blind image quality assessment (BIQA) models, with the help of full-reference metrics. When applying this type of approach to troubleshoot "best-performing" BIQA models in the wild, we are faced with a practical challenge: it is highly nontrivial to obtain stronger competing models for efficient failure-spotting. Inspired by recent findings that difficult samples of deep models may be exposed through network pruning, we construct a set of "self-competitors," as random ensembles of pruned versions of the target model to be improved. Diverse failures can then be efficiently identified via self-gMAD competition. Next, we fine-tune both the target and its pruned variants on the human-rated gMAD set. This allows all models to learn from their respective failures, preparing themselves for the next round of self-gMAD competition. Experimental results demonstrate that our method efficiently troubleshoots BIQA models in the wild with improved generalizability. Zhihua Wang 0002, Haotao Wang, Tianlong Chen 0001, Zhangyang Wang, Kede Ma |
CVPR | 5 |
| 2021 | Locally Adaptive Structure and Texture Similarity for Image Quality AssessmentabstractThe latest advances in full-reference image quality assessment (IQA) involve unifying structure and texture similarity based on deep representations. The resulting Deep Image Structure and Texture Similarity (DISTS) metric, however, makes rather global quality measurements, ignoring the fact that natural photographic images are locally structured and textured across space and scale. In this paper, we describe a locally adaptive structure and texture similarity index for full-reference IQA, which we term A-DISTS. Specifically, we rely on a single statistical feature, namely the dispersion index, to localize texture regions at different scales. The estimated probability (of one patch being texture) is in turn used to adaptively pool local structure and texture measurements. The resulting A-DISTS is adapted to local image content, and is free of expensive human perceptual scores for supervised training. We demonstrate the advantages of A-DISTS in terms of correlation with human data on ten IQA databases and optimization of single image super-resolution methods. Keyan Ding, Xueyi Zou, Shiqi Wang 0001, Kede Ma |
ACM Multimedia | 5 |
| 2021 | Image Quality Assessment in the Modern AgeabstractThis tutorial provides the audience with the basic theories, methodologies, and current progresses of image quality assessment (IQA). From an actionable perspective, we will first revisit several subjective quality assessment methodologies, with emphasis on how to properly select visual stimuli. We will then present in detail the design principles of objective quality assessment models, supplemented by an in-depth analysis of their advantages and disadvantages. Both hand-engineered and (deep) learning-based methods will be covered. Moreover, the limitations with the conventional model comparison methodology for objective quality models will be pointed out, and novel comparison methodologies such as those based on the theory of "analysis by synthesis" will be introduced. We will last discuss the real-world multimedia applications of IQA, and give a list of open challenging problems, in the hope of encouraging more and more talented researchers and engineers devoting to this exciting and rewarding research field. Kede Ma, Yuming Fang 0001 |
ACM Multimedia | 1 |
| 2021 | Comparison of Full-Reference Image Quality Models for Optimization of Image Processing Systems
Keyan Ding, Kede Ma, Shiqi Wang 0001, Eero P. Simoncelli |
Int. J. Comput. Vis. | 2 |
| 2021 | Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
Jiebin Yan, Yuming Fang 0001, Zhangyang Wang, Kede Ma |
Int. J. Comput. Vis. | 5 |
| 2021 | Uncertainty-Aware Blind Image Quality Assessment in the Laboratory and WildabstractPerformance of blind image quality assessment (BIQA) models has been significantly boosted by end-to-end optimization of feature engineering and quality regression. Nevertheless, due to the distributional shift between images simulated in the laboratory and captured in the wild, models trained on databases with synthetic distortions remain particularly weak at handling realistic distortions (and vice versa). To confront the cross-distortion-scenario challenge, we develop a unified BIQA model and an approach of training it for both synthetic and realistic distortions. We first sample pairs of images from individual IQA databases, and compute a probability that the first image of each pair is of higher quality. We then employ the fidelity loss to optimize a deep neural network for BIQA over a large number of such image pairs. We also explicitly enforce a hinge constraint to regularize uncertainty estimation during optimization. Extensive experiments on six IQA databases show the promise of the learned method in blindly assessing image quality in the laboratory and wild. In addition, we demonstrate the universality of the proposed training strategy by using it to improve existing BIQA models. Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Perceptual Quality Assessment of Smartphone PhotographyabstractAs smartphones become people's primary cameras to take photos, the quality of their cameras and the associated computational photography modules has become a de facto standard in evaluating and ranking smartphones in the consumer market. We conduct so far the most comprehensive study of perceptual quality assessment of smartphone photography. We introduce the Smartphone Photography Attribute and Quality (SPAQ) database, consisting of 11,125 pictures taken by 66 smartphones, where each image is attached with so far the richest annotations. Specifically, we collect a series of human opinions for each image, including image quality, image attributes (brightness, colorfulness, contrast, noisiness, and sharpness), and scene category labels (animal, cityscape, human, indoor scene, landscape, night scene, plant, still life, and others) in a well-controlled laboratory environment. The exchangeable image file format (EXIF) data for all images are also recorded to aid deeper analysis. We also make the first attempts using the database to train blind image quality assessment (BIQA) models constructed by baseline and multi-task deep neural networks. The results provide useful insights on how EXIF data, image attributes and high-level semantics interact with image quality, how next-generation BIQA models can be designed, and how better computational photography systems can be optimized on mobile devices. The database along with the proposed BIQA models are available at https://github.com/h4nwei/SPAQ. Yuming Fang 0001, Hanwei Zhu, Yan Zeng 0001, Kede Ma, Zhou Wang 0001 |
CVPR | 4 |
| 2020 | Learning To Blindly Assess Image Quality In The Laboratory And WildabstractComputational models for blind image quality assessment (BIQA) are typically trained in well-controlled laboratory environments with limited generalizability to realistically distorted images. Similarly, BIQA models optimized for images captured in the wild cannot adequately handle synthetically distorted images. To face the cross-distortion-scenario challenge, we develop a BIQA model and an approach of training it on multiple IQA databases (of different distortion scenarios) simultaneously. A key step in our approach is to create and combine image pairs within individual databases as the training set, which effectively bypasses the issue of perceptual scale realignment. We compute a continuous quality annotation for each pair from the corresponding human opinions, indicating the probability of one image having better perceptual quality. We train a deep neural network for BIQA over the training set of massive image pairs by minimizing the fidelity loss. Experiments on six IQA databases demonstrate that the optimized model by the proposed training strategy is effective in blindly assessing image quality in the laboratory and wild, outperforming previous BIQA methods by a large margin. Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001 |
ICIP | 2 |
| 2020 | I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifiers Adaptively
Haotao Wang, Tianlong Chen 0001, Zhangyang Wang, Kede Ma |
ICLR | 4 |
| 2020 | Group Maximum Differentiation Competition: Model Comparison with Few SamplesabstractIn many science and engineering fields that require computational models to predict certain physical quantities, we are often faced with the selection of the best model under the constraint that only a small sample set can be physically measured. One such example is the prediction of human perception of visual quality, where sample images live in a high dimensional space with enormous content variations. We propose a new methodology for model comparison named group maximum differentiation (gMAD) competition. Given multiple computational models, gMAD maximizes the chances of falsifying a "defender" model using the rest models as "attackers". It exploits the sample space to find sample pairs that maximally differentiate the attackers while holding the defender fixed. Based on the results of the attacking-defending game, we introduce two measures, aggressiveness and resistance, to summarize the performance of each model at attacking other models and defending attacks from other models, respectively. We demonstrate the gMAD competition using three examples-image quality, image aesthetics, and streaming video quality-of-experience. Although these examples focus on visually discriminable quantities, the gMAD methodology can be extended to many other fields, and is especially useful when the sample space is large, the physical measurement is expensive and the cost of computational prediction is low. Kede Ma, Zhengfang Duanmu, Zhou Wang 0001, Qingbo Wu 0001, Wentao Liu 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Blind Image Quality Assessment Using a Deep Bilinear Convolutional Neural NetworkabstractWe propose a deep bilinear model for blind image quality assessment that works for both synthetically and authentically distorted images. Our model constitutes two streams of deep convolutional neural networks (CNNs), specializing in two distortion scenarios separately. For synthetic distortions, we first pre-train a CNN to classify the distortion type and the level of an input image, whose ground truth label is readily available at a large scale. For authentic distortions, we make use of a pre-train CNN (VGG-16) for the image classification task. The two feature sets are bilinearly pooled into one representation for a final quality prediction. We fine-tune the whole network on the target databases using a variant of stochastic gradient descent. The extensive experimental results show that the proposed model achieves state-of-the-art performance on both synthetic and authentic IQA databases. Furthermore, we verify the generalizability of our method on the large-scale Waterloo Exploration Database, and demonstrate its competitiveness using the group maximum differentiation competition methodology. Weixia Zhang, Kede Ma, Jia Yan 0006, Dexiang Deng, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Characterizing Generalized Rate-Distortion Performance of Video Coding: An Eigen Analysis ApproachabstractRate-distortion (RD) theory is at the heart of lossy data compression. Here we aim to model the generalized RD (GRD) trade-off between the visual quality of a compressed video and its encoding profiles (e.g., bitrate and spatial resolution). We first define the theoretical functional space W of the GRD function by analyzing its mathematical properties. We show that W is a convex set in a Hilbert space, inspiring a computational model of the GRD function, and a method of estimating model parameters from sparse measurements. To demonstrate the feasibility of our idea, we collect a large-scale database of real-world GRD functions, which turn out to live in a low-dimensional subspace of W. Combining the GRD reconstruction framework and the learned low-dimensional space, we create a low-parameter eigen GRD method to accurately estimate the GRD function of a source video content from only a few queries. Experimental results on the database show that the learned GRD method significantly outperforms state-of-the-art empirical RD estimation methods both in accuracy and efficiency. Last, we demonstrate the promise of the proposed model in video codec comparison. Zhengfang Duanmu, Wentao Liu 0001, Kede Ma, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Perceptual Evaluation for Multi-Exposure Image Fusion of Dynamic ScenesabstractA common approach to high dynamic range (HDR) imaging is to capture multiple images of different exposures followed by multi-exposure image fusion (MEF) in either radiance or intensity domain. A predominant problem of this approach is the introduction of the ghosting artifacts in dynamic scenes with camera and object motion. While many MEF methods (often referred to as deghosting algorithms) have been proposed for reduced ghosting artifacts and improved visual quality, little work has been dedicated to perceptual evaluation of their deghosting results. Here we first construct a database that contains 20 multiexposure sequences of dynamic scenes and their corresponding fused images by nine MEF algorithms. We then carry out a subjective experiment to evaluate fused image quality, and find that none of existing objective quality models for MEF provides accurate quality predictions. Motivated by this, we develop an objective quality model for MEF of dynamic scenes. Specifically, we divide the test image into static and dynamic regions, measure structural similarity between the image and the corresponding sequence in the two regions separately, and combine quality measurements of the two regions into an overall quality score. Experimental results show that the proposed method significantly outperforms the state-of-the-art. In addition, we demonstrate the promise of the proposed model in parameter tuning of MEF methods.1. Yuming Fang 0001, Hanwei Zhu, Kede Ma, Zhou Wang 0001, Shutao Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Fast Multi-Scale Structural Patch Decomposition for Multi-Exposure Image FusionabstractExposure bracketing is crucial to high dynamic range imaging, but it is prone to halos for static scenes and ghosting artifacts for dynamic scenes. The recently proposed structural patch decomposition for multi-exposure fusion (SPD-MEF) has achieved reliable performance in deghosting, but suffers from visible halo artifacts and is computationally expensive. In addition, its relationship to other MEF methods is unclear. We show that without explicitly performing structural patch decomposition, we arrive at an unnormalized version of SPD-MEF, which enjoys an order of 30× speed-up, and is closely related to pixel-level MEF methods as well as the standard two-layer decomposition method for MEF. Moreover, we develop a fast multi-scale SPD-MEF method, which can effectively reduce halo artifacts. Experimental results demonstrate the effectiveness of the proposed MEF method in terms of speed and quality. Hui Li 0029, Kede Ma, Hongwei Yong, Lei Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2020 | Efficient and Effective Context-Based Convolutional Entropy Modeling for Image CompressionabstractPrecise estimation of the probabilistic structure of natural images plays an essential role in image compression. Despite the recent remarkable success of end-to-end optimized image compression, the latent codes are usually assumed to be fully statistically factorized in order to simplify entropy modeling. However, this assumption generally does not hold true and may hinder compression performance. Here we present contextbased convolutional networks (CCNs) for efficient and effective entropy modeling. In particular, a 3D zigzag scanning order and a 3D code dividing technique are introduced to define proper coding contexts for parallel entropy decoding, both of which boil down to place translation-invariant binary masks on convolution filters of CCNs. We demonstrate the promise of CCNs for entropy modeling in both lossless and lossy image compression. For the former, we directly apply a CCN to the binarized representation of an image to compute the Bernoulli distribution of each code for entropy estimation. For the latter, the categorical distribution of each code is represented by a discretized mixture of Gaussian distributions, whose parameters are estimated by three CCNs. We then jointly optimize the CCNbased entropy model along with analysis and synthesis transforms for rate-distortion performance. Experiments on the Kodak and Tecnick datasets show that our methods powered by the proposed CCNs generally achieve comparable compression performance to the state-of-the-art while being much faster. Mu Li 0005, Kede Ma, Jane You, David Zhang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2020 | Deep Guided Learning for Fast Multi-Exposure Image FusionabstractWe propose a fast multi-exposure image fusion (MEF) method, namely MEF-Net, for static image sequences of arbitrary spatial resolution and exposure number. We first feed a low-resolution version of the input sequence to a fully convolutional network for weight map prediction. We then jointly upsample the weight maps using a guided filter. The final image is computed by a weighted fusion. Unlike conventional MEF methods, MEF-Net is trained end-to-end by optimizing the perceptually calibrated MEF structural similarity (MEF-SSIM) index over a database of training sequences at full resolution. Across an independent set of test sequences, we find that the optimized MEF-Net achieves consistent improvement in visual quality for most sequences, and runs 10 to 1000 times faster than state-of-the-art methods. The code is made publicly available at. Kede Ma, Zhengfang Duanmu, Hanwei Zhu, Yuming Fang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Blind Image Quality Assessment by Learning from Multiple AnnotatorsabstractModels for image quality assessment (IQA) are generally optimized and tested by comparing to human ratings, which are expensive to obtain. Here, we develop a blind IQA (BIQA) model, and a method of training it without human ratings. We first generate a large number of corrupted image pairs, and use a set of existing IQA models to identify which image of each pair has higher quality. We then train a convolutional neural network to estimate perceived image quality along with the uncertainty, optimizing for consistency with the binary labels. The reliability of each IQA annotator is also estimated during training. Experiments demonstrate that our model outperforms state-of-the-art BIQA models in terms of correlation with human ratings in existing databases, as well in group maximum differentiation (gMAD) competition. Kede Ma, Xuelin Liu, Yuming Fang 0001, Eero P. Simoncelli |
ICIP | 1 |
| 2019 | Intrinsic Image Popularity AssessmentabstractThe goal of research in automatic image popularity assessment (IPA) is to develop computational models that can accurately predict the potential of a social image to go viral on the Internet. Here, we aim to single out the contribution of visual content to image popularity, \ie, intrinsic image popularity. Specifically, we first describe a probabilistic method to generate massive popularity-discriminable image pairs, based on which the first large-scale image database for intrinsic IPA (I$^2$PA) is established. We then develop computational models for I$^2$PA based on deep neural networks, optimizing for ranking consistency with millions of popularity-discriminable image pairs. Experiments on Instagram and other social platforms demonstrate that the optimized model performs favorably against existing methods, exhibits reasonable generalizability on different databases, and even surpasses human-level performance on Instagram. In addition, we conduct a psychophysical experiment to analyze various aspects of human behavior in I$^2$PA. Keyan Ding, Kede Ma, Shiqi Wang 0001 |
ACM Multimedia | 2 |
| 2018 | Geometric Transformation Invariant Image Quality Assessment Using Convolutional Neural NetworksabstractMost existing full-reference (FR) image quality assessment (IQA) models assume that the reference and distorted images are perfectly aligned, and fail dramatically when the assumption does not hold. In this study, we first show that pre-registration, especially feature-based (as opposed to area-based) registration, is effective at reducing the performance drop of FR-IQA models. However, registration is an expensive process that often slows down the speed of the IQA algorithms by several orders of magnitude. This motivates us to construct an end-to-end convolutional neural network (CNN) for direct image quality prediction, which contains built-in invariance to geometric distortions. Our results show that when the training images are augmented by their geometrically transformed versions, the learned network performs at a high level without image registration, resulting in a fast and effective approach for geometric transformation invariant IQA. Kede Ma, Zhengfang Duanmu, Zhou Wang 0001 |
ICASSP | 1 |
| 2018 | A Hybrid Quality Metric for Non-Integer Image InterpolationabstractA great need of High-Resolution (HR) images has boosted the development of interpolation techniques. However, it is still a challenging task to objectively evaluate the perceptual quality of interpolated images, especially when the interpolation factor is a non-integer. To address this issue, we propose a hybrid quality metric for non-integer image interpolation that combines both reduced-reference and no-reference philosophies. To validate the proposed metric, we construct a non-integer interpolated image database and conduct a subjective user study to collect subjective opinions for each image. Experiments on the new database show that the proposed metric outperforms previous methods by a large margin. Jinling Chen, Kede Ma, Huiwen Huang, Tiesong Zhao |
QoMEX | 3 |
| 2018 | Blind Image Quality Assessment Using Local Consistency Aware Retriever and Uncertainty Aware EvaluatorabstractBlind image quality assessment (BIQA) aims to automatically predict the perceptual quality of a digital image without accessing its pristine reference. Previous studies mainly focus on extracting various quality-relevant image features. By contrast, the explorations on highly efficient learning model are still very limited. Motivated by the fact that it is difficult to approximate a complex and large data set via a global parametric model, we propose a novel local learning method for BIQA to improve quality prediction performance. More specifically, we search for the perceptually similar neighbors of a test image to serve as its unique training set. Unlike the widely used k nearest neighbors principle, which only measures the similarity between the testing and training samples, the local consistency of the selected training data is also considered to generate smoother sample space. The image quality is estimated via a sparse Gaussian process. As an additional benefit, the uncertainty of the predicted score is jointly inferred, which can subsequently drive more robust perceptual image processing applications, such as deblocking investigated in this paper. Extensive experiments demonstrate that the proposed learning model leads to consistent quality prediction improvements over many state-of-the-art BIQA algorithms. Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Quality-of-Experience for Adaptive Streaming Videos: An Expectation Confirmation Theory Motivated ApproachabstractThe dynamic adaptive streaming over HTTP (DASH) provides an inter-operable solution to overcome volatile network conditions, but how the human visual quality-ofexperience (QoE) changes with time-varying video quality is not well-understood. Here, we build a large-scale video database of time-varying quality and design a series of subjective experiments to investigate how humans respond to compression level, spatial and temporal resolution adaptations. Our path-analytic results show that quality adaptations influence the QoE by modifying the perceived quality of subsequent video segments. Specifically, the quality deviation introduced by quality adaptations is asymmetric with respect to the adaptation direction, which is further influenced by other factors such as compression level and content. Furthermore, we propose an objective QoE model by integrating the empirical findings from our subjective experiments and the expectation confirmation theory (ECT). Experimental results show that the proposed ECT-QoE model is in close agreement with subjective opinions and significantly outperforms existing QoE models. The video database together with the code are available online at https://ece.uwaterloo.ca/~zduanmu/tip2018ectqoe/. Zhengfang Duanmu, Kede Ma, Zhou Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Deep Blur Mapping: Exploiting High-Level Semantics by Deep Neural NetworksabstractThe human visual system excels at detecting local blur of visual images, but the underlying mechanism is not well understood. Traditional views of blur such as reduction in energy at high frequencies and loss of phase coherence at localized features have fundamental limitations. For example, they cannot well discriminate flat regions from blurred ones. Here we propose that high-level semantic information is critical in successfully identifying local blur. Therefore, we resort to deep neural networks that are proficient at learning high-level features and propose the first end-to-end local blur mapping algorithm based on a fully convolutional network. By analyzing various architectures with different depths and design philosophies, we empirically show that high-level features of deeper layers play a more important role than low-level features of shallower layers in resolving challenging ambiguities for this task. We test the proposed method on a standard blur detection benchmark and demonstrate that it significantly advances the state-of-the-art (ODS F-score of 0.853). Furthermore, we explore the use of the generated blur maps in three applications, including blur region segmentation, blur degree estimation, and blur magnification. Kede Ma, Huan Fu, Tongliang Liu, Zhou Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2018 | End-to-End Blind Image Quality Assessment Using Deep Neural NetworksabstractWe propose a multi-task end-to-end optimized deep neural network (MEON) for blind image quality assessment (BIQA). MEON consists of two sub-networks-a distortion identification network and a quality prediction network-sharing the early layers. Unlike traditional methods used for training multi-task networks, our training process is performed in two steps. In the first step, we train a distortion type identification sub-network, for which large-scale training samples are readily available. In the second step, starting from the pre-trained early layers and the outputs of the first sub-network, we train a quality prediction sub-network using a variant of the stochastic gradient descent method. Different from most deep neural networks, we choose biologically inspired generalized divisive normalization (GDN) instead of rectified linear unit as the activation function. We empirically demonstrate that GDN is effective at reducing model parameters/layers while achieving similar quality prediction performance. With modest model complexity, the proposed MEON index achieves state-of-the-art performance on four publicly available benchmarks. Moreover, we demonstrate the strong competitiveness of MEON against state-of-the-art BIQA models using the group maximum differentiation competition methodology. Kede Ma, Wentao Liu 0001, Kai Zhang 0008, Zhengfang Duanmu, Zhou Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2017 | Perceptual quality assessment of HDR deghosting algorithmsabstractHigh dynamic range (HDR) imaging techniques aim to extend the dynamic range of images that cannot be well captured using conventional camera sensors. A common practice is to take a stack of pictures with different exposure levels and fuse them to produce a final image with more details. However, a small displacement between images caused by either camera or scene motion would void the benefits and cause the so-called ghosting artifacts. Over the past decade, many HDR deghosting algorithms have been proposed, but little work has been dedicated to evaluate HDR deghosting results either subjectively or objectively. In this work, we present a comprehensive subjective study for HDR deghosting. Specifically, we create a database that contains 20 dynamic image sequences and their corresponding deghosting results by 9 deghosting algorithms. A subjective user study is then carried out to evaluate the perceptual quality of deghosted images. The experimental results demonstrate the performance and limitations of existing HDR deghosting algorithm as well as no-reference image quality assessment models. In the future, we will make the database available to the public. Yuming Fang 0001, Hanwei Zhu, Kede Ma, Zhou Wang 0001 |
ICIP | 3 |
| 2017 | Quality-of-Experience of Adaptive Video Streaming: Exploring the Space of AdaptationsabstractWith the remarkable growth of adaptive streaming media applications, especially the wide usage of dynamic adaptive streaming schemes over HTTP (DASH), it becomes ever more important to understand the perceptual quality-of-experience (QoE) of end users, who may be constantly experiencing adaptations (switchings) of video bitrate, spatial resolution, and frame-rate from one time segment to another in a scale of a few seconds. This is a sophisticated and challenging problem, for which existing visual studies provide very limited guidance. Here we build a new adaptive streaming video database and carry out a series of subjective experiments to understand human QoE behaviors in this multi-dimensional adaptation space. Our study leads to several useful findings. First, our path-analytic results show that quality deviation introduced by quality adaptation is asymmetric with respect to the adaptation direction (positive or negative), and is further influenced by the intensity of quality change (intensity), dimension of adaptation (type), intrinsic video quality (level), content, and the interactions between them. Second, we find that for the same intensity of quality adaptation, a positive adaptation occurred in the low-quality range has more impact on QoE, suggesting an interesting Weber's law effect; while such phenomenon is reversed for a negative adaptation. Third, existing objective video quality assessment models are very limited in predicting time-varying video quality. Zhengfang Duanmu, Kede Ma, Zhou Wang 0001 |
ACM Multimedia | 2 |
| 2017 | Waterloo Exploration Database: New Challenges for Image Quality Assessment ModelsabstractThe great content diversity of real-world digital images poses a grand challenge to image quality assessment (IQA) models, which are traditionally designed and validated on a handful of commonly used IQA databases with very limited content variation. To test the generalization capability and to facilitate the wide usage of IQA techniques in real-world applications, we establish a large-scale database named the Waterloo Exploration Database, which in its current state contains 4744 pristine natural images and 94 880 distorted images created from them. Instead of collecting the mean opinion score for each image via subjective testing, which is extremely difficult if not impossible, we present three alternative test criteria to evaluate the performance of IQA models, namely, the pristine/distorted image discriminability test, the listwise ranking consistency test, and the pairwise preference consistency test (P-test). We compare 20 well-known IQA models using the proposed criteria, which not only provide a stronger test in a more challenging testing environment for existing models, but also demonstrate the additional benefits of using the proposed database. For example, in the P-test, even for the best performing no-reference IQA model, more than 6 million failure cases against the model are "discovered" automatically out of over 1 billion test pairs. Furthermore, we discuss how the new database may be exploited using innovative approaches in the future, to reveal the weaknesses of existing IQA models, to provide insights on how to improve the models, and to shed light on how the next-generation IQA models may be developed. The database and codes are made publicly available at: https://ece.uwaterloo.ca/~k29ma/exploration/. Kede Ma, Zhengfang Duanmu, Qingbo Wu 0001, Zhou Wang 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2017 | dipIQ: Blind Image Quality Assessment by Learning-to-Rank Discriminable Image PairsabstractObjective assessment of image quality is fundamentally important in many image processing tasks. In this paper, we focus on learning blind image quality assessment (BIQA) models, which predict the quality of a digital image with no access to its original pristine-quality counterpart as reference. One of the biggest challenges in learning BIQA models is the conflict between the gigantic image space (which is in the dimension of the number of image pixels) and the extremely limited reliable ground truth data for training. Such data are typically collected via subjective testing, which is cumbersome, slow, and expensive. Here, we first show that a vast amount of reliable training data in the form of quality-discriminable image pairs (DIPs) can be obtained automatically at low cost by exploiting large-scale databases with diverse image content. We then learn an opinion-unaware BIQA (OU-BIQA, meaning that no subjective opinions are used for training) model using RankNet, a pairwise learning-to-rank (L2R) algorithm, from millions of DIPs, each associated with a perceptual uncertainty level, leading to a DIP inferred quality (dipIQ) index. Extensive experiments on four benchmark IQA databases demonstrate that dipIQ outperforms the state-of-the-art OU-BIQA models. The robustness of dipIQ is also significantly improved as confirmed by the group MAximum Differentiation competition method. Furthermore, we extend the proposed framework by learning models with ListNet (a listwise L2R algorithm) on quality-discriminable image lists (DIL). The resulting DIL inferred quality index achieves an additional performance gain. Kede Ma, Wentao Liu 0001, Tongliang Liu, Zhou Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2017 | Robust Multi-Exposure Image Fusion: A Structural Patch Decomposition ApproachabstractWe propose a simple yet effective structural patch decomposition approach for multi-exposure image fusion (MEF) that is robust to ghosting effect. We decompose an image patch into three conceptually independent components: signal strength, signal structure, and mean intensity. Upon fusing these three components separately, we reconstruct a desired patch and place it back into the fused image. This novel patch decomposition approach benefits MEF in many aspects. First, as opposed to most pixel-wise MEF methods, the proposed algorithm does not require post-processing steps to improve visual quality or to reduce spatial artifacts. Second, it handles RGB color channels jointly, and thus produces fused images with more vivid color appearance. Third and most importantly, the direction of the signal structure component in the patch vector space provides ideal information for ghost removal. It allows us to reliably and efficiently reject inconsistent object motions with respect to a chosen reference image without performing computationally expensive motion estimation. We compare the proposed algorithm with 12 MEF methods on 21 static scenes and 12 deghosting schemes on 19 dynamic scenes (with camera and object motion). Extensive experimental results demonstrate that the proposed algorithm not only outperforms previous MEF algorithms on static scenes but also consistently produces high quality fused images with little ghosting artifacts for dynamic scenes. Moreover, it maintains a lower computational cost compared with the state-of-the-art deghosting schemes. Kede Ma, Hui Li 0029, Hongwei Yong, Zhou Wang 0001, Deyu Meng, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2017 | Unified Blind Quality Assessment of Compressed Natural, Graphic, and Screen Content ImagesabstractDigital images in the real world are created by a variety of means and have diverse properties. A photographical natural scene image (NSI) may exhibit substantially different characteristics from a computer graphic image (CGI) or a screen content image (SCI). This casts major challenges to objective image quality assessment, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. To tackle this problem, we first construct a cross-content-type (CCT) database, which contains 1,320 distorted NSIs, CGIs, and SCIs, compressed using the high efficiency video coding (HEVC) intra coding method and the screen content compression (SCC) extension of HEVC. We then carry out a subjective experiment on the database in a well-controlled laboratory environment. Moreover, we propose a unified content-type adaptive (UCA) blind image quality assessment model that is applicable across content types. A key step in UCA is to incorporate the variations of human perceptual characteristics in viewing different content types through a multi-scale weighting framework. This leads to superior performance on the constructed CCT database. UCA is training-free, implying strong generalizability. To verify this, we test UCA on other databases containing JPEG, MPEG-2, H.264, and HEVC compressed images/videos, and observe that it consistently achieves competitive performance. Xiongkuo Min, Kede Ma, Ke Gu 0001, Guangtao Zhai, Zhou Wang 0001, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2017 | Perceptual Depth Quality in Distorted Stereoscopic ImagesabstractSubjective and objective measurement of the perceptual quality of depth information in symmetrically and asymmetrically distorted stereoscopic images is a fundamentally important issue in stereoscopic 3D imaging that has not been deeply investigated. Here, we first carry out a subjective test following the traditional absolute category rating protocol widely used in general image quality assessment research. We find this approach problematic, because monocular cues and the spatial quality of images have strong impact on the depth quality scores given by subjects, making it difficult to single out the actual contributions of stereoscopic cues in depth perception. To overcome this problem, we carry out a novel subjective study where depth effect is synthesized at different depth levels before various types and levels of symmetric and asymmetric distortions are applied. Instead of following the traditional approach, we ask subjects to identify and label depth polarizations, and a depth perception difficulty index (DPDI) is developed based on the percentage of correct and incorrect subject judgements. We find this approach highly effective at quantifying depth perception induced by stereo cues and observe a number of interesting effects regarding image content dependency, distortion-type dependence, and the impact of symmetric versus asymmetric distortions. Furthermore, we propose a novel computational model for DPDI prediction. Our results show that the proposed model, without explicitly identifying image distortion types, leads to highly promising DPDI prediction performance. We believe that these are useful steps toward building a comprehensive understanding on 3D quality-of-experience of stereoscopic images. Jiheng Wang, Shiqi Wang 0001, Kede Ma, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Group MAD Competition? A New Methodology to Compare Objective Image Quality ModelsabstractObjective image quality assessment (IQA) models aim to automatically predict human visual perception of image quality and are of fundamental importance in the field of image processing and computer vision. With an increasing number of IQA models proposed, how to fairly compare their performance becomes a major challenge due to the enormous size of image space and the limited resource for subjective testing. The standard approach in literature is to compute several correlation metrics between subjective mean opinion scores (MOSs) and objective model predictions on several well-known subject-rated databases that contain distorted images generated from a few dozens of source images, which however provide an extremely limited representation of real-world images. Moreover, most IQA models developed on these databases often involve machine learning and/or manual parameter tuning steps to boost their performance, and thus their generalization capabilities are questionable. Here we propose a novel methodology to compare IQA models. We first build a database that contains 4,744 source natural images, together with 94,880 distorted images created from them. We then propose a new mechanism, namely group MAximum Differentiation (gMAD) competition, which automatically selects subsets of image pairs from the database that provide the strongest test to let the IQA models compete with each other. Subjective testing on the selected subsets reveals the relative performance of the IQA models and provides useful insights on potential ways to improve them. We report the gMAD competition results between 16 well-known IQA models, but the framework is extendable, allowing future IQA models to be added into the competition. Kede Ma, Qingbo Wu 0001, Zhou Wang 0001, Zhengfang Duanmu, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
CVPR | 1 |
| 2015 | Perceptual evaluation of single image dehazing algorithmsabstractImages captured in outdoor scenes often suffer from poor visibility and color shift due to the presence of haze. Although many algorithms have been proposed to remove the haze, not much effort has been made on quality assessment of dehazed images. In this paper, we first build a database that contains 25 hazy images as well as dehazed images created by eight dehazing algorithms. A subjective user study is then carried out based on the database, from which we have several useful findings. First, considerable agreement between human subjects on the perceived quality of hazy and dehazed images is observed. Second, not a single dehazing algorithm performs the best for all test images. Third, existing objective image quality assessment (IQA) models are very limited in providing proper quality predictions of dehazed images. Kede Ma, Wentao Liu 0001, Zhou Wang 0001 |
ICIP | 1 |
| 2015 | Multi-exposure image fusion: A patch-wise approachabstractWe propose a patch-wise approach for multi-exposure image fusion (MEF). A key step in our approach is to decompose each color image patch into three conceptually independent components: signal strength, signal structure and mean intensity. Upon processing the three components separately based on patch strength and exposedness measures, we uniquely reconstruct a color image patch and place it back into the fused image. Unlike most pixel-wise MEF methods in the literature, the proposed algorithm does not require significant pre/postprocessing steps to improve visual quality or to reduce spatial artifacts. Moreover, the novel patch decomposition allows us to handle RGB color channels jointly and thus produces fused images with more vivid color appearances. Extensive experiments demonstrate the superiority of the proposed algorithm both qualitatively and quantitatively. Kede Ma, Zhou Wang 0001 |
ICIP | 1 |
| 2015 | No-Reference Quality Assessment of Contrast-Distorted Images Based on Natural Scene StatisticsabstractContrast distortion is often a determining factor in human perception of image quality, but little investigation has been dedicated to quality assessment of contrast-distorted images without assuming the availability of a perfect-quality reference image. In this letter, we propose a simple but effective method for no-reference quality assessment of contrast distorted images based on the principle of natural scene statistics (NSS). A large scale image database is employed to build NSS models based on moment and entropy features. The quality of a contrast-distorted image is then evaluated based on its unnaturalness characterized by the degree of deviation from the NSS models. Support vector regression (SVR) is employed to predict human mean opinion score (MOS) from multiple NSS features as the input. Experiments based on three publicly available databases demonstrate the promising performance of the proposed method. Yuming Fang 0001, Kede Ma, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001, Guangtao Zhai |
IEEE Signal Process. Lett. | 2 |
| 2015 | A Patch-Structure Representation Method for Quality Assessment of Contrast Changed ImagesabstractContrast is a fundamental attribute of images that plays an important role in human visual perception of image quality. With numerous approaches proposed to enhance image contrast, much less work has been dedicated to automatic quality assessment of contrast changed images. Existing approaches rely on global statistics to estimate contrast quality. Here we propose a novel local patch-based objective quality assessment method using an adaptive representation of local patch structure, which allows us to decompose any image patch into its mean intensity, signal strength and signal structure components and then evaluate their perceptual distortions in different ways. A unique feature that differentiates the proposed method from previous contrast quality models is the capability to produce a local contrast quality map, which predicts local quality variations over space and may be employed to guide contrast enhancement algorithms. Validations based on four publicly available databases show that the proposed patch-based contrast quality index (PCQI) method provides accurate predictions on the human perception of contrast variations. Shiqi Wang 0001, Kede Ma, Hojatollah Yeganeh, Zhou Wang 0001, Weisi Lin |
IEEE Signal Process. Lett. | 2 |
| 2015 | High Dynamic Range Image Compression by Optimizing Tone Mapped Image Quality IndexabstractTone mapping operators (TMOs) aim to compress high dynamic range (HDR) images to low dynamic range (LDR) ones so as to visualize HDR images on standard displays. Most existing TMOs were demonstrated on specific examples without being thoroughly evaluated using well-designed and subject-validated image quality assessment models. A recently proposed tone mapped image quality index (TMQI) made one of the first attempts on objective quality assessment of tone mapped images. Here, we propose a substantially different approach to design TMO. Instead of using any predefined systematic computational structure for tone mapping (such as analytic image transformations and/or explicit contrast/edge enhancement), we directly navigate in the space of all images, searching for the image that optimizes an improved TMQI. In particular, we first improve the two building blocks in TMQI—structural fidelity and statistical naturalness components—leading to a TMQI-II metric. We then propose an iterative algorithm that alternatively improves the structural fidelity and statistical naturalness of the resulting image. Numerical and subjective experiments demonstrate that the proposed algorithm consistently produces better quality tone mapped images even when the initial images of the iteration are created by the most competitive TMOs. Meanwhile, these results also validate the superiority of TMQI-II over TMQI. Kede Ma, Hojatollah Yeganeh, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Perceptual Quality Assessment for Multi-Exposure Image FusionabstractMulti-exposure image fusion (MEF) is considered an effective quality enhancement technique widely adopted in consumer electronics, but little work has been dedicated to the perceptual quality assessment of multi-exposure fused images. In this paper, we first build an MEF database and carry out a subjective user study to evaluate the quality of images generated by different MEF algorithms. There are several useful findings. First, considerable agreement has been observed among human subjects on the quality of MEF images. Second, no single state-of-the-art MEF algorithm produces the best quality for all test images. Third, the existing objective quality models for general image fusion are very limited in predicting perceived quality of MEF images. Motivated by the lack of appropriate objective models, we propose a novel objective image quality assessment (IQA) algorithm for MEF images based on the principle of the structural similarity approach and a novel measure of patch structural consistency. Our experimental results on the subjective database show that the proposed model well correlates with subjective judgments and significantly outperforms the existing IQA models for general image fusion. Finally, we demonstrate the potential application of the proposed model by automatically tuning the parameters of MEF algorithms. Kede Ma, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Objective Quality Assessment for Color-to-Gray Image ConversionabstractColor-to-gray (C2G) image conversion is the process of transforming a color image into a grayscale one. Despite its wide usage in real-world applications, little work has been dedicated to compare the performance of C2G conversion algorithms. Subjective evaluation is reliable but is also inconvenient and time consuming. Here, we make one of the first attempts to develop an objective quality model that automatically predicts the perceived quality of C2G converted images. Inspired by the philosophy of the structural similarity index, we propose a C2G structural similarity (C2G-SSIM) index, which evaluates the luminance, contrast, and structure similarities between the reference color image and the C2G converted image. The three components are then combined depending on image type to yield an overall quality measure. Experimental results show that the proposed C2G-SSIM index has close agreement with subjective rankings and significantly outperforms existing objective quality metrics for C2G conversion. To explore the potentials of C2G-SSIM, we further demonstrate its use in two applications: 1) automatic parameter tuning for C2G conversion algorithms and 2) adaptive fusion of C2G converted images. Kede Ma, Tiesong Zhao, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | High dynamic range image tone mapping by optimizing tone mapped image quality indexabstractAn active research topic in recent years is to design tone mapping operators (TMOs) that convert high dynamic range (H-DR) to low dynamic range (LDR) images, so that HDR images can be visualized on standard displays. Nevertheless, most existing work has been done in the absence of a well-established and subject-validated image quality assessment (IQA) model, without which fair comparisons and further improvement are difficult. Recently, a tone mapped image quality index (TMQI) was proposed, which has shown to have good correlation with subjective evaluations of tone mapped images. Here we propose a substantially different approach to design TMO, where instead of using any pre-defined systematic computational structure (such as image transformation or contrast/edge enhancement) for tone mapping, we navigate in the space of all images, searching for the image that optimizes TMQI. The navigation involves an iterative process that alternately improves the structural fidelity and statistical naturalness of the resulting image, which are the two fundamental building blocks in TMQI. Experiments demonstrate the superior performance of the proposed method. Kede Ma, Hojatollah Yeganeh, Kai Zeng 0003, Zhou Wang 0001 |
ICME | 1 |
| 2014 | Recursive code construction for reversible data hiding in DCT domain
Weiming Zhang 0001, Kede Ma, Nenghai Yu |
Multim. Tools Appl. | 3 |
| 2014 | Reversibility improved data hiding in encrypted images
Weiming Zhang 0001, Kede Ma, Nenghai Yu |
Signal Process. | 2 |
| 2013 | Reversible Data Hiding in Encrypted Images by Reserving Room Before EncryptionabstractRecently, more and more attention is paid to reversible data hiding (RDH) in encrypted images, since it maintains the excellent property that the original cover can be losslessly recovered after embedded data is extracted while protecting the image content's confidentiality. All previous methods embed data by reversibly vacating room from the encrypted images, which may be subject to some errors on data extraction and/or image restoration. In this paper, we propose a novel method by reserving room before encryption with a traditional RDH algorithm, and thus it is easy for the data hider to reversibly embed data in the encrypted image. The proposed method can achieve real reversibility, that is, data extraction and image recovery are free of any error. Experiments show that this novel method can embed more than 10 times as large payloads for the same image quality as the previous methods, such as for PSNR=40 dB. Kede Ma, Weiming Zhang 0001, Xianfeng Zhao, Nenghai Yu, Fenghua Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |