Qiong Yan

dblp:122/4814 · DBLP profile ↗
← Back
43ranked-venue papers
2as first author
27since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 21 since 2021Artificial intelligence and machine learning · 28 · 2 first-author · 16 since 2021
YearPublicationVenuePosition
2025 Enhancing HDR Imaging with Joint Denoising and Deblurring
abstract
Abstract Considerable progress has been made in high dynamic range (HDR) image reconstruction from multi-exposure low dynamic range (LDR) frames in recent years. Despite achieving satisfactory HDR results while handling the misalignment among LDR frames, previous studies pay less attention to a prevalent but more crucial concern: the significant noise and motion blur that exist in multi-exposure frames captured by a handheld camera. To overcome these extensive and challenging corruptions, we achieve the HDR imaging task from two key aspects: First, due to the absence of related datasets, previous learning-based methods struggle with performing high-quality HDR imaging when the input LDR frames are confronted with real-world image noise and motion blur. Recognizing the importance of this aspect, we propose the first available realistic dataset based on the real-world burst imaging pipeline for training and evaluating different methods in the joint HDR imaging, denoising, and deblurring task. Second, due to the corruption-insensitive of previous network architectures, we propose a novel and efficient attention-based multi-exposure HDR imaging method, which skillfully selects the optimal information (clean or sharp) from corrupted LDR inputs by our customized cross-attention mechanism to generate HDR information. Furthermore, to enhance the robustness of our cross-attention mechanism, we introduce a novel E ntropy D ecreasing loss (ED loss) which decreases the entropy of the calculated attention map to alleviate the ghosting artifacts during multi-exposure information fusion. Extensive experimental results demonstrate that the proposed method trained on the proposed dataset surpasses related state-of-the-art methods with outstanding real-world photography quality. Codes and the dataset will be available.
Zhefan Rao, Chenyang Lei, Wenxiu Sun, Qiong Yan, Qifeng Chen 0001
Int. J. Comput. Vis.5
2025 Modeling Dual-Exposure Quad-Bayer Patterns for Joint Denoising and Deblurring
abstract
Image degradation caused by noise and blur remains a persistent challenge in imaging systems, stemming from limitations in both hardware and methodology. Single-image solutions face an inherent tradeoff between noise reduction and motion blur. While short exposures can capture clear motion, they suffer from noise amplification. Long exposures reduce noise but introduce blur. Learning-based single-image enhancers tend to be over-smooth due to the limited information. Multi-image solutions using burst mode avoid this tradeoff by capturing more spatial-temporal information but often struggle with misalignment from camera/scene motion. To address these limitations, we propose a physical-model-based image restoration approach leveraging a novel dual-exposure Quad-Bayer pattern sensor. By capturing pairs of short and long exposures at the same starting point but with varying durations, this method integrates complementary noise-blur information within a single image. We further introduce a Quad-Bayer synthesis method (B2QB) to simulate sensor data from Bayer patterns to facilitate training. Based on this dual-exposure sensor model, we design a hierarchical convolutional neural network called QRNet to recover high-quality RGB images. The network incorporates input enhancement blocks and multi-level feature extraction to improve restoration quality. Experiments demonstrate superior performance over state-of-the-art deblurring and denoising methods on both synthetic and real-world datasets. The code, model, and datasets are publicly available at https://github.com/zhaoyuzhi/QRNet.
Yuzhi Zhao, Lai-Man Po, Yongzhe Xu, Qiong Yan
IEEE Trans. Image Process.5
2024 Iterative Token Evaluation and Refinement for Real-World Super-resolution
abstract
Real-world image super-resolution (RWSR) is a long-standing problem as low-quality (LQ) images often have complex and unidentified degradations. Existing methods such as Generative Adversarial Networks (GANs) or continuous diffusion models present their own issues including GANs being difficult to train while continuous diffusion models requiring numerous inference steps. In this paper, we propose an Iterative Token Evaluation and Refinement (ITER) framework for RWSR, which utilizes a discrete diffusion model operating in the discrete token representation space, i.e., indexes of features extracted from a VQGAN codebook pre-trained with high-quality (HQ) images. We show that ITER is easier to train than GANs and more efficient than continuous diffusion models. Specifically, we divide RWSR into two sub-tasks, i.e., distortion removal and texture generation. Distortion removal involves simple HQ token prediction with LQ images, while texture generation uses a discrete diffusion model to iteratively refine the distortion removal output with a token refinement network. In particular, we propose to include a token evaluation network in the discrete diffusion process. It learns to evaluate which tokens are good restorations and helps to improve the iterative refinement results. Moreover, the evaluation network can first check status of the distortion removal output and then adaptively select total refinement steps needed, thereby maintaining a good balance between distortion removal and texture generation. Extensive experimental results show that ITER is easy to train and performs well within just 8 iterative steps.
Chaofeng Chen, Shangchen Zhou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin
AAAI6
2024 Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation Models
abstract
Multi-modality large language models (MLLMs), as represented by GPT-4V, have introduced a paradigm shift for visual perception and understanding tasks, that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the identification of low-level visual attributes (e.g., clarity, brightness) to the evaluation on image quality, there's still an imperative to further improve the accuracy of MLLMs to substantially alleviate human burdens. To address this, we collect the first dataset consisting of human natural language feedback on low-level vision. Each feedback offers a comprehensive description of an image's low-level visual attributes, culminating in an overall quality assessment. The constructed Q-Pathway dataset includes 58K detailed human feedbacks on 18,973 multi-sourced images with diverse low-level appearance. To ensure MLLMs can adeptly handle diverse queries, we further propose a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct. Experimental results indicate that the Q-Instruct consistently elevates various low-level visual capabilities across multiple base models. We anticipate that our datasets can pave the way for a future that foundation models can assist humans on low-level visual tasks.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Kaixin Xu, Chunyi Li 0001, Jingwen Hou, Guangtao Zhai, Geng Xue, Wenxiu Sun, Qiong Yan, Weisi Lin
CVPR13
2024 Boosting Image Quality Assessment Through Efficient Transformer Adaptation with Local Feature Enhancement
abstract
Image Quality Assessment (IQA) constitutes a funda-mental task within the field of computer vision, yet it re-mains an unresolved challenge, owing to the intricate dis-tortion conditions, diverse image contents, and limited availability of data. Recently, the community has wit-nessed the emergence of numerous large-scale pretrained foundation models. However, it remains an open problem whether the scaling law in high-level tasks is also appli-cable to IQA tasks which are closely related to low-level clues. In this paper, we demonstrate that with a proper in-jection of local distortion features, a larger pretrained vision transformer (ViT) foundation model performs better in IQA tasks. Specifically, for the lack of local distortion structure and inductive bias of the large-scale pretrained ViT, we use another pretrained convolution neural networks (CNNs), which is well known for capturing the local structure, to extract multi-scale image features. Further, we propose a local distortion extractor to obtain local distortion features from the pretrained CNNs and a local distortion in-jector to inject the local distortion features into ViT. By only training the extractor and injector, our method can benefit from the rich knowledge in the powerful foundation models and achieve state-of-the-art performance on popular IQA datasets, indicating that IQA is not only a low-level problem but also benefits from stronger high-level features drawn from large-scale pretrained models. Codes are publicly available at: https://github.com/NeosXu/LoDa.
Kangmin Xu, Jing Xiao 0004, Chaofeng Chen, Haoning Wu 0001, Qiong Yan, Weisi Lin
CVPR6
2024 Enhancing Diffusion Models with Text-Encoder Reinforcement Learning
Chaofeng Chen, Annan Wang, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin
ECCV (25)6
2024 Towards Open-Ended Visual Quality Comparison
Haoning Wu 0001, Hanwei Zhu, Erli Zhang 0001, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Xiaohong Liu 0001, Guangtao Zhai, Shiqi Wang 0001, Weisi Lin
ECCV (3)10
2024 Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
abstract
The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding**. To address this gap, we present **Q-Bench**, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment. **_a)_** To evaluate the low-level **_perception_** ability, we construct the **LLVisionQA** dataset, consisting of 2,990 diverse-sourced images, each equipped with a human-asked question focusing on its low-level attributes. We then measure the correctness of MLLMs on answering these questions. **_b)_** To examine the **_description_** ability of MLLMs on low-level information, we propose the **LLDescribe** dataset consisting of long expert-labelled *golden* low-level text descriptions on 499 images, and a GPT-involved comparison pipeline between outputs of MLLMs and the *golden* descriptions. **_c)_** Besides these two tasks, we further measure their visual quality **_assessment_** ability to align with human opinion scores. Specifically, we design a softmax-based strategy that enables MLLMs to predict *quantifiable* quality scores, and evaluate them on various existing image quality assessment (IQA) datasets. Our evaluation across the three abilities confirms that MLLMs possess preliminary low-level visual skills. However, these skills are still unstable and relatively imprecise, indicating the need for specific enhancements on MLLMs towards these abilities. We hope that our benchmark can encourage the research community to delve deeper to discover and enhance these untapped potentials of MLLMs.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Chunyi Li 0001, Wenxiu Sun, Qiong Yan, Guangtao Zhai, Weisi Lin
ICLR9
2024 Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
abstract
The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released.
Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ICML11
2024 Q-Ground: Image Quality Grounding with Large Multi-modality Models
abstract
Recent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual content. However, these advancements are mostly focused on overall quality assessment, and the detailed examination of local quality, which is crucial for comprehensive visual understanding, is still largely unexplored. In this work, we introduce Q-Ground, the first framework aimed at tackling fine-scale visual quality grounding by combining large multi-modality models with detailed visual quality analysis. Cen- tral to our contribution is the introduction of the QGround-100K dataset, a novel resource containing 100k triplets of (image, quality text, distortion segmentation) to facilitate deep investigations into visual quality. The dataset comprises two parts: one with human- labeled annotations for accurate quality assessment, and another la- beled automatically by LMMs such as GPT4V, which helps improve the robustness of model training while also reducing the costs of data collection. With the QGround-100K dataset, we propose a LMM-based method equipped with multi-scale feature learning to learn models capable of performing both image quality answer- ing and distortion segmentation based on text prompts. This dual- capability approach not only refines the model’s understanding of region-aware image quality but also enables it to interactively re- spond to complex, text-based queries about image quality and spe- cific distortions. Q-Ground takes a step towards sophisticated vi- sual quality analysis in a finer scale, establishing a new benchmark for future research in the area. Codes and dataset are available at https://github.com/Q-Future/Q-Ground.
Chaofeng Chen, Sensen Yang, Haoning Wu 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ACM Multimedia8
2024 Esophageal cancer detection framework based on time series information from smear images
Chuanwang Zhang, Dongyao Jia, Nengkai Wu, Qiong Yan
Expert Syst. Appl.7
2024 TOPIQ: A Top-Down Approach From Semantics to Distortions for Image Quality Assessment
abstract
Image Quality Assessment (IQA) is a fundamental task in computer vision that has witnessed remarkable progress with deep neural networks. Inspired by the characteristics of the human visual system, existing methods typically use a combination of global and local representations (i.e., multi-scale features) to achieve superior performance. However, most of them adopt simple linear fusion of multi-scale features, and neglect their possibly complex relationship and interaction. In contrast, humans typically first form a global impression to locate important regions and then focus on local details in those regions. We therefore propose a top-down approach that uses high-level semantics to guide the IQA network to focus on semantically important local distortion regions, named as TOPIQ. Our approach to IQA involves the design of a heuristic coarse-to-fine network (CFANet) that leverages multi-scale features and progressively propagates multi-level semantic information to low-level representations in a top-down manner. A key component of our approach is the proposed cross-scale attention mechanism, which calculates attention maps for lower level features guided by higher level features. This mechanism emphasizes active semantic regions for low-level distortions, thereby improving performance. TOPIQ can be used for both Full-Reference (FR) and No-Reference (NR) IQA. We use ResNet50 as its backbone and demonstrate that TOPIQ achieves better or competitive performance on most public FR and NR benchmarks compared with state-of-the-art methods based on vision transformers, while being much more efficient (with only ∼ 13% FLOPS of the current best FR method). Codes are released at https://github.com/chaofengc/IQA-PyTorch.
Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin
IEEE Trans. Image Process.7
2024 Blind Video Quality Prediction by Uncovering Human Video Perceptual Representation
abstract
Blind video quality assessment (VQA) has become an increasingly demanding problem in automatically assessing the quality of ever-growing in-the-wild videos. Although efforts have been made to measure temporal distortions, the core to distinguish between VQA and image quality assessment (IQA), the lack of modeling of how the human visual system (HVS) relates to the temporal quality of videos hinders the precise mapping of predicted temporal scores to the human perception. Inspired by the recent discovery of the temporal straightness law of natural videos in the HVS, this paper intends to model the complex temporal distortions of in-the-wild videos in a simple and uniform representation by describing the geometric properties of videos in the visual perceptual domain. A novel videolet, with perceptual representation embedding of a few consecutive frames, is designed as the basic quality measurement unit to quantify temporal distortions by measuring the angular and linear displacements from the straightness law. By combining the predicted score on each videolet, a perceptually temporal quality evaluator (PTQE) is formed to measure the temporal quality of the entire video. Experimental results demonstrate that the perceptual representation in the HVS is an efficient way of predicting subjective temporal quality. Moreover, when combined with spatial quality metrics, PTQE achieves top performance over popular in-the-wild video datasets. More importantly, PTQE requires no additional information beyond the video being assessed, making it applicable to any dataset without parameter tuning. Additionally, the generalizability of PTQE is evaluated on video frame interpolation tasks, demonstrating its potential to benefit temporal-related enhancement tasks.
Kangmin Xu, Haoning Wu 0001, Chaofeng Chen, Wenxiu Sun, Qiong Yan, C.-C. Jay Kuo, Weisi Lin
IEEE Trans. Image Process.6
2024 HSGAN: Hyperspectral Reconstruction From RGB Images With Generative Adversarial Network
abstract
Hyperspectral (HS) reconstruction from RGB images denotes the recovery of whole-scene HS information, which has attracted much attention recently. State-of-the-art approaches often adopt convolutional neural networks to learn the mapping for HS reconstruction from RGB images. However, they often do not achieve high HS reconstruction performance across different scenes consistently. In addition, their performance in recovering HS images from clean and real-world noisy RGB images is not consistent. To improve the HS reconstruction accuracy and robustness across different scenes and from different input images, we present an effective HSGAN framework with a two-stage adversarial training strategy. The generator is a four-level top-down architecture that extracts and combines features on multiple scales. To generalize well to real-world noisy images, we further propose a spatial-spectral attention block (SSAB) to learn both spatial-wise and channel-wise relations. We conduct the HS reconstruction experiments from both clean and real-world noisy RGB images on five well-known HS datasets. The results demonstrate that HSGAN achieves superior performance to existing methods. Please visit https://github.com/zhaoyuzhi/HSGAN to try our codes.
Yuzhi Zhao, Lai-Man Po, Tingyu Lin 0002, Qiong Yan, Wei Liu 0004, Pengfei Xian
IEEE Trans. Neural Networks Learn. Syst.4
2023 Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives
abstract
The rapid increase in user-generated content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the $\color{Green}{\text{technical perspective}}$, measuring the perception of distortions; and the $\color{Blue}{\text{aesthetic perspective}}$, which relates to preference and recommendation on contents. To understand how these two perspectives affect overall subjective opinions in UGC-VQA, we conduct a large-scale subjective study to collect human quality opinions on the overall quality of videos as well as perceptions from aesthetic and technical perspectives. The collected Disentangled Video Quality Database (DIVIDE-3k) confirms that human quality opinions on UGC videos are universally and inevitably affected by both aesthetic and technical perspectives. In light of this, we propose the Disentangled Objective Video Quality Evaluator (DOVER) to learn the quality of UGC videos based on the two perspectives. The DOVER proves state-of-the-art performance in UGC-VQA under very high efficiency. With perspective opinions in DIVIDE-3k, we further propose DOVER++, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective. Code at https://github.com/VQAssessment/DOVER.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ICCV8
2023 Exploring Opinion-Unaware Video Quality Assessment with Semantic Affinity Criterion
abstract
Recent learning-based video quality assessment (VQA) algorithms are expensive to implement due to the cost of data collection of human quality opinions, and are less robust across various scenarios due to the biases of these opinions. This motivates our exploration on opinion-unaware (a.k.a zero-shot) VQA approaches. Existing approaches only considers low-level naturalness in spatial or temporal domain, without considering impacts from high-level semantics. In this work, we introduce an explicit semantic affinity index for opinion-unaware VQA using text-prompts in the contrastive language-image pre-training (CLIP) model. We also aggregate it with different traditional low-level naturalness indexes through gaussian normalization and sigmoid rescaling strategies. Composed of aggregated semantic and technical metrics, the proposed Blind Unified Opinion-Unaware Video Quality Index via Semantic and Technical Metric Aggregation (BUONA-VISTA) outperforms existing opinion-unaware VQA methods by at least 20% improvements, and is more robust than opinion-aware approaches.
Haoning Wu 0001, Jingwen Hou, Chaofeng Chen, Erli Zhang 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ICME8
2023 Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach
abstract
The proliferation of in-the-wild videos has greatly expanded the Video Quality Assessment (VQA) problem. Unlike early definitions that usually focus on limited distortion types, VQA on in-the-wild videos is especially challenging as it could be affected by complicated factors, including various distortions and diverse contents. Though subjective studies have collected overall quality scores for these videos, how the abstract quality scores relate with specific factors is still obscure, hindering VQA methods from more concrete quality evaluations (e.g. sharpness of a video). To solve this problem, we collect over two million opinions on 4,543 in-the-wild videos on 13 dimensions of quality-related factors, including in-capture authentic distortions (e.g. motion blur, noise, flicker), errors introduced by compression and transmission, and higher-level experiences on semantic contents and aesthetic issues (e.g. composition, camera trajectory), to establish the multi-dimensional Maxwell database. Specifically, we ask the subjects to label among a positive, a negative, and a neutral choice for each dimension. These explanation-level opinions allow us to measure the relationships between specific quality factors and abstract subjective quality ratings, and to benchmark different categories of VQA algorithms on each dimension, so as to more comprehensively analyze their strengths and weaknesses. Furthermore, we propose the MaxVQA, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in our analyses. The MaxVQA can jointly evaluate various specific quality factors and final quality scores with state-of-the-art accuracy on all dimensions, and superb generalization ability on existing datasets. Code and data available at https://github.com/VQAssessment/MaxVQA.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ACM Multimedia8
2023 Neighbourhood Representative Sampling for Efficient End-to-End Video Quality Assessment
abstract
The increased resolution of real-world videos presents a dilemma between efficiency and accuracy for deep Video Quality Assessment (VQA). On the one hand, keeping the original resolution will lead to unacceptable computational costs. On the other hand, existing practices, such as resizing or cropping, will change the quality of original videos due to difference in details or loss of contents, and are henceforth harmful to quality assessment. With obtained insight from the studies of spatial-temporal redundancy in the human visual system, visual quality around a neighbourhood has high probability to be similar, and this motivates us to investigate an effective quality-sensitive neighbourhood representative sampling scheme for VQA. In this work, we propose a unified scheme, spatial-temporal grid mini-cube sampling (St-GMS), and the resultant samples are namedfragments. In St-GMS, full-resolution videos are first divided into mini-cubes with predefined spatial-temporal grids, then the temporal-aligned quality representatives are sampled to compose the fragments that serve as inputs for VQA. In addition, we design the Fragment Attention Network (FANet), a network architecture tailored specifically for fragments. With fragments and FANet, the proposedFAST-VQAandFasterVQA(with an improved sampling scheme) achieves up to 1612× efficiency than the existing state-of-the-art, meanwhile achieving significantly better performance on all relevant VQA benchmarks.
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Jinwei Gu, Weisi Lin
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment
abstract
Compared with spatial counterparts, temporal relationships between frames and their influences on video quality assessment (VQA) are still relatively under-studied in existing works. These relationships lead to two important types of effects for video quality. Firstly, some meaningless temporal variations (such as shaking, flicker, and unsmooth scene transitions) cause temporal distortions that degrade quality of videos. Secondly, the human visual system often has different attention to frames with different contents, resulting in their different importance to the overall video quality. Based on prominent time-series modeling ability of transformers, we propose a novel and effective transformer-based VQA method to tackle these two issues. To better differentiate temporal variations and thus capture the temporal distortions, we design the Spatial-Temporal Distortion Extraction (STDE) module that extracts multi-level spatial-temporal features with a video swin transformer tiny (Swin-T) backbone and uses temporal difference layer to further capture these distortions. To tackle with temporal quality attention, we propose the encoder-decoder-like temporal content transformer (TCT). We also introduce the temporal sampling on features to reduce the input length for the TCT, so as to improve the learning effectiveness and efficiency of this module. Consisting of the STDE and the TCT, the proposed Temporal Distortion-Content Transformers for Video Quality Assessment (DisCoVQA) reaches state-of-the-art performance on several VQA benchmarks without any extra pre-training datasets and up to 10% better generalization ability than existing methods. We also conduct extensive ablation experiments to prove the effectiveness of each part in our proposed model, and provide visualizations to prove that the proposed modules achieve our intention on modeling these temporal issues. Our code is published athttps://github.com/QualityAssessment/DisCoVQA.
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.6
2023 Fine-Grained Face Editing via Personalized Spatial-Aware Affine Modulation
abstract
Fine-grained face editing, as a special case of image translation task, aims at modifying face attributes according to users' preference. Although generative adversarial networks (GANs) have achieved great success in general image translation tasks, these models cannot be directly applied in the face editing problem. Ideal face editing is challenging as it has two special requirements -- personalization and spatial-awareness. To address these issues, we propose a novel Personalized Spatial-aware Affine Modulation (PSAM) method based on a general GAN structure. The key idea is to modulate the intermediate features in a personalized and spatial-aware manner, which corresponds to the face editing procedure. Specifically, for personalization, we adopt both the face image and the desired attribute as input to generate the modulation tensors. For spatial-aware, we set these tensors to be of the same size as the input image, allowing pixel-wise modulation. Extensive experiments in four fine-grained face editing tasks, i.e., makeup, expression, illumination and aging, demonstrate the effectiveness of the proposed PSAM method. The synthesis results of PSAM can be further boosted by a new transferable training strategy. To facilitate research of face editing, we also construct a new large-scale makeup dataset. The code and dataset are available at http://colalab.org/projects/PSAM.
Si Liu 0001, Renda Bao, Defa Zhu, Shaofei Huang 0001, Qiong Yan, Liang Lin 0004, Chao Dong 0005
IEEE Trans. Multim.5
2023 ChildPredictor: A Child Face Prediction Framework With Disentangled Learning
abstract
The appearances of children are inherited from their parents, which makes it feasible to predict them. Predicting realistic children's faces may help settle many social problems, such as age-invariant face recognition, kinship verification, and missing child identification. It can be regarded as an image-to-image translation task. Existing approaches usually assume domain information in the image-to-image translation can be interpreted by “style”, i.e., the separation of image content and style. However, such separation is improper for the child face prediction, because the facial contours between children and parents are not the same. To address this issue, we propose a new disentangled learning strategy for children's face prediction. We assume that children's faces are determined by genetic factors (compact family features, e.g., face contour), external factors (facial attributes irrelevant to prediction, such as moustaches and glasses), and variety factors (individual properties for each child). On this basis, we formulate predictions as a mapping from parents’ genetic factors to children's genetic factors, and disentangle them from external and variety factors. In order to obtain accurate genetic factors and perform the mapping, we propose a ChildPredictor framework. It transfers human faces to genetic factors by encoders and back by generators. Then, it learns the relationship between the genetic factors of parents and children through a mapping function. To ensure the generated faces are realistic, we collect a large Family Face Database to train ChildPredictor and evaluate it on the FF-Database validation set. Experimental results demonstrate that ChildPredictor is superior to other well-known image-to-image translation methods in predicting realistic and diverse child faces. Implementation codes can be found athttps://github.com/zhaoyuzhi/ChildPredictor.
Yuzhi Zhao, Lai-Man Po, Qiong Yan, Wei Shen 0002, Yujia Zhang 0002, Wei Liu 0004, Chun Kit Wong, Chiu-Sing Pang, Weifeng Ou, Wing Yin Yu, Buhua Liu
IEEE Trans. Multim.4
2022 MODNet: Real-Time Trimap-Free Portrait Matting via Objective Decomposition
abstract
Existing portrait matting methods either require auxiliary inputs that are costly to obtain or involve multiple stages that are computationally expensive, making them less suitable for real-time applications. In this work, we present a light-weight matting objective decomposition network (MODNet) for portrait matting in real-time with a single input image. The key idea behind our efficient design is by optimizing a series of sub-objectives simultaneously via explicit constraints. In addition, MODNet includes two novel techniques for improving model efficiency and robustness. First, an Efficient Atrous Spatial Pyramid Pooling (e-ASPP) module is introduced to fuse multi-scale features for semantic estimation. Second, a self-supervised sub-objectives consistency (SOC) strategy is proposed to adapt MODNet to real-world data to address the domain shift problem common to trimap-free methods. MODNet is easy to be trained in an end-to-end manner. It is much faster than contemporaneous methods and runs at 67 frames per second on a 1080Ti GPU. Experiments show that MODNet outperforms prior trimap-free methods by a large margin on both Adobe Matting Dataset and a carefully designed photographic portrait matting (PPM-100) benchmark proposed by us. Further, MODNet achieves remarkable results on daily photos and videos.
Zhanghan Ke, Kaican Li, Qiong Yan, Rynson W. H. Lau
AAAI4
2022 FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ECCV (6)7
2022 D2HNet: Joint Denoising and Deblurring with Hierarchical Network for Robust Night Image Restoration
Yuzhi Zhao, Yongzhe Xu, Qiong Yan, Dingdong Yang, Lai-Man Po
ECCV (7)3
2022 Exploring the Effectiveness of Video Perceptual Representation in Blind Video Quality Assessment
abstract
With the rapid growth of in-the-wild videos taken by non-specialists, blind video quality assessment (VQA) has become a challenging and demanding problem. Although lots of efforts have been made to solve this problem, it remains unclear how the human visual system (HVS) relates to the temporal quality of videos. Meanwhile, recent work has found that the frames of natural video transformed into the perceptual domain of the HVS tend to form a straight trajectory of the representations. With the obtained insight that distortion impairs the perceived video quality and results in a curved trajectory of the perceptual representation, we propose a temporal perceptual quality index (TPQI) to measure the temporal distortion by describing the graphic morphology of the representation. Specifically, we first extract the video perceptual representations from the lateral geniculate nucleus (LGN) and primary visual area (V1) of the HVS, and then measure the straightness and compactness of their trajectories to quantify the degradation in naturalness and content continuity of video. Experiments show that the perceptual representation in the HVS is an effective way of predicting subjective temporal quality, and thus TPQI can, for the first time, achieve comparable performance to the spatial quality metric and be even more effective in assessing videos with large temporal variations. We further demonstrate that by combining with NIQE, a spatial quality metric, TPQI can achieve top performance over popular in-the-wild video datasets. More importantly, TPQI does not require any additional information beyond the video being evaluated and thus can be applied to any datasets without parameter tuning. Source code is available at https://github.com/UoLMM/TPQI-VQA.
Kangmin Xu, Haoning Wu 0001, Chaofeng Chen, Wenxiu Sun, Qiong Yan, Weisi Lin
ACM Multimedia6
2021 Dual-Camera Super-Resolution with Aligned Attention Modules
abstract
We present a novel approach to reference-based super-resolution (RefSR) with the focus on dual-camera super-resolution (DCSR), which utilizes reference images for high-quality and high-fidelity results. Our proposed method generalizes the standard patch-based feature matching with spatial alignment operations. We further explore the dual-camera super-resolution that is one promising application of RefSR, and build a dataset that consists of 146 image pairs from the main and telephoto cameras in a smartphone. To bridge the domain gaps between real-world images and the training images, we propose a self-supervised domain adaptation strategy for real-world images. Extensive experiments on our dataset and a public benchmark demonstrate clear improvement achieved by our method over state of the art in both quantitative evaluation and visual comparisons. Our code and data are available at https://tengfei-wang.github.io/Dual-Camera-SR/index.html.
Tengfei Wang 0002, Jiaxin Xie, Wenxiu Sun, Qiong Yan, Qifeng Chen 0001
ICCV4
2021 Semasuperpixel: A Multi-Channel Probability-Driven Superpixel Segmentation Method
abstract
Superpixel, an efficient image segmentation approach, aggregates a group of similar pixels into the same cluster. Existing superpixel algorithms still mainly focus on the color information while ignoring the semantic distribution knowledge. In this paper, we propose a semantic information-driven method that adopts multi-channel semantic probabilities for superpixel segmentation. By conducting statistical analysis on the semantic output and then formulating the distance measure, the prior knowledge of the semantic with a dynamic confidence value could be utilized by our method during the global update effectively. Extensive experimental evaluations show that our method achieves a leading segmentation quality and convergence speed, compared to other five state-of-the-art algorithms, as measured by boundary recall, undersegmentation error, and explained variation.
Qingyun Zhao, Lei Fan 0005, Yuzhi Zhao, Qiong Yan, Long Chen 0005
ICIP6
2020 Polarized Reflection Removal With Perfect Alignment in the Wild
abstract
We present a novel formulation to removing reflection from polarized images in the wild. We first identify the misalignment issues of existing reflection removal datasets where the collected reflection-free images are not perfectly aligned with input mixed images due to glass refraction. Then we build a new dataset with more than 100 types of glass in which obtained transmission images are perfectly aligned with input mixed images. Second, capitalizing on the special relationship between reflection and polarized light, we propose a polarized reflection removal model with a two-stage architecture. In addition, we design a novel perceptual NCC loss that can improve the performance of reflection removal and general image decomposition tasks. We conduct extensive experiments, and results suggest that our model outperforms state-of-the-art methods on reflection removal.
Chenyang Lei, Xuhua Huang, Qiong Yan, Wenxiu Sun, Qifeng Chen 0001
CVPR4
2020 Guided Collaborative Training for Pixel-Wise Semi-Supervised Learning
Zhanghan Ke, Di Qiu, Kaican Li, Qiong Yan, Rynson W. H. Lau
ECCV (13)4
2019 Deep Surface Normal Estimation With Hierarchical RGB-D Fusion
abstract
The growing availability of commodity RGB-D cameras has boosted the applications in the field of scene understanding. However, as a fundamental scene understanding task, surface normal estimation from RGB-D data lacks thorough investigation. In this paper, a hierarchical fusion network with adaptive feature re-weighting is proposed for surface normal estimation from a single RGB-D image. Specifically, the features from color image and depth are successively integrated at multiple scales to ensure global surface smoothness while preserving visually salient details. Meanwhile, the depth features are re-weighted with a confidence map estimated from depth before merging into the color branch to avoid artifacts caused by input depth corruption. Additionally, a hybrid multi-scale loss function is designed to learn accurate normal estimation given noisy ground-truth dataset. Extensive experimental results validate the effectiveness of the fusion strategy and the loss design, outperforming state-of-the-art normal estimation schemes.
Jin Zeng 0004, Yanfeng Tong, Yunmu Huang, Qiong Yan, Wenxiu Sun, Jing Chen 0018, Yongtian Wang
CVPR4
2019 Dual Student: Breaking the Limits of the Teacher in Semi-Supervised Learning
abstract
Recently, consistency-based methods have achieved state-of-the-art results in semi-supervised learning (SSL). These methods always involve two roles, an explicit or implicit teacher model and a student model, and penalize predictions under different perturbations by a consistency constraint. However, the weights of these two roles are tightly coupled since the teacher is essentially an exponential moving average (EMA) of the student. In this work, we show that the coupled EMA teacher causes a performance bottleneck. To address this problem, we introduce Dual Student, which replaces the teacher with another student. We also define a novel concept, stable sample, following which a stabilization constraint is designed for our structure to be trainable. Further, we discuss two variants of our method, which produce even higher performance. Extensive experiments show that our method improves the classification performance significantly on several main SSL benchmarks. Specifically, it reduces the error rate of the 13-layer CNN from 16.84% to 12.39% on CIFAR-10 with 1k labels and from 34.10% to 31.56% on CIFAR-100 with 10k labels. In addition, our method also achieves a clear improvement in domain adaptation.
Zhanghan Ke, Daoye Wang, Qiong Yan, Jimmy S. J. Ren, Rynson W. H. Lau
ICCV3
2018 DSR: Direct Self-Rectification for Uncalibrated Dual-Lens Cameras
abstract
With the developments of dual-lens camera modules, depth information representing the third dimension of the captured scenes becomes available for smartphones. It is estimated by stereo matching algorithms, taking as input the two views captured by dual-lens cameras at slightly different viewpoints. Depth-of-field rendering (also be referred to as synthetic defocus or bokeh) is one of the trending depth-based applications. However, to achieve fast depth estimation on smartphones, the stereo pairs need to be rectified in the first place. In this paper, we propose a cost-effective solution to perform stereo rectification for dual-lens cameras called direct self-rectification, short for DSR. It removes the need of individual offline calibration for every pair of dual-lens cameras. In addition, the proposed solution is robust to the slight movements, {\it e.g.}, due to collisions, of the dual-lens cameras after fabrication. Different with existing self-rectification approaches, our approach computes the homography in a novel way with zero geometric distortions introduced to the master image. It is achieved by directly minimizing the vertical displacements of corresponding points between the original master image and the transformed slave image. Our method is evaluated on both realistic and synthetic stereo image pairs, and produces superior results compared to the calibrated rectification or other self-rectification approaches.
Ruichao Xiao, Wenxiu Sun, Jiahao Pang, Qiong Yan, Jimmy S. J. Ren
3DV4
2018 BeautyGAN: Instance-level Facial Makeup Transfer with Deep Generative Adversarial Network
abstract
Facial makeup transfer aims to translate the makeup style from a given reference makeup face image to another non-makeup one while preserving face identity. Such an instance-level transfer problem is more challenging than conventional domain-level transfer tasks, especially when paired data is unavailable. Makeup style is also different from global styles (e.g., paintings) in that it consists of several local styles/cosmetics, including eye shadow, lipstick, foundation, and so on. Extracting and transferring such local and delicate makeup information is infeasible for existing style transfer methods. We address the issue by incorporating both global domain-level loss and local instance-level loss in an dual input/output Generative Adversarial Network, called BeautyGAN. Specifically, the domain-level transfer is ensured by discriminators that distinguish generated images from domains' real samples. The instance-level loss is calculated by pixel-level histogram loss on separate local facial regions. We further introduce perceptual loss and cycle consistency loss to generate high quality faces and preserve identity. The overall objective function enables the network to learn translation on instance-level through unsupervised adversarial learning. We also build up a new makeup dataset that consists of 3834 high-resolution face images. Extensive experiments show that BeautyGAN could generate visually pleasant makeup faces and accurate transferring results. Data and code are available at http://liusi-group.com/projects/BeautyGAN.
Ruihe Qian, Chao Dong 0005, Si Liu 0001, Qiong Yan, Wenwu Zhu 0001, Liang Lin 0004
ACM Multimedia5
2017 Accurate Single Stage Detector Using Recurrent Rolling Convolution
abstract
Most of the recent successful methods in accurate object detection and localization used some variants of R-CNN style two stage Convolutional Neural Networks (CNN) where plausible regions were proposed in the first stage then followed by a second stage for decision refinement. Despite the simplicity of training and the efficiency in deployment, the single stage detection methods have not been as competitive when evaluated in benchmarks consider mAP for high IoU thresholds. In this paper, we proposed a novel single stage end-to-end trainable object detection network to overcome this limitation. We achieved this by introducing Recurrent Rolling Convolution (RRC) architecture over multi-scale feature maps to construct object classifiers and bounding box regressors which are deep in context. We evaluated our method in the challenging KITTI dataset which measures methods under IoU threshold of 0.7. We showed that with RRC, a single reduced VGG-16 based model already significantly outperformed all the previously published results. At the time this paper was written our models ranked the first in KITTI car detection (the hard level), the first in cyclist detection and the second in pedestrian detection. These results were not reached by the previous single stage methods. The code is publicly available.
Jimmy S. J. Ren, Xiaohao Chen, Wenxiu Sun, Jiahao Pang, Qiong Yan, Yu-Wing Tai, Li Xu 0001
CVPR6
2016 Look, Listen and Learn - A Multimodal LSTM for Speaker Identification
abstract
Speaker identification refers to the task of localizing the face of a person who has the same identity as the ongoing voice in a video. This task not only requires collective perception over both visual and auditory signals, the robustness to handle severe quality degradations and unconstrained content variations are also indispensable. In this paper, we describe a novel multimodal Long Short-Term Memory (LSTM) architecture which seamlessly unifies both visual and auditory modalities from the beginning of each sequence input. The key idea is to extend the conventional LSTM by not only sharing weights across time steps, but also sharing weights across modalities. We show that modeling the temporal dependency across face and voice can significantly improve the robustness to content quality degradations and variations. We also found that our multimodal LSTM is robustness to distractors, namely the non-speaking identities. We applied our multimodal LSTM to The Big Bang Theory dataset and showed that our system outperforms the state-of-the-art systems in speaker identification with lower false alarm rate and higher recognition accuracy.
Jimmy S. J. Ren, Yongtao Hu 0001, Yu-Wing Tai, Li Xu 0001, Wenxiu Sun, Qiong Yan
AAAI7
2016 Hierarchical Image Saliency Detection on Extended CSSD
abstract
Complex structures commonly exist in natural images. When an image contains small-scale high-contrast patterns either in the background or foreground, saliency detection could be adversely affected, resulting erroneous and non-uniform saliency assignment. The issue forms a fundamental challenge for prior methods. We tackle it from a scale point of view and propose a multi-layer approach to analyze saliency cues. Different from varying patch sizes or downsizing images, we measure region-based scales. The final saliency values are inferred optimally combining all the saliency cues in different scales using hierarchical inference. Through our inference model, single-scale information is selected to obtain a saliency map. Our method improves detection quality on many images that cannot be handled well traditionally. We also construct an extended Complex Scene Saliency Dataset (ECSSD) to include complex but general natural images.
Jianping Shi, Qiong Yan, Li Xu 0001, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Deep Edge-Aware Filters
abstract
There are many edge-aware filters varying in their construction forms and filtering properties. It seems impossible to uniformly represent and accelerate them in a single framework. We made the attempt to learn a big and important family of edge-aware operators from data. Our method is based on a deep convolutional neural network with a gradient domain training procedure, which gives rise to a powerful tool to approximate various filters without knowing the original models and implementation details. The only difference among these operators in our system becomes merely the learned parameters. Our system enables fast approximation for complex edge-aware filters and achieves up to 200x acceleration, regardless of their originally very different implementation. Fast speed can also be achieved when creating new effects using spatially varying filter or filter combination, bearing out the effectiveness of our deep edge-aware filters.
Li Xu 0001, Jimmy S. J. Ren, Qiong Yan, Renjie Liao 0001, Jiaya Jia
ICML3
2015 Shepard Convolutional Neural Networks
abstract
Deep learning has recently been introduced to the field of low-level computer vision and image processing. Promising results have been obtained in a number of tasks including super-resolution, inpainting, deconvolution, filtering, etc. However, previously adopted neural network approaches such as convolutional neural networks and sparse auto-encoders are inherently with translation invariant operators. We found this property prevents the deep learning approaches from outperforming the state-of-the-art if the task itself requires translation variant interpolation (TVI). In this paper, we draw on Shepard interpolation and design Shepard Convolutional Neural Networks (ShCNN) which efficiently realizes end-to-end trainable TVI operators in the network. We show that by adding only a few feature maps in the new Shepard layers, the network is able to achieve stronger results than a much deeper architecture. Superior performance on both image inpainting and super-resolution is obtained where our system outperforms previous ones while keeping the running time competitive.
Jimmy S. J. Ren, Li Xu 0001, Qiong Yan, Wenxiu Sun
NIPS3
2015 Multispectral Joint Image Restoration via Optimizing a Scale Map
abstract
Color, infrared and flash images captured in different fields can be employed to effectively eliminate noise and other visual artifacts. We propose a two-image restoration framework considering input images from different fields, for example, one noisy color image and one dark-flashed near-infrared image. The major issue in such a framework is to handle all structure divergence and find commonly usable edges and smooth transitions for visually plausible image reconstruction. We introduce a novel scale map as a competent representation to explicitly model derivative-level confidence and propose new functions and a numerical solver to effectively infer it following our important structural observations. Multispectral shadow detection is also used to make our system more robust. Our method is general and shows a principled way to solve multispectral restoration problems.
Xiaoyong Shen, Qiong Yan, Li Xu 0001, Lizhuang Ma, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Hierarchical Saliency Detection
abstract
When dealing with objects with complex structures, saliency detection confronts a critical problem - namely that detection accuracy could be adversely affected if salient foreground or background in an image contains small-scale high-contrast patterns. This issue is common in natural images and forms a fundamental challenge for prior methods. We tackle it from a scale point of view and propose a multi-layer approach to analyze saliency cues. The final saliency map is produced in a hierarchical model. Different from varying patch sizes or downsizing images, our scale-based region handling is by finding saliency values optimally in a tree model. Our approach improves saliency detection on many images that cannot be handled well traditionally. A new dataset is also constructed.
Qiong Yan, Li Xu 0001, Jianping Shi, Jiaya Jia
CVPR1
2013 Cross-Field Joint Image Restoration via Scale Map
abstract
Color, infrared, and flash images captured in different fields can be employed to effectively eliminate noise and other visual artifacts. We propose a two-image restoration framework considering input images in different fields, for example, one noisy color image and one dark-flashed near infrared image. The major issue in such a framework is to handle structure divergence and find commonly usable edges and smooth transition for visually compelling image reconstruction. We introduce a scale map as a competent representation to explicitly model derivative-level confidence and propose new functions and a numerical solver to effectively infer it following new structural observations. Our method is general and shows a principled way for cross-field restoration.
Qiong Yan, Xiaoyong Shen, Li Xu 0001, Shaojie Zhuo, Xiaopeng Zhang 0001, Liang Shen 0007, Jiaya Jia
ICCV1
2013 A sparse control model for image and video editing
abstract
It is common that users draw strokes, as control samples, to modify color, structure, or tone of a picture. We discover inherent limitation of existing methods for their implicit requirement on where and how the strokes are drawn, and present a new system that is principled on minimizing the amount of work put in user interaction. Our method automatically determines the influence of edit samples across the whole image jointly considering spatial distance, sample location, and appearance. It greatly reduces the number of samples that are needed, while allowing for a decent level of global and local manipulation of resulting effects and reducing propagation ambiguity. Our method is broadly beneficial to applications adjusting visual content.
Li Xu 0001, Qiong Yan, Jiaya Jia
ACM Trans. Graph.2
2012 Structure extraction from texture via relative total variation
abstract
It is ubiquitous that meaningful structures are formed by or appear over textured surfaces. Extracting them under the complication of texture patterns, which could be regular, near-regular, or irregular, is very challenging, but of great practical importance. We propose new inherent variation and relative total variation measures, which capture the essential difference of these two types of visual forms, and develop an efficient optimization system to extract main structures. The new variation measures are validated on millions of sample patches. Our approach finds a number of new applications to manipulate, render, and reuse the immense number of "structure with texture" images and drawings that were traditionally difficult to be edited properly.
Li Xu 0001, Qiong Yan, Jiaya Jia
ACM Trans. Graph.2