VLDB 2026 Research / reviewers in the wild / expert
Chaofeng Chen
dblp:198/2537
· DBLP profile ↗
47ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0001-6137-5162ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 9 first-author · 34 since 2021Artificial intelligence and machine learning · 28 · 5 first-author · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text4Seg++: Advancing Image Segmentation via Generative Language ModelingabstractMultimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains a significant challenge.In this work, we propose a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. We first introduce image-wise semantic descriptors, a patch-aligned textual representation of segmentation masks that integrates naturally into the language modeling pipeline. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptorsby 74% and accelerating inference by $3\times$3×, without compromising performance. Building upon this, our initial framework Text4Segachieves strong segmentation performance across a wide range of vision tasks. To further improve granularity and compactness, we propose box-wise semantic descriptors, which localizes regions of interest using bounding boxes and represents region masks via structured mask tokens called semantic bricks. This leads to our refined model, Text4Seg++, which formulates segmentation as a next-brick prediction task, combining precision, scalability, and generative efficiency. Comprehensive experiments on natural and remote sensing datasets show that Text4Seg++consistently outperforms state-of-the-art models across diverse benchmarks without any task-specific fine-tuning, while remaining compatible with existing MLLM backbones. Our work highlights the effectiveness, scalability, and generalizability of text-driven image segmentation within the MLLM framework. Mengcheng Lan, Chaofeng Chen, Jiaxing Xu, Zongrui Li 0001, Yiping Ke, Xudong Jiang 0001, Yingchen Yu, Yunqing Zhao, Song Bai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Denoised Semantic Features for Local Consistent No-Reference Image Quality AssessmentabstractMulti-dataset no-reference image quality assessment (NR-IQA) aims to deliver consistent image quality evaluation across a variety of contexts, empowering platform developers to optimize image processing pipelines while maintaining acceptable visual quality. Human vision, when observing images, tends to prioritize local semantics, for example, a blurry sky is perceived differently than a blurry face. This insight forms the basis of many multi-dataset NR-IQA models, which commonly rely on pretrained deep networks to extract semantic information that is crucial for assessing perceptual quality. Vision Transformer-based pre-trained models often exhibit persistent noise artifacts, as demonstrated by previous studies such as Denoising Vision Transformers; many existing IQA approaches fail to appropriately address these local semantic artifacts, leading to inconsistent local IQA score maps, even when overall performance appears satisfactory. To tackle this, we introduce DINO-IQA, a novel dual-branch network architecture designed for NR-IQA to multi-dataset. The first branch focuses on extracting local distortion features, effectively capturing image degradation, while the second branch utilizes denoised DINOv2 from ViT decomposition to extract refined semantic features, free from local artifacts. By enabling visual interaction between distortion and semantic features, our method generates locally consistent quality maps that align more closely with human perception. This approach achieves remarkable accuracy and sets a new benchmark for state-of-the-art multi-dataset NR-IQA performance. Our findings underscore the critical need to address semantic noise in pre-trained networks for enhancing NR-IQA, demonstrating that our dual-branch framework offers a robust solution to this previously underexplored challenge. Hui Li 0035, Chaofeng Chen, Xiaopeng Fan 0001, Wangmeng Zuo, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2026 | Prompt-Image-Caption Consistency for AI-Generated Image Quality AssessmentabstractWith the advancement of generative AI, generative models such as Generative Adversarial Networks (GANs) and diffusion-based models have significantly enhanced the capabilities in generating photorealistic images. Despite significant progress in AI-Generated Images (AIGIs), the evaluation of AIGI quality remains underexplored in multiple aspects, including AI artifacts, unnatural content, and the availability of input text prompts. However, existing methods are insufficient to capture both feature representations and correlations between image and text for AIGI quality assessment in aspects of perceptual image quality and text-to-image alignment. To address this, we proposed a novel Prompt-Image-Caption Consistency (PICC) framework, designed to adapt a pre-trained vision-language model for AIGI quality assessment by considering both perceptual image quality and text-to-image alignment at the same time. This framework effectively explores the prompt-image-caption triplet by adapting image embedding and multimodal embedding via Quality-Aware Attention to capture quality-aware features and calculating consistency scores to build the correlations among the prompt, image, and caption. Additionally, we propose a multilevel strategy to integrate both local and global information from multiple prompt-image-caption triplets, further enhancing prediction performance. Extensive experiments on benchmark datasets, including AIGIQA-20K and AGIQA-3K, demonstrate that the proposed PICC achieves state-of-the-art performance compared to baseline methods. The code will be made publicly available. Wen Sun 0015, Chaofeng Chen, Weisi Lin |
IEEE Trans. Multim. | 2 |
| 2025 | MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentabstractThe rapid growth of long-duration, high-definition videos has made efficient video quality assessment (VQA) a critical challenge. Existing research typically tackles this problem through two main strategies: reducing model parameters and resampling inputs. However, light-weight Convolution Neural Networks (CNN) and Transformers often struggle to balance efficiency with high performance due to the requirement of long-range modeling capabilities. Recently, the state-space model, particularly Mamba, has emerged as a promising alternative, offering linear complexity with respect to sequence length. Meanwhile, efficient VQA heavily depends on resampling long sequences to minimize computational costs, yet current resampling methods are often weak in preserving essential semantic information. In this work, we present MVQA, a Mamba-based model designed for efficient VQA along with a novel Unified Semantic and Distortion Sampling (USDS) approach. USDS combines semantic patch sampling from low-resolution videos and distortion patch sampling from original-resolution videos. The former captures semantically dense regions, while the latter retains critical distortion details. To prevent computation increase from dual inputs, we propose a fusion mechanism using pre-defined masks, enabling a unified sampling strategy that captures both semantic and quality information without additional computational burden. Experiments show that the proposed MVQA, equipped with USDS, achieve comparable performance to state-of-the-art methods while being $2\times$ as fast and requiring only $1/5$ GPU memory. Yachun Mi, Weicheng Meng, Chaofeng Chen, Shaohui Liu |
ICCV | 4 |
| 2025 | Text4Seg: Reimagining Image Segmentation as Text GenerationabstractMultimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74\% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework. Mengcheng Lan, Chaofeng Chen, Yue Zhou 0005, Jiaxing Xu, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001 |
ICLR | 2 |
| 2025 | Toward Transparent Deep Image Aesthetics Assessment With Tag-Based Content DescriptorsabstractDeep learning approaches for Image Aesthetics Assessment (IAA) have shown promising results in recent years, but the internal mechanisms of these models remain unclear. Previous studies have demonstrated that image aesthetics can be predicted using semantic features, such as pre-trained object classification features. However, these semantic features are learned implicitly, and therefore, previous works have not elucidated what the semantic features are representing. In this work, we aim to create a more transparent deep learning framework for IAA by introducing explainable semantic features. To achieve this, we propose Tag-based Content Descriptors (TCDs), where each value in a TCD describes the relevance of an image to a human-readable tag that refers to a specific type of image content. This allows us to build IAA models from explicit descriptions of image contents. We first propose the explicit matching process to produce TCDs that adopt predefined tags to describe image contents. We show that a simple MLP-based IAA model with TCDs only based on predefined tags can achieve an SRCC of 0.767, which is comparable to most state-of-the-art methods. However, predefined tags may not be sufficient to describe all possible image contents that the model may encounter. Therefore, we further propose the implicit matching process to describe image contents that cannot be described by predefined tags. By integrating components obtained from the implicit matching process into TCDs, the IAA model further achieves an SRCC of 0.817, which significantly outperforms existing IAA methods. Both the explicit matching process and the implicit matching process are realized by the proposed TCD generator. To evaluate the performance of the proposed TCD generator in matching images with predefined tags, we also labeled 5101 images with photography-related tags to form a validation set. And experimental results show that the proposed TCD generator can meaningfully assign photography-related tags to images. Jingwen Hou, Weisi Lin, Yuming Fang 0001, Haoning Wu 0001, Chaofeng Chen, Weide Liu |
IEEE Trans. Image Process. | 5 |
| 2024 | Iterative Token Evaluation and Refinement for Real-World Super-resolutionabstractReal-world image super-resolution (RWSR) is a long-standing problem as low-quality (LQ) images often have complex and unidentified degradations. Existing methods such as Generative Adversarial Networks (GANs) or continuous diffusion models present their own issues including GANs being difficult to train while continuous diffusion models requiring numerous inference steps. In this paper, we propose an Iterative Token Evaluation and Refinement (ITER) framework for RWSR, which utilizes a discrete diffusion model operating in the discrete token representation space, i.e., indexes of features extracted from a VQGAN codebook pre-trained with high-quality (HQ) images. We show that ITER is easier to train than GANs and more efficient than continuous diffusion models. Specifically, we divide RWSR into two sub-tasks, i.e., distortion removal and texture generation. Distortion removal involves simple HQ token prediction with LQ images, while texture generation uses a discrete diffusion model to iteratively refine the distortion removal output with a token refinement network. In particular, we propose to include a token evaluation network in the discrete diffusion process. It learns to evaluate which tokens are good restorations and helps to improve the iterative refinement results. Moreover, the evaluation network can first check status of the distortion removal output and then adaptively select total refinement steps needed, thereby maintaining a good balance between distortion removal and texture generation. Extensive experimental results show that ITER is easy to train and performs well within just 8 iterative steps. Chaofeng Chen, Shangchen Zhou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin |
AAAI | 1 |
| 2024 | Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation ModelsabstractMulti-modality large language models (MLLMs), as represented by GPT-4V, have introduced a paradigm shift for visual perception and understanding tasks, that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the identification of low-level visual attributes (e.g., clarity, brightness) to the evaluation on image quality, there's still an imperative to further improve the accuracy of MLLMs to substantially alleviate human burdens. To address this, we collect the first dataset consisting of human natural language feedback on low-level vision. Each feedback offers a comprehensive description of an image's low-level visual attributes, culminating in an overall quality assessment. The constructed Q-Pathway dataset includes 58K detailed human feedbacks on 18,973 multi-sourced images with diverse low-level appearance. To ensure MLLMs can adeptly handle diverse queries, we further propose a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct. Experimental results indicate that the Q-Instruct consistently elevates various low-level visual capabilities across multiple base models. We anticipate that our datasets can pave the way for a future that foundation models can assist humans on low-level visual tasks. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Kaixin Xu, Chunyi Li 0001, Jingwen Hou, Guangtao Zhai, Geng Xue, Wenxiu Sun, Qiong Yan, Weisi Lin |
CVPR | 4 |
| 2024 | Boosting Image Quality Assessment Through Efficient Transformer Adaptation with Local Feature EnhancementabstractImage Quality Assessment (IQA) constitutes a funda-mental task within the field of computer vision, yet it re-mains an unresolved challenge, owing to the intricate dis-tortion conditions, diverse image contents, and limited availability of data. Recently, the community has wit-nessed the emergence of numerous large-scale pretrained foundation models. However, it remains an open problem whether the scaling law in high-level tasks is also appli-cable to IQA tasks which are closely related to low-level clues. In this paper, we demonstrate that with a proper in-jection of local distortion features, a larger pretrained vision transformer (ViT) foundation model performs better in IQA tasks. Specifically, for the lack of local distortion structure and inductive bias of the large-scale pretrained ViT, we use another pretrained convolution neural networks (CNNs), which is well known for capturing the local structure, to extract multi-scale image features. Further, we propose a local distortion extractor to obtain local distortion features from the pretrained CNNs and a local distortion in-jector to inject the local distortion features into ViT. By only training the extractor and injector, our method can benefit from the rich knowledge in the powerful foundation models and achieve state-of-the-art performance on popular IQA datasets, indicating that IQA is not only a low-level problem but also benefits from stronger high-level features drawn from large-scale pretrained models. Codes are publicly available at: https://github.com/NeosXu/LoDa. Kangmin Xu, Jing Xiao 0004, Chaofeng Chen, Haoning Wu 0001, Qiong Yan, Weisi Lin |
CVPR | 4 |
| 2024 | Enhancing Diffusion Models with Text-Encoder Reinforcement Learning
Chaofeng Chen, Annan Wang, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin |
ECCV (25) | 1 |
| 2024 | ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001 |
ECCV (47) | 2 |
| 2024 | ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001 |
ECCV (68) | 2 |
| 2024 | Towards Open-Ended Visual Quality Comparison
Haoning Wu 0001, Hanwei Zhu, Erli Zhang 0001, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Xiaohong Liu 0001, Guangtao Zhai, Shiqi Wang 0001, Weisi Lin |
ECCV (3) | 5 |
| 2024 | Combining Generative and Geometry Priors for Wide-Angle Portrait Correction
Lan Yao, Chaofeng Chen, Xiaoming Li 0002, Zifei Yan, Wangmeng Zuo |
ECCV (29) | 2 |
| 2024 | Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionabstractThe rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding**. To address this gap, we present **Q-Bench**, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment. **_a)_** To evaluate the low-level **_perception_** ability, we construct the **LLVisionQA** dataset, consisting of 2,990 diverse-sourced images, each equipped with a human-asked question focusing on its low-level attributes. We then measure the correctness of MLLMs on answering these questions. **_b)_** To examine the **_description_** ability of MLLMs on low-level information, we propose the **LLDescribe** dataset consisting of long expert-labelled *golden* low-level text descriptions on 499 images, and a GPT-involved comparison pipeline between outputs of MLLMs and the *golden* descriptions. **_c)_** Besides these two tasks, we further measure their visual quality **_assessment_** ability to align with human opinion scores. Specifically, we design a softmax-based strategy that enables MLLMs to predict *quantifiable* quality scores, and evaluate them on various existing image quality assessment (IQA) datasets. Our evaluation across the three abilities confirms that MLLMs possess preliminary low-level visual skills. However, these skills are still unstable and relatively imprecise, indicating the need for specific enhancements on MLLMs towards these abilities. We hope that our benchmark can encourage the research community to delve deeper to discover and enhance these untapped potentials of MLLMs. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Chunyi Li 0001, Wenxiu Sun, Qiong Yan, Guangtao Zhai, Weisi Lin |
ICLR | 4 |
| 2024 | Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsabstractThe explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released. Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ICML | 4 |
| 2024 | T2I-Scorer: Quantitative Evaluation on Text-to-Image Generation via Fine-Tuned Large Multi-Modal ModelsabstractText-to-image (T2I) generation is a pivotal and core interest within the realm of AI content generation. Amid the swift advancements of both open-source (such as Stable Diffusion) and proprietary (for example, DALLE, MidJourney) T2I models, there is a notable absence of a comprehensive and robust quantitative framework for evaluating their output quality. Traditional methods of quality assessment overlook the textual prompts when judging images; meanwhile, the advent of large multi-modal models (LMMs) introduces the capability to incorporate text prompts in evaluations, yet the challenge of fine-tuning these models for precise T2I quality assessment remains unresolved. In our study, we introduce the T2I-Scorer, a novel two-stage training methodology aimed at fine-tuning LMMs for T2I evaluation. For the first stage, we collect 397K GPT-4V-labeled question-answer pairs related to T2I evaluation. Termed as T2I-ITD, the pseudo-labeled dataset is analyzed and examined by human, and used for instruction tuning to improve the LMM's low-level quality perception. The first stage model, T2I-Scorer-IT, has reached superior accuracy on T2I evaluation than all kinds of existing T2I metrics under zero-shot settings. For the second stage, we define an explicit multi-task training scheme to further align the LMM with human opinion scores, and the fine-tuned T2I-Scorer can reach state-of-the-art accuracy on both image quality and image-text alignment perspectives with significant improvements. We anticipate the proposed metrics can serve as a reliable metric to gauge the ability of T2I generation models in the future. We will make code, data, and weights publicly available. Haoning Wu 0001, Xiele Wu, Chunyi Li 0001, Chaofeng Chen, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 5 |
| 2024 | Q-Ground: Image Quality Grounding with Large Multi-modality ModelsabstractRecent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual content. However, these advancements are mostly focused on overall quality assessment, and the detailed examination of local quality, which is crucial for comprehensive visual understanding, is still largely unexplored. In this work, we introduce Q-Ground, the first framework aimed at tackling fine-scale visual quality grounding by combining large multi-modality models with detailed visual quality analysis. Cen- tral to our contribution is the introduction of the QGround-100K dataset, a novel resource containing 100k triplets of (image, quality text, distortion segmentation) to facilitate deep investigations into visual quality. The dataset comprises two parts: one with human- labeled annotations for accurate quality assessment, and another la- beled automatically by LMMs such as GPT4V, which helps improve the robustness of model training while also reducing the costs of data collection. With the QGround-100K dataset, we propose a LMM-based method equipped with multi-scale feature learning to learn models capable of performing both image quality answer- ing and distortion segmentation based on text prompts. This dual- capability approach not only refines the model’s understanding of region-aware image quality but also enables it to interactively re- spond to complex, text-based queries about image quality and spe- cific distortions. Q-Ground takes a step towards sophisticated vi- sual quality analysis in a finer scale, establishing a new benchmark for future research in the area. Codes and dataset are available at https://github.com/Q-Future/Q-Ground. Chaofeng Chen, Sensen Yang, Haoning Wu 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ACM Multimedia | 1 |
| 2024 | G-Refine: A General Quality Refiner for Text-to-Image Generation
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Tengchuan Kou, Chaofeng Chen, Lei Bai 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 6 |
| 2024 | LMM-PCQA: Assisting Point Cloud Quality Assessment with LMMabstractAlthough large multi-modality models (LMMs) have seen extensive exploration and application in various quality assessment studies, their integration into Point Cloud Quality Assessment (PCQA) remains unexplored. Given LMMs' exceptional performance and robustness in low-level vision and quality assessment tasks, this study aims to investigate the feasibility of imparting PCQA knowledge to LMMs through text supervision. To achieve this, we transform quality labels into textual descriptions during the fine-tuning phase, enabling LMMs to derive quality rating logits from 2D projections of point clouds. To compensate for the loss of perception in the 3D domain, structural features are extracted as well. These quality logits and structural features are then combined and regressed into quality scores. Our experimental results affirm the effectiveness of our approach, showcasing a novel integration of LMMs into PCQA that enhances model understanding and assessment accuracy. We hope our contributions can inspire subsequent investigations into the fusion of LMMs with PCQA, fostering advancements in 3D visual quality analysis and beyond. The code is available at https://github.com/zzc-1998/LMM-PCQA. Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Chaofeng Chen, Xiongkuo Min, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 6 |
| 2024 | Continual Learning of No-Reference Image Quality Assessment With Channel Modulation KernelabstractNo-Reference Image Quality Assessment (NR-IQA), a subset of IQA techniques, is critical in scenarios where reference images are unavailable. With advancements in camera technology and computer vision, IQA datasets have evolved significantly in distortion types, image contents, and domains. This highlights the need for a broad study of NR-IQA continual learning, optimizing on a sequence of tasks, in both in-domain and domain-transfer settings. In this paper, we introduce the Channel Modulation Kernel (CMKernel) as a solution to enhance NR-IQA continual learning from two perspectives. Firstly, CMKernel encodes channel attention information for both in-domain and domain-transfer scenarios. By imposing constraints on CMKernels of successive models, the channel attention distillation loss effectively mitigates the divergence between old and new models. Secondly, in the context of the domain-transfer setting, a significant challenge lies in training a robust and transferable base model from the general domain for subsequent continual learning across specific domains. To tackle this, we introduce CMKernel-based multi-dataset learning to acquire a generative model. By dynamically weighting convolutional channels, the base model learns more equally from mixed datasets, enhancing its performance for subsequent incremental tasks. Comprehensive experiments validate the superiority of CMKernel in both in-domain and domain-transfer continual learning settings, showcasing its efficacy in addressing the evolving challenges of NR-IQA in diverse image contexts. Hui Li 0035, Chaofeng Chen, Xiaopeng Fan 0001, Wangmeng Zuo, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | TOPIQ: A Top-Down Approach From Semantics to Distortions for Image Quality AssessmentabstractImage Quality Assessment (IQA) is a fundamental task in computer vision that has witnessed remarkable progress with deep neural networks. Inspired by the characteristics of the human visual system, existing methods typically use a combination of global and local representations (i.e., multi-scale features) to achieve superior performance. However, most of them adopt simple linear fusion of multi-scale features, and neglect their possibly complex relationship and interaction. In contrast, humans typically first form a global impression to locate important regions and then focus on local details in those regions. We therefore propose a top-down approach that uses high-level semantics to guide the IQA network to focus on semantically important local distortion regions, named as TOPIQ. Our approach to IQA involves the design of a heuristic coarse-to-fine network (CFANet) that leverages multi-scale features and progressively propagates multi-level semantic information to low-level representations in a top-down manner. A key component of our approach is the proposed cross-scale attention mechanism, which calculates attention maps for lower level features guided by higher level features. This mechanism emphasizes active semantic regions for low-level distortions, thereby improving performance. TOPIQ can be used for both Full-Reference (FR) and No-Reference (NR) IQA. We use ResNet50 as its backbone and demonstrate that TOPIQ achieves better or competitive performance on most public FR and NR benchmarks compared with state-of-the-art methods based on vision transformers, while being much more efficient (with only ∼ 13% FLOPS of the current best FR method). Codes are released at https://github.com/chaofengc/IQA-PyTorch. Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2024 | Blind Video Quality Prediction by Uncovering Human Video Perceptual RepresentationabstractBlind video quality assessment (VQA) has become an increasingly demanding problem in automatically assessing the quality of ever-growing in-the-wild videos. Although efforts have been made to measure temporal distortions, the core to distinguish between VQA and image quality assessment (IQA), the lack of modeling of how the human visual system (HVS) relates to the temporal quality of videos hinders the precise mapping of predicted temporal scores to the human perception. Inspired by the recent discovery of the temporal straightness law of natural videos in the HVS, this paper intends to model the complex temporal distortions of in-the-wild videos in a simple and uniform representation by describing the geometric properties of videos in the visual perceptual domain. A novel videolet, with perceptual representation embedding of a few consecutive frames, is designed as the basic quality measurement unit to quantify temporal distortions by measuring the angular and linear displacements from the straightness law. By combining the predicted score on each videolet, a perceptually temporal quality evaluator (PTQE) is formed to measure the temporal quality of the entire video. Experimental results demonstrate that the perceptual representation in the HVS is an efficient way of predicting subjective temporal quality. Moreover, when combined with spatial quality metrics, PTQE achieves top performance over popular in-the-wild video datasets. More importantly, PTQE requires no additional information beyond the video being assessed, making it applicable to any dataset without parameter tuning. Additionally, the generalizability of PTQE is evaluated on video frame interpolation tasks, demonstrating its potential to benefit temporal-related enhancement tasks. Kangmin Xu, Haoning Wu 0001, Chaofeng Chen, Wenxiu Sun, Qiong Yan, C.-C. Jay Kuo, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2023 | MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video PredictionabstractThe mainstream of the existing approaches for video prediction builds up their models based on a Single-In-Single-Out (SISO) architecture, which takes the current frame as input to predict the next frame in a recursive manner. This way often leads to severe performance degradation when they try to extrapolate a longer period of future, thus limiting the practical use of the prediction model. Alternatively, a Multi-In-Multi-Out (MIMO) architecture that outputs all the future frames at one shot naturally breaks the recursive manner and therefore prevents error accumulation. However, only a few MIMO models for video prediction are proposed and they only achieve inferior performance due to the date. The real strength of the MIMO model in this area is not well noticed and is largely under-explored. Motivated by that, we conduct a comprehensive investigation in this paper to thoroughly exploit how far a simple MIMO architecture can go. Surprisingly, our empirical studies reveal that a simple MIMO model can outperform the state-of-the-art work with a large margin much more than expected, especially in dealing with long-term error accumulation. After exploring a number of ways and designs, we propose a new MIMO architecture based on extending the pure Transformer with local spatio-temporal blocks and a new multi-output decoder, namely MIMO-VP, to establish a new standard in video prediction. We evaluate our model in four highly competitive benchmarks. Extensive experiments show that our model wins 1st place on all the benchmarks with remarkable performance gains and surpasses the best SISO model in all aspects including efficiency, quantity, and quality. A dramatic error reduction is achieved when predicting 10 frames on Moving MNIST and Weather datasets respectively. We believe our model can serve as a new baseline to facilitate the future research of video prediction tasks. The code will be released. Shuliang Ning, Mengcheng Lan, Yanran Li, Chaofeng Chen, Xunlai Chen, Xiaoguang Han 0001, Shuguang Cui |
AAAI | 4 |
| 2023 | Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesabstractThe rapid increase in user-generated content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the $\color{Green}{\text{technical perspective}}$, measuring the perception of distortions; and the $\color{Blue}{\text{aesthetic perspective}}$, which relates to preference and recommendation on contents. To understand how these two perspectives affect overall subjective opinions in UGC-VQA, we conduct a large-scale subjective study to collect human quality opinions on the overall quality of videos as well as perceptions from aesthetic and technical perspectives. The collected Disentangled Video Quality Database (DIVIDE-3k) confirms that human quality opinions on UGC videos are universally and inevitably affected by both aesthetic and technical perspectives. In light of this, we propose the Disentangled Objective Video Quality Evaluator (DOVER) to learn the quality of UGC videos based on the two perspectives. The DOVER proves state-of-the-art performance in UGC-VQA under very high efficiency. With perspective opinions in DIVIDE-3k, we further propose DOVER++, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective. Code at https://github.com/VQAssessment/DOVER. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ICCV | 4 |
| 2023 | Exploring Opinion-Unaware Video Quality Assessment with Semantic Affinity CriterionabstractRecent learning-based video quality assessment (VQA) algorithms are expensive to implement due to the cost of data collection of human quality opinions, and are less robust across various scenarios due to the biases of these opinions. This motivates our exploration on opinion-unaware (a.k.a zero-shot) VQA approaches. Existing approaches only considers low-level naturalness in spatial or temporal domain, without considering impacts from high-level semantics. In this work, we introduce an explicit semantic affinity index for opinion-unaware VQA using text-prompts in the contrastive language-image pre-training (CLIP) model. We also aggregate it with different traditional low-level naturalness indexes through gaussian normalization and sigmoid rescaling strategies. Composed of aggregated semantic and technical metrics, the proposed Blind Unified Opinion-Unaware Video Quality Index via Semantic and Technical Metric Aggregation (BUONA-VISTA) outperforms existing opinion-unaware VQA methods by at least 20% improvements, and is more robust than opinion-aware approaches. Haoning Wu 0001, Jingwen Hou, Chaofeng Chen, Erli Zhang 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ICME | 4 |
| 2023 | Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted ApproachabstractThe proliferation of in-the-wild videos has greatly expanded the Video Quality Assessment (VQA) problem. Unlike early definitions that usually focus on limited distortion types, VQA on in-the-wild videos is especially challenging as it could be affected by complicated factors, including various distortions and diverse contents. Though subjective studies have collected overall quality scores for these videos, how the abstract quality scores relate with specific factors is still obscure, hindering VQA methods from more concrete quality evaluations (e.g. sharpness of a video). To solve this problem, we collect over two million opinions on 4,543 in-the-wild videos on 13 dimensions of quality-related factors, including in-capture authentic distortions (e.g. motion blur, noise, flicker), errors introduced by compression and transmission, and higher-level experiences on semantic contents and aesthetic issues (e.g. composition, camera trajectory), to establish the multi-dimensional Maxwell database. Specifically, we ask the subjects to label among a positive, a negative, and a neutral choice for each dimension. These explanation-level opinions allow us to measure the relationships between specific quality factors and abstract subjective quality ratings, and to benchmark different categories of VQA algorithms on each dimension, so as to more comprehensively analyze their strengths and weaknesses. Furthermore, we propose the MaxVQA, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in our analyses. The MaxVQA can jointly evaluate various specific quality factors and final quality scores with state-of-the-art accuracy on all dimensions, and superb generalization ability on existing datasets. Code and data available at https://github.com/VQAssessment/MaxVQA. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ACM Multimedia | 4 |
| 2023 | Semi-supervised Cycle-GAN for face photo-sketch translation in the wild
Chaofeng Chen, Wei Liu 0091, Xiao Tan 0001, Kwan-Yee Kenneth Wong |
Comput. Vis. Image Underst. | 1 |
| 2023 | Neighbourhood Representative Sampling for Efficient End-to-End Video Quality AssessmentabstractThe increased resolution of real-world videos presents a dilemma between efficiency and accuracy for deep Video Quality Assessment (VQA). On the one hand, keeping the original resolution will lead to unacceptable computational costs. On the other hand, existing practices, such as resizing or cropping, will change the quality of original videos due to difference in details or loss of contents, and are henceforth harmful to quality assessment. With obtained insight from the studies of spatial-temporal redundancy in the human visual system, visual quality around a neighbourhood has high probability to be similar, and this motivates us to investigate an effective quality-sensitive neighbourhood representative sampling scheme for VQA. In this work, we propose a unified scheme, spatial-temporal grid mini-cube sampling (St-GMS), and the resultant samples are namedfragments. In St-GMS, full-resolution videos are first divided into mini-cubes with predefined spatial-temporal grids, then the temporal-aligned quality representatives are sampled to compose the fragments that serve as inputs for VQA. In addition, we design the Fragment Attention Network (FANet), a network architecture tailored specifically for fragments. With fragments and FANet, the proposedFAST-VQAandFasterVQA(with an improved sampling scheme) achieves up to 1612× efficiency than the existing state-of-the-art, meanwhile achieving significantly better performance on all relevant VQA benchmarks. Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Jinwei Gu, Weisi Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | DisCoVQA: Temporal Distortion-Content Transformers for Video Quality AssessmentabstractCompared with spatial counterparts, temporal relationships between frames and their influences on video quality assessment (VQA) are still relatively under-studied in existing works. These relationships lead to two important types of effects for video quality. Firstly, some meaningless temporal variations (such as shaking, flicker, and unsmooth scene transitions) cause temporal distortions that degrade quality of videos. Secondly, the human visual system often has different attention to frames with different contents, resulting in their different importance to the overall video quality. Based on prominent time-series modeling ability of transformers, we propose a novel and effective transformer-based VQA method to tackle these two issues. To better differentiate temporal variations and thus capture the temporal distortions, we design the Spatial-Temporal Distortion Extraction (STDE) module that extracts multi-level spatial-temporal features with a video swin transformer tiny (Swin-T) backbone and uses temporal difference layer to further capture these distortions. To tackle with temporal quality attention, we propose the encoder-decoder-like temporal content transformer (TCT). We also introduce the temporal sampling on features to reduce the input length for the TCT, so as to improve the learning effectiveness and efficiency of this module. Consisting of the STDE and the TCT, the proposed Temporal Distortion-Content Transformers for Video Quality Assessment (DisCoVQA) reaches state-of-the-art performance on several VQA benchmarks without any extra pre-training datasets and up to 10% better generalization ability than existing methods. We also conduct extensive ablation experiments to prove the effectiveness of each part in our proposed model, and provide visualizations to prove that the proposed modules achieve our intention on modeling these temporal issues. Our code is published athttps://github.com/QualityAssessment/DisCoVQA. Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Deep Face Video Inpainting via UV MappingabstractThis paper addresses the problem of face video inpainting. Existing video inpainting methods target primarily at natural scenes with repetitive patterns. They do not make use of any prior knowledge of the face to help retrieve correspondences for the corrupted face. They therefore only achieve sub-optimal results, particularly for faces under large pose and expression variations where face components appear very differently across frames. In this paper, we propose a two-stage deep learning method for face video inpainting. We employ 3DMM as our 3D face prior to transform a face between the image space and the UV (texture) space. In Stage I, we perform face inpainting in the UV space. This helps to largely remove the influence of face poses and expressions and makes the learning task much easier with well aligned face features. We introduce a frame-wise attention module to fully exploit correspondences in neighboring frames to assist the inpainting task. In Stage II, we transform the inpainted face regions back to the image space and perform face video refinement that inpaints any background regions not covered in Stage I and also refines the inpainted face regions. Extensive experiments have been carried out which show our method can significantly outperform methods based merely on 2D information, especially for faces under large pose and expression variations. Project page: https://ywq.github.io/FVIP. Wenqi Yang, Zhenfang Chen, Chaofeng Chen, Guanying Chen, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 3 |
| 2022 | From Face to Natural Image: Learning Real Degradation for Blind Image Super-Resolution
Xiaoming Li 0002, Chaofeng Chen, Xianhui Lin, Wangmeng Zuo, Lei Zhang 0006 |
ECCV (18) | 2 |
| 2022 | FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ECCV (6) | 2 |
| 2022 | PS-NeRF: Neural Inverse Rendering for Multi-view Photometric Stereo
Wenqi Yang, Guanying Chen, Chaofeng Chen, Zhenfang Chen, Kwan-Yee Kenneth Wong |
ECCV (1) | 3 |
| 2022 | A Unified Framework for Masked and Mask-Free Face Recognition Via Feature RectificationabstractFace recognition under ideal conditions is now considered a well-solved problem with advances in deep learning. Recognizing faces under occlusion, however, still remains a challenge. Existing techniques often fail to recognize faces with both the mouth and nose covered by a mask, which is now very common under the COVID-19 pandemic. Common approaches to tackle this problem include 1) discarding information from the masked regions during recognition and 2) restoring the masked regions before recognition. Very few works considered the consistency between features extracted from masked faces and from their mask-free counterparts. This resulted in models trained for recognizing masked faces of-ten showing degraded performance on mask-free faces. In this paper, we propose a unified framework, named Face Feature Rectification Network (FFR-Net), for recognizing both masked and mask-free faces alike. We introduce rectification blocks to rectify features extracted by a state-of-the-art recognition model, in both spatial and channel dimensions, to minimize the distance between a masked face and its mask-free counterpart in the rectified feature space. Experiments show that our unified framework can learn a rectified feature space for recognizing both masked and mask-free faces effectively, achieving state-of-the-art results. Project code: https://github.com/haoosz/FFR-Net Shaozhe Hao, Chaofeng Chen, Zhenfang Chen, Kwan-Yee Kenneth Wong |
ICIP | 2 |
| 2022 | Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution PriorsabstractA key challenge of real-world image super-resolution (SR) is to recover the missing details in low-resolution (LR) images with complex unknown degradations (\eg, downsampling, noise and compression). Most previous works restore such missing details in the image space. To cope with the high diversity of natural images, they either rely on the unstable GANs that are difficult to train and prone to artifacts, or resort to explicit references from high-resolution (HR) images that are usually unavailable. In this work, we propose Feature Matching SR (FeMaSR), which restores realistic HR images in a much more compact feature space. Unlike image-space methods, our FeMaSR restores HR images by matching distorted LR image features to their distortion-free HR counterparts in our pretrained HR priors, and decoding the matched features to obtain realistic HR images. Specifically, our HR priors contain a discrete feature codebook and its associated decoder, which are pretrained on HR images with a Vector Quantized Generative Adversarial Network (VQGAN). Notably, we incorporate a novel semantic regularization in VQGAN to improve the quality of reconstructed images. For the feature matching, we first extract LR features with an LR encoder consisting of several Swin Transformer blocks and then follow a simple nearest neighbour strategy to match them with the pretrained codebook. In particular, we equip the LR encoder with residual shortcut connections to the decoder, which is critical to the optimization of feature matching loss and also helps to complement the possible feature matching errors.Experimental results show that our approach produces more realistic HR images than previous methods. Code will be made publicly available. Chaofeng Chen, Yipeng Qin, Xiaoming Li 0001, Xiaoguang Han 0001, Shihui Guo |
ACM Multimedia | 1 |
| 2022 | Exploring the Effectiveness of Video Perceptual Representation in Blind Video Quality AssessmentabstractWith the rapid growth of in-the-wild videos taken by non-specialists, blind video quality assessment (VQA) has become a challenging and demanding problem. Although lots of efforts have been made to solve this problem, it remains unclear how the human visual system (HVS) relates to the temporal quality of videos. Meanwhile, recent work has found that the frames of natural video transformed into the perceptual domain of the HVS tend to form a straight trajectory of the representations. With the obtained insight that distortion impairs the perceived video quality and results in a curved trajectory of the perceptual representation, we propose a temporal perceptual quality index (TPQI) to measure the temporal distortion by describing the graphic morphology of the representation. Specifically, we first extract the video perceptual representations from the lateral geniculate nucleus (LGN) and primary visual area (V1) of the HVS, and then measure the straightness and compactness of their trajectories to quantify the degradation in naturalness and content continuity of video. Experiments show that the perceptual representation in the HVS is an effective way of predicting subjective temporal quality, and thus TPQI can, for the first time, achieve comparable performance to the spatial quality metric and be even more effective in assessing videos with large temporal variations. We further demonstrate that by combining with NIQE, a spatial quality metric, TPQI can achieve top performance over popular in-the-wild video datasets. More importantly, TPQI does not require any additional information beyond the video being evaluated and thus can be applied to any datasets without parameter tuning. Source code is available at https://github.com/UoLMM/TPQI-VQA. Kangmin Xu, Haoning Wu 0001, Chaofeng Chen, Wenxiu Sun, Qiong Yan, Weisi Lin |
ACM Multimedia | 4 |
| 2022 | S3-NeRF: Neural Reflectance Field from Shading and Shadow under a Single ViewpointabstractIn this paper, we address the "dual problem" of multi-view scene reconstruction in which we utilize single-view images captured under different point lights to learn a neural scene representation. Different from existing single-view methods which can only recover a 2.5D scene representation (i.e., a normal / depth map for the visible surface), our method learns a neural reflectance field to represent the 3D geometry and BRDFs of a scene. Instead of relying on multi-view photo-consistency, our method exploits two information-rich monocular cues, namely shading and shadow, to infer scene geometry. Experiments on multiple challenging datasets show that our method is capable of recovering 3D geometry, including both visible and invisible parts, of a scene from single-view images. Thanks to the neural reflectance field representation, our method is robust to depth discontinuities. It supports applications like novel-view synthesis and relighting. Our code and model can be found at https://ywq.github.io/s3nerf. Wenqi Yang, Guanying Chen, Chaofeng Chen, Zhenfang Chen, Kwan-Yee Kenneth Wong |
NeurIPS | 3 |
| 2021 | Progressive Semantic-Aware Style Transformation for Blind Face RestorationabstractFace restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transformation framework, named PSFR-GAN, for face restoration. Specifically, instead of using an encoder-decoder framework as previous methods, we formulate the restoration of LQ face images as a multi-scale progressive restoration procedure through semantic-aware style transformation. Given a pair of LQ face image and its corresponding parsing map, we first generate a multi-scale pyramid of the inputs, and then progressively modulate different scale features from coarse-to-fine in a semantic-aware style transfer way. Compared with previous networks, the proposed PSFR-GAN makes full use of the semantic (parsing maps) and pixel (LQ images) space information from different scales of input pairs. In addition, we further introduce a semantic aware style loss which calculates the feature style loss for each semantic region individually to improve the details of face textures. Finally, we pretrain a face parsing network which can generate decent parsing maps from real-world LQ face images. Experiment results show that our model trained with synthetic data can not only produce more realistic high-resolution results for synthetic LQ inputs but also generalize better to natural LQ face images compared with state-of-the-art methods. Chaofeng Chen, Xiaoming Li 0002, Lingbo Yang, Xianhui Lin, Lei Zhang 0006, Kwan-Yee Kenneth Wong |
CVPR | 1 |
| 2021 | HDR Video Reconstruction: A Coarse-to-fine Network and A Real-world Benchmark DatasetabstractHigh dynamic range (HDR) video reconstruction from sequences captured with alternating exposures is a very challenging problem. Existing methods often align low dynamic range (LDR) input sequence in the image space using optical flow, and then merge the aligned images to produce HDR output. However, accurate alignment and fusion in the image space are difficult due to the missing details in the over-exposed regions and noise in the under-exposed regions, resulting in unpleasing ghosting artifacts. To enable more accurate alignment and HDR fusion, we introduce a coarse-to-fine deep learning framework for HDR video reconstruction. Firstly, we perform coarse alignment and pixel blending in the image space to estimate the coarse HDR video. Secondly, we conduct more sophisticated alignment and temporal fusion in the feature space of the coarse HDR video to produce better reconstruction. Considering the fact that there is no publicly available dataset for quantitative and comprehensive evaluation of HDR video reconstruction methods, we collect such a benchmark dataset, which contains 97 sequences of static scenes and 184 testing pairs of dynamic scenes. Extensive experiments show that our method outperforms previous state-of-the-art methods. Our code and dataset can be found at https://guanyingc.github.io/DeepHDRVideo. Guanying Chen, Chaofeng Chen, Shi Guo, Zhetong Liang, Kwan-Yee Kenneth Wong, Lei Zhang 0006 |
ICCV | 2 |
| 2021 | Learning Spatial Attention for Face Super-ResolutionabstractGeneral image super-resolution techniques have difficulties in recovering detailed face structures when applying to low resolution face images. Recent deep learning based methods tailored for face images have achieved improved performance by jointly trained with additional task such as face parsing and landmark prediction. However, multi-task learning requires extra manually labeled data. Besides, most of the existing works can only generate relatively low resolution face images (e.g., 128×128 ), and their applications are therefore limited. In this paper, we introduce a novel SPatial Attention Residual Network (SPARNet) built on our newly proposed Face Attention Units (FAUs) for face super-resolution. Specifically, we introduce a spatial attention mechanism to the vanilla residual blocks. This enables the convolutional layers to adaptively bootstrap features related to the key face structures and pay less attention to those less feature-rich regions. This makes the training more effective and efficient as the key face structures only account for a very small portion of the face image. Visualization of the attention maps shows that our spatial attention network can capture the key face structures well even for very low resolution faces (e.g., 16×16 ). Quantitative comparisons on various kinds of metrics (including PSNR, SSIM, identity similarity, and landmark detection) demonstrate the superiority of our method over current state-of-the-arts. We further extend SPARNet with multi-scale discriminators, named as SPARNetHD, to produce high resolution results (i.e., 512×512 ). We show that SPARNetHD trained with synthetic data can not only produce high quality and high resolution outputs for synthetically degraded face images, but also show good generalization ability to real world low quality face images. Codes are available at https://github.com/chaofengc/Face-SPARNet. Chaofeng Chen, Dihong Gong, Hao Wang 0050, Zhifeng Li 0001, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 1 |
| 2020 | Blind Face Restoration via Deep Multi-scale Component Dictionaries
Xiaoming Li 0002, Chaofeng Chen, Shangchen Zhou, Xianhui Lin, Wangmeng Zuo, Lei Zhang 0006 |
ECCV (9) | 2 |
| 2018 | Char-Net: A Character-Aware Neural Network for Distorted Scene Text RecognitionabstractIn this paper, we present a Character-Aware Neural Network (Char-Net) for recognizing distorted scene text. Our Char-Net is composed of a word-level encoder, a character-level encoder, and a LSTM-based decoder. Unlike previous work which employed a global spatial transformer network to rectify the entire distorted text image, we take an approach of detecting and rectifying individual characters. To this end, we introduce a novel hierarchical attention mechanism (HAM) which consists of a recurrent RoIWarp layer and a character-level attention layer. The recurrent RoIWarp layer sequentially extracts a feature region corresponding to a character from the feature map produced by the word-level encoder, and feeds it to the character-level encoder which removes the distortion of the character through a simple spatial transformer and further encodes the character region. The character-level attention layer then attends to the most relevant features of the feature map produced by the character-level encoder and composes a context vector, which is finally fed to the LSTM-based decoder for decoding. This approach of adopting a simple local transformation to model the distortion of individual characters not only results in an improved efficiency, but can also handle different types of distortion that are hard, if not impossible, to be modelled by a single global transformation. Experiments have been conducted on six public benchmark datasets. Our results show that Char-Net can achieve state-of-the-art performance on all the benchmarks, especially on the IC-IST which contains scene text with large distortion. Code will be made available. Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong |
AAAI | 2 |
| 2018 | Semi-supervised Learning for Face Sketch Synthesis in the Wild
Chaofeng Chen, Wei Liu 0091, Xiao Tan 0001, Kwan-Yee Kenneth Wong |
ACCV (1) | 1 |
| 2018 | SAFE: Scale Aware Feature Encoder for Scene Text Recognition
Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong |
ACCV (2) | 2 |
| 2018 | Face Sketch Synthesis with Style Transfer Using Pyramid Column FeatureabstractIn this paper, we propose a novel framework based on deep neural networks for face sketch synthesis from a photo. Imitating the process of how artists draw sketches, our framework synthesizes face sketches in a cascaded manner. A content image is first generated that outlines the shape of the face and the key facial features. Textures and shadings are then added to enrich the details of the sketch. We utilize a fully convolutional neural network (FCNN) to create the content image, and propose a style transfer approach to introduce textures and shadings based on a newly proposed pyramid column feature. We demonstrate that our style transfer approach based on the pyramid column feature can not only preserve more sketch details than the common style transfer method, but also surpasses traditional patch based methods. Quantitative and qualitative evaluations suggest that our framework outperforms other state-of-the-arts methods, and can also generalize well to different test images. Chaofeng Chen, Xiao Tan 0001, Kwan-Yee Kenneth Wong |
WACV | 1 |
| 2016 | STAR-Net: A SpaTial Attention Residue Network for Scene Text Recognition
Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong, Zhizhong Su, Junyu Han |
BMVC | 2 |