VLDB 2026 Research / reviewers in the wild / expert
Lingyu Zhu 0006
dblp:282/1603
· DBLP profile ↗
29ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0001-7608-7913ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 25 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality AssessmentabstractDespite the impressive performance of large multimodal models (LMMs) in high-level visual tasks, their capacity for image quality assessment (IQA) remains limited. One main reason is that LMMs are primarily trained for high-level tasks (e.g., image captioning), emphasizing unified image semantics extraction under varied quality. Such semantic-aware yet quality-insensitive perception bias inevitably leads to a heavy reliance on image semantics when those LMMs are forced for quality rating. In this paper, instead of retraining or tuning an LMM costly, we propose a training-free debiasing framework, in which the image quality prediction is rectified by mitigating the bias caused by image semantics. Specifically, we first explore several semantic-preserving distortions that can significantly degrade image quality while maintaining identifiable semantics. By applying these specific distortions to the query/test images, we ensure that the degraded images are recognized as poor quality while their semantics remain. During quality inference, both a query image and its corresponding degraded version are fed to the LMM along with a prompt indicating that the query image quality should be inferred under the condition that the degraded one is deemed poor quality. This prior condition effectively aligns the LMM’s quality perception, as all degraded images are consistently rated as poor quality, regardless of their semantic difference. Finally, the quality scores of the query image inferred under different prior conditions (degraded versions) are aggregated using a conditional probability model. Extensive experiments on various IQA datasets show that our debiasing framework could consistently enhance the LMM performance and the code will be publicly available. Baoliang Chen, Siyi Pan, Dongxu Wu, Liang Xie 0013, Xiangjie Sui, Lingyu Zhu 0006, Hanwei Zhu |
AAAI | 6 |
| 2026 | Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality AssessmentabstractRecent efforts have repurposed the Contrastive Language-Image Pre-training (CLIP) model for No-Reference Image Quality Assessment (NR-IQA) by measuring the cosine similarity between the image embedding and textual prompts such as "a good photo" or "a bad photo." However, this semantic similarity overlooks a critical yet underexplored cue: the magnitude of the CLIP image features, which we empirically find to exhibit a strong correlation with perceptual quality. In this work, we introduce a novel adaptive fusion framework that complements cosine similarity with a magnitude-aware quality cue. Specifically, we first extract the absolute CLIP image features and apply a Box-Cox transformation to statistically normalize the feature distribution and mitigate semantic sensitivity. The resulting scalar summary serves as a semantically-normalized auxiliary cue that complements cosine-based prompt matching. To integrate both cues effectively, we further design a confidence-guided fusion scheme that adaptively weighs each term according to its relative strength. Extensive experiments on multiple benchmark IQA datasets demonstrate that our method consistently outperforms standard CLIP-based IQA and state-of-the-art baselines, without any task-specific training. Zhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai, Hanwei Zhu, Lingyu Zhu 0006, Yuncheng Jiang 0004, Baoliang Chen |
AAAI | 6 |
| 2026 | Temporal Quality Aggregation for VQA: Benchmark and Psychology-Inspired Model
Baoliang Chen, Changsheng Gao, Lingyu Zhu 0006, Liang Xie 0013, Hanwei Zhu, Zhijian Hao |
QoMEX | 3 |
| 2026 | Low-Light Image Enhancement via Diffusion Models With Semantic Priors of Any Region
Lingyu Zhu 0006, Wenhan Yang, Howard Leung, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Optimizing Fidelity-Perception Tradeoff via Large Vision-Language Model Prior for Image CompressionabstractCurrent neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity. Yudong Mao, Peilin Chen 0001, Lingyu Zhu 0006, Yung-Hui Li, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual TokensabstractRecently, deep generative models have greatly advanced the progress of face video coding towards promising rate-distortion performance and diverse application functionalities. Beyond traditional hybrid video coding paradigms, Generative Face Video Compression (GFVC) relying on the strong capabilities of deep generative models and the philosophy of early Model-Based Coding (MBC) can facilitate the compact representation and realistic reconstruction of visual face signal, thus achieving ultra-low bitrate face video communication. However, these GFVC algorithms are sometimes faced with unstable reconstruction quality and limited bitrate ranges. To address these problems, this paper proposes a novel Progressive Face Video Compression framework, namely PFVC, that utilizes adaptive visual tokens to realize exceptional trade-offs between reconstruction robustness and bandwidth intelligence. In particular, the encoder of the proposed PFVC projects the high-dimensional face signal into adaptive visual tokens in a progressive manner, whilst the decoder can further reconstruct these adaptive visual tokens for motion estimation and signal synthesis with different granularity levels. Experimental results demonstrate that the proposed PFVC framework can achieve better coding flexibility and superior rate-distortion performance in comparison with the latest Versatile Video Coding (VVC) codec and the state-of-the-art GFVC algorithms. The project page can be found at https://github.com/Berlin0610/PFVC. Shanzhi Yin, Jie Chen 0006, Ru-Ling Liao, Lingyu Zhu 0006, Shiqi Wang 0001, Yan Ye 0003 |
DCC | 6 |
| 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionabstractThis work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate geometric and contextual information within adaptive regions in each image and dynamically adjust the contributions from different views, enhancing the representation capability of voxel features. Furthermore, we propose a sparse volume construction strategy that adaptively identifies and selects voxels with high occupancy probabilities for feature refinement, minimizing redundant computation in free space. Benefiting from the above designs, our framework achieves effective and efficient volume construction in an adaptive way. Better still, our network can be supervised using only 3D bounding boxes, eliminating the dependence on ground-truth scene geometry. Experimental results demonstrate that SGCDet achieves state-of-the-art performance on the ScanNet, ScanNet200 and ARKitScenes datasets. The source code is available at https://github.com/RM-Zhang/SGCDet. Runmin Zhang, Zhu Yu 0001, Si-Yuan Cao, Lingyu Zhu 0006, Guangyi Zhang 0005, Xiaokai Bai |
ICCV | 4 |
| 2025 | The Loop Game: Quality Assessment and Optimization for Low-Light Image Enhancement
Danni Huang, Lingyu Zhu 0006, Hanwei Zhu, Shiqi Wang 0001, Baoliang Chen |
ICIC (3) | 2 |
| 2025 | Compressing Human Body Video with Interactive Semantics: A Generative ApproachabstractIn this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body signal into a series of configurable embeddings, which are controllably edited, compactly compressed, and efficiently transmitted. Moreover, the proposed decoder can evolve the mesh-based motion fields from these decoded semantics to realize the high-quality human body video reconstruction. Experimental results illustrate that the proposed framework can achieve promising compression performance for human body videos at ultra-low bitrate ranges compared with the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes. Furthermore, the proposed framework enables interactive human body video coding without any additional pre-/post-manipulation processes, which is expected to shed light on metaverse-related digital human communication in the future. Shanzhi Yin, Hanwei Zhu, Lingyu Zhu 0006, Jie Chen 0006, Ru-Ling Liao, Shiqi Wang 0001, Yan Ye 0003 |
ICIP | 4 |
| 2025 | Exploiting Long and Short Temporal Dependence for Low-Light Video EnhancementabstractExisting learning-based methods often lack temporal coherence in low-light video enhancement due to rarely considering intrinsic temporal dependence. To address this issue, we propose the Long-short Temporal Filtering Network (TFNet) to learn the mapping from low-light videos to normal-light ones, utilizing the well-considered data-centric strategy and a refined architecture. From the data-centric temporal strategy, we incorporate both long-range and short-range temporal dependence into TFNet, effectively capturing the temporal information. From the model design perspective, the TFNet incorporates the Temporal-aware Attentional Filtering (TAF) module, which aims to estimate and adaptively combine filtering kernels for guided filtering towards features of the middle frame. To further refine the filtered features, the cascaded Grouped Attention (GA) blocks are presented in a grouped attention strategy. Experimental results on benchmark datasets have demonstrated the superiority of our TFNet against the state-of-the-art methods in terms of video frame quality and brightness consistency. Lingyu Zhu 0006, Yudong Mao, Zhiwei Zhong 0001, Shanshe Wang, Shiqi Wang 0001 |
ICME | 2 |
| 2025 | MS-MoE: Multi-modal Structural Mixture of Experts Framework for Pan-SharpeningabstractPan-sharpening aims to generate the high-resolution (HR) multi-spectral (MS) target image from its low-resolution (LR) counterpart, which is guided by corresponding HR panchromatic (PAN) image with abundant texture structural details. Although the existing state-of-the-art methods have made remarkable progress, they are still struggling with integrating inherent structural correlation between PAN and MS images through the early or late-stage fusion alone. This would lead to texture-less pan-sharpening reconstruction due to the insufficient learning of complementary features from PAN image. To address this issue, we propose the Multi-modal Structural Mixture of Experts (MS-MoE) framework for pan-sharpening. Specifically, given the upsampled LRMS and PAN images spatially rotated at various angles, we design a set of structural experts to extract the complementary spatial and spectral features between them, in which the Texture Enhancement Module (TEM) is introduced to extract and enhance texture-structural features from different modalities. Subsequently, we introduce an additional expert network to perform feature fusion by integrating the outputs from multiple experts. To reconstruct the high-frequency information, we further leverage the Frequency feature Refinement Module (FRM) to aggregate and refine the fused features in the frequency domain. Experimental results on the benchmark pan-sharpening datasets demonstrate that the proposed MS-MoE framework achieves more competitive performance than recent state-of-the-art methods. Zhiwei Zhong 0001, Lingyu Zhu 0006, Yudong Mao, Shiqi Wang 0001 |
IJCNN | 3 |
| 2025 | Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
Jiaxi Huang, Dongxu Wu, Hanwei Zhu, Lingyu Zhu 0006, Jun Xing, Xu Wang 0006, Baoliang Chen |
PRCV (8) | 4 |
| 2025 | Simple Lines, Big Ideas: Towards Interpretable Assessment of Human Creativity from Drawings
Zhenshan Shi, Sasa Zhao, Hanwei Zhu, Lingyu Zhu 0006, Baoliang Chen, Lei Mo |
PRCV (9) | 5 |
| 2025 | DeepDC: Deep Distance Correlation as a Perceptual Image Quality EvaluatorabstractDeep neural networks pre-trained on ImageNet have demonstrated remarkable transferability for developing effective full-reference image quality assessment (FR-IQA) models. However, existing approaches typically demand pixel-level alignment between reference and distorted images-a requirement that poses significant challenges in practical scenarios involving natural photography and texture similarity evaluation. To address this limitation, we propose a novel FR-IQA model leveraging deep statistical similarity derived from pre-trained features without relying on spatial co-location of these features or requiring fine-tuning with mean opinion scores. Specifically, we employ distance correlation, a potent yet relatively underexplored statistical measure, to quantify similarity between reference and distorted images within a deep feature space. The distance correlation is computed via the ratio of the distance covariance to the product of their respective distance standard deviations, for which we derive a closed-form solution using the inner product of deep double-centered distance matrices. Extensive experimental evaluations across diverse IQA benchmarks demonstrate the superiority and robustness of the proposed model. Furthermore, we demonstrate the utility of our model for optimizing texture synthesis and neural style transfer tasks, achieving state-of-the-art performance in both quantitative measures and qualitative assessments. The implementation is publicly available at https://github.com/h4nwei/DeepDC. Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Shiqi Wang 0001, Weisi Lin |
IEEE Trans. Image Process. | 3 |
| 2025 | Debiased Mapping for Full-Reference Image Quality AssessmentabstractAn ideal full-reference image quality (FR-IQA) model should exhibit both high separability for images with different quality and compactness for images with the same or indistinguishable quality. However, existing learning-based FR-IQA models that directly compare images in deep-feature space, usually overly emphasize the quality separability, neglecting to maintain the compactness when images are of similar quality. In our work, we identify that the perception bias mainly stems from an inappropriate subspace where images are projected and compared. For this issue, we propose a Debiased Mapping based quality Measure (DMM), leveraging orthonormal bases formed by singular value decomposition (SVD) in the deep features domain. The SVD effectively decomposes the quality variations into singular values and mapping bases, enabling quality inference with more reliable feature difference measures. Extensive experimental results reveal that our proposed measure could mitigate the perception bias effectively and demonstrates excellent quality prediction performance on various IQA datasets. Baoliang Chen, Hanwei Zhu, Lingyu Zhu 0006, Shanshe Wang, Jingshan Pan, Shiqi Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | RCNet: Deep Recurrent Collaborative Network for Multi-View Low-Light Image EnhancementabstractScene observation from multiple perspectives would bring a more comprehensive visual experience. However, in the context of acquiring multiple views in the dark, the highly correlated views are seriously alienated, making it challenging to improve scene understanding with auxiliary views. Recent single image-based enhancement methods may not be able to provide consistently desirable restoration performance for all views due to the ignorance of potential feature correspondence among different views. To alleviate this issue, we make the first attempt to investigate multi-view low-light image enhancement. First, we construct a new dataset called Multi-View Low-light Triplets (MVLT), including 1,860 pairs of triple images with large illumination ranges and wide noise distribution. Each triplet is equipped with three different viewpoints towards the same scene. Second, we propose a deep multi-view enhancement framework based on the Recurrent Collaborative Network (RCNet). Specifically, in order to benefit from similar texture correspondence across different views, we design the recurrent feature enhancement, alignment and fusion (ReEAF) module, in which intra-view feature enhancement (Intra-view EN) followed by inter-view feature alignment and fusion (Inter-view AF) is performed to model the intra-view and inter-view feature propagation sequentially via multi-view collaboration. In addition, two different modules from enhancement to alignment (E2A) and from alignment to enhancement (A2E) are developed to enable the interactions between Intra-view EN and Inter-view AF, which explicitly utilize attentive feature weighting and sampling for enhancement and alignment, respectively. Experimental results demonstrate that our RCNet significantly outperforms other state-of-the-art methods. All of our dataset, code, and model will be available athttps://github.com/hluo29/RCNet. Baoliang Chen, Lingyu Zhu 0006, Peilin Chen 0001, Shiqi Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Zhangkai Ni, Qi Mao 0002, Shiqi Wang 0001 |
ECCV (23) | 1 |
| 2024 | Learned Image Compression for Both Humans and Machines via Dynamic AdaptationabstractRecent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs. Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001 |
ICIP | 1 |
| 2024 | Diffusion-Based Bit-Depth ExpansionabstractDiffusion-based generative models have achieved remarkable success across a variety of applications. However, the potential application for bit-depth expansion has not been extensively studied. This paper introduces a wavelet-based diffusion model for the bit-depth expansion task. In this method, the image is first decomposed into low and high-frequency components via wavelet transformation. This decomposition allows for targeted processing by specialized modules and reduces computational complexity by lowering the image resolution. The low-frequency component is processed in both the forward diffusion and reverse denoising stages. Meanwhile, the high-frequency components are filtered by the High Frequency Denoising Filter (HFDF) to eliminate noise and artifacts. Finally, the low and high-frequency components are recombined into a predicted high-bit-depth image through inverse wavelet transformation. Experimental results demonstrate the superiority of the proposed method in producing perceptually compelling outputs that outperform previous methods. Riyu Lu, Lingyu Zhu 0006, Baoliang Chen, Xiaopeng Fan 0001, Shiqi Wang 0001 |
MMSP | 2 |
| 2024 | Adaptive Image Quality Assessment via Teaching Large Multimodal Model to CompareabstractWhile recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplored. To address this gap, we introduce an all-around LMM-based NR-IQA model, which is capable of producing qualitatively comparative responses and effectively translating these discrete comparison outcomes into a continuous quality score. Specifically, during training, we present to generate scaled-up comparative instructions by comparing images from the same IQA dataset, allowing for more flexible integration of diverse IQA datasets. Utilizing the established large-scale training corpus, we develop a human-like visual quality comparator. During inference, moving beyond binary choices, we propose a soft comparison method that calculates the likelihood of the test image being preferred over multiple predefined anchor images. The quality score is further optimized by maximum a posteriori estimation with the resulting probability matrix. Extensive experiments on nine IQA datasets validate that the Compare2Score effectively bridges text-defined comparative levels during training with converted single image quality scores for inference, surpassing state-of-the-art IQA models across diverse scenarios. Moreover, we verify that the probability-matrix-based inference conversion not only improves the rating accuracy of Compare2Score but also zero-shot general-purpose LMMs, suggesting its intrinsic effectiveness. Hanwei Zhu, Haoning Wu 0001, Baoliang Chen, Lingyu Zhu 0006, Yuming Fang 0001, Guangtao Zhai, Weisi Lin, Shiqi Wang 0001 |
NeurIPS | 6 |
| 2024 | Temporally Consistent Enhancement of Low-Light Videos via Spatial-Temporal Compatible LearningabstractAbstract Temporal inconsistency is the annoying artifact that has been commonly introduced in low-light video enhancement, but current methods tend to overlook the significance of utilizing both data-centric clues and model-centric design to tackle this problem. In this context, our work makes a comprehensive exploration from the following three aspects. First, to enrich the scene diversity and motion flexibility, we construct a synthetic diverse low/normal-light paired video dataset with a carefully designed low-light simulation strategy, which can effectively complement existing real captured datasets. Second, for better temporal dependency utilization, we develop a Temporally Consistent Enhancer Network (TCE-Net) that consists of stacked 3D convolutions and 2D convolutions to exploit spatial-temporal clues in videos. Last, the temporal dynamic feature dependencies are exploited to obtain consistency constraints for different frame indexes. All these efforts are powered by a Spatial-Temporal Compatible Learning (STCL) optimization technique, which dynamically constructs specific training loss functions adaptively on different datasets. As such, multiple-frame information can be effectively utilized and different levels of information from the network can be feasibly integrated, thus expanding the synergies on different kinds of data and offering visually better results in terms of illumination distribution, color consistency, texture details, and temporal coherence. Extensive experimental results on various real-world low-light video datasets clearly demonstrate the proposed method achieves superior performance to state-of-the-art methods. Our code and synthesized low-light video database will be publicly available at https://github.com/lingyzhu0101/low-light-video-enhancement.git . Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Xiandong Meng, Shiqi Wang 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Video Quality Assessment for Spatio-Temporal Resolution Adaptive CodingabstractSpatio-temporal resolution adaptive (STRA) coding has been repeatedly proven to be a promising way to improve coding efficiency and reduce coding complexity. The wide consensus is that the optimal subsampled resolution and frame rate should be governed by so- called generalized rate-distortion performance based on the ultimately perceived distortion. However, it is non-trivial to accurately predict the quality of reconstructed videos due to the fact that the distortion originates from both subsampling and compression. To address this issue, we propose a novel video quality assessment model that is fully aware of the information available in downsampled videos for compression, such as resolution and frame rate. More specifically, the proposed model relies on quality-aware spatial features that are extracted by an image quality fine-tuned backbone. Subsequently, the spatio-temporal quality is modeled based on the transformer encoder, which is adaptive to the downsampling spatial and temporal resolutions. This enables the transformer encoder to produce discriminative features that capture long-range temporal dependencies related to the current context. The quality score, which is the output of the transformer encoder, thus reflects both the influence of the subsampling and compression. We conduct extensive experiments that demonstrate the superiority of the proposed model over state-of-the-art methods on four subsampling and compression video quality datasets. Furthermore, we apply the proposed model to bitrate ladder optimization, leading to a perceptual-aware spatial and temporal downsampling strategy that yields promising bitrate savings. The source codes of the proposed model will be publicly available athttps://github.com/h4nwei/STRA-VQA. Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Peilin Chen 0001, Linqi Song, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Deep Feature Statistics Mapping for Generalized Screen Content Image Quality AssessmentabstractThe statistical regularities of natural images, referred to as natural scene statistics, play an important role in no-reference image quality assessment. However, it has been widely acknowledged that screen content images (SCIs), which are typically computer generated, do not hold such statistics. Here we make the first attempt to learn the statistics of SCIs, based upon which the quality of SCIs can be effectively determined. The underlying mechanism of the proposed approach is based upon the mild assumption that the SCIs, which are not physically acquired, still obey certain statistics that could be understood in a learning fashion. We empirically show that the statistics deviation could be effectively leveraged in quality assessment, and the proposed method is superior when evaluated in different settings. Extensive experimental results demonstrate the Deep Feature Statistics based SCI Quality Assessment (DFSS-IQA) model delivers promising performance compared with existing NR-IQA models and shows a high generalization capability in the cross-dataset settings. The implementation of our method is publicly available at https://github.com/Baoliang93/DFSS-IQA. Baoliang Chen, Hanwei Zhu, Lingyu Zhu 0006, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2024 | Gap-Closing Matters: Perceptual Quality Evaluation and Optimization of Low-Light Image EnhancementabstractThere is a growing consensus in the research community that the optimization of low-light image enhancement approaches should be guided by the visual quality perceived by end users. Despite the substantial efforts invested in the design of low-light enhancement algorithms, there has been comparatively limited focus on assessing subjective and objective quality systematically. To mitigate this gap and provide a clear path towards optimizing low-light image enhancement for better visual quality, we propose a gap-closing framework. In particular, our gap-closing framework starts with the creation of a large-scale dataset for Subjective QUality Assessment of REconstructed LOw-Light Images (SQUARE-LOL). This database serves as the foundation for studying the quality of enhanced images and conducting a comprehensive subjective user study. Subsequently, we propose an objective quality assessment measure that plays a critical role in bridging the gap between visual quality and enhancement. Finally, we demonstrate that our proposed objective quality measure can be incorporated into the process of optimizing the learning of the enhancement model toward perceptual optimality. We validate the effectiveness of our proposed framework through both the accuracy of quality prediction and the perceptual quality of image enhancement. Baoliang Chen, Lingyu Zhu 0006, Hanwei Zhu, Wenhan Yang, Linqi Song, Shiqi Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Learning Spatiotemporal Interactions for User-Generated Video Quality AssessmentabstractDistortions from spatial and temporal domains have been identified as the dominant factors that govern the visual quality. Though both have been studied independently in deep learning-based user-generated content (UGC) video quality assessment (VQA) by frame-wise distortion estimation and temporal quality aggregation, much less work has been dedicated to the integration of them with deep representations. In this paper, we propose a SpatioTemporal Interactive VQA (STI-VQA) model based upon the philosophy that video distortion can be inferred from the integration of both spatial characteristics and temporal motion, along with the flow of time. In particular, for each timestamp, both the spatial distortion explored by the feature statistics and local motion captured by feature difference are extracted and fed to a transformer network for the motion aware interaction learning. Meanwhile, the information flow of spatial distortion from the shallow layer to the deep layer is constructed adaptively during the temporal aggregation. The transformer network enjoys an advanced advantage for long-range dependencies modeling, leading to superior performance on UGC videos. Experimental results on five UGC video benchmarks demonstrate the effectiveness and efficiency of our STI-VQA model, and the source code will be available online athttps://github.com/h4nwei/STI-VQA. Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality AssessmentabstractIn this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method. The codes are released athttps://github.com/Baoliang93/GSTVQA Baoliang Chen, Lingyu Zhu 0006, Fangbo Lu, Hongfei Fan, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Enlightening Low-Light Images With Dynamic Guidance for Context EnrichmentabstractImages acquired in low-light conditions suffer from a series of visual quality degradations,e.g., low visibility, degraded contrast, and intensive noise. These complicated degradations based on various contexts (e.g., noise in smooth regions, over-exposure in well-exposed regions and low contrast around edges) cast major challenges to the low-light image enhancement. Herein, we propose a new methodology by imposing a learnable guidance map from the signal and deep priors, making the deep neural network adaptively enhance low-light images in a region-dependent manner. The enhancement capability of the learnable guidance map is further exploited with the multi-scale dilated context collaboration, leading to contextually enriched feature representations extracted by the model with various receptive fields. Through assimilating the intrinsic perceptual information from the learned guidance map, richer and more realistic textures are generated. Extensive experiments on real low-light images demonstrate the effectiveness of our method, which delivers superior results quantitatively and qualitatively. The code is available athttps://github.com/lingyzhu0101/GEMSCto facilitate future research. Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Fangbo Lu, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | No-Reference Image Quality Assessment by Hallucinating Pristine FeaturesabstractIn this paper, we propose a no-reference (NR) image quality assessment (IQA) method via feature level pseudo-reference (PR) hallucination. The proposed quality assessment framework is rooted in the view that the perceptually meaningful features could be well exploited to characterize the visual quality, and the natural image statistical behaviors are exploited in an effort to deliver the accurate predictions. Herein, the PR features from the distorted images are learned by a mutual learning scheme with the pristine reference as the supervision, and the discriminative characteristics of PR features are further ensured with the triplet constraints. Given a distorted image for quality inference, the feature level disentanglement is performed with an invertible neural layer for final quality prediction, leading to the PR and the corresponding distortion features for comparison. The effectiveness of our proposed method is demonstrated on four popular IQA databases, and superior performance on cross-database evaluation also reveals the high generalization capability of our method. The implementation of our method is publicly available on https://github.com/Baoliang93/FPR. Baoliang Chen, Lingyu Zhu 0006, Chenqi Kong, Hanwei Zhu, Shiqi Wang 0001, Zhu Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | PUGCQ: A Large Scale Dataset for Quality Assessment of Professional User-Generated ContentabstractRecent years have witnessed a surge of professional user-generated content (PUGC) based video services, coinciding with the accelerated proliferation of video acquisition devices such as mobile phones, wearable cameras, and unmanned aerial vehicles. Different from traditional UGC videos by impromptu shooting, PUGC videos produced by professional users tend to be carefully designed and edited, receiving high popularity with a relatively satisfactory playing count. In this paper, we systematically conduct the comprehensive study on the perceptual quality of PUGC videos and introduce a database consisting of 10,000 PUGC videos with subjective ratings. In particular, during the subjective testing, we collect the human opinions based upon not only the MOS, but also the attributes that could potentially influence the visual quality including face, noise, blur, brightness, and color. We make the attempt to analyze the large-scale PUGC database with a series of video quality assessment (VQA) algorithms and a dedicated baseline model based on pretrained deep neural network is further presented. The cross-dataset experiments reveal a large domain gap between the PUGC and the traditional user-generated videos, which are critical in learning based VQA. These results shed light on developing next-generation PUGC quality assessment algorithms with desired properties including promising generalization capability, high accuracy, and effectiveness in perceptual optimization. The dataset and the codes are released at https://github.com/wlkdb/pugcq_create. Baoliang Chen, Lingyu Zhu 0006, Qingwen He, Hongfei Fan, Shiqi Wang 0001 |
ACM Multimedia | 3 |