EDBT 2026 Demo / reviewers in the wild / expert
Jingchao Cao
dblp:250/6882
· DBLP profile ↗
16ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-8944-8033ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UQ-Bench: A Benchmark for Evaluating Multimodal LLMs on Underwater Image Quality AssessmentabstractDespite the rapid progress of multimodal large language models (MLLMs), their capacity for low-level visual perception in underwater environments remains underexplored. To address this gap, we present UQ-Bench, the first systematically designed benchmark for evaluating the ability of MLLMs to perceive and assess underwater image quality at the low-level visual attribute level. UQ-Bench comprises three components: (1) UW-Perception, a dataset of 3,000 underwater images paired with targeted questions on key degradations such as color cast, blur, contrast, and exposure, covering both global and local perceptual dimensions; (2) UW-Describe, a dataset of 500 images with expert-annotated gold-standard descriptions for assessing the accuracy of model-generated text; and (3) UW-Eval, an evaluation protocol employing human mean opinion scores (MOS) for quantitative quality assessment. To ensure rigorous and reproducible benchmarking, we propose a GPT-assisted evaluation framework that aligns model outputs with expert references and enables fine-grained analysis of distortion perception. Experimental results demonstrate that while MLLMs exhibit preliminary competence in underwater low-level visual tasks, they still fall short in capturing subtle degradations and achieving human-level consistency, highlighting the need for further advances in foundation models for marine vision. Jingchao Cao, Guo An, Feng Gao 0005, Ke Gu 0001, Yutao Liu 0002 |
AAAI | 1 |
| 2026 | SEA-PACE: Semi-Supervised Underwater Image Enhancement via Gaussian Process-Assisted Self-Paced LearningabstractThe scarcity of paired data severely limits the performance and generalization of learning-based underwater image enhancement (UIE) methods. This challenge is particularly prominent in scenes with complex degradations. Semi-supervised learning has emerged as a promising solution by enabling the utilization of large-scale unlabeled data. However, its effectiveness is limited by the use of static, model-agnostic metrics for pseudo-label reliability assessment. To address this, we propose SEA-PACE, a novel semi-supervised framework that integrates model-aware uncertainty modeling and self-paced consistency learning to fully exploit unlabeled data for UIE. Specifically, we design a Model-Aware Reliability Estimator (MARE) that quantifies the uncertainty of the teacher model's predictions through Gaussian Process Regression in latent feature space. The resulting uncertainty is then transformed into reliability weights via a rank-based mapping. Additionally, we apply the Self-Paced Consistency Learning (SPCL) strategy that employs a loss-aware schedule to dynamically prioritize high-confidence pseudo-labels, gradually incorporating more challenging samples during training. Extensive experiments on several public UIE benchmarks demonstrate that SEA-PACE consistently surpasses state-of-the-art methods in both visual quality and generalization capability. Hengyue Bi, Jingchao Cao, Feng Gao 0005, Junyu Dong |
AAAI | 3 |
| 2026 | Deep Feature Prior-Guided Conditional Diffusion Model for Underwater Image EnhancementabstractUnderwater imaging always suffers from color distortion and reduced visibility due to light absorption and scattering, severely hindering visual perception and analysis. In this letter, we propose an underwater image enhancement framework based on diffusion model augmented with two lightweight guidance modules. The first module is a conditional branch that extracts structural features from a coarsely enhanced version to guide the denoising process toward more faithful restoration. While the second module retrieves high-quality features from a pre-constructed feature dictionary as priors, effectively restoring colors and fine details in degraded regions. Extensive experiments on public underwater image datasets demonstrate that our proposed method outperforms the state-of-the-art approaches both quantitatively and visually. It also generalizes well across various underwater environments, highlighting the effectiveness of incorporating structural and feature-level guidance into the diffusion process. The source code and pre-trained model are available at https://github.com/Juneit/PGUIE. Linwei Zhu, Tao Tian, Wenhui Wu 0001, Jingchao Cao |
IEEE Signal Process. Lett. | 5 |
| 2026 | A Lightweight Deep and Wide Network for Image-Based Detection of Industrial Waste GasabstractDue to inadequate monitoring, key pollutants (e.g., PM2.5, VOCs, etc) very possibly leak into atmosphere, thus to endanger the long-term and short-term life safety of people that work and live in the environment. Therefore, it is imperative to effectively and efficiently detect the leakage of industrial waste gas, for the purpose of timely lowering the risk of pollution and explosions. To solve such a problem, we in this paper propose a new lightweight deep and wide network (LdwNet) for detecting the leakage of industrial waste gas from an image, which brings about the two main merits: 1) Compensating for the deficiencies of sensor-based detection methods, which can accurately detect the leakage of waste gas and even measure its concentrations but require to seek leakage sources beforehand; 2) Overcoming the shortcomings of image-based detection methods, which leverage DNN-based recognition technologies and usually suffer from low efficacy, low efficiency and high energy consumption during the model training and inference. To specify, the proposed LdwNet is developed by simulating human perception, motivated by the method which detects the leakage of industrial waste gas from surveillance images with the human observation and judgement. First, based on the inspiration that the human eyes are highly sensitive to horizontal and vertical stimuli, we construct a novel lightweight parallel-series-stripe (PS2) module to validly extract features with very few parameters. Second, to fully exploit deep and shallow features for fusing the global and local information, we extend the PS2 module as a backbone along both the deep and wide directions to build the multi-channel network. Third, to achieve effective, efficient and low-carbon detection in model running, we constraint the extended PS2 modules with parameter sharing to prodigiously reduce the model parameters and thus to make the proposed model ultra-lightweight. Experiments on the datasets of carbon particulate matters and ethylene leakage prove that our LdwNet with ten thousand parameters outperforms the state-of-the-art models with millions of parameters in detection accuracy and implementation cost, and this renders our proposed LdwNet more suitable for real industrial applications. Ke Gu 0001, Hongyan Liu 0004, Jingchao Cao, Lai-Kuan Wong, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001, Weisi Lin, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | SeaDiff: Underwater Image Enhancement With Degradation-Aware Diffusion ModelabstractLight propagation in underwater scenes is significantly hindered by wavelength- and distance-dependent attenuation and scattering, leading to low contrast and severe color distortion in underwater images. Recent advancements in diffusion models have shown impressive performance in image restoration by learning data distribution prior knowledge (diffusion prior) from large amounts of paired data. However, due to the difficulties in collecting paired underwater images, the available data for underwater image enhancement is limited in both quality and quantity. This scarcity leads to a biased diffusion prior and suboptimal performance of diffusion models. To address this issue, we propose a novel method, termed SeaDiff, to learn underwater diffusion prior with wavelength- and distance-dependent degradation awareness. Specifically, we introduce a Prior Knowledge Mining Model (PKMM), which includes two key components: (1) the Physical Prior Embedding Module (PPEM) that simulates the underwater imaging process through a distance-dependent physical model and embeds physical prior by incorporating generalizable distance-aware cues from a large vision foundation model; and (2) the Color Prior Embedding Module (CPEM) that extracts wavelength-dependent color distribution prior from a log-chroma color space. Additionally, we propose a Degradation-Aware Diffusion Model (DADM) that seamlessly integrates degradation prior with diffusion prior and enhances the underwater images with high visual quality. Extensive experiments on popular UIE benchmarks and downstream tasks demonstrate that the proposed SeaDiff achieves state-of-the-art performance in terms of both visual quality and quantitative metrics. The code will be released at https://github.com/Henry-Bi/SeaDiff. Hengyue Bi, Long Chen 0019, Jingchao Cao, Jinghao Sun, Yuan Rao 0001, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ERD: Encoder-Residual-Decoder Neural Network for Underwater Image EnhancementabstractIn underwater environments, the absorption and scattering of light often result in various types of degradation in captured images, including color cast, low contrast, low brightness, and blurriness. These undesirable effects pose significant challenges for both underwater photography and downstream tasks such as object detection, recognition, and navigation. To address these challenges, we propose a novel end-to-end underwater image enhancement (UIE) network via the multistage and mixed attention mechanism and a residual-based feature refinement module, called ERD. Specifically, our network includes an encoder stage for extracting features from input underwater images with channel, spatial, and patch attention modules to emphasize degraded channels and regions for restoration; a residual stage for further purification of informative features through sufficient feature learning; and a decoder stage for effective image reconstruction. Inspired by visual perception mechanism, we design the frequency domain loss and edge details loss to retain more high-frequency information and object details while ensuring that the enhanced image approximates the reference image in terms of color tone while preserving content and structure. To comprehensively evaluate our proposed UIE model, we also curated three additional underwater image datasets through online collection and generation using Cycle-GAN. Rigorous experiments conducted on a total of eight underwater image datasets demonstrate that the proposed ERD model outperforms state-of-the-art methods in enhancing both real-world and generated underwater images. Our code and datasets are available athttps://github.com/fansuregrin/ERD. Jingchao Cao, Wangzhen Peng, Yutao Liu 0002, Junyu Dong, Patrick Le Callet, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Multi-Scale Local and Global Feature Fusion for Blind Quality Assessment of Enhanced ImagesabstractImage enhancement plays a crucial role in computer vision by improving visual quality while minimizing distortion. Traditional methods enhance images through pixel value transformations, yet they often introduce new distortions. Recent advancements in deep learning-based techniques promise better results but challenge the preservation of image fidelity. Therefore, it is essential to evaluate the visual quality of enhanced images. However, existing quality assessment methods frequently encounter difficulties due to the unique distortions introduced by these enhancements, thereby restricting their effectiveness. To address these challenges, this paper proposes a novel blind image quality assessment (BIQA) method for enhanced natural images, termed multi-scale local feature fusion and global feature representation-based quality assessment (MLGQA). This model integrates three key components: a multi-scale Feature Attention Mechanism (FAM) for local feature extraction, a Local Feature Fusion (LFF) module for cross-scale feature synthesis, and a Global Feature Representation (GFR) module using Vision Transformers to capture global perceptual attributes. This synergistic framework effectively captures both fine-grained local distortions and broader global features that collectively define the visual quality of enhanced images. Furthermore, in the absence of a dedicated benchmark for enhanced natural images, we design the Natural Image Enhancement Database (NIED), a large-scale dataset consisting of 8,581 original images and 102,972 enhanced natural images generated through a wide array of traditional and deep learning-based enhancement techniques. Extensive experiments on NIED demonstrate that the proposed MLGQA model significantly outperforms current state-of-the-art BIQA methods in terms of both prediction accuracy and robustness. Jingchao Cao, Yutao Liu 0002, Feng Gao 0005, Ke Gu 0001, Guangtao Zhai, Junyu Dong, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Real-World Multi-View Stereo via Learning RGB-D Structural Consistency From Depth Super-ResolutionabstractLearning-based Multi-View Stereo (MVS) methods, typically reliant on cascaded cost volume formulations, perform well on small-scale scenes. However, as the depth range of captured images becomes broader and more varied, the coarse-to-fine depth sampling process, which depends solely on feature matching, is increasingly prone to local optima. Despite recent advancements in feature representation, depth sampling patterns, and cost aggregation techniques, challenges related to model generalization and computational efficiency persist. In this paper, we propose SR-MVSNet, a novel framework that integrates multi-view feature matching and RGB-D cross-modal structural consistency learning to achieve high-quality 3D reconstruction. Our approach begins with the construction of Low-Resolution (LR) cost volumes for initial LR depth estimation, which are then enhanced to full-resolution via a tailored uncertainty-aware guided depth super-resolution module. To ensure cross-view consistency, the depth maps undergo further refinement through multi-view feature matching. By avoiding high-resolution cost volume processing, our framework improves depth estimation robustness and efficiency. Additionally, we introduce an iterative depth fusion post-processing strategy during inference to improve reconstruction in ambiguous matching regions, a critical challenge for MVS methods. Experiments show that our method achieves top-3 performance on the DTU and Tanks & Temples datasets and ranks first on the ETH3D dataset. Furthermore, it uses significantly fewer GPU resources than most high performing methods, offering a favorable trade-off between reconstruction quality and computational efficiency. Yimei Liu, Jingchao Cao, Hao Fan 0004, Junyu Dong, Sheng Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Wavelet-Assisted Mamba for Satellite-Derived Sea Surface Temperature Super-ResolutionabstractSea surface temperature (SST) is an essential indicator of global climate change and one of the most intuitive factors reflecting ocean conditions. Obtaining high-resolution SST data remains challenging due to limitations in physical imaging, and super-resolution via deep neural networks is a promising solution. Recently, Mamba-based approaches leveraging State Space Models (SSM) have demonstrated significant potential for long-range dependency modeling with linear complexity. However, their application to SST data super-resolution remains largely unexplored. To this end, we propose the Wavelet-assisted Mamba Super-Resolution (WMSR) framework for satellite-derived SST data. The WMSR includes two key components: the Low-Frequency State Space Module (LFSSM) and High-Frequency Enhancement Module (HFEM). The LFSSM uses 2D-SSM to capture global information of the input data, and the robust global modeling capabilities of SSM are exploited to preserve the critical temperature information in the low-frequency component. The HFEM employs the pixel difference convolution to match and correct the high-frequency feature, achieving accurate and clear textures. Through comprehensive experiments on three SST datasets, our WMSR demonstrated superior performance over state-of-the-art methods. Our codes and datasets will be made publicly available at https://github.com/oucailab/WMSR. Wankun Chen, Feng Gao 0005, Yanhai Gan, Jingchao Cao, Junyu Dong, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Adaptive Frequency Enhancement Network for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation of high-resolution remote sensing images plays a crucial role in land-use monitoring and urban planning. Recent remarkable progress in deep learning-based methods makes it possible to generate satisfactory segmentation results. However, existing methods still face challenges in adapting network parameters to various land cover distributions and enhancing the interaction between spatial and frequency domain features. To address these challenges, we propose the Adaptive Frequency Enhancement Network (AFENet), which integrates two key components: the Adaptive Frequency and Spatial feature Interaction Module (AFSIM) and the Selective feature Fusion Module (SFM). AFSIM dynamically separates and modulates high- and low-frequency features according to the content of the input image. It adaptively generates two masks to separate high- and low-frequency components, therefore providing optimal details and contextual supplementary information for ground object feature representation. SFM selectively fuses global context and local detailed features to enhance the network’s representation capability. Hence, the interactions between frequency and spatial features are further enhanced. Extensive experiments on three publicly available datasets demonstrate that the proposed AFENet outperforms state-of-the-art methods. In addition, we also validate the effectiveness of AFSIM and SFM in managing diverse land cover types and complex scenarios. Our codes are available at https://github.com/oucailab/AFENet. Feng Gao 0005, Miao Fu, Jingchao Cao, Junyu Dong, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Attention and Mamba-Driven Quality Assessment for Underwater ImagesabstractUnderwater imaging is essential in a variety of fields, including resource exploration, marine observation, and scientific research. However, the quality of underwater images is often compromised by environmental factors such as light scattering, absorption, and the presence of fog, leading to distortions such as color shifts, low contrast, and blurriness. To address these challenges, we propose a novel underwater image quality assessment (UIQA) method, the Attention and Mamba-driven Quality Index (AMQI). The AMQI model employs a multi-stage architecture designed to capture both local and global image features critical for underwater quality evaluation. First, a Shallow Feature Extractor (SFE) captures essential spatial details. Next, the Local Information Representation Network (LIR-Net), equipped with Channel Attention (CA) and Large Kernel-guided Spatial (LKS) mechanisms, enhances fine details and captures long-range dependencies to address underwater-specific distortions. The Global Information Representation Network (GIR-Net) further processes the features using a combination of the Visual State-Space Model (VSSM) and ResNet-50 to capture high-level semantic and contextual information. Finally, the Feature-Quality Mapping Network (FQM) converts the learned features into a quality score, ensuring precise predictions of image quality. Extensive experiments on the Underwater Image Quality Database (UIQD) demonstrate that AMQI outperforms current state-of-the-art IQA and UIQA models in terms of accuracy and correlation with human subjective evaluations. The model's robustness and generalization capabilities are further validated through detailed ablation studies and cross-database evaluations, showcasing its strong performance across diverse underwater environments. The source code is available athttps://github.com/ibaochao/AMQI. Jingchao Cao, Baochao Zhang, Yutao Liu 0002, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 1 |
| 2024 | UIQI: A Comprehensive Quality Evaluation Index for Underwater ImagesabstractDue to the light absorption and scattering in waterbodies, acquired underwater images frequently suffer from color cast, blur, low contrast, noise, etc., which seriously degrade the image quality and affect their subsequent applications. Therefore, it is necessary to propose a reliable and practical underwater image quality assessment (IQA) model that can faithfully evaluate underwater image quality. To this end, in this article, we establish a novel quality assessment model for underwater images by in-depth analysis and characterization of multiple image properties. Specifically, we propose characterizing the image luminance, color cast, sharpness, contrast, fog density and noise to comprehensively describe the image quality to evaluate the underwater image quality more accurately. Dedicated features are elaborately investigated to characterize those quality-aware image properties. After feature extraction, we employ support vector regression (SVR) to integrate all the quality-aware features and regress them onto the underwater image quality score. Extensive tests performed on standard underwater image quality databases demonstrate the superior prediction performance of the proposed underwater IQA model to state-of-the-art congeneric quality assessment models. Yutao Liu 0002, Ke Gu 0001, Jingchao Cao, Shiqi Wang 0001, Guangtao Zhai, Junyu Dong, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2021 | HOCA: Higher-Order Channel Attention for Single Image Super-ResolutionabstractConvolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich feature statistics higher than second-order, thus hindering the representation ability of CNNs. To address this issue, we propose a higher-order channel attention (HOCA) module to enhance the representation ability of CNNs. In our HOCA module, to capture different types of semantic information, we first compute k-order of feature statistics, followed by channel attention to learn the feature interdependencies. Considering the diversity of input contents, we design a gate mechanism to adaptively select a specific k-order channel attention. Besides, our HOCA module serves as a plug-and-play module and can be easily plugged into existing state-of-art CNN-based SR methods. Extensive experiments on public benchmarks show that our HOCA module effectively improves the performance of various CNN-based SR methods. Yalei Lv, Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia, Jingchao Cao |
ICASSP | 6 |
| 2021 | No-reference image quality assessment for contrast-changed images via a semi-supervised robust PCA model
Jingchao Cao, Ran Wang 0001, Yuheng Jia, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong |
Inf. Sci. | 1 |
| 2021 | Reinforcement learning-based QoE-oriented dynamic adaptive streaming framework
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Shiqi Wang 0001, Guopu Zhu, Jingchao Cao |
Inf. Sci. | 7 |
| 2019 | Content-oriented image quality assessment with multi-label SVM classifier
Jingchao Cao, Shiqi Wang 0001, Ran Wang 0001, Xinfeng Zhang 0001, Sam Kwong |
Signal Process. Image Commun. | 1 |