Yudong Mao

dblp:320/3547 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-4718-8592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Codebook-Empowered Analysis-Friendly Extreme Underwater Image Compression
abstract
While existing underwater image compression (UIC) methods optimize for human perception or basic redundancies, they neglect inter-image correlations and fail to prioritize machine-friendly features essential for automated analysis. This paper introduces a novel -quantized (VQ) codebook-driven framework for machine-centric UIC. We leverage VQ codebooks -- pre-trained as external priors on diverse underwater data -- to unify three critical stages: (1) Machine-friendly feature extraction via contrastive learning with high/low-quality codebooks, enhancing degradation robustness; (2) Compact compression using variable-size codebooks to map discriminative features to entropy-coded indices, enabling ultra-low bitrates (less than 0.04bpp); and (3) Feature refinement at the decoder, restoring semantic fidelity for downstream tasks. In addition, we contribute the first Underwater Visual Question Answering (UVQA) benchmark to holistically evaluate machine perception across object presence, counting, and localization. Extensive experiments demonstrate that our framework significantly outperforms state-of-the-art codecs in machine vision task performance at ultra-low bitrates. The VQ-codebook effectively harnesses inter-image redundancy, combats joint degradation, and delivers compact, analysis-friendly representations, establishing a new paradigm for machine-centric UIC.
Jianhao Wu, Yudong Mao, Qiuping Jiang
AAAI2
2026 Underwater image compression for human and machine visions with hybrid priors embedding
Jianhao Wu, Zhanhong Lu, Yudong Mao, Qiuping Jiang
Pattern Recognit.3
2026 Optimizing Fidelity-Perception Tradeoff via Large Vision-Language Model Prior for Image Compression
abstract
Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity.
Yudong Mao, Peilin Chen 0001, Lingyu Zhu 0006, Yung-Hui Li, Shiqi Wang 0001
IEEE Trans. Image Process.1
2025 Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration
Yudong Mao, Zhiwei Zhong 0001, Peilin Chen 0001, Zhijiang Zhang, Shiqi Wang 0001
CVPR1
2025 Exploiting Long and Short Temporal Dependence for Low-Light Video Enhancement
abstract
Existing learning-based methods often lack temporal coherence in low-light video enhancement due to rarely considering intrinsic temporal dependence. To address this issue, we propose the Long-short Temporal Filtering Network (TFNet) to learn the mapping from low-light videos to normal-light ones, utilizing the well-considered data-centric strategy and a refined architecture. From the data-centric temporal strategy, we incorporate both long-range and short-range temporal dependence into TFNet, effectively capturing the temporal information. From the model design perspective, the TFNet incorporates the Temporal-aware Attentional Filtering (TAF) module, which aims to estimate and adaptively combine filtering kernels for guided filtering towards features of the middle frame. To further refine the filtered features, the cascaded Grouped Attention (GA) blocks are presented in a grouped attention strategy. Experimental results on benchmark datasets have demonstrated the superiority of our TFNet against the state-of-the-art methods in terms of video frame quality and brightness consistency.
Lingyu Zhu 0006, Yudong Mao, Zhiwei Zhong 0001, Shanshe Wang, Shiqi Wang 0001
ICME3
2025 Deep Reinforcement Learning-Based Semi-Autonomous Control for Magnetic Micro-Robot Navigation with Immersive Manipulation
abstract
Magnetic micro-robots have demonstrated immense potential in biomedical applications, such as in vivo drug delivery, non-invasive diagnostics, and cell-based therapies, owing to their precise maneuverability and small size. However, current micromanipulation techniques often rely solely on a two-dimensional (2D) microscopic view as sensory feedback, while traditional control interfaces do not provide an intuitive manner for operators to manipulate micro-robots. These limitations increase the cognitive load on operators, who must interpret limited feedback and translate it into effective control actions. To address these challenges, we propose a Deep Re-inforcement Learning-Based Semi-Autonomous Control (DRL-SC) framework for magnetic micro-robot navigation in a simulated microvascular system. Our framework integrates Mixed Reality (MR) to facilitate immersive manipulation of micro-robots, thereby enhancing situational awareness and control precision. Simulation and experimental results demonstrate that our approach significantly improves navigation efficiency, reduces control errors, and enhances the overall robustness of the system in simulated microvascular environments.
Yudong Mao
ICRA1
2025 MS-MoE: Multi-modal Structural Mixture of Experts Framework for Pan-Sharpening
abstract
Pan-sharpening aims to generate the high-resolution (HR) multi-spectral (MS) target image from its low-resolution (LR) counterpart, which is guided by corresponding HR panchromatic (PAN) image with abundant texture structural details. Although the existing state-of-the-art methods have made remarkable progress, they are still struggling with integrating inherent structural correlation between PAN and MS images through the early or late-stage fusion alone. This would lead to texture-less pan-sharpening reconstruction due to the insufficient learning of complementary features from PAN image. To address this issue, we propose the Multi-modal Structural Mixture of Experts (MS-MoE) framework for pan-sharpening. Specifically, given the upsampled LRMS and PAN images spatially rotated at various angles, we design a set of structural experts to extract the complementary spatial and spectral features between them, in which the Texture Enhancement Module (TEM) is introduced to extract and enhance texture-structural features from different modalities. Subsequently, we introduce an additional expert network to perform feature fusion by integrating the outputs from multiple experts. To reconstruct the high-frequency information, we further leverage the Frequency feature Refinement Module (FRM) to aggregate and refine the fused features in the frequency domain. Experimental results on the benchmark pan-sharpening datasets demonstrate that the proposed MS-MoE framework achieves more competitive performance than recent state-of-the-art methods.
Zhiwei Zhong 0001, Lingyu Zhu 0006, Yudong Mao, Shiqi Wang 0001
IJCNN4
2025 Multiple Glancing at Quality: Benchmark Dataset and Objective Quality Assessment Metric for Low-light Image Enhancement
abstract
Quality metrics play a crucial role in guiding the development of image enhancement algorithms, which have consistently sought effective quality assessment methodologies and comprehensive datasets. To address this need, we first built a large-scale dataset for low-light enhanced image quality assessment and gathered the corresponding subjective evaluation scores. Recently, vision-language pre-training models have demonstrated considerable potential in the realm of quality assessment. However, its efficacy is limited by the fine-grained perception in low-level quality assessment. As such, we further propose a novel quality assessment framework using contrastive prompt learning, which harnesses the robust priors of vision-language pre-training models to improve the perceptual capacity of deep networks for low-level quality features. Experiments on the proposed RSLE dataset show that our method outperforms existing SOTA image quality assessment methods. Our database and the source code will be made publicly available.
Yudong Mao, Peilin Chen 0001, Zhao Wang 0004, Qiuping Jiang, Shiqi Wang 0001
ISCAS1
2023 Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine Perception
abstract
We propose a novel face compression framework that leverages the external priors for joint human and machine perception under ultra-low bitrate scenarios. The proposed framework leverages the semantic richness of face images by representing the faces into sketches and thumbnails, resulting in improved bitrate utility for both human and machine vision. At the decoder side, the framework introduces a two-stage generative reconstruction, which faithfully enhances the reconstructed image via semi-parametric modeling and retrieved guidance from the external database. In particular, this coarse-to-fine strategy also results in improved identity consistency and analysis performance of the reconstructed image. Extensive evaluations of the proposed method have been conducted on the public face dataset by comparing it with end-to-end image compression techniques as well as traditional image compression standards. The experimental results demonstrate the effectiveness of the proposed method via superior perceptual and analytical performance under ultra-low bitrate conditions.
Yudong Mao, Peilin Chen 0001, Shurun Wang, Shiqi Wang 0001, Dapeng Oliver Wu
ACM Multimedia1
2022 Unsupervised Decomposition and Correction Network for Low-Light Image Enhancement
abstract
Vision-based intelligent driving assistance systems and transportation systems can be improved by enhancing the visibility of the scenes captured in extremely challenging conditions. In particular, many low-image image enhancement (LIE) algorithms have been proposed to facilitate such applications in low-light conditions. While deep learning-based methods have achieved substantial success in this field, most of them require paired training data, which is difficult to be collected. This paper advocates a novel Unsupervised Decomposition and Correction Network (UDCN) for LIE without depending on paired data for training. Inspired by the Retinex model, our method first decomposes images into illumination and reflectance components with an image decomposition network (IDN). Then, the decomposed illumination is processed by an illumination correction network (ICN) and fused with the reflectance to generate a primary enhanced result. In contrast with fully supervised learning approaches, UDCN is an unsupervised one which is trained only with low-light images and corresponding histogram equalized (HE) counterparts (can be derived from the low-light image itself) as input. Both the decomposition and correction networks are optimized under the guidance of hybrid no-reference quality-aware losses and inter-consistency constraints between the low-light image and its HE counterpart. In addition, we also utilize an unsupervised noise removal network (NRN) to remove the noise previously hidden in the darkness for further improving the primary result. Qualitative and quantitative comparison results are reported to demonstrate the efficacy of UDCN and its superiority over several representative alternatives in the literature. The results and code will be made public available athttps://github.com/myd945/UDCN.
Qiuping Jiang, Yudong Mao, Runmin Cong, Wenqi Ren, Chao Huang 0008, Feng Shao 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Cross-Modality Fusion and Progressive Integration Network for Saliency Prediction on Stereoscopic 3D Images
abstract
Traditional 2D image-based saliency prediction models suffer from unsatisfactory performance when dealing with stereoscopic 3D (S3D) images because eye movements in the case of freely viewing S3D images are demonstrated to be guided by both RGB and depth features. This paper studies the problem of saliency prediction on S3D images, where the interactions between RGB and depth modalities are both taken into account. Specifically, we design a novel deep neural network named Cross-modality Fusion and Progressive Integration Network (CFPI-Net) to address this problem. It consists of a Multi-level Cross-modality Feature Fusion (MCFF) module and a Multi-stage Progressive Feature Integration (MPFI) module. The MCFF module first captures hierarchical contexture features from each modality and then effectively fuses the hierarchical contexture features from different modalities at each level. The MPFI module involves multiple cascaded deeply supervised feature integration (DSFI) blocks in which the low-level and high-level cross-modality features are progressively integrated using the integrated features in the previous stage as a guidance. Our proposed CFPI-Net benefits from the advantages of multi-level feature representation, cross-modality feature fusion, and multi-stage progressive feature integration, which hereby fully boost the performance. Experimental results on two benchmark datasets demonstrate that CFPI-Net outperforms state-of-the-art saliency prediction methods both quantitatively and qualitatively. All the results and relevant codes will be made available to the public.
Yudong Mao, Qiuping Jiang, Runmin Cong, Wei Gao 0003, Feng Shao 0001, Sam Kwong
IEEE Trans. Multim.1