Chenxi Ma

dblp:246/5932 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Efficient ALF Algorithm for Eliminating Inter-Frame Parallelism Barrier in VVC Encoder
abstract
The Adaptive Loop Filter (ALF) in Versatile Video Coding (VVC) provides significant compression gains. However, its high computational complexity and the inter-frame parallelism barrier introduced by the frame-wise derivation of filter coefficients significantly reduce the overall encoding speed. To address these issues, this paper proposes an efficient ALF Parallel Acceleration scheme.
Mingrui Wang, Chenxi Ma, Xiangxu Wu
DCC2
2026 Prompt-based cross-domain graph distillation network for smart tourism recommendation
Dongmei Yuan, Chenxi Ma, Lina Ni, Chunbo Xu
Expert Syst. Appl.2
2025 Scaling Laws for Data-Efficient Visual Transfer Learning
Wenxuan Yang, Qingqv Wei, Chenxi Ma, Weimin Tan, Bo Yan 0001
ACM Multimedia3
2025 MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
abstract
Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis remains underdeveloped, primarily due to less specialized text descriptions in current dermatology multimodal datasets. To address this issue, we propose MM-Skin, the first large-scale multimodal dermatology dataset that encompasses 3 imaging modalities, including clinical, dermoscopic, and pathological and nearly 10k high-quality image-text pairs collected from professional textbooks. In addition, we generate over 27k diverse, instruction-following vision question answering (VQA) samples (9× the size of current largest dermatology VQA dataset). Leveraging public datasets and MM-Skin, we developed SkinVL, a dermatology-specific VLM designed for precise and nuanced skin disease interpretation. Comprehensive benchmark evaluations of SkinVL on VQA, supervised fine-tuning (SFT) and zero-shot classification tasks across 8 datasets, reveal its exceptional performance for skin diseases in comparison to both general and medical VLM models. The introduction of MM-Skin and SkinVL offers a meaningful contribution to advancing the development of clinical dermatology VLM assistants. Code and dataset are available at https://github.com/ZwQ803/MM-Skin.
Chenxi Ma, Weimin Tan, Bo Yan 0001
ACM Multimedia3
2025 TabiMed: Tabularizing Medical Images for Few-Shot In-Context Diagnosis
abstract
Achieving accurate predictions with limited samples is a key challenge in biomedical image artificial intelligence. Previous methods rely on pre-trained image foundation models with supervised fine-tuning (SFT) or zero-shot inference to enhance small-data performance. However, SFT is time-consuming and prone to overfitting, whereas zero-shot inference fails to fully exploit available data. Inspired by recent tabular foundation models, which show superior performance on small-sample tasks with in-context learning (ICL), we propose TabiMed, a novel framework that transforms visual representations into structured tabular data, leveraging pre-trained tabular models for fast and accurate analysis on small data. TabiMed consists of three key components: dynamic modality-aware representation engine, tabularization adapter and in-context inference module. Experiments on 10 datasets from different fields demonstrate three major advantages of TabiMed: 1) excellent performance on small datasets, with an average AUC of 14.1% higher than zero-shot; 2) high efficiency, with a training time 250x faster than SFT; 3) scalability to larger datasets through our tabularization adapter. TabiMed proposes a novel pathway to address the challenges of analyzing biomedical images with few samples.
Wanying Zhou, Yu Ling, Chenxi Ma, Weimin Tan, Bo Yan 0001
ACM Multimedia5
2025 Efficient yet secure: An archive knowledge graph-enhanced native sparse attention network for lightweight privacy-preserving recommendation
Chenxi Ma, Yaobin Wang, Limei Sun
Knowl. Based Syst.2
2025 Spatiotemporal-Aware Self-Supervised Fluorescence Microscopy Image Denoising
abstract
Fluorescence microscopy has been an indispensable tool in many scientific disciplines. However, the expensive imaging cost and the photo-toxicity problem make it difficult to obtain high-quality images. The independent shot noise in fluorescence microscopy images always overwhelms signals and limits the imaging resolution, hindering progress in related research. Recently, self-supervised image denoising has received wide attention for its ability to train a denoiser without paired low Signal-to-Noise Ratio (SNR) and high SNR images. Existing self-supervised fluorescence microscopy image denoising works either suffer from high imaging/computational cost or large training difficulty. Here, we propose a Single-image based Self-Supervised Denoising approach (TriS-D) by utilizing the spatiotemporal redundancy of the fluorescence microscopy imaging data, which facilitates the low-cost and convenient training. The TriS-D can generate the training data from a raw image, not only releasing the demand for multiple low SNR time-lapse imaging data but also enabling the building of a 2D convolution-based model. Comprehensive experiments across different imaging modalities and biological samples verify the effectiveness of the TriS-D.
Chenxi Ma, Weimin Tan, Zhaohui Zhou, Bo Yan 0001
IEEE Signal Process. Lett.1
2024 Uncertainty-Aware GAN for Single Image Super Resolution
abstract
Generative adversarial network (GAN) has become a popular tool in the perceptual-oriented single image super-resolution (SISR) for its excellent capability to hallucinate details. However, the performance of most GAN-based SISR methods is impeded due to the limited discriminative ability of their discriminators. In specific, these discriminators only focus on the global image reconstruction quality and ignore the more fine-grained reconstruction quality for constraining the generator, as they predict the overall realness of an image instead of the pixel-level realness. Here, we first introduce the uncertainty into the GAN and propose an Uncertainty-aware GAN (UGAN) to regularize SISR solutions, where the challenging pixels with large reconstruction uncertainty and importance (e.g., texture and edge) are prioritized for optimization. The uncertainty-aware adversarial training strategy enables the discriminator to capture the pixel-level SR uncertainty, which constrains the generator to focus on image areas with high reconstruction difficulty, meanwhile, it improves the interpretability of the SR. To balance weights of multiple training losses, we introduce an uncertainty-aware loss weighting strategy to adaptively learn the optimal loss weights. Extensive experiments demonstrate the effectiveness of our approach in extracting the SR uncertainty and the superiority of the UGAN over the state-of-the-arts in terms of the reconstruction accuracy and perceptual quality.
Chenxi Ma
AAAI1
2024 Learning Cross-Spectral Prior for Image Super-Resolution
abstract
With the rising interest in multi-camera cross-spectral systems, cross-spectral images have been widely used in computer vision and image processing. Therefore, an effective super-resolution (SR) method provides high-resolution (HR) cross-spectral images for different research and applications. However, existing SR methods rarely consider utilizing cross-spectral information to assist the SR of visible images. They cannot handle complex degradation (noise, high brightness, low light) and misalignment problems in low-resolution (LR) cross-spectral images. Here, we first explore the potential of using near-infrared (NIR) image guidance for better SR, based on the observation that NIR images can preserve valuable information for recovering adequate image details. To take full advantage of the cross-spectral prior, we propose a novel Cross-Spectral Prior guided image SR approach (CSPSR). The cross-view matching (CVM) module and the dynamic multi-modal fusion (DMF) module can enhance the spatial correlation between cross-spectral images and bridge the multi-modal feature gap, respectively. Extensive experiments demonstrate the effectiveness of our CSPSR.
Chenxi Ma, Weimin Tan, Shili Zhou, Bo Yan 0001
ACM Multimedia1
2023 Multi-Modality Deep Network for JPEG Artifacts Reduction
abstract
In recent years, many convolutional neural network-based models are designed for JPEG artifacts reduction, and have achieved notable progress. However, few methods are suitable for extreme low-bitrate image compression artifacts reduction. The main challenge is that the highly compressed image loses too much information, resulting in reconstructing high-quality image difficultly. To address this issue, we propose a multimodal fusion learning method for text-guided JPEG artifacts reduction, in which the corresponding text description not only provides the potential prior information of the highly compressed image, but also serves as supplementary information to assist in image deblocking. We fuse image features and text semantic features from the global and local perspectives respectively, and design a contrastive loss built upon contrastive learning to produce visually pleasing results. Extensive experiments, including a user study, prove that our method can obtain better deblocking results compared to the state-of-the-art methods.
Xuhao Jiang, Weimin Tan, Chenxi Ma, Bo Yan 0001, Liquan Shen
IJCAI4
2022 Geometry-Aware Reference Synthesis for Multi-View Image Super-Resolution
abstract
Recent multi-view multimedia applications struggle between high-resolution (HR) visual experience and storage or bandwidth constraints. Therefore, this paper proposes a Multi-View Image Super-Resolution (MVISR) task. It aims to increase the resolution of multi-view images captured from the same scene. One solution is to apply image or video super-resolution (SR) methods to reconstruct HR results from the low-resolution (LR) input view. However, these methods cannot handle large-angle transformations between views and leverage information in all multi-view images. To address these problems, we propose the MVSRnet, which uses geometry information to extract sharp details from all LR multi-view to support the SR of the LR input view. Specifically, the proposed Geometry-Aware Reference Synthesis module in MVSRnet uses geometry information and all multi-view LR images to synthesize pixel-aligned HR reference images. Then, the proposed Dynamic High-Frequency Search network fully exploits the high-frequency textural details in reference images for SR. Extensive experiments on several benchmarks show that our method significantly improves over the state-of-the-art approaches.
Ri Cheng, Bo Yan 0001, Weimin Tan, Chenxi Ma
ACM Multimedia5
2022 Rethinking Super-Resolution as Text-Guided Details Generation
abstract
Deep neural networks have greatly promoted the performance of single image super-resolution (SISR). Conventional methods still resort to restoring the single high-resolution (HR) solution only based on the input of image modality. However, the image-level information is insufficient to predict adequate details and photo-realistic visual quality facing large upscaling factors (×8, ×16). In this paper, we propose a new perspective that regards the SISR as a semantic image detail enhancement problem to generate semantically reasonable HR image that are faithful to the ground truth. To enhance the semantic accuracy and the visual quality of the reconstructed image, we explore the multi-modal fusion learning in SISR by proposing a Text-Guided Super-Resolution (TGSR) framework, which can effectively utilize the information from the text and image modalities. Different from existing methods, the proposed TGSR could generate HR image details that match the text descriptions through a coarse-to-fine process. Extensive experiments and ablation studies demonstrate the effect of the TGSR, which exploits the text reference to recover realistic images.
Chenxi Ma, Bo Yan 0001, Weimin Tan, Siming Chen 0001
ACM Multimedia1
2022 Prior embedding multi-degradations super resolution network
Chenxi Ma, Weimin Tan, Bo Yan 0001, Shili Zhou
Neurocomputing1
2021 Perception-Oriented Stereo Image Super-Resolution
abstract
Recent studies of deep learning based stereo image super-resolution (StereoSR) have promoted the development of StereoSR. However, existing StereoSR models mainly concentrate on improving quantitative evaluation metrics and neglect the visual quality of super-resolved stereo images. To improve the perceptual performance, this paper proposes the first perception-oriented stereo image super-resolution approach by exploiting the feedback, provided by the evaluation on the perceptual quality of StereoSR results. To provide accurate guidance for the StereoSR model, we develop the first special stereo image super-resolution quality assessment (StereoSRQA) model, and further construct a StereoSRQA database. Extensive experiments demonstrate that our StereoSR approach significantly improves the perceptual quality and enhances the reliability of stereo images for disparity estimation.
Chenxi Ma, Bo Yan 0001, Weimin Tan, Xuhao Jiang
ACM Multimedia1
2021 Motion Blur Removal With Quality Assessment Guidance
abstract
Non-uniform blind motion deblurring is a challenging yet fundamental task in the computer vision field, which aims to restore the latent sharp image from the blurry input. Recently, deep-learning-based methods have made significant improvement and progress, on the metric of PSNR. They achieve good results mainly because they adopt Mean Squared Error (MSE) as the optimization objective, in addition to their good model design. However, simple adoption of the PSNR metric and the MSE loss, has non-ignorable disadvantages. PSNR cannot always succeed in assessing the deblurred quality in accordance with the human visual system (HVS), and MSE guides the network to generate over-smoothed images. To address these problems, we are the first to propose the deep-learning-based multi-scale non-reference quality assessment network (Deep DEBLUR-IQA) for assessing the quality of deblurred results. Moreover, a deblurring network of high efficiency is presented. It is more than 50 times faster than other SOTA multi-scale Convolution Neural Network (CNN) methods, with the newly propose Residual Dilated Block (RDB) and Light ResBlock (LRB). The deblurring network's performance can be further boosted with Multiple Dilation Block (MDB), with an acceptable speed decrease. Finally, and most importantly, we are the first to let Deep DEBLUR-IQA guide the deblurring network's optimization. This IQA-guided enhancement paradigm can significantly improve the deblurring results’ subjective quality while achieving excellent PSNR. Experimental results demonstrate that the proposed method performs favorably against state-of-the-art methods quantitatively and qualitatively.
Bo Yan 0001, Chenxi Ma
IEEE Trans. Multim.5
2020 Disparity-Aware Domain Adaptation in Stereo Image Restoration
abstract
Under stereo settings, the problems of disparity estimation, stereo magnification and stereo-view synthesis have gathered wide attention. However, the limited image quality brings non-negligible difficulties in developing related applications and becomes the main bottleneck of stereo images. To the best of our knowledge, stereo image restoration is rarely studied. Towards this end, this paper analyses how to effectively explore disparity information, and proposes a unified stereo image restoration framework. The proposed framework explicitly learn the inherent pixel correspondence between stereo views and restores stereo image with the cross-view information at image and feature level. A Feature Modulation Dense Block (FMDB) is introduced to insert disparity prior throughout the whole network. The experiments in terms of efficiency, objective and perceptual quality, and the accuracy of depth estimation demonstrates the superiority of the proposed framework on various stereo image restoration tasks.
Bo Yan 0001, Chenxi Ma, Bahetiyaer Bare, Weimin Tan, Steven C. H. Hoi
CVPR2
2020 Deep Image Quality Assessment Driven Single Image Deblurring
abstract
Motion deblurring is a challenging task in computer vision, which aims to recover the sharp image from the blurry one. Recently, deep learning based methods have made a significant improvement in the metric of PSNR due to the optimazation of the Mean Squared Error (MSE) loss function between deblurring results and sharp images. However, PSNR prefers smooth images and fails to evaluate the sharpness of deblurred images. To solve this problem, we firstly propose the deep learning based multi-scale non-reference quality assessment network (Deep DEBLUR-IQA) for assessing deblurred results. Moreover, we propose an efficient deblurring network, which is over 50 times faster than SOTA multi-scale networks. Finally, the combination of our Deep DEBLUR-IQA network and novel single image deblurring network can significantly increase subjective quality while maintaining satisfactory PSNR. Experiments show that the proposed method outperforms state-of-the-art methods, both qualitatively and quantitatively.
Chenxi Ma, Bo Yan 0001
ICME4
2019 A Multi-level Aggregated Network for Image Restoration
abstract
Recently, significant progress has been witnessed in image restoration benefited from the development of deep convolutional neural networks (CNN). However, we note that many state-of-the-art image restoration networks can be unfolded as a one-level architecture, which is constructed by stacking multiple convolution layers or blocks. As the depth of network grows, the information flow is weakened. And, existing skip connection has restricted ability to pass the previous state to the latter layers in the network. Based on above observations, we explore a multi-level aggregated network (MLAN) to fully exploit features of deeper layers. The proposed MLAN can extract more features by merging layers at different levels progressively, and can better aggregate the features by the augmented connection manner. Experimental results demonstrate a satisfactory performance of the proposed model on different image restoration tasks.
Chenxi Ma, Weimin Tan, Bahetiyaer Bare, Bo Yan 0001
ICME1
2019 Real-time video super-resolution via motion convolution kernel estimation
Bahetiyaer Bare, Bo Yan 0001, Chenxi Ma, Ke Li 0010
Neurocomputing3
2019 Deep Objective Quality Assessment Driven Single Image Super-Resolution
abstract
Single-image super-resolution (SISR) is a classic problem in the image processing community, which aims at generating a high-resolution image from a low-resolution one. In recent years, deep learning based SISR methods emerged and achieved a performance leap than previous methods. However, because the evaluation metrics of SISR methods is peak signal-to-noise ratio (PSNR), previous methods usually choose L2-norm as the loss function. This leads to a significant improvement in the final PSNR value but little improvement in perceptual quality. In this paper, in order to achieve better results in both perceptual quality and PSNR values, we propose an objective quality assessment driven SISR method. First, we propose a novel full-reference image quality assessment approach for SISR and employ it as a loss function, namely super-resolution image quality assessment (SR-IQA) loss. Then, we combine SR-IQA loss with L2-norm to guide our proposed SISR method to achieve better results. Besides that, our proposed SISR method consists of several proposed highway units. Furthermore, in order to verify the generalization ability of our new kind of loss function, we integrate SR-IQA loss to generative adversarial networks based SR method and achieve better perceptual quality. Experimental results prove that our proposed SISR method achieves better performance than other methods both qualitatively and quantitatively in most of the cases.
Bo Yan 0001, Bahetiyaer Bare, Chenxi Ma, Ke Li 0010, Weimin Tan
IEEE Trans. Multim.3