Zhuoqi Ma

dblp:211/5867 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-0729-9706ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 RSA-CR: Resisting Shilling Attacks in Citation Recommendation via Dumbbell Inductive Learning
abstract
Citation recommendation aims to provide researchers with the most relevant references for their manuscripts, helping them swiftly discover pertinent studies and bolster the reliability of their arguments. However, some individuals manipulate these recommendation systems by injecting false information, such as deliberately inflating the citation count of their own papers, to obtain favorable recommendations and ratings. This form of attack, commonly termed “shilling attack”, is not only highly concealed but also has an unimaginable impact on all scientific research. To address this problem, we theoretically reveal the impact of shilling attacks on citation recommendation and propose three feasible resistance strategies: historical collaborations, significant citations and content constraints. Based on these insights, we introduce RSA-CR, a robust and hybrid citation recommendation algorithm resistant to shilling attacks. The algorithm constructs a two-layer academic graph and uses random and content generation strategies to initialize author and paper embeddings. Confidence-guided inductive aggregations based on collaboration and citation relationships are then performed at the author and paper sides, where author aggregation results directly influences the paper aggregation strength. Finally, recommendations are made by measuring the distances between the fused paper embeddings. The entire learning process resembles a dumbbell, hence termed “dumbbell inductive learning”. Experiments on four academic datasets demonstrate that our method outperforms baselines in both effectiveness and robustness.
Xiyue Gao, Zhuoqi Ma, Xiaotian Qiao, Hui Li 0005, Kunhua Zhang, Jiangtao Cui
AAAI3
2026 PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
abstract
Chest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge---including clinical context (e.g., symptoms, medical history) and the most recent prior image---which radiologists routinely rely on for diagnostic reasoning. Most existing methods generate reports from single images, neglecting this essential prior information and thus failing to capture diagnostic intent or disease progression. To bridge this gap, we propose PriorRG, a novel chest X-ray report generation framework that emulates real-world clinical workflows via a two-stage training pipeline. In Stage 1, we introduce a prior-guided contrastive pre-training scheme that leverages clinical context to guide spatiotemporal feature extraction, allowing the model to align more closely with the intrinsic spatiotemporal semantics in radiology reports. In Stage 2, we present a prior-aware coarse-to-fine decoding for report generation that progressively integrates patient-specific prior knowledge with the vision encoder's hidden states. This decoding allows the model to align with diagnostic focus and track disease progression, thereby enhancing the clinical accuracy and fluency of the generated reports. Extensive experiments on MIMIC-CXR and MIMIC-ABN datasets demonstrate that PriorRG outperforms state-of-the-art methods, achieving a 3.6% BLEU-4 and 3.8% F1 score improvement on MIMIC-CXR, and a 5.9% BLEU-1 gain on MIMIC-ABN.
Kang Liu 0025, Zhuoqi Ma, Zikang Fang, Yunan Li 0001, Kun Xie 0011, Qiguang Miao
AAAI2
2026 Patient-specific multimodal learning with multi-view contrastive alignment for chest X-ray report generation
abstract
MOTIVATION: Radiology reports play a pivotal role in guiding treatment planning and enabling effective doctor-patient communication. However, their manual composition imposes a substantial workload on radiologists. Although automatic radiology report generation has emerged as a promising alternative, existing approaches predominantly rely on single-view chest X-rays and fail to adequately leverage patient-specific context, thereby limiting diagnostic accuracy. RESULTS: To address this challenge, we propose EVOKE, a novel chest X-ray report generation framework that incorporates multi-view contrastive learning and patient-specific knowledge. Specifically, we introduce a multi-view contrastive learning method that captures semantic correspondences both among multi-view radiographs within a study and between these radiographs and their associated report, thereby improving visual representation learning. We further present a knowledge-guided report generation module that integrates available patient-specific knowledge (i.e., indication, which includes symptom descriptions) to facilitate the generation of accurate and coherent radiology reports. To support research in multi-view report generation, we construct Multiview CXR and Two-view CXR datasets using publicly available sources. Our proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, achieving a 2.9% F1 RadGraph improvement on MIMIC-CXR, a 5.0% BLEU-1 improvement on MIMIC-ABN, a 1.5% BLEU-4 improvement on Multi-view CXR, and an 8.2% F1,mic-14 CheXbert improvement on Two-view CXR. AVAILABILITY: Code is publicly available at https://github.com/mk-runner/EVOKE, with an archived release available on Zenodo (doi:10.5281/zenodo.21000219).
Qiguang Miao, Kang Liu 0025, Zhuoqi Ma, Yunan Li 0001, Xiaolu Kang, Ruixuan Liu, Kun Xie 0011
Bioinform.3
2026 Factual serialization enhancement: A key innovation for chest X-ray report generation
Kang Liu 0025, Zhuoqi Ma, Zhicheng Jiao, Xiaolu Kang, Qiguang Miao, Kun Xie 0011
Expert Syst. Appl.2
2026 Abn-BLIP: Abnormality-aligned Bootstrapping Language-Image Pre-training for pulmonary embolism diagnosis and report generation from CTPA
abstract
Medical imaging plays a pivotal role in modern healthcare, with computed tomography pulmonary angiography (CTPA) being a critical tool for diagnosing pulmonary embolism and other thoracic conditions. However, the complexity of interpreting CTPA scans and generating accurate radiology reports remains a significant challenge. This paper introduces Abn-BLIP (Abnormality-aligned Bootstrapping Language-Image Pretraining), an advanced diagnosis model designed to align abnormal findings to generate the accuracy and comprehensiveness of radiology reports. By leveraging learnable queries and cross-modal attention mechanisms, our model demonstrates superior performance in detecting abnormalities, reducing missed findings, and generating structured reports compared to existing methods. Our experiments show that Abn-BLIP outperforms state-of-the-art medical vision-language models and 3D report generation methods in both accuracy and clinical relevance. These results highlight the potential of integrating multimodal learning strategies for improving radiology reporting. The source code is available at https://github.com/zzs95/abn-blip.
Zhusi Zhong, Yuli Wang, Lulu Bi, Zhuoqi Ma, Sun Ho Ahn, Christopher J. Mullin, Colin Greineder, Michael Atalay, Scott Collins, Grayson Baird, Cheng Ting Lin, J. Webster Stayman, Todd M. Kolb, Ihab Kamel, Harrison X. Bai, Zhicheng Jiao
Medical Image Anal.4
2026 No blind alignment but generation: A different view of continuous sign language recognition based on diffusion
Xi Geng, Yunan Li 0001, Zhuoqi Ma, Zixiang Lu, Qiguang Miao
Pattern Recognit.3
2025 UrbanWaste: In-the-Bin Dataset for Waste Disposal Inspection with Multi-Granularity Hierarchical Labels
abstract
Our world faces the challenge of efficiently and responsibly managing the ever-growing volume of urban waste. Many countries and regions have implemented categorized trash bins and require residents to sort their waste according to specified criteria. Proper waste classification by residents significantly reduces the workload in the waste disposal process. However, due to the lack of effective supervision during classification, the quality of waste sorting is often compromised. This misclassification can lead to higher pollution risks, lower recycling rates, and increased waste management costs and difficulties. To address this issue, we propose using images captured from within trash bins to supervise garbage delivery. We introduce UrbanWaste, an image dataset specifically designed for in-the-bin waste detection and segmentation. The dataset includes 25,254 RGB images and 140,008 annotated items, featuring dense annotations and multi-granularity labels across 193 distinct waste categories. We evaluated state-of-the-art segmentation models to understand their generalization and performance on UrbanWaste. Based on this dataset, we developed a comprehensive workflow for waste classification inspection, which has been deployed in real-world districts to assess the system's effectiveness. We hope UrbanWaste will inspire new directions in AI research for environmental sustainability.
Zhuoqi Ma, Zejun You, Xiyue Gao, Qiguang Miao
AAAI1
2025 Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation
abstract
Automated radiology report generation offers an effective solution to alleviate radiologists’ workload. However, most existing methods focus primarily on single or fixed-view images to model current disease conditions, which limits diagnostic accuracy and overlooks disease progression. Although some approaches utilize longitudinal data to track disease progression, they still rely on single images to analyze current visits. To address these issues, we propose enhanced contrastive learning with Multi-view Longitudinal data to facilitate chest X-ray Report Generation, named MLRG. Specifically, we introduce a multi-view longitudinal contrastive learning method that integrates spatial information from current multi-view images and temporal information from longitudinal data. This method also utilizes the inherent spatiotemporal information of radiology reports to supervise the pre-training of visual and textual representations. Subsequently, we present a tokenized absence encoding technique to flexibly handle missing patient-specific prior knowledge, allowing the model to produce more accurate radiology reports based on available prior knowledge. Extensive experiments on MIMIC-CXR, MIMIC-ABN, and Two-view CXR datasets demonstrate that our MLRG outperforms recent state-of-the-art methods, achieving a 2.3% BLEU-4 improvement on MIMIC-CXR, a 5.5% F1 score improvement on MIMIC-ABN, and a 2.7% F1 RadGraph improvement on Two-view CXR.
Kang Liu 0025, Zhuoqi Ma, Xiaolu Kang, Yunan Li 0001, Kun Xie 0011, Zhicheng Jiao, Qiguang Miao
CVPR2
2025 RetouchDiffusion: Unsupervised Personalized Image Retouching via Diffusion Models
abstract
Image retouching aims to enhance the visual quality of images, but existing methods based on style-specific learning often lack the flexibility to cater to individual preferences. To address this limitation, we propose RetouchDiffusion, a personalized image retouching framework that utilizes a user-selected reference image as a guide and leverages the powerful representational capabilities of a pre-trained diffusion model to refine the retouching process. To enable more precise and adaptable adjustments, we introduce the Retouch Network, a dedicated retouching network that preprocesses brightness and tonality, providing controllable auxiliary guidance for the diffusion procedure. Experimental results demonstrate that our method produces high-quality, personalized retouched images more closely aligned with users’ aesthetic preferences across various scenarios. Our code is available at https://github.com/SuperOptimalZ/RetouchDiffusion.
Zhuoqi Ma, Zejun You, Yunan Li 0001, Qiguang Miao
ICME2
2025 Multi-scale information sharing and selection network with boundary attention for polyp segmentation
Xiaolu Kang, Zhuoqi Ma, Kang Liu 0025, Yunan Li 0001, Qiguang Miao
Eng. Appl. Artif. Intell.2
2025 Fooling human detectors via robust and visually natural adversarial patches
Dawei Zhou 0004, Hongbin Qu, Nannan Wang 0001, Chunlei Peng, Zhuoqi Ma, Xi Yang 0011, Xinbo Gao 0001
Neurocomputing5
2025 Modeling multi-scale uncertainty with evidence integration for reliable polyp segmentation
Xiaolu Kang, Zhuoqi Ma, Kang Liu 0025, Yunan Li 0001, Qiguang Miao
Neural Networks2
2025 Boundary-Aware Sentence-Gloss Alignment With Semantic Similarity Measurement for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) plays a crucial role in facilitating communication between deaf and hearing individuals. A key aspect of achieving precise CSLR is the alignment of the video segment of each sign with its gloss, namely its corresponding text representation in natural language. However, the coarticulation phenomenon, where contextual dependencies between adjacent signs blur the boundaries of individual signs, poses a significant challenge to this task. In this paper, we propose a novel boundary-aware sentence-gloss alignment network for CSLR to address this challenge. Our network first designs a task-relevant boundary-aware similarity measurement, evaluating sign frames by both appearance and their recognition contribution, mitigating coarticulation-induced transition noise to restore precise boundaries. For enhanced alignment, we propose a hierarchical sentence-gloss alignment: coarse sentence-level alignment reduces cross-modal disparity, while fine-grained gloss-level alignment refines video-to-token mapping. Finally, an adaptive class-divergence loss sharpens gloss decoding by maximizing inter-class discrimination. Our proposed framework provides a simple and effective solution to mitigate the boundary ambiguity caused by coarticulation, optimizing continuous sign language recognition algorithms from a new perspective. Extensive experiments conducted on four public sign language recognition (SLR) datasets demonstrate that our proposed boundary-aware sentence-gloss alignment network learns precise alignments and achieves state-of-the-art performance.
Yunan Li 0001, Xi Geng, Zhuoqi Ma, Qiguang Miao, Chi-Man Pun
IEEE Trans. Circuits Syst. Video Technol.3
2025 Seeking a Hierarchical Prototype for Multimodal Gesture Recognition
abstract
Gesture recognition has drawn considerable attention from many researchers owing to its wide range of applications. Although significant progress has been made in this field, previous works always focus on how to distinguish between different gesture classes, ignoring the influence of inner-class divergence caused by gesture-irrelevant factors. Meanwhile, for multimodal gesture recognition, feature or score fusion in the final stage is a general choice to combine the information of different modalities. Consequently, the gesture-relevant features in different modalities may be redundant, whereas the complementarity of modalities is not exploited sufficiently. To handle these problems, we propose a hierarchical gesture prototype framework to highlight gesture-relevant features such as poses and motions in this article. This framework consists of a sample-level prototype and a modal-level prototype. The sample-level gesture prototype is established with the structure of a memory bank, which avoids the distraction of gesture-irrelevant factors in each sample, such as the illumination, background, and the performers' appearances. Then the modal-level prototype is obtained via a generative adversarial network (GAN)-based subnetwork, in which the modal-invariant features are extracted and pulled together. Meanwhile, the modal-specific attribute features are used to synthesize the feature of other modalities, and the circulation of modality information helps to leverage their complementarity. Extensive experiments on three widely used gesture datasets demonstrate that our method is effective to highlight gesture-relevant features and can outperform the state-of-the-art methods.
Yunan Li 0001, Tianyu Qi, Zhuoqi Ma, Dou Quan, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.3
2024 Structural Entities Extraction and Patient Indications Incorporation for Chest X-Ray Report Generation
Kang Liu 0025, Zhuoqi Ma, Xiaolu Kang, Zhusi Zhong, Zhicheng Jiao, Grayson Baird, Harrison X. Bai, Qiguang Miao
MICCAI (3)2
2023 Hierarchical Category-Enhanced Prototype Learning for Imbalanced Temporal Recommendation
abstract
Temporal recommendation systems aim to suggest items to users at the optimal time. However, the significant imbalance of items in the training data poses a major challenge to predictive accuracy. Existing approaches attempt to alleviate this issue by modifying the loss function or utilizing resampling techniques, but such approaches may inadvertently amplify the specificity of certain behaviors.
Xiyue Gao, Zhuoqi Ma, Jiangtao Cui, Xiaofang Xia
ACM Multimedia2
2023 Image Hazing and Dehazing: From the Viewpoint of Two-Way Image Translation With a Weakly Supervised Framework
abstract
Image dehazing is an important task since it is the prerequisite for many downstream high-level computer vision tasks. Previous dehazing methods depend on either the hand-designed priors/assumptions or supervised learning with plenty of data, which are not easy to implement in practice. Meanwhile, synthesizing hazy images is also significant in many scenes like multi-weather image generation. In this paper, we change the viewpoint of this task to image translation and develop a weakly supervised framework to achieve it. Instead of simply considering the hazy image as the source domain and the haze-free image as the target domain for translation, we design a feature representation scheme that generates a domain indicator, and embed it into the decoder to achieve both hazing and dehazing within one network. This design significantly reduces the complexity of network and can be more easily extended to multi-domain translation tasks than the previous methods, which need one pair of generator-discriminator for each direction of the translation. Meanwhile, aiming at solving the haze-relevant task, we design a haze attention module, which takes the local entropy map as the input. Unlike the previous weakly supervised dehazing methods, our approach only requires unpaired hazy and haze-free images rather than any intermediate supervising data like the transmission map or atmospheric light defined in the atmospheric scattering model. Experimental results on synthetic datasets show our method can achieve competitive results when compared with the state-of-the-art methods and yield more appealing dehazing and hazing results on real-world images.
Yunan Li 0001, Huizhou Chen, Qiguang Miao, Siyu Liang 0002, Zhuoqi Ma, Bocheng Zhao
IEEE Trans. Multim.6
2023 Dual-Affinity Style Embedding Network for Semantic-Aligned Image Style Transfer
abstract
Image style transfer aims at synthesizing an image with the content from one image and the style from another. User studies have revealed that the semantic correspondence between style and content greatly affects subjective perception of style transfer results. While current studies have made great progress in improving the visual quality of stylized images, most methods directly transfer global style statistics without considering semantic alignment. Current semantic style transfer approaches still work in an iterative optimization fashion, which is impractically computationally expensive. Addressing these issues, we introduce a novel dual-affinity style embedding network (DaseNet) to synthesize images with style aligned at semantic region granularity. In the dual-affinity module, feature correlation and semantic correspondence between content and style images are modeled jointly for embedding local style patterns according to semantic distribution. Furthermore, the semantic-weighted style loss and the region-consistency loss are introduced to ensure semantic alignment and content preservation. With the end-to-end network architecture, DaseNet can well balance visual quality and inference efficiency for semantic style transfer. Experimental results on different scene categories have demonstrated the effectiveness of the proposed method.
Zhuoqi Ma, Xin Li 0106, Fu Li 0003, Dongliang He, Errui Ding, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer
abstract
Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize complex style, especially when holistic global and local patterns exist. Inspired by the common painting process of drawing a draft and revising the details, we introduce a novel feed-forward method named Laplacian Pyramid Network (LapStyle). LapStyle first transfers global style patterns in low-resolution via a Drafting Network. It then revises the local details in high-resolution via a Revision Network, which hallucinates a residual image according to the draft and the image textures extracted by Laplacian filtering. Higher resolution details can be easily generated by stacking Revision Networks with multiple Laplacian pyramid levels. The final stylized image is obtained by aggregating outputs of all pyramid levels. Experiments demonstrate that our method can synthesize high quality stylized images in real time, where holistic style patterns are properly transferred.
Zhuoqi Ma, Fu Li 0003, Dongliang He, Xin Li 0106, Errui Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
CVPR2
2020 Semantic-related image style transfer with dual-consistency loss
Zhuoqi Ma, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neurocomputing1
2020 Image style transfer with collection representation space and semantic-guided reconstruction
Zhuoqi Ma, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks1
2018 From Reality to Perception: Genre-Based Neural Image Style Transfer
abstract
We introduce a novel thought for integrating artists’ perceptions on the real world into neural image style transfer process. Conventional approaches commonly migrate color or texture patterns from style image to content image, but the underlying design aspect of the artist always get overlooked. We want to address the in-depth genre style, that how artists perceive the real world and express their perceptions in the artwork. We collect a set of Van Gogh’s paintings and cubist artworks, and their semantically corresponding real world photos. We present a novel genre style transfer framework modeled after the mechanism of actual artwork production. The target style representation is reconstructed based on the semantic correspondence between real world photo and painting, which enable the perception guidance in style transfer. The experimental results demonstrate that our method can capture the overall style of a genre or an artist. We hope that this work provides new insight for including artists’ perceptions into neural style transfer process, and helps people to understand the underlying characters of the artist or the genre.
Zhuoqi Ma, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
IJCAI1
2017 Semantic Segmentation Based Automatic Two-Tone Portrait Synthesis
Zhuoqi Ma, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
ICIG (3)1