Dong Zhang 0009

dblp:68/3245-9 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-2948-1384ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation
abstract
Multi-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and response generation, which limits the explainability of the results. To address this gap, we propose to introduce reinforcement learning into multi-modal retrieval-augmented generation, enhancing the reasoning capabilities of multi-modal large language models through a two-stage reinforcement fine-tuning framework to achieve explainable multi-modal retrieval-augmented generation. Specifically, in the first stage, rule-based reinforcement fine-tuning is employed to perform coarse-grained point-wise ranking of multi-modal documents, effectively filtering out those that are significantly irrelevant. In the second stage, reasoning-based reinforcement fine-tuning is utilized to jointly optimize fine-grained list-wise ranking and answer generation, guiding multi-modal large language models to output explainable reasoning logic in the MMRAG process. Our method achieves state-of-the-art results on WebQA and MultimodalQA, two benchmark datasets for multi-modal retrieval-augmented generation, and its effectiveness is validated through comprehensive ablation experiments.
Shengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu, Yuying Liu 0007, Dong Zhang 0009, Shaoyi Du
AAAI6
2026 HOBN: A general multi-view high-order brain network learning framework for brain disease diagnosis
Rundong Xue, Shaoyi Du, Xiangmin Han, Dong Zhang 0009, Junchang Li
Expert Syst. Appl.4
2026 AsyCMST: Asymmetric cross-modal spatio-temporal learning for multimodal ultrasound nodule recognition
Hongcheng Han, Dong Zhang 0009, Qinbo Guo, Jue Jiang, Shaoyi Du
Medical Image Anal.5
2026 Uncertainty-guided and reliable collaborative perception for open heterogeneous systems
Yihan Tian, ShuaiChen Zhu, Dong Zhang 0009, Yuying Liu 0007, Shaoyi Du
Pattern Recognit. Lett.4
2026 Keypoint-Guided Medical Video Segmentation Model With Spatiotemporal Feature Fusion
abstract
Atrial fibrillation, characterized by high prevalence and poor prognosis, presents a significant global health burden. Accurate segmentation and measurement of left ventricular and left atrial appendage morphology and function are essential for reliable risk assessment. However, these tasks are hindered by ambiguous boundaries, complex cardiac motion, and sparse annotations. To address these challenges, we propose a Keypoint-Guided Medical Video Segmentation Model with Spatiotemporal Feature Fusion (KG-STS). First, we propose a shape-constrained point encoder that explicitly encodes boundary points to improve the representation of ambiguous boundaries. Next, we introduce a motion-aware alignment module that models cardiac motion by forming coherent motion information across frames. Building on these two modules, we develop a keypoint-guided spatiotemporal feature fusion module that integrates spatial boundary representations with temporal motion cues to enhance decoding features under sparse annotations, enabling temporally consistent segmentation and supporting morphological measurement. We evaluate the segmentation and measurement performance of our method on a self-constructed multi-view transesophageal echocardiography dataset and two publicly available transthoracic echocardiography datasets. The results demonstrate that KG-STS achieves superior temporal consistency in segmentation and higher accuracy in morphological measurements compared to competing methods.
Shaoyi Du, Huanhuan Huo, Jue Jiang, Dong Zhang 0009, Hongcheng Han, Shengdi Hou
IEEE Trans. Medical Imaging5
2026 Self-Supervised T2WI-Bridged Framework for Liver Segmentation and PDFF Prediction From US Images
abstract
Proton Density Fat Fraction (PDFF) is the gold standard for non-invasive fatty liver diagnosis, but its reliance on Magnetic Resonance Imaging (MRI) limits broad clinical applicability. Motivated by the accessibility of B-mode Ultrasound (US) in fatty liver assessment, we propose a novel framework for liver segmentation and PDFF prediction from US images. To enhance generalization ability despite limited paired US-PDFF data, our framework integrates a cross-task self-supervised pretext task that extracts semantic features to guide echo intensity capture, benefiting both liver segmentation and PDFF prediction. To address the noise and artifacts inherent in US images, our framework leverages T2-weighted imaging (T2WI) exclusively during training to establish a feature bridge between US and PDFF, thereby enhancing PDFF prediction. Once trained, the model relies solely on US for inference, making it a practical and cost-effective alternative to MRI-based PDFF estimation. Additionally, our framework introduces an uncertainty-augmented adversarial loss function to refine liver boundary delineation, further improving segmentation and PDFF prediction accuracy. Experimental results demonstrate that our method outperforms state-of-the-art methods in liver segmentation and PDFF prediction; and in a specific application study, our predicted PDFF achieves accuracy comparable to real PDFF for hepatic steatosis classification, highlighting its clinical potential. The full source code and detailed documentation are publicly available at https://github.com/D0ngZhang/SSTB.
Dong Zhang 0009, Qi Zeng 0004, Tim Salcudean, Z. Jane Wang 0001
IEEE Trans. Medical Imaging1
2025 Cross-Modal Brain Graph Transformer via Function-Structure Connectivity Network for Brain Disease Diagnosis
Jingxi Feng, Heming Xu, Junhao Cai, Yujie Chang, Dong Zhang 0009, Shaoyi Du
MICCAI (12)5
2025 Hierarchical candidate recursive network for highlight restoration in endoscopic videos
Chenchu Xu, Jiangnan Wu, Dong Zhang 0009, Longfei Han, Dingwen Zhang, Junwei Han 0001
Expert Syst. Appl.3
2024 Self-supervised anatomical continuity enhancement network for 7T SWI synthesis from 3T SWI
Dong Zhang 0009, Caohui Duan, Udunna Anazodo, Z. Jane Wang 0001
Medical Image Anal.1
2024 Deep Generative Adversarial Reinforcement Learning for Semi-Supervised Segmentation of Low-Contrast and Small Objects in Medical Images
abstract
Deep reinforcement learning (DRL) has demonstrated impressive performance in medical image segmentation, particularly for low-contrast and small medical objects. However, current DRL-based segmentation methods face limitations due to the optimization of error propagation in two separate stages and the need for a significant amount of labeled data. In this paper, we propose a novel deep generative adversarial reinforcement learning (DGARL) approach that, for the first time, enables end-to-end semi-supervised medical image segmentation in the DRL domain. DGARL ingeniously establishes a pipeline that integrates DRL and generative adversarial networks (GANs) to optimize both detection and segmentation tasks holistically while mutually enhancing each other. Specifically, DGARL introduces two innovative components to facilitate this integration in semi-supervised settings. First, a task-joint GAN with two discriminators links the detection results to the GAN's segmentation performance evaluation, allowing simultaneous joint evaluation and feedback. This ensures that DRL and GAN can be directly optimized based on each other's results. Second, a bidirectional exploration DRL integrates backward exploration and forward exploration to ensure the DRL agent explores the correct direction when forward exploration is disabled due to lack of explicit rewards. This mitigates the issue of unlabeled data being unable to provide rewards and rendering DRL unexplorable. Comprehensive experiments on three generalization datasets, comprising a total of 640 patients, demonstrate that our novel DGARL achieves 85.02% Dice and improves at least 1.91% for brain tumors, achieves 73.18% Dice and improves at least 4.28% for liver tumors, and achieves 70.85% Dice and improves at least 2.73% for pancreas compared to the ten most recent advanced methods, our results attest to the superiority of DGARL. Code is available at GitHub.
Chenchu Xu, Dong Zhang 0009, Dingwen Zhang, Junwei Han 0001
IEEE Trans. Medical Imaging3
2023 Color-Difference Correntropy Guided Convolution Network for Point Cloud Semantic Segmentation
abstract
With the development of data acquisition technology, RGB color information is widely collected to strengthen the 3D point cloud. To some extent, RGB color information contains the prior relation about the spatial position of objects. However, the existing point cloud segmentation networks based on deep learning do not pay much attention to it. To remedy the lack in this area, we propose a color-difference correntropy guided convolution network, which introduces correntropy to optimize the measurement of color-difference. Meanwhile, we select points in the local neighborhood via the color-difference guided module, and construct an ordered sequence of points with correlation information, which not only facilitates the feature extraction by directly applying the convolution but also fully studies the correlation between the color information and spatial position of the point cloud. Moreover, we fuse the sequence features extracted by convolution with the geometric features acquired by MLP to get new features with more abundant semantic information, thus improving the segmentation performance. On both indoor and outdoor datasets, the experimental results demonstrate the effectiveness and superiority of the proposed method by the comparison experiments and ablation experiments.
Zhou Jiang 0001, Jing Yang 0014, Chunyu Xuan, Dong Zhang 0009, Shaoyi Du
IJCNN4
2023 Heuristic multi-modal integration framework for liver tumor detection from multi-modal non-enhanced MRIs
Dong Zhang 0009, Chenchu Xu, Shuo Li 0001
Expert Syst. Appl.1
2023 Spatiotemporal knowledge teacher-student reinforcement learning to detect liver tumors without contrast agents
Chenchu Xu, Yuhong Song, Dong Zhang 0009, Leonardo Kayat Bittencourt, Sree Harsha Tirumani, Shuo Li 0001
Medical Image Anal.3
2023 Coarse-to-fine feature representation based on deformable partition attention for melanoma identification
Dong Zhang 0009, Jing Yang 0014, Shaoyi Du, Hongcheng Han, Yuyan Ge, Longfei Zhu, Ce Li 0001, Meifeng Xu, Nanning Zheng 0001
Pattern Recognit.1
2023 Dual Uncertainty-Guided Mixing Consistency for Semi-Supervised 3D Medical Image Segmentation
abstract
3D semi-supervised medical image segmentation is extremely essential in computer-aided diagnosis, which can reduce the time-consuming task of performing annotation. The challenges with current 3D semi-supervised segmentation algorithms includes the methods, limited attention to volume-wise context information, their inability to generate accurate pseudo labels and a failure to capture important details during data augmentation. This article proposes a dual uncertainty-guided mixing consistency network for accurate 3D semi-supervised segmentation, which can solve the above challenges. The proposed network consists of a Contrastive Training Module which improves the quality of augmented images by retaining the invariance of data augmentation between original data and their augmentations. The Dual Uncertainty Strategy calculates dual uncertainty between two different models to select a more confident area for subsequent segmentation. The Mixing Volume Consistency Module that guides the consistency between mixing before and after segmentation for final segmentation, uses dual uncertainty and can fully learn volume-wise context information. Results from evaluative experiments on brain tumor and left atrial segmentation shows that the proposed method outperforms state-of-the-art 3D semi-supervised methods as confirmed by quantitative and qualitative analysis on datasets. This effectively demonstrates that this study has the potential to become a medical tool for accurate segmentation. Code is available at:https://github.com/yang6277/DUMC.
Chenchu Xu, Zhiqiang Xia, Dong Zhang 0009, Yanping Zhang 0001, Shu Zhao 0005
IEEE Trans. Big Data5
2023 BMAnet: Boundary Mining With Adversarial Learning for Semi-Supervised 2D Myocardial Infarction Segmentation
abstract
Automatic segmentation of myocardial infarction (MI) regions in late gadolinium-enhanced cardiac magnetic resonance images is an essential step in the computed diagnosis of myocardial infarction. Most of the current myocardial infarction region segmentation methods are based on fully supervised deep learning. However, cardiologists' annotation of myocardial infarction regions in cardiac magnetic resonance images during the diagnosis process is time-consuming and expensive. This paper proposes a semi-supervised myocardial infarction segmentation. It consists of two models: 1) a boundary mining model and 2) an adversarial learning model. The boundary mining model can solve the boundary ambiguity problem by enlarging the gap between the foreground and background features, thus segmenting the myocardial infarction region accurately. The adversarial learning model can make the boundary mining model learn from additional unlabeled data by evaluating the segmentation performance and providing pseudo supervision, which significantly increases the robustness of the boundary mining model. We conduct extensive experiments on an in-house myocardial magnetic resonance dataset. The experimental results on six evaluation metrics demonstrate that our method achieves excellent results in myocardial infarction segmentation and outperforms the state-of-the-art semi-supervised methods.
Chenchu Xu, Dong Zhang 0009, Longfei Han, Yanping Zhang 0001, Jie Chen 0025, Shuo Li 0001
IEEE J. Biomed. Health Informatics3
2023 An Uncertainty-Aware and Sex-Prior Guided Biological Age Estimation From Orthopantomogram Images
abstract
Bone age, as a measure of biological age (BA), plays an important role in a variety of fields, including forensics, orthodontics, sports, and immigration. Despite its significance, accurate estimation of BA remains a challenge due to the uncertainty error between BA and chronological age (CA) caused by individual diversity and the difficult integration of multiple factors, such as sex, and identified or measured anatomical structures, into the estimation process. To address problems, we propose an uncertainty-aware and sex-prior guided biological age estimation from orthopantomogram images (OPGs), named UASP-BAE, which models uncertainty errors while setting sex dimorphism as tractive features to enhance age-related specific features, aiming to improve the accuracy of BA estimation. Furthermore, considering the global relevance of the anatomic structure, such as the mandible, teeth, maxillary sinus, etc., a cross-attention module based on CNN and self-attention is proposed to mine the local texture and global semantic features of OPGs. Moreover, we design a novel age composition loss by cross-entropy, probability bias, and regression functions, aiming at evaluating BA's uncertainty errors and results to obtain an accurate and robust model. On 10703 OPGs from 5.00 to 25.00 years of age, our model had a best MAE value of 0.8005 years and higher than the comparison popular algorithms, which also demonstrates the method's potential for improved accuracy in BA estimation.
Dong Zhang 0009, Jing Yang 0014, Shaoyi Du, Wenqing Bu, Yu-Cheng Guo
IEEE J. Biomed. Health Informatics1
2022 Contrast-Free Liver Tumor Detection Using Ternary Knowledge Transferred Teacher-Student Deep Reinforcement Learning
Chenchu Xu, Dong Zhang 0009, Yuhui Song, Leonardo Kayat Bittencourt, Sree Harsha Tirumani, Shuo Li 0001
MICCAI (5)2
2022 Stroke Lesion Segmentation from Low-Quality and Few-Shot MRIs via Similarity-Weighted Self-ensembling Framework
Dong Zhang 0009, Raymond Confidence, Udunna Anazodo
MICCAI (5)1
2022 Region-aware network: Model human's Top-Down visual perception mechanism for crowd counting
Yuehai Chen, Jing Yang 0014, Dong Zhang 0009, Badong Chen, Shaoyi Du
Neural Networks3
2021 Applying Cross-Modality Data Processing for Infarction Learning in Medical Internet of Things
abstract
Cross-modality data processing is critical for the Internet-of-Things (IoT) deployment in healthcare. It can convert the innumerable raw day-to-day medical big data from massive IoT-based medical devices to diagnostic valuable data so that they can be feed to clinical routine. In this article, we propose a novel spatiotemporal two-streams generative adversarial network (SpGAN) as a cross-modality data processing approach to deploy the medical IoT in infarction learning. Our SpGAN remotely converts diagnostic valuable contrast-enhanced images (the “gold standard” for infarction learning, but it requires the injection of contrast agents) directly from raw nonenhanced cine MR images. This converting allows physicians to remotely perform infarction observation and analysis to break through the limitations of time and space by building a cloud computing platform of IoT-based MRI devices. Importantly, this converting offers a low-risk IoT-based manner to eliminate the potential fatal risk caused by contrast agent injection in the current infarction learning workflow. Specifically, SpGAN consists of: 1) a spatiotemporal two-stream framework as an encoding–decoding model to achieve data converting and 2) a spatiotemporal pyramid network enhances those features that are responsible to the infarction learning during encoding to improve decoding performance. Real IoT-based remote diagnosis experiments performed on 230 patients demonstrate that SpGAN provides high-quality converted images for infarction learning and promotes the in-depth application and deployment of IoT in the medical field.
Chenchu Xu, Zhifan Gao, Dong Zhang 0009, Jinglin Zhang 0003, Lei Xu 0037, Shuo Li 0001
IEEE Internet Things J.3
2021 Weakly-Supervised teacher-Student network for liver tumor segmentation from non-enhanced images
Dong Zhang 0009, Bo Chen 0013, Jaron Chong, Shuo Li 0001
Medical Image Anal.1
2021 Synthesis of gadolinium-enhanced liver tumors on nonenhanced liver MR images using pixel-level graph reinforcement learning
Chenchu Xu, Dong Zhang 0009, Jaron Chong, Bo Chen 0013, Shuo Li 0001
Medical Image Anal.2
2021 Sequential conditional reinforcement learning for simultaneous vertebral body detection and segmentation with modeling the spine anatomy
Dong Zhang 0009, Bo Chen 0013, Shuo Li 0001
Medical Image Anal.1
2020 Enhanced image no-reference quality assessment based on colour space distribution
abstract
In this study, the authors investigate the problem of enhanced image no‐reference (NR) quality assessment. For resolving the problem of the enhanced images, it is difficult to obtain reference images, this study proposes an NR image quality assessment (IQA) model based on colour space distribution. Given an enhanced image, our method first uses a gist to select a clear target image in which the scene, colour and quality are similar to the hypothetical reference images. And then, the colour transfer is used between the input images and target images to construct the reference image. Next, the appropriate IQA method is used to assess enhanced image quality. The absolute colour difference and feature similarity (FSIM) are used to measure the colour and grey‐scale image quality, respectively. Extensive experiments demonstrate that the proposed method is good at evaluating enhanced image quality for X‐ray, dust, underwater and low‐light images. The experimental results are consistent with human subjective evaluation and achieve good assessment effects.
Hao Liu 0060, Ce Li 0001, Dong Zhang 0009, Yannan Zhou, Shaoyi Du
IET Image Process.3
2020 Holistic multitask regression network for multiapplication shape regression segmentation
Clara M. Tam, Dong Zhang 0009, Bo Chen 0013, Terry M. Peters, Shuo Li 0001
Medical Image Anal.2
2020 Predicting COVID-19 in China Using Hybrid AI Model
abstract
The coronavirus disease 2019 (COVID-19) breaking out in late December 2019 is gradually being controlled in China, but it is still spreading rapidly in many other countries and regions worldwide. It is urgent to conduct prediction research on the development and spread of the epidemic. In this article, a hybrid artificial-intelligence (AI) model is proposed for COVID-19 prediction. First, as traditional epidemic models treat all individuals with coronavirus as having the same infection rate, an improved susceptible-infected (ISI) model is proposed to estimate the variety of the infection rates for analyzing the transmission laws and development trend. Second, considering the effects of prevention and control measures and the increase of the public's prevention awareness, the natural language processing (NLP) module and the long short-term memory (LSTM) network are embedded into the ISI model to build the hybrid AI model for COVID-19 prediction. The experimental results on the epidemic data of several typical provinces and cities in China show that individuals with coronavirus have a higher infection rate within the third to eighth days after they were infected, which is more in line with the actual transmission laws of the epidemic. Moreover, compared with the traditional epidemic models, the proposed hybrid AI model can significantly reduce the errors of the prediction results and obtain the mean absolute percentage errors (MAPEs) with 0.52%, 0.38%, 0.05%, and 0.86% for the next six days in Wuhan, Beijing, Shanghai, and countrywide, respectively.
Nanning Zheng 0001, Shaoyi Du, Jianji Wang 0001, Wenting Cui, Zijian Kang, Tao Yang 0032, Bin Lou, Yuting Chi, Hong Long, Mei Ma, Dong Zhang 0009, Jingmin Xin
IEEE Trans. Cybern.14
2020 A Multi-Label Classification Method Using a Hierarchical and Transparent Representation for Paper-Reviewer Recommendation
abstract
The paper-reviewer recommendation task is of significant academic importance for conference chairs and journal editors. It aims to recommend appropriate experts in a discipline to comment on the quality of papers of others in that discipline. How to effectively and accurately recommend reviewers for the submitted papers is a meaningful and still tough task. Generally, the relationship between a paper and a reviewer often depends on the semantic expressions of them. Creating a more expressive representation can make the peer-review process more robust and less arbitrary. So the representations of a paper and a reviewer are very important for the paper-reviewer recommendation. Actually, a reviewer or a paper often belongs to multiple research fields, which increases difficulty in paper-reviewer recommendation. In this article, we propose a Multi-Label Classification method using a HIErarchical and transPArent Representation named Hiepar-MLC . First, we introduce HIErarchical and transPArent Representation (Hiepar) to express the semantic information of the reviewer and the paper. Hiepar is learned from a two-level bidirectional gated recurrent unit based network applying the attention mechanism. It is capable of capturing the two-level hierarchical information (word-sentence-document) and highlighting the elements in reviewers or papers to support the labels. This word-sentence-document information mirrors the hierarchical structure of a reviewer or a paper and captures the exact semantics of them. Then we transform the paper-reviewer recommendation problem into a multi-level classification issue, whose multiple research labels exactly guide the learning process. It is flexible in that we can select any multi-label classification method to solve the paper-reviewer recommendation problem. Further, we propose a simple multi-label-based reviewer assignment (MLBRA) strategy to select the appropriate reviewers. It is interesting in that we also explore the paper-reviewer recommendation in the coarse-grain granularity. Extensive experiments on the real-world dataset consisting of the papers in the ACM Digital Library show that Hiepar-MLC achieves better label prediction performance than the existing representation alternatives. In addition, with the MLBRA strategy, we show the effectiveness and the feasibility of our transformation from paper-reviewer recommendation to multi-label classification.
Dong Zhang 0009, Shu Zhao 0005, Zhen Duan, Jie Chen 0025, Yanping Zhang 0001, Jie Tang 0001
ACM Trans. Inf. Syst.1
2019 A Deep Reinforcement Learning Framework for Frame-by-Frame Plaque Tracking on Intravascular Optical Coherence Tomography Image
Gongning Luo, Suyu Dong, Kuanquan Wang, Dong Zhang 0009, Yue Gao 0002, Xin Chen 0025, Henggui Zhang, Shuo Li 0001
MICCAI (1)4
2018 Segmentation in Weakly Labeled Videos via a Semantic Ranking and Optical Warping Network
abstract
Weakly supervised video object segmentation (WSVOS) focuses on generating pixel-level object masks for videos only tagged with class labels, which is an essential yet challenging task. For WSVOS, the algorithm is just aware of rough category information rather than the concrete object size and location cues, besides it lacks reliable annotated exemplars to learn temporal evolution in the investigated videos. Basically, there are three challenging factors which may influence the performance of WSVOS: foreground object discovery in each frame, coarse object semantic consistency within each video, and fine-grained segmentation smoothness within neighbor frames. In this paper, we establish a semantic ranking and optical warping network (SROWN) to simultaneously solve these three challenges in a unified framework. For the first challenge, we apply the still image saliency detection method and discover the foreground object for each frame via a segmentation network. Due to the huge discrepancies between the image saliency and the video object segmentation, we step further and propose two subnetworks to solve the other two challenges. For the second one, we propose an attentive semantic ranking subnetwork to mine video-level tags, which can learn discriminative features for semantic ranking and lead to semantic consistent segmentation masks. For the third one, we propose an optical flow warping subnetwork to constrain fine-grained segmentation smoothness within neighbor frames, which can suppress the large deformation and thus obtain smooth object boundaries for adjacent frames. Experiments on two benchmark datasets, i.e., DAVIS dataset and YouTube-Objects dataset, demonstrate the effectiveness of the proposed approach for segmenting out video objects under weak supervision.
Le Yang 0008, Junwei Han 0001, Dingwen Zhang, Nian Liu 0002, Dong Zhang 0009
IEEE Trans. Image Process.5