Wenji Wang

dblp:233/1510 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation
abstract
Personalized image generation is crucial for improving the user experience, as it renders reference images into preferred ones according to user visual preferences. Although effective, existing methods face two main issues. First, existing methods treat all items in the user's historical sequence equally when extracting user preferences, overlooking the varying semantic similarities between historical items and the reference item. Disproportionately high weights for low-similarity items distort user visual preferences for the reference item. Second, existing methods heavily rely on consistency between generated and reference images to optimize generation, which leads to underfitting user preferences and hinders personalization. To address these issues, we propose Retrieval Augmented Personalized Image GenerAtion guided by Recommendation (RAGAR). Our approach uses a retrieval mechanism to assign different weights to historical items according to their similarities to the reference item, thereby extracting more refined users' visual preferences for the reference item. Then we introduce a novel rank task based on the multi-modal ranking model to optimize the personalization of the generated images instead of forcing depend on consistency. Extensive experiments and human evaluations on three real-world datasets demonstrate that RAGAR achieves significant improvements in both personalization and semantic metrics compared to five baselines.
Run Ling, Wenji Wang, Yuting Liu 0003, Guibing Guo, Quanwei Zhang, Yexing Xu, Shuo Lu, Yihua Shao, Linying Jiang, Xingwei Wang 0001
AAAI2
2025 Privacy-preserved federated clustering with Non-IID data via GANs
Jianzhe Zhao, Wenji Wang, Zhelin Fan, Stan Matwin
J. Supercomput.2
2025 MSDUNet: A Model Based on Feature Multi-Scale and Dual-Input Dynamic Enhancement for Skin Lesion Segmentation
abstract
Melanoma is a malignant tumor originating from the lesions of skin cells. Medical image segmentation tasks for skin lesion play a crucial role in quantitative analysis. Achieving precise and efficient segmentation remains a significant challenge for medical practitioners. Hence, a skin lesion segmentation model named MSDUNet, which incorporates multi-scale deformable block (MSD Block) and dual-input dynamic enhancement module(D2M), is proposed. Firstly, the model employs a hybrid architecture encoder that better integrates global and local features. Secondly, to better utilize macroscopic and microscopic multiscale information, improvements are made to skip connection and decoder block, introducing D2M and MSD Block. The D2M leverages large kernel dilated convolution to draw out attention bias matrix on the decoder features, supplementing and enhancing the semantic features of the decoder's lower layers transmitted through skip connection features, thereby compensating semantic gaps. The MSD Block uses channel-wise split and deformable convolutions with varying receptive fields to better extract and integrate multi-scale information while controlling the model's size, enabling the decoder to focus more on task-relevant regions and edge details. MSDUNet attains outstanding performance with Dice scores of 93.08% and 91.68% on the ISIC-2016 and ISIC-2018 datasets, respectively. Furthermore, experiments on the HAM10000 dataset demonstrate its superior performance with a Dice score of 95.40%. External validation experiments based on the ISIC-2016, ISIC-2018, and HAM10000 experimental weights on the PH2 dataset yield Dice scores of 92.67%, 92.31%, and 93.46%, respectively, showcasing the exceptional generalization capability of MSDUNet. Our code implementation is publicly available at the Github.
Xiaosen Li, Linli Li, Xinlong Xing, Huixian Liao, Wenji Wang, Qiutong Dong, Xiao Qin 0005, Chang-an Yuan 0001
IEEE Trans. Medical Imaging5
2025 Slice2Mesh: 3D Surface Reconstruction From Sparse Slices of Images for the Left Ventricle
abstract
Cine MRI is a widely used technique to evaluate left ventricular function and motion, as it captures temporal information. However, due to the limited spatial resolution, cine MRI only provides a few sparse scans at regular positions and orientations, which poses challenges for reconstructing dense 3D cardiac structures, which is essential for better understanding the cardiac structure and motion in a dynamic 3D manner. In this study, we propose a novel learning-based 3D cardiac surface reconstruction method, Slice2Mesh, which directly predicts accurate and high-fidelity 3D meshes from sparse slices of cine MRI images under partial supervision of sparse contour points. Slice2Mesh leverages a 2D UNet to extract image features and a graph convolutional network to predict deformations from an initial template to various 3D surfaces, which enables it to produce topology-consistent meshes that can better characterize and analyze cardiac movement. We also introduce As Rigid As Possible energy in the deformation loss to preserve the intrinsic structure of the predefined template and produce realistic left ventricular shapes. We evaluated our method on 150 clinical test samples and achieved an average chamfer distance of 3.621 mm, outperforming traditional methods by approximately 2.5 mm. We also applied our method to produce 4D surface meshes from cine MRI sequences and utilized a simple SVM model on these 4D heart meshes to identify subjects with myocardial infarction, and achieved a classification sensitivity of 91.8% on 99 test subjects, including 49 abnormal patients, which implies great potential of our method for clinical use.
Wenji Wang, Qing Xia 0002, Zhennan Yan, Xiao Wang 0004, Shaoping Nie, Shaoting Zhang 0001
IEEE Trans. Medical Imaging3
2024 EPSViTs: A hybrid architecture for image classification based on parameter-shared multi-head self-attention
Huixian Liao, Xiaosen Li, Xiao Qin 0005, Wenji Wang, Guodui He, Xin Chun, Jinyong Zhang, Yunqin Fu, Zhengyou Qin
Image Vis. Comput.4
2024 AVDNet: Joint coronary artery and vein segmentation with topological consistency
abstract
Coronary CT angiography (CCTA) is an effective and non-invasive method for coronary artery disease diagnosis. Extracting an accurate coronary artery tree from CCTA image is essential for centerline extraction, plaque detection, and stenosis quantification. In practice, data quality varies. Sometimes, the arteries and veins have similar intensities and locate closely, which may confuse segmentation algorithms, even deep learning based ones, to obtain accurate arteries. However, it is not always feasible to re-scan the patient for better image quality. In this paper, we propose an artery and vein disentanglement network (AVDNet) for robust and accurate segmentation by incorporating the coronary vein into the segmentation task. This is the first work to segment coronary artery and vein at the same time. The AVDNet consists of an image based vessel recognition network (IVRN) and a topology based vessel refinement network (TVRN). IVRN learns to segment the arteries and veins, while TVRN learns to correct the segmentation errors based on topology consistency. We also design a novel inverse distance weighted dice (IDD) loss function to recover more thin vessel branches and preserve the vascular boundaries. Extensive experiments are conducted on a multi-center dataset of 700 patients. Quantitative and qualitative results demonstrate the effectiveness of the proposed method by comparing it with state-of-the-art methods and different variants. Prediction results of the AVDNet on the Automated Segmentation of Coronary Artery Challenge dataset are avaliabel at https://github.com/WennyJJ/Coronary-Artery-Vein-Segmentation for follow-up research.
Wenji Wang, Qing Xia 0002, Zhennan Yan, Xiao Wang 0004, Shaoping Nie, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.1
2023 Local differentially private federated learning with homomorphic encryption
Jianzhe Zhao, Chenxi Huang 0002, Wenji Wang, Rulin Xie, Rongrong Dong, Stan Matwin
J. Supercomput.3
2021 A Deep Reinforced Tree-Traversal Agent for Coronary Artery Centerline Extraction
Zhuowei Li 0002, Qing Xia 0002, Wenji Wang, Lijian Xu, Shaoting Zhang 0001
MICCAI (5)4
2021 Few-Shot Learning by a Cascaded Framework With Shape-Constrained Pseudo Label Assessment for Whole Heart Segmentation
abstract
Automatic and accurate 3D cardiac image segmentation plays a crucial role in cardiac disease diagnosis and treatment. Even though CNN based techniques have achieved great success in medical image segmentation, the expensive annotation, large memory consumption, and insufficient generalization ability still pose challenges to their application in clinical practice, especially in the case of 3D segmentation from high-resolution and large-dimension volumetric imaging. In this paper, we propose a few-shot learning framework by combining ideas of semi-supervised learning and self-training for whole heart segmentation and achieve promising accuracy with a Dice score of 0.890 and a Hausdorff distance of 18.539 mm with only four labeled data for training. When more labeled data provided, the model can generalize better across institutions. The key to success lies in the selection and evolution of high-quality pseudo labels in cascaded learning. A shape-constrained network is built to assess the quality of pseudo labels, and the self-training stages with alternative global-local perspectives are employed to improve the pseudo labels. We evaluate our method on the CTA dataset of the MM-WHS 2017 Challenge and a larger multi-center dataset. In the experiments, our method outperforms the state-of-the-art methods significantly and has great generalization ability on the unseen data. We also demonstrate, by a study of two 4D (3D+T) CTA data, the potential of our method to be applied in clinical practice.
Wenji Wang, Qing Xia 0002, Zhennan Yan, Zhuowei Li 0002, Yue Gao 0002, Dimitris N. Metaxas, Shaoting Zhang 0001
IEEE Trans. Medical Imaging1