Yuli Wang

dblp:91/7695 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Deep Cross-Branch Multi-Modal Fusion Network for early Alzheimer's diagnosis
Jiaqiang Li, Yian Gao, Zhenghua Guan, Teng Cheng, Rengmin Wu, Aocai Yang, Manxi Xu, Yuli Wang, Peng Yang 0011, Tianfu Wang 0001, Guolin Ma, Bai Ying Lei
Artif. Intell. Medicine9
2026 Cross-chain identity privacy protection scheme based on oblivious transfer protocol and key agreement
Yuli Wang, Zhichao Cai, Bin Ma 0003
Inf. Sci.1
2026 Abn-BLIP: Abnormality-aligned Bootstrapping Language-Image Pre-training for pulmonary embolism diagnosis and report generation from CTPA
abstract
Medical imaging plays a pivotal role in modern healthcare, with computed tomography pulmonary angiography (CTPA) being a critical tool for diagnosing pulmonary embolism and other thoracic conditions. However, the complexity of interpreting CTPA scans and generating accurate radiology reports remains a significant challenge. This paper introduces Abn-BLIP (Abnormality-aligned Bootstrapping Language-Image Pretraining), an advanced diagnosis model designed to align abnormal findings to generate the accuracy and comprehensiveness of radiology reports. By leveraging learnable queries and cross-modal attention mechanisms, our model demonstrates superior performance in detecting abnormalities, reducing missed findings, and generating structured reports compared to existing methods. Our experiments show that Abn-BLIP outperforms state-of-the-art medical vision-language models and 3D report generation methods in both accuracy and clinical relevance. These results highlight the potential of integrating multimodal learning strategies for improving radiology reporting. The source code is available at https://github.com/zzs95/abn-blip.
Zhusi Zhong, Yuli Wang, Lulu Bi, Zhuoqi Ma, Sun Ho Ahn, Christopher J. Mullin, Colin Greineder, Michael Atalay, Scott Collins, Grayson Baird, Cheng Ting Lin, J. Webster Stayman, Todd M. Kolb, Ihab Kamel, Harrison X. Bai, Zhicheng Jiao
Medical Image Anal.2
2026 Can Watermarks Be Removed Like Noise? A Watermarking Attack Network Using Residual Diffusion Model
abstract
Digital image watermarking is a critical technology for image copyright protection. The concurrent evolution of watermarking attacks and defenses has spurred rapid advancements in the field. However, watermarking attack methods have lagged behind, often facing two primary challenges: limited watermark removal ability and quality degradation of the attacked image. In this paper, we introduce a Watermarking Attack method based on the Residual Diffusion Model, termed WARDM. Our WARDM treats watermark information as noise and leverages the powerful image reconstruction capabilities of the diffusion model to effectively remove the watermark. Specifically, we construct a Markov chain based on the residuals between the host and watermarked images, and employ reverse propagation to reconstruct the original host image. To optimally balance watermark removal ability and image quality, we incorporate a noise schedule into WARDM that controls both the velocity and intensity of noise at each stage of the Markov chain. Extensive experiments demonstrate the superior performance of WARDM in both watermark removal capability and visual quality preservation, achieving an improvement of 5.39% in PSNR over state-of-the-art methods. Moreover, our method demonstrates strong generalization, effectively executing attacks across a variety of watermarking techniques.
Chunpeng Wang 0001, Shanshan Zhang 0001, Yunan Liu 0001, Yuli Wang, Qi Li 0029
IEEE Trans. Dependable Secur. Comput.5
2025 Active Defense Against Deepfakes: An Integrated Framework of Adversarial Data Embedding and Blockchain Authentication
Yuli Wang, Yiyan Liang, Bin Ma 0003
ICONIP (3)1
2025 Enhancing Radiology Report Interpretation through Modality-Specific RadGraph Fine-Tuning
Haoyue Guan, Yuwei Dai, Shadi Afyouni, Alec Kain, Wen-Chi Hsu, Jiashu Cheng, Sophie Yao, Yuli Wang, Rishitha Pulakhandam, Lin-Mei Zhao, Chengzhang Zhu, Zhicheng Jiao, Craig Jones, Harrison Bai
MICCAI (7)8
2024 DTBNet: medical image segmentation model based on dual transformer bridge
abstract
Accurate medical image segmentation [1] techniques help doctors make better diagnoses. The traditional U-Net [2] model usually uses skip connections to directly connect the corresponding layers of the encoder and decoder, which can induce the model to retain more spatial information. Therefore, a series of ways to optimise skip connections, such as the introduction of residual connections, have appeared to optimise the convergence speed of the model and to improve the model’s performance. Although skip connections have achieved remarkable success in the field of medical image segmentation, since the structure of medical images is more complex than that of general images, skip connections may cause the model to learn some irrelevant features when performing image segmentation, and the transfer of the bottom layer information to the top layer will also cause some interference, in order to better achieve the multi-scale fusion of the bottom layer features and the top layer features and to reduce the interference between them, we proposed DTBNet for medical image segmentation, which is a U-shaped network with a bridge structure over the skip connections, and is characterised by the fact that the features output from the different layers of the encoder are first processed by the bridge structure before being transmitted to the corresponding layers of the decoder. Next, we design a SE_ASPP module in the encoder that can capture information using different scales of receptive fields and adaptively adjusts the weights for each channel. We also propose a neighbour fusion module(NFM) in the encoder for fusing features from adjacent layers of the encoder to enhance the representation of features at different scales. In addition, we construct a perceptual loss function [3] to comprehensively supervise the task-aware features from the bottom to the top layer, which helps the model to better capture the structures and textures in medical images. Our method achieves sota performance compared to previous work under different evaluation metrics on two medical image segmentation datasets including abdomen and heart.
Kequan Chen, Yuli Wang, Jinyong Cheng
IJCNN2
2024 CEC-YOLO: An Improved Steel Defect Detection Algorithm Based on YOLOv5
abstract
Steel surface defect detection is a critical step in the steel manufacturing process and an important guarantee to improve the quality of steel production. However, the contrast of steel surface defect images is poor, the defects are complex and irregularly distributed, and the existing steel surface defect detection algorithms have some problems, such as low detection accuracy and slow detection speed. In this paper, a steel surface defect detection algorithm named CEC-YOLO is proposed. Building upon YOLOv5, we first replaced the C3 module with the C2f and C2f-DSC modules to enhance the detection accuracy of slender and weak local structural features while maintaining a lightweight model. Additionally, we introduce an Explicit Vision Center (EVC) block in the backbone network to capture global remote dependencies among top-level features and extract feature representations of local corner regions for comprehensive analysis. Finally, we integrate Coordinate Attention (CA) into the detection head to fuse feature information effectively and improve the network’s ability to locate defects accurately. We conducted extensive experiments on two real-world datasets: NEU-DET and Micro surface defect database. The experimental results demonstrate that the improved model achieves average accuracies of 80.8% and 91.3%, respectively, at mAP@IoU = 0.5, surpassing the baseline by 4.6% and 5.5%. Furthermore, comparative experiments highlight the superiority of our improved model over other state-of-the-art approaches.
Yuli Wang, Qiliang Gu
IJCNN2
2024 Car-Dcros: A Dataset and Benchmark for Enhancing Cardiovascular Artery Segmentation Through Disconnected Components Repair and Open Curve Snake
Yuli Wang, Wen-Chi Hsu, Victoria Shi, Gigin Lin, Cheng Ting Lin, Harrison X. Bai
MICCAI (1)1
2024 Enhancing vision-language models for medical imaging: bridging the 3D gap with innovative slice selection
abstract
Recent approaches to vision-language tasks are built on the remarkable capabilities of large vision-language models (VLMs). These models excel in zero-shot and few-shot learning, enabling them to learn new tasks without parameter updates. However, their primary challenge lies in their design, which primarily accommodates 2D input, thus limiting their effectiveness for medical images, particularly radiological images like MRI and CT, which are typically 3D. To bridge the gap between state-of-the-art 2D VLMs and 3D medical image data, we developed an innovative, one-pass, unsupervised representative slice selection method called Vote-MI, which selects representative 2D slices from 3D medical imaging. To evaluate the effectiveness of vote-MI when implemented with VLMs, we introduce BrainMD, a robust, multimodal dataset comprising 2,453 annotated 3D MRI brain scans with corresponding textual radiology reports and electronic health records. Based on BrainMD, we further develop two benchmarks, BrainMD-select (including the most representative 2D slice of 3D image) and BrainBench (including various vision-language downstream tasks). Extensive experiments on the BrainMD dataset and its two corresponding benchmarks demonstrate that our representative selection method significantly improves performance in zero-shot and few-shot learning tasks. On average, Vote-MI achieves a 14.6\% and 16.6\% absolute gain for zero-shot and few-shot learning, respectively, compared to randomly selecting examples. Our studies represent a significant step toward integrating AI in medical imaging to enhance patient care and facilitate medical research. We hope this work will serve as a foundation for data selection as vision-language models are increasingly applied to new tasks.
Yuli Wang, Peng jian, Yuwei Dai, Craig K. Jones, Haris I. Sair, Jinglai Shen, Nicolas Loizou, Wen-Chi Hsu, Maliha R. Imami, Zhicheng Jiao, Harrison X. Bai
NeurIPS1
2024 Optimized Landslide Segmentation From Satellite Imagery Based on Multiresolution Fusion and Attention Mechanism
abstract
Due to the irregular distribution of landslides and their overlap with other land features, such as vegetation, achieving precise segmentation of landslides in images remains a challenge. This letter presents an optimized landslide segmentation network to detect landslide, which includes a multiresolution fusion module (MRFM) to extract and integrate features from coarse to fine details and a selective kernel attention module (SKAM) to optimize the fused multiresolution information using progressively increasing receptive fields. This configuration effectively captures the distribution and geometric information of landslides. Extensive experiments conducted on the Bijie and Landslide4Sense datasets demonstrate that the proposed network outperforms four other networks. Among them, it exceeds SegFormer by over 7.61% in mean intersection over union (mIoU) and 8.55% in F1-score on the Bijie dataset and by over 4.05% and 7.4%, respectively, on the Landslide4Sense dataset. Ablation experiments confirm the effectiveness of the two novel modules. Additionally, by integrating terrain factors with RGB in various networks, a general improvement in landslide segmentation is observed.
Yibo Ling, Yuli Wang, Yi Lin 0002, Ting On Chan, Joseph L. Awange
IEEE Geosci. Remote. Sens. Lett.2
2022 Study on Path Planning of Multi-storey Parking Lot Based on Combined Loss Function
Zhongtian Hu, Yuli Wang, Qiming Fu 0001, Weizhong Lu, Hongjie Wu
ICIC (3)3
2022 A high-performance insulators location scheme based on YOLOv4 deep learning network with GDIoU loss function
abstract
Abstract This paper proposes a Gaussian Distance Intersection over Union (GDIoU) loss function‐based YOLOv4 deep learning network to solve the problem of slow speed and low accuracy insulator location in power facilities health inspection. In the scheme, A GDIoU loss function is designed to accelerate the convergence speed of the YOLOv4 deep learning network; at the same time, the GDIoU loss is added as one part of the network propagation loss, and the insulator's location accuracy is accordingly improved. Moreover, a re‐location scheme for tilt insulators correction is proposed to enhance the location accuracy of the insulators in different spatial angle states. Large amounts of field insulator images were gathered as training and testing samples to evaluate the performance of the proposed scheme. The experimental results have demonstrated that the GDIoU‐based YOLOv4 deep learning network combined with the tilt correction scheme can improve the insulator location speed by three times compared with the peer schemes, and the average precision is increased by 7.37% compared with the naive YOLOv4 network. The performance of the proposed scheme meets the requirement of online insulator location adequately.
Bin Ma 0003, Yongkang Fu, Chunpeng Wang 0001, Jian Li 0034, Yuli Wang
IET Image Process.5
2022 An ECG Signal Denoising Method Using Conditional Generative Adversarial Net
abstract
In this paper, a novel denoising method for electrocardiogram (ECG) signal is proposed to improve performance and availability under multiple noise cases. The method is based on the framework of conditional generative adversarial network (CGAN), and we improved the CGAN framework for ECG denoising. The proposed framework consists of two networks: a generator that is composed of the optimized convolutional auto-encoder (CAE) and a discriminator that is composed of four convolution layers and one full connection layer. As the convolutional layers of CAE can preserve spatial locality and the neighborhood relations in the latent higher-level feature representations of ECG signal, and the skip connection facilitates the gradient propagation in the denoising training process, the trained denoising model has good performance and generalization ability. The extensive experimental results on MIT-BIH databases show that for single noise and mixed noises, the average signal-to-noise ratio (SNR) of denoised ECG signal is above 39 dB, and it is better than that of the state-of-the-art methods. Furthermore, the denoised classification results of four cardiac diseases show that the average accuracy increased above 32 % under multiple noises under SNR=0 dB. So, the proposed method can remove noise effectively as well as keep the details of the features of ECG signals.
Bingchu Chen, Yuli Wang, Hui Liu 0046, Ruixia Liu, Lan Tian, Xiaoshan Lu
IEEE J. Biomed. Health Informatics4
2021 Super-Large Medical Image Storage and Display Technology Based on Concentrated Points of Interest
Yuli Wang, Haiou Li, Weizhong Lu, Hongjie Wu
ICIC (1)2
2021 Medical Image Key Area Protection Scheme Based on QR Code and Reversible Data Hiding
abstract
Medical image data, like most patient information, has high requirements for privacy and confidentiality. To improve the security of medical image transmission within the open network, we proposed a medical image key area protection algorithm based on reversible data hiding. First, the coefficient of variation is used to identify the key area, that is, the lesion area of the image. Then, the other regions are divided into blocks to analyze the texture complexity. Next, we propose a new reversible data hiding algorithm, which embeds the content of the key area into the high-texture regions. On this basis, a quick response (QR) code is generated using the ciphertext of the basic image information to replace the original lesion area. Experimental results show that this method can not only safely transmit sensitive patient information by hiding the content of the lesion, it can also store copyright information through QR code and achieve accurate image retrieval.
Jian Xu 0025, Bin Ma 0003, Chunpeng Wang 0001, Jian Li 0034, Yuli Wang
Secur. Commun. Networks6