Ziyun Li 0002

dblp:164/4195-2 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0003-4286-9030ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LotBoNC: Novel Botnet Traffic Classification under Long-tailed Distributions
abstract
Botnets are a persistent cyber threat, leveraging global infrastructures to launch large-scale attacks. Yet, most existing classification methods are evaluated under balanced and closed-set assumptions, which fail to capture real-world conditions. In practice, botnet traffic is both long-tailed and open-world: unknown variants continually emerge, and rare threats are buried under dominant traffic, often evading detection. To reflect real-world conditions, we define a deployment-oriented setting where unlabeled traffic follows a long-tailed distribution, with dominant known classes in the head and rare novel botnet variants in the tail. We propose LotBoNC, a unified framework for encrypted traffic classification under long-tailed open-world conditions. LotBoNC first performs self-supervised pre-training to learn transferable representations, then applies entropy-regularized optimal transport to assign pseudo-labels aligned with estimated class priors. An EM-style loop iteratively refines prototypes and priors, improving class separation between frequent and rare categories. We evaluate LotBoNC on three public encrypted traffic datasets with diverse long-tailed scenarios. LotBoNC consistently outperforms prior state-of-the-art methods and accurately classifies known botnets and discovers unseen botnet variants in diverse, umbalanced open-world scenarios.
Huancheng Hu, Ziyun Li 0002, Christian Doerr
AsiaCCS2
2025 QuARF: Quality-Adaptive Receptive Fields for Degraded Image Perception
abstract
Advanced Deep Neural Networks (DNNs) perform well for high-quality images, but their performance dramatically decreases for degraded images. Data augmentation is commonly used to alleviate this problem, but using too much perturbed data might seriously decrease the performance on pristine images. To tackle this challenge, we take our cue from the assumption of spatial coincidence in human visual perception, i.e. multiscale and varying receptive fields are required for understanding pristine and degraded images. Correspondingly, we propose a novel plug-and-play network architecture, dubbed Quality-Adaptive Receptive Fields (QuARF), to automatically select the optimal receptive fields based on the quality of the input image. To this end, we first design a multi-kernel convolutional block, which comprises multiscale continuous receptive fields. Afterward, we design a quality-adaptive routing network to predict the significance of each kernel, based on the quality features extracted from the input image. In this way, QuARF automatically selects the optimal inference route for each image. To further boost efficiency and effectiveness, the input feature map is split into multiple groups, with each group independently learning its quality-adaptive routing parameters. We apply QuARF to a variety of DNNs and conduct experiments in both discriminative and generation tasks, including semantic segmentation, image translation, and restoration. Thorough experimental results show that QuARF significantly and robustly improves the performance for degraded images, and outperforms data augmentation in most cases.
Fei Gao 0006, Ziyun Li 0002, Wenwang Han, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
AAAI3
2025 Q-Norm: Robust Representation Learning via Quality-Adaptive Normalization
Lanning Zhang, Fei Gao 0006, Ziyun Li 0002, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
ICCV4
2025 BI-RADS Boosted Breast Cancer Diagnosis With Masked Pretraining On Imbalanced Ultrasound Data
abstract
In clinical diagnosis, the Breast Imaging Reporting and Data System (BI-RADS) levels are highly correlated with pathological categories (benign or malignant). Thus, in this paper, we propose a BI-RADS Boosted Breast Cancer Diagnosis (B3CD) method, for joint predicting both the BI-RADS levels and pathological categories. Specifically, we first train two networks for each task for learning task specific features, and then fuse them through dual spatial attention. Besides, the network backbones are initialized through masked pretraining, due to the limited amount of labeled data. A balanced cross-entropy loss is used for the BI-RADS prediction branch to combat the extremely imbalanced distribution of BI-RADS levels. Experimental results demonstrate that B3CD achieves remarkably superior performance in both breast cancer diagnosis and BI-RADS prediction tasks, across the GDPH&SYSUCC, BUSBRA, and Breast-Lesions-USG datasets. Our code has been released at: https://github.com/AiArt-Gao/B3CD.
Xueqian Pang, Ziyun Li 0002, Junhui Lv, Ruiquan Ge, Zhuoxuan Wu, Fei Gao 0006
ICME2
2025 CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable Quantization
abstract
Face cartoonization remains a challenging task due to significant geometric deformations between facial photos and cartoons, as well as the absence of paired training data for supervised learning. Existing methods struggle to generate high-quality cartoonized avatars with attribute consistency. To address this challenge, this paper proposes an unsupervised facial cartoonization method based on cross-domain aligned and deformable vector quantization (CADQ). Firstly, we construct textual descriptions with facial attributes for both photo datasets and cartoon collections. Attribute consistency during transformation is enforced through individually contrastive learning between image-text cross-modal features and globally distribution alignment across photo-cartoon domains. Secondly, a deformable Transformer with dual attention is introduced during the transformation process, which queries corresponding cartoon codebook entries based on image features to simulate cross-domain geometric deformations. Experimental results demonstrate that the proposed method can convert facial photos into high-quality cartoons with attribute consistency, outperforming existing state-of-the-art approaches. Furthermore, the method can be effectively extended to unsupervised cross-domain generation of other artistic portrait styles, achieving superior or highly competitive performance. Our code has been released at: https://github.com/IIP-Lab-XDU/CADQ.
Yongjie Hu, Ziyun Li 0002, Fei Gao 0006, Henrik Boström, Nannan Wang 0001
ACM Multimedia3
2025 S3OIL: Semi-Supervised SAR-to-Optical Image Translation via Multi-Scale and Cross-Set Matching
abstract
Image-to-image translation has achieved great success, but still faces the significant challenge of limited paired data, particularly in translatingSynthetic Aperture Radar(SAR) images to optical images. Furthermore, most existing semi-supervised methods place limited emphasis on leveraging the data distribution. To address those challenges, we propose aSemi-Supervised SAR-to-Optical Image Translation(S3OIL) method that achieves high-quality image generation using minimal paired data and extensive unpaired data while strategically exploiting the data distribution. To this end, we first introduce aCross-Set Alignment Matching(CAM) mechanism to create local correspondences between the generated results of paired and unpaired data, ensuring cross-set consistency. In addition, for unpaired data, we apply weak and strong perturbations and establish intra-setMulti-Scale Matching(MSM) constraints. For paired data, intra-modal semantic consistency (ISC) is presented to ensure alignment with the ground truth. Finally, we propose local and global cross-modal semantic consistency (CSC) to boost structural identity during translation. We conduct extensive experiments on SAR-to-optical datasets and another sketch-to-anime task, demonstrating that S3OIL delivers competitive performance compared to state-of-the-art unsupervised, supervised, and semi-supervised methods, both quantitatively and qualitatively. Ablation studies further reveal that S3OIL can ensure the preservation of both semantic content and structural integrity of the generated images. Our code is available at: https://github.com/XduShi/SOIL.
Xi Yang 0011, Ziyun Li 0002, Maoying Qiao, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Image Process.3
2020 Style-adaptive photo aesthetic rating via convolutional neural networks and multi-task learning
Fei Gao 0006, Ziyun Li 0002, Jun Yu 0002, Junze Yu, Qingming Huang, Qi Tian 0001
Neurocomputing2
2017 Convolutional neural networks for intestinal hemorrhage detection in wireless capsule endoscopy images
abstract
Wireless capsule endoscopy (WCE) can painlessly capture a large number of images inside the intestine. However, only a small portion of these WCE images contain hemorrhage. It is thus critical to develop automated hemorrhage detection method to facilitate the diagnosis of intestinal diseases. However, automated hemorrhage detection is complicated by 1) the extreme imbalance between the amount of hemorrhage images and that of normal images; and 2) the variety of the appearance, texture, and luminance inside the intestine. In this paper, we proposed to learn a robust intestinal hemorrhage detection model via Convolutional Neural Networks (CNNs), because of CNNs' extraordinary performance in solving various image understanding tasks. Specially, we explored different CNN architectures and data augmentation methods. Besides, we investigated the correlation between hemorrhage detection accuracy and image quality. Across about 1.3k hemorrhage images and 40k normal images, the learned CNN model achieves an F-measure of 98.87%.
Panpeng Li, Ziyun Li 0002, Fei Gao 0006, Jun Yu 0002
ICME2