Yu Zhu 0005

dblp:38/5267-5 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0003-1535-6520ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SFJD-Net: spatial-frequency domain joint feature enhancement with differential learning for brain stroke segmentation
Yu Zhu 0005, YaTong Liu
Appl. Intell.2
2026 FCL: frequency-based contrastive learning for generalizable face forgery detection
Yu Zhu 0005, Shengze Wang 0008, Yufeng Gu, Nan Wang 0003
Multim. Syst.1
2025 Machine Learning-Based Real-Time Detection of Power Analysis Attacks Using Supply Voltage Comparisons
abstract
Modern power analysis attacks (PAAs) pose significant threats to hardware security, and reliably securing integrated systems against advanced PAAs has become a significant design target in integrated circuits. However, the detection accuracy of most countermeasures to PAAs significantly decreases when power side-channel information is mixed with voltage noise. In this paper, a real-time PAA detection technique is proposed to achieve high detection accuracy even with large voltage noise. The voltage drops of certain power grid (PG) nodes caused by PAA are evaluated by a number of voltage comparisons between PG nodes, which compensate for the effects of voltage noise. These voltage comparison results are analyzed by machine learning algorithms, and a linear support vector machine (SVM) model is selected as the PAA detection model. The PAAs on an IBM benchmarked microprocessor are applied to evaluate the detection accuracy, and our proposed PAA detection method achieves 93.11% accuracy in detecting a resistance of 1 Ω with noise equal to 20% of Vdd. Furthermore, the power and area overheads (evaluated using a 65-nm CMOS process) of this method are reduced by 68% and 75%, respectively, compared to those of the existing machine learning-based PAA detection techniques.
Nan Wang 0003, Ruichao Liu, Yufeng Shan, Yu Zhu 0005, Song Chen 0001
ASP-DAC4
2025 CSV-Occ: Fusing Multi-frame Alignment for Occupancy Prediction with Temporal Cross State Space Model and Central Voting Mechanism
abstract
Recently, image-based 3D semantic occupancy prediction has become a hot topic in 3D scene understanding for autonomous driving. Compared with the bounding box form of 3D object detection, the ability to describe the fine-grained contours of any obstacles in the scene is the key insight of voxel occupancy representation, which facilitates subsequent tasks of autonomous driving. In this work, we propose CSV-Occ to address the following two challenges: (1) Existing methods fuse temporal information based on the attention mechanism, but are limited by high complexity. We extend the state space model to support multi-input sequence interaction and conduct temporal modeling in a cascaded architecture, thereby reducing the computational complexity from quadratic to linear. (2) Existing methods are limited by semantic ambiguity, resulting in the centers of foreground objects often being predicted as empty voxels. We enable the model to explicitly vote for the instance center to which the voxels belong and spontaneously learn to utilize the other voxel features of the same instance to update the semantics of the internal vacancies of the objects from coarse to fine. Experiments on the Occ3D-nuScenes dataset show that our method achieves state-of-the-art in camera-based 3D semantic occupancy prediction and also performs well on lidar point cloud semantic segmentation on the nuScenes dataset. Therefore, we believe that CSV-Occ is beneficial to the community and industry of autonomous vehicles.
Yu Zhu 0005, Xiaofeng Ling, Huanlei Chen, Lihua Sun
ICML2
2025 Hybrid CNN-RWKV with high-frequency enhancement for real-world chinese-english scene text image super-resolution
Yu Zhu 0005, Xiaofeng Ling
Appl. Intell.2
2025 Fine-grained grading network based on sparse transformer and spectral attention for multiparametric MR image segmentation
Yatong Liu, Yu Zhu 0005, Zeyan Zeng
Neurocomputing3
2025 Dual attention transformer with adaptive frequency enhancement for real-world Chinese-English scene text image super-resolution
Xiaofeng Ling, Yu Zhu 0005
Multim. Syst.5
2025 Occluded person re-identification based on parallel triplet augmentation and parameter-free token spatial attention
Yu Zhu 0005, Shengze Wang 0008, Jiongyao Ye, Xiaofeng Ling
Multim. Tools Appl.2
2024 MDHT-Net: Multi-scale Deformable U-Net with Cos-spatial and Channel Hybrid Transformer for pancreas segmentation
Yu Zhu 0005, Yatong Liu, Jiajun Lin
Appl. Intell.3
2024 Dynamic facial expression recognition based on spatial key-points optimized region feature fusion and temporal self-attention
Zhiwei Huang 0010, Yu Zhu 0005
Eng. Appl. Artif. Intell.2
2024 FDTNet: Enhancing frequency-aware representation for prohibited object detection from X-ray images via dual-stream transformers
Yu Zhu 0005, Nan Wang 0003, Jiongyao Ye, Xiaofeng Ling
Eng. Appl. Artif. Intell.2
2024 Local attention and long-distance interaction of rPPG for deepfake detection
Yu Zhu 0005, Xiaoben Jiang, Yatong Liu, Jiajun Lin
Vis. Comput.2
2023 Perceiving Multiple Representations for scene text image super-resolution guided by text recognizer
Yu Zhu 0005, Yatong Liu, Jiongyao Ye
Eng. Appl. Artif. Intell.2
2023 TLWSR: Weakly supervised real-world scene text image super-resolution using text label
abstract
Abstract Scene text image super‐resolution (STISR) has recently received considerable attention. Existing STISR methods are applicable to the situation that all the LR‐HR pairs are available. However, in real‐world scenarios, it is difficult and expensive to collect ground‐truth HR labels and align them with LR images, and thus it is essential to find a way to implement weakly supervised learning. We investigate the STISR problem in the situation that only a subset of HR labels is available and design a weak supervision framework using coarse‐grained text labels named TLWSR, which combines incomplete supervision and inexact supervision. Specifically, a lightweight text recognition network and connectionist temporal classification loss are used to guide the super‐resolution of text images during training. Extensive experiments on the benchmark TextZoom demonstrate that TLWSR generates distinguishable text images and exceeds the fully supervised baseline TSRN in boosting text recognition accuracywith only 50% HR labels available. Meanwhile, TLWSR can be applied to different super‐resolution backbones and significantly improves their performance. Furthermore, TLWSR shows good generalization capability to low‐quality images on scene text recognition benchmarks, which verifies the effectiveness of this framework. To the authors' knowledge, this is the first work exploring the problem of STISR in weakly supervised scenarios.
Yu Zhu 0005, Chuantao Fang
IET Image Process.2
2023 TD-Net: Trans-Deformer network for automatic pancreas segmentation
Shunbo Dai, Yu Zhu 0005, Xiaoben Jiang, Fuli Yu, Jiajun Lin
Neurocomputing2
2023 A multimodal transformer to fuse images and metadata for skin disease classification
Gan Cai, Yu Zhu 0005, Xiaoben Jiang, Jiongyao Ye
Vis. Comput.2
2022 RAOD: refined oriented detector with augmented feature in remote sensing images object detection
Yu Zhu 0005, Chuantao Fang, Nan Wang 0003, Jiajun Lin
Appl. Intell.2
2022 MA-Net: Mutex attention network for COVID-19 diagnosis on CT images
Bingbing Zheng, Yu Zhu 0005, Yanmei Shao, Tao Xu 0035
Appl. Intell.2
2022 FERGCN: facial expression recognition based on graph convolution network
Yu Zhu 0005, Bingbing Zheng, Xiaoben Jiang, Jiajun Lin
Mach. Vis. Appl.2
2022 Cross-Domain Attention and Center Loss for Sketch Re-Identification
abstract
Matching all RGB photos of the target person in the gallery database with the full-body sketch image drawn by the professional is defined as Sketch re-identification (Sketch Re-id). The big gap between the sketch domain and RGB domain makes Sketch Re-id challenging. This paper addresses the problem by proposing a new framework to obtain domain-invariant features, which uses CNN as the backbone. To make the model focus more on the regions related to the sketch image in the RGB photo, we propose a novel cross-domain attention (CDA) mechanism. It uses different ways of splitting feature maps in its two branches and calculates the relationship between different parts in the sketch images and RGB photos. Moreover, we designed the cross-domain center loss (CDC), which breaks through the limitations that datasets need to be in the same domain in the traditional center loss. It effectively reduces the gap between two domains and makes the features with the same ID closer. The experiment is performed on the Sketch Re-id dataset. Each person has one sketch image and two RGB photos. To evaluate the generalization, we also experimented on two popular sketch-photo face datasets. The result in the Sketch Re-id dataset shows the model performs 3.7% higher than the previous methods. And the result in the CUHK student dataset performs 0.38% higher than the state-of-the-art methods.
Fengyao Zhu, Yu Zhu 0005, Xiaoben Jiang, Jiongyao Ye
IEEE Trans. Inf. Forensics Secur.2
2021 A multi-class COVID-19 segmentation network with pyramid attention and edge loss in CT images
abstract
At the end of 2019, a novel coronavirus COVID-19 broke out. Due to its high contagiousness, more than 74 million people have been infected worldwide. Automatic segmentation of the COVID-19 lesion area in CT images is an effective auxiliary medical technology which can quantitatively diagnose and judge the severity of the disease. In this paper, a multi-class COVID-19 CT image segmentation network is proposed, which includes a pyramid attention module to extract multi-scale contextual attention information, and a residual convolution module to improve the discriminative ability of the network. A wavelet edge loss function is also proposed to extract edge features of the lesion area to improve the segmentation accuracy. For the experiment, a dataset of 4369 CT slices is constructed, including three symptoms: ground glass opacities, interstitial infiltrates, and lung consolidation. The dice similarity coefficients of three symptoms of the model achieve 0.7704, 0.7900, 0.8241 respectively. The performance of the proposed network on public dataset COVID-SemiSeg is also evaluated. The results demonstrate that this model outperforms other state-of-the-art methods and can be a powerful tool to assist in the diagnosis of positive infection cases, and promote the development of intelligent technology in the medical field.
Fuli Yu, Yu Zhu 0005, Xiangxiang Qin, Tao Xu 0035
IET Image Process.2
2021 TSRGAN: Real-world text image super-resolution based on adversarial learning and triplet attention
Chuantao Fang, Yu Zhu 0005, Xiaofeng Ling
Neurocomputing2
2021 Images denoising for COVID-19 chest X-ray based on multi-resolution parallel residual CNN
Xiaoben Jiang, Yu Zhu 0005, Bingbing Zheng
Mach. Vis. Appl.2
2020 Pulmonary nodule risk classification in adenocarcinoma from CT images using deep CNN with scale transfer module
abstract
Pulmonary nodules risk classification in adenocarcinoma is essential for early detection of lung cancer and clinical treatment decision. Improving the level of early diagnosis and the identification of small lung adenocarcinoma has been always an important topic for imaging studies. In this study, the authors propose a deep convolutional neural network (CNN) with scale‐transfer module (STM) and incorporate multi‐feature fusion operation, named STM‐Net. This network can amplify small targets and adapt to different resolution images. The evaluation data were obtained from the computed tomography (CT) database provided by Zhongshan Hospital Fudan University (ZSDB). All data have a pathological label and their lung adenocarcinomas risk are classified into four categories: atypical adenomatous hyperplasia, adenocarcinoma in situ, minimally invasive adenocarcinoma, and invasive adenocarcinoma. The authors’ deep learning network STM‐Net was trained and tested for the risk stage prediction. The accuracy and the average area under the receiver operating characteristic curve achieved by their method are 95.455% and 0.987 for the ZSDB dataset. The experimental results show that STM‐Net largely boosts classification accuracy on the pulmonary nodules classification compared with state‐of‐the‐art approaches. The proposed method will be an effective auxiliary to help physicians diagnosis pulmonary nodules risk classification in adenocarcinoma in early‐stage.
Yu Zhu 0005, Wang Huan Gu, Bingbing Zheng, Chunxue Bai, Hongcheng Shi, Jie Hu 0015, Shaohua Lu, Weibin Shi, Ningfang Wang
IET Image Process.3
2020 Learning multi-level domain invariant features for sketch re-identification
Shaojun Gui, Yu Zhu 0005, Xiangxiang Qin, Xiaofeng Ling
Neurocomputing2
2020 3D multi-scale discriminative network with multi-directional edge loss for prostate zonal segmentation in bi-parametric MR images
Xiangxiang Qin, Yu Zhu 0005, Shaojun Gui, Bingbing Zheng, Peijun Wang
Neurocomputing2
2019 Integrating operation scheduling and binding for functional unit power-gating in high-level synthesis
Nan Wang 0003, Song Chen 0001, Zhiyuan Ma 0001, Xiaofeng Ling, Yu Zhu 0005
Integr.5
2018 Power-gating-aware scheduling with effective hardware resources optimization
Nan Wang 0003, Song Chen 0001, Zhiyuan Ma 0001, Xiaofeng Ling, Yu Zhu 0005
Integr.6
2016 Image analysis by generalized Chebyshev-Fourier and generalized pseudo-Jacobi-Fourier moments
Hongqing Zhu, Zhiguo Gui, Yu Zhu 0005
Pattern Recognit.4
2016 Leakage-Power-Aware Scheduling With Dual-Threshold Voltage Design
abstract
The exponential increase in leakage power and the substantial power-saving opportunities provided by scheduling have made dual-threshold voltage (dual-Vth) an attractive choice for low-leakage-power designs. In this paper, we work under the assumption that functional units (FUs) are allocated after scheduling, and fully explore the solution space of scheduling with dual-Vthoperations to optimize the leakage power of the FUs. First, a binding conflict graph (BCG)-based scheduling method is presented to minimize the number of FUs. Second, the BCG-based method is extended to allow scheduling with dual-Vthoperation targeting the minimization of leakage power. In timing-constrained scheduling, each operation in the data flow is initialized with low-Vth. Then, starting from an operation schedule with the timing constraint satisfied, we scale the sets of low-Vthoperations in the off-critical paths with high-Vthso as to reduce the number of low-VthFUs without increasing the total delay. Finally, a scheduling method for minimizing the leakage power under both timing and resource constraints is presented. The results of benchmark tests show that the proposed algorithms can reduce the leakage power reported in previous works by 10.2% while maintaining high circuit performance.
Nan Wang 0003, Cong Hao, Song Chen 0001, Takeshi Yoshimura, Yu Zhu 0005
IEEE Trans. Very Large Scale Integr. Syst.6