Yuzhong Zhang

dblp:59/5685 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 14 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accuracy improvement of temperature measurement using a wide-spectrum visible camera based on deep learning network denoising
Shuangbao Shu, Yufeng Fu, Xiaoyue Hu, Yuzhong Zhang, Junjie Niu
Eng. Appl. Artif. Intell.5
2026 HCRT: Hybrid network with correlation-aware region transformer for breast tumor segmentation in DCE-MRI
Lei Zheng 0017, Yuzhong Zhang, Tao Zhou 0002, Lei Zhou 0003, Dinggang Shen
Pattern Recognit.2
2025 Enhancing interpretability in video-based personality trait recognition using SHAP analysis
Yang Liu 0025, Wenyi Zhu, Linyu Dong, Yuzhong Zhang
Multim. Syst.4
2025 Prototype Learning Guided Hybrid Network for Breast Tumor Segmentation in DCE-MRI
abstract
Automated breast tumor segmentation on the basis of dynamic contrast-enhancement magnetic resonance imaging (DCE-MRI) has shown great promise in clinical practice, particularly for identifying the presence of breast disease. However, accurate segmentation of breast tumor is a challenging task, often necessitating the development of complex networks. To strike an optimal trade-off between computational costs and segmentation performance, we propose a hybrid network via the combination of convolution neural network (CNN) and transformer layers. Specifically, the hybrid network consists of a encoder-decoder architecture by stacking convolution and deconvolution layers. Effective 3D transformer layers are then implemented after the encoder subnetworks, to capture global dependencies between the bottleneck features. To improve the efficiency of hybrid network, two parallel encoder subnetworks are designed for the decoder and the transformer layers, respectively. To further enhance the discriminative capability of hybrid network, a prototype learning guided prediction module is proposed, where the category-specified prototypical features are calculated through online clustering. All learned prototypical features are finally combined with the features from decoder for tumor mask prediction. The experimental results on private and public DCE-MRI datasets demonstrate that the proposed hybrid network achieves superior performance than the state-of-the-art (SOTA) methods, while maintaining balance between segmentation accuracy and computation cost. Moreover, we demonstrate that automatically generated tumor masks can be effectively applied to identify HER2-positive subtype from HER2-negative subtype with the similar accuracy to the analysis based on manual tumor segmentation. The source code is available at https://github.com/ZhouL-lab/PLHN.
Lei Zhou 0003, Yuzhong Zhang, Xuejun Qian, Chen Gong 0002, Zhongxiang Ding, Zhenhui Li, Zaiyi Liu, Dinggang Shen
IEEE Trans. Medical Imaging2
2024 CINDA: Don't Ignore Instructions When Cloning Memory Access Behavior
abstract
Existing workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error.
Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen
CCGrid7
2024 BlueJay: A Platform to Quantifying the Impact of Memory Latency on Datacenter Application Performance
abstract
Understanding the impact of memory latency on datacenter application performance can provide decision support to memory subsystem designers. Currently, various methods are available to quantify this impact, including cycle-accurate simulators, memory-level parallelism models, and software delay injection techniques. However, these methods suffer from several limitations, such as slow simulation speed, inaccuracy, and insufficient compatibility that requires application modification.This paper proposes BlueJay, a novel platform to quantify the impact of memory latency on the end-to-end performance of datacenter applications, avoiding slow simulation and providing high accuracy and compatibility. The key idea of BlueJay is to control the consumed memory bandwidth and read/write ratio, thereby manipulating memory latency to achieve quantification. Experiment shows that BlueJay provides accurate quantification with an average error of 3.04%. In addition, we built regression models for five applications deployed at scale in Alibaba data centers. The results reveal that a 10 ns increase in memory latency results in a performance decrease of 2.61%-3.31% for enterprise Java applications and databases, while the elastic block storage service experiences a more modest performance decrease of 0.73%-0.92%.
Jingchang Qin, Yiquan Chen, Shishun Cai, Wenhai Lin, Jiexiong Xu, Zhen Jin 0008, Lifa Cao, Yuzhong Zhang, Wenzhi Chen
CCGrid9
2024 Nash Equilibrium and Price of Anarchy for Scheduling Games Based on a Mixed Coordination Mechanism
Wei Sheng, Yuzhong Zhang
COCOON (1)5
2024 PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region Sizes
abstract
Hardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%).
Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2023 JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application Services
abstract
Many Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead.
Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen
CLUSTER6
2023 Development of a cross-scale weighted feature fusion network for hot-rolled steel surface defect detection
Yuzhong Zhang, Zhaoming Li, Shuangbao Shu, Xianli Lang, Tengda Zhang, Jingtao Dong
Eng. Appl. Artif. Intell.1
2023 Efficiency and inefficiency of Nash equilibrium for scheduling games on batching-machines with activation cost
Jiguo Yu, Yuzhong Zhang, Donglei Du
Theor. Comput. Sci.3
2021 On Workload-Aware DRAM Failure Prediction in Large-Scale Data Centers
abstract
DRAM failures are one of the major hardware threats to the reliability of large-scale data centers since the uncorrectable errors in DRAMs may cause servers to shut down. Existing works try to solve this problem by predicting DRAM failures in advance with Machine Learning models. In these works, correctable errors (CEs) are generally deemed as the most important feature. The major reason behind CEs' emergence is the accumulated stress caused by intensive workloads. Moreover, defective DRAMs will not manifest themselves as system errors until the defective cells are accessed by some specific workloads. Therefore, the running workloads on a server are also important for DRAM failure prediction. In this paper, we focus on the impact of workloads on DRAM failures. We design the workload features from both macroscopical and microscopical aspects, i.e. node-level performance metrics and cell-level DRAM access pattern, respectively. Furthermore, we propose Hierarchical DRAM Error Code (HiDEC) to represent the DRAM access pattern. We leverage several Decision Tree-based models for DRAM failure prediction to highlight the generality of our designed features. Experiments are carried out based on the dataset collected from a real-world commercial data center. The results show that both macroscopic and microscopic features can bring significant improvements to the prediction performance.
Xingyi Wang, Yu Li 0007, Yiquan Chen, Yin Du, Yuzhong Zhang, Pinan Chen, Wenjun Song, Qiang Xu 0001, Li Jiang 0002
VTS7
2018 An implicit degree condition for k-connected 2-heavy graphs to be hamiltonian
Junqing Cai, Hao Li 0002, Yuzhong Zhang
Inf. Process. Lett.3
2016 Fan-type implicit-heavy subgraphs for hamiltonicity of implicit claw-heavy graphs
Junqing Cai, Yuzhong Zhang
Inf. Process. Lett.2
2016 Inefficiency analysis of the scheduling game on limited identical machines with activation costs
Yuzhong Zhang, Qingguo Bai
Inf. Process. Lett.2
2015 Scheduling games on uniform machines with activation cost
Yuzhong Zhang, Qingguo Bai
Theor. Comput. Sci.3
2012 Scheduling of deteriorating jobs with release dates to minimize the maximum lateness
Cuixia Miao, Yuzhong Zhang, Cuilian Wu
Theor. Comput. Sci.2
2011 Bounded parallel-batch scheduling on single and multi machines for deteriorating jobs
Cuixia Miao, Yuzhong Zhang
Inf. Process. Lett.2
2010 Bounded Parallel-Batch Scheduling on Unrelated Parallel Machines
Cuixia Miao, Yuzhong Zhang, Chengfei Wang
AAIM2
2009 Approximation Algorithm for Minimizing the Weighted Number of Tardy Jobs on a Batch Machine
Jianfeng Ren, Yuzhong Zhang, Xianzhao Zhang, Guo Sun
COCOA2
2009 Scheduling with Rejection to Minimize the Makespan
Yuzhong Zhang, Jianfeng Ren, Chengfei Wang
COCOA1
2007 An Asymptotic PTAS for Batch Scheduling with Nonidentical Job Sizes to Minimize Makespan
Yuzhong Zhang
COCOA1
2006 On Several Scheduling Problems with Rejection or Discretely Compressible Processing Times
Yuzhong Zhang, Shoupeng Liu
TAMC3
2005 A PTAS for Scheduling on Agreeable Unrelated Parallel Batch Processing Machines with Dynamic Job Arrivals
Yuzhong Zhang, Qingguo Bai
AAIM1
2004 Minimizing Mean Completion Time in a Batch Processing System
Xiaotie Deng, Haodi Feng, Pixing Zhang, Yuzhong Zhang, Hong Zhu 0004
Algorithmica4
1999 Minimizing Mean Response Time in Batch Processing System
Xiaotie Deng, Yuzhong Zhang
COCOON2
1999 Approximation Algorithms in Batch Processing
Xiaotie Deng, Chung Keung Poon, Yuzhong Zhang
ISAAC3