EDBT 2026 Demo / reviewers in the wild / expert
Yicong Zhang
dblp:182/8992
· DBLP profile ↗
9ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CCacheSim: A Circuit-Architecture Cross-Level Simulation Framework for SRAM-Based in-Cache Computing System EvaluationabstractSRAM-based Compute-In-Memory (CIM) circuits have demonstrated significant performance and energy efficiency advantages. Although numerous frameworks or tools have emerged for simulating CIM-based systems, most frameworks are tailored for specific DNN accelerators and rarely consider SRAM-CIM solutions in general processor systems because it is difficult to establish an effective mechanism to build the CIM data path to integrate the SRAM-CIM module that is tightly coupled with the cache hierarchy into the system. To address this problem, we propose a circuit-architecture cross-level simulation framework named CCacheSim for in-cache computing system. CCacheSim integrates the simulation of SRAM-CIM circuit timing and energy consumption characteristics, providing circuit-level accuracy evaluation support for in-cache computing system simulations. For the circuit level, the SRAM-CIM model can automatically generate a circuit-level netlist and conduct accurate simulation through corresponding configurations, thus balancing accuracy and agility for early design exploration. For the architectural level, to efficiently support the portable integration of SRAM-CIM module to in-cache computing system, a configurable hardware programming interface is implemented within the cache model to manage the interaction of the control stream between processor and cache for CIM tasks. Moreover, a request queue based access mechanism is proposed to ensure the completeness of the operands required by CIM tasks. To validate the proposed framework, CCacheSim is implemented to simulate varying configurations of in-cache computing systems and SRAM-CIM modules. The results prove that CCacheSim can conduct accurate performance and energy consumption evaluation for given processor architecture with given CIM module. CCacheSim can support flexible and effective design space exploration for in-cache computing system. Baiqing Zhong, Mingyu Wang 0003, Yicong Zhang, Zhiyi Yu |
ICCD | 3 |
| 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsabstractGeneral-purpose graphics processing unit (GPGPU), widely recognized as an exceptional computing platform for de-ploying emerging parallel applications, requires strict adherence to atomicity and memory consistency models for shared variable synchronization. This is crucial to ensure deterministic execution and leverage the performance advantages of the GPGPU single-instruction -multiple-threads architecture. However, the escalating demand for shared variable updates across thread blocks, notably in applications like deep neural networks and graph analysis, significantly exacerbates the serialization overhead of atomic operations due to the von Neumann bottleneck. Additionally, the overhead introduced by memory fences supporting the memory consistency model further complicates this fine-grained synchronization requirement. To address these challenges, this paper proposes Atomic Cache, facilitating an In-Cache computing hardware-software co-design for GPGPUs. At the software level, we propose relaxed memory consistency based on non-ordering commutativity to alleviate the execution of in-cache atomic operations, thereby mitigating the performance overhead of memory fences. At the hardware level, we present the In-Situ Store Atomic Cache Macro, which empowers the Atomic Cache to efficiently execute atomic logic and arithmetic operations within the cache array. This innovation alleviates the von Neumann bottleneck associated with serialized execution of atomic operations. The experimental evaluation results demonstrate that the Atomic Cache can save more than 60% of memory access energy while incurring only 9.42% chip area overhead. Furthermore, it not only delivers an average speedup ratio of 2.59 × and an IPC performance improvement of 1.48× for RISC-V GPGPUs, but also achieves an average speedup ratio of 1.31 × and an IPC performance improvement of 39.92% when compared to state-of-the-art designs employing local atomic buffers. Yicong Zhang, Mingyu Wang 0003, Wangguang Wang, Yangzhan Mai, Haiqiu Huang, Zhiyi Yu |
MICRO | 1 |
| 2023 | LWSDP: Locality-Aware Warp Scheduling and Dynamic Data Prefetching Co-design in the Per-SM Private Cache of GPGPUsabstractGeneral Purpose Graphics Processing Units (GPG-PUs) employ frequent context switching to mask the long-latency of memory operations. However, GPGPUs still suffer from stagnation due to the incomplete overlapping of memory operations. To alleviate this stagnation and enhance Memory-Level Parallelism (MLP), it is crucial to overlap and minimize memory operations. This paper conducts a comprehensive analysis of data locality in GPGPUs and proposes an approach called Locality-Aware Warp Scheduling and Dynamic Data Prefetching (LWSDP) Co-design in the Per-SM Private Cache of GPGPUs, which effectively utilizes data locality to improve MLP. In addition to employing a coordinated scheduler and dynamic data prefetching, we incorporate Prefetching Requests Admitted Cache Access Re-execution (PRA-CAR) to mitigate the adverse impact of excessive prefetching memory requests on memory saturation. Experimental results demonstrate that LWSDP achieves an average 33.02% performance improvement and an average 28.16% miss rate reduction compared to the previous schedulers on data locality-sensitive kernels. Wangguang Wang, Mingyu Wang 0003, Yicong Zhang, Yukun Wei, Zhiyi Yu |
ICPADS | 3 |
| 2023 | A Scalable Deadlock-Free Static Routing Algorithm for Chiplet-Based SystemsabstractThe utilization of the Chiplet methodology can accelerate VLSI system development and provide better flexibility. Building interconnection networks across multiple Chiplets and ensuring high-performance deadlock-free routing in systems with diverse irregular topologies is a challenging task.To avoid the reordering introduced by adaptive routing algorithms, a scalable static deadlock-free routing algorithm specifically designed for Chiplets is proposed. This approach capitalizes on static routing, a feature that ensures consistent message order due to fixed paths, fundamentally averting reordering issues. By adaptively configuring the state of the local router, it is possible to proactively initiate turns or exit detour loops, thereby effectively preventing deadlocks. Furthermore, by employing state configuration, the routing remains deadlock-free even when there are variations in the number of Chiplets, scale, or internal topology. This showcases the system’s scalability.Due to limited wiring resources, we chose classical routing algorithms up*/down* to compare, and the results showed a significant latency advantages, with a saturation injection rate approximately 1.5 to 2 times higher. Mingyu Wang 0003, Yicong Zhang, Tao Lu 0012, Zhiyi Yu |
ICPADS | 3 |
| 2023 | TensorCache: Reconstructing Memory Architecture With SRAM-Based In-Cache Computing for Efficient Tensor Computations in GPGPUsabstractGeneral purpose graphics processing units (GPGPUs) have emerged as a convincing and pivotal computing platform for deep learning applications. However, the fundamental tensor computations for neural networks on GPGPUs are still restricted by the von Neumann bottleneck. The memory bandwidth and energy consumption of moving a large amount of neural network data between the memory hierarchy and computational units of GPGPUs dominate the overall computational cost. To address these challenges, this article proposes TensorCache to reconstruct memory architecture with static random-access memory (SRAM)-based In-Cache Computing for efficient tensor computations in GPGPUs. It provides an innovative digital SRAM processing-in-memory (PIM) solution by transforming the cache array into large-scale PIM units, effectively mitigating the significant performance and energy consumption losses caused by data movement. To enable efficient hardware-software co-design for TensorCache, a decoupled architecture-based SRAM-PIM macro (SPM) is introduced at the hardware level, supporting in-memory bit-parallel comparison (IMBC) and near-memory radix-4 booth encoder (NRBE) for efficient mixed-precision floating-point (FP) tensor computations. At the software level, a programming model leveraging the GPGPU’s flexible programmability is proposed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that TensorCache achieves up to$38.59\times $speedup and$16.26\times $throughput enhancement compared to GPU CUDA Cores. Furthermore, it attains an acceleration of up to$1.78\times $and$3.87\times $throughput improvement compared to GPU Tensor Cores, while saving power consumption in tensor computations by over 90% with a mere 21% chip area overhead. Yicong Zhang, Mingyu Wang 0003, Yangzhan Mai, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Modulation of the Transmission Spectra of the Double-Ring Structure by Surface Plasmonic PolaritonsabstractThis paper proposes a new structural design to excite surface plasmonic polaritons to enhance the double‐ring interference structure. The double‐ring structure was etched into a thin film to form fundamental interference patterns, and periodic concentric‐ring grooves were employed to gather energy from the surrounding regions through the excitation of surface plasmonic polaritons. Accordingly, the energy of the incident light can be concentrated at the center. The surface plasmon modulates the interference pattern and the transmission spectra. The transmission peak position and its intensity can be tuned by changing the alignment of the grooves. The proposed structure can be applied for designing plasmonic devices as useful components of the plasmonic toolbox. Senfeng Lai, Yanpei Guo, Guiyang Liu, Chun Shan, Lixin Huang, Yicong Zhang, Yanghui Wu, Wenhua Gu, Wen Wu 0005 |
Wirel. Commun. Mob. Comput. | 6 |
| 2020 | An efficient object detection framework with modified dense connections for small objects optimizationsabstractObject detection frameworks for small objects are increasingly demanded in some specific fields such as high-speed object tracking and remote sensing image recognition. In this paper, we propose an efficient object detection framework with modified dense connections for small objects. In order to improve both the detection accuracy and speed for small objects, the proposed framework constructs a convolutional neural network by using modified dense and residual cross-layer connections between multi-scale convolutional layers to extract deep features effectively. Based on the modified dense structure, a hybrid-scale feature fusion method is proposed to concatenate the multi-channel high-dimensional features and performs cross-entropy calculation and regression prediction. By using this method, this framework not only improves the detection accuracy for small objects significantly, but also improves the overall detection accuracy and optimizes the network parameters to reduce the detection time greatly. The experimental results show that the proposed framework achieves 90.6% mAP for small objects on a public ship dataset which is 25.2% more than SSD-VGGNet. Due to the detection efficiency for small objects, it improves the overall detection accuracy and detection speed by 9% and 40% respectively while about 70% network parameters are reduced. Yicong Zhang, Zhaolin Li |
CF | 1 |
| 2020 | Atomic Predicates-Based Data Plane Properties Verification in Software Defined Networking Using SparkabstractSoftware-Defined Networking (SDN) is an innovational network architecture which gives network administrators the ability to directly control the whole network by programming on a centralized controller. Due to network complexity, networks are unlikely to be bug-free. The ability to verify data plane properties will make network management easier for network administrators in SDN. In this paper, we present a novel atomic predicates based data plane properties verification method for SDN using Spark which is a big data processing framework. First, we verify packet reachability which is a fundamental data plane property. Then, we verify other data plane properties such as loop-freedom and nonexistence of black holes. In addition, the proposed method can detect a security threat existing in SDN called firewall bypass threat with packet reachability verification. By adopting atomic predicates, we achieve less computational and storage overhead. We implement the methods and study the performance. The results of experiments show that we can efficiently and accurately detect loops, black holes and firewall bypass threats. Yicong Zhang, Jie Li 0002, Shigetomo Kimura, Wei Zhao 0001, Sajal K. Das 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | QoS Analysis and Optimization of Business ProcessesabstractWith the development of modern society, business processes are widely used, and they play an important role in peo-ple's daily life. In government agencies, there are many business processes employed to deal with different kinds of transactions. Such as one of the medical insurance refund processes is used to claim for maternity insurance. In order to analyze, evaluate and optimize the business process in different situations, this paper develops to use the queueing network theory to establish the corresponding business process model. Then apply this model to analyze the performance of the business process by using four performance metrics i.e., average queue length, average job sojourn time, job arrival intensity and throughput. Finally, this paper optimizes the business process based on three situations. It has been proven that the optimization methods are scientific and the business process after the optimization is more efficient. Yicong Zhang, Yongqing Zheng, Shidong Zhang |
CSCWD | 1 |