Jizeng Wei

dblp:45/7702 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-3040-6859ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 5 since 2021Security and privacy · 3Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 BASH: A Bandwidth Sharing Mechanism for Subpartition Scheduling in GPU L2 Caches
abstract
Modern GPUs employ a multi-level cache hierarchy where several requests missing in L1 Caches can be redirected to L2 Cache simultaneously. Therefore, the memory space of L2 Cache is divided into multiple partitions. A partition corresponding to a DRAM channel is composed of two subpartitions, each of which featuring a L2 Cache bank with a 32-byte/cycle data port. Given that the NoC datapath between L2 and L1 Caches also has a 32-byte bandwidth. Therefore, a request associating 128-byte data that are obtained from the L2 cache-line requires four cycles for completion, leading to bandwidth fragmentation and low resource utilization when request distribution is uneven. To address this issue, we propose the BASH architecture, which aggregates four subpartition data ports into a logical 128-byte channel and incorporates intra-group round-robin scheduling. This enables full processing of a request data transmission within a single cycle for a subpartition. Without modifying the existing model, our solution significantly improves L2 bandwidth utilization and system performance. Experimental results demonstrate that BASH can increase data port utilization by 15.7% compared to the baseline GPU, with an average performance improvement of 23.1%.
Bingchao Li, Jizeng Wei
HPCC4
2025 3D-DP: A Practical DRL-Based Data Prefetcher with a Dynamic Prefetch Degree and Direction
Jizeng Wei, Shuangsheng Li, Yaogong Yang
ICA3PP (1)2
2025 Solving online resource-constrained scheduling for follow-up observation in astronomy: A reinforcement learning approach
Ce Yu, Chao Sun 0008, Jizeng Wei, Junhan Ju, Shanjiang Tang
Future Gener. Comput. Syst.4
2024 Pseudo-Cache: Extending the Access Scope of Requests with Global Perspective in GPUs
abstract
GPUs are comprised of numerous streaming multiprocessors (SMs) tailored for high performance computing. SMs incorporate private L1 caches to facilitate swift data access for thousands of threads running concurrently inside SMs. Requests that miss in L1 caches are directed to L2 caches that are shared by all SMs to retrieve the desired data, which takes much longer time compared with the delay of accessing L1 cache. Due to the locality among tasks running on SMs, the same data block might be accessed by requests from various SMs, leading to data replication across multiple L1 caches. Notably, data replication is prevalent in applications exhibiting high data locality. In order to capitalize on the data replications among L1 caches, we propose routing requests destined for L2 cache to other L1 caches that are predicted to contain the desired data, extending the access scope of requests. Consequently, we introduce a Pseudo-Cache positioned adjacent to the network on chip on L2 cache side, offering a global perspective of all SMs to manage these predictions effectively. Moreover, we implement the data path for request forwarding in a cost-effective manner, leveraging the existing network structure. Experimental results underscore the efficacy of our approach, showcasing an average 16.3% performance enhancement for applications characterized by substantial data replication on GPUs.
Bingchao Li, Jizeng Wei
HPCC2
2022 REMOC: efficient request managements for on-chip memories of GPUs
abstract
The on-chip memories of GPUs, including the register file, shared memory and L1 cache, can provide high bandwidth and low latency access for the temporary storage of data. The capacity of L1 cache can be increased by using the registers/shared memory that are unassigned to any warps/thread blocks or released after warps/thread blocks are finished as cache-lines. In this paper, we propose two techniques to manage requests for on-chip memories to improve the efficiency of L1 cache on the base of leveraging registers and shared memory as cache-lines. Specifically, we develop a data transferring policy which is triggered when cache-lines are recalled by the first register or shared memory accesses of warps that are newly launched to prevent the data locality from being destroyed. Additionally, we design a parallel issue scheme by exploring the parallel feature of requests of an instruction accessing the register file, shared memory and L1 cache to decrease the processing latency and hence increase the throughput of instructions. The experimental results demonstrate that our approach improves the performance by 15% over prior work.
Bingchao Li, Jizeng Wei
CF2
2019 A Systematic Approach to Horizontal Clustering Analysis on Embedded RSA Implementation
abstract
Side-channel attacks are a threat to cryptographic algorithms running on embedded devices. The exponent blinding is the main countermeasure which resists the classical forms of side-channel attacks. Horizontal side-channel analysis is an effective method to overcome this kind of countermeasures and recover the secret key through the analysis of a single trace. However, it is very difficult to exploit leakage from a single trace in practice, because the segment borders of crucial operation are hard to find and the single execution possibly contains a lot of noise in a real application environment. In this paper, a framework for systematic approach is presented to overcome these drawbacks. Firstly, we propose a method to use the Hilbert-Huang transform for the segment borders detection and the noise reduction. Then, an effective clustering method is used to recover the private key. The proposed horizontal clustering approach is applied to a protected implementation on embedded devices. The results show that the success rate of the key recovery could be up to 98%. Finally, we analyze the mechanism of leakage exploration.
Fanyu Shi, Jizeng Wei, Da-Zhi Sun, Guo Wei 0005
ICPADS2
2019 Highly-parallel hardware implementation of optimal ate pairing over Barreto-Naehrig curves
Jizeng Wei
Integr.3
2019 An Efficient GPU Cache Architecture for Applications with Irregular Memory Access Patterns
abstract
GPUs provide high-bandwidth/low-latency on-chip shared memory and L1 cache to efficiently service a large number of concurrent memory requests. Specifically, concurrent memory requests accessing contiguous memory space are coalesced into warp-wide accesses. To support such large accesses to L1 cache with low latency, the size of L1 cache line is no smaller than that of warp-wide accesses. However, such L1 cache architecture cannot always be efficiently utilized when applications generate many memory requests with irregular access patterns especially due to branch and memory divergences that make requests uncoalesced and small. Furthermore, unlike L1 cache, the shared memory of GPUs is not often used in many applications, which essentially depends on programmers. In this article, we propose Elastic-Cache, which can efficiently support both fine- and coarse-grained L1 cache line management for applications with both regular and irregular memory access patterns to improve the L1 cache efficiency. Specifically, it can store 32- or 64-byte words in non-contiguous memory space to a single 128-byte cache line. Furthermore, it neither requires an extra memory structure nor reduces the capacity of L1 cache for tag storage, since it stores auxiliary tags for fine-grained L1 cache line managements in the shared memory space that is not fully used in many applications. To improve the bandwidth utilization of L1 cache with Elastic-Cache for fine-grained accesses, we further propose Elastic-Plus to issue 32-byte memory requests in parallel, which can reduce the processing latency of memory instructions and improve the throughput of GPUs. Our experiment result shows that Elastic-Cache improves the geometric-mean performance of applications with irregular memory access patterns by 104% without degrading the performance of applications with regular memory access patterns. Elastic-Plus outperforms Elastic-Cache and improves the performance of applications with irregular memory access patterns by 131%.
Bingchao Li, Jizeng Wei, Murali Annavaram, Nam Sung Kim
ACM Trans. Archit. Code Optim.2
2018 Practical chosen-message CPA attack on message blinding exponentiation algorithm and its efficient countermeasure
Wei Guo 0005, Jizeng Wei
World Wide Web3
2015 Analyzing graphics processor unit (GPU) instruction set architectures
abstract
Because of their high throughput and power efficiency, massively parallel architectures like graphics processing units (GPUs) become a popular platform for generous purpose computing. However, there are few studies and analyses on GPU instruction set architectures (ISAs) although it is wellknown that the ISA is a fundamental design issue of all modern processors including GPUs.
Kothiya Mayank, Hongwen Dai, Jizeng Wei, Huiyang Zhou
ISPASS3
2014 Further Research on N-1 Attack against Exponentiation Algorithms
Zhaojing Ding, Wei Guo 0005, Liangjian Su, Jizeng Wei, Haihua Gu
ACISP4
2014 The Micro-architectural Support Countermeasures against the Branch Prediction Analysis Attack
abstract
Recently, a kind of micro-architectural side-channel analysis attacks, Branch Prediction Analysis (BPA), has been demonstrated to be practically feasible on the popular commodity PC platform. This attack extracts the secret information based on monitoring the branch target buffers (BTB). Some cryptography algorithms, such as RSA, ECC are naturally vulnerable to BPA because of the key-centric sequence of conditional branches. BPA attack can successfully steal almost all of the security key bits during one single encryption process in virtue of an elaborately designed and "legitimate" spy-process. Although there are some countermeasures existing in the state-of-art literatures, all of them are software-based methods, which lead to a series of design challenges. This paper proposes an architectural support scheme against the BPA attack comprehensively. A well-customized surveillance table with limited size is appended to record each process in order to dynamically recognize which one is malicious in time. And then a lock-based BTB scheme is utilized to protect the BTB visiting from BPA attack efficiently to ensure the sensitive information not be leaked due to the conditional branches loophole. Experimental results show that the proposed anti-BPA attack scheme not only leverages approximate 8KB area cost to provide strong security protection but also incurs slight performance improvement about 0.12% on average about the benchmarks. Meanwhile, it is transparent on the application level to alleviate the difficulties of the programmers.
Ya Tan, Jizeng Wei, Wei Guo 0005
TrustCom2
2011 An Efficient Stereoscopic Game Conversion System for Embedded Platform
abstract
A low cost embedded auto-stereoscopic game conversion system is presented in this paper. The system is implemented both in software and hardware. The software is run on embedded GPU (E-GPU) to generate the depth map from a traditional 3D game which only has 2D images by extracting and processing the hidden information which is used to build the scene. Then multi-view image is generated by depth image based rendering (DIBR) accelerator which is implemented by hardware. To improve the efficiency, the parallel and pipelined architecture is proposed. The design is verified on FPGA platform. The result shows that the proposed system can convert traditional 3D game to stereoscopic game with low cost and low power. Therefore it is suitable for hand held devices.
Wei Guo 0005, Jizeng Wei
TrustCom4
2011 An optimized TTA-like vertex shader datapath for embedded 3D graphics processing unit
abstract
An alternative VLIW architecture of vertex shader datapath based on transport triggered architecture (TTA) is proposed in details. This architecture can exploit more instruction level parallelism (ILP) than traditional VLIW architecture by the fine-grained data transport. The proposed vertex shader architecture can also provide a simple and user-optimized inter-connection network which can efficiently reduce the complexity of interconnections design. The evaluation results show that the proposed architecture can achieve almost 18% reduction in interconnection number and 1.4 times improvement in code density compared with the multi-threaded expanded VLIW architecture (MT-eVLIW).
Jizeng Wei, Yisong Chang, Wei Guo 0005
VLSI-SoC1