VLDB 2026 Research / reviewers in the wild / expert
Xin Xin 0008
dblp:35/1895-8
· DBLP profile ↗
16ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-0952-2115ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 5 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReliaFHE: Resilient Design for Fully Homomorphic Encryption AcceleratorsabstractThe significant computational complexity of Fully Homomorphic Encryption (FHE) has prompted numerous accelerator designs. However, existing FHE accelerators often implicitly assume that all computations are executed reliably, overlooking the fact that modern FHE schemes can be highly vulnerable to hardware faults: even a single-bit error in a ciphertext can cascade into widespread plaintext corruption. Ruizhi Zhu, Mengxin Zheng, Qian Lou, Xin Xin 0008 |
ASPLOS (2) | 6 |
| 2026 | ASPA: Reassigning DDR5 Parity BandwidthabstractRecent memory advancements, such as DDR5, HBM3, and emerging memory-centric accelerators, primarily focus on increasing bandwidth capacity, yet often omit the significance of effective bandwidth utilization, i.e., bandwidth efficiency. Motivated by the suboptimal channel allocation in DDR5, where parity accounts for 25% bandwidth overhead, we propose ASPA, an efficiency-oriented solution that reallocates parity bandwidth to boost regular data transfer. The objective of ASPA is to enhance bandwidth efficiency without compromising reliability while maintaining low hardware overhead. In particular, we leverage existing CRC (Cyclic Redundancy Check) units in DRAM chips to opportunistically generate second-tier parity for the existing ECC (Error Correction Code) parity. For bulk-sized memory accesses, only the small second-tier parity is transmitted. This reduces the parity bandwidth consumption, allowing data chips to reuse the freed parity bandwidth for data transfer. Our observation indicates that the protection capability is sufficient when the second-tier parity (with a 64-bit size) is used exclusively for error detection. Furthermore, by exploiting underutilized resources in high-performance memory systems, ASPA is implemented with negligible hardware overhead. Qiufeng Li, Yanan Guo 0002, Weidong Cao 0001, Xin Xin 0008 |
HPCA | 5 |
| 2026 | DANMP: Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing ArchitectureabstractMulti-Scale Deformable Attention (MSDAttn) has become a fundamental component in various vision tasks due to its effective multi-scale grid sampling (MSGS). However, its reliance on random sampling results in highly irregular memory access patterns, making it a memory-intensive operation inefficient for GPUs. Near-memory processing (NMP) offers a promising solution for accelerating memory-bound kernels, yet existing NMP-based attention accelerators remain suboptimal for MSDAttn due to incompatible load balancing and data reuse strategies. Specifically, current NMP solutions uniformly distribute processing elements (PEs) across all banks, leading to significant PE underutilization and excessive cross-bank data transfers. Moreover, most rely on locality-based reuse, which fails under MSDAttn’s unpredictable sampling patterns. Huize Li, Qinggang Wang, Bin Gao 0013, Dan Chen 0006, Yu Huang 0013, Xin Xin 0008 |
ICS | 6 |
| 2026 | HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECC
Ruizhi Zhu, Yanan Guo 0002, Huize Li, Weidong Cao 0001, Qian Lou, Xin Xin 0008 |
ISCA | 6 |
| 2025 | SCREME: A Scalable Framework for Resilient Memory DesignabstractThe continuing advancement of memory technology has not only fueled a surge in performance, but also substantially exacerbated reliability challenges. Traditional solutions have primarily focused on improving the efficiency of protection schemes, i.e., Error Correction Codes (ECC), under the assumption that allocating additional memory space for ECC parity is always costly and therefore unsustainable as parity scales. We break the stereotype by proposing an orthogonal approach that provides additional, cost-effective memory space for resilient memory design. In particular, we recognize that ECC chips (used for parity storage) do not necessarily require the same performance level as regular data chips. This offers two-fold benefits: First, the bandwidth originally provisioned for a regularperformance ECC chip can instead be used to accommodate multiple low-performance chips. Second, the cost of ECC chips can be effectively reduced, as lower performance often correlates with lower expense. In addition, we observe that server-class memory chips are often provisioned with ample, yet underutilized I/O resources. This suggests an opportunity to repurpose these resources for flexible on-DIMM interconnections. Based on the above two insights, we finally propose SCREME, a scalable memory framework that leverages cost-effective yet slower chips - byproducts of rapid technology evolution - to meet the growing reliability demands driven by this evolution. Mimi Xie, Yanan Guo 0002, Huize Li, Xin Xin 0008 |
PACT | 5 |
| 2025 | PIM-SUM: Fast and Reliable In-Memory Summation for Recommendation SystemsabstractEmbedding aggregation in large-scale recommendation systems creates severe memory bandwidth bottlenecks, as each query sums many high-dimensional vectors. Bitwise-operation-based PIM can exploit subarray bandwidth, but traditional designs struggle with summation because long carry propagation limits parallelism. We propose PIM-SUM, an in-DRAM summation primitive for Sparse Length Sum (SLS) in recommendation workloads. PIM-SUM reformulates vector summation as column-wise accumulation via popcount, truncating carry propagation and avoiding redundant in-DRAM computation. This enables highthroughput summation using native DRAM bitwise primitives. PIM-SUM integrates a Reed-Solomon-inspired error correction at the DRAM row level. Its linearity supports parity propagation, enabling integrity checking and multi-bit error correction with low overhead. On DLRM workloads, PIM-SUM achieves up to$5.14 \times$speedup, logarithmic I/O reduction, and large energy savings. It also reduces silent data corruption by$1778 \times$and improves detection by over$170 \times$. These results show PIM-SUM is a scalable, fault-tolerant, and energy-efficient solution for memory-bound inference at data center scale. Ruizhi Zhu, Huize Li, Di Wu 0016, Xin Xin 0008 |
ICCD | 5 |
| 2025 | LightML: A Photonic Accelerator for Efficient General Purpose Machine LearningabstractThe rapid integration of AI technologies into everyday life across sectors such as healthcare, autonomous driving, and smart home applications requires extensive computational resources, placing strain on server infrastructure and incurring significant costs.We present LightML, the first system-level photonic crossbar design, optimized for high-performance machine learning applications.This work provides the first complete memory and buffer architecture carefully designed to support the high-speed photonic crossbar, achieving over 80% utilization.LightML also introduces solutions for key ML functions, including large-scale matrix multiplication (MMM), element-wise operations, non-linear functions, and convolutional layers.Delivering 325 TOP/s at only 3 watts, LightML offers significant improvements in speed and power efficiency, making it ideal for both edge devices and dense data center workloads. Sadra Rahimi Kari, Xin Xin 0008, Nathan Youngblood, Youtao Zhang, Jun Yang 0002 |
ISCA | 3 |
| 2024 | Addition is Most You Need: Efficient Floating-Point SRAM Compute-in-Memory by Harnessing Mantissa AdditionabstractThe compute-in-memory (CIM) paradigm holds great promise to efficiently accelerate machine learning workloads. Among memory devices, static random-access memory (SRAM) stands out as a practical choice for its exceptional reliability in the digital domain and excellent scalability. Recently, there has been a growing interest in accelerating floating-point (FP) deep neural networks (DNNs) with SRAM CIM due to their critical importance in DNN training and high-accurate inference. This paper proposes an energy-efficient SRAM CIM macro for FP DNNs. To achieve the design, we identify a lightweight approach that decomposes conventional FP mantissa multiplication into two parts: mantissa sub-addition (sub-ADD) and mantissa sub-multiplication (sub-MUL). Our study shows that while mantissa sub-MUL is compute-intensive, it only contributes to the minority of FP products, whereas mantissa sub-ADD, although compute-light, accounts for the majority of FP products. Recognizing "Addition is Most You Need", we develop a novel hybrid-domain SRAM CIM macro to accurately handle mantissa sub-ADD in the digital domain while improving the energy efficiency of mantissa sub-MUL using analog computing. Experiments with the MLPerf benchmark show its remarkable improvement in energy efficiency on average by 3×~ 3.6× (2.5×~3.1×) in inference (training) compared to a fully digital baseline without any accuracy loss, showcasing its great potential for FP DNN acceleration. Weidong Cao 0001, Xin Xin 0008, Xuan Zhang 0001 |
DAC | 3 |
| 2023 | Uncore Encore: Covert Channels Exploiting Uncore Frequency ScalingabstractModern processors dynamically adjust clock frequencies and voltages to reduce energy consumption. Recent Intel processors separate the uncore frequency from the core frequency, using Uncore Frequency Scaling (UFS) to adapt the uncore frequency to various workloads. While UFS improves power efficiency, it also introduces security vulnerabilities. In this paper, we study the feasibility of covert channels exploiting UFS. First, we conduct a series of experiments to understand the details of UFS, such as the factors that can cause uncore frequency variations. Then, based on the results, we build the first UFS-based covert channel, UF-variation, which works both across-cores and across-processors. Finally, we analyze the robustness of UF-variation under known defense mechanisms against uncore covert channels, and show that UF-variation remains functional even with those defenses in place. Yanan Guo 0002, Dingyuan Cao 0002, Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
MICRO | 3 |
| 2022 | Architecting DDR5 DRAM caches for non-volatile memory systemsabstractWith the release of Intel's Optane DIMM, Non-Volatile Memories (NVMs) are emerging as viable alternatives to DRAM memories because of the advantage of higher capacity. However, the higher latency and lower bandwidth of Optane prevent it from outright replacing DRAM. A prevailing strategy is to employ existing DRAM as a data cache for Optane, thereby achieving overall benefit in capacity, bandwidth, and latency. Xin Xin 0008, Wanyi Zhu |
DAC | 1 |
| 2022 | Leaky Way: A Conflict-Based Cache Covert Channel Bypassing Set AssociativityabstractModern $\times$86 processors feature many prefetch instructions that developers can use to enhance performance. However, with some prefetch instructions, users can more directly manipulate cache states which may result in powerful cache covert channel and side channel attacks. In this work, we reverse-engineer the detailed cache behavior of PREFETCHNTA on various Intel processors. Based on the results, we first propose a new conflict-based cache covert channel named NTP+NTP. Prior conflict-based channels often require priming the cache set in order to cause cache conflicts. In contrast, in NTP+NTP, the data of the sender and receiver can compete for one specific way in the cache set, achieving cache conflicts without cache set priming for the first time. As a result, NTP+NTP has higher bandwidth than prior conflict-based channels such as Prime+Probe. The channel capacity of NTP+NTP is 302 KB/s. Second, we found that PREFETCHNTA can also be used to boost the performance of existing side channel attacks that utilize cache replacement states, making those attacks much more efficient than before. Yanan Guo 0002, Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
MICRO | 2 |
| 2021 | ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory based SSDsabstractProcessing-in-memory (PIM) and in-storage-computing (ISC) architectures have been constructed to implement computation inside memory and near storage, respectively. While effectively mitigating the overhead of data movement from memory and storage to the processor, due to the limited bandwidth of existing systems, these architectures still suffer from the large data movement overhead between storage and memory, in particular, if the amount of required data is large. It has become a major constraint for further improving the computation efficiency in PIM and ISC architectures. Congming Gao, Xin Xin 0008, Youyou Lu, Youtao Zhang, Jun Yang 0002, Jiwu Shu |
MICRO | 2 |
| 2021 | SAM: Accelerating Strided Memory AccessesabstractStrided memory accesses are an important type of operations for In-Memory Databases (IMDB) applications. Strided memory accesses often demand data at word granularity with fixed strides. Hence, they tend to produce sub-optimal performance on DRAM memory (the de facto standard memory in modern computer systems) that accesses data at cacheline granularity. Recently proposed optimizations either introduce significant reliability degradation or are limited to non-volatile crossbar memory structures. Xin Xin 0008, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
MICRO | 1 |
| 2020 | Reducing DRAM Access Latency via Helper RowsabstractThe DRAM technology advancement has seen success in memory density and throughput improvement, but less in access latency reduction. This is mainly due to the intrinsic limitation of capacitance based bit store and access mechanism. The reduction of access latency has been well explored in literature. However, the recently proposed DRAM techniques, such as RowClone and Half-DRAM, offer new opportunities to further optimise the access latency.In this paper, we propose an efficient access strategy to improve the performance of DRAM by optionally discarding the restore. When activating a new row, our technique makes a copy of the row leveraging the RowClone method. Next time when accessing the same row, the cloned row is opened for sensing but it is not restored as the data is preserved in the original row. To improve the efficiency of our proposed strategy, we further exploit three schemes to minimize the copy overhead and increase the reuse of the cloned row. Experimental results show that our proposed strategy can achieve 11% performance improvement on average. Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
DAC | 1 |
| 2020 | ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMabstractRecently proposed DRAM based memory-centric architectures have demonstrated their great potentials in addressing the memory wall challenge of modern computing systems. Such architectures exploit charge sharing of multiple rows to enable in-memory bitwise operations. However, existing designs rely heavily on reserved rows to implement computation, which introduces high data movement overhead, large operation latency, large energy consumption, and low operation reliability. In this paper, we propose ELP2IM, an efficient and low power processing in-memory architecture, to address the above issues. ELP2IM utilizes two stable states of sense amplifiers in DRAM subarrays so that it can effectively reduce the number of intra-subarray data movements as well as the number of concurrently opened DRAM rows, which exhibits great performance and energy consumption advantages over existing designs. Our experimental results show that the power efficiency of ELP2IM is more than 2X improvement over the state-of-the-art DRAM based memory-centric designs in real application. Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
HPCA | 1 |
| 2019 | ROC: DRAM-based Processing with Reduced Operation CyclesabstractDRAM based memory-centric computing architectures are promising solutions to tackle the challenges of memory wall. In this paper, we develop a novel design of DRAM-based processing-in-memory (PIM) architecture which achieves lower cycles in every basic operation than prior arts. Our small yet fast in-memory computing units support basic logic operations including NOT, AND, and OR. Using those operations, along with shift and propagation, bitwise operations can be extended to word-wise operations, e.g. increment and comparison, with high efficiency. We also optimize the designs to exploit parallelism and data reuse to further improve the performance of compound operations. Compared with the most powerful state-of-the-art PIM architecture, we can achieve comparable or even better performance while consuming only 6% of its area overhead. Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
DAC | 1 |