EDBT 2026 Demo / reviewers in the wild / expert
Jindong Zhou
dblp:234/0258
· DBLP profile ↗
12ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 7 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CNN-Assisted Low-Power Clock Tree Synthesis for 3D ICsabstractIn this work, a convolutional neural network (CNN)-assisted low-power clock topology generation method for 3D clock tree synthesis (CTS) is proposed. Our approach considers both local and global costs in each merging step to prevent getting stuck in local optima. To explore the trade-off between local and global costs, we use CNN to set two important weighting factors to obtain an optimized clock tree that reduces power consumption. Compared with the conventional NNG-based method, experimental results on ISPD09 benchmarks show that our approach can reduce wirelength by 5.83% and power consumption by 4.73% on average. We also demonstrate the transferability of our method on larger-scale ISPD10 benchmarks. Chenbo Xi, Jindong Zhou, Pingqiang Zhou |
ASP-DAC | 2 |
| 2025 | Clock-Wirelength-Driven Detailed Placement
Ziang Ge, Yikai Liu, Jindong Zhou, Pingqiang Zhou |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Inductance-aware Clock Network Synthesis Considering Hierarchical Interconnects in 3D ICs
Jindong Zhou, Ziang Ge, Chenbo Xi, Pingqiang Zhou |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Spiking-NeRF: Spiking Neural Network for Energy-Efficient Neural RenderingabstractArtificial Neural Networks (ANNs) have achieved remarkable performance in many artificial intelligence tasks. As the application scenarios become more sophisticated, the computation and energy consumption of ANNs are also constantly increasing, which poses a challenge for deploying ANNs on energy-constrained devices. Spiking Neural Networks (SNNs) provide a promising solution to build energy-efficiency neural networks. However, the current training methods of SNNs cannot output values as precise as ANNs. This limits the applications of SNNs to relatively simple image classification tasks. In this article, we extend the application of SNNs to neural rendering tasks and propose an energy-efficient spiking neural rendering model, called Spiking-NeRF (Spiking Neural Radiance Fields). We first analyze the ANN-to-SNN conversion theory and propose an output scheme for SNNs to obtain the precise scene property values. Then we customize the parameter normalization method for the special network architecture of neural rendering. Furthermore, we present an early termination strategy (ETS) based on the discrete nature of spikes to reduce energy consumption. We evaluate the performance of Spiking-NeRF on both realistic and synthetic scenes. Experimental results show that Spiking-NeRF can achieve comparable rendering performance to ANN-based NeRF with up to \(2.27\times\) energy reduction. Ziwen Li 0004, Jindong Zhou, Pingqiang Zhou |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2023 | Exploring Remote Power Attacks Targeting Parallel Data Encryption On Multi-Tenant FPGAsabstractCloud service providers (CSPs) are increasingly incorporating Field Programmable Gate Arrays (FPGAs) into their cloud data centers due to the benefits of their flexibility and high performance in heterogeneous designs. However, the optimization of hardware resource utilization through multi-tenancy presents new security concerns. Prior research has demonstrated that remote side-channel attacks represent a significant security threat in the case of a single Advanced Encryption Standard (AES) module. However, it remains an open question whether parallel encryption can offer natural protection against Correlation Power Analysis (CPA). Our research focuses on side-channel attacks on parallel data encryption modules. We implemented delay-line based power sensors to collect mixed power traces and conducted CPA to steal the cipher key. Our results show that clocking methodology would have a significant influence on data protection. If parallel modules work at the same frequency without difference in clocking phase, the mixed voltage drops would contain sufficient information for attackers to decrypt the cipher key. Nevertheless, once the victim applies unique clocking phase to each module, he would convert voltage fluctuations from other modules into noises that offer a natural protection mechanism for parallel data encryption. Yankun Zhu, Jindong Zhou, Pingqiang Zhou |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | The study of TSV-induced and strained silicon-enhanced stress in 3D-ICs
Jindong Zhou, Youliang Jing, Pingqiang Zhou |
Integr. | 1 |
| 2023 | EaD: ECC-Assisted Deduplication With High Performance and Low Memory Overhead for Ultra-Low Latency Flash StorageabstractData deduplication has become a commodity feature in flash storage products to effectively reduce redundant write data and improve space efficiency. However, it also introduces computing and memory overhead to generate and store the cryptographic hash (fingerprint) in face of the moderate data redundancy in primary storage. With the advent of 3D XPoint and Z-NAND technologies, and the stronger cryptographic hash functions in use, such as SHA-256, both the computing and memory overheads are increasingly serious performance bottlenecks for inline data deduplication in these ultra-low latency flash storage. To address these problems, we propose an ECC-assisted Deduplication approach, called EaD, which exploits the ECC property and the asymmetric read-write performance characteristics of modern flash storage. EaD first identifies data similarity by leveraging the device-generated ECC values of data chunks as their fingerprints, significantly reducing the costly MD5/SHA-based cryptographic hash computing and alleviating the memory space overhead. Based on the identification results, similar data chunks and their ECCs are read from the flash to perform a byte-by-byte comparison in memory to definitively identify and remove redundant data chunks. Our experiments show that the EaD approach significantly increases I/O performance by up to 4.2${\times }$, with an average of 2.5${\times }$, compared with the existing MD5/SHA- and sampling-based deduplication approaches. Suzhen Wu, Chunfeng Du, Weidong Zhu 0002, Jindong Zhou, Hong Jiang 0001, Bo Mao 0003, Lingfang Zeng |
IEEE Trans. Computers | 4 |
| 2022 | ICARUS: A Specialized Architecture for Neural Radiance Fields RenderingabstractThe practical deployment of Neural Radiance Fields (NeRF) in rendering applications faces several challenges, with the most critical one being low rendering speed on even high-end graphic processing units (GPUs). In this paper, we present ICARUS, a specialized accelerator architecture tailored for NeRF rendering. Unlike GPUs using general purpose computing and memory architectures for NeRF, ICARUS executes the complete NeRF pipeline using dedicated plenoptic cores (PLCore) consisting of a positional encoding unit (PEU), a multi-layer perceptron (MLP) engine, and a volume rendering unit (VRU). A PLCore takes in positions & directions and renders the corresponding pixel colors without any intermediate data going off-chip for temporary storage and exchange, which can be time and power consuming. To implement the most expensive component of NeRF, i.e., the MLP, we transform the fully connected operations to approximated reconfigurable multiple constant multiplications (MCMs), where common subexpressions are shared across different multiplications to improve the computation efficiency. We build a prototype ICARUS using Synopsys HAPS-80 S104, a field programmable gate array (FPGA)-based prototyping system for large-scale integrated circuits and systems design. We evaluate the power-performancearea (PPA) of a PLCore using 40nm LP CMOS technology. Working at 400 MHz, a single PLCore occupies 16.5 mm 2 and consumes 282.8 mW, translating to 0.105 uJ/sample. The results are compared with those of GPU and tensor processing unit (TPU) implementations. Chaolin Rao, Huangjie Yu, Haochuan Wan, Jindong Zhou, Yueyang Zheng, Minye Wu, Anpei Chen, Binzhe Yuan, Pingqiang Zhou, Xin Lou 0001, Jingyi Yu 0001 |
ACM Trans. Graph. | 4 |
| 2020 | EaD: a Collision-free and High Performance Deduplication Scheme for Flash Storage SystemsabstractInline deduplication is a popular technique to effectively reduce the write traffic and improve the space efficiency for flash-based storage. However, it also introduces computing and memory overhead to generate and store the cryptographic hash (fingerprint). Along the advent of 3D XPoint and Z-NAND technologies with vastly improved latency and bandwidth, both the computing and memory overheads are becoming much more pronounced in deduplication-based flash storage with cryptographic hash functions in use. To address these problems, we propose an ECC (Error Correcting Code) assisted deduplication approach, called EaD, which exploits the ECC property and the asymmetric read-write performance characteristics of modern flash-based storage. EaD first identifies data similarity based on the fingerprints of data chunks represented by their ECC values, thus significantly reducing the costly cryptographic hash computing and alleviating the memory space overhead. Based on the identification results, similar data chunks and their ECCs are read from the flash to perform a byte-by-byte comparison in memory to definitively identify and remove redundant data chunks. Our experiments show that the EaD approach significantly reduces the I/O latency by an average of 1.92× and 1.86×, and reduces the memory consumption by an average of 35.0% and 21.9%, compared with the existing SHA- and sampling-based deduplication approaches, respectively. Suzhen Wu, Jindong Zhou, Weidong Zhu 0002, Hong Jiang 0001, Zhirong Shen, Bo Mao 0003 |
ICCD | 2 |
| 2019 | Improving Flash Memory Performance and Reliability for Smartphones With I/O DeduplicationabstractFlash-based storage subsystem is the key component that affects the system performance, reliability, and cost efficiency of Android-based smartphones. In this paper, we first introduce a trace collection tool specifically designed to capture the I/O requests with important content features in Android-based smartphones, which are critically important but rarely available in content-aware designs and optimizations, such as JProbe and Netlink. Based on the analysis of the traces collected from 15 popular mobile applications, we find that 20%-40% of the I/O requests on the I/O critical path of the storage stack are redundant and this data redundancy is minimally shared among different applications. Based on this key observation, we propose a content-aware optimization, called APP-Dedupe, that applies data deduplication on the I/O critical path to improve both performance and efficiency by reducing write amplification and improving GC efficiency of the flash storage on Android smartphones. The evaluation results show that APP-Dedupe reduces the GC overhead by an average of 41.5%, reduces the response times by up to 15.4% and reduces the amount of write data by an average of 45.2%. Bo Mao 0003, Jindong Zhou, Suzhen Wu, Hong Jiang 0001, Weijian Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | PFP: Improving the Reliability of Deduplication-based Storage Systems with Per-File ParityabstractData deduplication weakens the reliability of storage systems since by design it removes duplicate data chunks common to different files and forces these files to share a single physical date chunk, or critical chunk, after deduplication. Thus, the loss of a single such critical data chunk can potentially render all referencing (sharing) files unavailable. However, the reliability issue in deduplication-based storage systems has not received adequate attention. Existing approaches introduce data redundancy after files have been deduplicated, either by replication on critical data chunks, i.e., chunks with high reference count, or RAID schemes on unique data chunks, which means that these schemes are based on individual unique data chunks rather than individual files. This can leave individual files vulnerable to losses, particularly in the presence of transient and unrecoverable data chunk errors such as latent sector errors. To address this file reliability issue, this paper proposes a Per-File Parity (short for PFP) scheme to improve the reliability of deduplication-based storage systems. PFP computes the XOR parity within parity groups of data chunks of each file after the chunking process but before the data chunks are deduplicated. Therefore, PFP can provide parity redundancy protection for all files by intra-file recovery and a higher-level protection for data chunks with high reference counts by inter-file recovery. Our reliability analysis and extensive data-driven, failure-injection based experiments conducted on a prototype implementation of PFP show that PFP significantly outperforms the existing redundancy solutions, DTR and RCR, in system reliability, tolerating multiple data chunk failures and guaranteeing file availability upon multiple data chunk failures. Moreover, a performance evaluation shows that PFP only incurs an average of 5.7 percent performance degradation to the deduplication-based storage system. Suzhen Wu, Bo Mao 0003, Hong Jiang 0001, Huagao Luan, Jindong Zhou |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Improving Reliability of Deduplication-Based Storage Systems with Per-File ParityabstractThe reliability issue in deduplication-based storage systems has not received adequate attention. Existing approaches introduce data redundancy after files have been deduplicated, either by replication on critical data chunks, i.e., chunks with high reference count, or RAID schemes on unique data chunks, which means that these schemes are based on individual unique data chunks rather than individual files. This can leave individual files vulnerable to losses, particularly in the presence of transient and unrecoverable data chunk errors such as latent sector errors. To address this file reliability issue, this paper proposes a Per-File Parity (short for PFP) scheme to improve the reliability of deduplication-based storage systems. PFP computes the XOR parity within parity groups of data chunks of each file after the chunking process but before the data chunks are deduplicated. Therefore, PFP can provide parity redundancy protection for all files by intra-file recovery and a higher-level protection for data chunks with high reference counts by inter-file recovery. Our reliability analysis and extensive data-driven, failure-injection based experiments conducted on a prototype implementation of PFP show that PFP significantly outperforms the existing redundancy solutions, DTR and RCR, in system reliability, tolerating multiple data chunk failures and guaranteeing file availability upon multiple data chunk failures. Moreover, a performance evaluation shows that PFP only incurs an average of 5.7% performance degradation to the deduplication-based storage system. Suzhen Wu, Huagao Luan, Bo Mao 0003, Hong Jiang 0001, Gen Niu, Hui Rao, Jindong Zhou |
SRDS | 8 |