EDBT 2026 Demo / reviewers in the wild / expert
Dapeng Ju
dblp:86/6462
· DBLP profile ↗
7ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0005-6495-2275ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pitfall: Uncovering and Exploiting the Store Forwarding Predictor on Intel CPUs
Dapeng Ju, Yongqiang Lyu 0001, Dongsheng Wang 0002 |
APPT | 3 |
| 2026 | SSBleed: Non-Speculative Side-Channel Attacks via Speculative Store Bypass on Armv9 CPUsabstractModern CPUs employ Speculative Store Bypass (SSB) to reduce load latency and improve performance. In response to transient attacks such as Spectre, CPU vendors have also introduced mitigations to prevent incorrect speculation from leaking data. In this work, we show that the SSB on Armv9 CPUs introduces a previously unexplored form of non-speculative data leakage. Specifically, we find that the SSB on Armv9 performance cores is governed by an undocumented predictor. Through reverse engineering, we uncover the design of this predictor and show that it lacks isolation across security domains. Furthermore, existing mitigations such as SSBS are insufficient to prevent leaks. Based on this, we present SSBleed, the first non-speculative side-channel attack via SSB on Armv9 CPUs. We validate the practicality of SSBleed through five case studies, including crossprocess RSA signature and key generation attacks on the latest version of MbedTLS and WolfSSL, interrupt detection, and improved data transmission in two transient attacks. Finally, we propose a flush-based mitigation through a kernel patch, which incurs an average performance overhead of 0.46 %. Chang Liu 0117, Hongpei Zheng, Xin Zhang 0110, Dapeng Ju, Dongsheng Wang 0002, Yinqian Zhang, Trevor E. Carlson |
HPCA | 4 |
| 2026 | OCCUPY+PROBE: Cross-Privilege Branch Target Buffer Side-Channel Attacks at Instruction Granularity
Kaiyuan Rong, Junqi Fang, Haixia Wang 0001, Dapeng Ju, Dongsheng Wang 0002 |
NDSS | 4 |
| 2023 | Leaky MDU: ARM Memory Disambiguation Unit Uncovered and Vulnerabilities ExposedabstractMemory Disambiguation Unit (MDU) is widely used on modern processors to speculatively execute load instructions and improve pipeline performance. Given that the MDU design details on ARM processors are not available to the public, it is unclear whether there are any security vulnerabilities associated with its MDU. In this paper, we first reverse engineer the undocumented features of ARM MDU, then we discover three potential user-privilege attacks to leak secret data via MDU: cross-process attack that allows users to communicate through a convert channel, cross-domain attack that leaks kernel information and a new variant of inner-process and inter-processes Spectre attacks. These attacks pose serious security challenges as they can bypass both all the known countermeasures against cache side-channel attacks and those against transient execution attacks. Potential mitigation against the proposed MDU-based attacks are also discussed. Chang Liu 0117, Yongqiang Lyu 0001, Haixia Wang 0001, Pengfei Qiu, Dapeng Ju, Gang Qu 0001, Dongsheng Wang 0002 |
DAC | 5 |
| 2014 | Möbius: A high performance transactional SSD with rich primitivesabstractProviding transactional primitives of NAND flash based solid state disks (SSDs) have demonstrated a great potential for high performance transaction processing and relieving software complexity. Similar with software solutions like write-ahead logging (WAL) and shadow paging, transactional SSD has two parts of overhead which include: 1) write overhead under normal condition, and 2) recovery overhead after power failures. Prior transactional SSD designs utilize out-of-band (OOB) area in flash pages to store transaction information to reduce the first part of overhead. However, they are required to scan a large part of or even whole SSD after power failures to abort unfinished transactions. Another limitation of prior approaches is the unicity of transactional primitive they provided. In this paper, we propose a new transactional SSD design named Möbius. Möbius provides different types of transactional primitives to support static and dynamic transactions separately. Möbius flash translation layer (mFTL), which combines normal FTL with transaction processing by storing mapping and transaction information together in a physical flash page as atom inode. By amortizing the cost of transaction processing with FTL persistence, MFTL achieve high performance in normal condition and does not increase write amplification ratio. After power failures, Möbius can leverage atom inode to eliminate unnecessary scanning and recover quickly. We implemented a prototype of Möbius and compare it with other state-of-art transactional SSD designs. Experimental results show that Möbius can at most 67% outperform in transaction throughput (TPS) and 29 times outperform in recovery time while still have similar or even better write amphfication ratio comparing with prior hardware approaches. Wei Shi 0001, Dongsheng Wang 0002, Zhanye Wang, Dapeng Ju |
MSST | 4 |
| 2009 | A Novel Optimization Method to Improve De-duplication Storage System PerformanceabstractData De-duplication has become a commodity component in data-intensive storage systems. But compared with other traditional storage paradigms, de-duplication system achieves elimination of data duplications or redundancies at the cost of bringing several additional layers or function components into the I/O path, and these additional components are either CPU-intensive or I/O intensive, largely hindering the overall system performance. Direct against the above potential system bottlenecks, this paper quantitatively analyzes the overhead of each main component introduced by de-duplication, and then proposes two performance optimization methods. The one is parallel calculation of content aware chunk identifiers, which fully utilizes the parallelism both inter and intra chunks by using a certain task partition and chunk content distribution algorithm. Experiments demonstrate that it can improve up to 150% of the system throughput, and at the same time much better utilize the multiprocessor resources. The other one is storage pipelining, which overlaps the CPU-bound, I/O-bound and network communication tasks. Through a dedicated five-stage storage pipeline design for file archival operations, experimental results show that the system throughput can increase up to 25% according to our workloads. Chuanyi Liu, Yibo Xue, Dapeng Ju, Dongsheng Wang 0002 |
ICPADS | 3 |
| 2009 | TH-CDP: An Efficient Block Level Continuous Data Protection SystemabstractTraditional data protection technologies, such as remote mirroring, snapshot and backup, cannot completely solve virus attack, user error problems. Continuous data protection (CDP), capturing all data writes at file or block level, is an enabling technology to storage systems against malicious attacks or user mistakes, because it allows each block data to be undoable. Most of existing data protection systems or prototypes are not real CDP ones because there is a data exposure between subsequent snapshots. Therefore they provide less granular recovery points. In addition, some products work either at file system level or at application level which lack general purpose. This paper presents the design, implementation, and performance of a new block level continuous data protection system, TH-CDP. Besides providing the basic functions of true CDP, TH-CDP provides virtual volume image at any point in time without affecting the production system. Another distinctive feature of TH-CDP is its checkpointing mechanism. By encapsulating the checkpoint information into I/O request packet queue, TH-CDP can take checkpoints without temporarily halting normal business processing or any incoming request. In addition, TH-CDP uses log-structured technique to store changed block data on raw disk, thereby speeding up both data writing and space recycling. Extensive experiments on file systems, databases using IOzone, and TPC-C benchmark show that TH-CDP can effectively improve the convenience of checkpointing and recovery assurance process. Under the circumstances of the changed block data up to 20 times of the original data size, the 100% sequence read speed of the oldest virtual volume image version is nearly 1/3 to 1/4 compared to the normal iSCSI volume. Yonghong Sheng, Dongsheng Wang 0002, Jinyang He, Dapeng Ju |
NAS | 4 |