EDBT 2026 Demo / reviewers in the wild / expert
Tingji Zhang
dblp:274/7011
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Look Before You Leap : Precision Instruction Supply via SmartScoutabstractModern high-performance processors extensively employ Fetch-Directed Instruction Prefetching (FDIP) to mitigate instruction supply bottlenecks. However, the efficacy of FDIP is fundamentally constrained by the accuracy of the Branch Prediction Unit (BPU). As the critical component within the BPU, the Branch Target Buffer (BTB) faces severe capacity bottlenecks. While pre-decoding-based prefetching offers a remedy, existing approaches suffer from two critical impediments: (1) The Noise Dilemma: Suboptimal trade-off between coverage and accuracy. (2) Inefficient Miss Resolution: Current designs rely on reactive recovery or stalls, failing to leverage available front-end slack for proactive correction. Peng Qu 0001, Tingji Zhang, Fang Su, Zhe Pan 0001, Youhui Zhang |
ICS | 3 |
| 2025 | Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server ApplicationsabstractThe large working set of instructions in server-side applications causes a significant bottleneck in the front-end, even for high-performance processors equipped with fetch-directed instruction prefetching (FDIP). Prefetchers specifically designed for server scenarios typically rely on a record-and-replay mechanism that exploits the repetitiveness of instruction sequences. However, the efficacy of these techniques is compromised by discrepancies between actual and predicted control flows, resulting in loss of coverage and timeliness. This paper proposes Hierarchical Prefetching, a novel approach that tackles the limitations of existing prefetchers. It identifies common coarse-grained functionality blocks (called Bundles) within the server code and prefetches them as a whole. Bundles are significantly larger than typical prefetch targets, encompassing tens to hundreds of kilobytes of code. The approach combines simple software analysis of code for bundle formation and light-weight hardware for record-and-replay prefetching. The prefetcher requires under 2KB of on-chip storage by keeping most of the metadata in main memory. Experiments with 11 popular server workloads reveal that Hierarchical Prefetching significantly improves miss coverage and timeliness over prior techniques, achieving a 6.6% average performance gain over FDIP. Tingji Zhang, Boris Grot, Wenjian He, Yashuai Lv, Peng Qu 0001, Fang Su, Guowei Zhang 0002, Youhui Zhang |
ASPLOS (2) | 1 |
| 2025 | SoftGuide: A Hardware-Software Co-design Predictor for Data-Dependent BranchesabstractBranch predictors based on historical information exhibit exemplary performance. However, a subset of data-dependent branches still pose a significant challenge, often resulting in severe mispredictions. Such branches are commonly encountered during the processing of various data structures, and increasing the capacity of predictors has shown limited effective-ness. Thus, a dedicated branch predictor is necessary to enhance performance for varying data structures while maintaining lower hardware complexity. This paper proposes SoftGuide, a novel hardware-software cooperative branch predictor designed to tackle two main chal-lenges inherent in data-dependent branch prediction: (1) the detection and identification of branch dependencies(software-friendly) and (2) data prefetching and dependency chain pre-execution(hardware-friendly). SoftGuide leverages software to convey the memory access patterns and the dependency chains associated with the branch, thereby circumventing the overhead of hardware-based detection. Utilizing the information provided by the software, the enhanced hardware prefetches data and triggers pre-execution in advance. SoftGuide can perfectly unify prefetching and prediction tasks. For SPEC2006 and GAP benchmarks with the method of SimPoint, SoftGuide realizes a decrease in branch mispredictions per 1K instructions (MPKI) by 46.4% and an increase in Instructions Per Cycle (IPC) by 1.25x average on the processor equipped with the state-of-the-art branch predictor. Moreover, the storage overhead is just 3.88KB. Peng Qu 0001, Tingji Zhang, Youhui Zhang |
CCGrid | 3 |