EDBT 2026 Demo / reviewers in the wild / expert
Fuxin Zhang
dblp:49/2756
· DBLP profile ↗
40ranked-venue papers
4as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 23 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Computer networks · 4 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CRAFT: Exploiting Precharge Cost Asymmetry for Adaptive DRAM Row Buffer Management
Weitong Wang, Yuyangheng Wang, Fuxin Zhang |
APPT | 5 |
| 2026 | CUTE-XS: Bottleneck Analysis and Collaborative Optimization for Integrating CUTE Into XiangShan
Junyu Yue, Chongxi Wang, Jinpeng Ye, Jianan Xie, Longbing Zhang, Fuxin Zhang |
APPT | 9 |
| 2026 | Fast compiler autotuning framework using design of experiments
Chenghua Xu, Jingwei Sun 0001, Mengna Sai, Fuxin Zhang, Guangzhong Sun, Weiwu Hu |
CCF Trans. High Perform. Comput. | 4 |
| 2026 | Modeling and Optimization of a Share-a-Ride Problem With Flexible Pick-Up and Drop-Off PointsabstractA share-a-ride problem (SARP), which integrates the transportation of both passengers and parcels by the ride-hailing platforms such as Uber and Lyft, has drawn considerable attention. This work introduces a novel share-a-ride problem with flexible pick-up and drop-off points (SARP-FUO) with the objectives of maximizing the total revenue of the ride-hailing platforms and minimizing the total travel distance of vehicles. A mixed integer programming model is developed to formulate SARP-FUO. Then, a knowledge-based multi-objective brain storm optimization algorithm (KM-BSO) is proposed to solve it. Two knowledge-based local search operators are specifically designed to enhance the exploration capability of KM-BSO for identifying potential nondominated solutions. The first operator employs a dynamic programming algorithm to readjust pick-up and drop-off points, while the second modifies vehicle routes based on four derived properties. Extensive experiments are conducted to compare KM-BSO with nondominated sorting genetic algorithm II, multi-objective evolutionary algorithm based on decomposition, multi-objective artificial bee colony algorithm, and a mathematical programming solver CPLEX. The results and statistical analysis demonstrate the superiority of the proposed approach in solving the studied problem. Finally, a sensitivity analysis is performed with and without flexible pick-up and drop-off points, demonstrating the advantages of the proposed model in developing intelligent public transportation systems. Liang Qi 0001, Quanlu Xie, Wenjing Luan, Fuxin Zhang, Yangming Zhou, Xiwang Guo 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | AdaTP: Enhancing Temporal Prefetching with Adaptive Metadata FilteringabstractTemporal prefetching is a promising technology to predict the memory addresses of irregular memory accesses.It retains correlations of cache miss addresses in metadata, which can be stored either on-chip or off-chip.Recent advancements have favoured on-chip metadata storage within portions of the last level cache, making the optimization of metadata storage effectiveness crucial, because the benefits brought by temporal prefetching can be easily offset by the reduced capacity for data in the last level cache.However, current state-of-the-art temporal prefetchers employ static strategies to filter the metadata, which often results in suboptimal performance gains.In this study, we introduce AdaTP, a novel method that dynamically adjusts metadata filtering strategy based on the runtime measurement of data and metadata demands.This adaptive filtering strategy leverages criticality and reuse conditions of load instructions.Specifically, it permits only the most critical loads with repetitive access patterns to store correlations when metadata storage is limited, and allows all loads to store correlations when sufficient storage is available.Our evaluations show that AdaTP achieves a 22.1% speedup compared to baseline stride prefetch in irregular memory intensive benchmarks in SPEC CPU2006 and SPEC CPU2017, and outperforms state-of-the-art temporal prefetcher Triage and Triangel by 4.0% and 6.5% respectively. Junliang Wu, Fuxin Zhang |
CF | 3 |
| 2025 | Constructing Quantum Implementations with the Minimal T-depth or Minimal Width and Their Applications
Zhenyu Huang 0004, Fuxin Zhang, Dongdai Lin |
EUROCRYPT (1) | 2 |
| 2025 | MeMo: Enhancing Representative Sampling via Mechanistic Micro-Model SignaturesabstractRepresentative Sampling, exemplified by SimPoint is widely utilized in pre-silicon performance evaluation. It relies on the code signature to characterize programs and select representative simulation points to estimate the performance. Several code signatures have been proposed, including BBV, MIC, BBV-LDV, and PMC. Although these code signatures can achieve low average performance estimation errors for benchmark suites, they still yield high errors for specific programs within those suites, due to their inherent limitations of incomplete characteristic representations and indirect performance correlations. In this work, we propose the utilization of mechanistic micromodel signatures, MeMo, to enhance representative sampling. MeMo contains various signatures sourced from a series of micromodels under different configurations. These models encompass fetch, issue, and cache models, each concentrating on certain$\mu$Arch structure constraints while idealizing the others. MeMo enables the comprehensive depiction of different program characteristics and their performance responses while not confined to specific$\mu$Arch. We thoroughly evaluated MeMo against current code signatures on SPEC CPU2017. The experiment demonstrated that compared to the most commonly utilized BBV, MeMo can substantially reduce the average CPI estimation error from 3.96% to 1.63% and significantly suppress the maximum estimation error from 19.94% to 6.49%. Chenji Han, Huai Xu, Guangyao Guo, Fuxin Zhang |
ISPASS | 5 |
| 2025 | Non-cooperative multi-agent deep reinforcement learning for channel resource allocation in vehicular networks
Fuxin Zhang, Sihan Yao, Wei Liu 0051, Liang Qi 0001 |
Comput. Networks | 1 |
| 2025 | SnsBooster: Enhancing Sampling-based μArch Evaluation Efficiency through Online Performance Sensitivity AnalysisabstractSampling-based methods, such as SimPoint, are widely used for efficient pre-silicon μ Arch evaluations, where the costs are the number of simulation points multiplied by the number of evaluated μ Arch designs. However, these costs keep growing with an increasing number of simulation points and expanding μ Arch design space. Although techniques have been developed to accelerate the μ Arch design space exploration, less attention has been given to further reducing the simulation budget of each μ Arch evaluation. Common strategies like reducing simulation coverage or sampling fewer simulation points typically compromise estimation accuracy. Therefore, further reducing the simulation budget without compromising estimation accuracy remains a critical research problem. In this work, we propose SnsBooster to enhance sampling-based μ Arch evaluation efficiency, based on two insights: (a) large portions of simulation points’ performance changes are typically insensitive to the evaluated μ Arch changes, and (b) simulation points’ performance sensitivities under specific μ Arch change correlate with their inherent characteristics. By online building a μ Arch-specific performance sensitivity classifier via progressive simulation and continuous validation, SnsBooster can identify and selectively evaluate only performance-sensitive points, thus reducing the simulation budget without compromising estimation accuracy. When applied across various μ Arch changes, SnsBooster achieves an average simulation budget reduction of 39.04% with an accuracy loss of only 0.14%, compared to simulating all the sampled points. Under the same accuracy loss, SnsBooster’s simulation budgets are only 64.73% and 65.60% of those required by methods of reducing simulation coverage or sampling fewer points. Besides, under identical simulation budgets, the average accuracy losses of these methods are 1.41% and 1.23%, which is substantially higher than that of SnsBooster. Chenji Han, Zifei Zhang 0001, Xinyu Li 0010, Qi Guo 0001, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 9 |
| 2025 | Tiaozhuan: A General and Efficient Indirect Branch Optimization for Binary TranslationabstractBinary translation enables transparent execution, analysis, and modification of the binary program, serving as a core technology that facilitates instruction set emulation, cross-platform compatibility of software, and program instrumentation. Handling indirect branch instructions is widely recognized as a significant performance bottleneck in binary translation. While the target of a direct branch can be determined during the translation phase, an indirect branch requires a runtime lookup from the guest program counter to the host program counter, significantly influencing the performance of translator. Although several methods have been proposed to accelerate this process, each guest indirect branch instruction still translates into approximately 10 host instructions, resulting in considerable overhead. This article introduces Tiaozhuan, which addresses this issue by employing two optimization schemes. First, full address mapping uses a larger address space to store address mappings from guest to host, effectively reducing the number of instructions required to lookup the target of an indirect branch. Second, exceptionassisted branch elimination further eliminates branch instructions that check target correctness of targets in the lookup process. These two approaches enable indirect branches target lookup to be completed within one to two instructions, noticeably decreasing the overhead of indirect branches. Compared to state-of-the-art mechanisms, the SPEC CPU2006 benchmark suite showed a reduction in the number of instructions by an average of 4.2%, with the highest observed performance improvement reaching 19.4% and an average increase of 3.9%. Xinyu Li 0010, Guangyao Guo, Yanzhi Lan, Chenji Han, Gen Niu, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 7 |
| 2025 | Augur: Semantics-Aware Temporal Prefetching for Linked Data StructureabstractLinked data structures (LDS), such as lists and trees, are widely used in modern applications. Traversing LDS typically involves a significant amount of pointer chasing. Due to the serial nature of memory access in pointer chasing, the incurred long memory latency of traversing LDS has become a critical performance bottleneck. Furthermore, the poor spatial locality in LDS makes it difficult for spatial prefetchers to predict access addresses. Although temporal prefetchers can handle irregular memory access patterns, hindered by the challenges of collecting semantic information, current state-of-the-art temporal prefetchers suffer from significant metadata redundancy and frequent metadata conflicts. Consequently, there remain substantial opportunities to enhance the LDS prefetching. To solve this problem, we propose Augur, a semantics-aware temporal prefetcher to enhance LDS performance. Augur utilizes a novel pruning method to obtain semantic information and effectively extracts node address correlations from the perspective of nodes in LDS, thereby diminishing the metadata redundancy and conflicts. Additionally, Augur employs efficient metadata management strategies that guarantee a minimal storage overhead. Evaluated on LDS workloads, Augur achieves an average performance speedup of 17.8% and 11.7% over the baseline stride prefetcher and state-of-the-art spatial prefetcher Berti, respectively. Furthermore, Augur outperforms the state-of-the-art temporal prefetcher MISB, Triage, and Triangel, by 17.4%, 12.8%, and 6.3%, respectively, with a significantly lower storage overhead of only 1.26 KB. Junliang Wu, Chenji Han, Xinyu Li 0010, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 7 |
| 2025 | ETBench: Characterizing Hybrid Vision Transformer Workloads Across Edge DevicesabstractLightweight Convolution and Vision Transformer hybrid models have increasingly dominated the frontiers of deep learning (DL) on edge devices; however, to the best of our knowledge, no prior work has provided comprehensive evaluation on hybrid models’ performance and analyzed their characteristics by diving deep into the edge ecosystem with diversified modern DL inference engines and heterogeneous hardware. This paper proposes a comprehensive open-source benchmark suite,ETBench, to allow power-efficiency, performance and accuracy assessment for state-of-the-art (SOTA) hybrid models across 11 most widely-used DL engines deployed on diverse edge devices. After building ETBench that satisfies 6 design requirements proposed in our work, we conduct extensive experiments on 14 devices including 19 CPUs, 11 GPUs and 5 NPUs, and obtain benchmark results from all deployment scenarios (combinations of models, quantization formats, software engines, and hardware platforms). Valuable observations and insightful implications are finally summarized. For example, within current DL engines, the INT8 quantization is significantly underperformed in terms of accuracy and speed against FP16 for hybrid models. Overall, ETBench serves as a collaborative platform that assists model architects in better evaluating their models and makes it possible for future co-optimizations of DL engines and hardware accelerators. Yingkun Zhou, Zhengshuyuan Tian, Jinpeng Ye, Chenji Han, Fuxin Zhang |
IEEE Trans. Computers | 8 |
| 2024 | OmniCache: A Unified Cache for Efficient Query Handling in LSM-tree Based Key-Value StoresabstractKey-value (KV) stores built on Log-Sturctured Merge (LSM) trees have become fundamental in handling large-scale data in modern applications. While LSM-trees are optimized for high write throughput, they suffer from degraded read performance. Caching is one of the main techniques to improve the performance of read operations. Traditional block caches have several issues: 1. They store redundant data for point lookups by caching entire blocks even when only a single KV pair is needed. 2. Frequent compaction operations cause cached blocks to become invalid, leading to increased disk I/Os and higher costs. 3. Block cache does not store global ordering information, so for range queries, the system must repeatedly fetch, sort, and merge data from multiple blocks to generate results, leading to substantial overhead.In this paper, we introduce OmniCache, a unified cache system that stores the results of both point lookups and range queries in LSM-tree based KV stores to optimize read performance. Omni-Cache employs a hybrid data structure, combining a hash table and skip-list, to dynamically store fine-grained results with global ordering. Our system reduces redundant data, mitigates cache invalidation from compaction, and improves query performance. Experimental evaluations show that OmniCache avoids 39% disk read traffic by reducing block cache invalidation. OmniCache also improves throughput up to 2.72x in read-heavy workloads and up to 1.85x in write-heavy workloads. Yiyang Geng, Huai Xu, Yanyong Zhang, Fuxin Zhang |
HPCC | 4 |
| 2024 | AVM-BTB: Adaptive and Virtualized Multi-level Branch Target BufferabstractBranch Target Buffer (BTB) plays an important role in modern processors. It is used to identify branches in the instruction stream and predict branch targets. The accuracy of BTB is highly impacted by BTB capacity. However, expanding BTB capacity using traditional methods requires valuable on-chip SRAM. Both timing and area restriction make these approaches unsustainable. Moreover, these methods overlook the different demands of various applications, leading to increased power consumption and resource waste in some cases. To address this problem, we propose AVM-BTB. The key observations behind AVM-BTB come from three aspects: 1) BTB requirements vary over different applications and even over different running stages of the same application. 2) Micro-operation Cache (Uop Cache) and ICache exhibit inefficiency when confronted with instruction footprints that greatly exceed their capacity. 3) In specific scenarios of frontend overload, reducing cache capacity and increasing BTB size can effectively mitigate expensive branch prediction errors. Simultaneously, the implementation of Fetch Directed Instruction Prefetching (FDIP) can offset the limitations in cache capacity to some extent. These observations reveal the feasibility of dynamically borrowing cache capacity as temporary BTB and returning these BTB to cache when they are not needed, further resulting in an adaptive and virtualized multi-level BTB scheme. However, such a BTB structure is non-trivial. In this work, from the perspective of instructions, the cache hierarchy stores instruction data, while the BTB stores metadata used for branch prediction and instruction prefetch. Targeting high performance, AVM-BTB maintains a dynamic balance between data and metadata by monitoring the BTB error rate and effective accesses. Evaluation with 1253 traces shows that AVM-BTB is suitable for both frontend-bound and frontend-friendly scenarios, without consuming additional SRAM and with reasonable implementation efforts. Compared to baseline, AVM-BTB delivers an average performance boost of $18.22\%$ and a power consumption reduction of $2.77\%$. It also outperforms the five state-of-the-art solutions by $6.26 \%-18.26 \%$ on average in terms of IPC. Yunzhe Liu 0005, Xinyu Li 0010, Qi Guo 0001, Fuxin Zhang |
ISCA | 6 |
| 2024 | High-Utilization GPGPU Design for Accelerating GEMM Workloads: An Incremental ApproachabstractGeneral Purpose Graphics Processing Units (GPGPUs) have been employed primarily in domains such as graphics acceleration and high-performance computing in the past. However, the rise of artificial intelligence (AI), particularly the computational demands associated with matrix multiplications in AI models, has presented formidable demands on the computational power of GPGPUs. Consequently, the design of matrix multiplication units within GPGPUs and ensuring their utilization have become key issues in optimizing AI workloads. This paper explores an incremental design approach, building upon a Single Instruction Multiple Threads (SIMT) GPGPU architecture, to facilitate General Matrix Multiply (GEMM) acceleration. This approach encompasses not only the design of matrix units within the stream processors but, more crucially, the optimization of the data path within the GPGPU to maximize the utilization of the matrix units. We present a practical demonstration of our approach through the fabrication of a GPGPU on a 12 nm CMOS process node, achieving a core clock speed of 1 GHz and INT8 peak performance of 8 TOPS with memory bandwidth limited to 32 GB/s LPDDR4-4000. Notably, this design results in only a 6.57% increase in chip area compared to the original GPGPU design. In a series of fair GEMM workload tests, the GPGPU implemented in this work outperforms the recent three generations of NVIDIA GPGPUs—V100, T4, and A100—in terms of matrix unit utilization. Chongxi Wang, Penghao Song, Fuxin Zhang, Longbing Zhang |
ISCAS | 4 |
| 2024 | BTBench: A Benchmark for Comprehensive Binary Translation Performance EvaluationabstractBinary translation serves as a fundamental technol-ogy for instruction set emulation, system virtualization, runtime instrumentation, and numerous other applications. Many techniques have been proposed to enhance the efficiency of binary translation systems. However, imprecise performance evaluation leads to potential performance shortcomings of binary translators in real-world applications. Previous studies primarily employ CPU benchmarks, which may overlook performance issues spe-cific to binary translators and fail to guide for optimizing binary translation. To address this issue, we propose a new benchmark suite named BTBench(Binary Translation Benchmark), which provides a convenient, portable, and comprehensive solution. BT-Bench takes into account the inherent attributes of binary trans-lators, such as translation and code-cache lookup overhead. We carefully select benchmarks that offer comprehensive coverage and align with real-world application scenarios. To validate the effectiveness of BTBench, we conducted rigorous experimentation on four widely-used binary translators. The analysis of the results reveals that, compared to existing CPU benchmarks, BTBench is better suited for identifying potential performance shortcomings, providing invaluable insights for future optimization efforts. The BTBench benchmark suite is publicly available1• Xinyu Li 0010, Yanzhi Lan, Gen Niu, Fuxin Zhang |
ISPASS | 5 |
| 2024 | Joint Awareness and Congestion Control in Vehicular NetworksabstractIn vehicular networks, vehicles frequently broad-cast vehicle states information to track the movement of their neighbors. A large number of vehicles get access to the shared channel resources to broadcast their states information, which may lead to channel congestion. The existing channel congestion solutions mainly focus on the media access control layer state to optimize channel resource utilization, without considering the impact of the physical layer on successful packet reception performance. In this article, we propose a packet reception model to represent the successful reception rate of state packets under the influence of interference signals and noise, and design an application-specific utility function. This utility considers the states of the physical layer and the media access control layer to balance reducing the physical layer interference and meeting application requirements. We then propose a distributed joint power and rate control algorithm that uses on-demand mode to allocate channel resources to meet the safety application requirements of each vehicle. The simulation results show that our work can effectively prevent channel congestion and improve the safety performance of vehicles in different driving scenarios. Fuxin Zhang |
SMC | 2 |
| 2024 | Distributed and Adaptive Message Dissemination for Vehicle Platooning in Hybrid TrafficabstractIn vehicular networks, vehicle in the platooning relies on dissemination of beacons to perceive the status of neighbor vehicles and then take control low to maintain a constant inter-vehicle distance. Vehicle platooning communication has stringent high-reliability and low-latency requirements. In this paper, we focus on resource allocation strategies for vehicle platoons in hybrid scenarios. First, we propose a two-dimensional Markov model to describe the channel contention (back-off) processes for platoons and individual vehicles. Based on the model, we then can derive the transmission probabilities of beacons and event messages for platoons and individual vehicles. Finally, we design a distributed and adaptive beacon control scheme to determine the optimal beacon rates for vehicles in the platoons. The simulation results demonstrate that our efforts substantially enhance platooning communication reliability and ensure optimal performance for vehicle safety applications. Fuxin Zhang |
SMC | 2 |
| 2024 | Context-aware resource allocation for vehicle-to-vehicle communications in cellular-V2X networks
Fuxin Zhang, Guangping Wang |
Ad Hoc Networks | 1 |
| 2024 | Distributed and Coordinated Model Predictive Control for Channel Resource Allocation in Cooperative Vehicle Safety SystemsabstractCooperative vehicle safety systems rely on periodic broadcasts of beacons to track positions and movements of concerned vehicles. In vehicular networking, vehicle driving environment is changing rapidly. This unique characteristic can cause dynamic network topology and heavy traffic conditions. In scenarios where traffic density is high, a large number of beacons could cause channel congestion, and the tracking performance of safety applications can thus be seriously impacted. To maintain high tracking accuracy for each node under varying traffic situations, this paper presents a distributed and coordinated channel access control strategy based on Multi-agent Model Predictive Control theory. First, we propose a multi-dimensional and hybrid Petri net model to characterize the interactions among multiple vehicles. The interaction model describes the possibility of collisions among vehicles. We then propose an application-dependent utility function that incorporates inter-vehicle collision behaviour. A model predictive control problem for beaconing rate adaption is formulated based on the function. Next, a distributed and coordinated decision-making scheme is designed. In this scheme, each node is treated as an agent. Each agent uses a model predictive control controller and coordinates with its neighboring agents to take channel access control actions. Simulation results validate that it improves channel resource utilization and tracking accuracy under dynamic driving situations. Fuxin Zhang, MengChu Zhou, Liang Qi 0001 |
IEEE Internet Things J. | 1 |
| 2024 | CUTE: A scalable CPU-centric and Ultra-utilized Tensor Engine for convolutions
Jinpeng Ye, Fuxin Zhang |
J. Syst. Archit. | 3 |
| 2024 | A dependence graph pattern mining method for processor performance analysis
Chenji Han, Fuxin Zhang |
Perform. Evaluation | 4 |
| 2024 | An Instruction Inflation Analyzing Framework for Dynamic Binary TranslatorsabstractDynamic binary translators (DBTs) are widely used to migrate applications between different instruction set architectures (ISAs). Despite extensive research to improve DBT performance, noticeable overhead remains, preventing near-native performance, especially when translating from complex instruction set computer (CISC) to reduced instruction set computer (RISC). For computational workloads, the main overhead stems from translated code quality. Experimental data show that state-of-the-art DBT products have dynamic code inflation of at least 1.46. This indicates that on average, more than 1.46 host instructions are needed to emulate one guest instruction. Worse, inflation closely correlates with translated code quality. However, the detailed sources of instruction inflation remain unclear. To understand the sources of inflation, we present Deflater , an instruction inflation analysis framework comprising a mathematical model, a collection of black-box unit tests called BenchMIAOes , and a trace-based simulator called InflatSim . The mathematical model calculates overall inflation based on the inflation of individual instructions and translation block optimizations. BenchMIAOes extract model parameters from DBTs without accessing DBT source code. InflatSim implements the model and uses the extracted parameters from BenchMIAOes to simulate a given DBT’s behavior. Deflater is a valuable tool to guide DBT analysis and improvement. Using Deflater, we simulated inflation for three state-of-the-art CISC-to-RISC DBTs: ExaGear, Rosetta2, and LATX, with inflation errors of 5.63%, 5.15%, and 3.44%, respectively for SPEC CPU 2017, gaining insights into these commercial DBTs. Deflater also efficiently models inflation for the open source DBT QEMU and suggests optimizations that can substantially reduce inflation. Implementing the suggested optimizations confirms Deflater’s effective guidance, with 4.65% inflation error, and gains 5.47x performance improvement. Benyi Xie, Chenghao Yan, Sicheng Tao, Xinyu Li 0010, Yanzhi Lan, Xiang Wu 0016, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 11 |
| 2024 | Tyche: An Efficient and General Prefetcher for Indirect Memory AccessesabstractIndirect memory accesses (IMAs, i.e., A [ f ( B [ i ])]) are typical memory access patterns in applications such as graph analysis, machine learning, and database. IMAs are composed of producer-consumer pairs, where the consumers’ memory addresses are derived from the producers’ memory data. Due to the built-in value-dependent feature, IMAs exhibit poor locality, making prefetching ineffective. Hindered by the challenges of recording the potentially complex graphs of instruction dependencies among IMA producers and consumers, current state-of-the-art hardware prefetchers either (a) exhibit inadequate IMA identification abilities or (b) rely on the run-ahead mechanism to prefetch IMAs intermittently and insufficiently. To solve this problem, we propose Tyche, 1 an efficient and general hardware prefetcher to enhance IMA performance. Tyche adopts a bilateral propagation mechanism to precisely excavate the instruction dependencies in simple chains with moderate length (rather than complex graphs). Based on the exact instruction dependencies, Tyche can accurately identify various IMA patterns, including nonlinear ones, and generate accurate prefetching requests continuously. Evaluated on broad benchmarks, Tyche achieves an average performance speedup of 16.2% over the state-of-the-art spatial prefetcher Berti. More importantly, Tyche outperforms the state-of-the-art IMA prefetchers IMP, Gretch, and Vector Runahead, by 15.9%, 12.8%, and 10.7%, respectively, with a lower storage overhead of only 0.57 KB. Chenji Han, Xinyu Li 0010, Junliang Wu, Yifan Hao 0001, Zidong Du, Qi Guo 0001, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 10 |
| 2023 | SCFM: A Statistical Coarse-to-Fine Method to Select Cross-Microarchitecture Reliable Simulation Points
Chenji Han, Hongze Tan 0001, Xinyu Li 0010, Ruiyang Wu 0001, Fuxin Zhang |
APPT | 6 |
| 2023 | On-Demand Triggered Memory Management Unit in Dynamic Binary Translator
Benyi Xie, Xinyu Li 0010, Chenghao Yan, Fuxin Zhang |
APPT | 8 |
| 2023 | MFHBT: Hybrid Binary Translation System with Multi-stage Feedback Powered by LLVM
Zhaoxin Yang, Xuehai Chen, Liangpu Wang, Weiming Guo, Dongru Zhao, Fuxin Zhang |
APPT | 7 |
| 2023 | RBGC: Repurpose the Buffer of Fixed Graphics Pipeline to Enhance GPU CacheabstractThe limited cache size of GPU in general-purpose computing hinders the execution efficiency of thousands of concurrent threads. Several techniques have been proposed to increase the cache size per thread, such as repurposing shared memory and register files as a cache to reduce contention. However, these studies only focus on improving the general-purpose computing hardware structure, ignoring the fixed graphics pipeline hardware structure. To solve this issue, we propose repurposing the buffer of the fixed graphics pipeline as a victim cache to enhance the GPU cache. This strategy utilizes the idle fixed graphics pipeline buffer as a victim cache for general-purpose computing tasks. The victim cache returns data to the L1 data cache when a load is located at the victim cache. We also optimize the victim cache management strategy by monitoring load reusability and only allocating cache lines to loads with higher reusability. This optimization improves the victim cache efficiency. Our experimental results show that the RBGC achieves a 39.1% performance improvement with minimal hardware overhead compared to the baseline GPU architecture. Longbing Zhang, Fuxin Zhang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | LAST: An Efficient In-place Static Binary Translator for RISC Architectures
Yanzhi Lan, Gen Niu, Xinyu Li 0010, Liangpu Wang, Fuxin Zhang |
ICA3PP (2) | 6 |
| 2023 | Randomized Testing Framework for Dissecting NVIDIA GPGPU Thread Block-To-SM Scheduling MechanismsabstractNVIDIA General Purpose Graphics Processing Units (GPGPUs) have been extensively utilized in scientific computing, artificial intelligence acceleration, and numerous other fields. The compute unified device architecture (CUDA) programming interface facilitates access to the enormous computing power of NVIDIA GPGPUs. However, the thread block-to-stream multiprocessor (SM) scheduler of NVIDIA GPGPUs has remained a black box, presenting challenges for performance modeling of CUDA programs, especially those featuring concurrent kernel execution, as well as hindering further exploration of the full potential of NVIDIA GPGPUs. To address these issues, we have devised a randomized testing framework targeting thread block-to-SM scheduling. In this framework, concurrent kernel sequences are randomly generated and assessed by a self-implemented scheduling predictor along with an actual NVIDIA GPGPU. Through iteratively scrutinizing the discrepancies between predictions and actual results, we dissect the key thread block-to-SM scheduling mechanisms intrinsic to NVIDIA GPGPUs, revealing their underlying logic and considerations. This process continuously refines the scheduling predictor until it accurately aligns with real-world NVIDIA GPGPU scheduling. Furthermore, we investigate the monitoring and allocation mechanisms for distinct resources in NVIDIA GPGPUs and SMs, elucidating the essential connection between resource management and thread block scheduling. Our comprehensive and precise analysis of Nvidia GPGPU thread block-to-SM scheduling contributes significantly to constructing more accurate simulators for GPGPU performance modeling. Moreover, it facilitates optimizing CUDA programs based on the scheduling mechanisms to exploit more parallelism among thread blocks of concurrent kernels. Chongxi Wang, Penghao Song, Fuxin Zhang, Longbing Zhang |
ICPADS | 4 |
| 2022 | Automatic ICD Coding Exploiting Discourse Structure and Reconciled Code EmbeddingsabstractThe International Classification of Diseases (ICD) is the foundation of global health statistics and epidemiology. The ICD is designed to translate health conditions into alphanumeric codes. A number of approaches have been proposed for automatic ICD coding, since manual coding is labor-intensive and there is a global shortage of healthcare workers. However, existing studies did not exploit the discourse structure of clinical notes, which provides rich contextual information for code assignment. In this paper, we exploit the discourse structure by leveraging section type classification and section type embeddings. We also focus on the class-imbalanced problem and the heterogeneous writing style between clinical notes and ICD code definitions. The proposed reconciled embedding approach is able to tackle them simultaneously. Experimental results on the MIMIC dataset show that our model outperforms all previous state-of-the-art models by a large margin. The source code is available at https://github.com/discnet2022/discnet Shurui Zhang 0002, Bozheng Zhang, Fuxin Zhang, Bo Sang, Wanchun Yang |
COLING | 3 |
| 2022 | Eliminate the overhead of interrupt checking in full-system dynamic binary translatorabstractDynamic binary translation (DBT) is a ubiquitous technique for program emulation, instrumentation and debugging. Full-system dynamic binary translators, which can run operating systems, are required to emulate interrupt delivery. Existing full-system dynamic binary translators use a simple scheme to do so, by attaching to each translated code block a prologue that checks for pending interrupts. However, this approach is inefficient, as interrupts are delivered infrequently, relatively to the execution of translated blocks, and therefore most of the interrupt checks are unnecessary and wasteful. Gen Niu, Fuxin Zhang, Xinyu Li 0010 |
SYSTOR | 2 |
| 2021 | BTMMU: an efficient and versatile cross-ISA memory virtualizationabstractFull system dynamic binary translation (DBT) has many important applications, but it is typically much slower than the native host. One major overhead in full system DBT comes from cross-ISA memory virtualization, where multi-level memory address translation is needed to map guest virtual address into host physical address. Like the SoftMMU used in the popular open-source emulator QEMU, software-based memory virtualization solutions are not efficient. Meanwhile, mature techniques for same-ISA virtualization such as shadow page table or second level address translation are not directly applicable due to cross-ISA difficulties. Some previous studies achieved significant speedup by utilizing existing hardware (TLB or virtualization hardware) of the host. However, since the hardware is not designed with cross-ISA in mind, those solutions had some limitations that were hard to overcome. Most of them only supported guests with smaller virtual address space than the host. Some supported only guests with the same page size. And some did not support privileged memory accesses. Kele Huang, Fuxin Zhang, Gen Niu, Junrong Wu |
VEE | 2 |
| 2020 | A Game Theoretic Approach for Distributed and Coordinated Channel Access Control in Cooperative Vehicle Safety SystemsabstractFairness and efficiency are two key requirements that have to be guaranteed in channel resource allocation in cooperative vehicle safety systems. Existing channel access control strategies, however, rely on each individual node to adjust networking parameters independently according to its locally measured state information, thus leading to unfairness. Although some coordinated strategies have been proposed to resolve this issue, they pay little attention to the efficiency. In order to achieve the tradeoff between fairness and efficiency, in this paper, we propose a utility function in terms of inter-packet reception time required to capture the performance of consecutive successful packets' reception under various vehicle densities. A channel access control problem among vehicles is then formulated as a non-cooperation game model. This model utilizes a punishment function to penalize a node that monopolizes the channel resources and hence can enable nodes to coordinate with each other to achieve desired fairness and efficiency. Next, a distributed decision-making scheme for channel access control is designed. It adjusts transmission rate in a coordinated manner and can guide each node to reach a Pareto-optimal Nash equilibrium point. The experimental results validate that the proposed strategy can result in fair channel resource allocation and ensure high-tracking accuracy for each vehicle under dynamic traffic conditions. Fuxin Zhang, MengChu Zhou, Liang Qi 0001, Yuyue Du, Haichun Sun |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | APDM: An adaptive multi-priority distributed multichannel MAC protocol for vehicular ad hoc networks in unsaturated conditions
Caixia Song, Guozhen Tan, Chao Yu 0004, Nan Ding 0001, Fuxin Zhang |
Comput. Commun. | 5 |
| 2007 | Code reordering on limited branch offsetabstractSince the 1980's code reordering has gained popularity as an important way to improve the spatial locality of programs. While the effect of the processor's microarchitecture and memory hierarchy on this optimization technique has been investigated, little research has focused on the impact of the instruction set. In this paper, we analyze the effect of limited branch offset of the MIPS-like instruction set [Hwu et al. 2004, 2005] on code reordering, explore two simple methods to handle the exceeded branches, and propose the bidirectional code layout (BCL) algorithm to reduce the number of branches exceeding the offset limit. The BCL algorithm sorts the chains according to the position of related chains, avoids cache conflict misses deliberately and lays out the code bidirectionally. It strikes a balance among the distance of related blocks, the instruction cache miss rate, the memory size required, and the control flow transfer. Experimental results show that BCL can effectively reduce exceeded branches by 50.1%, on average, with up to 100% for some programs. Except for some programs with little spatial locality, the BCL algorithm can achieve the performance, as the case with no branch offset limitation. Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 2 |
| 2005 | Microarchitecture of the Godson-2 Processor
Weiwu Hu, Fuxin Zhang, Zusong Li |
J. Comput. Sci. Technol. | 2 |
| 2001 | Communication with Threads in Software DSMabstractMost software DSMs use the interrupt mechanism for asynchronous message arrival notification. Interrupt, however, is expensive and may not be available in user-level communication protocols. This paper studies the performance of software DSMs using threads instead of interrupt for communication. We implement three thread communication methods in the JIAJIA software DSM. The first method separates a communication thread from the main thread, the second one divides the communication thread into a sending thread and a receiving thread, and the third method further splits a serving thread from the receiving thread. The effect of the thread communication methods is evaluated in a PC cluster and a cluster of PowerPC workstations with some well-accepted benchmarks. Evaluation results show that the performance of thread communication is slightly worse than that of interrupt communication when there is only one processor to run a process, but is better than that of interrupt communication when there are more than one processors in a node to run a process. The thread communication method with one sending thread and one receiving thread achieves the best performance inthethree thread communication methods. Besides, the effect of the thread communication is also related to the operating system. Analysis of the evaluation results reveals that the waiting time constitutes the major overhead of the execution, and the dominant factor that in uence waiting time is neither the interrupt overhead nor the thread switching overhead but the responsiveness of requests (i.e., whether requests can be processed and replied promptly). Weiwu Hu, Fuxin Zhang |
CLUSTER | 3 |
| 2001 | Dynamic Data Prefetching in Home-Based Software DSMs
Weiwu Hu, Fuxin Zhang |
J. Comput. Sci. Technol. | 2 |
| 2000 | A New Home-Based Software DSM Protocol for SMP Clusters
Weiwu Hu, Fuxin Zhang |
Euro-Par | 2 |