EDBT 2026 Demo / reviewers in the wild / expert
Xianhua Liu 0001
dblp:95/1574
· DBLP profile ↗
16ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-4777-3847ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMIX: Schedulable Instruction Set Architecture Extension Interface for Multi-Operand OperatorsabstractIntegrating domain-specific operators into processor cores is essential for performance scaling. However, multi-operand operators often face a semantic gap with conventional ISAs, which are limited in operand capacity and scheduling flexibility. This paper presents SMIX, a schedulable instruction set extension interface for multi-operand operators. SMIX decouples execution into three stages: out-of-order input filling, computation, and out-of-order result picking. By employing explicit encoding and counter-based dependency management, SMIX enables both efficient static scheduling by compilers and dynamic out-of-order execution in hardware. Experimental results demonstrate the high schedulability of SMIX, where static scheduling provides an average 12% performance gain on the Rocket core and dynamic out-of-order scheduling contributes an additional 9.2% speedup on the BOOM core, all while maintaining minimal hardware overhead. Shufan He, Hanmo Wang, Kefa Chen, Xuyin Chen, Xianhua Liu 0001 |
DATE | 5 |
| 2025 | Software-Style Hardware Debugging: A Hardware Generation, Simulation and Debugging FrameworkabstractHardware generation frameworks (HGFs) leverage high-level descriptions to introduce agile methods into hardware design. However, HGFs still rely on traditional RTL verification workflows, limiting design-verification iteration efficiency. This paper presents ComoPy, an HGF supporting host language native simulation and debugging, introducing software debugging techniques to hardware verification. The framework enables fine-grained tracing and breakpoints directly on HDL source lines, and implements software-style hot-reloading for incremental debug-fix iterations. Furthermore, it supports switching between fast low-level RTL simulation and high-level HDL execution. Case studies on the Sodor-clone processor and a database accelerator featuring systolic array show timely debugging responses, with only a 6% performance loss compared to Verilator simulation. Shixuan Chen, Xianhua Liu 0001 |
ICCAD | 4 |
| 2024 | Multi: Reduce Energy Overhead of Criticality-Aware Dynamic Instruction Scheduling for Energy EfficiencyabstractCriticality-aware dynamic instruction scheduling (DIS) focuses on prioritizing the execution of critical instructions, thereby significantly improving performance. However, criticality-aware DIS is energy-intensive. Energy overhead mainly comes from the design of DIS strategy and the multi-port tables used for extracting critical instruction slices. In this study, Multi is proposed, which develops an energy-efficient DIS strategy and an adaptive multi-port table coordinated with the DIS strategy to reduce the energy overhead of criticality-aware DIS. Evaluation on the SPEC CPU2017 benchmark shows that Multi outperforms the state-of-the-art work in terms of core energy consumption with a geometric mean reduction of 5.9%. Moreover, post-layout simulation results demonstrate a 61.2% reduction in energy consumption of the proposed adaptive multi-port table. It also proves the applicability of Multi in ultra-wide processors with unified or distributed schedulers, contributing to enhanced energy efficiency. Honglan Zhan, Xianhua Liu 0001, Xu Cheng 0001 |
ICCD | 5 |
| 2023 | MBAPIS: Multi-Level Behavior Analysis Guided Program Interval Selection for Microarchitecture StudiesabstractUnderstanding program behavior is crucial in computer architecture research, but the growing size of benchmarks makes analyzing and simulating entire programs increasingly challenging. In practice, researchers often select representative program intervals for analysis and testing. These intervals are different sections of continuous execution of a program. SimPoint is a well-known method for selecting representative intervals using hardware-independent information. However, when focusing on a specific microarchitecture study, it is desirable to select intervals that are more relevant to that study. For instance, intervals with more branch mispredictions are more appropriate for branch prediction studies. We refer to these intervals as “tailored intervals” for branch prediction studies. This paper presents a Multi-level Behavior Analysis guided Program Interval Selection (MBAPIS) for selecting tailored intervals. For a given microarchitecture study, the first level of MBAPIS uses hardware performance counters to prioritize selecting the intervals that exhibit clearer microarchitectural characteristics relevant to that study. The second level analyzes the processor performance bottlenecks to further select the intervals where the concerned microarchitecture design more strongly impacts performance. Finally, MBAPIS performs clustering analysis with the basic block information of each interval selected by the first two levels, and selects the representative intervals among them while preserving the diverse software behavior. Additionally, we present a general and extensible interval-replaying design to accurately re-execute selected intervals. The SPEC CPU2006 and CPU2017 benchmarks are used for evaluation. The results demonstrate that MBAPIS can select representative and tailored intervals for two typical microarchi-tecture studies and deliver accurate estimates of the concerned hardware events for all tailored intervals in each benchmark, with an average error rate of less than 1.5%. Moreover, the interval-replaying design effectively restores the hardware behavior of the intervals selected by MBAPIS, with an average relative error rate of annroximately 1.4%. Hongwei Cui, Honglan Zhan, Shuhao Liang, Xianhua Liu 0001, Xu Cheng 0001 |
PACT | 5 |
| 2023 | A Hardware-Software Cooperative Interval-Replaying for FPGA-based Architecture EvaluationabstractOpen-source processors and FPGA provide more real and accurate results of the new microarchitecture design, but the long execution time for running large benchmarks on FPGA boards still hinders researchers. This paper proposes a hardware-software cooperative interval-replaying. It uses simula-tors to create checkpoints for arbitrary program intervals and provides an extensible and portable checkpoint loader to re-execute selected intervals. In addition, this paper extends RISC-VISA and proposes an event-based sampling design to find hot program intervals with more representative microarchitecture characteristics. By using checkpoints in hot regions, researchers can quickly verify the effectiveness of microarchitecture designs on FPGA and alleviate the speed bottleneck of FPGA. The correctness and effectiveness of the checkpoint scheme and the event-based sampling design are evaluated on FPGA. The experimental results show that the solution is effective. Hongwei Cui, Shuhao Liang, Honglan Zhan, Xianhua Liu 0001, Xu Cheng 0001 |
DATE | 7 |
| 2023 | High-Speed and Energy-Efficient Single-Port Content Addressable Memory to Achieve Dual-Port OperationabstractHigh-speed and energy-efficient multi-port content addressable memory (CAM) is very important to modern superscalar processors. In order to overcome the disadvantages of multi-port CAM and improve the performance of searching stage, a high-speed and energy-efficient single-port (SP) CAM is introduced to achieve dual-port (DP) operation. For different bit cell topologies - the traditional 9T CAM cell and 6T SRAM cell, two novel peripheral schemes - CShare and VClamp are proposed. The proposed schemes are verified using all possible corners, a wide range of temperature and detailed Monte-Carlo variation analysis. With 65-nm process and 1.2 V supply, the search delay of CShare and VClamp is 0.55 ns and 0.6 ns, respectively, a reduction of approximately 87% compared to the state-of-the-art works. In addition, compared with the recently proposed IOT BCAM, CShare and VClamp can provide 84.9% and 85.1% energy reduction in the TT corner, respectively. Experimental results in an 8 Kb CAM at 1.2 V supply and across different corners show that the energy efficiency is improved by 45.56% (CShare) and 45.64% (VClamp) on average in comparison with DP CAM. Honglan Zhan, Hongwei Cui, Xianhua Liu 0001, Xu Cheng 0001 |
DATE | 4 |
| 2020 | Efficient arithmetic expression optimization with weighted adjoint matrixabstractPolynomial arithmetic expressions are frequently used in encryption, decryption, digital signal processing and many other embedded applications. Compiler optimization for polynomial expressions can improve the performance of the embedded applications. This article presents an improved compiler optimization method for multiple arithmetic polynomial expressions. Considering the semantic and timing information of arithmetic instructions, the algorithm uses a canonical representation for all expressions, taking consideration of times of execution, architecture feature and control-flow information. It calculates sub-expressions’ weights based on the target architecture description and heuristically choose the sub-expressions. It achieves better instruction level parallelism from the consideration of sub-expressions’ weights, which contains architecture information. Experiment results show that compared to traditional optimization methods, this algorithm further improves the code density and the performance of the generated binary codes. Xianhua Liu 0001, Zixin Guan |
IPCCC | 1 |
| 2020 | FTCLNet: Convolutional LSTM with Fourier Transform for Vulnerability DetectionabstractAs software vulnerabilities become increasingly serious, it is necessary to detect them efficiently and accurately. However, vulnerabilities are diverse and context sensitive. Previous solutions either rely on features defined by experts, or use only recurrent neural networks on code sequence. It is difficult to extract complex features of vulnerabilities in traditional code space. This article proposes a deep convolutional LSTM neural network with Fourier transform for vulnerability detection. The discrete Fourier transform method convert code space into frequency domain, which significantly helps deep models learn remarkable patterns. This article combines convolutional neural network (CNN) with long short term memory (LSTM) network to extract local and global features in frequency domain, and utilize attention mechanism to decide the weight of each element in code space. Besides, this method rewrite the source code and convert them to vectors without guidance from the specified domain knowledge. Experiments on Buffer Error dataset (CWE-119) and Resource Management Error dataset (CWE-399) show that this new method achieves a significantly improved results. Defu Cao, Xianhua Liu 0001 |
TrustCom | 4 |
| 2017 | Content Look-Aside Buffer for Redundancy-Free Virtual Disk I/O and CachingabstractStorage consolidation in a virtualized environment introduces numerous duplications in virtual disks and imposes considerable pressure on disk I/O and caching. In this paper, we present a content look-aside buffer (CLB) approach for simultaneously providing redundancy-free virtual disk I/O and caching. CLB attaches persistent fingerprints to virtual disk blocks, which enables detection of I/O redundancy before disk access. At run time, CLB exploits content pages already present in the guest disk caches to service the redundant reads through page sharing, thus eliminating both redundant I/O requests and redundant disk cache copies. For write requests, CLB uses a group invalidating writeback protocol for updating fingerprints to support crash consistency while minimizing disk write overhead. By implementing and evaluating a CLB prototype on KVM hypervisor, we demonstrate that CLB delivers considerably improved I/O performance with realistic workloads. Our CLB prototype improves the throughput of sequential and random read on duplicate data by 4.1x and 26.2x, respectively. For typical read-intensive workloads, such as booting VM and launching application, CLB's I/O deduplication and cache deduplication eliminates 94.9%--98.5% of read requests and saves 50%--100% cache memory in each VM, respectively. Compared with the QEMU's raw virtual disk format, CLB improves the per-disk VM density by 8x--16x. For mixed read-write workloads, the cost of on-line fingerprint updating offsets the read benefit; nevertheless, CLB substantially improves overall performance. Xianhua Liu 0001, Xu Cheng 0001 |
VEE | 2 |
| 2016 | MFAP: Fair Allocation between fully backlogged and non-fully backlogged applicationsabstractIn this paper, we consider the problem of ensuring fairness in systems serving a mixture of fully backlogged applications, which continuously demand resources, and non-fully backlogged applications. We introduce a fairness metric, called interference fairness, the basic idea underlying which is that the interference caused by application A for another application B should be equal to that caused by B for A. To effectively and efficiently guarantee this fairness metric, we propose Mutual Fair Allocation Policy (MFAP), a simple and powerful resource sharing policy, and show how it guarantees interference fairness between any pair of applications. We also show that MFAP, unlike other viable policies, satisfies several highly desirable properties, including some from game theory, as well as common sense intuitions. As a use case, we implemented MFAP on a disk scheduling framework. The experimental results based on synthetic and real workloads show how our implementation achieved interference fairness and improved non-fully backlogged applications performance. Yan Sui, Dong Tong 0001, Xianhua Liu 0001, Xu Cheng 0001 |
ICCD | 4 |
| 2015 | Exploration of the Relationship Between Just-in-Time Compilation Policy and Number of Cores
Mingkai Huang, Xianhua Liu 0001, Xu Cheng 0001 |
ICA3PP (4) | 2 |
| 2015 | An Energy-Efficient Branch Prediction with Grouped Global HistoryabstractBranch prediction has been playing an increasingly important role in improving the performance and energy efficiency for modern microprocessors. The state-of-the-art branch predictors, such as the perceptron and TAGE predictors, leverage novel prediction algorithms to explore longer branch history for higher prediction accuracy. We observe that as the branch history is becoming longer, the efficiency of global history is degraded by the interference of different branch instructions. In order to mitigate the excessive influence of the branch history interference, we propose the Grouped Global History (GGH) based branch predictor, a lightweight yet efficient branch predictor. Unlike existing branch predictors that make use of a unified global history for prediction, GGH divides the global history into a set of subgroups such that the interference resulted by frequently executed branch instructions could be restricted. With subgroups of global history, GGH also enables us to track even longer effective branch correlation without introducing hardware storage overhead. Our experimental results based on SPEC CINT 2006 workloads demonstrate that our approach can significantly reduce the branch mispredictions per kilo instructions (MPKI) by 4.76 over the baseline perceptron predictor, with a simple control logic extension. Mingkai Huang, Xianhua Liu 0001, Mingxing Tan, Xu Cheng 0001 |
ICPP | 3 |
| 2012 | Energy-efficient branch prediction with Compiler-guided History StackabstractBranch prediction is critical in exploring instruction level parallelism for modern processors. Previous aggressive branch predictors generally require significant amount of hardware storage and complexity to pursue high prediction accuracy. This paper proposes the Compiler-guided History Stack (CHS), an energy-efficient compiler-microarchitecture cooperative technique for branch prediction. The key idea is to track very-long-distance branch correlation using a low-cost compiler-guided history stack. It relies on the compiler to identify branch correlation based on two program substructures: loop and procedure, and feed the information to the predictor by inserting guiding instructions. At runtime, the processor dynamically saves and restores the global history using a low-cost history stack structure according to the compiler-guided information. The modification on the global history enables the predictor to track very-long-distance branch correlation and thus improves the prediction accuracy. We show that CHS can be combined with most of existing branch predictors and it is especially effective with small and simple predictors. Our evaluations show that the CHS technique can reduce the average branch mispredictions by 28.7% over gshare predictor, resulting in average performance improvement of 10.4%. Furthermore, it can also improve those aggressive perceptron, OGEHL and TAGE predictors. Mingxing Tan, Xianhua Liu 0001, Zichao Xie, Dong Tong 0001, Xu Cheng 0001 |
DATE | 2 |
| 2012 | CVP: an energy-efficient indirect branch prediction with compiler-guided value patternabstractIndirect branch prediction is becoming increasingly important in modern high-performance processors. However, previous indirect branch predictors either require a significant amount of hardware storage and complexity, or heavily rely on the expensive manual profiling. Mingxing Tan, Xianhua Liu 0001, Xu Cheng 0001 |
ICS | 2 |
| 2010 | Bit-level optimization for high-level synthesis and FPGA-based accelerationabstractAutomated hardware design from behavior-level abstraction has drawn wide interest in FPGA-based acceleration and configurable computing research field. However, for many high-level programming languages, such as C/C++, the description of bitwise access and computation is not as direct as hardware description languages, and high-level synthesis of algorithmic descriptions may generate suboptimal implementations for bitwise computation-intensive applications. In this paper we introduce a bit-level transformation and optimization approach to assisting high-level synthesis of algorithmic descriptions. We introduce a bit-flow graph to capture bit-value information. Analysis and optimizing transformations can be performed on this representation, and the optimized results are transformed back to the standard data-flow graphs extended with a few instructions representing bitwise access. This allows high-level synthesis tools to automatically generate circuits with higher quality. Experiments show that our algorithm can reduce slice usage by 29.8% on average for a set of real-life benchmarks on Xilinx Virtex-4 FPGAs. In the meantime, the clock period is reduced by 13.6% on average, with an 11.4% latency reduction. Jiyu Zhang, Zhiru Zhang, Mingxing Tan, Xianhua Liu 0001, Xu Cheng 0001, Jason Cong |
FPGA | 5 |
| 2010 | Research Progress of UniCore CPUs and PKUnity SoCs
Xu Cheng 0001, Xiaoyin Wang, Junlin Lu, Jiangfang Yi, Dong Tong 0001, Xuetao Guan, Xianhua Liu 0001, Yi Feng 0003 |
J. Comput. Sci. Technol. | 8 |