EDBT 2026 Demo / reviewers in the wild / expert
Yanqiang Liu
dblp:194/3970
· DBLP profile ↗
8ranked-venue papers
4as first author
3since 2021 · last 2021
0000-0002-8343-7923ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | MEGATRON: Software-Managed Device TLB for Shared-Memory FPGA VirtualizationabstractFPGAs are being virtualized to improve resource utilization in data centers. Memory access performance is essential to FPGA hypervisors for shared-memory FPGA platform, where accelerators access memory spontaneously. DMA remapping with IOMMU provides a handy solution; however, fixed IOMMU can not benefit from the reconfigurability of FPGAs. In this work, we propose MEGATRON, a hybrid address translation service consisting of a hardware TLB and a software page table walker. By integrating MEGATRON into an existing FPGA hypervisor, we conduct a comprehensive analysis of link performance of a multi-link CPU-FPGA platform, and demonstrate the competitiveness of the customizable translation service. Yanqiang Liu, Jiacheng Ma 0001, Zhengjun Zhang, Linsheng Li, Zhengwei Qi, Haibing Guan |
DAC | 1 |
| 2021 | Performance Analysis of Open-Source Hypervisors for Automotive SystemsabstractNowadays, automotive products are intelligence intensive and thus inevitably handle multiple functionalities under the current high-speed networking environment. The embedded virtualization has high potentials in the automotive industry, thanks to its advantages in function integration, resource utilization, and security. The invention of ARM virtualization extensions has made it possible to run open-source hypervisors, such as Xen and KVM, for embedded applications. Nevertheless, there is little work to investigate the performance of these hypervisors on automotive platforms. This paper presents a detailed analysis of different types of open-source hypervisors that can be applied in the ARM platform. We carry out the virtualization performance experiment from the perspectives of CPU, memory, file I/O, and some OS operation performance on Xen and Jailhouse. A series of microbenchmark programs have been designed, specifically to evaluate the real-time performance of various hypervisors and the relevant overhead. Compared with Xen, Jailhouse has better latency performance, stable latency, and little interference jitter. The performance experiment results help us summarize the advantages and disadvantages of these hypervisors in automotive applications. Zhengjun Zhang, Yanqiang Liu, Jiangtao Chen, Zhengwei Qi, Huai Liu |
ICPADS | 2 |
| 2021 | Efficient shuffle management for DAG computing frameworks based on the FRQ model
Chunghsuan Wu, Zhouwang Fu, Tao Song 0003, Yanqiang Liu, Zhengwei Qi, Haibing Guan |
J. Parallel Distributed Comput. | 5 |
| 2020 | A Hypervisor for Shared-Memory FPGA PlatformsabstractCloud providers widely deploy FPGAs as application-specific accelerators for customer use. These providers seek to multiplex their FPGAs among customers via virtualization, thereby reducing running costs. Unfortunately, most virtualization support is confined to FPGAs that expose a restrictive, host-centric programming model in which accelerators cannot issue direct memory accesses (DMAs). The host-centric model incurs high runtime overhead for workloads that exhibit pointer chasing. Thus, FPGAs are beginning to support a shared-memory programming model in which accelerators can issue DMAs. However, virtualization support for shared-memory FPGAs is limited. This paper presents Optimus, the first hypervisor that supports scalable shared-memory FPGA virtualization. Optimus offers both spatial multiplexing and temporal multiplexing to provide efficient and flexible sharing of each accelerator on an FPGA. To share the FPGA-CPU interconnect at a high clock frequency, Optimus implements a multiplexer tree. To isolate each guest's address space, Optimus introduces the technique of page table slicing as a hardware-software co-design. To support preemptive temporal multiplexing, Optimus provides an accelerator preemption interface. We show that Optimus supports eight physical accelerators on a single FPGA and improves the aggregate throughput of twelve real-world benchmarks by 1.98x-7x. Jiacheng Ma 0001, Gefei Zuo, Kevin Loughlin, Xiaohe Cheng, Yanqiang Liu, Abel Mulugeta Eneyew, Zhengwei Qi, Baris Kasikci |
ASPLOS | 5 |
| 2020 | OPS: Optimized Shuffle Management System for Apache SparkabstractIn recent years, distributed computing frameworks, such as Hadoop MapReduce and Spark, are widely used for big data processing. With the explosive growth of the amount of data, companies tend to store intermediate data of the shuffle phase on disk instead of memory. Therefore, intensive network and disk I/O are both involved in the shuffle phase. To optimize the overhead of the shuffle phase, we propose OPS, an open-source distributed computing shuffle management system based on Spark, which provides an independent shuffle service for Spark. By using early-merge and early-shuffle strategy, OPS alleviates the I/O overhead in the shuffle phase and efficiently schedules the I/O and computing resources. OPS also proposes a slot-based scheduling algorithm to predict and calculate the optimal scheduling result of the reduce task. Besides, OPS provides a taint-redo strategy to ensure the fault tolerance of computing jobs. We evaluated the performance of OPS on a 100-node Amazon AWS EC2 cluster. Overall, OPS optimizes the overhead of shuffle by nearly 50%. In the test cases of HiBench, OPS improves end-to-end completion time by nearly 30% on average. Yuchen Cheng, Chunghsuan Wu, Yanqiang Liu, Hong Xu 0001, Zhengwei Qi |
ICPP | 3 |
| 2020 | TimelyRep: Timing deterministic replay for Android web applicationsabstractSummary With the constantly growing and changing requirements of app users, web techniques are used in mobile application development for better cross‐platform compatibility and online update. As the embedded web contents gain complexity, debugging web apps become a critical demand. Web replay tools can record program inputs and reproduce the same execution for debugging and performance tuning. However, traditional replay approaches are largely intended for apps with desktop interaction methods (keyboard, mouse) and require modification to the browser, which limits their applicability in mobile platforms. In this paper, we develop TimelyRep, which provides deterministic record‐and‐replay as a software library, running on commodity Android. TimelyRep can be used for app development with unmodified Android devices and for production to collect faulty execution from users. Also, we propose an efficient replay timing control mechanism and achieve higher timing precision as facing higher event rate on touchscreen devices. TimelyRep also supports cross‐device replay and can replay logged event traces on different devices, which is useful for developers to reproduce user inputs on their own devices. We evaluate TimelyRep with real‐world web applications. The results show that TimelyRep is useful for recreating program bugs and maintaining low delays for touch‐intensive web games. Yanqiang Liu, Fangge Yan, Mingyuan Xia 0001, Zhengwei Qi, Xue (Steve) Liu |
Softw. Test. Verification Reliab. | 1 |
| 2019 | A scala based framework for developing acceleration systems with FPGAs
Yanqiang Liu, Yao Li 0004, Zhengwei Qi, Haibing Guan |
J. Syst. Archit. | 1 |
| 2017 | Scala Based FPGA Design Flow (Abstract Only)
Yanqiang Liu, Yao Li 0004, Weilun Xiong, Meng Lai, Zhengwei Qi, Haibing Guan |
FPGA | 1 |