EDBT 2026 Demo / reviewers in the wild / expert
Qixiao Liu
dblp:149/8095
· DBLP profile ↗
8ranked-venue papers
7as first author
1since 2021 · last 2026
0000-0002-8196-7584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 40% Energy-efficient computing · 28% Memory systems · 27% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › concurrency control
distributed lock management |
1.0 | 1 | 2026 | CXLock: Efficient and Scalable Lock Management for CXL-Enabled Distributed Systems · IEEE Trans. Computers 2026 |
Memory systems › memory interconnect
CXL |
0.3 | 1 | 2026 | CXLock: Efficient and Scalable Lock Management for CXL-Enabled Distributed Systems · IEEE Trans. Computers 2026 |
Memory systems › memory management
memory sharing |
0.3 | 1 | 2026 | CXLock: Efficient and Scalable Lock Management for CXL-Enabled Distributed Systems · IEEE Trans. Computers 2026 |
Energy-efficient computing
energy-aware scheduling |
0.3 | 2 | 2016 | Sensible Energy Accounting with Abstract Metering for Multicore Systems · ACM Trans. Archit. Code Optim. 2016 Hardware support for accurate per-task energy metering in multicore systems · ACM Trans. Archit. Code Optim. 2013 |
Energy-efficient computing
energy accounting |
0.2 | 1 | 2016 | Sensible Energy Accounting with Abstract Metering for Multicore Systems · ACM Trans. Archit. Code Optim. 2016 |
Energy-efficient computing
energy measurement |
0.2 | 1 | 2013 | Hardware support for accurate per-task energy metering in multicore systems · ACM Trans. Archit. Code Optim. 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2016 | Sensible Energy Accounting with Abstract Metering for Multicore Systems · ACM Trans. Archit. Code Optim. 2016 |
Memory systems › cache
shared last-level cache |
0.1 | 1 | 2016 | Sensible Energy Accounting with Abstract Metering for Multicore Systems · ACM Trans. Archit. Code Optim. 2016 |
Methods — techniques the papers use, named apart from their topics
software-based coherency · 1.0manager-side arbitration · 1.0client-side queuing · 1.0hardware energy metering · 0.4abstract metering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CXLock: Efficient and Scalable Lock Management for CXL-Enabled Distributed SystemsabstractEfficient lock management is critical in distributed systems with shared resources, especially in the context of enhanced computational power and reduced processing time. While numerous studies have leveraged Remote Direct Memory Access (RDMA) for distributed locking, these network-based approaches suffer from high network latency and jitter. The emergence of the Compute Express Link (CXL) protocol provides high-speed, low-latency connections between host processors and memory devices, presenting a compelling alternative for the lock system design. This paper introduces CXLock, the industry’s first distributed lock management system based on CXL 2.0 switches. CXLock decouples lock management into client-side queuing and manager-side arbitration, resolving race conditions of multiple clients without using remote atomics. To handle the lack of hardware-enforced coherency in CXL 2.0, CXLock employs a software-based coherency model that enables memory sharing across multiple hosts with cacheable and write-back attributes. CXLock is implemented and evaluated on a real rack-scale hardware platform containing a CXL switch and multiple host nodes. The results show that CXLock delivers an average throughput improvement of 3.7× over the RDMA-based solutions. Moreover, CXLock reduces the lock grant latency by up to 76% in both contention-free and contended scenarios. Yudi Qiu, Qixiao Liu, Wenpu Hu, Xinjun Yang, Yingqiang Zhang, Hao Chen 0080, Zipeng Ouyang, Yuemin Wu |
IEEE Trans. Computers | 4 |
| 2019 | MiC: Multi-level Characterization and Optimization of GPGPU KernelsabstractGraphics processing units (GPUs) 1 have enjoyed increasing popularity in recent years, which benefits from, for example, general-purpose GPU (GPGPU) for parallel programs and new computing paradigms, such as the Internet of Things (IoT). GPUs hold great potential in providing effective solutions for big data analytics while the demands for processing large quantities of data in real time are also increasing. However, the pervasive presence of GPUs on mobile devices presents great challenges for GPGPU, mainly because GPGPU integrates a large amount of processor arrays and concurrent executing threads (up to hundreds of thousands). In particular, the root causes of performance loss in a GPGPU program can not be revealed in detail by current approaches. In this article, we propose MiC (Multi-level Characterization), a framework that comprehensively characterizes GPGPU kernels at the instruction, Basic Block (BBL), and thread levels. Specifically, we devise Instruction Vectors (IV) and Basic Blocks Vectors (BBV), a Thread Similarity Matrix (TSM), and a Divergence Flow Statistics Graph (DFSG) to profile information in each level. We use MiC to provide insights into GPGPU kernels through the characterizations of 34 kernels from popular GPGPU benchmark suites such as Compute Unified Device Architecture (CUDA) Software Development Kit (SDK), Rodinia, and Parboil. In comparison with Central Processing Unit (CPU) workloads, we conclude the key findings as follows: (1) There are comparable Instruction-Level Parallelism (ILP); (2) The BBL count is significantly smaller than CPU workloads—only 22.8 on average; (3) The dynamic instruction count per thread varies from dozens to tens of thousands and it is extremely small compared to CPU benchmarks; (4) The Pareto principle (also called 90/10 rule) does not apply to GPGPU kernels while it pervasively exists in CPU programs; (5) The loop patterns are dramatically different from those in CPU workloads; (6) The branch ratio is lower than that of CPU programs but higher than pure GPU workloads. In addition, we have also shown how TSM and DFSG are used to characterize the branch divergence in a visual way, to enable the analysis of thread behavior in GPGPU programs. In addition, we show an optimization case for a GPGPU kernel from the bottleneck identified through its characterization result, which improves 16.8% performance. Qixiao Liu, Zhibin Yu 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2018 | The Elasticity and Plasticity in Semi-Containerized Co-locating Cloud Workload: a View from Alibaba TraceabstractCloud computing with large-scale datacenters provides great convenience and cost-efficiency for end users. However, the resource utilization of cloud datacenters is very low, which wastes a huge amount of infrastructure investment and energy to operate. To improve resource utilization, cloud providers usually co-locate workloads of different types on shared resources. However, resource sharing makes the quality of service (QoS) unguaranteed. In fact, improving resource utilization (IRU) and guaranteeing QoS at the same time in cloud has been a dilemma which we name an IRU-QoS curse. To tackle this issue, characterizing the workloads from real production cloud computing platforms is extremely important. Qixiao Liu, Zhibin Yu 0001 |
SoCC | 1 |
| 2017 | SEDEA: A Sensible Approach to Account DRAM Energy in Multicore SystemsabstractAs the energy cost in todays computing systems keeps increasing, measuring the energy becomes crucial in many scenarios. For instance, due to the fact that the operational cost of datacenters largely depends on the energy consumed by the applications executed, end users should be charged for the energy consumed, which requires a fair and consistent energy measuring approach. However, the use of multicore system complicates per-task energy measurement as the increased Thread Level Parallelism (TLP) allows several tasks to run simultaneously sharing resources. Therefore, the energy usage of each task is hard to determine due to interleaved activities and mutual interferences. To this end, Per-Task Energy Metering (PTEM) has been proposed to measure the actual energy of each task based on their resource utilization in a workload. However, the measured energy depends on the interferences from co-running tasks sharing the resources, and thus fails to provide the consistency across executions. Therefore, Sensible Energy Accounting (SEA) has been proposed to deliver an abstraction of the energy consumption based on a particular allocation of resources to a task. In this work we provide a realization of SEA for the DRAM memory system, SEDEA, where we account a task for the DRAM energy it would have consumed when running in isolation with a fraction of the on-chip shared cache. SEDEA is a mechanism to sensibly account for the DRAM energy of a task based on predicting its memory behavior. Our results show that SEDEA provides accurate estimates, yet with low-cost, beating existing per-task energy models, which do not target accounting energy in multicore system. We also provide a use case showing that SEDEA can be used to guide shared cache and memory bank partition schemes to save energy. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
SBAC-PAD | 1 |
| 2016 | Sensible Energy Accounting with Abstract Metering for Multicore SystemsabstractChip multicore processors (CMPs) are the preferred processing platform across different domains such as data centers, real-time systems, and mobile devices. In all those domains, energy is arguably the most expensive resource in a computing system. Accurately quantifying energy usage in a multicore environment presents a challenge as well as an opportunity for optimization. Standard metering approaches are not capable of delivering consistent results with shared resources, since the same task with the same inputs may have different energy consumption based on the mix of co-running tasks. However, it is reasonable for data-center operators to charge on the basis of estimated energy usage rather than time since energy is more correlated with their actual cost. This article introduces the concept of Sensible Energy Accounting (SEA). For a task running in a multicore system, SEA accurately estimates the energy the task would have consumed running in isolation with a given fraction of the CMP shared resources. We explain the potential benefits of SEA in different domains and describe two hardware techniques to implement it for a shared last-level cache and on-core resources in SMT processors. Moreover, with SEA, an energy-aware scheduler can find a highly efficient on-chip resource assignment, reducing by up to 39% the total processor energy for a 4-core system. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Daniel A. Jiménez, Mateo Valero |
ACM Trans. Archit. Code Optim. | 1 |
| 2016 | DReAM: An Approach to Estimate per-Task DRAM Energy in Multicore SystemsabstractAccurate per-task energy estimation in multicore systems would allow performing per-task energy-aware task scheduling and energy-aware billing in data centers, among other applications. Per-task energy estimation is challenged by the interaction between tasks in shared resources, which impacts tasks’ energy consumption in uncontrolled ways. Some accurate mechanisms have been devised recently to estimate per-task energy consumed on-chip in multicores, but there is a lack of such mechanisms for DRAM memories. This article makes the case for accurate per-task DRAM energy metering in multicores, which opens new paths to energy/performance optimizations. In particular, the contributions of this article are (i) an ideal per-task energy metering model for DRAM memories; (ii) DReAM, an accurate yet low cost implementation of the ideal model (less than 5% accuracy error when 16 tasks share memory); and (iii) a comparison with standard methods (even distribution and access-count based) proving that DReAM is much more accurate than these other methods. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2014 | DReAM: Per-Task DRAM Energy Metering in Multicore Systems
Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
Euro-Par | 1 |
| 2013 | Hardware support for accurate per-task energy metering in multicore systemsabstractAccurately determining the energy consumed by each task in a system will become of prominent importance in future multicore-based systems because it offers several benefits, including (i) better application energy/performance optimizations, (ii) improved energy-aware task scheduling, and (iii) energy-aware billing in data centers. Unfortunately, existing methods for energy metering in multicores fail to provide accurate energy estimates for each task when several tasks run simultaneously. This article makes a case for accurate Per-Task Energy Metering (PTEM) based on tracking the resource utilization and occupancy of each task. Different hardware implementations with different trade-offs between energy prediction accuracy and hardware-implementation complexity are proposed. Our evaluation shows that the energy consumed in a multicore by each task can be accurately measured. For a 32-core, 2-way, simultaneous multithreaded core setup, PTEM reduces the average accuracy error from more than 12% when our hardware support is not used to less than 4% when it is used. The maximum observed error for any task in the workload we used reduces from 58% down to 9% when our hardware support is used. Qixiao Liu, Miquel Moretó, Víctor Jiménez, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Archit. Code Optim. | 1 |