VLDB 2026 Research / reviewers in the wild / expert
Xingyue Qian
dblp:319/2811
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2025
0009-0003-6162-2778ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MASIM: An Energy-Efficient Multi-Array Scheduler for SIMD Logic-in-Memory ArchitecturesabstractSingle instruction, multiple data (SIMD) is a popular design style of logic-in-memory (LiM) architectures, which enables memory arrays to perform logic operations to achieve low energy consumption and high throughput. To implement a target function on the data stored in memory, the function is first transformed into a netlist of the supported logic operations by logic synthesis. Then, a scheduler transforms the netlist into an instruction sequence given to the architecture, where an instruction either performs a logic operation in the netlist on memory rows within a single array or copies the data from one array to another. Most existing schedulers focus on optimizing the execution sequence of the operations to minimize the number of memory rows needed, neglecting the energy-consuming copy instructions that cannot be avoided when working with arrays with limited sizes. In this work, we focus on reducing the number of copy instructions to decrease the total energy consumption. We propose MASIM, a multi-array scheduler for SIMD logic-in-memory architectures. It consists of a priority-based scheduling algorithm and an iterative improvement process. Compared to the best existing scheduler, MASIM reduces the number of copy instructions by 63.2% on average, which leads to a 28.0% reduction in energy. The experiment also shows that MASIM can be applied to various SIMD LiM architectures, showing its wide applicability. Xingyue Qian, Chen Nie, Zhezhi He, Weikang Qian |
ICCAD | 1 |
| 2025 | A Recursive Partition-Based In-Memory SIMD Computation Scheduler for Memory Footprint MinimizationabstractIn-memory computing (IMC) is a technique that enables memory to perform computation so that data transfer between processor and memory can be reduced, improving energy efficiency. A popular IMC design style is based on the single-instruction-multiple-data (SIMD) concept. The SIMD IMC can implement a high-level function by two steps: 1) synthesis and 2) scheduling. The former converts the high-level function into a netlist of the supported primitive logic operations, while the latter determines the execution sequence of the operations. To fully exploit the advantage of SIMD IMC, it is crucial to find a schedule for the given netlist with less memory usage, known as memory footprint (MF). In this work, we first propose an optimal scheduler that can minimize the MF for small netlists. It is at least$8\times $faster than the state-of-the-art optimal method. For large netlists, we propose a recursive partition-based scheduler consisting of a scheduling-friendly bipartition algorithm and our optimal scheduler. Compared to four state-of-the-art heuristic methods, ours reduces the MF by 54.7%, 48.9%, 44.0%, and 25.5%, respectively, under the same runtime. Our experiments also demonstrate that our scheduler achieves good end-to-end performance when applied to various IMC architectures. The code of our scheduler is made open-source. Xingyue Qian, Chenyang Lv, Zhezhi He, Weikang Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Efficient Approximate Decomposition Solver using Ising ModelabstractComputing with memory is an energy-efficient computing approach. It pre-computes a function and stores its values in a lookup table (LUT), which can be retrieved at runtime. Approximate Boolean decomposition reduces the LUT size for implementing complex functions, but it takes a long time to find a decomposition with a minimal error. In this work, to address this issue, we propose an efficient Ising model-based approximate Boolean decomposition solver. First, a new column-based approximate disjoint decomposition method is proposed to fit the Ising model. Then, it is adapted to the Ising model-based optimization solver. Moreover, two improvement techniques are developed for an efficient search of the approximate disjoint decomposition when using simulated bifurcation to solve the Ising model. Experimental results show that compared to the state-of-the-art work, our approach achieves a 11% smaller mean error distance with an average 1.16× speedup when approximately decomposing 16-input Boolean functions. Weihua Xiao, Xingyue Qian, Jie Han 0001, Weikang Qian |
DAC | 3 |
| 2024 | An Efficient Logic Operation Scheduler for Minimizing Memory Footprint of In-Memory SIMD ComputationabstractMany in-memory computing (IMC) designs based on single instruction multiple data (SIMD) concept have been proposed in recent years to perform primitive logic operations within memory, for improving energy efficiency. To fully exploit the advantage of SIMD IMC, it is crucial to identify an optimized schedule for the operations with less intermediate memory usage, known as memory footprint (MF). In this work, we implement a recursive partition-based scheduler which consists of our scheduler-friendly partition algorithm and a modified optimal scheduler. Compared to three state-of-the-art heuristic strategies, ours can reduce MF by 56.9%, 46.0%, and 31.9%, respectively. Xingyue Qian, Zhezhi He, Weikang Qian |
DATE | 1 |
| 2023 | High-accuracy Low-power Reconfigurable Architectures for Decomposition-based Approximate Lookup TableabstractStoring pre-computed results of frequently-used functions into lookup table (LUT) is a popular way to improve energy efficiency, but its advantage diminishes as the number of input bits increases. A recent work shows that by decomposing the target function approximately, the total LUT entries can be dramatically reduced, leading to significant energy saving. However, its heuristic approximate decomposition algorithm leads to sub-optimal approximation quality. Also, its rigid hardware architecture only supports disjoint decomposition and may have unnecessary extra power consumption sometimes. To address these issues, we develop a novel approximate decomposition algorithm based on beam search and simulated annealing, which can reduce 11.1% approximation error. We also propose a non-disjoint approximate decomposition method and two reconfigurable architectures. The first has 10.4% less error using 19.2% less energy and the second has 23.0% less error with same energy consumption compared to the state-of-the-art design. Xingyue Qian, Chang Meng, Xiaolong Shen, Junfeng Zhao 0003, Leibin Ni, Weikang Qian |
DATE | 1 |
| 2022 | Exploiting Scheduling Information for Efficient High-Level Synthesis Design Space ExplorationabstractHigh-level synthesis (HLS) automatically transforms highlevel programming language into RTL design. It is widely used to program FPGAs as accelerators. HLS tools have many knobs that can be controlled by users to produce designs with different area-latency trade-offs. Xingyue Qian, Lijian Bian, Weikang Qian |
FCCM | 1 |
| 2022 | Scheduling Information-Guided Efficient High-Level Synthesis Design Space ExplorationabstractHigh-level synthesis (HLS) transforms designs specified by high-level programming language into RTL designs. In order to get the optimal designs, many design space exploration (DSE) methods are proposed. However, most of them consider the HLS tool as a black box, ignoring crucial information from the synthesis process, particularly the scheduling step. In this work, we propose to extract some useful information from scheduling to guide the DSE and develop a genetic algorithm (GA)-based DSE method based on our in-house HLS tool. The experimental results show that our method can obtain more Pareto-optimal points than the counterpart without using the scheduling information. It also outperforms a traditional GA-based HLS DSE method by using only a quarter of the total run time. For a large benchmark, our method finds 95.7% Pareto-optimal designs by visiting only 0.18% total promising design points. Xingyue Qian, Lijian Bian, Weikang Qian |
ICCD | 1 |
| 2022 | GiantVM: A Novel Distributed Hypervisor for Resource Aggregation with DSM-aware OptimizationsabstractWe present GiantVM, 1 an open-source distributed hypervisor that provides the many-to-one virtualization to aggregate resources from multiple physical machines. We propose techniques to enable distributed CPU and I/O virtualization and distributed shared memory (DSM) to achieve memory aggregation. GiantVM is implemented based on the state-of-the-art type-II hypervisor QEMU-KVM, and it can currently host conventional OSes such as Linux. (1) We identify the performance bottleneck of GiantVM to be DSM, through a top-down performance analysis. Although GiantVM offers great opportunities for CPU-intensive applications to enjoy the aggregated CPU resources, memory-intensive applications could suffer from cross-node page sharing, which requires frequent DSM involvement and leads to performance collapse. We design the guest-level thread scheduler, DaS (DSM-aware Scheduler), to overcome the bottleneck. When benchmarking with NAS Parallel Benchmarks, the DaS could achieve a performance boost of up to 3.5×, compared to the default Linux kernel scheduler. (2) While evaluating DaS, we observe the advantage of GiantVM as a resource reallocation facility. Thanks to the SSI abstraction of GiantVM, migration could be done by guest-level scheduling. DSM allows standby pages in the migration destination, which need not be transferred through the network. The saved network bandwidth is 68% on average, compared to VM live migration. Resource reallocation with GiantVM increases the overall CPU utilization by 14.3% in a co-location experiment. Xingguo Jia, Boshi Yu, Xingyue Qian, Zhengwei Qi, Haibing Guan |
ACM Trans. Archit. Code Optim. | 4 |