EDBT 2026 Demo / reviewers in the wild / expert
Yuqi Xue
dblp:229/0374
· DBLP profile ↗
18ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Integrative Prompt Learning for Continual Defect Detection in Industrial Scenarios
Jiayuan Xie, Yuqi Xue, Yi Cai 0001, Qing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-designabstractThe CXL-based solid-state drive (CXL-SSD) provides a promising approach towards scaling the main memory capacity at low cost. However, the CXL-based SSD faces performance challenges due to the long flash access latency and unpredictable events such as garbage collection in the SSD device, stalling the host processor and wasting compute cycles. Although the CXL interface enables the byte-granular data access to the SSD, accessing flash chips is still at page granularity due to physical limitations. The mismatch of access granularity causes significant unnecessary I/O traffic to flash chips, worsening the suboptimal end-to-end data access performance. In this paper, we present SkyByte, an efficient CXL-based SSD that employs a holistic approach to address the aforementioned challenges by co-designing the host operating system (OS) and SSD controller. To alleviate the long memory stall when accessing the CXL-SSD, SkyByte revisits the OS context switch mechanism and enables opportunistic context switches upon the detection of long access delays. To accommodate byte-granular data accesses, SkyByte architects the internal DRAM of the SSD controller into a cacheline-level write $\log$ and a page-level data cache, and enables data coalescing upon log cleaning to reduce the I/O traffic to flash chips. SkyByte also employs optimization techniques that include adaptive page migration for exploring the performance benefits of fast host memory by promoting hot pages in CXL-SSD to the host. We implement SkyByte with a CXL-SSD simulator and evaluate its efficiency with various data-intensive applications. Our experiments show that SkyByte outperforms current CXL-based SSD by $6.11 \times$, and reduces the I/O traffic to flash chips by $\mathbf{2 3. 0 8} \times$ on average. SkyByte also reaches $\mathbf{7 5 \%}$ of the performance of the ideal case that assumes unlimited DRAM capacity in the host, which offers an attractive cost-effective solution. Yuqi Xue, Yirui Eric Zhou, Shaobo Li 0005, Jian Huang 0006 |
HPCA | 2 |
| 2025 | Elk: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
Yuqi Xue, Noelle Crawford, Jilong Xue, Jian Huang 0006 |
MICRO | 2 |
| 2025 | ReGate: Enabling Power Gating in Neural Processing UnitsabstractThe energy efficiency of neural processing units (NPU) plays a critical role in developing sustainable data centers.Our study with different generations of NPU chips reveals that 30%-72% of their energy consumption is contributed by static power dissipation, due to the lack of power management support in modern NPU chips.In this paper, we present ReGate, which enables fine-grained power-gating of each hardware component in NPU chips with hardware/software co-design.Unlike conventional power-gating techniques for generic processors, enabling power-gating in NPUs faces unique challenges due to the fundamental difference in hardware architecture and program execution model.To address these challenges, we carefully investigate the power-gating opportunities in each component of NPU chips and decide the best-fit power management scheme (i.e., hardware-vs.software-managed power gating).Specifically, for systolic arrays (SAs) that have deterministic execution patterns, ReGate enables cycle-level power gating at the granularity of processing elements (PEs) following the inherent dataflow execution in SAs.For inter-chip interconnect (ICI) and HBM controllers that have long idle intervals, ReGate employs a lightweight hardware-based idle-detection mechanism.For vector units and SRAM whose idle periods vary significantly depending on workload patterns, ReGate extends the NPU ISA and allows software (e.g., compilers) to manage the power gating.With implementation on a production-level NPU simulator, we show that ReGate can reduce the energy consumption of NPU chips by up to 32.8% (15.5% on average), with negligible impact on AI workload performance.The hardware implementation of power-gating logic introduces less than 3.3% overhead in NPU chips. Yuqi Xue, Jian Huang 0006 |
MICRO | 1 |
| 2025 | Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMsabstractModern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI). Xingang Guo, Xiangyi Kong, Yilan Jiang, Xiayu Zhao, Zhihua Gong, Daixuan Li, Tianle Sang, Beixiao Zhu, Gregory Jun, Yingbing Huang, Yuqi Xue, Rahul Dev Kundu, Qi Jian Lim, Luke Alexander Granger, Mohamed Badr Younis, Darioush Keivan, Nippun Sabharwal, Shreyanka Sinha, Prakhar Agarwal, Kojo Vandyck, Hanlin Mai, Aditya Venkatesh, Ayush Barik, Jiankun Yang, Chongying Yue, Jingjie He, Licheng Xu, Liujun Xu, Rushabh Shetty, Ziheng Guo, Dahui Song, Manvi Jha, Weijie Liang, Weiman Yan, Bryan Zhang, Sahil Bhandary Karnoor, Rutva Pandya, Xinyi Gong, Mithesh Ballae Ganesh, Feize Shi, Ruiling Xu, Yanfeng Ouyang, Lianhui Qin, Elyse Rosenbaum, Corey Snyder, Peter J. Seiler, Geir E. Dullerud, Xiaojia Shelly Zhang, Zuofu Cheng, Pavan Kumar Hanumolu, Mayank Kulkarni, Mahdi Namazifar, Bin Hu 0002 |
NeurIPS | 14 |
| 2025 | Managing Scalable Direct Storage Accesses for GPUs with GoFSabstractAs we shift from CPU-centric computing to GPU-accelerated computing for supporting intelligent data processing at scale, the storage bottleneck has been exacerbated. To bypass the host CPUand alleviate unnecessary data movements, modern GPUs enable direct storage access to SSDs (i.e., GPUDirect Storage). However, current GPUDirect Storage solutions still rely on the host file system to manage the storage device, direct storage accesses are still bottlenecked by the host. Shaobo Li 0005, Yirui Eric Zhou, Yuqi Xue, Yuan Xu 0020, Jian Huang 0006 |
SOSP | 3 |
| 2025 | Deep Learning based Unified CSI Feedback for TDD and FDD Massive MIMO SystemsabstractAccurate channel state information (CSI) is essential for maximizing massive MIMO throughput. While time-division duplexing (TDD) systems exploit channel reciprocity for easier downlink CSI (DL-CSI) acquisition from uplink CSI (UL-CSI), moderate channel differences still arise due to environmental factors and mobility. Frequency-division duplexing (FDD) systems also exhibit some channel reciprocity, allowing both TDD and FDD to reduce CSI feedback overhead. For devices supporting both modes, a unified CSI feedback framework can lower cost and energy consumption. This paper proposes a deep learning method based on a cascaded convolutional long short-term memory network, consisting of a main prediction module and a TDD preprocessing module. By leveraging the correlation between TDD and FDD channels under the same spatiotemporal conditions, the proposed method enhances TDD prediction performance and achieves unified CSI feedback for both TDD and FDD systems. Experimental results show that the scheme improves the squared generalized cosine similarity by 5% in TDD mode and by 10% in FDD mode. Mingyu Jia, Shaoli Kang, Yuqi Xue, Yekang Wang, Xianjun Yang |
VTC2025-Fall | 3 |
| 2025 | AI-Enhanced CSI Feedback via Exploiting Multi-user Shared Information in mMIMO SystemsabstractIn massive multiple-input multiple-output (mMIMO) frequency division duplex (FDD) systems, user equipment (UEs) must feed back the channel state information (CSI) to the base station (BS) over the uplink to enhance spectral efficiency. However, in the millimeter wave (mmWave) or higher frequency bands, the significant increase in the number of transmit antennas results in prohibitive feedback overhead. To tackle this issue, we propose a deep learning based multi-user CSI compression framework called CoTransNet. For users in a neighboring area, CSI can be decomposed into two components: one is shared by all users in the area and another is unique to each individual UE. By leveraging shared information, CoTransNet extracts correlations among channels and eliminates redundant feedback information. Experimental results indicate that employing the CoTransNet architecture leads to an average improvement of approximately 4.98% in squared generalized cosine similarity (SGCS) across diverse feedback budgets and a reduction of about 3.54 dB in normalized mean-squared error (NMSE). Yekang Wang, Shanzhi Chen, Shaoli Kang, Xianjun Yang, Mingyu Jia, Yuqi Xue |
VTC2025-Fall | 6 |
| 2025 | Deep Learning-Based Joint Prediction and Compression with Optimal Prediction Interval for High-speed CSI FeedbackabstractIn 6G-oriented ultra-large-scale multiple-input multiple-output (MIMO) systems, accurate channel state information (CSI) is crucial in achieving extremely high data rates. To address the challenge of poor timeliness in CSI feedback for high-speed mobile scenarios, this paper proposes a joint CSI prediction and compression scheme for CSI feedback. Traditional methods improve timeliness by increasing the feedback frequency but suffer from additional overhead. The core innovation of this proposal is the integration of CSI prediction and compression functions on the user equipment (UE) side using deep learning methods. Considering channel estimation errors, the paper investigates the optimal prediction time interval to maximize the system’s channel capacity. ConvLSTM network is used to predict the future CSI when PDSCH is transmitted with historical CSIs, followed by a Transformer model that performs adaptive compression under dynamic channel conditions. Experimental results demonstrate that the proposed method achieves a 7.06% performance improvement over the eType II codebook-based feedback scheme. It ensures both timeliness and accuracy of CSI feedback in high-speed mobility scenarios while maintaining low feedback overhead. Yuqi Xue, Shaoli Kang, Xianjun Yang, Mingyu Jia, Yekang Wang |
VTC2025-Fall | 1 |
| 2025 | Optimization of multi-prior 3D human reconstruction methods based on single-view images
Yuqi Xue, Shida Gao |
Comput. Graph. | 3 |
| 2024 | Hardware-Assisted Virtualization of Neural Processing Units for Cloud PlatformsabstractCloud platforms today have been deploying hardware accelerators like neural processing units (NPUs) for powering machine learning (ML) inference services. To maximize the resource utilization while ensuring reasonable quality of service, a natural approach is to virtualize NPUs for efficient resource sharing for multi-tenant ML services. However, virtualizing NPUs for modern cloud platforms is not easy. This is not only due to the lack of system abstraction support for NPU hardware, but also due to the lack of architectural and ISA support for enabling fine-grained dynamic operator scheduling for virtualized NPUs. We present Neu10, a holistic NPU virtualization framework. We investigate virtualization techniques for NPUs across the entire software and hardware stack. Neul0 consists of (1) a flexible NPU abstraction called vNPU, which enables fine-grained virtualization of the heterogeneous compute units in a physical NPU (pNPU); (2) a vNPU resource allocator that enables pay-as-you-go computing model and flexible vNPU-to-pNPU mappings for improved resource utilization and cost-effectiveness; (3) an ISA extension of modern NPU architecture for facilitating fine-grained tensor operator scheduling for multiple vNPUs. We implement Neu10 based on a production-level NPU simulator. Our experiments show that Neul0 improves the throughput of ML inference services by up to 1.4 × and reduces the tail latency by up to 4.6 ×, while improving the NPU utilization by 1.2 × on average, compared to state-of-the-art NPU sharing approaches. Yuqi Xue, Lifeng Nai, Jian Huang 0006 |
MICRO | 1 |
| 2024 | Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning DistractorabstractLarge multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the ability of LMMs to correct potential visual commonsense errors in the distractor upon their occurrence is yet under-explored. Drawing inspiration from how a human teacher crafts challenging distractors to test students' comprehension of the concepts or skills and assists them in identifying and correcting errors toward the answer, we are the pioneering research for LMMs to simulate this error correction process. To this end, we employ GPT-4 as a ''teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers. In addition, we propose an LMM-based Pedagogical Expert Instructed Feedback Generation (PEIFG) model to incorporate the learnable expert prompts and multimodal instruction as guidance for feedback generation. Experimental results show that our PEIFG significantly outperforms existing LMMs. We believe that our benchmark provides a new direction for evaluating the capabilities of LMMs. Xusen Hei, Yuqi Xue, Yuancheng Wei, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
ACM Multimedia | 3 |
| 2024 | Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10abstractAs AI chips incorporate numerous parallelized cores to scale deep learning (DL) computing, inter-core communication is enabled recently by employing high-bandwidth and low-latency interconnect links on the chip (e.g., Graphcore IPU). It allows each core to directly access the fast scratchpad memory in other cores, which enables new parallel computing paradigms. However, without proper support for the scalable inter-core connections in current DL compilers, it is hard for developers to exploit the benefits of this new architecture. Yuqi Xue, Yu Cheng 0030, Lingxiao Ma, Ziming Miao, Jilong Xue, Jian Huang 0006 |
SOSP | 2 |
| 2023 | System Virtualization for Neural Processing UnitsabstractModern cloud platforms have been employing hardware accelerators such as neural processing units (NPUs) to meet the increasing demand for computing resources for AI-based application services. However, due to the lack of system virtualization support, the current way of using NPUs in cloud platforms suffers from either low resource utilization or poor isolation between multi-tenant application services. In this paper, we investigate the system virtualization techniques for NPUs across the entire software and hardware stack, and present our NPU virtualization solution named NeuCloud. We propose a flexible NPU abstraction named vNPU that allows fine-grained NPU virtualization and resource management. We leverage this abstraction and design the vNPU allocation, mapping, and scheduling policies to maximize the resource utilization, while achieving both performance and security isolation for vNPU instances at runtime. Yuqi Xue, Jian Huang 0006 |
HotOS | 1 |
| 2023 | V10: Hardware-Assisted NPU Multi-tenancy for Improved Resource Utilization and FairnessabstractModern cloud platforms have deployed neural processing units (NPUs) like Google Cloud TPUs to accelerate online machine learning (ML) inference services. To improve the resource utilization of NPUs, they allow multiple ML applications to share the same NPU, and developed both time-multiplexed and preemptive-based sharing mechanisms. However, our study with real-world NPUs discloses that these approaches suffer from surprisingly low utilization, due to the lack of support for fine-grained hardware resource sharing in the NPU. Specifically, its separate systolic array and vector unit cannot be fully utilized at the same time, which requires fundamental hardware assistance for supporting multi-tenancy. Yuqi Xue, Lifeng Nai, Jian Huang 0006 |
ISCA | 1 |
| 2023 | G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor MigrationsabstractTo break the GPU memory wall for scaling deep learning workloads, a variety of architecture and system techniques have been proposed recently. Their typical approaches include memory extension with flash memory and direct storage access. However, these techniques still suffer from suboptimal performance and introduce complexity to the GPU memory management, making them hard to meet the scalability requirement of deep learning workloads today. Yirui Eric Zhou, Yuqi Xue, Jian Huang 0006 |
MICRO | 3 |
| 2023 | RackBlox: A Software-Defined Rack-Scale Storage System with Network-Storage Co-DesignabstractSoftware-defined networking (SDN) and software-defined flash (SDF) have been serving as the backbone of modern data centers. They are managed separately to handle I/O requests. At first glance, this is a reasonable design by following the rack-scale hierarchical design principles. However, it suffers from suboptimal end-to-end performance, due to the lack of coordination between SDN and SDF. Benjamin Reidys, Yuqi Xue, Daixuan Li, Bharat Sukhwani, Wen-Mei W. Hwu, Deming Chen, Sameh W. Asaad, Jian Huang 0006 |
SOSP | 2 |
| 2021 | IceClave: A Trusted Execution Environment for In-Storage ComputingabstractIn-storage computing with modern solid-state drives (SSDs) enables developers to offload programs from the host to the SSD. It has been proven to be an effective approach to alleviate the I/O bottleneck. To facilitate in-storage computing, many frameworks have been proposed. However, few of them treat the in-storage security as the first citizen. Specifically, since modern SSD controllers do not have a trusted execution environment, an offloaded (malicious) program could steal, modify, and even destroy the data stored in the SSD. Luyi Kang, Yuqi Xue, Weiwei Jia 0001, Xiaohao Wang, Jongryool Kim, Changhwan Youn, Myeong Joon Kang, Hyung Jin Lim, Bruce L. Jacob, Jian Huang 0006 |
MICRO | 2 |