EDBT 2026 Demo / reviewers in the wild / expert
Xuanyi Li
dblp:213/1886
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WSGraph: A Framework for Tackling Redundant and Irregular Data Access in Streaming Graph ProcessingabstractThe demand for real-time streaming graph analysis has grown significantly, as hundreds of thousands of updates come every second. Monotonic graph algorithms such as Shortest Path are widely used in real-time analytics, but there are two bottlenecks that limit their performance, specially on planar graphs. One is massive redundant data accesses due to irregular state propagations and the other is high memory latency caused by irregular data accesses. We observe that existing systems mainly focus on general-purpose graph algorithms. If the properties of specific graph algorithms are exploited, the analysis performance can be further improved. Moreover, these systems typically tackle these two bottlenecks separately through either software or hardware mechanisms, but not both. However, both bottlenecks need to be addressed simultaneously in real scenarios such as road navigation. This article proposes WSGraph, a software-hardware co-design framework for high-performance streaming graph processing. WSGraph tackles these two challenges by enforcing regularized processing orders and enabling precise data prefetching. Specifically, at the software level, WSGraph integrates a priority-based work scheduler with sliding-window bucket mapping scheme to regulate state propagations, thereby drastically reducing redundant data accesses. At the hardware level, WSGraph incorporates a lightweight in-core Proactive Data Engine (PDE). By exploiting intra-vertex access regularity, the PDE accurately prefetches relevant graph data to effectively hide the high latency of irregular memory accesses. Experimental results demonstrate that WSGraph achieves significant performance improvements over existing systems. Compared with the state-of-the-art software system KickStarter, WSGraph gains a 2.13× speedup primarily by reducing graph data accesses by an average of 78.6%. Xuanyi Li, Chen Li 0015, Zhengyi Dai, Jianzhuang Lu, Yang Guo 0003 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | LAMP: A Locality-Aware TLB Sharing Framework for Multi-Tenant GPUs via TLB Comprehensive ProfilingabstractWith the growing prevalence of cloud services, GPUs are increasingly shared across multiple tenants, making address translation a performance-critical path. Our study makes three key observations: 1) Shared L2 TLB becomes the bottleneck in multi-tenant systems. Intense competition for entries can lead to high miss rates and thrashing under heavy pressure. 2) Benchmarks vary significantly in TLB sensitivity due to differing access patterns and memory reuse potential. 3) The prevailing management strategies, including free-for-all sharing and static partitioning, fail to account for tenant access behavior. And the state-of-the-art token-based strategy relies solely on miss rate, an inadequate metric that fails to capture per-tenant access patterns. Based on these observations, we propose LAMP, a localityaware TLB sharing framework. LAMP comprehensively profiles the access pattern considering three key aspects: spatial locality, temporal locality, and access density. These aspects allow LAMP to assess each tenant's sensitivity to TLB capacity and their reuse potential. Based on these insights, LAMP optimizes the sharing by dynamically partitioning L2 TLB, granting more entries to reuse-efficient tenants while limiting wasteful occupancy. This reduces thrashing and cross-tenant interference. Experimental results show that LAMP improves system performance by$\mathbf{1 2. 3 \%}$on average across a range of multi-tenant workloads, with negligible hardware overhead. Chen Li 0015, Xuanyi Li, Jianzhuang Lu, Yang Guo 0003 |
HPCC | 3 |
| 2025 | An Edge Morphology-Aware Self-Correcting Framework for Precise Infrared UAV Detection
Xuanyi Li, Yiyao Wan, Fuhui Zhou |
IEEE Internet Things J. | 1 |
| 2024 | Pyramid quaternion discrete cosine transform based ConvNet for cancelable face recognition
Zhuhong Shao, Leding Li, Xuanyi Li, Bicao Li |
Image Vis. Comput. | 5 |
| 2023 | Residual shuffle attention network for image super-resolution
Xuanyi Li, Zhuhong Shao, Bicao Li, Jiasong Wu, Yuping Duan |
Mach. Vis. Appl. | 1 |
| 2022 | Few-Shot Relational Triple Extraction with Perspective Transfer NetworkabstractFew-shot Relational Triple Extraction (RTE) aims at detecting emerging relation types along with their entity pairs from unstructured text with the support of a few labeled samples. Prior arts use conditional random field or nearest-neighbor matching strategy to extract entities and use prototypical networks for extracting relations from sentences. Nevertheless, they fail to utilize the triple-level information to verify the plausibility of extracted relational triples, and ignore the proper transfer among the perspectives of entity, relation and triple. To fill in these gaps, in this work, we put forward a novel perspective transfer network (PTN) to address few-shot RTE. Specifically, PTN starts from the relation perspective by checking the existence of a given relation. Then, it transfers to the entity perspective to locate entity spans with relation-specific support sets. Next, it transfers to the triple perspective to validate the plausibility of extracted relational triples. Finally, it transfers back to the relation perspective to check the next relation, and repeats the aforementioned procedure. By transferring among the perspectives of relation, entity, and triple, PTN not only validates the extracted elements at both local and global levels, but also effectively handles more realistic and difficult few-shot RTE scenarios such as multiple triple extraction and nonexistence of triples. Extensive experimental results on existing dataset and new datasets demonstrate that our approach can significantly improve performance over the state-of-the-arts. Junbo Fei, Weixin Zeng, Xiang Zhao 0002, Xuanyi Li, Weidong Xiao 0003 |
CIKM | 4 |
| 2021 | Improving Inter-kernel Data Reuse With CTA-Page Coordination in GPGPUabstractAlthough modern GPUs are equipped with expanding memory, accommodating the entire working set of large-scale workloads can still be a challenge. With the support of unified virtual memory and demand paging, programmers can transparently oversubscribe the main memory. However, this transparent management still comes at a severe performance cost, especially for applications with inter-kernel data sharing. While there have been many efforts to reduce additional data migrations caused by the memory oversubscription, few consider the reuse of shared data during the boundary of adjacent kernels. Due to limited memory capacity, we observe that adjacent kernel often demands shared pages that were evicted by the previous kernel, resulting in a significant number of costly data migrations. In this paper, we propose a CTA-Page collaborative framework, called CPC, that transparently reduces the impact of memory oversubscription using CTA dispatch switching and page replacement switching coordinately to reuse inter-kernel shared data. We evaluate CPC with a variety of GPGPU benchmark suites. Experimental results show that the system performance is improved by 65 % compared with the state-of-the-art technique for applications with inter-kernel data sharing. Xuanyi Li, Chen Li 0015, Yang Guo 0003, Rachata Ausavarungnirun |
ICCAD | 1 |
| 2017 | Convolutional Neural Networks Based Multi-task Deep Learning for Movie Review ClassificationabstractDeep learning has achieved impressive success in natural language processing. However, most previous models are learned on the specific single tasks, suffering from insufficient training set. Multi-task deep learning can solve this dilemma by sharing the part of parameters, improving generalization. The common multi-task deep learning model consists of the shared layer and the task specific layer. In this paper, we attempt to enhance the performance of shared layer and proposed two variants based on convolutional neural networks. The first model is Agent Model-Direct Concatenate, where each task is assigned with a separate convolutional neural network for extracting the common and task specific features simultaneously. The second model is Agent Model-Gating Concatenation, where the task specific layer could automatically decide the information flow of each element of the output of shared layer. The two networks are trained jointly over three pair-wise groups of movie review data sets. Experiments show the effectiveness of our two networks, inspiring a potential direction for the related research of multi-task deep learning. Xuanyi Li, Weimin Wu 0002 |
DSAA | 1 |