EDBT 2026 Demo / reviewers in the wild / expert
Youxuan Xu
dblp:231/1925
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Omnia: Efficient RAG Serving through Speculative SchedulingabstractRetrieval-Augmented Generation (RAG) has emerged for enhancing Large Language Models (LLMs) by improving factual accuracy and mitigating hallucinations. A typical RAG pipeline executes in three cascaded stages: retrieval, reranking, and generation. The existing serving systems suffer from two critical system-level bottlenecks when applying to RAG serving: the first is the cumulative latency caused by rigid sequential dependencies between reranking and generation, and the second is the system saturation triggered by bursty, high fan-in reranking workloads. Rongtian Fu, Shigang Li 0002, Youxuan Xu, Tong Wu 0024, Jinliang Shi |
HPDC | 3 |
| 2025 | FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresabstractSparse Matrix-matrix Multiplication (SpMM) and Sampled Dense-dense Matrix Multiplication (SDDMM) are important sparse operators in scientific computing and deep learning. Tensor Core Units (TCUs) enhance modern accelerators with superior computing power, which is promising to boost the performance of matrix operators to a higher level. However, due to the irregularity of unstructured sparse data, it is difficult to deliver practical speedups on TCUs. To this end, we propose FlashSparse, a novel approach to bridge the gap between sparse workloads and the TCU architecture. Specifically, FlashSparse minimizes the sparse granularity for SpMM and SDDMM on TCUs through a novel swap-and-transpose matrix multiplication strategy. Benefiting from the minimum sparse granularity, the computation redundancy is remarkably reduced while the computing power of TCUs is fully utilized. Besides, FlashSparse is equipped with a memory-efficient thread mapping strategy for coalesced data access and a sparse matrix storage format to save memory footprint. Extensive experimental results on H100 and RTX 4090 GPUs show that FlashSparse sets a new state-of-the-art for sparse matrix multiplications (geometric mean 5.5x speedup over DTC-SpMM and 3.22x speedup over RoDe). Jinliang Shi, Shigang Li 0002, Youxuan Xu, Rongtian Fu, Xueying Wang 0003, Tong Wu 0024 |
PPoPP | 3 |
| 2025 | SparkAttention: high-performance multi-head attention for large models on Volta GPU architecture
Youxuan Xu, Tong Wu 0024, Shigang Li 0002, Xueying Wang 0003 |
CCF Trans. High Perform. Comput. | 1 |
| 2018 | Fused Text Segmentation Networks for Multi-oriented Scene Text DetectionabstractIn this paper, we introduce a novel end-end framework for multi-oriented scene text detection from an instance-aware semantic segmentation perspective. We present Fused Text Segmentation Networks, which combine multi-level features during the feature extracting as text instance may rely on finer feature expression compared to general objects. It detects and segments the text instance jointly and simultaneously, leveraging merits from both semantic segmentation task and region proposal based object detection task. Not involving any extra pipelines, our approach surpasses the current state of the art on multi-oriented scene text detection benchmarks: ICDAR2015 Incidental Scene Text and MSRA-TD500 reaching Hmean 84.1 % and 82.0 % respectively. Morever, we report a baseline on total-text containing curved text which suggests effectiveness of the proposed approach. Yuchen Dai, Youxuan Xu, Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICPR | 4 |