Jilong Xue

dblp:06/10336 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-4495-1997ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 10 · 9 since 2021Computer networks · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 MetaAttention: A Unified and Performant Attention Framework across Hardware Backends
abstract
Computing attention is the backbone of transformer-based models like large language models. However, the increasing diversity of attention algorithms presents significant challenges for unleashing hardware performance. State-of-the-art variants like FlashAttention target a specific attention algorithm or hardware platform, which fail to generalize to other algorithms and platforms.
Yu Cheng 0030, Lei Wang 0222, Yuqing Xia, Ziming Miao, Lingxiao Ma, Fan Yang 0024, Jilong Xue, Zhi Yang 0001, Mao Yang 0004, Xingda Wei, Haibo Chen 0001
PPoPP8
2025 NeuStream: Bridging Deep Learning Serving and Stream Processing
abstract
Modern Deep Neural Network (DNN) exhibits a pattern where multiple sub-models are executed, guided by control flows such as loops and switch/merge operations. This dynamic nature introduces complexities in batching the requests of such DNNs for efficient execution on GPUs. In this paper, we present NeuStream, a programming model and runtime system for serving deep learning workloads using stream processing. NeuStream decomposes the inference workflow into modules and forms them into a streaming processing system where a request flows through. Based on such abstraction, NeuStream is able to batch requests at fine-grained module granularity. To maximize serving goodput, NeuStream exploits a two-level scheduling approach to decide the best batching requests and resource allocation for each module while satisfying service level objectives (SLOs). Our evaluation of NeuStream on a set of modern DNNs like Large Language Models (LLM) and diffusion models, etc., shows that NeuStream significantly improves goodput compared to state-of-the-art DNN serving systems.
Yu Cheng 0030, Ziming Miao, Lingxiao Ma, Jilong Xue, Zhi Yang 0001
EuroSys7
2025 LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
abstract
Large Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency.Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations.Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains.To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference.LUT Tensor Core differentiates itself from conventional LUT designs through: 1) * Work is done during internship at Microsoft Research.
Zhiwen Mo, Lei Wang 0222, Jianyu Wei, Zhichen Zeng 0002, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao 0003, Jilong Xue, Fan Yang 0024, Mao Yang 0004
ISCA9
2025 Elk: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
Yuqi Xue, Noelle Crawford, Jilong Xue, Jian Huang 0006
MICRO4
2025 MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
abstract
The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third—a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics—Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)—to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios. This benchmark is available on Github: https://github.com/sparse-generative-ai/MoE-CAP.
Yinsicheng Jiang, Yao Fu 0013, Yeqi Huang, Ping Nie, Zhan Lu, Leyang Xue, Congjie He, Man-Kit Sit, Jilong Xue, Ziming Miao, Dayou Du, Tairan Xu, Edoardo Maria Ponti, Luo Mai
NeurIPS9
2025 PipeThreader: Software-Defined Pipelining for Efficient DNN Execution
Yu Cheng 0030, Lei Wang 0222, Yining Shi 0001, Yuqing Xia, Lingxiao Ma, Jilong Xue, Yang Wang 0053, Zhiwen Mo, Fan Yang 0024, Mao Yang 0004, Zhi Yang 0001
OSDI6
2025 WaferLLM: Large Language Model Inference at Wafer Scale
Congjie He, Yeqi Huang, Pei Mu 0003, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang 0024, Luo Mai
OSDI5
2025 Revealing Floating-Point Accumulation Orders in Software/Hardware Implementations
Peichen Xie, Yanjie Gao, Jilong Xue
USENIX ATC4
2025 BitNet: 1-bit Pre-training for Large Language Models
abstract
The increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs.
Hongyu Wang 0009, Shuming Ma, Lingxiao Ma, Lei Wang 0222, Wenhui Wang 0003, Li Dong 0004, Shaohan Huang, Huaijie Wang, Jilong Xue, Yi Wu 0013, Furu Wei
J. Mach. Learn. Res.9
2024 Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Lei Wang 0222, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi 0001, Ningxin Zheng, Ziming Miao, Fan Yang 0024, Ting Cao 0003, Yuqing Yang 0001, Mao Yang 0004
OSDI5
2024 Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10
abstract
As AI chips incorporate numerous parallelized cores to scale deep learning (DL) computing, inter-core communication is enabled recently by employing high-bandwidth and low-latency interconnect links on the chip (e.g., Graphcore IPU). It allows each core to directly access the fast scratchpad memory in other cores, which enables new parallel computing paradigms. However, without proper support for the scalable inter-core connections in current DL compilers, it is hard for developers to exploit the benefits of this new architecture.
Yuqi Xue, Yu Cheng 0030, Lingxiao Ma, Ziming Miao, Jilong Xue, Jian Huang 0006
SOSP6
2023 Welder: Scheduling Deep Learning Memory Access via Tile-graph
Yining Shi 0001, Zhi Yang 0001, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo 0001, Fan Yang 0024, Lidong Zhou
OSDI3
2023 Optimizing Dynamic Neural Networks with Brainstorm
Weihao Cui, Zhenhua Han, Lingji Ouyang, Yichuan Wang 0002, Ningxin Zheng, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Jilong Xue, Lili Qiu, Lidong Zhou, Quan Chen 0002, Haisheng Tan, Minyi Guo
OSDI9
2023 Cocktailer: Analyzing and Optimizing Dynamic Control Flow in Deep Learning
Chen Zhang 0001, Lingxiao Ma, Jilong Xue, Yining Shi 0001, Ziming Miao, Fan Yang 0024, Jidong Zhai, Zhi Yang 0001, Mao Yang 0004
OSDI3
2023 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
abstract
With the increasing data volume, there is a trend of using large-scale pre-trained models to store the knowledge into an enormous number of model parameters. The training of these models is composed of lots of dense algebras, requiring a huge amount of hardware resources. Recently, sparsely-gated Mixture-of-Experts (MoEs) are becoming more popular and have demonstrated impressive pretraining scalability in various downstream tasks. However, such a sparse conditional computation may not be effective as expected in practical systems due to the routing imbalance and fluctuation problems. Generally, MoEs are becoming a new data analytics paradigm in the data life cycle and suffering from unique challenges at scales, complexities, and granularities never before possible. In this paper, we propose a novel DNN training framework, FlexMoE, which systematically and transparently address the inefficiency caused by dynamic dataflow. We first present an empirical analysis on the problems and opportunities of training MoE models, which motivates us to overcome the routing imbalance and fluctuation problems by a dynamic expert management and device placement mechanism. Then we introduce a novel scheduling module over the existing DNN runtime to monitor the data flow, make the scheduling plans, and dynamically adjust the model-to-hardware mapping guided by the real-time data traffic. A simple but efficient heuristic algorithm is exploited to dynamically optimize the device placement during training. We have conducted experiments on both NLP models (e.g., BERT and GPT) and vision models (e.g., Swin). And results show FlexMoE can achieve superior performance compared with existing systems on real-world workloads --- FlexMoE outperforms DeepSpeed by 1.70x on average and up to 2.10x, and outperforms FasterMoE by 1.30x on average and up to 1.45x.
Xiaonan Nie, Xupeng Miao, Zilong Wang 0033, Jilong Xue, Lingxiao Ma, Gang Cao 0003, Bin Cui 0001
Proc. ACM Manag. Data5
2022 ROLLER: Fast and Efficient Tensor Compilation for Deep Learning
Hongyu Zhu 0003, Yijia Diao, Shanbin Ke, Chen Zhang 0001, Jilong Xue, Lingxiao Ma, Yuqing Xia, Fan Yang 0024, Mao Yang 0004, Lidong Zhou, Asaf Cidon, Gennady Pekhimenko
OSDI7
2020 Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Youshan Miao, Wenxiang Hu, Fan Yang 0024, Lidong Zhou
OSDI4
2020 Distributed Graph Computation Meets Machine Learning
abstract
TuX2is a new distributed graph engine that bridges graph computation and distributed machine learning.TuX2inherits the benefits of elegant graph computation model, efficient graph layout, and balanced parallelism to scale to billion-edge graphs, while extended and optimized for distributed machine learning to support heterogeneity in data model, Stale Synchronous Parallel in scheduling, and a new Mini-batch, Exchange, GlobalSync, and Apply (MEGA) model for programming.TuX2further introduces a hybrid vertex-cut graph optimization and supports various consistency models in fault tolerance for machine learning. We have developed a set of representative distributed machine learning algorithms inTuX2, covering both supervised and unsupervised learning. Compared to the implementations on distributed machine learning platforms, writing those algorithms inTuX2takes only about 25 percent of the code: our graph computation model hides the detailed management of data layout, partitioning, and parallelism from developers. The extensive evaluation ofTuX2, using large datasets with up to 64 billion of edges, shows thatTuX2outperforms PowerGraph/PowerLyra, the state-of-the-art distributed graph engines, by an order of magnitude, while beating two state-of-the-art distributed machine learning systems by at least 60 percent.
Wencong Xiao, Jilong Xue, Youshan Miao, Ming Wu 0007, Wei Li 0022, Lidong Zhou
IEEE Trans. Parallel Distributed Syst.2
2019 Fast Distributed Deep Learning over RDMA
abstract
Deep learning emerges as an important new resource-intensive workload and has been successfully applied in computer vision, speech, natural language processing, and so on. Distributed deep learning is becoming a necessity to cope with growing data and model sizes. Its computation is typically characterized by a simple tensor data abstraction to model multi-dimensional matrices, a dataflow graph to model computation, and iterative executions with relatively frequent synchronizations, thereby making it substantially different from Map/Reduce style distributed big data computation.
Jilong Xue, Youshan Miao, Ming Wu 0007, Lidong Zhou
EuroSys1
2019 NeuGraph: Parallel Deep Neural Network Computation on Large Graphs
Lingxiao Ma, Zhi Yang 0001, Youshan Miao, Jilong Xue, Ming Wu 0007, Lidong Zhou, Yafei Dai
USENIX ATC4
2017 Tux2: Distributed Graph Computation for Machine Learning
Wencong Xiao, Jilong Xue, Youshan Miao, Ming Wu 0007, Wei Li 0022, Lidong Zhou
NSDI2
2017 Garaph: Efficient GPU-accelerated Graph Processing on a Single Machine with Balanced Replication
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Yafei Dai
USENIX ATC4
2017 Processing Concurrent Graph Analytics with Decoupled Computation Model
abstract
Graph processing systems have been widely used in enterprises like online social networks to process their daily jobs. With the fast growing of social applications, they have to efficiently handle massive concurrent jobs. However, due to the inherent design for a single job, existing systems incur great inefficiency in terms of memory usage, execution and fault tolerance. Motivated by this issue, in this paper we introduce Seraph, a graph processing system that enables efficient job-level parallelism. Seraph is designed based on a decoupled computation model, which decouples both the runtime job data and the computation logic. Decoupling the runtime data allows multiple concurrent jobs to share graph structure data in memory, which fundamentally increases job-level concurrency and reduces fault tolerance overhead. Decoupling computation logic could extend scheduling space, which benefits both execution performance and memory consumption. Seraph adopts a copy-on-write semantic to isolate the graph mutation of concurrent jobs, and a lazy snapshot protocol to generate consistent graph snapshots for jobs submitted at different time. Based on the decoupled model, it provides unified programming interfaces for both synchronous and asynchronous graph applications. Moreover, Seraph implements a lightweight checkpoint mechanism which can tremendously reduce the fault tolerance overhead. The evaluation results show that Seraph significantly outperforms popular systems (such as Giraph, Spark, GraphX and PowerLyra) in both memory usage and job completion time, when executing concurrent graph jobs.
Jilong Xue, Zhi Yang 0001, Shian Hou, Yafei Dai
IEEE Trans. Computers1
2016 Efficient Distributed Machine Learning with Trigger Driven Parallel Training
abstract
Distributed machine learning is becoming increasingly popular for large scale data mining on large scale cluster. To mitigate the interference of straggler machines, recent distributed machine learning systems support flexible model consistency, which allows worker using a local stale model to compute model update without waiting for the newest model, while limiting the asynchronous step in a certain bound to guarantee the algorithm correctness. However, bounded asynchronous computing can not tolerate consistent straggler. We explore that the root cause of this problem derives from the worker driven parallel training mechanism in existing systems. To address the straggler problem fundamentally and fully leverage the asynchronous efficiency, we propose a novel trigger driven parallel training mechanism, where model server proactively triggers to collect updates from workers instead of passively receiving them, which can inherently avoid the coordinating issue among workers. Besides, we devise a dynamic load balancing strategy to make the sampling frequency of each data equal. Furthermore, bounded asynchronous computing is introduced to achieve the algorithm efficiency, as well as the convergence guarantee. Finally, we integrate the above techniques into a distributed machine learning system called Squirrel. Squirrel provides simple programming interface and can easily deploy machine learning algorithms on distributed cluster. In comparison with traditional worker driven parallel training mechanism, trigger driven mechanism can improve up to 4x faster convergence speed of machine learning algorithm.
Shenglong Li, Jilong Xue, Zhi Yang 0001, Yafei Dai
GLOBECOM2
2016 VoteTrust: Leveraging Friend Invitation Graph to Defend against Social Network Sybils
abstract
Online social networks (OSNs) suffer from the creation of fake accounts that introduce fake product reviews, malware and spam. Existing defenses focus on using the social graph structure to isolate fakes. However, our work shows that Sybils could befriend a large number of real users, invalidating the assumption behind social-graph-based detection. In this paper, we present VoteTrust, a scalable defense system that further leverages user-level activities. VoteTrust models the friend invitation interactions among users as a directed, signed graph, and uses two key mechanisms to detect Sybils over the graph: a voting-based Sybil detection to find Sybils that users vote to reject, and a Sybil community detection to find other colluding Sybils around identified Sybils. Through evaluating on Renren social network, we show that VoteTrust is able to prevent Sybils from generating many unsolicited friend requests. We also deploy VoteTrust in Renen, and our real experience demonstrates that VoteTrust can detect large-scale collusion among Sybils.
Zhi Yang 0001, Jilong Xue, Xiaoyong Yang, Xiao Wang 0018, Yafei Dai
IEEE Trans. Dependable Secur. Comput.2
2015 When computing meets heterogeneous cluster: Workload assignment in graph computation
abstract
In order to process very large graphs, existing graph processing systems, such as Pregel and Giraph, usually partition and distribute the graph computation on large number of nodes (i.e., workers). However, due to the heterogeneity of computing clusters (e.g., nodes with various bandwidth or CPU resource), blindly increasing the number of workers for a job may even degrade the overall performance. In this paper, we address the question of how to distribute the graph computation over the heterogeneous cluster to maximize performance. Based on the practical constraints of current systems, we address this problem in two scenarios. For systems using hash-based partition method (for avoiding the overhead of indexing and searching vertex), we propose a coarse-grained mechanism to greedily select suitable worker set to execute the job. For systems allowing arbitrary graph partition, we further propose a heterogeneity-aware streaming graph partitioning model that can assign workload in fine-grained level. We implement the scheduling mechanisms as a general middleware which can be easily adopted in existing graph computing systems. Our experiments on both university lab cluster (46 machines) and EC2 cluster (100 instances) show that, the proposed framework can significantly improve the execution performance. Compared with the default configurations (i.e., using the whole set of workers and hash-based graph partition), our framework can reduce the overall execution time by 55.9% for lab cluster and 44.7% for EC2 cluster respectively.
Jilong Xue, Zhi Yang 0001, Shian Hou, Yafei Dai
IEEE BigData1
2015 GraM: scaling graph computation to the trillions
abstract
GraM is an efficient and scalable graph engine for a large class of widely used graph algorithms. It is designed to scale up to multicores on a single server, as well as scale out to multiple servers in a cluster, offering significant, often over an order-of-magnitude, improvement over existing distributed graph engines on evaluated graph algorithms. GraM is also capable of processing graphs that are significantly larger than previously reported. In particular, using 64 servers (1,024 physical cores), it performs a PageRank iteration in 140 seconds on a synthetic graph with over one trillion edges, setting a new milestone for graph engines.
Ming Wu 0007, Fan Yang 0024, Jilong Xue, Wencong Xiao, Youshan Miao, Haoxiang Lin, Yafei Dai, Lidong Zhou
SoCC3
2015 Uncovering User Interaction Dynamics in Online Social Networks
Zhi Yang 0001, Jilong Xue, Christo Wilson, Ben Y. Zhao, Yafei Dai
ICWSM2
2015 Understanding the performance of offline download in real p2p networks
Zhi Yang 0001, Yuanjian Xing, Jilong Xue, Yafei Dai
Peer-to-Peer Netw. Appl.4
2014 Seraph: an efficient, low-cost system for concurrent graph processing
abstract
Graph processing systems have been widely used in enterprises like online social networks to process their daily jobs. With the fast growing of social applications, they have to efficiently handle massive concurrent jobs. However, due to the inherent design for single job, existing systems incur great inefficiency in memory use and fault tolerance. Motivated by this, in this paper we introduce Seraph, a graph processing system that enables efficient job-level parallelism. Seraph is designed based on a decoupled data model, which allows multiple concurrent jobs to share graph structure data in memory. Seraph adopts a copy-on-write semantic to isolate the graph mutation of concurrent jobs, and a lazy snapshot protocol to generate consistent graph snapshots for jobs submitted at different time. Moreover, Seraph adopts an incremental checkpoint/regeneration model which can tremendously reduce the overhead of checkpointing. We have implemented Seraph, and the evaluation results show that Seraph significantly outperforms popular systems (such as Giraph and Spark) in both memory usage and job completion time, when executing concurrent graph jobs.
Jilong Xue, Zhi Yang 0001, Shian Hou, Yafei Dai
HPDC1
2013 VoteTrust: Leveraging friend invitation graph to defend against social network Sybils
abstract
Online social networks (OSNs) currently face a significant challenge by the existence and continuous creation of fake user accounts (Sybils), which can undermine the quality of social network service by introducing spam and manipulating online rating. Recently, there has been much excitement in the research community over exploiting social network structure to detect Sybils. However, they rely on the assumption that Sybils form a tight-knit community, which may not hold in real OSNs. In this paper, we present VoteTrust, a Sybil detection system that further leverages user interactions of initiating and accepting links. VoteTrust uses the techniques of trust-based vote assignment and global vote aggregation to evaluate the probability that the user is a Sybil. Using detailed evaluation on real social network (Renren), we show VoteTrust's ability to prevent Sybils gathering victims (e.g., spam audience) by sending a large amount of unsolicited friend requests and befriending many normal users, and demonstrate it can significantly outperform traditional ranking systems (such as TrustRank or BadRank) in Sybil detection.
Jilong Xue, Zhi Yang 0001, Xiaoyong Yang, Xiao Wang 0018, Lijiang Chen, Yafei Dai
INFOCOM1
2011 On the QoS of Offline Download in Retrieving Peer-Side File Resource
abstract
P2P file-sharing systems have been suffering from file unavailability or poor download speed due to lack of file replicas. As a result, to download rare files, users are typically forced to pay long online waiting time (i.e., always-on strategy), which is unacceptable for most users. To get over this, several commercial P2P systems launch an offline download service, employing stable and high-capacity servers to take over the downloads. Thus users can log off and wait offline. While offline download shows a notable trend of growing popular, there has been little insight into its detailed QoS (i.e., success ratio and download time) provision and QoS characteristics. Through trace-driven simulation study, this paper quantitatively reveals the potential QoS provision of offline download, confirming that it can achieve high download success ratio (over 90%) with tolerable waiting time (less than 2 days in most cases). We further find that the retrieval of rare file is dramatically affected by any single replica's participation or leave, resulting in wildly fluctuated download time. The wild fluctuation prevents users from well perceiving the expected QoS of their offline downloads, leading to poor user experience. Motivated by this, we develop a QoS prediction method for offline download. Experiments show that the prediction achieves high accuracy and precision. Finally, we implement a prototype of offline download service, called SimpleOD, and present our experience of the QoS provision and the prediction performance in practice.
Yuanjian Xing, Zhi Yang 0001, Jilong Xue, Yafei Dai
ICPP4