Tianjing Xu

dblp:10/10636 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 50% Parallel and multicore computing · 26% Storage systems · 19%
Network and information security
2 papers
Blockchain and cryptocurrency security · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Blockchain and cryptocurrency security › smart contract
smart contract execution
0.912025
Vegeta: Enabling Parallel Smart Contract Execution in Leaderless Blockchains · NSDI 2025
Parallel and multicore computing
distributed deep learning training
0.812024
Kspeed: Beating I/O Bottlenecks of Data Provisioning for RDMA Training Clusters · ICNP 2024
Blockchain and cryptocurrency security
consensus protocol
0.712023
Flexible Advancement in Asynchronous BFT Consensus · SOSP 2023
Distributed systems › consensus › fault-tolerant consensus
asynchronous consensus
0.712023
Flexible Advancement in Asynchronous BFT Consensus · SOSP 2023
Distributed systems › consensus
byzantine agreement
0.712023
Flexible Advancement in Asynchronous BFT Consensus · SOSP 2023
Distributed systems
consensus
0.712023
Flexible Advancement in Asynchronous BFT Consensus · SOSP 2023
Parallel and multicore computing
parallel programming models
0.312025
Vegeta: Enabling Parallel Smart Contract Execution in Leaderless Blockchains · NSDI 2025

Methods — techniques the papers use, named apart from their topics

asynchronous BFT protocol · 1.3multi-rail RDMA · 0.8disaggregated memory pooling · 0.8
YearPublicationVenuePosition
2026 Integrating 2PC with Consensus for Fast Replication
abstract
Fast consensus protocols, such as CURP-Q and NOPaxos, have coupled objectives of reducing latency and tolerating faults, but doing so incurs considerable processing overhead and/or necessitates complex changes to existing architectures. Our insight is that when no adverse execution conditions (including operation conflicts, node failures, and network failures) occur, replication can be done safely without consensus at all. Rather than pursuing one single consensus protocol for fast and fault-tolerant replication, we propose to decouple low latency from fault tolerance, and design a hybrid solution that seamlessly integrates client-coordinated two-phase commit (2PC) with consensus, respectively to reduce latency in normal situations and tolerate faults in faulty situations. In the absence of faults, the system enters the Fast mode, where the client directly broadcasts its requests to all replicas with 2PC completing all updates in the first phase, i.e., in one RTT. Otherwise, the system enters the Consensus mode, which resorts to a consensus protocol to tolerate faults. We have applied the hybrid solution to the widely-used Raft consensus protocol to design xRaft, which can adaptively switch between the Fast and Consensus modes, ensuring correctness with minimal switching overhead. Evaluation shows that xRaft significantly outperforms state-of-the-art consensus protocols.
Shengyun Liu, Ruofan Xiong, Tianjing Xu, Yongwei Wu 0001, Yiming Zhang 0003
SIGCOMM5
2025 Vegeta: Enabling Parallel Smart Contract Execution in Leaderless Blockchains
Tianjing Xu, Yongqi Zhong, Yiming Zhang 0003, Ruofan Xiong, Jingjing Zhang 0002, Guangtao Xue, Shengyun Liu
NSDI1
2024 Kspeed: Beating I/O Bottlenecks of Data Provisioning for RDMA Training Clusters
abstract
The rapidly-increasing computing power of GPUs has rendered the I/O subsystem a bottleneck for distributed deep learning (DL) training. Currently, substantial data preprocessing work (e.g., decoding) has to be conducted on CPUs for a wide range of training scenarios such as computer vision (CV) and audio. Unfortunately, the involvement of training nodes' host memory and/or CPUs on the critical path of loading data to GPUs incurs significant GPU stalls in modern RDMA training clusters, because CPUs are much slower than GPUs and the connection from PCIe switches to host memory tends to suffer from incast problems. Moreover, this also incurs high CPU usage and resource contention, which consequently causes data loading performance variation and stragglers. This paper presents KSpeed, a novel data provisioning framework for large-scale RDMA training clusters. As many data preprocessing tasks need to be done by CPUs, KSpeed organizes host memory and CPU resources in the cluster to build a disaggregated memory/CPU pool, where the nodes can read raw input data from backend storage to their host memory, preprocess the data by their CPUs if necessary, and write cached/preprocessed data (on demand) directly to the training workers' GPU memory to minimize GPU stalls. KSpeed leverages the multi-rail RDMA network to eliminate unnecessary memory copies, interference, and congestion. Evaluation on a 96-GPU cluster shows that KSpeed delivers$5.4 \times \sim 100 \times$higher data loading performance over the state-of-the-art designs (DPP and Alluxio). KSpeed achieves near-linear scalability as the GPU number increases from 8 to 512.
Jianbo Dong, Hao Qi 0008, Tianjing Xu, Xiaoli Liu 0002, Rongyao Wang, Xiaoyi Lu 0001, Zheng Cao 0003, Binzhang Fu
ICNP3
2023 Flexible Advancement in Asynchronous BFT Consensus
abstract
Byzantine fault tolerant (BFT) consensus protocols are becoming an appealing solution to blockchains. As most blockchain systems are deployed on Wide Area Networks (WANs), with each node acting on behalf of its entity, partially synchronous BFT protocols that rely on network synchrony to elect a single leader can be ill-suited. In contrast, asynchronous protocols have no such timing assumptions. Existing asynchronous protocols confront challenges in terms of both flexibility and performance.
Shengyun Liu, Wenbo Xu 0002, Chen Shan, Xiaofeng Yan, Tianjing Xu, Bo Wang 0116, Lei Fan 0002, Fuxi Deng, Ying Yan 0002, Hui Zhang 0002
SOSP5
2021 VPC: Pruning connected components using vector-based path compression for Graph500
Xinbiao Gan, Tianjing Xu, Menghan Jia, Juan Chen 0001, Yiming Zhang 0003
CCF Trans. High Perform. Comput.3
2021 Correction to: VPC: Pruning connected components using vector-based path compression for Graph500
Xinbiao Gan, Tianjing Xu, Menghan Jia, Juan Chen 0001, Yiming Zhang 0003
CCF Trans. High Perform. Comput.3