Tingxuan Zhong

dblp:359/0378 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0000-5086-4649ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 56% GPUs and heterogeneous computing · 28% Parallel and multicore computing · 17%
Artificial intelligence
1 paper
Efficient and distributed learning · 50% Language models and text generation · 50%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.912025
Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling · SC 2025
High-performance computing
sparse linear algebra
0.912025
Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling · SC 2025
High-performance computing › sparse linear solver
sparse LU factorization
0.912025
Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling · SC 2025
Natural language and speech › Language models and text generation
large language model inference
0.812024
QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models · EMNLP 2024
Machine learning › Efficient and distributed learning
model quantization
0.812024
QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models · EMNLP 2024
Parallel and multicore computing › task scheduling
fine-grain scheduling
0.312025
Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling · SC 2025
Parallel and multicore computing
task scheduling
0.312025
Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling · SC 2025

Methods — techniques the papers use, named apart from their topics

static scheduling · 0.9memory caching · 0.91d block-cyclic distribution · 0.94-bit quantization · 0.8
YearPublicationVenuePosition
2025 Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling
abstract
We address inefficiencies in task scheduling, memory management, and scalability in GPU-resident sparse LU factorization with a two-level approach of sequentially scheduled coarse-grained blocks containing multiple fine-grained blocks managed with a lightweight static scheduler enabling multi-stream parallelism. Additionally, we design an intelligent memory caching mechanism for the fine-grained scheduler, which retains frequently accessed data in GPU memory. To further enhance scalability, we introduce a distributed memory design that partitions the input matrix using a 1D block-cyclic distribution and optimizes inter-GPU communication via NVLink. The multi-GPU design reaches a computational throughput of 6.46 TFLOP/s on four A100 GPUs, demonstrating promising scalability. This is up to 7x speedup over the latest SuperLU_DIST with 3D communication, 94x speedup over PanguLU, 16x speedup over PasTiX, and 10x speedup over our own coarse-grained dynamic scheduling implementation while reaching up to 21% of the A100’s theoretical peak performance.
Tingxuan Zhong, Yuxi Hong 0001, Guofeng Feng, Weile Jia, Hatem Ltaief, David E. Keyes
SC2
2024 QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models
abstract
Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Xincheng Wang, Jie Ren, Torsten Hoefler, Dan Alistarh. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Torsten Hoefler, Dan Alistarh
EMNLP4