Tailing Yuan

dblp:217/1542 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-6119-8829ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 83% Language models and text generation · 17%
Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 55% Geometric modeling and processing · 31% Visual content generation and editing · 14%
Network and information security
2 papers
Blockchain and cryptocurrency security · 60% Digital forensics and information hiding · 40%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 69% Memory systems · 31%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.622025
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training · SC 2025
Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism · USENIX ATC 2024
Natural language and speech › Language models and text generation
large language model training
0.912025
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training · SC 2025
Machine learning › Efficient and distributed learning › efficient training
long-context training
0.912025
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training · SC 2025
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
0.912025
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training · SC 2025
Computer animation and physical simulation
deformable body simulation
0.912025
Fast Galerkin Multigrid Method for Unstructured Meshes · ACM Trans. Graph. 2025
Geometric modeling and processing
multigrid solver
0.912025
Fast Galerkin Multigrid Method for Unstructured Meshes · ACM Trans. Graph. 2025
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training
0.812024
Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism · USENIX ATC 2024
Blockchain and cryptocurrency security
consensus protocol
0.612022
Meta-Regulation: Adaptive Adjustment to Block Size and Creation Interval for Blockchain Systems · IEEE J. Sel. Areas Commun. 2022
Distributed systems
consensus
0.612022
Meta-Regulation: Adaptive Adjustment to Block Size and Creation Interval for Blockchain Systems · IEEE J. Sel. Areas Commun. 2022
Visual content generation and editing
QR code generation
0.412019
Two-Layer QR Codes · IEEE Trans. Image Process. 2019
Digital forensics and information hiding
information hiding
0.412019
Two-Layer QR Codes · IEEE Trans. Image Process. 2019
Computer animation and physical simulation
fluid simulation
0.312018
Real-Time High-Fidelity Surface Flow Simulation · IEEE Trans. Vis. Comput. Graph. 2018
Machine learning › Efficient and distributed learning › memory-efficient training
re-materialization
0.212024
Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism · USENIX ATC 2024

Methods — techniques the papers use, named apart from their topics

micro-batch scheduling · 1.7matrix-free vertex block jacobi smoothing · 0.9galerkin multigrid · 0.9full approximation scheme · 0.9error correction coding · 0.8triangle mesh discretization · 0.3shallow water equations · 0.3bottom friction model · 0.3
YearPublicationVenuePosition
2025 SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
abstract
Pipeline Parallelism serves as a crucial technique for training Large Language Models, as it alleviates memory pressure from model states with relatively low communication overhead. However, in long-context scenarios, existing pipeline parallelism methods fail to address the substantial activation memory pressure, primarily due to the peak memory consumption resulting from the accumulation of activations across multiple microbatches. Moreover, these approaches inevitably introduce considerable pipeline bubbles, further hindering efficiency.
Zhouyang Li, Tailing Yuan, Chengru Song
SC4
2025 Fast Galerkin Multigrid Method for Unstructured Meshes
abstract
We present a novel multigrid solver framework that significantly advances the efficiency of physical simulation for unstructured meshes. While multi-grid methods theoretically offer linear scaling, their practical implementation for deformable body simulations faces substantial challenges, particularly on GPUs. Our framework achieves up to 6.9× speedup over traditional methods through an innovative combination of matrix-free vertex block Jacobi smoothing with a Full Approximation Scheme (FAS), enabling both piecewise constant and linear Galerkin formulations without the computational burden of dense coarse matrices. Our approach demonstrates superior performance across varying mesh resolutions and material stiffness values, maintaining consistent convergence even under extreme deformations and challenging initial configurations. Comprehensive evaluations against state-of-the-art methods confirm our approach achieves lower simulation error with reduced computational cost, enabling simulation of tetrahedral meshes with over one million vertices at approximately one frame per second on modern GPUs.
Jia-Ming Lu, Tailing Yuan, Zhe-Han Mo, Shi-Min Hu 0001
ACM Trans. Graph.2
2024 Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism
Tailing Yuan, Xucheng Ye, Shenglong Zhang, Jianchao Tan, Chengru Song
USENIX ATC1
2022 Meta-Regulation: Adaptive Adjustment to Block Size and Creation Interval for Blockchain Systems
abstract
Once deployed, a decentralized blockchain system ensures that it will operate faithfully so that no one can interfere with or manipulate its predefined regulations, such as block size and block creation interval investigated in this paper. However, fixed regulations prevent that system from adapting to the change of the environment, such as increasing the underlying network capacity, and result in sub-optimal performance. For example, Bitcoin remains at 7 TPS (transactions per second), even operating over the current Internet. In this paper, we propose a new paradigm for defining the behavior of a consensus system, named as Meta-Regulation, which allows autonomous evolution of the system behavior. A meta-regulation adjusts the actual behavior of a consensus system in response to the changing capacity of the underlying infrastructure and the community of participants. We demonstrate the effectiveness of the proposed meta-regulation by achieving significantly improved throughput and latency for Bitcoin, adapted to the current capacity of the Internet. Our experimental results show that Meta-Regulation can achieve at least$7\times $performance improvement over Bitcoin network deployed in 2009, resulting in 49.7 TPS or 68% reduction confirmation latency by fully utilizing the bandwidth and the computing power of average network nodes.
Mingpei Cao, Hao Wang 0002, Tailing Yuan, Kun Xu 0003, Kai Lei, Jiaping Wang
IEEE J. Sel. Areas Commun.3
2019 A Large Chinese Text Dataset in the Wild
Tailing Yuan, Zhe Zhu, Kun Xu 0003, Cheng-Jun Li, Tai-Jiang Mu, Shi-Min Hu 0001
J. Comput. Sci. Technol.1
2019 Two-Layer QR Codes
abstract
A quick-response code (QR code) is a two-dimensional code akin to a barcode that encodes a message of limited length. In this paper, we present a variant of QR code, a two-layer QR code. Its two-layer structure can display two alternative messages when scanned from two different directions. We propose a method to generate such two-layer QR codes encoding two given messages in a few seconds. We also demonstrate the robustness of our method on both synthetic and fabricated examples. All source code will be made publicly available (https://github.com/yuantailing/two-layer-qrcode).
Tailing Yuan, Yili Wang 0003, Kun Xu 0003, Ralph R. Martin, Shi-Min Hu 0001
IEEE Trans. Image Process.1
2018 Real-Time High-Fidelity Surface Flow Simulation
abstract
Surface flow phenomena, such as rain water flowing down a tree trunk and progressive water front in a shower room, are common in real life. However, compared with the 3D spatial fluid flow, these surface flow problems have been much less studied in the graphics community. To tackle this research gap, we present an efficient, robust and high-fidelity simulation approach based on the shallow-water equations. Specifically, the standard shallow-water flow model is extended to general triangle meshes with a feature-based bottom friction model, and a series of coherent mathematical formulations are derived to represent the full range of physical effects that are important for real-world surface flow phenomena. In addition, by achieving compatibility with existing 3D fluid simulators and by supporting physically realistic interactions with multiple fluids and solid surfaces, the new model is flexible and readily extensible for coupled phenomena. A wide range of simulation examples are presented to demonstrate the performance of the new approach.
Bo Ren 0003, Tailing Yuan, Chenfeng Li, Kun Xu 0003, Shi-Min Hu 0001
IEEE Trans. Vis. Comput. Graph.2