Ruibai Tang

dblp:412/7467 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0003-5935-4217ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 87% Parallel and multicore computing · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
automatic differentiation
1.012026
ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing · PPoPP 2026
Compilers and program optimization › automatic differentiation
reverse-mode automatic differentiation
1.012026
ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing · PPoPP 2026
Reconfigurable computing and FPGAs
FPGA accelerator
1.012026
Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Reconfigurable computing and FPGAs
FPGA architecture
1.012026
Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Algorithms and data structures › numerical linear algebra
eigenvalue computation
0.312026
Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Algorithms and data structures
numerical linear algebra
0.312026
Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026

Methods — techniques the papers use, named apart from their topics

tape restructuring · 2.0single-shift QR iteration · 2.0hessenberg matrix optimization · 2.0givens rotations · 2.0direct indexing · 2.0
YearPublicationVenuePosition
2026 ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing
abstract
Automatic Differentiation (AD) is a technique that computes the derivatives of numerical programs by systematically applying the chain rule, playing a critical role in domains such as machine learning, simulation, and control systems. However, parallelizing differentiated programs remains a significant challenge due to the conflict between tapes (a data structure for intermediate variable storage) and summations: the differentiation process inherently introduces inter-thread summation patterns, which require prohibitively expensive atomic operations; and traditional tape designs tightly couple data retrieval with the program’s control flow, preventing code restructuring needed to eliminate these costly dependencies.
Shuhong Huang, Shizhi Tang, Yuan Wen, Huanqi Cao, Ruibai Tang, Yidong Chen 0003, Jiping Yu, Jidong Zhai
PPoPP5
2026 Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting
abstract
Solving and visualizing the potential roots of complex functions is essential in both theoretical and applied domains, yet often computationally intensive. We present a hardware-accelerated algorithm for complex function roots density graph plotting by approximating functions with polynomials and solving their roots using single-shift QR iteration. By leveraging the Hessenberg structure of companion matrices and optimizing QR decomposition with Givens rotations, we design a pipelined FPGA architecture capable of processing a large amount of polynomials with high throughput. Our implementation achieves up to 65× higher energy efficiency than CPU-based approaches, and while it trails modern GPUs in performance. Compared with state-of-the-art QR decomposition solutions, our design specificly optimize QR decomposition for complex-valued Hessenberg matrices up to size 6x6, exhibiting a moderate throughput of 16.5M QR decompositions per second, while prior works have predominantly focused on 4x4 general matrices.
Ruibai Tang, Chengbin Quan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1