Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shihan Yuan

dblp:433/8544 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0005-9508-1660ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop optimization
1.012026
A Decoupled Analytical Model for Tile Size Selection in Affine Programs · ACM Trans. Archit. Code Optim. 2026
Compilers and program optimization › loop transformation
tile size selection
1.012026
A Decoupled Analytical Model for Tile Size Selection in Affine Programs · ACM Trans. Archit. Code Optim. 2026
Compilers and program optimization › loop transformation
tiling
1.012026
A Decoupled Analytical Model for Tile Size Selection in Affine Programs · ACM Trans. Archit. Code Optim. 2026

Methods — techniques the papers use, named apart from their topics

nonlinear optimization · 1.0linearization · 1.0analytical modeling · 1.0
YearPublicationVenuePosition
2026 A Decoupled Analytical Model for Tile Size Selection in Affine Programs
abstract
Existing tile size selection approaches are tightly coupled with compiler transformation pipelines, often leading to inaccurate modeling of cache behavior and limited effectiveness for non-rectangular tile shapes. This article presents TileMind , a decoupled analytical model that combines compile-time and runtime information for tile size selection in affine programs. It introduces a transformation-aware pre-tiling step that enables the decoupled selector to remain consistent with compiler transformations while extracting compile-time metadata. The extracted metadata is then combined with profiled runtime characteristics to construct a richer yet tractable feasible domain, within which a nonlinear objective for tile size selection is formulated. This objective is subsequently transformed into a binary product linearization problem, with its nonlinear constraints also linearized for efficient optimization. Finally, an intra-tile optimization aligns computation with data layout to enhance data reuse within tiles. Across two multi-core Intel CPUs, TileMind achieves 1.49× (sequential) and 1.33× (parallel) mean speedups on twenty PolyBench kernels, and 2.08–3.54× speedups on three deep learning workloads over the state-of-the-art analytical model Pluto-tss . Compared with TVM’s latest autotuner MetaSchedule, TileMind delivers 1.35–1.46× mean speedups while reducing tuning overhead by 2–4 orders of magnitude. While demonstrating effectiveness on selecting tile sizes for non-rectangular tile shapes and compatibility with PPCG, Pluto, and TVM, we further provide proof-of-concept results on GPUs, illustrating the potential portability of TileMind across architectures.
Shihan Yuan, Zuoyan Zhang, Guanghui Song, Junhui Peng, Feng Wang 0050, Zhuo Tang, Kenli Li 0001, Jie Zhao 0002
ACM Trans. Archit. Code Optim.1
2025 Scalable Detection of Floating-Point Errors via Adaptive Parallel Subdomain Search
abstract
Floating-point error detection is crucial in numerical computing, particularly for multi-parameter functions where even minor errors can propagate and significantly impact results. The sparse distribution of floating-point errors poses a significant detection challenge, as significant deviations are triggered by only rare inputs. Existing search algorithms face two major limitations: poor scalability for multi-parameter functions and insufficient utilization of floating-point representation characteristics. To address these challenges, we propose SDPS (Scalable Detection via Parallel Subdomain Search), a novel algorithm that combines adaptive domain partitioning with floating-point-specific heuristics. SDPS employs a multi-level error classification system and specialized point generation strategies, supported by efficient parallel processing through dynamic task allocation. Our comprehensive evaluation demonstrates that SDPS significantly outperforms state-of-the-art methods in both detection accuracy and computational efficiency, especially for multi-parameter functions where it effectively addresses the exponential growth of search space that limits existing approaches.
Zuoyan Zhang, Shihan Yuan, Hongru Yang, Jie Zhao 0002, Jinchen Xu
QRS2