Mengtian Yang

dblp:293/8205 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-7051-2250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 57% GPUs and heterogeneous computing · 36% Electronic design automation · 6%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › neural rendering accelerator
3d gaussian splatting accelerator
0.912025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
GPUs and heterogeneous computing
graphics accelerator
0.912025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
Rendering › gaussian splatting
3d gaussian splatting
0.312025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
Rendering › neural rendering
radiance field rendering
0.312025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
Electronic design automation
design space exploration
0.112021
NAAS: Neural Accelerator Architecture Search · DAC 2021

Methods — techniques the papers use, named apart from their topics

gaussian-centric rasterization · 1.7depth speculation · 1.7neural architecture search · 0.5compiler mapping co-search · 0.5
YearPublicationVenuePosition
2025 GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization
abstract
D Gaussian Splatting (3DGS) has emerged as a promising real-time photorealistic radiance field rendering technique. Existing GPU and hardware accelerators face limitations due to insufficient parallelism in sequential rendering pipeline stages and the memory overhead associated with interim results. This paper presents GSAcc, a hardware accelerator co-designed with dataflow to render compressed 3DGS models on edge platforms efficiently. GSAcc enhances 3DGS rendering performance through several key innovations. First, it introduces Gaussian depth speculation, parallelizing preprocessing and sorting tasks. Second, GSAcc adopts a Gaussian-centric dataflow that interleaves preprocessing and rasterization, allowing all rendering steps to execute concurrently without storing intermediate results. Finally, it employs dedicated hardware acceleration to address sorting and rasterization bottlenecks within the optimized dataflow. We implemented and synthesized GSAcc using Intel16 PDK and evaluated its performance on real-world 3DGS scenes. Compared with desktop GPUs, GSAcc achieves up to $1.66 \times 10^{4} \mathrm{x}$ Power-Performance-Area (PPA) improvement as well as 48.7 x energy savings. Additionally, GSAcc outperforms the state-of-the-art hardware accelerator GSCore with up to 2.3x PPA improvement and 2.9x energy savings.
Mengtian Yang, Yipeng Wang 0017, Chieh-Pu Lo, Xiuhao Zhang, Sirish Oruganti, Jaydeep P. Kulkarni
DAC1
2023 A 118 GOPS/mm23D eDRAM TensorCore Architecture for Large-scale Matrix Multiplication
abstract
The computational demands for recent large transformer- based language models and Neural Radiance Fields (NeRF) have rapidly increased, impacting applications like conversational AI and Mixed Reality (MR). Current accelerator architectures struggle to cope with the vast computational requirements, creating a gap with slowly growing hardware resources. This paper proposes repurposing memory components as high-density computational units, leveraging recent advancements in Back-End-Of-Line (BEOL) transistors and monolithic 3D integration techniques. An ultra-high density monolithic 3D eDRAM is presented as a reconfigurable matrix multiplication unit, co-designed with analog computation circuits, achieving energy efficiency up to 2.41 TOPS/W, performance up to 1.71 TOPS on bfloat16, and compute intensity up to 118 GOPS/mm2. A comprehensive multi-cube(core) architecture is also devised and optimized with bit stationary tensorcore dataflow. We evaluate the proposed architecture on state-of-the-art machine learning models: NeRF and LLaMa-7B, improving the computation density by up to 6.59x and 1.12x compared with GPU and state-of-the-art vector processor designs, respectively.
Mengtian Yang, Yipeng Wang 0017, Jaydeep P. Kulkarni
HiPC1
2021 NAAS: Neural Accelerator Architecture Search
abstract
Data-driven, automatic design space exploration of neural accelerator architecture is desirable for specialization and productivity. Previous frameworks focus on sizing the numerical architectural hyper-parameters while neglect searching the PE connectivities and compiler mappings. To tackle this challenge, we propose Neural Accelerator Architecture Search (NAAS) that holistically searches the neural network architecture, accelerator architecture and compiler mapping in one optimization loop. NAAS composes highly matched architectures together with efficient mapping. As a data-driven approach, NAAS rivals the human design Eyeriss by $4.4 \times$ EDP reduction with 2.7% accuracy improvement on ImageNet under the same computation resource, and offers $1.4 \times$ to $3.5 \times$ EDP reduction than only sizing the architectural hyper-parameters.
Yujun Lin 0001, Mengtian Yang, Song Han 0003
DAC2