Boyuan Tian

dblp:174/3004 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-1726-3248ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 48% Efficient and distributed learning · 48% Robot navigation and mapping · 4%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware accelerators and domain-specific architectures · 69% Energy-efficient computing · 31%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
compression
0.912025
FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting · ACM Multimedia 2025
Energy-efficient computing › energy-quality tradeoff
energy-latency-accuracy trade-off
0.812024
Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy · ISMAR 2024
Computer vision › 3D vision
point cloud processing
0.522020
Tigris: Architecture and Algorithms for 3D Perception in Point Clouds · MICRO 2019
Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-Aggregation · MICRO 2020
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design
0.412020
Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-Aggregation · MICRO 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-Aggregation · MICRO 2020
Hardware accelerators and domain-specific architectures › vision accelerator
point cloud accelerator
0.412020
Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-Aggregation · MICRO 2020
Computer vision › 3D vision
point cloud registration
0.412019
Tigris: Architecture and Algorithms for 3D Perception in Point Clouds · MICRO 2019
Robotics › Robot navigation and mapping
localization
0.112021
Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines Industry Track Paper · HPCA 2021

Methods — techniques the papers use, named apart from their topics

software-hardware co-design · 1.0FPGA prototyping · 1.0quantization · 0.9pruning · 0.9pareto optimization · 0.8energy-efficient architecture design · 0.8design space exploration · 0.8delayed-aggregation · 0.4delayed aggregation · 0.4
YearPublicationVenuePosition
2026 A comprehensive survey on encrypted network traffic classification
Shangbin Han, Han Zhang 0009, Mengmeng Lu, Sifang Guo, Boyuan Tian, Jilong Wang 0001
Comput. Networks5
2025 FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting
abstract
3D Gaussian Splatting has emerged as a prominent technique for representing and rendering complex 3D scenes, offering high fidelity and speed but resulting in large file sizes. Existing compression methods can reduce 3D Gaussian data size but often require costly retraining or refinement, which is memory- and compute-intensive and lacks flexibility for varying compression needs. This challenge grows as large-scale scenes become more common, increasing the demand for efficient, low-overhead compression methods - especially for resource-constrained mobile and edge devices.
Boyuan Tian, Qizhe Gao, Siran Xianyu, Xiaotong Cui, Minjia Zhang
ACM Multimedia1
2024 Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy
abstract
Extended Reality (XR) enables immersive experiences through untethered headsets but suffers from stringent battery and resource constraints. Energy-efficient design is crucial to ensure both longevity and high performance in XR devices. However, latency and accuracy are often prioritized over energy, leading to a gap in achieving energy efficiency. This paper examines scene reconstruction, a key building block for immersive XR experiences, and demonstrates how energy efficiency can be achieved by navigating the trilemma of energy, latency, and accuracy. We explore three classes of energy-oriented optimizations, covering the algorithm, execution, and data, that reveal a broad de-sign space through configurable parameters. Our resulting 72 designs expose a wide range of latency and energy trade-offs, with a smaller range of accuracy loss. We identify a Pareto-optimal curve and show that the designs on the curve are achievable only through synergistic co-optimization of all three optimization classes and by considering the latency and accuracy needs of downstream scene reconstruction consumers. Our analysis covering various use cases and measurements on an embedded class system shows that, relative to the baseline, our designs offer energy benefits of up to $60 \times$ with potential latency range of $4 \times$ slowdown to $2 \times$ speedup. Detailed exploration of a use case across representative data sequences from ScanNet showed about $25 \times$ energy savings with $1.5 \times$ latency reduction and negligible reconstruction quality loss.
Boyuan Tian, Yihan Pang, Muhammad Huzaifa, Shenlong Wang, Sarita V. Adve
ISMAR1
2021 Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines Industry Track Paper
abstract
We develop and commercialize autonomous machines, such as logistic robots and self-driving cars, around the globe. A critical challenge to our—and any—autonomous machine is accurate and efficient localization under resource constraints, which has fueled specialized localization accelerators recently. Prior acceleration efforts are point solutions in that they each specialize for a specific localization algorithm. In real-world commercial deployments, however, autonomous machines routinely operate under different environments and no single localization algorithm fits all the environments. Simply stacking together point solutions not only leads to cost and power budget overrun, but also results in an overly complicated software stack. This paper demonstrates our new software-hardware co-designed framework for autonomous machine localization, which adapts to different operating scenarios by fusing fundamental algorithmic primitives. Through characterizing the software framework, we identify ideal acceleration candidates that contribute significantly to the end-to-end latency and/or latency variation. We show how to co-design a hardware accelerator to systematically exploit the parallelisms, locality, and common building blocks inherent in the localization framework. We build, deploy, and evaluate an FPGA prototype on our next-generation self-driving cars. To demonstrate the flexibility of our framework, we also instantiate another FPGA prototype targeting drones, which represent mobile autonomous machines. We achieve about $2 \times$ speedup and $4 \times$ energy reduction compared to widely-deployed, optimized implementations on general-purpose platforms.
Yiming Gan, Bo Yu 0014, Boyuan Tian, Leimeng Xu, Shaoshan Liu, Qiang Liu 0011, Jie Tang 0003, Yuhao Zhu 0001
HPCA3
2020 Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-Aggregation
abstract
Point cloud analytics is poised to become a key workload on battery-powered embedded and mobile platforms in a wide range of emerging application domains, such as autonomous driving, robotics, and augmented reality, where efficiency is paramount. This paper proposes Mesorasi, an algorithm-architecture co-designed system that simultaneously improves the performance and energy efficiency of point cloud analytics while retaining its accuracy.Our extensive characterizations of state-of-the-art point cloud algorithms show that, while structurally reminiscent of convolutional neural networks (CNNs), point cloud algorithms exhibit inherent compute and memory inefficiencies due to the unique characteristics of point cloud data. We propose delayed-aggregation, a new algorithmic primitive for building efficient point cloud algorithms. Delayed-aggregation hides the performance bottlenecks and reduces the compute and memory redundancies by exploiting the approximately distributive property of key operations in point cloud algorithms. Delayed-aggregation let point cloud algorithms achieve 1.6× speedup and 51.1% energy reduction on a mobile GPU while retaining the accuracy (-0.9% loss to 1.2% gains). To maximize the algorithmic benefits, we propose minor extensions to contemporary CNN accelerators, which can be integrated into a mobile Systems-on-a-Chip (SoC) without modifying other SoC components. With additional hardware support, Mesorasi achieves up to 3.6× speedup.
Yu Feng 0007, Boyuan Tian, Tiancheng Xu, Paul N. Whatmough, Yuhao Zhu 0001
MICRO2
2019 Tigris: Architecture and Algorithms for 3D Perception in Point Clouds
abstract
Machine perception applications are increasingly moving toward manipulating and processing 3D point cloud. This paper focuses on point cloud registration, a key primitive of 3D data processing widely used in high-level tasks such as odometry, simultaneous localization and mapping, and 3D reconstruction. As these applications are routinely deployed in energy-constrained environments, real-time and energy-efficient point cloud registration is critical.
Tiancheng Xu, Boyuan Tian, Yuhao Zhu 0001
MICRO2