Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Qiuliang Wang

dblp:121/5583 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2026
0000-0001-5987-4069ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 56% Hardware accelerators and domain-specific architectures · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
analytical modeling
1.012026
NPUMeter: Automatic Operator Optimization for Ascend NPU with Accurate Analytical Performance Models · ACM Trans. Archit. Code Optim. 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
1.012026
NPUMeter: Automatic Operator Optimization for Ascend NPU with Accurate Analytical Performance Models · ACM Trans. Archit. Code Optim. 2026
Performance modeling and evaluation › performance prediction
latency prediction
0.312026
NPUMeter: Automatic Operator Optimization for Ascend NPU with Accurate Analytical Performance Models · ACM Trans. Archit. Code Optim. 2026

Methods — techniques the papers use, named apart from their topics

design space exploration · 1.0analytical performance modeling · 1.0
YearPublicationVenuePosition
2026 NPUMeter: Automatic Operator Optimization for Ascend NPU with Accurate Analytical Performance Models
abstract
With the rapid development of AI and deep learning, computational demands are increasing significantly. While GPUs excel in parallel computing, they fall short in terms of energy efficiency, specialization, and processing latency. In contrast, Neural Processing Units (NPUs), such as the Ascend NPUs, designed specifically for deep learning tasks, demonstrate superior performance. However, the architecture specialization makes operator development more challenging, leading to a reliance on manual tuning and optimization, which incurs significant time cost and developing effort. To address this issue, we propose NPUMeter, an automatic operator optimization framework for Ascend NPUs built upon accurate and comprehensive analytical performance models. NPUMeter comprises two components: (1) an analytical performance model that accurately estimates operator latency on NPU given different configurations of optimization parameters; (2) an efficient design space exploration (DSE) algorithm that automatically searches for the optimal parameter configuration in a large design space within minutes. Experimental results demonstrate that NPUMeter achieves high estimation accuracy, with an average error below 5%. It effectively generates near-optimal configurations for various operators, achieving up to a 1.46× performance speedup compared to the configuration generated by the Ascend C compiler while reducing the DSE time from hours to minutes.
Weichuang Zhang, Yufei Shangguan, Yuting Mai, Qiuliang Wang, Chen Chen 0067, Quan Chen 0002, Wenchao Ding 0001, Jieru Zhao, Minyi Guo
ACM Trans. Archit. Code Optim.5
2014 Multiple View Based Building Modeling with Multi-box Grammar
abstract
This paper describes a multiple view based approach for building modeling via a novel multi-box grammar, which represents an occlusion relationship among the projections of a set of buildings sharing a common Manhattan World coordinate system. We formulate the building modeling problem as an energy minimization to combine the constraints from the multi-box grammar with (1) the semantic labeling information from appearance models, (2) the directional information w.r.t the vanishing points in each single view, and (3) the planar homography correspondence among multiple views. We further propose a two-step coarse-to-fine approach to achieve the optimal solution. First we employ super-pixels and a simplified edition of the grammar to reduce the searching space, and obtain an initial layout to accelerate the convergence speed. At the second stage, the scene model is refined to achieve pixel-level accuracy by minimizing the energy using Random Walk. Experiments on street view images demonstrate the capability of our method in reconstructing multiple buildings at different distances, and also the robustness in handling occlusion.
Ruiling Deng, Qiuliang Wang, Rui Gan, Hongbin Zha
ICPR2