Mingzi Wang

dblp:385/0606 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 85% Electronic design automation · 15%
Artificial intelligence
3 papers
Efficient and distributed learning · 64% Trustworthy machine learning · 18% 3D vision · 18%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 77% Image and video processing · 23%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
1.012026
Oiso: Outlier-Isolated Data Format for Low-Bit Large Language Model Quantization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
Oiso: Outlier-Isolated Data Format for Low-Bit Large Language Model Quantization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Hardware accelerators and domain-specific architectures › quantization
outlier-aware quantization
1.012026
Oiso: Outlier-Isolated Data Format for Low-Bit Large Language Model Quantization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Hardware accelerators and domain-specific architectures
quantization
1.012026
Oiso: Outlier-Isolated Data Format for Low-Bit Large Language Model Quantization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Machine learning › Efficient and distributed learning › federated learning
client selection
0.912025
A Pricing Game for Federated Learning Supporting Lightweight Local Model Training · IEEE Trans. Mob. Comput. 2025
Machine learning › Efficient and distributed learning
federated learning
0.912025
A Pricing Game for Federated Learning Supporting Lightweight Local Model Training · IEEE Trans. Mob. Comput. 2025
Computer vision › 3D vision
implicit neural representation
0.912025
EVOS: Efficient Implicit Neural Training via EVOlutionary Selector · CVPR 2025
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
sample selection
0.912025
EVOS: Efficient Implicit Neural Training via EVOlutionary Selector · CVPR 2025
Machine learning › Efficient and distributed learning › efficient training
training acceleration
0.912025
EVOS: Efficient Implicit Neural Training via EVOlutionary Selector · CVPR 2025
Geometric modeling and processing
implicit neural representation
0.912025
Enhancing Implicit Neural Representations via Symmetric Power Transformation · AAAI 2025
Electronic design automation
hardware/software co-design
0.912025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural network accelerator design
0.912025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization
0.312025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Machine learning › Efficient and distributed learning
model compression
0.312025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Edge and fog computing
edge networks
0.312025
A Pricing Game for Federated Learning Supporting Lightweight Local Model Training · IEEE Trans. Mob. Comput. 2025

Methods — techniques the papers use, named apart from their topics

stackelberg game · 1.7hardware generation network · 1.7compiler mapping search · 1.7channel-wise sparse quantization · 1.7approximation algorithm · 1.7subblock alignment · 1.0block encoding · 1.0sparse fitness evaluation · 0.9frequency-guided crossover · 0.9evolutionary algorithm · 0.9data transformation · 0.9adaptive calibration · 0.9
YearPublicationVenuePosition
2026 Oiso: Outlier-Isolated Data Format for Low-Bit Large Language Model Quantization
abstract
The scale of large language models (LLMs) has steadily increased over time, leading to enhanced performance in multi-modal understanding and complex reasoning, but with significant execution overhead on hardware. Quantization is a promising approach to reduce computation and memory overhead for LLM deployment. However, maintaining accuracy and efficiency simultaneously is challenging due to the presence of outliers. Moreover, low-bit quantization tends to deteriorate accuracy due to its limited precision. Existing outlier-aware quantization/hardware co-design methods split the sparse outliers from the normal values with dedicated encoding schemes. However, such separation produces a non-uniform data format for normal values and outliers, leading to additional hardware design and inefficient memory access. This paper presents an outlier-isolated data format for low-bit LLM quantization called Oiso. Oiso is a unified representation for both outliers and normal values. It isolates the normal values from the outliers, which can reduce the impact of outliers on the normal values during the quantization process. Taking advantage of the uniform format, Oiso arithmetic can be performed using a homogeneous computational unit, and Oiso values can be stored in a standardized format. Hierarchical block encoding with a subblock alignment scheme is introduced to reduce the encoding cost and the hardware overhead. We introduce the Oiso architecture, equipped with Oiso processing elements and encoders tailored for Oiso arithmetic, realizing efficient low-bit LLM inference. Oiso quantization can push the limits of low-bit LLM quantization, and the Oiso accelerator outperforms the state-of-the-art outlieraware accelerator design with 1.26× performance improvement and 25% energy reduction.
Lancheng Zou, Mingzi Wang, Wenqian Zhao 0002, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration
abstract
The co-design of neural network architectures, quantization precisions, and hardware accelerators offers a promising approach to achieving an optimal balance between performance and efficiency, particularly for model deployment on resource-constrained edge devices. In this work, we propose the JAQ Framework, which jointly optimizes the three critical dimensions. However, effectively automating the design process across the vast search space of those three dimensions poses significant challenges, especially when pursuing extremely low-bit quantization. Specifical, the primary challenges include: (1) Memory overhead in software-side: Low-precision quantization-aware training can lead to significant memory usage due to storing large intermediate features and latent weights for backpropagation, potentially causing memory exhaustion. (2) Search time-consuming in hardware-side: The discrete nature of hardware parameters and the complex interplay between compiler optimizations and individual operators make the accelerator search time-consuming. To address these issues, JAQ mitigates the memory overhead through a channel-wise sparse quantization (CSQ) scheme, selectively applying quantization to the most sensitive components of the model during optimization. Additionally, JAQ designs BatchTile, which employs a hardware generation network to encode all possible tiling modes, thereby speeding up the search for the optimal compiler mapping strategy. Extensive experiments demonstrate the effectiveness of JAQ, achieving approximately 7% higher Top-1 accuracy on ImageNet compared to previous methods and reducing the hardware search time per iteration to 0.15 seconds.
Mingzi Wang, Weixiang Zhang, Yijian Qin, Yang Yao 0003, Yingxin Li, Tongtong Feng, Xin Wang 0019, Xun Guan, Zhi Wang 0001, Wenwu Zhu 0001
AAAI1
2025 Enhancing Implicit Neural Representations via Symmetric Power Transformation
abstract
We propose symmetric power transformation to enhance the capacity of Implicit Neural Representation (INR) from the perspective of data transformation. Unlike prior work utilizing random permutation or index rearrangement, our method features a reversible operation that does not require additional storage consumption. Specifically, we first investigate the characteristics of data that can benefit the training of INR, proposing the Range-Defined Symmetric Hypothesis, which posits that specific range and symmetry can improve the expressive ability of INR. Based on this hypothesis, we propose a nonlinear symmetric power transformation to achieve both range-defined and symmetric properties simultaneously. We use the power coefficient to redistribute data to approximate symmetry within the target range. To improve the robustness of the transformation, we further design deviation-aware calibration and adaptive soft boundary to address issues of extreme deviation boosting and continuity breaking. Extensive experiments are conducted to verify the performance of the proposed method, demonstrating that our transformation can reliably improve INR compared with other data transformations. We also conduct 1D audio, 2D image and 3D video fitting tasks to demonstrate the effectiveness and applicability of our method.
Weixiang Zhang, Shuzhao Xie, Chengwei Ren, Shijia Ge, Mingzi Wang
AAAI5
2025 EVOS: Efficient Implicit Neural Training via EVOlutionary Selector
abstract
We propose EVOlutionary Selector (EVOS), an efficient training paradigm for accelerating Implicit Neural Representation (INR). Unlike conventional INR training that feeds all samples through the neural network in each iteration, our approach restricts training to strategically selected points, reducing computational overhead by eliminating redundant forward passes. Specifically, we treat each sample as an individual in an evolutionary process, where only those fittest ones survive and merit inclusion in training, adaptively evolving with the neural network dynamics. While this is conceptually similar to Evolutionary Algorithms, their distinct objectives (selection for acceleration vs. iterative solution optimization) require a fundamental redefinition of evolutionary mechanisms for our context. In response, we design sparse fitness evaluation, frequency-guided crossover, and augmented unbiased mutation to comprise EVOS. These components respectively guide sample selection with reduced computational cost, enhance performance through frequency-domain balance, and mitigate selection bias from cached evaluation. Extensive experiments demonstrate that our method achieves approximately 48%-66% reduction in training time while ensuring superior convergence without additional cost, establishing state-of-the-art acceleration among recent sampling-based strategies. Our code is available at this link.
Weixiang Zhang, Shuzhao Xie, Chengwei Ren, Siyi Xie, Shijia Ge, Mingzi Wang
CVPR7
2025 A Pricing Game for Federated Learning Supporting Lightweight Local Model Training
abstract
The pervasive distribution of data across clients with privacy concerns and heterogeneous performance in edge networks presents a significant opportunity to enhance AI model performance. Federated learning (FL) enables a model owner (MO) to recruit these clients, offering compensation for their contributions, and to improve model quality by aggregating knowledge from their locally trained models. However, several challenges arise in this process. Clients may decline participation if they do not achieve positive utility. Moreover, due to constraints in memory, computing, and communication resources, some clients can only train lightweight models that represent partial versions of the global model. Importantly, the MO's pricing for client contributions and the proportions of local model training are interdependent, collectively influencing client utilities and participation decisions. To address these challenges, we first model the utility functions of both the MO and the clients, accommodating the support for lightweight local models. We then formulate their interactions as a Stackelberg game and theoretically prove the existence of a Nash equilibrium. Based on this equilibrium, we derive optimal collaboration strategies for both the MO and the clients. Additionally, we design an efficient approximation algorithm to enable the MO to maximize its utility by selecting suitable clients to participate in FL. Finally, extensive experiments validate our theoretical findings, demonstrating the superior performance and effectiveness of the proposed algorithms
Fengsen Tian, Mingzi Wang, Guoqiang Deng, Lingyu Liang, Xinglin Zhang 0001
IEEE Trans. Mob. Comput.2
2025 Rank-DSE: Neural Pareto Comparator of Microarchitecture Design Space Exploration
abstract
The complexity of microarchitecture design has surged due to the expanding design space and time-intensive verification processes. Existing regression-based machine learning methods struggle with inaccurate estimations because of limited training samples. To address these challenges, we propose Rank-DSE, a novel framework for microarchitecture design space exploration (DSE) that leverages a Neural Pareto Comparator (NPC) to directly model the comparative relationships between different architecture designs. Rank-DSE bypasses the inaccuracies of absolute PPA (performance, power, area) predictions by focusing on relative comparisons. The NPC computes the probability of one architecture dominating another and employs semi-supervised learning to reduce the reliance on labeled data. Additionally, a reinforcement-learning-based sampling scheme with an updating baseline Pareto set accelerates the exploration process. Experimental results on the ICCAD 2021 benchmark demonstrate that Rank-DSE achieves superior search quality and cost-efficiency compared to state-of-the-art methods. Specifically, Rank-DSE improves hypervolume by up to 7% while reducing exploration cost by 53.09% compared to cutting-edge approaches. These results highlight the advantages of Rank-DSE in terms of efficiency and effectiveness for microarchitecture DSE.
Peng Xu 0052, Su Zheng, Mingzi Wang, Ziyang Yu 0001, Shixin Chen, Tinghuan Chen, Keren Zhu 0001, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.3