Mingzhen He

dblp:313/1706 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-9214-4196ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Correction to: MUSO: achieving exact machine unlearning in over‑parameterized regimes
Ruikai Yang, Mingzhen He, Zhengbao He, Youmei Qiu, Xiaolin Huang
Mach. Learn.2
2026 Data imputation by pursuing better classification: A supervised kernel-based method
Ruikai Yang, Mingzhen He, Xiaolin Huang
Pattern Recognit.3
2025 Simulating Training Dynamics to Reconstruct Training Data from Deep Neural Networks
abstract
Whether deep neural networks (DNNs) memorize the training data is a fundamental open question in understanding deep learning. A direct way to verify the memorization of DNNs is to reconstruct training data from DNNs’ parameters. Since parameters are gradually determined by data throughout training, characterizing training dynamics is important for reconstruction. Pioneering works rely on the linear training dynamics of shallow NNs with large widths, but cannot be extended to more practical DNNs which have non-linear dynamics. We propose Simulation of training Dynamics (SimuDy) to reconstruct training data from DNNs. Specifically, we simulate the training dynamics by training the model from the initial parameters with a dummy dataset, then optimize this dummy dataset so that the simulated dynamics reach the same final parameters as the true dynamics. By incorporating dummy parameters in the simulated dynamics, SimuDy effectively describes non-linear training dynamics. Experiments demonstrate that SimuDy significantly outperforms previous approaches when handling non-linear training dynamics, and for the first time, most training samples can be reconstructed from a trained ResNet’s parameters.
Hanling Tian, Yuhang Liu 0003, Mingzhen He, Zhengbao He, Zhehao Huang, Ruikai Yang, Xiaolin Huang
ICLR3
2025 Primphormer: Efficient Graph Transformers with Primal Representations
abstract
Graph Transformers (GTs) have emerged as a promising approach for graph representation learning. Despite their successes, the quadratic complexity of GTs limits scalability on large graphs due to their pair-wise computations. To fundamentally reduce the computational burden of GTs, we propose a primal-dual framework that interprets the self-attention mechanism on graphs as a dual representation. Based on this framework, we develop Primphormer, an efficient GT that leverages a primal representation with linear complexity. Theoretical analysis reveals that Primphormer serves as a universal approximator for functions on both sequences and graphs, while also retaining its expressive power for distinguishing non-isomorphic graphs. Extensive experiments on various graph benchmarks demonstrate that Primphormer achieves competitive empirical results while maintaining a more user-friendly memory and computational costs.
Mingzhen He, Ruikai Yang, Hanling Tian, Youmei Qiu, Xiaolin Huang
ICML1
2025 MUSO: achieving exact machine unlearning in over-parameterized regimes
Ruikai Yang, Mingzhen He, Zhenghao He, Youmei Qiu, Xiaolin Huang
Mach. Learn.2
2025 Decentralized Kernel Ridge Regression Based on Data-Dependent Random Feature
abstract
Random feature (RF) has been widely used for node consistency in decentralized kernel ridge regression (KRR). Currently, the consistency is guaranteed by imposing constraints on coefficients of features, necessitating that the RFs on different nodes are identical. However, in many applications, data on different nodes vary significantly on the number or distribution, which calls for adaptive and data-dependent methods that generate different RFs. To tackle the essential difficulty, we propose a new decentralized KRR algorithm that pursues consensus on decision functions, which allows great flexibility and well adapts data on nodes. The convergence is rigorously given, and the effectiveness is numerically verified: by capturing the characteristics of the data on each node, while maintaining the same communication costs as other methods, we achieved an average regression accuracy improvement of 25.5% across six real-world datasets.
Ruikai Yang, Mingzhen He, Jie Yang 0002, Xiaolin Huang
IEEE Trans. Neural Networks Learn. Syst.3
2024 Kernel PCA for Out-of-Distribution Detection
abstract
Out-of-Distribution (OoD) detection is vital for the reliability of Deep Neural Networks (DNNs). Existing works have shown the insufficiency of Principal Component Analysis (PCA) straightforwardly applied on the features of DNNs in detecting OoD data from In-Distribution (InD) data. The failure of PCA suggests that the network features residing in OoD and InD are not well separated by simply proceeding in a linear subspace, which instead can be resolved through proper non-linear mappings. In this work, we leverage the framework of Kernel PCA (KPCA) for OoD detection, and seek suitable non-linear kernels that advocate the separability between InD and OoD data in the subspace spanned by the principal components. Besides, explicit feature mappings induced from the devoted task-specific kernels are adopted so that the KPCA reconstruction error for new test samples can be efficiently obtained with large-scale data. Extensive theoretical and empirical results on multiple OoD data sets and network structures verify the superiority of our KPCA detector in efficiency and efficacy with state-of-the-art detection performance.
Kun Fang 0004, Qinghua Tao, Kexin Lv, Mingzhen He, Xiaolin Huang, Jie Yang 0002
NeurIPS4
2024 Random fourier features for asymmetric kernels
Mingzhen He, Fanghui Liu 0001, Xiaolin Huang
Mach. Learn.1
2024 Global Search and Analysis for the Nonconvex Two-Level ℓ₁ Penalty
abstract
Imposing suitably designed nonconvex regularization is effective to enhance sparsity, but the corresponding global search algorithm has not been well established. In this article, we propose a global search algorithm for the nonconvex two-level$\ell _{1}$penalty based on its piecewise linear property and apply it to machine learning tasks. With the search capability, the optimization performance of the proposed algorithm could be improved, resulting in better sparsity and accuracy than most state-of-the-art global and local algorithms. Besides, we also provide an approximation analysis to demonstrate the effectiveness of our global search algorithm in sparse quantile regression.
Mingzhen He, Lei Shi 0010, Xiaolin Huang
IEEE Trans. Neural Networks Learn. Syst.2
2023 Diffusion Representation for Asymmetric Kernels via Magnetic Transform
abstract
As a nonlinear dimension reduction technique, the diffusion map (DM) has been widely used. In DM, kernels play an important role for capturing the nonlinear relationship of data. However, only symmetric kernels can be used now, which prevents the use of DM in directed graphs, trophic networks, and other real-world scenarios where the intrinsic and extrinsic geometries in data are asymmetric. A promising technique is the magnetic transform which converts an asymmetric matrix to a Hermitian one. However, we are facing essential problems, including how diffusion distance could be preserved and how divergence could be avoided during diffusion process. Via theoretical proof, we successfully establish a diffusion representation framework with the magnetic transform, named MagDM. The effectiveness and robustness for dealing data endowed with asymmetric proximity are demonstrated on three synthetic datasets and two trophic networks.
Mingzhen He, Ruikai Yang, Xiaolin Huang
NeurIPS1
2023 Learning With Asymmetric Kernels: Least Squares and Feature Interpretation
abstract
Asymmetric kernels naturally exist in real life, e.g., for conditional probability and directed graphs. However, most of the existing kernel-based learning methods require kernels to be symmetric, which prevents the use of asymmetric kernels. This paper addresses the asymmetric kernel-based learning in the framework of the least squares support vector machine named AsK-LS, resulting in the first classification method that can utilize asymmetric kernels directly. We will show that AsK-LS can learn with asymmetric features, namely source and target features, while the kernel trick remains applicable, i.e., the source and target features exist but are not necessarily known. Besides, the computational burden of AsK-LS is as cheap as dealing with symmetric kernels. Experimental results on various tasks, including Corel, PASCAL VOC, Satellite, directed graphs, and UCI database, all show that in the case asymmetric information is crucial, the proposed AsK-LS can learn with asymmetric kernels and performs much better than the existing kernel methods that rely on symmetrization to accommodate asymmetric kernels.
Mingzhen He, Lei Shi 0010, Xiaolin Huang, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Learning non-parametric kernel via matrix decomposition for logistic regression
Mingzhen He, Xiaolin Huang
Pattern Recognit. Lett.3
2022 Taylor Discrete Circadian Rhythms Neural Network for Resolving Bicriteria Optimization Problem of Redundant Robot Manipulators Perturbed by Periodic Noises
abstract
To solve the motion planning problems of redundant manipulators disturbed by the periodic noise from device hardwares or their surroundings, a Taylor-type discrete-time circadian rhythms neural network (TD-CRNN) method is proposed, developed, and studied in this article. First, a representative bicriteria optimization scheme combining torque criterion and acceleration criterion is presented for the redundant manipulator. Second, inspired by a continuous-time circadian rhythms model, the corresponding TD-CRNN model is derived based on the Taylor discrete formulation. Third, the 0-stability, convergence, and consistency of the proposed TD-CRNN model are analyzed theoretically and proved strictly. Finally, to confirm the capacity of the resisting periodic noise in the tracking problem of manipulators, two groups of comparative simulations and experiments conducted by the proposed TD-CRNN model are performed before the conclusion is given.
Zhijun Zhang 0003, Siyuan Chen 0006, Mingzhen He
IEEE Trans. Ind. Informatics3
2022 Runge-Kutta Type Discrete Circadian RNN for Resolving Tri-Criteria Optimization Scheme of Noises Perturbed Redundant Robot Manipulators
abstract
In order to resist periodic interfere in robot hardware or environment, a Runge–Kutta type discrete-time circadian rhythms neural network (RK-DCRNN) model is proposed, and investigated to plan the motion of redundant robot manipulators. To achieve the optimal control, a quadratic programming-based acceleration-level hybrid tri-criteria (ALHT) scheme is first designed, which simultaneously minimize the acceleration norm, torque norm, and joint-angle shift-free indices. Second, according to the neural dynamic design method, a continuous-time circadian rhythms neural network model is exploited, and then based on the Runge–Kutta numerical differential method, a discrete-time circadian rhythms neural network model is obtained. Third, the convergence of the proposed RK-DCRNN model is proved by detailed mathematical derivation. Fourth, comparative simulations and physical experiments verify that the proposed RK-DCRNN model can suppress the accumulation of position error in the motion planning of manipulators.
Zhijun Zhang 0003, Xianzhi Deng, Mingzhen He, Tao Chen 0025
IEEE Trans. Syst. Man Cybern. Syst.3