Mingzhi Mao

dblp:64/2306 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-9369-7828ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
abstract
Repository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source repository. Many benchmarks have been proposed to evaluate the performance of such code translators. However, previous benchmarks mostly provide fine-grained samples, focusing at either code snippet, function, or file-level code translation. Such benchmarks do not accurately reflect real-world demands, where entire repositories often need to be translated, involving longer code length and more complex functionalities. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code translation benchmark featuring 1,897 real-world repository samples across 13 language pairs with automatically executable test suites. Besides, we introduce RepoTransAgent, a general agent framework to perform repository-level code translation. We evaluate both our benchmark’s challenges and agent’s effectiveness using several methods and backbone LLMs, revealing that repository-level translation remains challenging, where the best-performing method achieves only a 32.8% success rate. Furthermore, our analysis reveals that translation difficulty varies significantly by language pair direction, with dynamic-to-static language translation being much more challenging than the reverse direction (achieving below 10% vs. static-to-dynamic at 45-63%). Finally, we conduct a detailed error analysis and highlight current LLMs’ deficiencies in repository-level code translation, which could provide a reference for further improvements. We provide the code and data athttps://github.com/DeepSoftwareAnalytics/RepoTransBench.
Yanli Wang 0001, Yanlin Wang 0001, Suiquan Wang, Daya Guo, Jiachi Chen, John C. Grundy, Xilin Liu 0001, Yuchi Ma, Mingzhi Mao, Hongyu Zhang 0002, Zibin Zheng
IEEE Trans. Software Eng.9
2025 NeRFSwap: A NeRF-Based Generative Model for Face Swapping
abstract
Most existing face swapping methods involve decomposing the identity component of a source image into a latent vector and integrating it with a target image at the feature level. However, such methods often encounter challenges in accurately transferring identity while preserving facial structure. In this work, we propose a NeRF-based generative model for face swapping (NeRFSwap). Our core idea is to acquire the 3D representation of an intermediate face whose identity is from the source image and the remaining attributes are from the target image. Subsequently, we employ a novel dual adaptive blending module to generate results conditioned on this 3D representation and the target image at the feature level, which effectively transfers the facial structure from the 3D representation to the target image. Extensive experiments demonstrate the effectiveness of our method and show that NeRFSwap surpasses other state-of-the-art methods in generating high-fidelity images with superior identity transfer.
Shuangyi Tan, Mingzhi Mao, Guanbin Li
ICME2
2025 A Quadratic Programming Framework Unifying Different Types of Visual Servoing with Obstacle Avoidance for Joint-Constrained Robots
Weibing Li, Zeyu Ping, Zhiping Tan, Mingzhi Mao, Zilian Yi, Wenjing Ouyang, Chia-Wen Liao
PRICAI (4)5
2025 Eleven-point discrete perturbation-handling ZNN algorithm applied to tracking control of MIMO nonlinear system under various disturbances
Meichun Huang, Mingzhi Mao, Yunong Zhang
Neural Comput. Appl.2
2024 Simplified Gradient-Zeroing Neuronet for Temporally-Variant Convex Objective Function Minimization
Qianlong Yu, Mingzhi Mao, Yunong Zhang
ISNN3
2024 A Fuzzy-Enhanced Robust DZNN Model for Future Multiconstrained Nonlinear Optimization With Robotic Manipulator Control
abstract
Different from the common static and continuous-time dynamic problems of unconstrained/constrained nonlinear optimization, this article aims to investigate a discrete-time dynamic problem of nonlinear optimization with multiple types of constraints, which can be succinctly termed as future multiconstrained nonlinear optimization (FMCNO) problem because of the unknown future. Considering the unique advantages of neural networks with parallelism and fuzzy control systems (FCSs) with adaptivity, a fuzzy-enhanced robust discretized zeroing neural network (FER-DZNN) model is proposed to address the FMCNO problem. Specifically, by introducing a fuzzy factor outputted from an FCS with dual inputs, the FER-DZNN model is designed on the basis of an FER evolution rule and a five-step look-ahead discretization rule. Moreover, theoretical results are provided to indicate the convergence and robustness of the FER-DZNN model under various noises. Finally, two illustrative examples, including an application example to robotic manipulator control, are presented to substantiate the superior convergent and robust performance of the FER-DZNN model under various noises for addressing the FMCNO problem.
Binbin Qiu, Jinjin Guo, Mingzhi Mao, Ning Tan 0003
IEEE Trans. Fuzzy Syst.3
2024 Uncertainty-Aware Active Domain Adaptive Salient Object Detection
abstract
Due to the advancement of deep learning, the performance of salient object detection (SOD) has been significantly improved. However, deep learning-based techniques require a sizable amount of pixel-wise annotations. To relieve the burden of data annotation, a variety of deep weakly-supervised and unsupervised SOD methods have been proposed, yet the performance gap between them and fully supervised methods remains significant. In this paper, we propose a novel, cost-efficient salient object detection framework, which can adapt models from synthetic data to real-world data with the help of a limited number of actively selected annotations. Specifically, we first construct a synthetic SOD dataset by copying and pasting foreground objects into pure background images. With the masks of foreground objects taken as the ground-truth saliency maps, this dataset can be used for training the SOD model initially. However, due to the large domain gap between synthetic images and real-world images, the performance of the initially trained model on the real-world images is deficient. To transfer the model from the synthetic dataset to the real-world datasets, we further design an uncertainty-aware active domain adaptive algorithm to generate labels for the real-world target images. The prediction variances against data augmentations are utilized to calculate the superpixel-level uncertainty values. For those superpixels with relatively low uncertainty, we directly generate pseudo labels according to the network predictions. Meanwhile, we select a few superpixels with high uncertainty scores and assign labels to them manually. This labeling strategy is capable of generating high-quality labels without incurring too much annotation cost. Experimental results on six benchmark SOD datasets demonstrate that our method outperforms the existing state-of-the-art weakly-supervised and unsupervised SOD methods and is even comparable to the fully supervised ones. Code will be released at: https://github.com/czh-3/UADA.
Guanbin Li, Zhuohua Chen, Mingzhi Mao, Liang Lin 0004, Chaowei Fang
IEEE Trans. Image Process.3
2023 Hybrid-Order Representation Learning for Electricity Theft Detection
abstract
Electricity theft is the primary cause of electrical losses in power systems, which severely harms the economic benefits of electricity providers and threatens the safety of the power supply. However, due to the inherent complex correlation and periodicity of electricity consumption and the low efficiency of large-scale data processing, detecting anomalies in electricity consumption data accurately and efficiently remains challenging. Existing methods usually focus on first-order information and ignore the second-order representation learning that can efficiently model global temporal dependency and facilitate discriminative representation learning of electricity consumption data. In this article, we propose a novel electricity theft detection framework named hybrid-order representation learning network (HORLN). Specifically, the sequential electricity consumption data is transformed into the matrix format containing weekly consumption records. Then, an inter-and-intra week convolution block is designed to capture multiscale features in a local-to-global manner. Meanwhile, a self-dependency modeling module is proposed to learn the second-order representations from self-correlation matrices, which are finally combined with the first-order representations to predict the anomaly scores of electricity consumers. Extensive experiments on a real-world benchmark demonstrate the advantages of our HORLN over state-of-the-art methods.
Yuying Zhu 0007, Lingbo Liu, Yang Liu 0084, Guanbin Li, Mingzhi Mao, Liang Lin 0004
IEEE Trans. Ind. Informatics6
2022 Dual Adversarial Adaptation for Cross-Device Real-World Image Super-Resolution
abstract
Due to the sophisticated imaging process, an identical scene captured by different cameras could exhibit distinct imaging patterns, introducing distinct proficiency among the super-resolution (SR) models trained on images from different devices. In this paper, we investigate a novel and practical task coded cross-device SR, which strives to adapt a real-world SR model trained on the paired images captured by one camera to low-resolution (LR) images captured by arbitrary target devices. The proposed task is highly challenging due to the absence of paired data from various imaging devices. To address this issue, we propose an unsupervised domain adaptation mechanism for real-world SR, named Dual ADversarial Adaptation (DADA), which only requires LR images in the target domain with available real paired data from a source camera. DADA employs the Domain-Invariant Attention (DIA) module to establish the basis of target model training even without HR supervision. Furthermore, the dual framework of DADA facilitates an Inter-domain Adversarial Adaptation (InterAA) in one branch for two LR input images from two domains, and an Intra-domain Adversarial Adaptation (IntraAA) in two branches for an LR input image. InterAA and IntraAA together improve the model transferability from the source domain to the target. We empirically conduct experiments under six$\text{Real} \rightarrow \text{Real}$adaptation settings among three different cameras, and achieve superior performance compared with existing state-of-the-art approaches. We also evaluate the proposed DADA to address the adaptation to the video camera, which presents a promising re-search topic to promote the wide applications of real-world super-resolution. Our source code is publicly available at https://github.com/lonelyhopeIDADA.
Xiaoqian Xu, Pengxu Wei, Weikai Chen 0001, Yang Liu 0267, Mingzhi Mao, Liang Lin 0004, Guanbin Li
CVPR5
2022 Multimodal Crowd Counting with Mutual Attention Transformers
abstract
Crowd counting is a fundamental yet challenging task that aims to automatically estimate the number of people in crowded scenes. Nowadays, with the rapid development of thermal and depth sensors, thermal images and depth maps become more accessible, which are proven to be beneficial information in boosting the performance of crowd counting. Consequently, we propose a Mutual Attention Transformer (MAT) module to fully leverage the complementary information of different modalities. Specifically, our MAT employs a cross-modal mutual attention mechanism to utilize the features of one modality to enhance the features of the other. Moreover, to improve performance by learning better visual representation and further exploiting modality-wise comple-mentarity, we design a self-supervised pre-training method based on cross-modal image reconstruction. Extensive experiments on two standard benchmarks (i.e., RGBT-CC and ShanghaiTechRGBD) show that the proposed method is effective and universal for multimodal crowd counting, outper-forming previous state-of-the-art methods.
Zhengtao Wu, Lingbo Liu, Mingzhi Mao, Liang Lin 0004, Guanbin Li
ICME4
2022 HairGAN: Spatial-Aware Palette GAN for Hair Color Transfer
abstract
Hair color transfer aims to transfer the hair color from a reference image to an original image while maintaining the hair structure of the original image. However, due to the complex hair structure and the misalignment of hair regions between the original image and the reference image, existing methods cannot complete this task well. To address these issues, we propose a hair color transfer GAN (HairGAN). It first utilizes off-the-shelf network to extract the hair region from the original image and transfer it into intermediary hair color. Then, Spatial-Aware Palette Alignment (SAPA) module is introduced to align the hair regions in the original image and the reference image. Furthermore, HairGAN employs cycle consistency reconstruction module to ensure the global consistency between original image and transferred image. We collect a dataset containing 2850 images and a few paired samples for hair color transfer. Extensive experiments show that HairGAN could generate high-quality transferred results.
Yinjia Huang, Ricong Huang, Zhengtao Wu, Mingzhi Mao, Guanbin Li
ICME5
2022 A Cerebellum-Inspired Model-Free Kinematic Control Method with RCM Constraint
Xin Wang 0164, Peng Yu 0003, Mingzhi Mao, Ning Tan 0003
ICONIP (2)3
2022 7-Instant Discrete-Time Synthesis Model Solving Future Different-Level Linear Matrix System via Equivalency of Zeroing Neural Network
abstract
Differing from the common linear matrix equation, the future different-level linear matrix system is considered, which is much more interesting and challenging. Because of its complicated structure and future-computation characteristic, traditional methods for static and same-level systems may not be effective on this occasion. For solving this difficult future different-level linear matrix system, the continuous different-level linear matrix system is first considered. On the basis of the zeroing neural network (ZNN), the physical mathematical equivalency is thus proposed, which is called ZNN equivalency (ZE), and it is compared with the traditional concept of mathematical equivalence. Then, on the basis of ZE, the continuous-time synthesis (CTS) model is further developed. To satisfy the future-computation requirement of the future different-level linear matrix system, the 7-instant discrete-time synthesis (DTS) model is further attained by utilizing the high-precision 7-instant Zhang et al. discretization (ZeaD) formula. For a comparison, three different DTS models using three conventional ZeaD formulas are also presented. Meanwhile, the efficacy of the 7-instant DTS model is testified by the theoretical analyses. Finally, experimental results verify the brilliant performance of the 7-instant DTS model in solving the future different-level linear matrix system.
Min Yang 0010, Yunong Zhang, Ning Tan 0003, Mingzhi Mao, Haifeng Hu 0001
IEEE Trans. Cybern.4
2022 VQAMix: Conditional Triplet Mixup for Medical Visual Question Answering
abstract
Medical visual question answering (VQA) aims to correctly answer a clinical question related to a given medical image. Nevertheless, owing to the expensive manual annotations of medical data, the lack of labeled data limits the development of medical VQA. In this paper, we propose a simple yet effective data augmentation method, VQAMix, to mitigate the data limitation problem. Specifically, VQAMix generates more labeled training samples by linearly combining a pair of VQA samples, which can be easily embedded into any visual-language model to boost performance. However, mixing two VQA samples would construct new connections between images and questions from different samples, which will cause the answers for those new fabricated image-question pairs to be missing or meaningless. To solve the missing answer problem, we first develop the Learning with Missing Labels (LML) strategy, which roughly excludes the missing answers. To alleviate the meaningless answer issue, we design the Learning with Conditional-mixed Labels (LCL) strategy, which further utilizes language-type prior to forcing the mixed pairs to have reasonable answers that belong to the same category. Experimental results on the VQA-RAD and PathVQA benchmarks show that our proposed method significantly improves the performance of the baseline by about 7% and 5% on the averaging result of two backbones, respectively. More importantly, VQAMix could improve confidence calibration and model interpretability, which is significant for medical VQA models in practical applications. All code and models are available at https://github.com/haifangong/VQAMix.
Haifan Gong, Guanqi Chen, Mingzhi Mao, Zhen Li 0026, Guanbin Li
IEEE Trans. Medical Imaging3
2021 MemoryPath: A deep reinforcement learning framework for incorporating memory component into knowledge graph reasoning
Shuangyin Li, Mingzhi Mao
Neurocomputing4
2020 Secure halftone image steganography with minimizing the distortion on pair swapping
Wanteng Liu, Xiaolin Yin, Wei Lu 0001, Junhong Zhang, Jinhua Zeng, Shaopei Shi, Mingzhi Mao
Signal Process.7
2020 A new goal ordering for incremental planning
Ruishi Liang, Mingzhi Mao
J. Supercomput.2
2020 Continuous and Discrete Zeroing Neural Network for Different-Level Dynamic Linear System With Robot Manipulator Control
abstract
Different-level dynamic linear system (DLDLS) is an interesting and challenging topic due to its complicated structure and time-variant characteristic. To solve this difficult problem, the equivalency of solutions at different levels is analyzed and obtained via zeroing neural network (ZNN) method. Based on the equivalency, a continuous ZNN model is proposed to solve the continuous DLDLS. For easier hardware realization, a new Zhang et al. discretization formula with high precision is proposed for the continuous ZNN model discretization, and the corresponding new discrete ZNN (NDZNN) model is proposed to solve discrete DLDLS. Note that the NDZNN model satisfies the requirement of real-time computation because it has the online ability to predict the solution for the future instant. Furthermore, the problems of robot manipulator control with additional restrictions (e.g., joint damage) are formulated as specific discrete DLDLS, and the proposed NDZNN model is employed to solve such problems. Simulation results substantiate the effectiveness of NDZNN model.
Jian Li 0018, Yunong Zhang, Mingzhi Mao
IEEE Trans. Syst. Man Cybern. Syst.3
2019 Incorporating Graph Attention Mechanism into Knowledge Graph Reasoning Based on Deep Reinforcement Learning
abstract
Heng Wang, Shuangyin Li, Rong Pan, Mingzhi Mao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shuangyin Li, Mingzhi Mao
EMNLP/IJCNLP (1)4
2019 Five-instant type discrete-time ZND solving discrete time-varying linear system, division and quadratic programming
Jian Li 0018, Yunong Zhang, Mingzhi Mao
Neurocomputing3
2019 General Square-Pattern Discretization Formulas via Second-Order Derivative Elimination for Zeroing Neural Network Illustrated by Future Optimization
abstract
Previous works provide a few effective discretization formulas for zeroing neural network (ZNN), of which the precision is a square pattern. However, those formulas are separately developed via many relatively blind attempts. In this paper, general square-pattern discretization (SPD) formulas are proposed for ZNN via the idea of the second-order derivative elimination. All existing SPD formulas in previous works are included in the framework of the general SPD formulas. The connections and differences of various general formulas are also discussed. Furthermore, the general SPD formulas are used to solve future optimization under linear equality constraints, and the corresponding general discrete ZNN models are proposed. General discrete ZNN models have at least one parameter to adjust, thereby determining their zero stability. Thus, the parameter domains are obtained by restricting zero stability. Finally, numerous comparative numerical experiments, including the motion control of a PUMA560 robot manipulator, are provided to substantiate theoretical results and their superiority to conventional Euler formula.
Jian Li 0018, Yunong Zhang, Mingzhi Mao
IEEE Trans. Neural Networks Learn. Syst.3
2018 New Discretization-Formula-Based Zeroing Dynamics for Real-Time Tracking Control of Serial and Parallel Manipulators
abstract
Improvement of the real-time performance of tracking control is increasingly desirable. It is a routine for most conventional algorithms that the control input at current time instant is to track the current desired output. However, lagging errors resulting from computational time and the fluctuation of the desired output exist for the tracking control. Different from conventional algorithms, a look-ahead scheme of zeroing dynamics (ZD) is established in this paper to achieve the real-time tracking control of both serial and parallel manipulators. With the exploitation of data at current time and that in history, the control inputs generated by the proposed ZD algorithms never lead to lagging errors with the source from the inevitable computational time. To tackle prediction errors for ZD algorithms, a new high-precision discretization formula, as an essential part of ZD algorithms, is presented to confine the prediction error in an ignorable range in comparison with lagging errors.
Jian Li 0018, Yunong Zhang, Shuai Li 0002, Mingzhi Mao
IEEE Trans. Ind. Informatics4
2017 Recurrent Attentional Topic Model
abstract
In a document, the topic distribution of a sentence depends on both the topics of preceding sentences and its own content, and it is usually affected by the topics of the preceding sentences with different weights. It is natural that a document can be treated as a sequence of sentences. Most existing works for Bayesian document modeling do not take these points into consideration. To fill this gap, we propose a Recurrent Attentional Topic Model (RATM) for document embedding. The RATM not only takes advantage of the sequential orders among sentence but also use the attention mechanism to model the relations among successive sentences. In RATM, we propose a Recurrent Attentional Bayesian Process (RABP) to handle the sequences. Based on the RABP, RATM fully utilizes the sequential information of the sentences in a document. Experiments on two copora show that our model outperforms state-of-the-art methods on document modeling and classification.
Shuangyin Li, Yu Zhang 0006, Mingzhi Mao, Yang Yang 0002
AAAI4
2016 Enhanced discrete-time Zhang neural network for time-variant matrix inversion in the presence of bias noises
Mingzhi Mao, Jian Li 0018, Long Jin 0001, Shuai Li 0002, Yunong Zhang
Neurocomputing1
2015 Common nature of learning between BP-type and Hopfield-type neural networks
Dongsheng Guo 0001, Yunong Zhang, Zhengli Xiao, Mingzhi Mao, Jianxi Liu
Neurocomputing4
2007 A Method of Software Structure Designing Based on Graph Planning
Mingzhi Mao, Yunfei Jiang, Xiaolong Chai
SoMeT1