Mingyu Fan

dblp:79/950 · DBLP profile ↗
← Back
63ranked-venue papers
13as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 2 since 2021Computer networks · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive drift aware continual learning with pool-based model reuse for time series applications
abstract
Unsupervised anomaly detection in streaming time series is challenging due to evolving data distributions and the lack of labeled anomalies, where concept drift can significantly degrade detection accuracy during long-term operation. Most existing adaptive methods rely on uniform update strategies or continuous incremental learning. Such designs can be slow to react to abrupt distribution shifts and are prone to overwriting previously learned regimes when patterns recur. This paper proposes ADAPTS (Adaptive Drift-Aware Pool-based framework for Time Series), an unsupervised and drift-aware framework for adaptive anomaly detection in non-stationary data streams. ADAPTS integrates statistical drift detection, explicit drift-type classification, drift-specific model adaptation, and a bounded model pool for knowledge reuse within a unified system. The framework distinguishes sudden, incremental, and recurrent concept drift and aligns adaptation strategies accordingly through retraining, controlled fine-tuning, or selective model reuse. Experiments on benchmark datasets show that ADAPTS can keep stable and reliable detection results under different types of drift. Experimental results show that ADAPTS achieves competitive performance across multiple datasets, with performance improvements observed under several drift scenarios, including up to 0.11 absolute gain in AUC on challenging scenarios such as the NAB dataset, while maintaining improved long-term stability under diverse drift conditions.
Danlei Li, Mingyu Fan, Nirmal-Kumar C. Nair, Kevin I-Kai Wang
Knowl. Based Syst.2
2026 Answer-syntax guided handwritten mathematical expression recognition for intelligent education
Mingyu Fan, Xiangshu Ruan, Xiangli Nie, Jianan Sun, Pengcheng Qian, Qianjin Chen, Matthias Rätsch
Pattern Recognit.1
2026 Collaborative Observation Imputation and Trajectory Prediction via Consistency Evaluation
abstract
Pedestrian trajectory prediction is a critical task in various mobile computing applications, such as video surveillance, robot navigation, autonomous driving, and human mobility analysis. Although significant progress has been made by current methods, the challenge of observation deficiency in pedestrian trajectory prediction remains largely unaddressed. Since most existing methods focus on optimizing prediction accuracy under the assumption of complete observations, while ignoring the potential for observation deficiency caused by failures in detection or tracking algorithms. To overcome this challenge, we propose a collaborative observation imputation and trajectory prediction framework, which employs consistency evaluation to jointly perform the imputation and prediction tasks. Specifically, we first build a consistency evaluation module to align features between observed and future trajectory pairs using contrastive learning. Then, we design a trajectory imputation and prediction baseline, which adopts a parallel paradigm, to mitigate the impact of coarse imputations on trajectory prediction when performing initial imputation and prediction. Next, we introduce a consistency-guided Skip-Diffusion module, which leverages consistency evaluation between initial imputations and ground truth future trajectories to refine the initial imputations. Finally, we propose a consistency-driven Cross-Mamba module, which uses consistency evaluation between ground-truth observations and initial predictions to refine the initial predictions. Extensive experiments demonstrate the effectiveness of the proposed framework in both imputation and prediction tasks.
Hao Zhou 0014, Mingyu Fan, Xu Yang 0004, Hai Huang 0004, Kaishun Wu, Lu Wang 0002, Fei Luo 0003
IEEE Trans. Mob. Comput.2
2025 AURORA: An Adaptive and Unsupervised Framework for Robust Anomaly Detection via Historical Model Reuse
Danlei Li, Mingyu Fan, Nirmal-Kumar C. Nair, Akshat Bisht, Kevin I-Kai Wang
IEEE Big Data2
2025 Three-Dimensional Trajectory Prediction with 3DMoTraj Dataset
abstract
With the growing interest in embodied and spatial intelligence, accurately predicting trajectories in 3D environments has become increasingly critical. However, no datasets have been explicitly designed to study 3D trajectory prediction. To this end, we contribute a 3D motion trajectory (3DMoTraj) dataset collected from unmanned underwater vehicles (UUVs) operating in oceanic environments. Mathematically, trajectory prediction becomes significantly more complex when transitioning from 2D to 3D. To tackle this challenge, we analyze the prediction complexity of 3D trajectories and propose a new method consisting of two key components: decoupled trajectory prediction and correlated trajectory refinement. The former decouples inter-axis correlations, thereby reducing prediction complexity and generating coarse predictions. The latter refines the coarse predictions by modeling their inter-axis correlations. Extensive experiments show that our method significantly improves 3D trajectory prediction accuracy and outperforms state-of-the-art methods. Both the 3DMoTraj dataset and the method are available at https://github.com/zhouhao94/3DMoTraj.
Hao Zhou 0014, Xu Yang 0004, Mingyu Fan, Lu Qi 0001, Xiangtai Li, Ming-Hsuan Yang 0001, Fei Luo 0003
ICML3
2025 STAF: Symbol-Targeted Adversarial Flow for Handwritten Mathematical Expression Recognition
Xiangshu Ruan, Mingyu Fan, Yijian Wu, Dan Cheng
PRCV (7)2
2025 NVP-HRI: Zero shot natural voice and posture-based human-robot interaction via large language model
abstract
Effective Human–Robot Interaction (HRI) is crucial for future service robots in aging societies . Existing solutions are biased towards only well-trained objects, creating a gap when dealing with new objects. Currently, HRI systems using predefined gestures or language tokens for pretrained objects pose challenges for all individuals, especially elderly ones. These challenges include difficulties in recalling commands, memorizing hand gestures, and learning new names. This paper introduces NVP-HRI, an intuitive multi-modal HRI paradigm that combines voice commands and deictic posture. NVP-HRI utilizes the Segment Anything Model (SAM) to analyze visual cues and depth data, enabling precise structural object representation. Through a pre-trained SAM network, NVP-HRI allows interaction with new objects via zero-shot prediction, even without prior knowledge . NVP-HRI also integrates with a large language model (LLM) for multimodal commands, coordinating them with object selection and scene distribution in real time for collision-free trajectory solutions. We also regulate the action sequence with the essential control syntax to reduce LLM hallucination risks. The evaluation of diverse real-world tasks using a Universal Robot showcased up to 59.2% efficiency improvement over traditional gesture control, as illustrated in the video https://youtu.be/EbC7al2wiAc . Our code and design will be openly available at https://github.com/laiyuzhi/NVP-HRI.git .
Yuzhi Lai, Shenghai Yuan 0001, Youssef Nassar, Mingyu Fan, Matthias Rätsch
Expert Syst. Appl.4
2024 Safety Classes Semantic Costmaps from RGBD Sensors for Risk-Aware Robot Navigation
abstract
For collision and obstacle avoidance in path planning, robots usually rely on basic 2D cost maps lacking semantic information about detected obstacles. As a result, the robot’s path planning follows an arbitrarily large safety margin around obstacles.We present a risk-aware 2D cost map for robot navigation that effectively mitigates potential risks, enabling the robot to navigate more confidently and efficiently while maintaining a safe distance from obstacles. It uses commonly available RGBD sensors, making it a practical and accessible option for many applications.Our approach employs a CNN to segment object instances into three distinct safety classes based on common characteristics. Using the abstract representation, we generate semantic occupancy grids, which are then inflated based on their safety classification. These semantic occupancy grids are merged into a final risk-aware 2D cost map. We provide feasible real world results.A robot’s path planner can use the merged cost map as a drop-in replacement for standard 2D cost map, enhancing established robot navigation algorithms based on cost maps without altering their fundamental algorithms and enabling risk-aware navigation in the real world.
Shival Indermun, Mingyu Fan, Lianghong Wu, Kristiaan Schreve, Matthias Rätsch
CoDIT3
2024 DoS/DDoS attacks in Software Defined Networks: Current situation, challenges and future directions
Mohamed Ali Setitra, Mingyu Fan, Ilyas Benkhaddra, Zine El Abidine Bensalem
Comput. Commun.2
2024 Learning to match features with discriminative sparse graph neural network
Junxiong Cai, Mingyu Fan, Wensen Feng
Pattern Recognit.3
2024 Online Active Continual Learning for Robotic Lifelong Object Recognition
abstract
In real-world applications, robotic systems collect vast amounts of new data from ever-changing environments over time. They need to continually interact and learn new knowledge from the external world to adapt to the environment. Particularly, lifelong object recognition in an online and interactive manner is a crucial and fundamental capability for robotic systems. To meet this realistic demand, in this article, we propose an online active continual learning (OACL) framework for robotic lifelong object recognition, in the scenario of both classes and domains changing with dynamic environments. First, to reduce the labeling cost as much as possible while maximizing the performance, a new online active learning (OAL) strategy is designed by taking both the uncertainty and diversity of samples into account to protect the information volume and distribution of data. In addition, to prevent catastrophic forgetting and reduce memory costs, a novel online continual learning (OCL) algorithm is proposed based on the deep feature semantic augmentation and a new loss-based deep model and replay buffer update, which can mitigate the class imbalance between the old and new classes and alleviate confusion between two similar classes. Moreover, the mistake bound of the proposed method is analyzed in theory. OACL allows robots to select the most representative new samples to query labels and continually learn new objects and new variants of previously learned objects from a nonindependent and identically distributed (i.i.d.) data stream without catastrophic forgetting. Extensive experiments conducted on real lifelong robotic vision datasets demonstrate that our algorithm, even trained with fewer labeled samples and replay exemplars, can achieve state-of-the-art performance on OCL tasks.
Xiangli Nie, Zhiguang Deng, Mingdong He, Mingyu Fan
IEEE Trans. Neural Networks Learn. Syst.4
2023 Static-dynamic global graph representation for pedestrian trajectory prediction
Hao Zhou 0014, Xu Yang 0004, Mingyu Fan, Hai Huang 0004, Dongchun Ren, Huaxia Xia
Knowl. Based Syst.3
2023 CSR: Cascade Conditional Variational Auto Encoder with Socially-aware Regression for Pedestrian Trajectory Prediction
Hao Zhou 0014, Dongchun Ren, Xu Yang 0004, Mingyu Fan, Hai Huang 0004
Pattern Recognit.4
2023 CSIR: Cascaded Sliding CVAEs With Iterative Socially-Aware Rethinking for Trajectory Prediction
abstract
Pedestrian trajectory prediction is a hot research topic in many applications, such as video surveillance and autonomous driving. Although many efforts have been done on this topic, there are still many challenges, including accumulated prediction errors, insufficient training data usage, and future-past incompatibility. To overcome these challenges, we propose a novel trajectory prediction method, called CSIR, which consists of a cascaded sliding conditional variational autoencoder (CS-CVAE) module and an iterative future-past social compatible rethinking (I-SCR) module. The CS-CVAE module reduces the accumulated prediction errors by using cascaded prediction models for the early future time steps. In this way, the training losses of the early time steps are separately considered and minimized from the later losses. For the following time steps in CS-CVAE, a sliding prediction model with a longer observation time span is used and additional data from the future time span can be collected for training. On the other hand, the I-SCR module generates offsets to improve the predictions iteratively by checking the interaction compatibility between the predicted trajectories and the past trajectories, which resembles with the human rethinking mechanism in motion planning. Experiments results on two widely explored pedestrian trajectory prediction datasets, Stanford Drone Dataset (SDD) and ETH/UCY, show that the proposed method surpasses previous state-of-the-art methods by notable margins.
Hao Zhou 0014, Xu Yang 0004, Dongchun Ren, Hai Huang 0004, Mingyu Fan
IEEE Trans. Intell. Transp. Syst.5
2022 Privacy Risk Assessment of Training Data in Machine Learning
abstract
With the wide application of machine learning technology, more and more sensitive data were used to develop the machine learning model. Unfortunately, machine learning systems have been shown to be vulnerable to various privacy leakage attacks, such as membership inference attacks and model inversion attacks. These privacy threats break down the confidentiality of the training data; what’s more, the leakage of sensitive privacy will reduce the enthusiasm of data owners for sharing data with the machine learning system. Without enough training data, more seriously, it will hinder the application of machine learning. Therefore, it requires analyzing strategies for privacy risks to help data owners pre-assess the candidate dataset and assist data owners in implementing reasonable privacy control. However, the systematic privacy risk assessment is still absent from the data owner’s perspective.This paper investigates and analyzes machine learning privacy risks to understand the relationship between training data properties and privacy leakage. Based on this recognition, we introduce a privacy risk assessment scheme based on the clustering distance of training data. Our clustering distance-based method can reflect the privacy risk level of a different individual data record. And then, we combine both existing privacy analysis based on data features and our clustering distance-based method to investigate privacy risks systematically. Our experiments showed that our clustering distance method and other set properties are tightly related to privacy leakage. And data owners can pre-assess the privacy risks of candidate datasets before uploading or sharing them to the machine learning model. If needed, decide to re-choose the dataset to reduce privacy risks to an acceptable level.
Mingyu Fan, Chuangmin Xie
ICC2
2022 Safety-based Reinforcement Learning Longitudinal Decision for Autonomous Driving in Crosswalk Scenarios
abstract
Autonomous vehicles (AVs) need to make driving decisions to interact with other traffic participants. By adapting to different scenarios with specific parameters, traditional strategies attempt to leverage rule-based methods to solve the decision problems. In this paper, we present a novel reinforcement learning method for resolving interaction uncertainty in the decision-making problem. We construct prior knowledge by introducing traffic regulations and constraints and then converting them into rules that govern the learning of driving policies. To promote safe driving, a safety-aware module equipped with a mathematical collision correlation analysis is developed to anticipate and handle dangerous traffic scenarios. A realistic scenario involving an AV approaching a crosswalk is used to validate the proposed method. The experimental results indicate that the proposed method improves driving safety and efficiency significantly when compared to alternative approaches and can be generalized to more difficult scenarios.
Fangzhou Xiong, Dongchun Ren, Mingyu Fan, Shuguang Ding, Zhiyong Liu 0001
IJCNN3
2022 Online Semisupervised Active Classification for Multiview PolSAR Data
abstract
Polarimetric synthetic aperture radar (PolSAR) data are sequentially acquired and have multiple views obtained from different feature extractors or multiple frequency bands. The fast and accurate classification of PolSAR data in dynamically changing environments is a critical and challenging task. Online learning can handle this task by learning a classifier incrementally from a stream of samples. In this article, we propose an online semisupervised active learning framework for multiview PolSAR data classification, called OSAM. First, a novel online active learning strategy is designed based on the relationships among multiple views and a randomized rule, which allows to only query the labels of some informative incoming samples. Then, in order to utilize both the incoming labeled and unlabeled samples to update the classifiers, a novel online semisupervised learning model is proposed based on co-regularized multiview learning and graph regularization. In addition, the proposed method can deal with the dynamic large-scale multifeature or multifrequency PolSAR data where not only the amount of data but also the number of classes gradually increases in the learning process. Moreover, the mistake bound of the proposed method is derived rigorously. Extensive experiments are conducted on real PolSAR data to evaluate the performance of our algorithm, and the results demonstrate the effectiveness of the proposed method.
Xiangli Nie, Mingyu Fan, Xiayuan Huang, Bo Zhang 0006, Xiaoshuang Ma
IEEE Trans. Cybern.2
2022 Simultaneous Past and Current Social Interaction-aware Trajectory Prediction for Multiple Intelligent Agents in Dynamic Scenes
abstract
Trajectory prediction of multiple agents in a crowded scene is an essential component in many applications, including intelligent monitoring, autonomous robotics, and self-driving cars. Accurate agent trajectory prediction remains a significant challenge because of the complex dynamic interactions among the agents and between them and the surrounding scene. To address the challenge, we propose a decoupled attention-based spatial-temporal modeling strategy in the proposed trajectory prediction method. The past and current interactions among agents are dynamically and adaptively summarized by two separate attention-based networks and have proven powerful in improving the prediction accuracy. Moreover, it is optional in the proposed method to make use of the road map and the plan of the ego-agent for scene-compliant and accurate predictions. The road map feature is efficiently extracted by a convolutional neural network, and the features of the ego-agent’s plan is extracted by a gated recurrent network with an attention module based on the temporal characteristic. Experiments on benchmark trajectory prediction datasets demonstrate that the proposed method is effective when the ego-agent plan and the the surrounding scene information are provided and achieves state-of-the-art performance with only the observed trajectories.
Yanliang Zhu, Dongchun Ren, Yi Xu 0005, Deheng Qian, Mingyu Fan, Huaxia Xia
ACM Trans. Intell. Syst. Technol.5
2022 Adaptive Data Structure Regularized Multiclass Discriminative Feature Selection
abstract
Feature selection (FS), which aims to identify the most informative subset of input features, is an important approach to dimensionality reduction. In this article, a novel FS framework is proposed for both unsupervised and semisupervised scenarios. To make efficient use of data distribution to evaluate features, the framework combines data structure learning (as referred to as data distribution modeling) and FS in a unified formulation such that the data structure learning improves the results of FS and vice versa. Moreover, two types of data structures, namely the soft and hard data structures, are learned and used in the proposed FS framework. The soft data structure refers to the pairwise weights among data samples, and the hard data structure refers to the estimated labels obtained from clustering or semisupervised classification. Both of these data structures are naturally formulated as regularization terms in the proposed framework. In the optimization process, the soft and hard data structures are learned from data represented by the selected features, and then, the most informative features are reselected by referring to the data structures. In this way, the framework uses the interactions between data structure learning and FS to select the most discriminative and informative features. Following the proposed framework, a new semisupervised FS (SSFS) method is derived and studied in depth. Experiments on real-world data sets demonstrate the effectiveness of the proposed method.
Mingyu Fan, Xiaoqin Zhang 0002, Jie Hu 0041, Nannan Gu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2021 GANMIA: GAN-based Black-box Membership Inference Attack
abstract
Membership inference attacks (MIAs) against machine learning systems have drawn tremendous attention from information security researchers. By MIA, an adversary can speculate whether an individual data record is a member of the training set or not. Existing black-box MIA assumes that much information about the training data is available. Specifically, the attacker assumes that (s)he has the ability to query the target model without limitations or can access a sufficient dataset whose distribution is the same as the training data set. However, in a realistic scenario, MIAs usually come up with the limited number and the imbalanced proportion of target training datasets which cause significant challenges for MIAs. To launch an MIA in the realistic scenario, in this paper, we present a novel method called GANMIA, which generates synthetic data to augment the training samples of the shadow model for the black-box MIA by a Generative Adversarial Network (GAN). GANMIA firstly augments synthesized samples and then uses the generated samples to train the given shadow model to increase the training efficiency, and additionally improve the MIA’s performance. The experimental results show that the accuracy of the black-box MIA increases by 23% with the help of our synthetic data.
Yang Bai 0011, Degang Chen 0003, Ting Chen 0002, Mingyu Fan
ICC4
2021 Visionnet: A Coarse-To-Fine Motion Prediction Algorithm Based On Active Interaction-Aware Drivable Space Learning
abstract
Trajectory prediction is a fundamental task in many real applications such as autonomous robotics and video surveillance. In this paper, we propose a novel vision-based trajectory prediction method which is able to extract the interactive features by active global interaction-aware drivable space learning. The learned global interaction-aware drivable spaces denote the areas with low occupation probabilities, which provide the regions and directions that the agents can move into. Specifically, our method describes a sequence of motion states, i.e. the location, the velocity and the acceleration, as occupancy grid maps, and then use them to train the deep learning model in the supervised manner. Moreover, an interactive loss for training the inference net of drivable spaces and trajectory prediction net simultaneously is introduced. Comparisons with state-of-the-art methods on benchmark datasets demonstrate the effectiveness of the proposed method.
Dongchun Ren, Yanliang Zhu, Mingyu Fan, Deheng Qian, Huaxia Xia, Zhuang Fu
ICME3
2021 Robust Trajectory Prediction of Multiple Interacting Pedestrians via Incremental Active Learning
Yi Xi, Dongchun Ren, Mingxia Li, Yuehai Chen, Mingyu Fan, Huaxia Xia
ICONIP (5)5
2021 Star Topology based Interaction for Robust Trajectory Forecasting in Dynamic Scene
abstract
Motion prediction of multiple agents in a dynamic scene is a crucial component in many real applications, including intelligent monitoring and autonomous driving. Due to the complex interactions among the agents and their interactions with the surrounding scene, accurate trajectory prediction is still a great challenge. In this paper, we propose a new method for robust trajectory prediction of multiple intelligent agents in a dynamic scene. The input of the method includes the observed trajectories of all agents, and optionally, the planning of the ego-agent and the surrounding high definition map at every time steps. Given observed trajectories, an efficient approach in a star computational topology is utilized to compute both the spatiotemporal interaction features and the current interaction features between the agents, where the time complexity scales linearly to the number of agents. Moreover, on an autonomous vehicle, the proposed prediction method can make use of the planning of ego-agent to improve the modeling of the interaction between surrounding agents. To increase the robustness to upstream perception noises, at the training stage, we randomly mask out the input data, a.k.a. the points on the observed trajectories of agents and the lane sequence. Experiments on autonomous driving and pedestrian-walking datasets demonstrate that the proposed method is not only effective when the planning of ego-agent and the high definition map are provided, but also achieves state-of-the-art performance with only the observed trajectories.
Yanliang Zhu, Dongchun Ren, Deheng Qian, Mingyu Fan, Huaxia Xia
ICRA4
2021 An Effective Dynamic Membership Authentication and Key Management Scheme in Wireless Sensor Networks
abstract
Wireless sensor networks (WSN) have been widely used in the field of industrial Internet of Things (IoT), one of the main security challenges is how to protect the transmission of sensitive data between sensors in the wireless channel. Key management is one of the most challenging work. Most IoT sensors' resources are limited, so the key management scheme should be designed to be as lightweight as possible. Generally, schemes can be divided into two categories: pairwise key management and group key management, but each of them has its limitations. In this paper, we propose a dynamic membership authentication and key management scheme for WSN. In our scheme, we add the access object authentication and key update mechanism, ensure the authenticity of the connected object and key freshness. Compared with other schemes, our solution ensures forward and backward secrecy and resists capture attacks. We finally demonstrate that our scheme is of confidentiality, integrity, and scalability for resource-constrained WSN.
Mingyu Fan, Maoyang Chen, Xingyu Liao, Wenqiang Hu
WCNC2
2021 Graph matching based point correspondence with alternating direction method of multipliers
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001, Mingyu Fan
Neurocomputing5
2021 AST-GNN: An attention-based spatio-temporal graph neural network for Interaction-aware pedestrian trajectory prediction
Hao Zhou 0014, Dongchun Ren, Huaxia Xia, Mingyu Fan, Xu Yang 0004, Hai Huang 0004
Neurocomputing4
2021 AdvSGAN: Adversarial image Steganography with adversarial networks
Lin Li 0082, Mingyu Fan
Multim. Tools Appl.2
2021 A Defense Framework for Privacy Risks in Remote Machine Learning Service
abstract
In recent years, machine learning approaches have been widely adopted for many applications, including classification. Machine learning models deal with collective sensitive data usually trained in a remote public cloud server, for instance, machine learning as a service (MLaaS) system. In this scene, users upload their local data and utilize the computation capability to train models, or users directly access models trained by MLaaS. Unfortunately, recent works reveal that the curious server (that trains the model with users’ sensitive local data and is curious to know the information about individuals) and the malicious MLaaS user (who abused to query from the MLaaS system) will cause privacy risks. The adversarial method as one of typical mitigation has been studied by several recent works. However, most of them focus on the privacy-preserving against the malicious user; in other words, they commonly consider the data owner and the model provider as one role. Under this assumption, the privacy leakage risks from the curious server are neglected. Differential privacy methods can defend against privacy threats from both the curious sever and the malicious MLaaS user by directly adding noise to the training data. Nonetheless, the differential privacy method will decrease the classification accuracy of the target model heavily. In this work, we propose a generic privacy-preserving framework based on the adversarial method to defend both the curious server and the malicious MLaaS user. The framework can adapt with several adversarial algorithms to generate adversarial examples directly with data owners’ original data. By doing so, sensitive information about the original data is hidden. Then, we explore the constraint conditions of this framework which help us to find the balance between privacy protection and the model utility. The experiments’ results show that our defense framework with the AdvGAN method is effective against MIA and our defense framework with the FGSM method can protect the sensitive data from direct content exposed attacks. In addition, our method can achieve better privacy and utility balance compared to the existing method.
Yang Bai 0011, Mingchuang Xie, Mingyu Fan
Secur. Commun. Networks4
2021 Top-k Feature Selection Framework Using Robust 0-1 Integer Programming
abstract
Feature selection (FS), which identifies the relevant features in a data set to facilitate subsequent data analysis, is a fundamental problem in machine learning and has been widely studied in recent years. Most FS methods rank the features in order of their scores based on a specific criterion and then select the k top-ranked features, where k is the number of desired features. However, these features are usually not the top- k features and may present a suboptimal choice. To address this issue, we propose a novel FS framework in this article to select the exact top- k features in the unsupervised, semisupervised, and supervised scenarios. The new framework utilizes thel0,2-norm as the matrix sparsity constraint rather than its relaxations, such as thel1,2-norm. Since thel0,2-norm constrained problem is difficult to solve, we transform the discretel0,2-norm-based constraint into an equivalent 0-1 integer constraint and replace the 0-1 integer constraint with two continuous constraints. The obtained top- k FS framework with two continuous constraints is theoretically equivalent to thel0,2-norm constrained problem and can be optimized by the alternating direction method of multipliers (ADMM). Unsupervised and semisupervised FS methods are developed based on the proposed framework, and extensive experiments on real-world data sets are conducted to demonstrate the effectiveness of the proposed FS framework.
Xiaoqin Zhang 0002, Mingyu Fan, Di Wang 0008, Peng Zhou 0006, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.2
2021 Self-Paced Clustering Ensemble
abstract
The clustering ensemble has emerged as an important extension of the classical clustering problem. It provides an elegant framework to integrate multiple weak base clusterings to generate a strong consensus result. Most existing clustering ensemble methods usually exploit all data to learn a consensus clustering result, which does not sufficiently consider the adverse effects caused by some difficult instances. To handle this problem, we propose a novel self-paced clustering ensemble (SPCE) method, which gradually involves instances from easy to difficult ones into the ensemble learning. In our method, we integrate the evaluation of the difficulty of instances and ensemble learning into a unified framework, which can automatically estimate the difficulty of instances and ensemble the base clusterings. To optimize the corresponding objective function, we propose a joint learning algorithm to obtain the final consensus clustering result. Experimental results on benchmark data sets demonstrate the effectiveness of our method.
Peng Zhou 0006, Liang Du 0003, Xinwang Liu 0002, Yidong Shen, Mingyu Fan, Xuejun Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2020 Knowledge-Experience Graph with Denoising Autoencoder for Zero-Shot Learning in Visual Cognitive Development
Xu Yang 0004, Zhiyong Liu 0001, Lu Zhang 0054, Dongchun Ren, Mingyu Fan
ICONIP (5)6
2020 An Attention-Based Interaction-Aware Spatio-Temporal Graph Neural Network for Trajectory Prediction
Hao Zhou 0014, Dongchun Ren, Huaxia Xia, Mingyu Fan, Xu Yang 0004, Hai Huang 0004
ICONIP (5)4
2020 Unsupervised feature selection for balanced clustering
Peng Zhou 0006, Jiangyong Chen, Mingyu Fan, Liang Du 0003, Yidong Shen, Xuejun Li 0001
Knowl. Based Syst.3
2020 Customized Network Security for Cloud Service
abstract
Modern cloud computing platforms based on virtual machine monitors (VMMs) host a variety of complex businesses which present many network security vulnerabilities. In order to protect network security for these businesses in cloud computing, nowadays, a number of middleboxes are deployed at front-end of cloud computing or parts of middleboxes are deployed in cloud computing. However, the former is leading to high cost and management complexity, and also lacking of network security protection between virtual machines while the latter does not effectively prevent network attacks from external traffic. To address the above-mentioned challenges, we introduce a novel customized network security for cloud service (CNS), which not only prevents attacks from external and internal traffic to ensure network security of services in cloud computing, but also affords customized network security service for cloud users. CNS is implemented by modifying the Xen hypervisor and proved by various experiments which showing the proposed solution can be directly applied to the extensive practical promotion in cloud computing.
Kaoru Ota, Mianxiong Dong, Laurence T. Yang, Mingyu Fan, Guangwei Wang, Stephen S. Yau
IEEE Trans. Serv. Comput.5
2019 Adaptive flow abnormity identification based on information entropy
abstract
Summary Along with the normal network flow, abnormal flow follows, threatening the security and normal use of the computer. Under the impact of the massive network flow, researchers were prompted to study identification methods based on machine learning. Currently, the existing identification methods based on machine learning still have deficiencies, ie, feature redundancy and deficiency in classifier. In order to deal with these problems, this paper proposes an adaptive approach (AEDs) to identify abnormal network flow. The AEDs utilizes the feature extraction algorithm proposed in this paper, which is based on information entropy to filter features and reduce the size of features. Then, a weight construction method based on rough set theory is implemented to construct the weights of given‐case‐based ensemble classifiers, and only the sub‐classifier with weight that is higher than the given threshold will be reserved. On this basis, in order to keep the effectiveness of the classifier, we utilize new flow data to update weights. Moreover, an adaptive detection algorithm is proposed to adaptively monitor the change of entropy of abnormal flow and enable the update of classifier. Finally, experiments have been conducted to illustrate the effectiveness and the feasibility of the approach.
Mingyu Fan, Guangwei Wang
Concurr. Comput. Pract. Exp.2
2019 Semantic Segmentation of Remote Sensing Images Using Multiscale Decoding Network
abstract
In this letter, we propose a practical convolutional neural network architecture for semantic pixelwise segmentation of remote sensing images, named Multiscale Decoding Network. The proposed method is built on the success of fully convolutional networks (FCNs) and the transfer of pretrained networks. The decoding network of our architecture utilizes the combination of three paths, namely, unpooling path, transposed convolution path, and dilated convolution path, in the form of an inception module. The whole network is trained in the end-to-end manner and the parameters of the three paths are learned automatically. Since the proposed method transfers the feature of pretrained networks and has three simplified decoding paths with fewer parameters, it requires less training data and training time. Compared with the classical networks FCN, SegNet, and U-net, our network shows better performance on remote sensing images segmentation.
Xiaoqin Zhang 0002, Zhiheng Xiao, Mingyu Fan, Li Zhao 0005
IEEE Geosci. Remote. Sens. Lett.4
2019 Structure regularized self-paced learning for robust semi-supervised pattern classification
Nannan Gu, Pengying Fan, Mingyu Fan, Di Wang 0008
Neural Comput. Appl.3
2018 Simultaneous Learning of Affinity Matrix and Laplacian Regularized Least Squares for Semi-Supervised Classification
abstract
Graph based Semi-Supervised Learning (G-SSL) methods usually include the stages of the construction of affinity matrix and the mechanism of inferring unknown labels. However, solving each of the stages individually does not fully exploit the potential relationship between the affinity matrix and the labels of samples. In this paper, we formulate the global self-expressiveness induced affinity and Laplacian Regularized Least Squares (LapRLS) into a single optimization model, called as Self-Taught LapRLS (ST-LapRLS). In the unified model, both the given labels and the estimated labels are used to build a better affinity matrix and to facilitate the LapRLS classifer. The proposed ST-LapRLS classifier is explicit, and can be easily extended to deal with out-of-sample problem. We propose an efficient algorithm which combines the alternating direction method of multiplier and LapRLS to solve the unified optimization problem. Experiments on several Benchmark datasets show the superior performance of our method in classification applications.
Di Wang 0008, Xiaoqin Zhang 0002, Nannan Gu, Mingyu Fan
ICIP5
2018 Semi-Supervised Learning Through Label Propagation on Geodesics
abstract
Graph-based semi-supervised learning (SSL) has attracted great attention over the past decade. However, there are still several open problems in this paper, including: 1) how to construct an effective graph over data with complex distribution and 2) how to define and effectively use pair-wise similarity for robust label propagation. In this paper, we utilize a simple and effective graph construction method to construct the graph over data lying on multiple data manifolds. The method can guarantee the connectiveness between pair-wise data points. Then, the global pair-wise data similarity is naturally characterized by geodesic distance-based joint probability, where the geodesic distance is approximated by the graph distance. The new data similarity is much more effective than previous Euclidean distance-based similarities. To apply data structure for robust label propagation, Kullback-Leibler divergence is utilized to measure the inconsistency between the input pair-wise similarity and the output similarity. In order to further consider intraclass and interclass variances, a novel regularization term on sample-wise margins is introduced to the objective function. This enables the proposed method fully utilizes the input data structure and the label information for classification. An efficient optimization method and the convergence analysis have been proposed for our problem. Besides, out-of-sample extension is discussed and addressed. Comparisons with the state-of-the-art SSL methods on image classification tasks have been presented to show the effectiveness of the proposed method.
Mingyu Fan, Xiaoqin Zhang 0002, Liang Du 0003, Liang Chen 0012, Dacheng Tao
IEEE Trans. Cybern.1
2017 Structure Regularized Unsupervised Discriminant Feature Analysis
abstract
Feature selection is an important technique in machine learning research. An effective and robust feature selection method is desired to simultaneously identify the informative features and eliminate the noisy ones of data. In this paper, we consider the unsupervised feature selection problem which is particularly difficult as there is not any class labels that would guide the search for relevant features. To solve this, we propose a novel algorithmic framework which performs unsupervised feature selection. Firstly, the proposed framework implements structure learning, where the data structures (including intrinsic distribution structure and the data segment) are found via a combination of the alternative optimization and clustering. Then, both the intrinsic data structure and data segmentation are formulated as regularization terms for discriminant feature selection. The results of the feature selection also affect the structure learning step in the following iterations. By leveraging the interactions between structure learning and feature selection, we are able to capture more accurate structure of data and select more informative features. Clustering and classification experiments on real world image data sets demonstrate the effectiveness of our method.
Mingyu Fan, Xiaojun Chang, Dacheng Tao
AAAI1
2017 Top-k Supervise Feature Selection via ADMM for Integer Programming
abstract
Recently, structured sparsity inducing based feature selection has become a hot topic in machine learning and pattern recognition. Most of the sparsity inducing feature selection methods are designed to rank all features by certain criterion and then select the k top ranked features, where k is an integer. However, the k top features are usually not the top k features and therefore maybe a suboptimal result. In this paper, we propose a novel supervised feature selection method to directly identify the top k features. The new method is formulated as a classic regularized least squares regression model with two groups of variables. The problem with respect to one group of the variables turn out to be a 0-1 integer programming, which had been considered very hard to solve. To address this, we utilize an efficient optimization method to solve the integer programming, which first replaces the discrete 0-1 constraints with two continuous constraints and then utilizes the alternating direction method of multipliers to optimize the equivalent problem. The obtained result is the top subset with k features under the proposed criterion rather than the subset of k top features. Experiments have been conducted on benchmark data sets to show the effectiveness of proposed method.
Mingyu Fan, Xiaojun Chang, Xiaoqin Zhang 0002, Di Wang 0008, Liang Du 0003
IJCAI1
2016 Semi-Supervised Dictionary Learning via Structural Sparse Preserving
abstract
While recent techniques for discriminative dictionary learning have attained promising results on the classification tasks, their performance is highly dependent on the number of labeled samples available for training. However, labeling samples is expensive and time consuming due to the significant human effort involved. In this paper, we present a novel semi- supervised dictionary learning method which utilizes the structural sparse relationships between the labeled and unlabeled samples. Specifically, by connecting the sparse reconstruction coefficients on both the original samples and dictionary, the unlabeled samples can be automatically grouped to the different labeled samples, and the grouped samples share a small number of atoms in the dictionary via mixed l2p- norm regularization. This makes the learned dictionary more representative and discriminative since the shared atoms are learned by using the labeled and unlabeled samples potentially from the same class. Minimizing the derived objective function is a challenging task because it is non-convex and highly non-smooth. We propose an efficient optimization algorithm to solve the problem based on the block coordinate descent method. Moreover, we have a rigorous proof of the convergence of the algorithm. Extensive experiments are presented to show the superior performance of our method in classification applications.
Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye
AAAI3
2016 Efficient isometric multi-manifold learning based on the self-organizing method
Mingyu Fan, Xiaoqin Zhang 0002, Hong Qiao, Bo Zhang 0006
Inf. Sci.1
2016 Hierarchical mixing linear support vector machines for nonlinear classification
Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye
Pattern Recognit.3
2016 Robust Semi-Supervised Classification for Noisy Labels Based on Self-Paced Learning
abstract
Data labeling is a tedious and subjective task that can be time consuming and error-prone; however, most learning algorithms are sensitive to noisy labels. This problem raises the need to develop algorithms that can exploit large amount of unlabeled data and also be robust to noisy label information. In this letter, we propose a novel semi-supervised classification framework that is robust to noisy labels, named self-paced manifold regularization. The proposed framework naturally integrates self-paced learning regime into the manifold regularization framework for selecting labeled training samples in a theoretically sound manner, and utilizes locally linear reconstructions to control the smoothness of the classifier with respect to the manifold structure of data. Finally, the alternative search strategy is adopted for the proposed framework to obtain the classifier. The proposed method can not only suppress the negative effect of noisy initial labels in semi-supervised learning, but also obtain an explicit multiclass classifier for newly coming data points. Experimental results demonstrate the effectiveness of the proposed method.
Nannan Gu, Mingyu Fan, Deyu Meng
IEEE Signal Process. Lett.2
2015 Robust Multiple Kernel K-means Using L21-Norm
Liang Du 0003, Peng Zhou 0006, Lei Shi 0015, Hanmo Wang, Mingyu Fan, Yidong Shen
IJCAI5
2015 An Efficient Classifier Based on Hierarchical Mixing Linear Support Vector Machines
Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye
IJCAI3
2015 Multi-Modality Tracker Aggregation: From Generative to Discriminative
Xiaoqin Zhang 0002, Wei Li 0034, Mingyu Fan, Di Wang 0008, Xiuzi Ye
IJCAI3
2015 An LLE based Heterogeneous Metric Learning for Cross-media Retrieval
abstract
With unstructured heterogeneous multimedia data such as texts, images being more and more widely used on the web, cross-media retrieval has become an increasingly important task. One of the key techniques in cross-media retrieval is how to compute distances or similarities among different types of media data. In this paper, we propose a novel heterogeneous metric learning method to compute distances between images and texts. We extend Locally Linear Embedding (LLE) to deal with heterogeneous data, so that we can not only preserve homogeneous local information but also capture heterogeneous constraints. In order to handle the out-of-sample problem, we learn two map functions from the embedding, and use them to transform heterogeneous data into a homogeneous space and do the retrieval in the new space. The experimental results on two real-world datasets show the effectiveness of our approach.
Peng Zhou 0006, Liang Du 0003, Mingyu Fan, Yidong Shen
SDM3
2015 Efficient sequential feature selection based on adaptive eigenspace model
Nannan Gu, Mingyu Fan, Liang Du 0003, Dongchun Ren
Neurocomputing2
2015 An Efficient Semi-Supervised Classifier Based on Block-Polynomial Mapping
abstract
In this paper, we propose a block-polynomial mapping for image feature learning, which can be efficiently represented by the matrix Khatri-Rao product. The block-polynomial mapping not only captures the local discriminative information within the image structure, but is also much more efficient than the traditional kernel mapping. Moreover, we embed the proposed mapping into the manifold regularization framework for semi-supervised image classification. Experimental results demonstrate that, while maintaining a comparable classification accuracy, the proposed algorithm performs much more efficient than the state-of-the-art methods.
Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye
IEEE Signal Process. Lett.3
2014 A kernel-based sparsity preserving method for semi-supervised classification
Nannan Gu, Di Wang 0008, Mingyu Fan, Deyu Meng
Neurocomputing3
2014 Dimensionality reduction: An interpretation from manifold regularization perspective
Mingyu Fan, Nannan Gu, Hong Qiao, Bo Zhang 0006
Inf. Sci.1
2014 A Regularized Approach for Geodesic-Based Semisupervised Multimanifold Learning
abstract
Geodesic distance, as an essential measurement for data dissimilarity, has been successfully used in manifold learning. However, most geodesic distance-based manifold learning algorithms have two limitations when applied to classification: 1) class information is rarely used in computing the geodesic distances between data points on manifolds and 2) little attention has been paid to building an explicit dimension reduction mapping for extracting the discriminative information hidden in the geodesic distances. In this paper, we regard geodesic distance as a kind of kernel, which maps data from linearly inseparable space to linear separable distance space. In doing this, a new semisupervised manifold learning algorithm, namely regularized geodesic feature learning algorithm, is proposed. The method consists of three techniques: a semisupervised graph construction method, replacement of original data points with feature vectors which are built by geodesic distances, and a new semisupervised dimension reduction method for feature vectors. Experiments on the MNIST, USPS handwritten digit data sets, MIT CBCL face versus nonface data set, and an intelligent traffic data set show the effectiveness of the proposed algorithm.
Mingyu Fan, Xiaoqin Zhang 0002, Zhouchen Lin, Zhongfei Zhang, Hujun Bao
IEEE Trans. Image Process.1
2013 Dimension estimation of image manifolds by minimal cover approximation
Mingyu Fan, Xiaoqin Zhang 0002, Shengyong Chen, Hujun Bao, Stephen J. Maybank
Neurocomputing1
2013 Regularized Discriminative Spectral Regression Method for Heterogeneous Face Matching
abstract
Face recognition is confronted with situations in which face images are captured in various modalities, such as the visual modality, the near infrared modality, and the sketch modality. This is known as heterogeneous face recognition. To solve this problem, we propose a new method called discriminative spectral regression (DSR). The DSR maps heterogeneous face images into a common discriminative subspace in which robust classification can be achieved. In the proposed method, the subspace learning problem is transformed into a least squares problem. Different mappings should map heterogeneous images from the same class close to each other, while images from different classes should be separated as far as possible. To realize this, we introduce two novel regularization terms, which reflect the category relationships among data, into the least squares approach. Experiments conducted on two heterogeneous face databases validate the superiority of the proposed method over the previous methods.
Xiangsheng Huang, Zhen Lei 0001, Mingyu Fan, Xiao Wang 0012, Stan Z. Li
IEEE Trans. Image Process.3
2012 Isometric Multi-manifold Learning for Feature Extraction
abstract
Manifold learning is an important topic in pattern recognition and computer vision. However, most manifold learning algorithms implicitly assume the data are aligned on a single manifold, which is too strict in actual applications. Isometric feature mapping (Isomap), as a promising manifold learning method, fails to work on data which distribute on clusters in a single manifold or manifolds. In this paper, we propose a new multi-manifold learning algorithm (M-Isomap). The algorithm first discovers the data manifolds and then reduces the dimensionality of the manifolds separately. Meanwhile, a skeleton representing the global structure of whole data set is built and kept in low-dimensional space. Secondly, by referring to the low-dimensional representation of the skeleton, the embeddings of the manifolds are relocated to a global coordinate system. Compared with previous methods, these algorithms can keep both of the intra and inter manifolds geodesics faithfully. The features and effectiveness of the proposed multi-manifold learning algorithms are demonstrated and compared through experiments.
Mingyu Fan, Hong Qiao, Bo Zhang 0006, Xiaoqin Zhang 0002
ICDM1
2012 Geodesic Based Semi-supervised Multi-manifold Feature Extraction
abstract
Manifold learning is an important feature extraction approach in data mining. This paper presents a new semi-supervised manifold learning algorithm, called Multi-Manifold Discriminative Analysis (Multi-MDA). The proposed method is designed to explore the discriminative information hidden in geodesic distances. The main contributions of the proposed method are: 1) we propose a semi-supervised graph construction method which can effectively capture the multiple manifolds structure of the data, 2) each data point is replaced with an associated feature vector whose elements are the graph distances from it to the other data points. Information of the nonlinear structure is contained in the feature vectors which are helpful for classification, 3) we propose a new semi-supervised linear dimension reduction method for feature vectors which introduces the class information into the manifold learning process and establishes an explicit dimension reduction mapping. Experiments on benchmark data sets are conducted to show the effectiveness of the proposed method.
Mingyu Fan, Xiaoqin Zhang 0002, Zhouchen Lin, Zhongfei Zhang, Hujun Bao
ICDM1
2012 Discriminative Sparsity Preserving Projections for Semi-Supervised Dimensionality Reduction
abstract
In this letter, we propose a semi-supervised dimensionality reduction method named Discriminative Sparsity Preserving Projection (DSPP). In order to get the feature mapping$f$which projects the high-dimensional data into a low-dimensional intrinsic space, DSPP attempts to maintain the prior low-dimensional representation constructed by the data points and the known class labels and, meanwhile, considers the complexity of$f$in the ambient space and the smoothness of$f$in preserving the sparse representation of data. On one hand, the DSPP method obtains an explicit nonlinear feature mapping for the out-of-sample extrapolation. On the other hand, the DSPP method has a high discriminative ability which is inherited from the sparse representation of data. Experiment results show the effectiveness of the proposed method.
Nannan Gu, Mingyu Fan, Hong Qiao, Bo Zhang 0006
IEEE Signal Process. Lett.2
2011 Sparse regularization for semi-supervised classification
Mingyu Fan, Nannan Gu, Hong Qiao, Bo Zhang 0006
Pattern Recognit.1
2010 Load Aware Routing with Delay Threshold for Vehicular Ad Hoc Networks
abstract
It is being envisaged that a wide variety of applications can be running on vehicular ad hoc networks (VANETs) in the near future. The coexistence of these applications suggests that they will be competing for the use of the wireless medium. That easily leads to severe congestion and even “hotspot” problem in high traffic density urban areas, and results in the failure of time-critical applications. To address this issue, we propose two carry-and-forward schemes called LARD-Greedy and LARD-Optimal respectively that attempt to deliver packets along the path with optimal performance. The proposed algorithms leverage local or global knowledge of traffic statistics to choose the path with optimal performance, and rationally alternate between the Carrying and Multihop Forwarding strategies according to the current load status in order to avoid congestion. Experimental results based on a real city map show the proposed solutions reach very good performance in terms of packet delivery ratio and delay.
Mingyu Fan, Lishu Wang
ICPADS2
2009 A trust model based on fuzzy recommendation for mobile ad-hoc networks
Junhai Luo, Xue (Steve) Liu, Mingyu Fan
Comput. Networks3
2009 Intrinsic dimension estimation of manifolds by incising balls
Mingyu Fan, Hong Qiao, Bo Zhang 0006
Pattern Recognit.1