Yang Gao 0027

dblp:89/4402-27 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
9since 2021 · last 2024
0000-0001-9328-1611ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Security and privacy · 4 · 3 since 2021
YearPublicationVenuePosition
2024 Heterogeneous Domain Adaptation for Multistream Classification on Cyber Threat Data
abstract
Under a newly introduced setting of multistream classification, two data streams are involved, which are referred to as source and target streams. The source stream continuously generates data instances from a certain domain with labels, while the target stream does the same task without labels from another domain. Existing approaches assume that domains for both data streams are identical, which is not quite true, since data streams from different sources may contain distinct features. Indeed, they may even have different numbers of features. Furthermore, obtaining labels for every instance in a data stream is often expensive and time-consuming. Therefore, it has become an important topic to explore if classes of labeled instances from other related streams are helpful to predict the classes of unlabeled instances in a different stream. Note that domains of source and target streams may have distinct feature spaces and data distributions. Our objective is to predict class labels of data instances in the target stream by using the classifiers trained by the source stream. We propose a framework of multistream classification by using projected data from a common latent feature space, which is embedded from both source and target domains. This framework is also crucial for enterprise system defenders to detect cross-platform attacks, such as Advanced Persistent Threats (APTs). Empirical valuation and analysis on both real-world and synthetic datasets are performed to validate the effectiveness of our proposed algorithm, comparing to state-of-the-art techniques. Experimental results show that our approach significantly outperforms other existing approaches.
Yifan Li 0003, Yang Gao 0027, Gbadebo Ayoade, Latifur Khan, Anoop Singhal, Bhavani Thuraisingham
IEEE Trans. Dependable Secur. Comput.2
2023 Advanced Persistent Threat Detection Using Data Provenance and Metric Learning
abstract
Advanced persistent threats (APT) have increased in recent times as a result of the rise in interest by nation-states and sophisticated corporations to obtain high-profile information. Typically, APT attacks are more challenging to detect since they leverage zero-day attacks and common benign tools. Furthermore, these attack campaigns are often prolonged to evade detection. We leverage an approach that uses a provenance graph to obtain execution traces of host nodes in order to detect anomalous behavior. By using the provenance graph, we extract features that are then used to train an online adaptive metric learning. Online metric learning is a deep learning method that learns a function to minimize the separation between similar classes and maximizes the separation between dis- similar instances. We compare our approach with baseline models and we show our method outperforms the baseline models by increasing detection accuracy on average by 11.3% and increases True positive rate (TPR) on average by 18.3%. We also show that our method outperforms several state-of-the-art models performances in comprehensive attack datasets in both binary and multi-class settings.
Khandakar Ashrafi Akbar, Yigong Wang, Gbadebo Ayoade, Yang Gao 0027, Anoop Singhal, Latifur Khan, Bhavani Thuraisingham, Kangkook Jee
IEEE Trans. Dependable Secur. Comput.4
2022 Unified Spatio-Temporal Graph Neural Networks: Data-Driven Modeling for Social Science
abstract
Time series forecasting with additional spatial de-pendencies has attracted a tremendous amount of research interest in social sciences, due to its importance in modern real-world applications. The Graph Neural Networks (GNN) is one of the most exciting deep learning techniques among these spatio-temporal modeling approaches. Most existing spatio-temporal GNN frameworks are based on a two-step modeling process. In such scenario, spatial and temporal dependencies are modeled in separate steps, which lead to problems such as complex architecture design, hard to scale, etc. Targeting the shortcomings of existing studies, we take both spatial and temporal dependencies from another perspective, and consider them as two heterogeneous types of edges in the graph. We propose a unified spatio-temporal GNN framework that captures both dependencies in a single step. More specifically, for each node in the graph, a unified neural network component is designed to simultaneously extract information from its sur-rounding neighbors (spatial) and its past records (temporal), which enables much easier dependency aggregation with faster execution. Experiment results demonstrate the superiority of the proposed framework over state-of-the-art (SOTA) baselines on various applications, including modeling smart cities and data-driven political science research.
Yifan Li 0003, Yu Lin 0002, Yang Gao 0027, Latifur Khan
IJCNN3
2022 SACCOS: A Semi-Supervised Framework for Emerging Class Detection and Concept Drift Adaption Over Data Streams
abstract
In this paper, we address challenges of detecting instances from emerging classes over a non-stationary data stream during data classification. In particular, data instances from an entirely unknown class may appear in a data stream over time. Existing classification techniques utilize unsupervised clustering to identify emergence of such data instances. Unfortunately, they make strong assumptions which are typically invalid in practice; (i) Most instances associated with a class are closer to each other in feature space than instances associated with different classes, (ii) Covariates of data are normalized through an oracle to overcome the effect of a few data instances having large feature values, and (iii) Labels of instances from emerging classes are readily available soon after detection. To address the challenges that occur in practice when the above assumptions are weak, i.e., instances of each class are scattered and the true labels of novel class instances are sparsely available, we propose a practical semi-supervised emerging class detection framework. Particularly, we aim to identify similar data instances within local regions in feature space by incorporating a mutual graph clustering mechanism. We also perform online normalization along the data stream instead of assuming an oracle, and propose a classification technique that uses only a small amount of true labels for training and emerging class detection. Our empirical evaluation of this framework on real-world datasets demonstrates its superiority of classification performance compared to existing methods, while using significantly fewer labeled instances.
Yang Gao 0027, Swarup Chandra, Yifan Li 0003, Latifur Khan, Bhavani Thuraisingham
IEEE Trans. Knowl. Data Eng.1
2021 Single View Point Cloud Generation via Unified 3D Prototype
abstract
As 3D point clouds become the representation of choice for multiple vision and graphics applications, such as autonomous driving, robotics, etc., the generation of them by deep neural networks has attracted increasing attention in the research community. Despite the recent success of deep learning models in classification and segmentation, synthesizing point clouds remains challenging, especially from a single image. State-of-the-art (SOTA) approaches can generate a point cloud from a hidden vector, however, they treat 2D and 3D features equally and disregard the rich shape information within the 3D data. In this paper, we address this problem by integrating image features with 3D prototype features. Specifically, we propose to learn a set of 3D prototype features from a real point cloud dataset and dynamically adjust them through the training. These prototypes are then integrated with incoming image features to guide the point cloud generation process. Experimental results show that our proposed method outperforms SOTA methods on single image based 3D reconstruction tasks.
Yu Lin 0002, Yigong Wang, Yifan Li 0003, Zhuoyi Wang, Yang Gao 0027, Latifur Khan
AAAI5
2021 An Episodic Learning based Geolocation Detection Framework for Imbalanced Data
abstract
A social media user's geographical location is vital to many applications like local search and event detection. The scarcity of publicly available location information motivates researchers to predict user geolocation based on information such as tweet text and social interaction data. In this paper, we investigate and improve on the task of predicting a Twitter user's city-level location based on the content of the user's historical tweets. In order to train a reliable location classifier, previous studies on this topic have typically assumed that there are sufficient amount of users living in each cities. However, they simply ignore the fact that different demographic groups may participate in social media platforms, which results in a highly imbalanced data distribution. Being aware of this population imbalance issue, we propose an episodic learning based framework to extract a single representative for each class (location), so that classifiers can later be trained on a balanced class distribution. To examine the effectiveness of our method, we design experiments which involve two kinds of baselines, the state-of-the-art geolocation detection methods and the well-known approaches handling imbalanced data in classification. The results of experiments on the data collected from Twitter demonstrated the superiority of our method when compared with baselines.
Hemeng Tao, Yang Gao 0027, Zhuoyi Wang, Latifur Khan, Bhavani Thuraisingham
IJCNN2
2021 Generating Point Cloud from Single Image in The Few Shot Scenario
abstract
Reconstructing point clouds from images would extremely benefit many practical CV applications, such as robotics, automated vehicles, and Augmented Reality. Fueled by the advances of deep neural network, many deep learning frameworks are proposed to address this problem recently. However, these frameworks generally rely on a large amount of labeled training data (e.g., image and point cloud pairs). Although we usually have numerous 2D images, corresponding 3D shapes are insufficient in practice. In addition, most available 3D data covers only a limited amount of classes, which further restricts the models' generalization ability to novel classes. To mitigate these issues, we propose a novel few-shot single-view point cloud generation framework by considering both class-specific and class-agnostic 3D shape priors. Specifically, we abstract each class by a prototype vector that embeds class-specific shape priors. Class-agnostic shape priors are modeled by a set of learnable shape primitives that encode universal 3D shape information shared across classes. Later, we combine the input image with class-specific prototypes and class-agnostic shape primitives to guide the point cloud generation process. Experiments on the popular ModelNet and ShapeNet datasets demonstrate that our method outperforms state-of-the-art methods in the few-shot setting.
Yu Lin 0002, Jinghui Guo, Yang Gao 0027, Yifan Li 0003, Zhuoyi Wang, Latifur Khan
ACM Multimedia3
2021 Attention-Based Spatial Guidance for Image-to-Image Translation
abstract
The aim of image-to-image translation algorithms is to tackle the challenges of learning a proper mapping function across different domains. Generative Adversarial Networks (GANs) have shown superior ability to handle this problem in both supervised and unsupervised ways. However, one critical problem of GAN in practice is that the discriminator is typically much stronger than the generator, which could lead to failures such as mode collapse, diminished gradient, etc. To address these shortcomings, we propose a novel framework, which incorporates a powerful spatial attention mechanism to guide the generator. Specifically, our designed discriminator estimates the probability of realness of a given image, and provides an attention map regarding this prediction. The generated attention map contains the informative regions to distinguish the real and fake images, from the perspective of the discriminator. Such information is particularly valuable for the translation because the generator is encouraged to focus on those areas and produce more realistic images. We conduct extensive experiments and evaluations, and show that our proposed method is both qualitatively and quantitatively better than other state-of-the-art image translation frameworks.
Yu Lin 0002, Yigong Wang, Yifan Li 0003, Yang Gao 0027, Zhuoyi Wang, Latifur Khan
WACV4
2021 Crook-sourced intrusion detection as a service
Frederico Araujo, Gbadebo Ayoade, Khaled Al-Naami, Yang Gao 0027, Kevin W. Hamlen, Latifur Khan
J. Inf. Secur. Appl.4
2020 A Simple, Effective and Extendible Approach to Deep Multi-task Learning
abstract
Existing solutions to multi-task learning typically rely on manually enumerating multiple network architectures to find the optimal structure, which incurs a heavy design workload. In addition, extending these models to new tasks is difficult in many cases, since it often requires a network redesign for achieving the best performance. To overcome these limitations, in this paper, we propose a novel principle for multitask learning, which focuses on the learning process itself. For each task, we consider its feature as a state, and treat the learning problem as a state transformation process, which is driven by a task-specific gradient. Specifically, we introduce a Gradient Modification Unit (GMU), which consists of a Gradient Estimation Network (GEN) (shared among all tasks) and multiple tiny task-specific gradient correctors (one for each task). At each iteration, for any task T, the corresponding gradient corrector modifies the gradient estimated by the GEN to produce the desired task-specific gradient tensor that is applied to update the state of task T. Our solution has several benefits including automatic end-to-end learning of the optimal transformation process, simplification of network design and high expansibility to new tasks. We demonstrate the superiority of our approach over existing solutions on a variety of datasets, across both image classification and image retrieval tasks.
Yang Gao 0027, Yifan Li 0003, Yu Lin 0002, Hemeng Tao, Latifur Khan
IEEE BigData1
2020 SetConv: A New Approach for Learning from Imbalanced Data
abstract
For many real-world classification problems, e.g., sentiment classification, most existing machine learning methods are biased towards the majority class when the Imbalance Ratio (IR) is high.To address this problem, we propose a set convolution (SetConv) operation and an episodic training strategy to extract a single representative for each class, so that classifiers can later be trained on a balanced class distribution.We prove that our proposed algorithm is permutation-invariant despite the order of inputs, and experiments on multiple large-scale benchmark text datasets show the superiority of our proposed framework when compared to other SOTA methods.
Yang Gao 0027, Yifan Li 0003, Yu Lin 0002, Charu C. Aggarwal, Latifur Khan
EMNLP (1)1
2020 Adaptive Multi-Region Network For Medical Image Analysis
abstract
Automated diagnosis of significant abnormalities (or lesions) from radiology images has been well exploited in Deep Learning (DL) because of the ability to model sophisticated features. However, a deep neural network should be trained on a huge amount of data to infer the parameter values. Unfortunately, for the problems in lesion diagnosis, there is only a limited amount of data annotated in a manner that is suitable to learn powerful deep models. Moreover, the lesion in the radiology image is often vague and hard to identify without expert knowledge. In this paper, we focus on previous challenges in the automated diagnosis and propose the approach named Adaptive Multi-region Network (AdapNet). The key idea is that we adaptively encode the similarity of lesions in different context regions through margin-max learning strategy, which incorporates the metrics learned on those regions to enhance the effectiveness of the model. Our experiments show that the proposed method can effectively obtain superior performance compared to the existing methods, on the DeepLesion data sets.
Hemeng Tao, Zhuoyi Wang, Yang Gao 0027, Yigong Wang, Latifur Khan
ICIP3
2020 Prediction of Plantar Shear Stress Distribution by Conditional GAN with Attention Mechanism
Jinghui Guo, Ali Ersen, Yang Gao 0027, Yu Lin 0002, Latifur Khan, Metin Yavuz
MICCAI (2)3
2019 Multistream Classification with Relative Density Ratio Estimation
abstract
In supervised learning, availability of sufficient labeled data is of prime importance. Unfortunately, they are sparingly available in many real-world applications. Particularly when performing classification over a non-stationary data stream, unavailability of sufficient labeled data undermines the classifier’s long-term performance by limiting its adaptability to changes in data distribution over time. Recently, studies in such settings have appealed to transfer learning techniques over a data stream while detecting drifts in data distribution over time. Here, the data stream is represented by two independent non-stationary streams, one containing labeled data instances (called source stream) having a biased distribution compared to the unlabeled data instances (called target stream). The task of label prediction under this representation is called Multistream Classification, where instances in the two streams occur independently. While these studies have addressed various challenges in the multistream setting, it still suffers from large computational overhead mainly due to frequent bias correction and drift adaptation methods employed. In this paper, we focus on utilizing an alternative bias correction technique, called relative density-ratio estimation, which is known to be computationally faster. Importantly, we propose a novel mechanism to automatically learn an appropriate mixture of relative density that adapts to changes in the multistream setting over time. We theoretically study its properties and empirically demonstrate its superior performance, within a multistream framework called MSCRDR, on benchmark datasets by comparing with other competing methods.
Yang Gao 0027, Swarup Chandra, Latifur Khan
AAAI2
2019 Improving intrusion detectors by crook-sourcing
abstract
Conventional cyber defenses typically respond to detected attacks by rejecting them as quickly and decisively as possible; but aborted attacks are missed learning opportunities for intrusion detection. A method of reimagining cyber attacks as free sources of live training data for machine learning-based intrusion detection systems (IDSes) is proposed and evaluated. Rather than aborting attacks against legitimate services, adversarial interactions are selectively prolonged to maximize the defender's harvest of useful threat intelligence. Enhancing web services with deceptive attack-responses in this way is shown to be a powerful and practical strategy for improved detection, addressing several perennial challenges for machine learning-based IDS in the literature, including scarcity of training data, the high labeling burden for (semi-)supervised learning, encryption opacity, and concept differences between honeypot attacks and those against genuine services. By reconceptualizing software security patches as feature extraction engines, the approach conscripts attackers as free penetration testers, and coordinates multiple levels of the software stack to achieve fast, automatic, and accurate labeling of live web data streams.
Frederico Araujo, Gbadebo Ayoade, Khaled Al-Naami, Yang Gao 0027, Kevin W. Hamlen, Latifur Khan
ACSAC4
2019 Regression Prediction For Geolocation Aware Through Relative Density Ratio Estimation
abstract
Typically, a traditional regression model over a data set is trained on instances by assuming a stationary distribution (the training data and test data has the same distribution). The model learned from the training set is later used to predict value of response-variable in future test instances. In this paper, we study the regression problem in a novelty problem setting, referred as a transfer learning involves training set and test set with different data distribution. Here, the training set has a biased data distribution with respect to the test set. The regression problem with the transfer learning approach aims to predict the dependent variable of test data, through utilizing the training data with the biased distribution. Availability of sufficient training data is of prime importance. Unfortunately, such problem settings are sparingly available in many real-word applications, such as you only have the flight delay information of DFW (DallasWorth) airport and prefer to predict the flight delay information of JFK (John F. Kennedy) airport. In our solution, we utilize relative density ratio to evaluate the difference between the data from different locations. We revise the relative density ratio estimation based on different location types, and propose an effective approach for predicting arrival delay time for commercial flights by geolocation aware relative kernel density estimation. Extensive experimental results on commercial flight datasets with different locations show that our approach effectively boosts the performance of other state-of-the art.
Jinghui Guo, Zhuoyi Wang, Yang Gao 0027, Latifur Khan
IEEE BigData5
2019 SIM: Open-World Multi-Task Stream Classifier with Integral Similarity Metrics
abstract
One of the key challenges of performing label predictions over a data stream is concerned with the emergence of instances belonging to unobserved (or novel) classes over time. Although existing studies have proposed various solutions to address this challenge, they mostly focus on streams with lowdimensional data and strongly rely on the intrinsic cohesion and separation data property, i.e., instances belonging to the same class are closer to each other (cohesion) than those belonging to different classes (separation) in the observed feature space, to detect instances from unknown classes. Unfortunately, such a property is typically not inherent in high-dimensional data such as images and texts. Thus, to perform classification and novel class detection on high-dimensional data streams, we need to address two main problems: 1) Finding a feature space that exhibit cohesion and separation properties, and 2) Training with limited amount of labeled data. In this paper, we propose a multi-task metric learning mechanism useful for identifying a latent space in which the cohesion and separation data property is valid and have designed a semi-supervised stream classifier called SIM based on this mechanism. We empirically measure the performance of SIM over multiple real-world image and text datasets, and demonstrate its superiority by comparing the performance with existing state-of-the-art frameworks.
Yang Gao 0027, Yifan Li 0003, Yu Lin 0002, Latifur Khan
IEEE BigData1
2019 Towards Self-Adaptive Metric Learning On the Fly
abstract
Good quality similarity metrics can significantly facilitate the performance of many large-scale, real-world applications. Existing studies have proposed various solutions to learn a Mahalanobis or bilinear metric in an online fashion by either restricting distances between similar (dissimilar) pairs to be smaller (larger) than a given lower (upper) bound or requiring similar instances to be separated from dissimilar instances with a given margin. However, these linear metrics learned by leveraging fixed bounds or margins may not perform well in real-world applications, especially when data distributions are complex. We aim to address the open challenge of “Online Adaptive Metric Learning” (OAML) for learning adaptive metric functions on-the-fly. Unlike traditional online metric learning methods, OAML is significantly more challenging since the learned metric could be non-linear and the model has to be self-adaptive as more instances are observed. In this paper, we present a new online metric learning framework that attempts to tackle the challenge by learning a ANN-based metric with adaptive model complexity from a stream of constraints. In particular, we propose a novel Adaptive-Bound Triplet Loss (ABTL) to effectively utilize the input constraints, and present a novel Adaptive Hedge Update (AHU) method for online updating the model parameters. We empirically validates the effectiveness and efficacy of our framework on various applications such as real-world image classification, facial verification, and image retrieval.
Yang Gao 0027, Yifan Li 0003, Swarup Chandra, Latifur Khan, Bhavani Thuraisingham
WWW1
2019 Multistream Classification for Cyber Threat Data with Heterogeneous Feature Space
abstract
Under a newly introduced setting of multistream classification, two data streams are involved, which are referred to as source and target streams. The source stream continuously generates data instances from a certain domain with labels, while the target stream does the same task without labels from another domain. Existing approaches assume that domains for both data streams are identical, which is not quite true in real world scenario, since data streams from different sources may contain distinct features. Furthermore, obtaining labels for every instance in a data stream is often expensive and time-consuming. Therefore, it has become an important topic to explore whether labeled instances from other related streams can be helpful to predict those unlabeled instances in a given stream. Note that domains of source and target streams may have distinct features spaces and data distributions. Our objective is to predict class labels of data instances in the target stream by using the classifiers trained by the source stream.
Yifan Li 0003, Yang Gao 0027, Gbadebo Ayoade, Hemeng Tao, Latifur Khan, Bhavani Thuraisingham
WWW2
2017 Multistream regression with asynchronous concept drift detection
abstract
A recently introduced problem setting, referred as multistream, involves two independent non-stationary data generating processes. One of them is called source stream, which generates continuous data instances with true output. And the other one called target stream, which generates data instances lacking of true output. Due to the nature of data streams, scholars have addressed prediction problems under scenarios such as covariate shift or concept drift in past studies by discussing one assumption while keeping others consistent. For example, it is assumed that the data distributions of training and testing data are similar, and true output values of the stream instances would be available soon. However, in practice these assumptions are not always valid. The multistream regression problem is to predict the output of target stream, using data instances and their true output from source stream. In this paper, we propose an approach of multistream regression by incorporating concept drift detection into covariate shift adaptation. Meanwhile, empirical evaluation on synthetic and real world datasets demonstrates the effectiveness of the proposed technique by competing with the state-of-the-art approaches. Experiment results indicate that our method significantly improved prediction performance compared to existing benchmark.
Yifan Li 0003, Yang Gao 0027, Ahsanul Haque, Latifur Khan, Mohammad M. Masud 0001
IEEE BigData3