Lianhua Chi

dblp:58/10365 · also Lian-Hua Chi, Lian-hua Chi · DBLP profile ↗
← Back
37ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-6851-0731ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 DWCL: Dual-Weighted Contrastive Learning for robust multi-view clustering
Hanning Yuan, Lianhua Chi, Sijie Ruan, Wei Zhou 0021, Jinhui Pang, Xiaoshuai Hao
Eng. Appl. Artif. Intell.4
2026 Tensor-Based Privacy-Aware Driving Route Navigation Based on Cloud-Fog-Edge Calculative User-Vehicle-Road Preferences
abstract
There are currently four key limitations in most existing privacy-aware driving route navigation methods: 1) lack of an offloading framework for massive driving route data, 2) potential privacy leakage even with virtual trajectories, 3) reliance on recommending the fastest route rather than the most suitable one, 4) lack of comprehensive consideration of multi-dimensional information, leading to inadequate navigation accuracy. The limitations restrict the further application of the navigation. To address these limitations, we propose a Tensor-based Privacy-aware driving route Navigation scheme based on Cloud-fog-edge calculative User-vehicle-road Preferences (named TPNCUP). First, TPNCUP includes an offloading framework based on cloud-fog-edge collaborative computing. Second, TPNCUP provides a route obfuscation method to enhance privacy level. Third, TPNCUP recommends routes based on user-vehicle-road preferences instead of simply selecting the fastest one. Finally, TPNCUP constructs a set of 3rd-order tensors to comprehensively consider multi-dimensional information for increasing navigation accuracy. From a philosophical perspective, the use of TPNCUP reconstructs the subject-object relationship, corrects technical determinism, and assumes ethical responsibility in intelligent navigation. Extensive experimental results show that our scheme significantly outperforms two state-of-the-art methods in the privacy level, navigation accuracy and computational overhead. Specifically, our scheme reduces the risk of privacy leakage by 70.53%, improves navigation accuracy by 8.01%, and decreases computational overhead by 77.3%.
Jing Yu 0012, Wei Wei 0006, Zongmin Cui, Lianhua Chi
IEEE Trans. Intell. Transp. Syst.4
2025 Tensor-based ranking-hiding privacy-preserving scheme for cloud-fog-edge cooperative cyber-physical-social systems
Jing Yu 0012, Lianhua Chi, Shunli Zhang 0003, Zongmin Cui
J. Netw. Comput. Appl.3
2025 ConEm: A novel framework for integrating external factors with inner and outer correlations in time series forecasting
abstract
Time series forecasting is pivotal in both academic research and practical applications across diverse industries. However, effectively leveraging external factors to enhance forecasting performance remains a significant challenge, necessitating further investigation. Current frameworks exhibit notable limitations in modeling the impact of external factors on both intrinsic and extrinsic correlations within time series data. To address these challenges, we propose a novel mechanism that systematically integrates contextual information from external factors with temporal dependencies, while maintaining compatibility with various encoder-decoder algorithms. This approach enables backbone models to embed dependent patterns from external factors across multiple correlated time series, effectively capturing their influence on both prior and adjacent timesteps. Our study centered on the application of time series forecasting for demand prediction, as sales forecasting poses unique challenges stemming from the complexity and variability of market conditions influenced by numerous external factors. We conducted extensive experiments on three real-world retail datasets, showcasing the substantial performance enhancement of backbone models when integrated with our proposed contextual embedding mechanism. Specifically, our approach achieves improvements of up to 26 % in Mean Squared Error (MSE) and 15 % in Mean Absolute Error (MAE) compared to both the original backbone models and other state-of-the-art (SOTA) baseline methods. The proposed mechanism is also evaluated on Weather and Energy datasets to further verify its generalization capability. We will release the source codes and experimental datasets at our GitHub1.
Hoang Nguyen Nguyen, Wei Xiang 0001, Lianhua Chi, Mike Da Gama, Sanjeevani Avashi, Michael Treloar
Knowl. Based Syst.3
2024 Deep Contrastive Multi-view Clustering Under Semantic Feature Guidance
Hanning Yuan, Ziqiang Yuan, Lianhua Chi, Jing Geng 0002, Shuliang Wang 0001
ADMA (1)4
2024 Validating Siamese embedded neural networks with identical representations for efficient model convergence
abstract
Deep learning, neural architecture search, reinforcement learning, and embedded learning use scaled data, hyperspace, reward, and similarity to produce efficient converging neural networks. However, all these state-of-the-art frameworks require training to evaluate whether the examples are identical. Therefore, we propose our Probabilistic Asymmetric Convergence Network with Validation and Transform-Learning (PACNVT) framework that learns with fewer data, reduces the hyperspace with our validation technique and modified tangent activation function, reinforces the learning with our transform learning algorithms, and ascertains similarity independent of spatial and mask for transformational consistency. Furthermore, our framework generates neural networks that yield state-of-the-art accuracies on MNIST, OMNIGLOT, and CIFAR datasets. Moreover, representing regression as a similarity problem unveils previously unseen residual fluid intelligence patterns, yielding 99.92% accuracy with as little as 20 pairs.
Mathias Hoy Talbo, Haishuai Wang, Lianhua Chi, Yi-Ping Phoebe Chen
Knowl. Based Syst.3
2024 Coresets for fast causal discovery with the additive noise model
Boxiang Zhao, Shuliang Wang 0001, Lianhua Chi, Hanning Yuan, Ye Yuan 0001, Qi Li 0022, Jing Geng 0002
Pattern Recognit.3
2024 Design and Robust Evaluation of Next Generation Node Authentication Approach
abstract
The flexibility of 5G-NGNs makes them an ideal infrastructure for supporting mission-critical IoT applications that require low latency and high bandwidth. However, due to the rapid proliferation and the integration of IoTs with 5 G, the threat surface has considerably expanded. Hence the security of IoT devices is a big concern. Unfortunately, IoT devices have limited resources, and the traditional security approaches (authentication and intrusion detection approaches) of cryptography do not work effectively on 5G-IoT ecosystems. Motivated from this, we leverage the distinctive RF (Radio Frequency) fingerprinting signatures of IoT devices and used them to train a Deep learning model, Mahalanobis Distance theory in addition to the Chi-square distribution theory, to authenticate the IoT nodes. Under robust scenarios we have tested the approach shows detection accuracy (99.35%) as well as significant amount of reduction in model's training time as these two metrics are one of the primary key performance indicators (KPIs). In order to evaluate the effectiveness of the proposed method in real-time scenarios, we tested the proposed solution with a real RF dataset and the OSM-MANO 5 G platform. The model underwent formal verification using the Tamarin Prover tool, and the proposal was also compared with recent research works.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.5
2024 Heterogeneous Network Motif Coding, Counting, and Profiling
abstract
Network motifs, as a fundamental higher-order structure in large-scale networks, have received significant attention over recent years. Particularly in heterogeneous networks, motifs offer a higher capacity to uncover diverse information compared to homogeneous networks. However, the structural complexity and heterogeneity pose challenges in coding, counting, and profiling heterogeneous motifs. This work addresses these challenges by first introducing a novel heterogeneous motif coding method, adaptable to homogeneous motifs as well. Building upon this coding framework, we then propose GIFT, a heterogeneous network motif counting algorithm. GIFT effectively leverages combined structures of heterogeneous motifs through three key procedures: neighborhood searching, motif combination, and redundant motif filtering. We apply GIFT to count three-order and four-order motifs across eight distinct heterogeneous networks. Subsequently, we profile these detected motifs using four classical motif-based indicators. Experimental results demonstrate that by appropriately selecting motifs tailored to specific networks, heterogeneous motifs emerge as significant features in characterizing the underlying network structure.
Shuo Yu 0001, Feng Xia 0001, Honglong Chen, Ivan Lee 0001, Lianhua Chi, Hanghang Tong
ACM Trans. Knowl. Discov. Data5
2024 Correlation-Aware Spatial-Temporal Graph Learning for Multivariate Time-Series Anomaly Detection
abstract
Multivariate time-series anomaly detection is critically important in many applications, including retail, transportation, power grid, and water treatment plants. Existing approaches for this problem mostly employ either statistical models which cannot capture the nonlinear relations well or conventional deep learning (DL) models e.g., convolutional neural network (CNN) and long short-term memory (LSTM) that do not explicitly learn the pairwise correlations among variables. To overcome these limitations, we propose a novel method, correlation-aware spatial-temporal graph learning (termed ), for time-series anomaly detection. explicitly captures the pairwise correlations via a correlation learning (MTCL) module based on which a spatial-temporal graph neural network (STGNN) can be developed. Then, by employing a graph convolution network (GCN) that exploits one-and multihop neighbor information, our STGNN component can encode rich spatial information from complex pairwise dependencies between variables. With a temporal module that consists of dilated convolutional functions, the STGNN can further capture long-range dependence over time. A novel anomaly scoring component is further integrated into to estimate the degree of an anomaly in a purely unsupervised manner. Experimental results demonstrate that can detect and diagnose anomalies effectively in general settings as well as enable early detection across different time delays. Our code is available at https://github.com/huankoh/CST-GL.
Yu Zheng 0013, Huan Yee Koh, Ming Jin 0005, Lianhua Chi, Khoa Tran Phan, Shirui Pan, Yi-Ping Phoebe Chen, Wei Xiang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Bio-Inspired Dual-Network Model to Tackle Statistical Heterogeneity in Federated Learning
abstract
The problem of statistical heterogeneity in Federated Learning has been a major challenge, with existing solutions making unrealistic assumptions about the availability of shared datasets and high bandwidth between clients and the server. Solving this problem is crucial for the success of Federated Learning in real-world scenarios. In this work, we propose a biologically inspired dual-network model FedDual, which mimics how the human brain learns and memorizes the information. The model consists of a neocortical and a hippocampal network similar to those in the human brain. The hippocampal network is comprised by an image classification model, while the neocortical network is a variational auto-encoder responsible for long-term and re-callable memory. In this manner, FedDual uses the neocortical network to generate pseudo-patterns (synthetic data) on the server (global model). This allows for the hippocampal network to be trained with these pseudo-patterns. The dual-network architecture allows devices to share information via the weight updates of the neocortical network to the server without sending the actual data. We compare FedDual against alternatives elsewhere in the literature when applied to widely available datasets. FedDual not only achieves a margin of accuracy improvement over the alternatives, but also converges faster, requiring less communication rounds.
Adnan Ahmad, Vinh Loi Chau, Antonio Robles-Kelly, Shang Gao 0003, Longxiang Gao, Lianhua Chi, Wei Luo 0001
IJCNN6
2023 DeepMNF: Deep Multimodal Neuroimaging Framework for Diagnosing Autism Spectrum Disorder
Syed Qasim Abbas, Lianhua Chi, Yi-Ping Phoebe Chen
Artif. Intell. Medicine2
2023 Toward IoT Node Authentication Mechanism in Next Generation Networks
abstract
Although the next generation networks (5G-NGNs) provide a flexible infrastructure to support latency-sensitive and bandwidth-hungry mission-critical Internet of Things (IoT) applications, however, the 5G-IoT integration in NGNs has increased the threat surface. Unfortunately, IoT devices are resource constrained, and the traditional intrusion detection systems (IDS) approaches based on cryptography are not effective on 5G-IoT ecosystems. In this article, we propose an effective 5G-IoT node authentication approach that leverages unique radio frequency (RF) fingerprinting data to train the Deep learning model to detect legitimate and nonlegitimate IoT nodes. Our approach is based on Mahalanobis Distance theory and Chi-square distribution theories. The proposed approach achieves a higher detection accuracy (99.35%) as well as lower training time compared to other existing approaches which is a key benefit of our approach in NGNs. The experiments are conducted using ETSI-open source NFV management and orchestration (OSM-MANO) platform on Amazon Web Services (AWSs) cloud platform to verify how the proposed approach would fit in real-life scenarios. The method can be used as a standalone security system or as a part of multifactor authentication.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi, Shui Yu 0001
IEEE Internet Things J.5
2023 Dual-branch cross-dimensional self-attention-based imputation model for multivariate time series
abstract
In real-world scenarios, partial information losses of multivariate time series degrade the time series analysis. Hence, the time series imputation technique has been adopted to compensate for the missing values. Existing methods focus on investigating temporal correlations, cross-variable correlations, and bidirectional dynamics of time series, and most of these methods rely on recurrent neural networks (RNNs) to capture temporal dependency. However, the RNN-based models suffer from the common problems of slow speed and high complexity when dealing with long-term dependency. While some self-attention-based models without any recurrent structures can tackle long-term dependency with parallel computing, they do not fully learn and utilize correlations across the temporal and cross-variable dimensions. To address the limitations of existing methods, we propose a novel so-called dual-branch cross-dimensional self-attention-based imputation (DCSAI) model for multivariate time series, which is capable of performing global and auxiliary cross-dimensional analyses when imputing the missing values. In particular, this model contains masked multi-head self-attention-based encoders aligned with auxiliary generators to obtain global and auxiliary correlations in two dimensions, and these correlations are then combined into one final representation through three weighted combinations. Extensive experiments are presented to show that our model performs better than other state-of-the-art benchmarkers on three real-world public datasets under various missing rates. Furthermore, ablation study results demonstrate the efficacy of each component of the model.
Le Fang 0001, Wei Xiang 0001, Yuan Zhou 0006, Juan Fang 0004, Lianhua Chi, ZongYuan Ge
Knowl. Based Syst.5
2023 Transformed domain convolutional neural network for Alzheimer's disease diagnosis using structural MRI
Syed Qasim Abbas, Lianhua Chi, Yi-Ping Phoebe Chen
Pattern Recognit.2
2023 eX-ViT: A Novel explainable vision transformer for weakly supervised semantic segmentation
abstract
Recently vision transformer models have become prominent models for a multitude of vision tasks. These models, however, are usually opaque with weak feature interpretability, making their predictions inaccessible to the users. While there has been a surge of interest in the development of post-hoc solutions that explain model decisions, these methods can not be broadly applied to different transformer architectures, as rules for interpretability have to change accordingly based on the heterogeneity of data and model structures. Moreover, there is no method currently built for an intrinsically interpretable transformer, which is able to explain its reasoning process and provide a faithful explanation. To close these crucial gaps, we propose a novel vision transformer dubbed the eXplainable Vision Transformer (eX-ViT), an intrinsically interpretable transformer model that is able to jointly discover robust interpretable features and perform the prediction. Specifically, eX-ViT is composed of the Explainable Multi-Head Attention (E-MHA) module, the Attribute-guided Explainer (AttE) module with the self-supervised attribute-guided loss. The E-MHA tailors explainable attention weights that are able to learn semantically interpretable representations from tokens in terms of model decisions with noise robustness. Meanwhile, AttE is proposed to encode discriminative attribute features for the target object through diverse attribute discovery, which constitutes faithful evidence for the model predictions. Additionally, we have developed a self-supervised attribute-guided loss for our eX-ViT architecture, which utilizes both the attribute discriminability mechanism and the attribute diversity mechanism to enhance the quality of learned representations. As a result, the proposed eX-ViT model can produce faithful and robust interpretations with a variety of learned attributes. To verify and evaluate our method, we apply the eX-ViT to several weakly supervised semantic segmentation (WSSS) tasks, since these tasks typically rely on accurate visual explanations to extract object localization maps. Particularly, the explanation results obtained via eX-ViT are regarded as pseudo segmentation labels to train WSSS models. Comprehensive simulation results illustrate that our proposed eX-ViT model achieves comparable performance to supervised baselines, while surpassing the accuracy and interpretability of state-of-the-art black-box methods using only image-level labels.
Wei Xiang 0001, Juan Fang 0004, Yi-Ping Phoebe Chen, Lianhua Chi
Pattern Recognit.5
2023 Privacy and Accuracy for Cloud-Fog-Edge Collaborative Driver-Vehicle-Road Relation Graphs
abstract
There are three key roles in Intelligent Transportation Systems (ITS): driver, vehicle and road. However, existing static interactions among Driver-Vehicle-Road (DVR) are too passive to reflect the change of driver preferences, vehicle conditions, road conditions, etc. Therefore, we provide a data-driven Cloud-Fog-Edge Collaborative Driver-Vehicle-Road (CFEC-DVR) framework. The framework could self-adaptively evolves through continuous iteration to provide better ITS services for humans. The collaboration among DVR creates a lot of relation data that construct our relation graphs. Cloud brings some privacy risks. Relation graphs have great analytic value. As DVR collaboration, privacy quality and analytic accuracy are three key issues in the framework, we propose a Relation Graph Privacy-Preserving scheme with High Accuracy in our framework, which is named as RGPP-HA. Based on machine learning, our method nearly maximizes the difficulty for attackers to know exactly how many other roles are connected to the attacked role, which enhances the privacy quality. Meanwhile, we find as much valuable information as possible from roles’ encrypted relations for more accurately analytic performance. Based on the experiments, we compare the proposed scheme RGPP-HA with existing classic and relevant schemes. The experimental results show that our scheme has the best privacy quality and analytic accuracy. This further verifies the feasibility of CFEC-DVR framework.
Zongmin Cui, Zhixing Lu, Laurence T. Yang, Jing Yu 0012, Lianhua Chi, Shunli Zhang 0003
IEEE Trans. Intell. Transp. Syst.5
2023 MIRROR: Mining Implicit Relationships via Structure-Enhanced Graph Convolutional Networks
abstract
Data explosion in the information society drives people to develop more effective ways to extract meaningful information. Extracting semantic information and relational information has emerged as a key mining primitive in a wide variety of practical applications. Existing research on relation mining has primarily focused on explicit connections and ignored underlying information, e.g., the latent entity relations. Exploring such information (defined as implicit relationships in this article) provides an opportunity to reveal connotative knowledge and potential rules. In this article, we propose a novel research topic, i.e., how to identify implicit relationships across heterogeneous networks. Specially, we first give a clear and generic definition of implicit relationships. Then, we formalize the problem and propose an efficient solution, namely MIRROR, a graph convolutional network (GCN) model to infer implicit ties under explicit connections. MIRROR captures rich information in learning node-level representations by incorporating attributes from heterogeneous neighbors. Furthermore, MIRROR is tolerant of missing node attribute information because it is able to utilize network structure. We empirically evaluate MIRROR on four different genres of networks, achieving state-of-the-art performance for target relations mining. The underlying information revealed by MIRROR contributes to enriching existing knowledge and leading to novel domain insights.
Jiaying Liu 0006, Feng Xia 0001, Jing Ren 0001, Bo Xu 0008, Guansong Pang, Lianhua Chi
ACM Trans. Knowl. Discov. Data6
2023 Causal Discovery via Causal Star Graphs
abstract
Discovering causal relationships among observed variables is an important research focus in data mining. Existing causal discovery approaches are mainly based on constraint-based methods and functional causal models (FCMs). However, the constraint-based method cannot identify the Markov equivalence class and the functional causal models cannot identify the complex interrelationships when multiple variables affect one variable. To address the two aforementioned problems, we propose a new graph structure Causal Star Graph (CSG) and a corresponding framework Causal Discovery via Causal Star Graphs (CD-CSG) to divide a causal directed acyclic graph into multiple CSGs for causal discovery. In this framework, we also propose a generalized learning in CSGs based on a variational approach to learn the representative intermediate variable of CSG’s non-central variables. Through the generalized learning in CSGs, the asymmetry in the forward and backward model of CD-CSG can be found to identify the causal directions in the directed acyclic graphs. We further divide the CSGs into three categories and provide the causal identification principle under each category in our proposed framework. Experiments using synthetic data show that the causal relationships between variables can be effectively identified with CD-CSG and the accuracy of CD-CSG is higher than the best existing model. By applying CD-CSG to real-world data, our proposed method can greatly augment the applicability and effectiveness of causal discovery.
Boxiang Zhao, Shuliang Wang 0001, Lianhua Chi, Qi Li 0022, Xiaojia Liu, Jing Geng 0002
ACM Trans. Knowl. Discov. Data3
2023 HANM: Hierarchical Additive Noise Model for Many-to-One Causality Discovery
abstract
Discovering causal relationships among observed variables is a new research focus in the area of data mining. Methods based on the additive noise model have been proved to be efficient in the identification of cause-effect pairs. However, when trying to determine many-to-one causality, additive noise models often fail to identify the causal direction due to the complex interrelationships and interactions even though the generation of each causal relation follows the additive noise model, and become unreliable in practical applications. In this work, to identify the causal direction, we propose a Hierarchical Additive Noise Model (HANM) to convert many-to-one causality into an approximate one-to-one causality by generalizing multiple factors into an intermediate variable with a variational approach, and use asymmetry in the forward model and backward model of HANM to identify causal direction. Experiments using synthetic data show that many-to-one causality can be effectively identified through asymmetry with our proposed HANM and the accuracy of HANM is higher than the best existing model. By applying the model to real-world data, it can be seen that HANM can greatly augment the application scope of functional causal models for causal discovery.
Boxiang Zhao, Shuliang Wang 0001, Lianhua Chi, Chuanfeng Zhao, Hanning Yuan, Qi Li 0022, Xiaojia Liu, Jing Geng 0002, Ye Yuan 0001
IEEE Trans. Knowl. Data Eng.3
2023 Generative and Contrastive Self-Supervised Learning for Graph Anomaly Detection
abstract
Anomaly detection from graph data has drawn much attention due to its practical significance in many critical applications including cybersecurity, finance, and social networks. Existing data mining and machine learning methods are either shallow methods that could not effectively capture the complex interdependency of graph data or graph autoencoder methods that could not fully exploit the contextual information as supervision signals for effective anomaly detection. To overcome these challenges, in this paper, we propose a novel method, Self-Supervised Learning for Graph Anomaly Detection (SL-GAD). Our method constructs different contextual subgraphs (views) based on a target node and employs two modules,generative attribute regressionandmulti-view contrastive learningfor anomaly detection. While thegenerative attribute regressionmodule allows us to capture the anomalies in the attribute space, themulti-view contrastive learningmodule can exploit richer structure information from multiple subgraphs, thus abling to capture the anomalies in the structure space, mixing of structure, and attribute information. We conduct extensive experiments on six benchmark datasets and the results demonstrate that our method outperforms state-of-the-art methods by a large margin.
Yu Zheng 0013, Ming Jin 0005, Yixin Liu 0001, Lianhua Chi, Khoa Tran Phan, Yi-Ping Phoebe Chen
IEEE Trans. Knowl. Data Eng.4
2022 Impersonation Attack Detection in IoT Networks
abstract
The deployment of Internet of Things (IoT) networks is growing at an extraordinary speed from last decade and has expanded the interconnection of billions of nodes, providing a range of flexible communication and computing services, etc. We note that this significant expansion of the IoT surface has expanded the attack surfaces and is a danger to companies of every size from security aspects. The IoT devices are easy to compromise and therefore the attacker can easily act as an impersonator to impersonate other legitimate IoT nodes. This is known as impersonation attacks or spoofing attacks in wireless IoT networks. In this paper, we propose a new methodology to detect an impersonation attack in IoT networks. We use Mahalanobis Distance correlation theory based two-stage attack detection model to resist IoT node spoofing. The approach is evaluated on cloud platforms and is compared with the recent state-of-the-art literature. The proposal is deployed as a pluggable module in cloud networks. The key metrics of our evaluation and comparisons are accuracy with respect to the varying size of the IoT network, classification metrics, attack detection time, and CPU utilization.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi
GLOBECOM5
2022 Advanced calibration of mortality prediction on cardiovascular disease using feature-based artificial neural network
Alessio Bonti, Lianhua Chi, Mohamed Almorsy, Yi-Ping Phoebe Chen
Expert Syst. Appl.3
2022 HashWalk: An efficient node classification method based on clique-compressed graph embedding
Shuliang Wang 0001, Xiaorui Qin, Lianhua Chi
Pattern Recognit. Lett.3
2021 ANEMONE: Graph Anomaly Detection with Multi-Scale Contrastive Learning
abstract
Anomaly detection on graphs plays a significant role in various domains, including cybersecurity, e-commerce, and financial fraud detection. However, existing methods on graph anomaly detection usually consider the view in a single scale of graphs, which results in their limited capability to capture the anomalous patterns from different perspectives. Towards this end, we introduce a novel graph anomaly detection framework, namely ANEMONE, to simultaneously identify the anomalies in multiple graph scales. Concretely, ANEMONE first leverages a graph neural network backbone encoder with multi-scale contrastive learning objectives to capture the pattern distribution of graph data by learning the agreements between instances at the patch and context levels concurrently. Then, our method employs a statistical anomaly estimator to evaluate the abnormality of each node according to the degree of agreement from multiple perspectives. Experiments on three benchmark datasets demonstrate the superiority of our method.
Ming Jin 0005, Yixin Liu 0001, Yu Zheng 0013, Lianhua Chi, Yuan-Fang Li, Shirui Pan
CIKM4
2020 Web of Scholars: A Scholar Knowledge Graph
abstract
In this work, we demonstrate a novel system, namely Web of Scholars, which integrates state-of-the-art mining techniques to search, mine, and visualize complex networks behind scholars in the field of Computer Science. Relying on the knowledge graph, it provides services for fast, accurate, and intelligent semantic querying as well as powerful recommendations. In addition, in order to realize information sharing, it provides open API to be served as the underlying architecture for advanced functions. Web of Scholars takes advantage of knowledge graph, which means that it will be able to access more knowledge if more search exist. It can be served as a useful and interoperable tool for scholars to conduct in-depth analysis within Science of Science.
Jiaying Liu 0006, Jing Ren 0001, Wenqing Zheng, Lianhua Chi, Ivan Lee 0001, Feng Xia 0001
SIGIR4
2020 Improving multi-label chest X-ray disease diagnosis by exploiting disease and health labels dependencies
ZongYuan Ge, Dwarikanath Mahapatra, Xiaojun Chang, Zetao Chen, Lianhua Chi, Huimin Lu 0001
Multim. Tools Appl.5
2018 Deep Learning Hash for Wireless Multimedia Image Content Security
abstract
With the explosive growth of the wireless multimedia data on the wireless Internet, a large number of illegal images have been widely disseminated in wireless networks, which seriously endangers the content security of wireless networks. However, how to identify and classify illegal images quickly, accurately, and in real time is a key challenge for wireless multimedia networks. To avoid illegal images circulating on the Internet, each image needs to be detected, extracted features, and compared with the image in the feature library to verify the legitimacy of the image. An improved image deep learning hash (IDLH) method to learn compact binary codes for image search is proposed in this paper. Specifically, there are three major processes of IDLH: the feature extraction, deep secondary search, and image classification. IDLH performs image retrieval by the deep neural networks (DNN) as well as image classification with the binary hash codes. Different from other deep learning-hash methods that often entail heavy computations by using a conventional classifier, exemplified by K nearest neighbor (K-NN) and support vector machines (SVM), our method learns classifiers using binary hash codes, which can be learned synchronously in training. Finally, comprehensive experiments are conducted to evaluate IDLH method by using CIFAR-10 and Caltech 256 image library datasets, and the results show that the retrieval performance of IDLH method can effectively identify illegal images.
Yu Zheng 0035, Jiezhong Zhu, Wei Fang 0007, Lianhua Chi
Secur. Commun. Networks4
2018 Hashing for Adaptive Real-Time Graph Stream Classification With Concept Drifts
abstract
Many applications involve processing networked streaming data in a timely manner. Graph stream classification aims to learn a classification model from a stream of graphs with only one-pass of data, requiring real-time processing in training and prediction. This is a nontrivial task, as many existing methods require multipass of the graph stream to extract subgraph structures as features for graph classification which does not simultaneously satisfy "one-pass" and "real-time" requirements. In this paper, we propose an adaptive real-time graph stream classification method to address this challenge. We partition the unbounded graph stream data into consecutive graph chunks, each consisting of a fixed number of graphs and delivering a corresponding chunk-level classifier. We employ a random hashing function to compress the original node set of graphs in each chunk for fast feature detection when training chunk-level classifiers. Furthermore, a differential hashing strategy is applied to map unlimited increasing features (i.e., cliques) into a fixed-size feature space which is then used as a feature vector for stochastic learning. Finally, the chunk-level classifiers are weighted in an ensemble learning model for graph classification. The proposed method substantially speeds up the graph feature extraction and avoids unbounded graph feature growth. Moreover, it effectively offsets concept drifts in graph stream classification. Experiments on real-world and synthetic graph streams demonstrate that our method significantly outperforms existing methods in both classification accuracy and learning efficiency.
Lianhua Chi, Bin Li 0015, Xingquan Zhu 0001, Shirui Pan, Ling Chen 0006
IEEE Trans. Cybern.1
2017 End-to-end Network for Twitter Geolocation Prediction and Hashing
abstract
We propose an end-to-end neural network to predict the geolocation of a tweet. The network takes as input a number of raw Twitter metadata such as the tweet message and associated user account information. Our model is language independent, and despite minimal feature engineering, it is interpretable and capable of learning location indicative words and timing patterns. Compared to state-of-the-art systems, our model outperforms them by 2%-6%. Additionally, we propose extensions to the model to compress representation learnt by the network into binary codes. Experiments show that it produces compact codes compared to benchmark hashing algorithms. An implementation of the model is released publicly.
Jey Han Lau, Lianhua Chi, Khoi-Nguyen Tran, Trevor Cohn
IJCNLP(1)2
2014 Context-Preserving Hashing for Fast Text Classification
abstract
There have been a number of approximate algorithms for text similarity computation, such as min-wise hashing, random projection, and feature hashing, which are based on the bag-of-words representation. A limitation of their “flat-set” representation is that context information and semantic hierarchy cannot be preserved. In this paper, we aim to fast compute similarities between texts while also preserving context information. To take into account semantic hierarchy, we consider a notion of “multi-level exchangeability” which can be applied at word-level, sentence-level, paragraph-level, etc. We employ a nested-set to represent a multi-level exchangeable object. To fingerprint nested-sets for fast comparison, we propose a Recursive Min-wise Hashing (RMH) algorithm at the same computational cost of the standard min-wise hashing algorithm. Theoretical study and bound analysis confirm that RMH is a highly-concentrated estimator. The empirical studies show that the proposed context-preserving hashing method can significantly outperform min-wise hashing and feature hashing in accuracy at the same (or less) computational cost.
Lianhua Chi, Bin Li 0015, Xingquan Zhu 0001
SDM1
2013 Graph hashing and factorization for fast graph stream classification
abstract
Graph stream classification concerns building learning models from continuously growing graph data, in which an essential step is to explore subgraph features to represent graphs for effective learning and classification. When representing a graph using subgraph features, all existing methods employ coarse-grained feature representation, which only considers whether or not a subgraph feature appears in the graph. In this paper, we propose a fine-grained graph factorization approach for Fast Graph Stream Classification (FGSC). Our main idea is to find a set of cliques as feature base to represent each graph as a linear combination of the base cliques. To achieve this goal, we decompose each graph into a number of cliques and select discriminative cliques to generate a transfer matrix called Clique Set Matrix (M). By using M as the base for formulating graph factorization, each graph is represented in a vector space with each element denoting the degree of the corresponding subgraph feature related to the graph, so existing supervised learning algorithms can be applied to derive learning models for graph classification.
Ting Guo 0005, Lianhua Chi, Xingquan Zhu 0001
CIKM2
2013 Lightweight Management of Authorization Update on Cloud Data
abstract
While outsourcing data to cloud, security and efficiency issues should be taken into account. However, it is very challenging to design a secure and efficient mechanism supporting authorization updates. In this paper, we aim to provide a mechanism supporting authorization updates which only incurs a lightweight cost of authorization updates and meanwhile supports a high level of security. This mechanism is consisted of two encryption schemes performed in different layers. The inner-layer encryption scheme is performed on the original plaintext and the generated cipher text is called inner-layer cipher text, while a part of the inner-layer cipher text is encrypted by the outer-layer encryption scheme to generate cipher text, called out-layer cipher text. These two encryption schemes are both performed by data owner. The inner-layer encryption realizes the initial authorization policy, while the outer-layer encryption reflects the updated authorization policy. We implement the proposed mechanism and conduct extensive experiments. The experimental results demonstrate that the proposed mechanism outperforms previous existing approaches, e.g. single-layer encryption and double-layer encryption.
Zongmin Cui, Hong Zhu 0003, Lianhua Chi
ICPADS4
2013 Fast Graph Stream Classification Using Discriminative Clique Hashing
Lianhua Chi, Bin Li 0015, Xingquan Zhu 0001
PAKDD (1)1
2013 Lightweight key management on sensitive data in the cloud
abstract
ABSTRACT As cloud servers may not be trusted, sensitive data have to be transmitted and stored in an encrypted form. Major challenges for users are from the management (storage, update, protection, backup, and recoverability) of keys that can help users to decrypt authorized data available on the servers. In this paper, we propose a versatile approach for extremely lightweight key management, which is one of the most basic security tasks in cloud systems. In the multiple data owners scenario, each user only needs to manage a single key by our approach. With the help of the single key and a set of public information stored on the server, users can decrypt all authorized data from different data owners. Specifically, our paper proposes a novel access control model, proves the correctness and security, and analyzes the complexity of the model. Experimental results show that our approach significantly outperforms the single‐layer derivation encryption and double‐layer derivation encryption on the lightweight performance. Copyright © 2013 John Wiley & Sons, Ltd.
Zongmin Cui, Hong Zhu 0003, Lianhua Chi
Secur. Commun. Networks3
2012 Nested Subtree Hash Kernels for Large-Scale Graph Classification over Streams
abstract
Most studies on graph classification focus on designing fast and effective kernels. Several fast subtree kernels have achieved a linear time-complexity w.r.t. the number of edges under the condition that a common feature space (e.g., a subtree pattern list) is needed to represent all graphs. This will be infeasible when graphs are presented in a stream with rapidly emerging subtree patterns. In this case, computing a kernel matrix for graphs over the entire stream is difficult since the graphs in the expired chunks cannot be projected onto the unlimitedly expanding feature space again. This leads to a big trouble for graph classification over streams -- Different portions of graphs have different feature spaces. In this paper, we aim to enable large-scale graph classification over streams using the classical ensemble learning framework, which requires the data in different chunks to be in the same feature space. To this end, we propose a Nested Subtree Hashing (NSH) algorithm to recursively project the multi-resolution subtree patterns of different chunks onto a set of common low-dimensional feature spaces. We theoretically analyze the derived NSH kernel and obtain a number of favorable properties: 1) The NSH kernel is an unbiased and highly concentrated estimator of the fast subtree kernel. 2) The bound of convergence rate tends to be tighter as the NSH algorithm steps into a higher resolution. 3) The NSH kernel is robust in tolerating concept drift between chunks over a stream. We also empirically test the NSH kernel on both a large-scale synthetic graph data set and a real-world chemical compounds data set for anticancer activity prediction. The experimental results validate that the NSH kernel is indeed efficient and robust for graph classification over streams.
Bin Li 0015, Xingquan Zhu 0001, Lianhua Chi, Chengqi Zhang
ICDM3
2011 Comprehensive and efficient discovery of time series motifs
abstract
Time series motifs are previously unknown, frequently occurring patterns in time series or approximately repeated subsequences that are very similar to each other. There are two issues in time series motifs discovery, the deficiency of the definition of K -motifs given by Lin et al. (2002) and the large computation time for extracting motifs. In this paper, we propose a relatively comprehensive definition of K -motifs to obtain more valuable motifs. To minimize the computation time as much as possible, we extend the triangular inequality pruning method to avoid unnecessary operations and calculations, and propose an optimized matrix structure to produce the candidate motifs almost immediately. Results of two experiments on three time series datasets show that our motifs discovery algorithm is feasible and efficient.
Lianhua Chi, He-hua Chi, Yucai Feng, Shuliang Wang 0001, Zhong-sheng Cao
J. Zhejiang Univ. Sci. C1