EDBT 2026 Demo / reviewers in the wild / expert
Guanggang Geng
dblp:57/1750 · also Guang-Gang Geng
· DBLP profile ↗
52ranked-venue papers
7as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 15 · 5 first-author · 2 since 2021Security and privacy · 11 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 4 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RNN-based optimal consensus of high-order heterogeneous nonlinear MAS with input constraints
Yinyan Zhang, Yuxuan Xiong, Jilian Zhang, Guanggang Geng, Shuai Li 0002 |
Neurocomputing | 4 |
| 2026 | A Cloud-Assisted Multi-Dimensional Reputation Management Scheme With Privacy Preservation for Emergency Message Dissemination in VANETsabstractThe proliferation of Vehicular Ad-hoc Networks (VANETs) has brought new security challenges, as malicious vehicles may disseminate fake messages, thereby threatening road safety. Therefore, it is crucial to ensure the credibility of emergency messages while enabling their rapid dissemination. Cloud computing can enhance the efficiency of VANETs, thus facilitating the application of cloud-assisted security mechanisms. Therefore, this paper proposes a Cloud-assisted Multi-dimensional Reputation Management (CMRM) scheme with privacy preservation for emergency message dissemination in VANETs. Specifically, vehicles are represented by multi-dimensional reputation vectors together with dynamically updated multi-dimensional threshold vectors, which enable more fine-grained and reliable reputation evaluation. A confusing process is employed to conceal vehicle identities, reputation vectors, and threshold vectors, thereby ensuring privacy preservation. Furthermore, we design a Reputation Judgment Vector Extraction (RJVE) algorithm. This algorithm, through the collaboration of two cloud servers, can obtain the reputation judgment vector without disclosing sensitive data. Based on this, receivers use the Credible Message Judgment (CMJ) algorithm to efficiently evaluate the credibility of emergency messages. Theoretical analysis demonstrates that the CMRM scheme achieves privacy preservation and resists common attacks with low computation and communication overhead on both the Trusted Authority (TA) and vehicle sides. Simulation results further verify the robustness and efficiency of the CMRM scheme. The probability that a true message is trusted can quickly converge to over 80%, and the probability that a fake message is trusted quickly drops to near 0. Senting Ma, Zhiquan Liu 0001, Feixiang Ye, Guanggang Geng, Yinbin Miao, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | DMFI: A Dual-Modality Log Analysis Framework for Insider Threat Detection with LoRA-Tuned Language ModelsabstractInsider threat detection (ITD) poses a persistent and high-impact challenge in cybersecurity due to the subtle, long-term, and context-dependent nature of malicious insider behaviors. Traditional models often struggle to capture semantic intent and complex behavior dynamics, while existing LLMbased solutions face limitations in prompt adaptability and modality coverage. To bridge this gap, we propose DMFI, a dual-modality framework that integrates semantic inference with behavior-aware fine-tuning. DMFI converts raw logs into two structured views: (1) a semantic view that processes content-rich artifacts (e.g., emails, https) using instruction-formatted prompts; and (2) a behavioral abstraction, constructed via a 4 W -guided (When-Where-What-Which) transformation to encode contextual action sequences. Two LoRA-enhanced LLMs are fine-tuned independently, and their outputs are fused via a lightweight MLP-based decision module. We further introduce DMFI-B, a discriminative adaptation strategy that separates normal and abnormal behavior representations, improving robustness under severe class imbalance. Experiments on CERT r4.2 and r5.2 datasets demonstrate that DMFI outperforms state-of-the-art methods in detection accuracy. Our approach combines the semantic reasoning power of LLMs with structured behavior modeling, offering a scalable and effective solution for real-world insider threat detection. Kaichuan Kong, Dongjie Liu, Xiao-Bo Jin, Guanggang Geng, Zhiying Li 0003, Jian Weng 0001 |
ICDM | 4 |
| 2025 | Twin Progressive Generative Adversarial Network For High-Resolution Image InpaintingabstractImage inpainting aims to generate content for missing regions while maintaining visual coherence in the reconstructed images. Generative Adversarial Networks (GANs) have received increasing attention for their ability to repair high-resolution images. However, existing methods often overly rely on surrounding pixel information, overlooking noisy or incomplete information in boundary regions, which leads to blurred and unnatural results. To address these limitations, we propose a Twin Progressive Generative Adversarial Network (TP-GAN), which leverages global visual features from distant image contexts to reconstruct the overall structure and texture, improving the quality of high-resolution image inpainting. TP-GAN incorporates two generators and one discriminator, where the generators collaborate via exponential moving average optimization, focusing respectively on capturing fine details and global information. A progressive learning strategy is employed, starting with low-resolution restoration and gradually increasing resolution to simplify tasks and enhance adaptability to boundary regions. Extensive experimental evaluations on popular datasets demonstrate the superiority of TP-GAN. Zhiying Li 0003, Zhaoxin Fan, Kaichuan Kong, Xiao-Bo Jin, Guanggang Geng |
ICME | 6 |
| 2025 | Log2Sig: Frequency-Aware Insider Threat Detection via Multivariate Behavioral Signal DecompositionabstractInsider threat detection presents a significant challenge due to the deceptive nature of malicious behaviors, which often resemble legitimate user operations. However, existing approaches typically model system logs as flat event sequences, thereby failing to capture the inherent frequency dynamics and multiscale disturbance patterns embedded in user behavior. To address these limitations, we propose Log2Sig, a robust anomaly detection framework that transforms user logs into multivariate behavioral frequency signals, introducing a novel representation of user behavior. Log2Sig employs Multivariate Variational Mode Decomposition (MVMD) to extract IMFs, which reveal behavioral fluctuations across multiple temporal scales. Based on this, the model further performs joint modeling of behavioral sequences and frequency-decomposed signals: the daily behavior sequences are encoded using a Mamba-Based temporal encoder to capture long-term dependencies, while the corresponding frequency components are linearly projected to match the encoder’s output dimension. These dual-view representations are then fused to construct a comprehensive user behavior profile, which is fed into a multilayer perceptron for precise anomaly detection. Experimental results on the CERT r4.2 and r5.2 datasets demonstrate that Log2Sig significantly outperforms state-of-the-art baselines in both accuracy and F1 score. Kaichuan Kong, Dongjie Liu, Xiao-Bo Jin, Zhiying Li 0003, Guanggang Geng |
TrustCom | 5 |
| 2025 | Unveiling traffic paths: Explainable path signature feature-based encrypted traffic classificationabstractEncryption technology ensures secure transmission for internet communications but poses significant challenges for effective encrypted traffic classification, which categorizes traffic into distinct groups, facilitating the process of monitoring network activities to uncover patterns and extract valuable information applicable in areas such as network management and anomaly detection. To this end, machine learning has emerged as a powerful technology for conducting encrypted traffic classification without compromising user data privacy. Machine learning-based classification demonstrates remarkable capabilities in processing vast amounts of data through sophisticated handcrafted features, with traffic path signature features representing the cutting edge of this field. This method shows stable performance improvements for common encrypted traffic types using only packet length information. However, it also yields a high dimensionality of path signature features, complicating the training of lightweight models and hindering further innovation due to a lack of model explainability. In this paper, we first propose leveraging feature selection to conduct feature dimensionality reduction, and then try to focus on the explanation of the model from both global and local perspectives. Performance comparisons indicate that our proposed method significantly reduces the number of path signature features while preserving classification performance, which enhances computational efficiency and meets the demand for lightweight models in various application scenarios. Furthermore, this significant reduction in the feature dimensionality allows for the interpretability of the model, which gives the user a clear understanding of the modeling decision-making process. Kai-Chuan Kong, Xiao-Bo Jin, Guanggang Geng |
Comput. Secur. | 4 |
| 2025 | Progressive Enhancement Dehazing for object detection in extreme weather
Zhiying Li 0003, Junhao Wu 0003, Shuyuan Lin, Zheng Wang 0013, Xiao-Bo Jin, Guanggang Geng, Feiran Huang, Jian Weng 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | DPI-ITD: A Dual-Perspective Information-Driven Framework for Insider Threat Detection in IoT SystemsabstractIn Internet of Things (IoT) environments, insider threat detection has advanced with the integration of deep learning techniques, which can effectively model complex behaviors and heterogeneous data. However, the fragmented nature of IoT logs, behavioral redundancy, and the sparsity of insider actions increase detection complexity. While fine-grained behavior classification can improve accuracy, it also raises computational overhead, limiting applicability in resource-constrained scenarios. To address these challenges, we propose dual-perspective information-driven framework for insider threat detection (DPI-ITD), which combines user-centric and behavior-centric analyses to enhance detection efficiency and accuracy. DPI-ITD introduces a symbolic tagging strategy guided by tagging scores (TS), derived from user action diversity and behavioral context, to filter redundant fragments and focus on high-impact behaviors. It further incorporates an adaptive embedding mechanism based on GloVe, which dynamically adjusts the context window for rare but critical actions. Experiments on multiple closed and open behavioral datasets demonstrate DPI-ITD’s superior detection performance, scalability, and efficiency, confirming its suitability for lightweight deployment in real-world IoT security systems. Kai-Chuan Kong, Xiao-Bo Jin, Dongjie Liu, Zhiquan Liu 0001, Guanggang Geng |
IEEE Internet Things J. | 6 |
| 2025 | BPSO-AHDL-IDS: Binary Particle Swarm Optimization-Based Automated Hybrid Deep Learning Model for Intrusion Detection of Internet of ThingsabstractThe pervasive adoption of Internet-of-Things (IoT) systems has exposed critical vulnerabilities in cyber-security frameworks due to their decentralized deployment in unattended environments. While deep learning-based intrusion detection systems (IDSs) offer promising solutions, the design of hyper-parameters and neural architectures in existing models imposes prohibitive computational costs and expert dependency. To address these limitations, this work proposes an innovative automated hybrid deep learning method for IDS by employing a binary particle swarm optimization (BPSO) algorithm called BPSO-AHDL-IDS to effectively address the intrusion detection tasks of IoT. In BPSO-AHDL-IDS, the combination of convolutional neural network and recurrent neural network is considered as the hybrid deep learning model to extract features of the IoT dataset for accurate detection of intrusions. First, an efficient binary encoding mechanism is developed to describe the hyper-parameters and neural architectures of the hybrid deep learning model. Then, an efficient BPSO-based evolutionary operation is introduced to evolve the hyper-parameters and neural architectures to discover the optimized hybrid deep learning model. The performance of the proposed BPSO-AHDL-IDS method is testified by employing four datasets gathered from different IoT scenarios. It achieves accuracy of 0.9832, 0.9959, and 0.9897, precision of 0.9700, 0.9788, and 0.9843, recall of 0.9724, 0.9822, and 0.965, andF1-score of 0.9712, 0.9804, and 0.9744 on Bot-IoT, ToN-IoT, and Gas Pipeline datasets, respectively. On SWaT dataset, it achieves a precision of 0.9962, a recall of 0.9969, and anF1-score of 0.9965, respectively. The experimental results show the superiority of the proposed BPSO-AHDL-IDS method to machine learning methods, state-of-the-art manually designed and automated deep learning-based IDSs in terms ofaccuracy, precision, recall, andF1-score. Kang-Di Lu, Yao-Wei Yang, Chen Peng 0001, Guanggang Geng, Jian Weng 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Multi-Objective Discrete Extremal Optimization of Variable-Length Blocks-Based CNN by Joint NAS and HPO for Intrusion Detection in IIoTabstractIndustrial Internet of Things (IIoT) is an important part of industrial infrastructure but facing serious and evolving security threats in recent years. Deep learning has been widely considered as a promising solution for enhancing the security of IIoT. However, these existing deep learning models utilized in the intrusion detection of IIoT are manually developed that not only greatly rely on the experience of the designers but also is lack of utility due to the high model complexity. By taking into account the trade-off between the model performance and model complexity, this article makes the first attempt to propose a multi-objective joint optimization method of neural architecture search (NAS) and hyper-parameter optimization (HPO) based on multi-objective discrete extremal optimization (MODEO) to automatically design a lightweight convolutional neural network (CNN) for the intrusion detection task of IIoT, abbreviated as MODEO-CNN. A novel hybrid variable-length encoding strategy is developed by combing binary and integer encoding to characterize both the neural architectures including the number of blocks, the blocks-based network topology and the corresponding architecture parameters in CNN block, and some important hyper-parameters including batch size, learning rate, weight optimizer and regularization. The individual-based discrete multi-objective evolutionary process of MODEO is designed to obtain the Pareto-optimal CNN models. Three widely-used IIoT intrusion detection datasets, including the Gas Pipeline, BoT-IoT, and Power System Attack datasets, have been used to illustrate the superiority of the proposed MODEO-CNN over the state-of-the-art hand-craft models and two single-objective fixed-length blocks-based NAS models in terms of accuracy, precision, recall,$F_{1}$-Score, and model's million floating point operations. Kang-Di Lu, Min-Rong Chen, Guanggang Geng, Jian Weng 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | MoCC-BD-FID: Multi-Objective Clustering Combination-Based Backdoor Defense for Federated Intrusion Detection of Industrial Control SystemsabstractDeep learning and federated learning (FL) play a crucial role in ensuring the security of industrial control systems (ICSs), but they also face severe security threats, especially the threat of backdoor attacks. Most FL backdoor defense methods primarily focus on a single clustering strategy, resulting in low true positive rates (TPR) and true negative rates (TNR) in the attack classification task. Due to the excessive combination scheme of currently available clustering strategies, it is difficult to manually select an appropriate combination scheme of clustering strategies to defense backdoor attacks in federated ICSs. This work is the first time to automatically design a multi-objective clustering combination-based backdoor defense for federated intrusion detection in ICSs, called MoCC-BD-FID. The automated design issue of clustering strategies combination for backdoor defense is formulated as a mixed-variable multi-objective optimization problem, which considers both combinatorial variables, i.e., the combination length and the specific combination of clustering strategies, and continuous variables, i.e., the confidence levels of each combined clustering as the decision variables, and considers maximization of both TPR and TNR as the two objectives. To describe and evolve the different combinations of 12 clustering strategies with confidence levels, we develop an efficient mixed and variable-length encoding mechanism, and the specifically tailored crossover operation and mutation operation under the framework of nondominated sorting genetic algorithm II. The experiments are conducted on the three widely-used ICS datasets including Secure Water Treatment, Water Distribution, and Power System Attack datasets under two different backdoor attacks. The experimental results demonstrate that MoCC-BDFID outperforms the single clustering strategy-based backdoor defense methods and five existing backdoor defense methods, i.e., Krum, Weak-DP, FoolsGold, DeepSight, and CrowdGuard, in terms of the classification accuracy of the poisoned model on regular samples and backdoor samples, TPR, and TNR. Jun-Min Shao, Kang-Di Lu, Guanggang Geng, Jian Weng 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Collaborative Edge-Cloud Data Transfer Optimization for Industrial Internet of ThingsabstractIn the Industrial Internet of Things, it is necessary to reserve enough bandwidth resources according to the maximum traffic peak. However, bandwidth reservation based on the maximum traffic peak leads to low resource utilization. In this paper, we propose a data transfer optimization solution, based on the cooperation of different entities in the local area, which strives to deliver data acquired by sensors to the cloud in a reliable manner and improve bandwidth utilization to save limited network resources. In our solution, the data transfers from the sensors in a local network are controlled by a local controller and some edge gateways with acceptable cost such that no congestion occurs in the path to the cloud and the bandwidth requirement of each flow can be met. To obtain a tradeoff between resource utilization and transfer delay, we study the problem of minimizing the maximum rate peak of periodic real-time traffic from distributed sensors and propose an algorithm to solve this problem with a desirable lower boundary of the performance. In addition, we design an application-level forwarding method that significantly improves resource utilization and a method of implementing reliable sampling instant adjustment. The experimental results show that our solution significantly improves resource utilization without producing network congestion. Xinchang Zhang 0001, Maoli Wang, Zhiwei Yan, Guanggang Geng |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | Evolutionary Adversarial Autoencoder for Unsupervised Anomaly Detection of Industrial Internet of ThingsabstractThe rapid growth of interconnected smart devices and advanced computing technologies in the industrial Internet of Things (IIoT) has significantly enhanced operational resilience and performance but also increased cybersecurity risks. While deep learning shows promise in IIoT security, it faces challenges due to the lack of labeled data and reliance on human expertise for unsupervised anomaly detection. To address these challenges, a novel automated adversarial deep learning-based unsupervised anomaly detection method called EvoAAE is proposed to optimize the hyperparameters and neural architectures of adversarial variational autoencoder (VAE) for securing IIoT. Specifically, a generative adversarial network-based VAE is employed to adversarially generate multivariate time series. Then, particle swarm optimization with an efficient binary encoding strategy is designed to evolve hyperparameters and neural architectures in adversarial VAE including batch size, learning rate, the type of optimizer, the number of convolutional layer, the number of kernels of convolutional layer, kernel size, the type of normalization layer, and the type of active function. The experimental results indicate that EvoAAE achieves notable performance across four IIoT datasets in industrial control domain, i.e., secure water treatment, water distribution, Mars Science Laboratory, and power system domain, i.e., power system attack with precision of 0.949, 0.8356, 0.972, and 0.981, recall of 0.971, 0.9214, 0.964, and 0.979, and$F_{1}$-score of 0.960, 0.8764, 0.968, and 0.980, respectively. Yao-Wei Yang, Kang-Di Lu, Guanggang Geng, Jian Weng 0001 |
IEEE Trans. Reliab. | 4 |
| 2024 | Automated federated learning for intrusion detection of industrial control systems based on evolutionary neural architecture search
Jun-Min Shao, Kang-Di Lu, Guanggang Geng, Jian Weng 0001 |
Comput. Secur. | 4 |
| 2024 | STFT-TCAN: A TCN-attention based multivariate time series anomaly detection architecture with time-frequency analysis for cyber-industrial systemsabstractNetworks and industrial systems play a pivotal role in modern society, and their security has garnered increasing attention. Anomalies within industrial equipment may propagate through fault transmission, leading to a cascade of failures. Additionally, cyberattacks on equipment can result in significant losses. Therefore, in the realm of industrial and cyberspace domains, an effective multivariate time series anomaly detection system for monitoring equipment is instrumental in ensuring the healthy operation of the machinery. Nevertheless, detecting anomalies in numerous time series remains challenging, stemming from the absence of anomaly labels and the complexity of the data patterns. Existing algorithms predominantly concentrate on modeling within the time domain, falling short in fully leveraging the informative features present in frequency domain data, resulting in diminished detection performance. This paper introduces STFT-TCAN, a model for anomaly detection in time series that seamlessly integrates information from both time and frequency domains for extracting data features. Sliding windows and the Short Time Fourier Transform (STFT) are utilized to construct a frequency matrix, effectively amalgamating the characteristics of both time and frequency domains within the time series. Furthermore, the model employs Temporal Convolutional Networks (TCN) and Transformer attention mechanisms (which combined to form the TCAN module) to capture the features of multivariate time series, thereby resulting in heightened detection accuracy. The proposed model undergoes validation on six publicly available datasets, showcasing the superior performance of the STFT-TCAN model in comparison to current baseline methods. It adeptly extracts features from both frequency and time domains in sequential data, thereby achieving state-of-the-art performance in tasks related to anomaly detection in multivariate time series. Fei-Fan Tu, Dongjie Liu, Zhiwei Yan, Xiao-Bo Jin, Guanggang Geng |
Comput. Secur. | 5 |
| 2024 | DoFA: Adversarial examples detection for SAR images by dual-objective feature attribution
Yu Zhang 0201, Min-Rong Chen, Guanggang Geng, Jian Weng 0001, Kang-Di Lu |
Expert Syst. Appl. | 4 |
| 2024 | Norm-Based Finite-Time Convergent Recurrent Neural Network for Dynamic Linear InequalityabstractVarious recurrent neural network (RNN) models, especially zeroing neutral network (ZNN) models, have been investigated to solve time-varying linear inequalities (TVLI) and applied to different important fields. Existing ZNN models can solve TVLI in finite time by using complicated elementwise nonlinear functions, which brings a concern about cost for hardware implementation. To achieve a balance between implementation cost and convergence performance, this article explores a new RNN model based on ZNN by using a two-norm method for solving TVLI, which is called norm-based ZNN (NBZNN), and the proposed model can achieve finite-time convergence without the assistance of elementwise nonlinear activation functions. Strict theoretical analysis is given on convergence properties of the proposed model, showing its global finite-time convergence, which is preserved under a class of bounded noises. For the first time, our work shows that a finite-time convergent RNN model can be designed for solving TVLI without using elementwise nonlinear activation functions. Computer simulation results further verify the effectiveness and superiority of the proposed NBZNN model for solving TVLI. An application to robotics further demonstrates the efficacy of the proposed NBZNN model. Linyan Dai 0001, Yinyan Zhang, Guanggang Geng |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | GNN Model With Robust Finite-Time Convergence for Time-Varying Systems of Linear EquationsabstractDynamic neural networks are considered as an effective method in the field of scientific computing, among which gradient neural networks (GNNs) are an efficient method for solving static problems. However, when solving dynamic problems, the current GNNs are often subject to lagging errors. In this article, we develop a finite-time convergent GNN (FTCGNN) model for solving static and time-varying systems of linear equations. Different from zeroing neural networks (ZNNs) dedicated to time-varying problem solving, the FTCGNN model has finite-time convergence regardless of the existence of time-varying noises. Simulation results show that the FTCGNN model is effective, among which the comparisons with existing GNNs and ZNNs validate the advantages of the FTCGNN model. Yinyan Zhang, Bolin Liao, Guanggang Geng |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | A Practical Framework of Blockchain in IoT Information ManagementabstractIn order to improve the security of IoT information systems, this paper proposes the Blockchain-based Framework for Securing IoT Information (BFSII), which is built on consortium blockchain and the edge IoT architectures. This paper addresses data security in smart hotels as a research scenario. The majority of data generated by IoT devices in smart hotels contains users' private information, which is susceptible to alteration and leakage during transmission and storage. The BFSII solution leverages the decentralized nature of blockchain to enhance data traceability and tamper-proof capabilities. And it uses edge IoT architecture and consortium blockchain to improve system operational efficiency. Sensitive data generated by IoT devices are protected in BFSII. The experiment's findings show that BFSII can boost smart hotel system security while maintaining operational effectiveness. The information management system of smart hotels is provided with an inventive and secure solution by the BFSII framework. Quanlong Guan, Jiawei Lei, Chaonan Wang, Guanggang Geng, Yuansheng Zhong, Liangda Fang, Xiujie Huang, Weiqi Luo 0002 |
SMC | 4 |
| 2023 | Differential evolution-based convolutional neural networks: An automatic architecture design method for intrusion detection in industrial control systems
Guanggang Geng, Jian Weng 0001, Kang-Di Lu, Yu Zhang 0201 |
Comput. Secur. | 3 |
| 2023 | Single-state distributed k-winners-take-all neural network modelabstractDistributed k-winners-takes-all (k-WTA) neural network (k-WTANN) models have better scalability than centralized ones. In this work, a distributed k-WTANN model with a simple structure is designed for the efficient selection of k winners among a group of more than k agents via competition based on their inputs. Unlike an existing distributed k-WTANN model, the proposed model does not rely on consensus filters, and only has one state variable. We prove that under mild conditions, the proposed distributed k-WTANN model has global asymptotic convergence. The theoretical conclusions are validated via numerical examples, which also show that our model is of better convergence speed than the existing distributed k-WTANN model. Yinyan Zhang, Shuai Li 0002, Xuefeng Zhou, Jian Weng 0001, Guanggang Geng |
Inf. Sci. | 5 |
| 2023 | Initialization-Based k-Winners-Take-All Neural Network Model Using Modified Gradient DescentabstractThe k -winners-take-all ( k -WTA) problem refers to the selection of k winners with the first k largest inputs over a group of n neurons, where each neuron has an input. In existing k -WTA neural network models, the positive integer k is explicitly given in the corresponding mathematical models. In this article, we consider another case where the number k in the k -WTA problem is implicitly specified by the initial states of the neurons. Based on the constraint conversion for a classical optimization problem formulation of the k -WTA, via modifying the traditional gradient descent, we propose an initialization-based k -WTA neural network model with only n neurons for n -dimensional inputs, and the dynamics of the neural network model is described by parameterized gradient descent. Theoretical results show that the state vector of the proposed k -WTA neural network model globally asymptotically converges to the theoretical k -WTA solution under mild conditions. Simulative examples demonstrate the effectiveness of the proposed model and indicate that its convergence can be accelerated by readily setting two design parameters. Yinyan Zhang, Shuai Li 0002, Guanggang Geng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Delay-Optimized Multicast Tree Packing in Software-Defined NetworksabstractIn traditional networks, the multicast tree packing solutions usually aim to minimize the overall multicast tree cost, which can effectively improve network accommodation capacity but is disadvantageous to fully use network resources. In this article, we propose a delay-optimized multicast tree packing problem called delivery delay minimized multicast tree packing (DDMMTP), which aims to minimize the average source-destination delay, under constraints on the bandwidth and maximum source-destination delay, according to available network resources. A low source-destination delay is desirable because it improves the service quality, especially for time-sensitive applications. In practice, the DDMMTP is highly valuable for the software-defined network (SDN) mainly because this new network paradigm has the ability to rapidly rearrange multicast routes on demand. The DDMMTP problem is NP-hard. We solve it approximately by a batched multicast tree packing algorithm and a network accommodation capacity improvement algorithm that adjusts existing multicast paths on demand. We also propose a source-destination delay improvement algorithm to further reduce source-destination delays based on new available network resources. Xinchang Zhang 0001, Yinglong Wang 0001, Guanggang Geng, Jiguo Yu |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Multi-scale semantic deep fusion models for phishing website detectionabstractIn view of semantic counterfeiting characteristics of phishing websites and their multi-scale composition, this paper fully considers the semantic information of different scales, and proposes three semantic-based phishing detection models at different depths using various deep learning methods. The proposed three models are Multi-scale Data-layer Fusion (MDF) model, Multi-scale Feature-layer Fusion (MFF) model and Multi-scale In-depth Fusion(MIF) model. Experimental results on a constructed complex dataset show that the three models all have good recognition capabilities and the MIF model achieves the best performance on a complex dataset, with an F1-Measure of 0.9830, AUC value of 0.9993 and a false positive rate of 0.0047. Then with further comparison with both visual and text methods and an active discovery experiment lasting for 6 months with 3016 phishing websites detected in the real network environment, it is found that the proposed model is both competitive and practical for real detection scenarios. Dongjie Liu, Guanggang Geng, Xinchang Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2022 | Sparse matrix factorization with L2, 1 norm for matrix completion
Xiao-Bo Jin, Jianyu Miao, Qiufeng Wang 0001, Guanggang Geng, Kaizhu Huang |
Pattern Recognit. | 4 |
| 2022 | Seeing Traffic Paths: Encrypted Traffic Classification With Path Signature FeaturesabstractAlthough many network traffic protection methods have been developed to protect user privacy, encrypted traffic can still reveal sensitive user information with sophisticated analysis. In this paper, we propose ETC-PS, a novel encrypted traffic classification method with path signature. We first construct the traffic path with a session packet length sequence to represent the interactions between the client and the server. Then, path transformations are conducted to exhibit its structure and obtain different information. A multiscale path signature is finally computed as a kind of distinctive feature to train the traditional machine learning classifier, which achieves highly robust accuracy and low training overhead. Six publicly available datasets with different traffic types of HTTPS/1, HTTPS/2, QUIC, VPN, non-VPN, Tor, and non-Tor are used to conduct closed-world and open-world evaluations to verify the effectiveness of ETC-PS. The experimental results demonstrate that ETC-PS is superior to the state-of-the-art methods in terms of accuracy, f1 score, time complexity, and stability. Guanggang Geng, Xiao-Bo Jin, Dongjie Liu, Jian Weng 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Dealing with the Domain Name Abuse: Issues and ApproachesabstractTo facilitate their worldwide access, online resources such as websites, mobile apps as well as email addresses would generally utilize domain names to label their logical locations on the Internet. Meanwhile, people who want to visit somewhere over the Internet would also tend to resort to a specific domain name first. Therefore, domain names are virtually playing an essential role during the interacting procedure between humans and online resources. However, domain names could sometimes be exploited maliciously and involved into some unexpected scenarios such as phishing, spamming, and harmful content distributing. We refer to these kinds of exploitation here as domain name abuse. Due to the explosive growth over the past years, domain name abuse has been drawing plenty of attentions both from academic and industrial areas. However, we note that there still exist a range of fundamental yet vital issues that have not been well addressed so far. In this paper, we conduct a comprehensive reexamination towards some of these major issues as well as relevant approaches regarding domain name abuse handling, and try further to set forth some of our own responses and corresponding updates towards them. We believe that these issues and relevant approaches could lead to extensive impact towards further progress in this area, and our work would help the community to deal with domain name abuse in a more accountable and practical manner. Xuebiao Yuchi, Zhiwei Yan, Kejun Dong, Guanggang Geng, Sachin Shetty |
EUC | 5 |
| 2021 | An efficient multistage phishing website detection model based on the CASE feature framework: Aiming at the real web environmentabstractPhishing has become a favorite method of hackers for committing data theft and continues to evolve. As long as phishing websites continue to operate, many more people and companies will suffer privacy leaks or financial losses. Therefore, the demand for fast and accurate phishing website detection grows stronger. However, the existing phishing detection methods do not fully analyze the features of phishing, and the performance and efficiency of the models only apply to certain limited datasets and need to be improved to be applied to the real web environment. This paper fully considers the social engineering principles of phishing, proposes a comprehensive and interpretable CASE feature framework and designs a multistage phishing detection model to effectively detect phishing sites, especially in the real web environment, where high efficiency and performance and extremely low false alarm rates are required. To fully verify the proposed method, two kinds of data experiments were carried out. One was the comparative experiments among different features and different detection models on CASE, which covers both classic machine learning and deep learning algorithms based on a constructed complex dataset. The other was a one-year phishing discovery experiment in the real web environment. The proposed method achieves better detection results under the premise of significantly shortening the execution time and works well in real phishing discovery, which proves its high practicability in reality. Dongjie Liu, Guanggang Geng, Xiao-Bo Jin, Wei Wang 0083 |
Comput. Secur. | 2 |
| 2019 | Stochastic Conjugate Gradient Algorithm With Variance ReductionabstractConjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction1and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functions. We experimentally demonstrate that the CG with variance reduction algorithm converges faster than its counterparts for four learning models, which may be convex, nonconvex or nonsmooth. In addition, its area under the curve performance on six large-scale data sets is comparable to that of the LIBLINEAR solver for the L2-regularized L2-loss but with a significant improvement in computational efficiency. Xiao-Bo Jin, Xu-Yao Zhang, Kaizhu Huang, Guanggang Geng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Approximately optimizing NDCG using pair-wise loss
Xiao-Bo Jin, Guanggang Geng, Guosen Xie, Kaizhu Huang |
Inf. Sci. | 2 |
| 2017 | A robust internet abuse detection methodabstractJavaScript can modify HyperText Markup Language(HTML) code.tag of HTML can load pages from other websites. These technologies are widely used nowadays for constructing flexible and robust web services. However, illegal websites also use these technologies to hide illegal contents. For traditional methods using web text to detect these illegal webpages (just getting initial source codes that aren't parsed by a browser), they can't give a precise judgment. In this paper, we solved the challenge and proposed a robust Internet abuse detection method. We get HTML codes that had been parsed by simulating the progress how browsers work in order to gather the necessary materials for text detection methods. Besides, we not only extracted the text features of HTML code, but also extracted structure features of HTML codes. With the experiments under detecting Internet abuse scenarios, we demonstrated that the proposal is efficient. Zhou Fa, Guanggang Geng, Zhiwei Yan, Xiaodong Lee |
IEEE BigData | 2 |
| 2017 | Boosting the phishing detection performance by semantic analysisabstractPhishing is increasingly severe in recent years, which seriously threatens the privacy and property security of netizens. Phishing is essentially a counterfeiting of brands. In order to effectively cheat the victim, phishing sites are visually and semantically highly similar to real sites. In recent years, anti-phishing methods based on machine learning are mainstream anti-phishing methods. The effectiveness of the machine learning models hinges on the extracted statistical features. However, the extracted statistical features mainly focus on visual similarity, stealing information and third-party services, which ignore the semantic information of web pages. Therefore, we extract a series of semantic features through word2vec to better describe the features of phishing sites, and further fuse them with other multi-scale statistical features to construct a more robust phishing detection model. The experimental results on the actual data sets show that the majority of phishing websites are effectively identified by only mining the semantic features of word embeddings. The phishing detection models based on fusion features obtained the best detection results, which shows that semantic features and other statistical features have good complementarity. The proposed method provides a promising way for phishing detection in actual Internet environment, which boosts the phishing detection performance effectively. Xiao-Bo Jin, Zhiwei Yan, Guanggang Geng |
IEEE BigData | 5 |
| 2017 | Towards tackling privacy disclosure issues in Domain Name ServiceabstractServing as the global Internet's phonebook, the Domain Name Service (DNS) helps to translate human-friendly domain names into machine-readable IP addresses, which makes DNS of great importance to the operation of the Internet and virtually relied on by today's almost all kinds of Internet-based activities. As such, people whoever want to go anywhere over the Internet will need to refer to the DNS first. Therefore, it has become an ideal way to conduct online privacy exploitations through the DNS due to people's pervasive usage of the Internet. However, the current DNS doesn't provide any countermeasure against this kind of exploitation, and thus risks severe privacy disclosure problems. In this paper, we give a comprehensive empirical analysis of DNS privacy disclosure problems by exploring potential privacy leaking paths in the DNS. Then we further identify and describe multiple criterions of validity systematically that are obligated when considering DNS privacy preservation. Finally, we propose a simple DNS privacy preserving solution with significant deployment potential in the current DNS, which can only lead to a moderate level of extra query latency perceived by end users. Xuebiao Yuchi, Guanggang Geng, Zhiwei Yan, Xiaodong Lee |
IM | 2 |
| 2017 | On-Demand DTN Communications in Heterogeneous Access Networks Based on NDNabstractWith the development of wireless communication technologies, the mobile terminal can be installed with more than one wireless interfaces to support its flexible roaming and vertical handover between heterogeneous access networks. Traditionally, the ongoing session has to be re-established when the multi-interface terminal hands over vertically while maintaining the Delay Tolerant Networking (DTN) communication. In this paper, a novel solution is proposed to efficiently support the on-demand DTN communication under the heterogeneous access networks based on Named Data Networking (NDN) architecture. Zhiwei Yan, Guanggang Geng, Hidenori Nakazato, Kashif Nisar, Ag Asri Ag Ibrahim |
VTC Spring | 2 |
| 2016 | Phishing detection based on newly registered domainsabstractPhishing is a security attack that involves the creation of websites that mimic legitimate websites, and these fraud websites bring Internet users a lot of loss. Traditional anti-phishing methods usually worked in a passive way by receiving report data of user. Due to the growing shorter survival time of phishing, this kind of methods is not efficient enough to find and take down new phishing attacks. In this paper, we propose an Intelligent Phishing Detection (IPD) system to address phishing detection problem actively. Specially, IPD first generates the detection dataset from the global massive domain name registration data automatically; then it applies the Naïve Bayes algorithm which is optimized by position-based features to achieve the high precision detection; finally, in order to find more phishing websites, IPD expands detection dataset by generating Uniform Resource Locator (URL) templates based on the detection results. The experimental results of IPD demonstrate the effectiveness and timeliness in detection phishing websites. Xueni Li, Guanggang Geng, Zhiwei Yan, Xiaodong Lee |
IEEE BigData | 2 |
| 2016 | A Cross-Domain Hidden Spam Detection Method Based on Domain Name Resolution
Cuicui Wang, Guanggang Geng, Zhiwei Yan |
QSHINE | 2 |
| 2016 | A Novel DMM Architecture Based on NDN
Zhiwei Yan, Jong-Hyouk Lee, Guanggang Geng, Xiaodong Lee |
QSHINE | 3 |
| 2015 | Combination of multiple bipartite ranking for multipartite web content quality evaluation
Xiao-Bo Jin, Guanggang Geng, Minghe Sun, Dexian Zhang |
Neurocomputing | 2 |
| 2015 | Combating phishing attacks via brand identity and authorization featuresabstractAbstract Phishing, also called brand spoofing, has become the most troubling scam on the Internet, which seriously threatens the Web security. The essence of phish is that “robbers” use false sites, which look like a trustworthy brand site, where favicon, logo and copyright notice are important brand identities. We analyzed 78‐day phishing data of PhishTank and Anti‐Phishing Working Group (APWG). The statistics show that more than 98.93% phishing sites contain at least one brand entity—favicon, logo or copyright notice. Indeed, only a few lowest‐quality phishing campaigns do not use such brand elements. Obviously, brand entities are powerful weapons of phishers to trick users. By analyzing the characteristics of brand entities in phishing sites, several brand identity features are extracted. However, only brand entities do not consider whether the Web page with brand entities belongs to the corresponding brand or has an authorization to use the brand entities. To solve this problem, redirection, incoming links and Domain Name System (DNS) information‐based brand authorization features are further extracted to discriminate the sites with branding rights from phishing sites. Based on extracted features, statistical anti‐phishing classification models are trained. We collected a diverse spectrum of corpora containing 3863 phishing cases from PhishTank and APWG, and 17 571 legitimate samples from DMOZ, Google and DNS resolution log. Experimental evaluations show that the model achieves 98.8% true positive rate and 0.09% false positive rate, which demonstrates the competitive performances of extracted features for statistical anti‐phishing in practice. Copyright © 2014 John Wiley & Sons, Ltd. Guanggang Geng, Xiaodong Lee, Yan-Ming Zhang 0001 |
Secur. Commun. Networks | 1 |
| 2015 | MTC: A Fast and Robust Graph-Based Transductive Learning MethodabstractDespite the great success of graph-based transductive learning methods, most of them have serious problems in scalability and robustness. In this paper, we propose an efficient and robust graph-based transductive classification method, called minimum tree cut (MTC), which is suitable for large-scale data. Motivated from the sparse representation of graph, we approximate a graph by a spanning tree. Exploiting the simple structure, we develop a linear-time algorithm to label the tree such that the cut size of the tree is minimized. This significantly improves graph-based methods, which typically have a polynomial time complexity. Moreover, we theoretically and empirically show that the performance of MTC is robust to the graph construction, overcoming another big problem of traditional graph-based methods. Extensive experiments on public data sets and applications on web-spam detection and interactive image segmentation demonstrate our method's advantages in aspect of accuracy, speed, and robustness. Yan-Ming Zhang 0001, Kaizhu Huang, Guanggang Geng, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | A Taxonomy of Hyperlink Hiding Techniques
Guanggang Geng, Xiu-Tao Yang, Wei Wang 0083, Chi-Jie Meng |
APWeb | 1 |
| 2013 | Fast kNN Graph Construction with Locality Sensitive Hashing
Yan-Ming Zhang 0001, Kaizhu Huang, Guanggang Geng, Cheng-Lin Liu 0001 |
ECML/PKDD (2) | 3 |
| 2012 | Multi-label learning vector quantization algorithm
Xiao-Bo Jin, Guanggang Geng, Dexian Zhang |
ICPR | 2 |
| 2012 | Statistical cross-language Web content quality assessment
Guanggang Geng, Wei Wang 0083, An-Lei Hu |
Knowl. Based Syst. | 1 |
| 2011 | Statistical feature extraction for cross-language web content quality assessmentabstractWeb content quality assessment is a typical static ranking problem. Heuristic content and TFIDF features based statistical systems have proven effective for Web content quality assessment. But they are all language dependent features, which are not suitable for cross-language ranking. In this paper, we fuse a series of language-independent features including hostname features, domain registration features, two-layer hyperlink analysis features and third-party Web service features to assess the Web content quality. The experiments on ECML/PKDD 2010 Discovery Challenge cross-language datasets show that the assessment is effective. Guanggang Geng, Xiaodong Li 0001, Wei Wang 0083 |
SIGIR | 1 |
| 2011 | A Hybrid System to Find & Fight Phishing Attacks ActivelyabstractTraditional anti-phishing methods and tools always worked in a passive way to receive users' submission and determine phishing URLs. Usually, they are not fast and efficient enough to find and take down phishing attacks. We analyze phishing reports from Anti-phishing Alliance of China(APAC) and propose a hybrid method to discover phishing attacks in an active way based on DNS query logs and known phishing URLs. We develop and deploy our system to report living phishing URLs automatically to APAC every day. Our system has become a main channel in supplying phishing reports to APAC in China and can be a good complement to traditional anti-phishing methods. Wei Wang 0083, Guanggang Geng, Yali Xiao, Xiaodong Li 0001 |
Web Intelligence | 4 |
| 2009 | Link based small sample learning for web spam detectionabstractRobust statistical learning based web spam detection system often requires large amounts of labeled training data. However, labeled samples are more difficult, expensive and time consuming to obtain than unlabeled ones. This paper proposed link based semi-supervised learning algorithms to boost the performance of a classifier, which integrates the traditional Self-training with the topological dependency based link learning. The experiments with a few labeled samples on standard WEBSPAM-UK2006 benchmark showed that the algorithms are effective. Guanggang Geng, Qiudan Li, Xinchang Zhang 0001 |
WWW | 1 |
| 2008 | Improving web spam detection with re-extracted featuresabstractWeb spam detection has become one of the top challenges for the Internet search industry. Instead of using some heuristic rules, we propose a feature re-extraction strategy to optimize the detection result. Based on the predicted spamicity obtained by the preliminary detection, through the host level web graph, three types of features are extracted. Experiments on WEBSPAM-UK2006 benchmark show that with this strategy, the performance of web spam detection can be improved evidently. Categories and Subject Descriptors Guanggang Geng, Chunheng Wang, Qiudan Li |
WWW | 1 |
| 2008 | Improving personalized services in mobile commerce by a novel multicriteria rating approachabstractWith the rapid growth of wireless technologies and mobile devices, there is a great demand for personalized services in m-commerce. Collaborative filtering (CF) is one of successful techniques to produce personalized recommendations for users. This paper proposes a novel approach to improve CF algorithms, where the contextual information of a user and the multicriteria ratings of an item are considered besides the typical information on users and items. The multilinear singular value decomposition (MSVD) technique is utilized to explore both explicit relations and implicit relations among user, item and criterion. We implement the approach in an existing m-commerce platform, and encouraging experimental results demonstrate its effectiveness. Qiudan Li, Chunheng Wang, Guanggang Geng |
WWW | 3 |
| 2007 | Fighting Link Spam with a Two-Stage Ranking Strategy
Guanggang Geng, Chunheng Wang, Qiudan Li, Yuanping Zhu |
ECIR | 1 |
| 2007 | A Concept Lattice-Based Kernel Method for Mining Knowledge in an M-Commerce System
Qiudan Li, Chunheng Wang, Guanggang Geng, Ruwei Dai |
ISNN (1) | 3 |
| 2007 | A novel collaborative filtering-based framework for personalized services in m-commerceabstractWith the rapid growth of wireless technologies and handheld devices, m-commerce is becoming a promising research area. Personalization is especially important to the success of m-commerce. This paper proposes a novel collaborative filtering-based framework for personalized services in m-commerce. The framework extends our previous work by using Online Analytical Processing (OLAP) to represent the relations among user, content and context information, and adopting a multi-dimensional collaborative filtering model to perform inference. It provides a powerful and well-founded mechanism to personalization for m-commerce. We implemented it in an existing m-commerce platform, and experimental results demonstrate its feasibility and correctness. Qiudan Li, Chunheng Wang, Guanggang Geng, Ruwei Dai |
WWW | 3 |