EDBT 2026 Demo / reviewers in the wild / expert
Huamin Feng
dblp:64/75 · also HuaMin Feng
· DBLP profile ↗
45ranked-venue papers
3as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MagicPaint: Operate Anything for Image Inpainting with Diffusion ModelabstractRecent diffusion-based models have significantly improved inpainting quality. However, existing methods struggle with multi-task inpainting due to conflicting optimization objectives, and current datasets are typically limited to task-specific scenarios, hindering joint training. To address these challenges, we propose MagicPaint, a unified diffusion-based inpainting model that supports object addition, removal, and unconditional inpainting across both text and image modalities. MagicPaint semantically decouples operation types and target content by learnable tokens in MMToken Module, effectively reconciling conflicting optimization objectives and enabling robust multi-task, multi-modal inpainting. Besides, a novel inpainting paradigm named MagicMask, encodes operating intent directly into the mask and applies a mask loss for spatially precise supervision. In addition, existing inpainting datasets are insufficient for multi-task and multi-modal scenarios, limiting the capability of inpainting models. Thus, we further introduce a new dataset comprising 2.1M image tuples. It is dedicatedly designed to support diverse inpainting scenarios and significantly improves upon existing datasets, particularly in object removal. Through efforts from both model and data perspectives, MagicPaint enables users to operate anything—add, remove or inpaint content which is specified through either text or image modalities in a seamless and unified manner. Extensive experiments demonstrate that MagicPaint achieves state-of-the-art performance across three key tasks (i.e., text-guided addition, image-guided addition, and object removal) and produces outputs with superior visual consistency and contextual fidelity compared to existing methods. Qinhong Yang, Dongdong Chen 0001, Qi Chu 0001, Qiankun Liu 0001, Zhentao Tan, Xulin Li, Huamin Feng, Nenghai Yu |
AAAI | 8 |
| 2026 | MGDA: A provenance graph-based framework for threat detection and attack scenario reconstructionabstractAdvanced persistent threat (APT) attacks are sophisticated, stealthy, and persistent, posing significant challenges to timely detection and investigation in modern network environments. Provenance graph analysis has become an important method for APT detection due to its ability to capture detailed causal relationships among system entities. However, existing methods suffer from several limitations: (1) lack of labeled attack data, (2) lack of high-level semantics in attack scenario reconstruction, and (3) high computational overhead limiting practical deployment. In this paper, we propose MGDA, a self-supervised method for effective and accurate threat detection as well as interpretable attack scenario reconstruction. MGDA introduces a multi-view masked graph autoencoder that jointly captures deep semantic features and structural patterns, enabling accurate detection of stealthy and unknown attacks. In the reconstruction phase, MGDA combines contextual analysis with rule-based attack pattern matching to produce attack scenario graphs that incorporate high-level semantics. We evaluate MGDA on three widely used datasets, including both real-world and simulated network attacks. The results demonstrate that MGDA achieves an average precision of 97.58% and F1-score of 98.03% in threat detection, outperforming state-of-the-art approaches. In addition, the automatically reconstructed scenario graphs help identify potential multi-step attacks and their stages, aiding analysts in conducting efficient network attack investigations. Mengjiao Cui, Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng |
Comput. Networks | 7 |
| 2026 | A Novel Privacy-Preserving user information queries scheme with functional policy
Yuhang Lei, Yang Yang 0026, Chunjie Cao, Huamin Feng |
J. Inf. Secur. Appl. | 5 |
| 2026 | FEAC: A New Construction of Fast and Expressive Anonymous Credential for Cloud ServiceabstractAnonymous credentials are an essential cryptography primitive to protect user privacy and provide fine-grained access control for proving ownership and rights of specific credentials. There are currently two roadmaps to designing anonymous credentials: one is signature credentials, which are constructed by signature with efficient protocols and non-interactive zero-knowledge proofs, and the other is functional credentials, which are transformed from predicate encryption schemes. However, none of the existing instances of anonymous credentials support$expressive$access policies expressed as conjunction, disjunction, or arbitrary Boolean formulas, which are particularly useful for cloud services. In this paper, we propose a new fast and expressive anonymous credential, called FEAC. It is constructed with the unique$dual$$randomness$$splitting$technique, which combines the most efficient anonymous key-policy attribute-based encryption (USENIX 24) and short randomizable signature (CT-RSA 18) to balance efficiency, expressiveness, and security, demonstrating a new way to instantiate anonymous credentials. Furthermore, our credential presentation protocol offloads most of the time-consuming computation to the cloud server (11 pairing) to reduce the computational burden on the user side (2 pairing). We propose formal definitions and formal security proofs of FEAC. We provide implementations and evaluate the performance of FEAC, comparing it to state-of-the-art work. Huamin Feng, Chunjie Cao, Yang Yang 0026, Baitao Zhang, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | LatInc: A Practical Lattice-Based Privacy-Preserving Incentive SystemabstractIncentive (or point) systems are widely deployed across industries such as retail, tourism, and finance to enhance customer loyalty and create benefits for service providers. However, their operation typically requires the collection and processing of sensitive customer data, leading to significant privacy concerns. Existing privacy-preserving incentive systems predominantly rely on bilinear pairings and the discrete logarithm assumption, which, while efficient in classical settings, are vulnerable to quantum adversaries and thus lack long-term security guarantees. To address this limitation, we present LatInc, a practical lattice-based privacy-preserving incentive system. LatInc integrates state-of-the-art lattice-based signatures with efficient protocols, the ABDLOP commitment, and efficient lattice zero-knowledge proofs, achieving a robust balance between post-quantum security and efficiency. Relying on the hardness of the MLWE and MSIS problems, we formally prove that LatInc achieves unforgeability, anonymity, and framing-resistance in the random oracle model. We implement a demo of the system and evaluate its performance on a standard laptop platform. Experimental results show that the communication overheads for the Earning and Spending protocols are approximately 99 KB and 140 KB, respectively, with execution times of 610 ms and 900 ms, highlighting significant efficiency gains over previous lattice-based incentive constructions. Huamin Feng, Yang Yang 0026, Zhen Guo 0003, Chunjie Cao, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | Hecate: Threshold Anonymous Credentials With Private Verifiers and Issuer-Hiding
Huamin Feng, Yang Yang 0026, Yingjiu Li, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | FlyCred: Contractual Anonymous Credentials Based on Oracles and EventsabstractIn a scenario where an issuer wishes to issue an attribute-based anonymous credential to a user, this issuance is conditional on a number of real-world outcomes. These outcomes involve multiple entrusted oracles confirming the occurrence of several events, after which the issuance can proceed successfully. Such contractual credentials can serve as an important building block for blockchain-based Web 3.0 systems and can be used in real-world applications that require privacy-preserving, prescheduled authorization. However, there is currently no work that enables the pre-issuance of credentials based on oracles and events. In this work, we propose contractual anonymous credentials, called FlyCred, to fill this gap. With FlyCred, the issuer can issue an encrypted credential to a user, controlled by a dual-layer authorization policy consisting of oracle-based and event-based expressive policies. As core building blocks, we introduce two novel cryptographic primitives: the Adaptor Anonymous Credential and ABE-based Signature Witness Encryption with Tags, which can serve as independent interests. We provide efficient instantiations of these primitives and evaluate their performance under different security levels and system parameters on a laptop, showing that the computation and communication overhead of the credential pre-issuance is less than 85.8 seconds and 8.7 MB, respectively. Yang Yang 0026, Huamin Feng, Yingjiu Li, Chunjie Cao, Robert H. Deng |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | AnoDS: A Blockchain-Based Anonymous Small Donation System for IoT DevicesabstractThis article introduces the anonymous small donation system (AnoDS), a blockchain-based AnoDS for Internet of Things (IoT) devices in industrial environments. Traditional donation models often rely on centralized charitable organizations, which can cause delays and reduce transparency. AnoDS aims to address these issues by leveraging blockchain's decentralized and tamper-resistant nature, allowing direct interaction between donors and donees via IoT devices. The system integrates seamlessly with industrial IoT networks, incorporating digital identity management and efficient pointcheval-sanders signature (PS) signatures to ensure real-time tracking of donations while preserving donor anonymity. We implement the core cryptographic components of AnoDS and conduct micro-benchmarks of the underlying cryptographic operations to estimate its computational and communication overhead. The results indicate that AnoDS achieves high efficiency and low communication overhead under realistic parameter settings, making it suitable for deployment in industrial IoT scenarios. In addition, we formalize the system and threat models, analyze the security properties under clearly stated trust assumptions, and compare AnoDS with representative donation schemes in terms of functionality and efficiency. This system offers a secure, transparent, and efficient solution for anonymous donations in IoT-driven industrial environments. Xiqin Yuan, Huamin Feng, Yani Sun |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Cross Page Recognition Methods for Encrypted Web Application FingerprintingabstractThe widespread implementation of the HTTPS protocol has greatly bolstered user privacy and data security. However, the widespread use of the HTTPS protocol has also provided criminals with a cloak to disseminate harmful content through websites, thereby undermining the integrity of the online environment. Web page fingerprinting has emerged as a highly popular method for web application identification. Yet, due to the frequent updates of web applications, existing methods struggle with accuracy issues. To tackle these challenges, this paper introduces a novel approach called CrossWP, which leverages cross-web page fingerprinting to enhance the security of the network environment. CrossWP aims at classifying web applications, which novelly constructs the cross web pages behavior sequences on handshake, request and response sequence. CrossWP uses the transformer model based on multi-sequence fusion to thoroughly learn and integrate these unique spatio-temporal sequence characteristics. This model can also capture the internal similarity of behavior sequence, and achieving high accuracy. The effectiveness of CrossWP is validated through closed-world and open-world evaluations, which involve identifying and classifying news websites, social websites, online video websites and e-commerce websites. The results indicate that CrossWP outperforms existing algorithms in terms of both robustness and accuracy. Zelin Cui, Pu Dong, Dongxu Han, Bo Jiang 0013, Zhigang Lu 0002, Huamin Feng |
CSCWD | 6 |
| 2025 | AGLHunter: Automated Threat Hunting Using In-Context Learning-Enhanced LLMabstractAdvanced Persistent Threats (APTs) are characterized by their persistence, sophistication, and stealth, posing significant challenges to network detection. Existing research on attack detection leveraging Provenance Graphs (PGs) has proven effective in correlating system entities and capturing persistence. However, the exponential growth of audit logs makes large-scale data storage and processing difficult. In addition, current threat hunting methods rely heavily on manually crafted attack query graphs, which are limited by expert knowledge and lack automated solutions. In this paper, we propose AGLHunter, an automated threat hunting system designed to enhance automation and efficiency while maintaining high detection accuracy. Our system leverages the In-Context Learning (ICL) capability of the Large Language Model (LLM) to automatically construct query graphs from Cyber Threat Intelligence (CTI) reports. Next, we extract suspicious subgraphs from PGs and employ graph representation learning to match these sub graphs with the query graphs, enabling efficient and accurate threat hunting. We use DARPA TC and OpTC datasets to evaluate AGLHunter's performance. The results show that AGLHunter not only achieves higher automation but also shows superior performance with reduced memory usage. AGLHunter, leveraging ICL-enhanced LLM, improved the F1 score for query graph construction by 13.6%, reduced the overall hunting time by more than 170 seconds, and maintained high detection accuracy. Mengjiao Cui, Zhengwei Jiang, Yepeng Yao, Qiying He, Peian Yang, Huamin Feng |
CSCWD | 7 |
| 2025 | Dynamic Behavior-Based Detection Techniques for Encrypted Variant WebshellsabstractWebshell, as a common type of malicious script, is frequently utilized by cyber attackers who execute unauthorized commands on the victim's server to carry out attacks. Strengthening research on Webshell detection techniques is crucial for building a robust cybersecurity defense. Despite significant progress in the field of Webshell detection, these techniques still face numerous challenges. Firstly, the continuous evolution of attacker techniques has enhanced the adversarial capabilities of Webshell, including rapid updates to version variants, as well as the use of advanced obfuscation techniques. Secondly, the use of HTTPS has grown dramatically from 40% in 2014 to 98% in 2023, rendering techniques based on plaintext rules ineffective for detecting Webshell. These technological updates make it difficult for traditional detection methods to effectively identify and defend against Webshell variants based on the HTTPS protocol. To address this problems, we propose a novel dynamic behavior-based detection techniques called DBBDdetect, aiming at detecting encrypted variant Webshell in order to protect critical infrastructure. DBBDdetect delves into the interaction process between the Webshell and the server, extracting three types of feature information. It utilizes CNNs to obtain vector features of the traffic payload and concat statistical features and similar sequential byte behavior features. Then, it uses DBSCAN clustering model to analyze behavioral similarities to detect variant Webshell attacks. This method captures the intrinsic similarities in behavior, and experiments have shown that it achieves a high level of accuracy. Zelin Cui, Pu Dong, Mengchuan Shang, Bo Jiang 0013, Zhigang Lu 0002, Huamin Feng |
CSCWD | 7 |
| 2025 | CotexFinger: Enhancing IoT Device Identification with Context-Packet Fingerprinting and Lightweight MambaabstractWith the growth of the Internet of Things (IoT) market, IoT devices have become major targets for attacks. For network administrators, effectively addressing these attacks requires accurately identifying IoT devices. Due to the rich communication protocols and complex network environments of IoT, current fingerprinting methods lack robustness. This paper proposes a fingerprinting method based on Context-Packet, which effectively reduces the impact of traffic noise and fingerprint redundancy by focusing only on session packets that carry payload data. Additionally, a lightweight Efficient VMamba model is trained to extract fingerprints from the packets. We design a method that utilizes the Context-Packet length sequences to classify traffic samples from the same device category, while retaining rare traffic samples during the random sampling process. Experiments show that our method outperforms existing approaches in classification performance, improving the F1 score to 97.08% ( 0.60% ↑ ), 98.07% ( 2.7% ↑ ), 95.55% ( 1.27% ↓ ), and 92.22% ( 3.71% ↑ ) in the AA, II, AI, and IA experiments, respectively. Moreover, it offers faster inference speed and lower resource overhead. Zelin Cui, Bo Jiang 0013, Zhigang Lu 0002, Huamin Feng |
IJCNN | 6 |
| 2025 | Automated Attack Graph Construction for Cross-host Threat Detection Using Cyber Threat IntelligenceabstractCyber Threat Intelligence (CTI) reports provide valuable insights into cyber threats. However, manually constructing attack graphs from the unstructured CTI reports requires significant human effort. With the development of Large Language Models (LLMs), researchers have begun to harness LLMs for attack graph construction from CTI reports. Nevertheless, existing works mainly focus on describing abstract and high-level attack behaviors through these graphs, which cannot be used as query graphs for threat detection based on graph matching, as provenance graphs are at the system level. Moreover, these works do not consider modeling cross-host attack behaviors. To address these problems, we propose a novel method for automatically constructing attack graphs from CTI reports. We utilize prompt engineering, and leverage the in-context learning ability of LLMs to generate attack graphs. In this method, we first restructure the CTI report by grouping continuous sentences with the same tactics, and then use multiple LLMs to extract entities and relations. Finally, we use an LLM to integrate all the results. We also design a cross-host threat detection algorithm using the generated attack graphs. The evaluation results show that our method constructs attack graphs with an average conversion rate of 72.4%. It also achieves nearly 94.8% precision and 96.5% recall for IoC and relation extraction compared with manually labeled results. Ziqing Feng, Qiuyun Wang, Liling Xin, Zhengwei Jiang, Huamin Feng |
TrustCom | 8 |
| 2025 | TIMFuser: A multi-granular fusion framework for cyber threat intelligence
Zhengwei Jiang, Kai Zhang 0035, Zhiting Ling, Yizhe You, Peian Yang, Huamin Feng |
Comput. Secur. | 8 |
| 2025 | Insider threat detection for specific threat scenariosabstractAbstract Insider threats pose significant challenges to network security due to their destructive and covert nature, often resulting in substantial losses for enterprises. Traditional methods mainly analyze user behavior patterns or convert behaviors into time sequences for further analysis. However, existing detection methods primarily focus on identifying abnormal users or behaviors, lacking the capability to pinpoint specific threats. Additionally, these methods struggle to accurately identify long-distance dependencies in behavior sequences, frequently increasing false positives. To address these issues, we introduce a scenario-oriented insider threat detection model. This model targets three specific threat scenarios-privilege abuse, identity theft, and data leakage-by analyzing user behavior patterns, extracting detailed behavioral characteristics, and constructing behavior sequences. Firstly, this paper serializes user behavior daily and vectorizes it using one-hot encoding. Then, it introduces contextual characteristic information and reconstructs the background of abnormal behavior through behavior vectorization, providing a comprehensive description of user behavior characteristics. This approach addresses the issue of behavior isolation, thereby improving the accuracy and robustness of anomaly detection. Subsequently, a time series analysis model based on a multi-head attention mechanism is employed to analyze long-distance dependencies in behavior sequences. The multi-head attention mechanism simultaneously attends to multiple positions in the behavior sequence, capturing potential correlations between behaviors and user behavior patterns. This mechanism can analyze local information and obtain long-distance dependencies, providing depth feature representation for anomaly detection. Ultimately, we achieve the goal of classifying abnormal behavior sequences. We conduct comprehensive tests on the CERT dataset, demonstrating that our method outperforms traditional deep learning approaches (LSTM, GNN, and GCN) in detecting abnormal sequences. Compared to the best results among the baseline methods, it shows an improvement in accuracy of approximately 2% for privilege abuse, 5% for identity theft, and 2% for data leakage. Bo Jiang 0013, Huamin Feng, Zhigang Lu 0002 |
Cybersecur. | 4 |
| 2024 | TiGNet: Joint entity and relation triplets extraction for APT campaign threat intelligenceabstractContemporary cybersecurity faces escalating challenges from sophisticated threats, notably Advanced Persistent Threats (APTs). Addressing these challenges necessitates a collaborative, multidisciplinary approach that transcends traditional boundaries. Gathering cyber threat intelligence (CTI) on APT campaigns and constructing a comprehensive knowledge graph empowers defenders to track the latest trends in these campaigns, update defense strategies, and attain crucial advantages in defense measures. Previous works used relation extraction techniques to obtain entity-relation triplets for constructing threat intelligence knowledge graphs. However, these works either rely on pipeline workflow, are susceptible to exposure errors and error propagation, or use sequence annotation method, which combines entity and relation labels but lacks the ability to extract single entity overlaps (SEO) or subject-object overlaps (SOO) triplets. This paper introduces TiGNet, a novel method that transforms the entity-relation triplet’s extraction task into multiple token-span recognition tasks utilizing token-pair matrices. Additionally, we integrate GlobalPointer to incorporate token position information into the token-pair matrix, significantly enhancing extraction performance. To facilitate method evaluation, we annotated a Chinese entity-relation triplets dataset about APT campaigns, named APT-Triplets, comprising 9711 triplets encompassing seven triplet types. Our evaluation demonstrates that TiGNet improves the F1 score of 3.79-5.59 compared to previous joint extraction methods. Furthermore, it outperforms methods based on large language models (LLMs) in terms of both extraction performance and inference time. These results underscore TiGNet’s capacity to accurately and swiftly extract threat intelligence, facilitating the construction of the APT campaign knowledge graphs, empowering defenders to track evolving trends and fortify defense strategies collaboratively. Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Huamin Feng, Peian Yang |
CSCWD | 4 |
| 2024 | CTIMiner: Cyber Threat Intelligence Mining Using Adaptive Multi-task Adversarial Active Learning
Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng |
ICDF2C (1) | 7 |
| 2024 | MAD-LLM: A Novel Approach for Alert-Based Multi-stage Attack Detection via LLMabstractIn the realm of cybersecurity, detecting multi-stage attacks is vital for uncovering the actual intentions and strategies of attackers. However, the detection of multi-stage attacks is fraught with challenges due to the proliferation of alerts from diverse sources, the heterogeneity of their formats, and the weak correlations among them. This study explores MAD-LLM, a novel approach for alert-based multi-stage attack detection using Large Language Model (LLM). Leveraging their advanced natural language processing capabilities, LLM demonstrates significant advantages in text comprehension and pattern recognition. This research attempts to aggregate and correlate security alerts using the prompt engineering capabilities of LLM to reconstruct multi-stage attack chain. Experimental results indicate that LLM exhibit excellent performance in the task of multi-stage attack detection, providing an innovative solution for cybersecurity defense. This study also offers insights and implications for the application of LLM in other fields. Dan Du, Xingmao Guan, Bo Jiang 0013, Huamin Feng |
ISPA | 6 |
| 2024 | CTIFuser: Cyber Threat Intelligence Fusion via Unsupervised Learning ModelabstractCyber attack campaigns are becoming increasingly complex and severe, causing significant impacts on institutions and individuals. Cyber Threat Intelligence (CTI) provides important evidential knowledge about attackers and is critical to the shift from reactive to proactive defense against cyber attacks. Attack detection based on Indicators of Compromise (IOCs), a type of CTI, is vulnerable to the limitation of insufficient context of attack scenarios. In contrast, attack behavior intelligence is associated with information on attackers’ techniques, targets, and intentions, providing a solid foundation for security practitioners to conduct attack investigations or other applications. Many current CTI mining systems are limited to extracting CTI from a single source, leading to challenges such as fragmented attack behavior view and low-value density. To address these issues, we propose an unsupervised fusion framework named CTIFuser, which includes a comprehensive pipeline of four subtasks aimed at mining and fusing multi-source attack behaviors at the attack technique level. In our evaluation of 739 real-world CTI reports from 542 sources, experimental results demonstrate that CTIFuser can obtain a complete view of the attack behaviors at the attack technique level. Zhengwei Jiang, Peian Yang, Mengjiao Cui, Fangming Dong, Huamin Feng |
ISPA | 7 |
| 2024 | LSTM-Diff: A Data Generation Method for Imbalanced Insider Threat Detection
Bo Jiang 0013, Huamin Feng, Zhigang Lu 0002 |
TrustCom | 5 |
| 2024 | Graph Anomaly Detection with Bi-level OptimizationabstractGraph anomaly detection (GAD) has various applications in finance, healthcare, and security. Graph Neural Networks (GNNs) are now the primary method for GAD, treating it as a task of semi-supervised node classification (normal vs. anomalous). However, most traditional GNNs aggregate and average embeddings from all neighbors, without considering their labels, which can hinder detecting actual anomalies. To address this issue, previous methods try to selectively aggregate neighbors. However, the same selection strategy is applied regardless of normal and anomalous classes, which does not fully solve this issue. This study discovers that nodes with different classes yet similar neighbor label distributions (NLD) tend to have opposing loss curves, which we term it as "loss rivalry". By introducing Contextual Stochastic Block Model (CSBM) and defining NLD distance, we explain this phenomenon theoretically and propose a Bi-level optimization Graph Neural Network (BioGNN), based on these observations. In a nutshell, the lower level of BioGNN segregates nodes based on their classes and NLD, while the upper level trains the anomaly detector using separation outcomes. Our experiments demonstrate that BioGNN outperforms state-of-the-art methods on four benchmarks and effectively mitigates "loss rivalry". Yuan Gao 0020, Junfeng Fang, Yongduo Sui, Xiang Wang 0010, Huamin Feng, Yongdong Zhang 0001 |
WWW | 6 |
| 2024 | P-TIMA: a framework of T witter threat intelligence mining and analysis based on a prompt-learning NER modelabstractAbstract Open-source information platforms such as Twitter continuously provide the latest threat intelligence, including new vulnerabilities and in-the-wild exploitations of advanced persistent threat (APT) groups. Automated extraction of threat intelligence from Twitter has become crucial for defenders to access up-to-date threat knowledge. However, existing studies mainly rely on supervised learning methods to extract threat intelligence knowledge, such as entities, which require a large amount of annotated data. This paper presents Threat Intelligence Mining and Analysis based on Prompt Learning (P-TIMA), a framework specifically crafted for extracting and analyzing threat intelligence from Twitter. P-TIMA employs our innovative few-shot entity recognition method, SecEntPrompt (SEP), built on prompt learning, to extract vulnerability intelligence from Twitter. Additionally, P-TIMA analyzes and profiles the overarching vulnerability intelligence obtained from Twitter, along with in-the-wild exploitation intelligence of APT groups. The SEP improves the average entity recognition F1 score by 3.62-4.40 compared with the best-performing comparison model and outperforms the method based on the large language model on recognition performance and inference time. To validate our framework, we apply P-TIMA to extract vulnerability-related threat intelligence from real Twitter data. Through case studies, we then analyze trends in vulnerability threats and the exploitation capabilities of APT groups. In conclusion, our framework provides a more efficient and accurate method for extracting threat intelligence from Twitter, enabling defenders to stay up-to-date with the latest threat trends and helping them improve their defense strategies against cyber attacks. Yizhe You, Zhengwei Jiang, Peian Yang, Kai Zhang 0035, Xuren Wang, Chenpeng Tu, Huamin Feng |
Comput. J. | 8 |
| 2024 | AnoPas: Practical anonymous transit pass from group signatures with time-bound keys
Yang Yang 0026, Yingjiu Li, Huamin Feng, HweeHwa Pang, Robert H. Deng |
J. Syst. Archit. | 4 |
| 2024 | Robust Model Watermarking for Image Processing Networks via Structure ConsistencyabstractThe intellectual property of deep networks can be easily "stolen" by surrogate model attack. There has been significant progress in protecting the model IP in classification tasks. However, little attention has been devoted to the protection of image processing models. By utilizing consistent invisible spatial watermarks, the work (Zhang et al. 2020) first considered model watermarking for deep image processing networks and demonstrated its efficacy in many downstream tasks. Its success depends on the hypothesis that if a consistent watermark exists in all prediction outputs, that watermark will be learned into the attacker's surrogate model. However, when the attacker uses common data augmentation attacks (e.g., rotate, crop, and resize) during surrogate model training, it will fail because the underlying watermark consistency is destroyed. To mitigate this issue, we propose a new watermarking methodology, "structure consistency", based on which a new deep structure-aligned model watermarking algorithm is designed. Specifically, the embedded watermarks are designed to be aligned with physically consistent image structures, such as edges or semantic regions. Experiments demonstrate that our method is more robust than the baseline in resisting data augmentation attacks. Besides that, we test the generalization ability and robustness of our method to a broader range of adaptive attacks. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Zehua Ma, Han Fang 0004, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Double Issuer-Hiding Attribute-Based Credentials From Tag-Based Aggregatable Mercurial SignaturesabstractAttribute-based anonymous credentials offer users fine-grained access control in a privacy-preserving manner. However, in such schemes obtaining a user's credentials requires knowledge of the issuer's public key, which obviously reveals the issuer's identity that must be hidden from users in certain scenarios. Moreover, verifying a user's credentials also requires the knowledge of issuer's public key, which may infer the user's private information from their choice of issuer. In this paper, we introduce the notion of double issuer-hiding attribute-based credentials (${\sf DIHAC}$) to tackle these two problems. In our model, a central authority can issue public-key credentials for a group of issuers, and users can obtain attribute-based credentials from one of the issuers without knowing which one it is. Then, a user can prove that their credential was issued by one of the authenticated issuers without revealing which one to a verifier. We provide a generic construction, as well as a concrete instantiation for${\sf DIHAC}$based on structure-preserving signatures on equivalence classes (JOC's 19) and a novel primitive which we calltag-based aggregatable mercurial signatures. Our construction is efficient without relying on zero-knowledge proofs. We provide rigorous evaluations on personal laptop and smartphone platforms, respectively, to demonstrate its practicability. Yang Yang 0026, Yingjiu Li, Huamin Feng, Guozhen Shi, HweeHwa Pang, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Revisiting Attack-Caused Structural Distribution Shift in Graph Anomaly DetectionabstractGraph anomaly detection (GAD) under semi-supervised setting poses a significant challenge due to the distinct structural distribution between anomalous and normal nodes. Specifically, anomalous nodes constitute a minority and exhibit high heterophily and low homophily compared to normal nodes, which makes the distribution of neighbors of the two types of nodes close, that is, most of them are composed of normal nodes, which causes the two types of nodes to be difficult to distinguish during the aggregation process. Furthermore, we discover that apart from various time factors and annotation preferences, graph adversarial attacks can lead to and amplify the heterophily difference across training and testing data, which is called structural distribution shift (SDS) in this paper. Current mainstream methods for GAD tend to overlook the SDS problem, resulting in poor generalization performance and limited effectiveness in detecting anomalies. This work solves the problem from a feature view. We observe that the degree of SDS varies between anomalies and normal nodes. Hence to address the issue, the key lies in resisting high heterophily for anomalies meanwhile benefiting the learning of normals from homophily. Since different labels correspond to the difference of critical anomaly features which make great contributions to the GAD, we tease out the anomaly features on which we constrain to mitigate the effect of heterophilous neighbors and make them invariant. However, the prior distribution of anomaly features is dynamic and hard to estimate, we thus devise a prototype vector to infer and update this distribution during training. For normal nodes, we constrain the remaining features to preserve the connectivity of nodes and reinforce the influence of the homophilous neighborhood. We term our proposed framework asGraphDecompositionNetwork(GDN). To demonstrate the effectiveness of the network, we explain the process of feature decomposition in the spectral domain. Extensive experiments are conducted on four benchmark datasets, including two additional datasets and two used in the preliminary work. To further validate our performance under SDS, we conduct an adversarial attack to incur different heterophily degrees for the training set and the test set. The proposed framework achieves remarkable accuracy and robustness boost in GAD, especially in an SDS environment where anomalies have largely different structural distribution across training and testing environments. Our code is open-sourced inhttps://github.com/fortunato-all/skl-GDN. Yuan Gao 0020, Jinghan Li, Xiang Wang 0010, Xiangnan He 0001, Huamin Feng, Yongdong Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | FineCTI: A Framework for Mining Fine-grained Cyber Threat Information from Twitter Using NER ModelabstractTo timely respond to cyber threats related to a specific IT infrastructure called fine-grained (e.g., Windows or Linux), security analysts need to require timely and comprehensive threat information. Twitter, as a vital source of real-time threat information, provides abundant but overwhelming information due to the increased data sources. Automatically mining and summarizing fine-grained threat information from Twitter can help security analysts maintain the infrastructure’s security. Most existing studies focus on classification, which carries less threat information. Some works use clustering based on text similarity relying on the embedding of text obtained from pre-trained models, which cannot be applied to short text, resulting in noisy clusters. Several works build topic models. However, the incoherent topic keywords are difficult to understand and analyze. To overcome these challenges, we design a FineCTI framework to mine the threat information related to the specific infrastructure on Twitter and generate a detailed threat information summary that is machine-readable and human-readable, efficiently reducing information overload. FineCTI optimizes the feature extraction part based on the named entity recognition model and performs clustering based on features extracted, thus effectively reducing the influence of sparsity of tweets on the clustering result and with the V-measure score improved by 7%. The cluster analysis results show that we can mine the fine-grained threats up to 15 days before the official disclosure date. Kai Zhang 0035, Zhengwei Jiang, Peian Yang, Xuren Wang, Huamin Feng |
TrustCom | 7 |
| 2023 | Insider Threat Detection Based On Heterogeneous Graph Neural NetworkabstractAs one of the most challenging threats in cyberspace, insider threats frequently lead to substantial losses for enterprises. Recently, there are many studies focus on user behavior analysis for insider threats detection. However, they ignore the underlying causes of insider threats and the implicit relationships between users, which is more critical for discover the insider threats. To address this gap, we propose the novel ITDE model in this paper, which applies a graph neural network approach based on two-layer attention. The core idea is to abstracting user features and potential relationships as heterogeneous graphs based on an analysis of user behavior and the causes of insider threats. Futhermore, we employ node-level attention and semantic-level attention to capture the complex graph structure information and generate node embedding by aggregating features from meta-path based neighbors. Finally, we use a cross-entropy loss function to implement insider threat detection. We verify the effectiveness of our model on the CERT r4.2 dataset and it outperforms state-of-the-art methods in insider threat detection. Yiru Gong, Bo Jiang 0013, Huamin Feng, Zhigang Lu 0002 |
TrustCom | 5 |
| 2023 | Alleviating Structural Distribution Shift in Graph Anomaly DetectionabstractGraph anomaly detection (GAD) is a challenging binary classification problem due to its different structural distribution between anomalies and normal nodes --- abnormal nodes are a minority, therefore holding high heterophily and low homophily compared to normal nodes. Furthermore, due to various time factors and the annotation preferences of human experts, the heterophily and homophily can change across training and testing data, which is called structural distribution shift (SDS) in this paper. The mainstream methods are built on graph neural networks (GNNs), benefiting the classification of normals from aggregating homophilous neighbors, yet ignoring the SDS issue for anomalies and suffering from poor generalization. Yuan Gao 0020, Xiang Wang 0010, Xiangnan He 0001, Zhenguang Liu, Huamin Feng, Yongdong Zhang 0001 |
WSDM | 5 |
| 2023 | Addressing Heterophily in Graph Anomaly Detection: A Perspective of Graph SpectrumabstractGraph anomaly detection (GAD) suffers from heterophily — abnormal nodes are sparse so that they are connected to vast normal nodes. The current solutions upon Graph Neural Networks (GNNs) blindly smooth the representation of neiboring nodes, thus undermining the discriminative information of the anomalies. To alleviate the issue, recent studies identify and discard inter-class edges through estimating and comparing the node-level representation similarity. However, the representation of a single node can be misleading when the prediction error is high, thus hindering the performance of the edge indicator. Yuan Gao 0020, Xiang Wang 0010, Xiangnan He 0001, Zhenguang Liu, Huamin Feng, Yongdong Zhang 0001 |
WWW | 5 |
| 2023 | SED-SGC: A scalable, efficient, and distributed secure group communication scheme based on superlattice PUF
Jianguo Xie, Huamin Feng |
Comput. Networks | 6 |
| 2023 | Rumor detection with self-supervised learning on texts and social graph
Yuan Gao 0020, Xiang Wang 0010, Xiangnan He 0001, Huamin Feng, Yongdong Zhang 0001 |
Frontiers Comput. Sci. | 4 |
| 2023 | PriRPT: Practical blockchain-based privacy-preserving reporting system with rewards
Yang Yang 0026, Huamin Feng, Huiqin Xie |
J. Syst. Archit. | 3 |
| 2023 | Non-transferable blockchain-based identity authentication
Yuxia Fu, Jun Shao 0001, Qingjia Huang, Qihang Zhou, Huamin Feng, Xiaoqi Jia, Ruiyi Wang, Wenzhi Feng |
Peer Peer Netw. Appl. | 5 |
| 2023 | Threshold Attribute-Based Credentials With Redactable SignatureabstractThreshold attribute-based credentials are suitable for decentralized systems such as blockchains as such systems generally assume that authenticity, confidentiality, and availability can still be guaranteed in the presence of a threshold number of dishonest or faulty nodes. Coconut (NDSS’19) was the first selective disclosure attribute-based credentials scheme supporting threshold issuance. However, it does not support threshold tracing of user identities and threshold revocation of user credentials, which is desired for internal governance such as identity management, data auditing, and accountability. The communication and computation complexities of Coconut for verifying credentials are linear in the number of each user's attributes and thus costly. Addressing these issues, we propose a novel efficient threshold attribute-based anonymous credential scheme. While retaining all the features of Coconut, our scheme supports threshold tracing of user identities and threshold revocation of user credentials, and it significantly reduces the computational and communication complexities of credential verification. In addition, we prove that our scheme enjoys strong security features, including anonymity, blindness, traceability, and non-frameability. Huamin Feng, Yang Yang 0026, Yingjiu Li, HweeHwa Pang, Robert H. Deng |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | TI-Prompt: Towards a Prompt Tuning Method for Few-shot Threat Intelligence Twitter Classification*abstractObtaining the latest Threat Intelligence (TI) via Twitter has become one of the most important methods for defenders to catch up with emerging cyber threats. Existing TI Twitter classification works mainly based on supervised learning methods. Such approaches require large amounts of annotated data and are difficult to be transferred to other TI Twitter classification tasks. This paper proposes a prompt-based method for classifying TI on Twitter, named TI-Prompt. TI-Prompt lever-ages the prompt-tuning method with two templates in different TI Twitter classification tasks. TI-Prompt also uses a semantic similarity-based approach to automatically enrich the prompt verbalizer without expert knowledge and a verbalizer refinement method to calibrate the verbalizer based on the training data. We evaluate TI-Prompt with binary and multi-classification tasks on two Twitter Threat Intelligence datasets. Evaluation results show that the proposed TI-Prompt improves 5-10% over the best performance of previous supervised learning methods under the few-shot settings. Compared to the general prompt-tuning methods, the proposed prompt-tuning templates can also improve the classification performance by 2–5%. Meanwhile, the proposed verbalizer enrichment method and refinement method improve classification accuracy by 1–4% compared with the general single-word verbalizer prompt method. Therefore, TI-Prompt can be extended to other Threat Intelligence classification tasks without requiring large amounts of training data, significantly reducing the annotation cost. Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Xuren Wang, Shirui Wang, Huamin Feng |
COMPSAC | 8 |
| 2022 | TIM: threat context-enhanced TTP intelligence mining on unstructured threat dataabstractAbstract TTPs (Tactics, Techniques, and Procedures), which represent an attacker’s goals and methods, are the long period and essential feature of the attacker. Defenders can use TTP intelligence to perform the penetration test and compensate for defense deficiency. However, most TTP intelligence is described in unstructured threat data, such as APT analysis reports. Manually converting natural language TTPs descriptions to standard TTP names, such as ATT&CK TTP names and IDs, is time-consuming and requires deep expertise. In this paper, we define the TTP classification task as a sentence classification task. We annotate a new sentence-level TTP dataset with 6 categories and 6061 TTP descriptions from 10761 security analysis reports. We construct a threat context-enhanced TTP intelligence mining (TIM) framework to mine TTP intelligence from unstructured threat data. The TIM framework uses TCENet (Threat Context Enhanced Network) to find and classify TTP descriptions, which we define as three continuous sentences, from textual data. Meanwhile, we use the element features of TTP in the descriptions to enhance the TTPs classification accuracy of TCENet. The evaluation result shows that the average classification accuracy of our proposed method on the 6 TTP categories reaches 0.941. The evaluation results also show that adding TTP element features can improve our classification accuracy compared to using only text features. TCENet also achieved the best results compared to the previous document-level TTP classification works and other popular text classification methods, even in the case of few-shot training samples. Finally, the TIM framework organizes TTP descriptions and TTP elements into STIX 2.1 format as final TTP intelligence for sharing the long-period and essential attack behavior characteristics of attackers. In addition, we transform TTP intelligence into sigma detection rules for attack behavior detection. Such TTP intelligence and rules can help defenders deploy long-term effective threat detection and perform more realistic attack simulations to strengthen defense. Yizhe You, Zhengwei Jiang, Peian Yang, Baoxu Liu, Huamin Feng, Xuren Wang |
Cybersecur. | 6 |
| 2022 | Deep Model Intellectual Property Protection via Deep WatermarkingabstractDespite the tremendous success, deep neural networks are exposed to serious IP infringement risks. Given a target deep model, if the attacker knows its full information, it can be easily stolen by fine-tuning. Even if only its output is accessible, a surrogate model can be trained through student-teacher learning by generating many input-output training pairs. Therefore, deep model IP protection is important and necessary. However, it is still seriously under-researched. In this work, we propose a new model watermarking framework for protecting deep networks trained for low-level computer vision or image processing tasks. Specifically, a special task-agnostic barrier is added after the target model, which embeds a unified and invisible watermark into its outputs. When the attacker trains one surrogate model by using the input-output pairs of the barrier target model, the hidden watermark will be learned and extracted afterwards. To enable watermarks from binary bits to high-resolution images, a deep invisible watermarking mechanism is designed. By jointly training the target model and watermark embedding, the extra barrier can even be absorbed into the target model. Through extensive experiments, we demonstrate the robustness of the proposed framework, which can resist attacks with different network structures and objective functions. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Poison Ink: Robust and Invisible Backdoor AttackabstractRecent research shows deep neural networks are vulnerable to different types of attacks, such as adversarial attacks, data poisoning attacks, and backdoor attacks. Among them, backdoor attacks are the most cunning and can occur in almost every stage of the deep learning pipeline. Backdoor attacks have attracted lots of interest from both academia and industry. However, most existing backdoor attack methods are visible or fragile to some effortless pre-processing such as common data transformations. To address these limitations, we propose a robust and invisible backdoor attack called "Poison Ink". Concretely, we first leverage the image structures as target poisoning areas and fill them with poison ink (information) to generate the trigger pattern. As the image structure can keep its semantic meaning during the data transformation, such a trigger pattern is inherently robust to data transformations. Then we leverage a deep injection network to embed such input-aware trigger pattern into the cover image to achieve stealthiness. Compared to existing popular backdoor attack methods, Poison Ink outperforms both in stealthiness and robustness. Through extensive experiments, we demonstrate that Poison Ink is not only general to different datasets and network architectures but also flexible for different attack scenarios. Besides, it also has very strong resistance against many state-of-the-art defense techniques. Jie Zhang 0073, Dongdong Chen 0001, Qidong Huang, Jing Liao 0001, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Image Process. | 6 |
| 2021 | Temporal ROI Align for Video Object RecognitionabstractVideo object detection is challenging in the presence of appearance deterioration in certain video frames. Therefore, it is a natural choice to aggregate temporal information from other frames of the same video into the current frame. However, ROI Align, as one of the most core procedures of video detectors, still remains extracting features from a single-frame feature map for proposals, making the extracted ROI features lack temporal information from videos. In this work, considering the features of the same object instance are highly similar among frames in a video, a novel Temporal ROI Align operator is proposed to extract features from other frames feature maps for current frame proposals by utilizing feature similarity. The proposed Temporal ROI Align operator can extract temporal information from the entire video for proposals. We integrate it into single-frame video detectors and other state-of-the-art video detectors, and conduct quantitative experiments to demonstrate that the proposed Temporal ROI Align operator can consistently and significantly boost the performance. Besides, the proposed Temporal ROI Align can also be applied into video instance segmentation. Kai Chen 0026, Xinjiang Wang, Qi Chu 0001, Feng Zhu 0006, Dahua Lin, Nenghai Yu, Huamin Feng |
AAAI | 8 |
| 2020 | Reversible Watermarking in Deep Convolutional Neural Networks for Integrity AuthenticationabstractDeep convolutional neural networks have made outstanding contributions in many fields such as computer vision in the past few years and many researchers published well-trained network for downloading. But recent studies have shown serious concerns about integrity due to model-reuse attacks and backdoor attacks. In order to protect these open-source networks, many algorithms have been proposed such as watermarking. However, these existing algorithms modify the contents of the network permanently and are not suitable for integrity authentication. In this paper, we propose a reversible watermarking algorithm for integrity authentication. Specifically, we present the reversible watermarking problem of deep convolutional neural networks and utilize the pruning theory of model compression technology to construct a host sequence used for embedding watermarking information by histogram shift. As shown in the experiments, the influence of embedding reversible watermarking on the classification performance is less than ±0.5% and the parameters of the model can be fully recovered after extracting the watermarking. At the same time, the integrity of the model can be verified by applying the reversible watermarking: if the model is modified illegally, the authentication information generated by original model will be absolutely different from the extracted watermarking information. Xiquan Guan, Huamin Feng, Weiming Zhang 0001, Hang Zhou 0007, Jie Zhang 0073, Nenghai Yu |
ACM Multimedia | 2 |
| 2004 | A bootstrapping framework for annotating and retrieving WWW imagesabstractMost current image retrieval systems and commercial search engines use mainly text annotations to index and retrieve WWW images. This research explores the use of machine learning approaches to automatically annotate WWW images based on a predefined list of concepts by fusing evidences from image contents and their associated HTML text. One major practical limitation of employing supervised machine learning approaches is that for effective learning, a large set of labeled training samples is needed. This is tedious and severely impedes the practical development of effective search techniques for WWW images, which are dynamic and fast-changing. As web-based images possess both intrinsic visual contents and text annotations, they provide a strong basis to bootstrap the learning process by adopting a co-training approach involving classifiers based on two orthogonal set of features -- visual and text. The idea of co-training is to start from a small set of labeled training samples, and successively annotate a larger set of unlabeled samples using the two orthogonal classifiers. We carry out experiments using a set of over 5,000 images acquired from the Web. We explore the use of different combinations of HTML text and visual representations. We find that our bootstrapping approach can achieve a performance comparable to that of the supervised learning approach with an F1 measure of over 54%. At the same time, it offers the added advantage of requiring only a small initial set of training samples. Huamin Feng, Tat-Seng Chua |
ACM Multimedia | 1 |
| 2004 | A Learning-based Approach for Annotating Large On-Line Image CollectionabstractSeveral recent works attempt to automatically annotate image collection by exploiting the links between visual information provided by segmented image features and semantic concepts provided by associated text. The main limitation of such approaches, however, is that semantically meaningful segmentation is in general unavailable. This paper proposes a novel statistical learning-based approach to overcome this problem. We employ two different segmentation methods to segment the image into two sets of regions and learn the association between each set of regions with text concepts. Given a new image, the idea is to first employ a greedy strategy to annotate the image with concepts derived from different sets of overlapping and possibly conflicting regions. We then incorporate a decision model to disambiguate the concepts learned using the visual features of the overlapping regions. Experiments on a mid-sized image collection demonstrate that the use of our disambiguation approach could improve the performance of the system by about 12-16% on average in terms of F/sub 1/ measures as compared to system that uses only one segmentation method. Huamin Feng, Tat-Seng Chua |
MMM | 1 |
| 2003 | An unified framework for shot boundary detection via active learningabstractVideo shot boundary detection is an important step in many video processing applications. We observe that video shot boundary is a multi-resolution edge phenomenon in the feature space. In this paper, we expanded our previous temporal multi-resolution analysis (TMRA) work by introducing the new feature vector based on motion. Further we employ the support vector machine (SVM) to refine the classification of shot boundaries. The resulting framework has been tested on the MPEG 7 video data set, and has been shown to have good accuracy for both the detection of abrupt and gradual transitions as well as their boundaries. It also has good noise tolerance characteristics. Tat-Seng Chua, Huamin Feng, A. Chandrashekhara |
ICASSP (2) | 2 |
| 2003 | ATMRA: An Automatic Temporal Multi-resolution Analysis Framework for Shot Boundary Detection
Huamin Feng, A. Chandrashekhara, Tat-Seng Chua |
MMM | 1 |