EDBT 2026 Demo / reviewers in the wild / expert
Linna Fan
dblp:238/8673
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0003-2269-1175ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 4 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhance CVE Severity Prediction From Vulnerability Description with Auxiliary SentenceabstractVulnerability severity assessment is of paramount importance for the cybersecurity defense strategy, as it enables the timely identification of threats and their underlying capabilities, thereby maximizing the efficacy of defense actions. And natural language processing (NLP)-based methods are widely adopted to construct an end-to-end CVE severity predictor based on existing vulnerability description. However, the long-tail distribution in the sentence length and word frequency of vulnerability description significantly hampers the performance of current methodologies. Coping with this challenge, we introduce CVSS-Predictor-AS, which mitigates the long-tail distribution by incorporating auxiliary sentences pertinent to the context of the text. The incorporation of auxiliary sentences rebalances the length distribution of descriptive sentences and addresses the information gaps, particularly for rare words. Comparative experiments conducted on vulnerability descriptions sourced from CVEdetail demonstrate that CVSS-Predictor-AS exhibits notable advantages over existing methods. Lin Ni, Linna Fan |
CSCWD | 5 |
| 2025 | MalImgDA: Diffusion-based Data Augmentation for Long-tailed Malware Family ClassificationabstractWith the rapid improvement of machine learning technology, leveraging machine learning methods for malware classification has emerged as a viable approach. However, under real-world circumstance, the imbalanced or long-tailed distribution among various malware families, poses a critical challenge to classify such few-shot malware families, eliminating the effectiveness of trained classifier. In this paper, we propose MalImgDA, a novel data augmentation framework that leverages diffusion model-based approach to tackle the long-tailed malware family classification problem. By fine-tuning the pre-trained diffusion model on few-shot data, we synthesize analogous malware images from existing samples with high resemblance to target family. And then we can mingle synthetic samples with existing data to build a re-balanced dataset for classifier training. Particularly, by utilizing MalImgDA, we can substantially enhance the diversity of data and generate plausible malware variants proactively while persevering the characteristics of target family. Experiments conducted on two publicly available datasets demonstrate the effectiveness of our proposed method in comparison to other commonly-used approaches. Linna Fan, Lin Ni |
ICASSP | 5 |
| 2025 | HGExplainer: Heterogeneous Graph Explainer for IoT Device IdentificationabstractIoT device identification is vital for network asset and security management. However, existing methods use statistical features that can not identify IoT devices accurately in complex network environments.GraphIoTproposes using non-statistical features and building a heterogeneous graph neural network to identify IoT devices accurately. However, heterogeneous graph neural networks lack interpretability, which reduces trust in the model. Besides, it is difficult to deploy on resource-constrained devices, limiting the broad application of IoT device identification. To make IoT device identification interpretable, easy to deploy, and with high accuracy, we get the interpretation results ofGraphIoTthrough interpretability and further build the rule set based on the interpretation results. Considering there is no suitable interpreter forGraphIoTwith many nodes and edges, we proposeHGExplainer, which reduces the time complexity by splitting the interpretation target into important relation solving and edge solving and uses a novel solution method, ExpandTree. Then, we also designed a rule extractor, which can build rule sets based on the interpretation results. Experimental results on Yourthings and UNSW datasets show thatHGExplainercan build high fidelity, concise sample-level explanations in less than 3 seconds, and the established rule set can precisely identify IoT devices. Linna Fan, Xuan Shen, Guanglei Song, Chaocan Xiang, Duohe Ma, Yongfeng Huang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | DAS-Gen: Continual Signature Generation for Evolving Malicious Traffic
Weifeng Mou, Linna Fan, Xuan Shen |
ICIC (9) | 4 |
| 2024 | AdvOcl: Naturalistic Clothing Pattern Adversarial to Person Detectors in OcclusionabstractAutomated surveillance cameras equipped with intelligent person detection systems are believed to have reached the maturity required for deployment in Intelligent Transport Systems, Intelligent Plants, and so on. However, recent studies have revealed that Deep Learning Neural Networks (DNN), on which mainstream person detection models are built, are vulnerable to adversarial attacks. Several methods have been proposed to generate adversarial patches that can evade person detectors. Nevertheless, these methods have limitations, as these adversarial patches are either restricted to being presented without any occlusion and placed in the center of the person, or they are too large in size and standing-out in pattern to be easily ignored by human eyes. Therefore, the adversarial patches in previous works did not consider both robustness and stealthiness when human posture changes and the patches are not in the center of person and partially occluded. In this paper, we propose AdvOcl that leverages the learned image manifold of the diffusion model to generate patterns that resemble one kind of the typical textures of daily clothes, such as common floral styles. Moreover, AdvOcl improved the adaptability and adversarial effectiveness by supporting changes in posture and partially occlusion during walking or running with warping and alignment module modeling deformation of clothes. Through extensive quantitative experiments, the results demonstrate the effectiveness of the proposed approach in generating more adversarially effective and naturalistic patterns in occluded scenarios compared to other state-of-the-art patch generation methods. Zhitong Lu, Duohe Ma, Linna Fan, Zhen Xu 0009, Kai Chen 0012 |
IH&MMSec | 3 |
| 2024 | CP-IoT: A Cross-Platform Monitoring System for Smart Home
Chenglong Li 0006, Jiahai Yang 0001, Linna Fan, Chenxin Duan |
NDSS | 5 |
| 2024 | KI-Mix: Enhancing Cyber Threat Detection in Incomplete Supervision Setting Through Knowledge-informed Pseudo-anomaly GenerationabstractData-driven methodologies have exhibited remarkable performance in identifying various cyber threats. However, obtaining well-labeled training samples is enormously expensive and often challenging when tackling practical cyber-security problems, due to the cost and difficulties in data annotation. To address this issue, we propose KI-Mix, a novel pseudo-anomaly generation algorithm for cyber threat detection on the basis of the limited labeled anomalies and a large volume of unlabeled data. In a nutshell, KI-Mix incorporates security domain knowledge into data interpolation to capture more labeled data to facilitate semi-supervised detection on cyber anomalies. We compare the performance of KI-Mix with several commonly applied augmentation techniques, such as Mixup and CutMix to evaluate its effectiveness in limited annotated data settings. Through extensive experiments on five security datasets covering various aspects of network threats, we demonstrate that KI-Mix outperforms other methods with equivalent baseline models. Notably, KI-Mix is model-agnostic to enable any data-driven threats detection models to handle incomplete supervision problems in real-world cyber threat detection. Linna Fan, Xia Tao |
SMC | 3 |
| 2024 | ProbeGeo: A Comprehensive Landmark Mining Framework Based on Web ContentabstractIP geolocation is essential for various location-aware Internet applications. High-quality IP geolocation landmarks play a decisive role in IP geolocation accuracy. However, the previous research works focusing on mining landmarks from the Internet are hampered by limited quantity, poor coverage, and insufficient landmark quality. In this paper, we present a new framework called ProbeGeo to mine high-quality landmarks automatically. We divide landmarks into common landmarks and probe landmarks, providing systematic mining methods based on online retrieval and web content. ProbeGeo expands traditional common landmarks by taking advantage of the exposure of multiple IoT (Internet of Things) devices on the Internet, mining them based on search engines and webpage contents. Common landmarks, consisting of multi-type devices, significantly improve landmark quantity and coverage. Furthermore, ProbeGeo establishes a methodology for acquiring new probe landmarks from Internet VPs (Vantage Points) webpages, extracting geographical locations from heterogeneous webpages and utilizing active probe functions. Probe landmarks enhance landmark quality and functions, bringing new geolocation frameworks and breaking through the geolocation accuracy bottleneck. We develop the ProbeGeo as a continuously running system and conduct real-world experiments to validate its efficacy. Our results show that ProbeGeo can detect 89,849 high-quality landmarks, including 6,874 probe landmarks and 82,975 common landmarks. ProbeGeo landmarks are about 10x more than existing work, distributed in 181 countries and 7,094 cities. ProbeGeo landmarks cover more than 8 types of devices, and more than 60% of them remain stable over one month. Moreover, the landmark accuracy of more than 58% of ProbeGeo landmarks is above street level, which has not been achieved in previous works. ProbeGeo can provide geolocation services with higher landmark accuracy and broader coverage by correlating a large scale of landmarks. Jinlei Lin, Chenglong Li 0006, Guanglei Song, Linna Fan, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | PMap: Reinforcement Learning-Based Internet-Wide Port ScanningabstractInternet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the most open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap performs better on hit rate, coverage, and intrusiveness. Our experiments over real networks show that PMap can find 90% open ports by only scanning 125 ports (90%@125) to each address, which is 99.3% less than the state-of-the-art port scanning methods. It reduces the number of scanned ports to decrease the intrusive nature of port scanning. In addition, PMap is highly parallel and lightweight. It scans 500 networks in parallel, achieving a port recommendation rate of up to 18 million per second, consuming only 7GB of memory. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research. Guanglei Song, Lin He 0004, Jinlei Lin, Linna Fan, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | AddrMiner: A Fast, Efficient, and Comprehensive Global Active IPv6 Address Detection SystemabstractFast Internet-wide scanning is essential for network situational awareness and asset evaluation. However, the vast IPv6 address space makes brute-force scanning infeasible. Despite advancements in state-of-the-art methods, they do not work in seedless regions and suffer low detection efficiency and speed in regions with known active IPv6 addresses (i.e., seed addresses). Moreover, the collected active address list (i.e., IPv6 hitlist) with low coverage cannot truly represent the active IPv6 address landscape of the Internet. This paper introduces AddrMiner, a fast, efficient, and comprehensive global active IPv6 address detection system. We design a systematic active IPv6 address detection strategy that divides the IPv6 space into two detection scenarios based on the presence or absence of seed addresses to discover active IPv6 addresses from scratch and from few to many. In the seedless regions, we present AddrMiner-N, leveraging a multi-level association policy to probe active addresses. It fills the gap of address detection in seedless regions and successfully discovers active addresses in 39,899 BGP prefixes without seed addresses, with a$1.03\times $higher hit rate,$30\sim 911\times $higher speed, and$2.7\times $broader coverage, compared to existing solutions. In the regions with seed addresses, our method AddrMiner-S dynamically generates target addresses using reinforcement learning. Compared to state-of-the-art methods, AddrMiner-S achieves an impressive 56.3% hit rate and a discovery speed of 839.0/s, which is$1.9\sim 2153\times $and$1.5\sim 755\times $of existing works, respectively. Finally, we deploy AddrMiner and discover 2.1B active IPv6 addresses, including 1.7B de-aliased active addresses and 0.4B aliased addresses, through continuous probing for three years. Guanglei Song, Lin He 0004, Feiyu Zhu 0002, Jinlei Lin, Wenjian Zhang, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | GraphIoT: Accurate IoT Identification based on Heterogeneous GraphabstractIoT devices deployed on campus and enterprise networks facilitate people's lives and work. However, these devices also bring serious network asset management and security management problems. IoT device identification is the premise to solve these problems. Although current IoT identification methods can identify devices with relatively high accuracy in ideal environments, it is difficult to accurately identify devices in real-world complex environments (e.g., campus networks, enterprise networks). Therefore, we propose to use exact features. To solve the problem of different dimensions of exact features, we creatively model the IoT identification problem as a heterogeneous graph representation learning problem and design a new representation learning algorithm. We are the first to propose an approach to accurately identify IoT devices in real-world complex environments and solve this problem through heterogeneous graphs. The evaluation shows that GraphIoT's macro F1 is on average 13.58% and 12.77% higher than the other methods on two public datasets. Linna Fan, Lin He 0004, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001, Jinlei Lin, Guanglei Song |
IWQoS | 1 |
| 2023 | Which Doors Are Open: Reinforcement Learning-based Internet-wide Port ScanningabstractInternet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the majority of open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap achieves better performance on hit rate, coverage, and intrusiveness. Our experiments over real-world networks show that PMap can find 90% open ports by only scanning 125 ports (90% @125) to each active address with 136× less than the state-of-the-art port probing methods. PMap reduces the number of scanned ports to decrease the intrusive nature of port scanning. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research. Guanglei Song, Lin He 0004, Tianyun Zhao, Yirui Luo, Yichao Wu, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IWQoS | 6 |
| 2023 | AutoIoT: Automatically Updated IoT Device Identification With Semi-Supervised LearningabstractIoT devices bring great convenience to a person's life and industrial production. However, their rapid proliferation also troubles device management and network security. Network administrators usually need to know how many IoT devices are in the network and whether they behave normally. IoT device identification is the first step to achieving these goals. Previous IoT device identification methods reach high accuracy in a closed environment. But they are not applicable in the continuously changing environment. When new types of devices are plugged in, they cannot update themselves automatically. Besides, they usually rely on supervised learning and need lots of labeled data, which is costly. To solve these problems, we propose a novel IoT device identification model namedAutoIoT, updating itself automatically when new types of devices are plugged in. Besides, it only needs a few labeled data and identifies IoT devices with high accuracy. The evaluation on two public datasets shows thatAutoIoTcan identify new device types only using 1.5$\sim$2.5 hours’ traffic and still have high accuracy after updating. Moreover, it has a better performance than other works when there are only a few labeled data, especially in an environment with scanning traffic. Linna Fan, Lin He 0004, Yichao Wu, Shize Zhang, Jia Li 0033, Jiahai Yang 0001, Chaocan Xiang, Xiaoqian Ma |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | EvoIoT: An evolutionary IoT and non-IoT classification model in open environments
Linna Fan, Lin He 0004, Enhuan Dong, Jiahai Yang 0001, Chenglong Li 0006, Jinlei Lin |
Comput. Networks | 1 |
| 2020 | An IoT Device Identification Method based on Semi-supervised LearningabstractWith the rapid proliferation of IoT devices, device management and network security are becoming significant challenges. Knowing how many IoT devices are in the network and whether they are behaving normally is significant. IoT device identification is the first step to achieve these goals. Previous IoT identification works mainly use supervised learning and need lots of labeled data. Considering collecting labeled data is time-consuming and cannot be scaled, in this paper, we propose an IoT identification model based on semi-supervised learning. The model can differentiate IoT and non-IoT and classify specific IoT devices based on time interval features, traffic volume features, protocol features and TLS related features. The evaluation in a public dataset shows that our model only needs 5% labeled data and gets accuracy over 99%. Linna Fan, Shize Zhang, Yichao Wu, Chenxin Duan, Jia Li 0033, Jiahai Yang 0001 |
CNSM | 1 |