EDBT 2026 Demo / reviewers in the wild / expert
Yuede Ji
dblp:138/4316
· DBLP profile ↗
29ranked-venue papers
10as first author
19since 2021 · last 2025
0000-0002-2419-6592ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 12 · 6 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HuggingGraph: Understanding the Supply Chain of LLM EcosystemabstractLarge language models (LLMs) leverage deep learning architectures to process and predict sequences of words, enabling them to perform a wide range of natural language processing tasks, such as translation, summarization, question answering, and content generation. As existing LLMs are often built from base models or other pre-trained models and use external datasets, they can inevitably inherit vulnerabilities, biases, or malicious components that exist in previous models or datasets. Therefore, it is critical to understand these components’ origin and development process to detect potential risks, improve model fairness, and ensure compliance with regulatory frameworks. Motivated by that, this project aims to study such relationships between models and datasets, which are the central parts of the LLM supply chain. First, we design a methodology to systematically collect LLMs’ supply chain information. Then, we design a new graph to model the relationships between models and datasets, which is a directed heterogeneous graph, having 402,654 nodes and 462,524 edges. Lastly, we perform different types of analysis and make multiple interesting findings. Mohammad Shahedur Rahman, Peng Gao 0008, Yuede Ji |
CIKM | 3 |
| 2025 | Bingo: Radix-based Bias Factorization for Random Walk on Dynamic GraphsabstractRandom walks are a primary means for extracting information from large-scale graphs. While most real-world graphs are inherently dynamic, state-of-the-art random walk engines failed to efficiently support such a critical use case. This paper takes the initiative to build a general random walk engine for dynamically changing graphs with two key principles: (i) This system should support both low-latency streaming updates and high-throughput batched updates. (ii) This system should achieve fast sampling speed while maintaining acceptable space consumption to support dynamic graph updates. Upholding both standards, we introduce Bingo, a GPU-based random walk engine for dynamically changing graphs. First, we propose a novel radix-based bias factorization algorithm to support constant time sampling complexity while supporting fast streaming updates. Second, we present a group-adaption design to reduce space consumption dramatically. Third, we incorporate GPU-aware designs to support high-throughput batched graph updates on massively parallel platforms. Together, Bingo outperforms existing efforts across various applications, settings, and datasets, achieving up to a 271.11x speedup compared to the state-of-the-art efforts. Pinhuan Wang, Chengying Huan, Zhibin Wang 0002, Chen Tian 0001, Yuede Ji, Hang Liu 0001 |
EuroSys | 5 |
| 2025 | TS-Net: Dual-Channel IoT Intrusion Detection with Temporal and Spatial ModelingabstractThe rapid growth of the Internet of Things (IoT) has introduced significant security challenges, particularly in detecting intrusions within complex IoT networks. This paper presents TS-Net, a robust dual-channel model that combines temporal and spatial feature learning to enhance IoT intrusion detection. By partitioning network traffic into temporal and spatial features, TS-Net processes them through separate channels. The temporal channel utilizes Bidirectional Gated Recurrent Units (BiGRU) paired with a self-attention mechanism to capture dynamic sequential dependencies, while the spatial channel employs multi-scale dilated convolutions to extract patterns from varying spatial perspectives. These two channels are then fused to improve the model accuracy in detecting anomalous traffic. Experimental results on three publicly available datasets demonstrate that TS-Net outperforms existing intrusion detection models, achieving higher precision, recall, and F1-scores, demonstrating its effectiveness in addressing the unique security needs of IoT networks. Haotian Chi, Haijun Geng, Xiaojiang Du, Yuede Ji |
GLOBECOM | 6 |
| 2025 | A Secret Sharing-Inspired Robust Distributed Backdoor Attack to Federated LearningabstractFederated Learning (FL) is vulnerable to backdoor attacks—especially distributed backdoor attacks (DBA) that are more persistent and stealthy than centralized backdoor attacks. However, we observe that the attack effectiveness of DBA can be largely reduced when encountering rebels, i.e., the agents promising to perform the attack but do not do so. To robustify DBAs, we present SSRDBA , a secret sharing-inspired robust DBA to FL. To be specific, given a same global trigger as DBA, SSRDBA carefully divides it into different shares based on secret sharing and exploits these shares to poison local data on malicious devices, respectively. SSRDBA enjoys several merits, e.g., only partial malicious agents guarantee the reconstruction of the global trigger. Extensive experimental results show that SSRDBA is more robust to rebels than DBA and can evade the state-of-the-art FL defenses mainly for centralized backdoor attacks. To mitigate SSRDBA , we further design a novel defense mechanism, termed NFDR, which shows great potential against SSRDBA on certain independent identically distributed datasets. Yuxin Yang 0003, Qiang Li 0008, Yuede Ji, Binghui Wang |
ACM Trans. Priv. Secur. | 3 |
| 2024 | CloudCover: Enforcement of Multi-Hop Network Connections in Microservice DeploymentsabstractMicroservices have emerged as a strong architecture for large-scale, distributed systems in the context of cloud computing and containerization. However, the size and complexity of microservice systems have strained current access control mechanisms. Intricate dependency structures, such as multi-hop dependency chains, go uncaptured by existing access control mechanisms and leave microservice deployments open to adversarial actions and influence.This work introduces CloudCover, an access control mechanism and enforcement framework for microservices. CloudCover provides holistic, deployment-wide analysis of microservice operations and behaviors. It implements a verification-in-the-loop access control approach, mitigating multi-hop microservice threats through control-flow integrity checks. We evaluate these domain-relevant multi-hop threats and CloudCover under existing, real-world scenarios such as Istio’s opensource microservice example and under theoretic and synthetic network loads of 10,000 requests per second. Our results show that CloudCover is appropriate for use in real deployments, requiring no microservice code changes by administrators. Dalton A. Brucker-Hahn, Shanchao Li, Matthew Petillo, Alexandru G. Bardas, Drew Davidson, Yuede Ji |
ACSAC | 7 |
| 2024 | Fine-Grained Geo-Obfuscation to Protect Workers' Location Privacy in Time-Sensitive Spatial Crowdsourcing
Chenxi Qiu, Yuede Ji, Anna Cinzia Squicciarini, Ram Dantu, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001 |
EDBT | 3 |
| 2024 | TrustEvent: Cross-Platform IoT Trigger Event Verification Using Edge ComputingabstractAs smart home IoT systems gain popularity, they inevitably become targets for security risks and concerns. Among various cyber-attacks targeting these systems, the fake event attack poses significant issues due to its ability to manipulate secure devices through automation rules. In response to this threat, we propose TrustEvent - a system designed to offer end-to-end event signature verification. By integrating TrustEvent with existing home automation platforms, event authenticity is verified against signatures generated from edge devices before these events trigger automation rule execution. Notably, we have developed a signature proxy module, enhancing our system's compatibility across various platform scenarios. We have implemented a TrustEvent prototype in conjunction with existing commercial smart home IoT platforms, evaluating its overhead in the process. Our experimentation demonstrates that our system only marginally increases the automation execution latency, by an average of 3.74 seconds, representing a acceptable compromise for enhanced security. Trent Reichenbach, Chenglong Fu 0002, Xiaojiang Du, Jia Di, Yuede Ji |
ICC | 5 |
| 2024 | Code is not Natural Language: Unlock the Power of Semantics-Oriented Graph Representation for Binary Code Similarity Detection
Haojie He, Xingwei Lin, Ziang Weng, Ruijie Zhao 0001, Shuitao Gan, Libo Chen 0001, Yuede Ji, Jiashui Wang, Zhi Xue |
USENIX Security Symposium | 7 |
| 2024 | API2Vec++: Boosting API Sequence Representation for Malware Detection and ClassificationabstractAnalyzing malware based on API call sequences is an effective approach, as these sequences reflect the dynamic execution behavior of malware. Recent advancements in deep learning have facilitated the application of these techniques to mine valuable information from API call sequences. However, these methods typically operate on raw sequences and may not effectively capture crucial information, especially in the case of multi-process malware, due to theAPI call interleaving problem. Furthermore, they often fail to capture contextual behaviors within or across processes, which is particularly important for identifying and classifying malicious activities. Motivated by this, we present API2Vec++, a graph-based API embedding method for malware detection and classification. First, we construct a graph model to represent the raw sequence. Specifically, we design the Temporal Process Graph (TPG) to model inter-process behaviors and the Temporal API Property Graph (TAPG) to model intra-process behaviors. Compared to our previous graph model, the TAPG model exposes operations with associated behaviors within the process through node properties and thus enhances detection and classification abilities. Using these graphs, we develop a heuristic random walk algorithm to generate numerous paths that can capture fine-grained malicious familial behavior. By pre-training these paths using the BERT model, we generate embeddings of paths and APIs, which can then be used for malware detection and classification. Experiments on a real-world malware dataset demonstrate that API2Vec++ outperforms state-of-the-art embedding methods and detection/classification methods in both accuracy and robustness, particularly for multi-process malware. Lei Cui 0003, Junnan Yin, Jiancong Cui, Yuede Ji, Peng Liu 0044, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Software Eng. | 4 |
| 2023 | API2Vec: Learning Representations of API Sequences for Malware DetectionabstractAnalyzing malware based on API call sequence is an effective approach as the sequence reflects the dynamic execution behavior of malware.Recent advancements in deep learning have led to the application of these techniques for mining useful information from API call sequences. However, these methods mainly operate on raw sequences and may not effectively capture important information especially for multi-process malware, mainly due to the API call interleaving problem. Lei Cui 0003, Jiancong Cui, Yuede Ji, Zhiyu Hao, Zhenquan Ding |
ISSTA | 3 |
| 2023 | TANGO: re-thinking quantization for graph neural network training on GPUsabstractGraph learning is becoming increasingly popular due to its superior performance in tackling many grand challenges. While quantization is widely used to accelerate Graph Neural Network (GNN) computation, quantized training faces remarkable roadblocks. Current quantized GNN training systems often experience longer training time than their full-precision counterparts for two reasons: (i) addressing the quantization accuracy challenge leads to excessive overhead, and (ii) the optimization potential exposed by quantization is not adequately leveraged. This paper introduces Tango which re-thinks quantization challenges and opportunities for graph neural network training on GPUs with three contributions: Firstly, we introduce efficient rules to maintain accuracy during quantized GNN training. Secondly, we design and implement quantization-aware primitives and inter-primitive optimizations to speed up GNN training. Finally, we integrate Tango with the popular Deep Graph Library (DGL) system and demonstrate its superior performance over the state-of-the-art approaches on various GNN models and datasets. Shiyang Chen 0004, Da Zheng 0004, Caiwen Ding, Chengying Huan, Yuede Ji, Hang Liu 0001 |
SC | 5 |
| 2023 | PeeK: A Prune-Centric Approach for K Shortest Path ComputationabstractThe K shortest path (KSP) algorithm, which finds the top K shortest simple paths from a source to a target vertex, has a wide range of real-world applications, e.g., routing, vulnerability detection, and biology analysis. While the top K shortest simple paths offer invaluable insights, computing them is time-consuming. For example, on a Twitter graph (61.6M vertices and 1.5B edges), the best parallel method needs about 20 minutes to get 128 shortest paths between two vertices. A key observation we made is existing works search K shortest paths from the original graph, while top K shortest paths only cover a meager portion of the original graph, e.g., less than 0.001% on a Twitter graph for K = 128. Shiyang Chen 0004, Hang Liu 0001, Yuede Ji |
SC | 4 |
| 2022 | Illuminati: Towards Explaining Graph Neural Networks for Cybersecurity AnalysisabstractGraph neural networks (GNNs) have been utilized to create multi-layer graph models for a number of cybersecurity applications from fraud detection to software vulnerability analysis. Unfortunately, like traditional neural networks, GNNs also suffer from a lack of transparency, that is, it is challenging to interpret the model predictions. Prior works focused on specific factor explanations for a GNN model. In this work, we have designed and implemented Illuminati, a comprehensive and accurate explanation framework for cybersecurity applications using GNN models. Given a graph and a pre-trained GNN model, Illuminati is able to identify the important nodes, edges, and attributes that are contributing to the prediction while requiring no prior knowledge of GNN models. We evaluate Illuminati in two cybersecurity applications, i.e., code vulnerability detection and smart contract vulnerability detection. The experiments show that Illuminati achieves more accurate explanation results than state-of-the-art methods, specifically, 87.6% of subgraphs identified by Illuminati are able to retain their original prediction, an improvement of 10.3% over others at 77.3%. Furthermore, the explanation of Illuminati can be easily understood by the domain experts, suggesting the significant usefulness for the development of cybersecurity applications. Yuede Ji, H. Howie Huang |
EuroS&P | 2 |
| 2022 | TLPGNN: A Lightweight Two-Level Parallelism Paradigm for Graph Neural Network Computation on GPUabstractGraph Neural Networks (GNNs) are an emerging class of deep learning models on graphs, with many successful applications, such as, recommendation systems, drug discovery, and social network analysis. The GNN computation includes both regular neural network operations and general graph convolution operations, which take the majority of the total computation time. Though several recent works have been proposed to accelerate the computation for GNNs, they face the limitations of heavy pre-processing, low efficient atomic operations, and unnecessary kernel launches. In this paper, we design TLPGNN, a lightweight two-level parallelism paradigm for GNN computation. First, we conduct a systematic analysis on the hardware resource usage of GNN workloads to deeply understand the specialties of GNN workloads. With the insightful observations, we then divide the GNN computation into two levels, i.e., vertex parallelism for the first level and feature par- allelism for the second. Next, we employ a novel hybrid dynamic workload assignment to address the imbalanced workload distribution. Furthermore, we fuse the kernels to reduce the number of kernel launches and cache the frequently accessed data into registers to avoid unnecessary memory traffics. Together, TLPGNN is able to significantly outperform existing GNN computation systems, such as DGL, GNNAdivsor, and FeatGraph, by 5.6×, 7.7×, and 3.3×, respectively, on the average. Yuede Ji, H. Howie Huang |
HPDC | 2 |
| 2022 | NestedGNN: Detecting Malicious Network Activity with Nested Graph Neural NetworksabstractNetwork attacks are dramatically increasing over the years. A graph can accurately model the network activities. Therefore, graph-based techniques are frequently used to detect network threats. Motivated by the strong representation of graph neural networks (GNNs), many GNN-based techniques have been proposed for various security problems, such as network threat detection, malware detection, insider threat detection, and fraud detection. Most GNNs work on the classical attributed graph structure, while we observe that a nested graph structure is a more accurate representation for modelling enterprise network, where the communications between hosts form a graph, while the local activities of each host, e.g., local event graph, form an inner graph. Observing no existing GNNs can directly learn on such a nested graph, in this paper, we designed NestedGNN, the first graph neural network for nested graphs. NestedGNN consists of three layers, i.e., inner GNN layers, nested graph layers, and outer GNN layers. We successfully applied it to compromised host detection. NestedGNN can significantly improve the performance over traditional methods on a publicly available cybersecurity dataset. Yuede Ji, H. Howie Huang |
ICC | 1 |
| 2021 | Vestige: Identifying Binary Code Provenance for Vulnerability Detection
Yuede Ji, Lei Cui 0003, H. Howie Huang |
ACNS (2) | 1 |
| 2021 | BugGraph: Differentiating Source-Binary Code Similarity with Graph Triplet-Loss NetworkabstractBinary code similarity detection, which answers whether two pieces of binary code are similar, has been used in a number of applications,such as vulnerability detection and automatic patching. Existing approaches face two hurdles in their efforts to achieve high accuracy and coverage: (1) the problem of source-binary code similarity detection, where the target code to be analyzed is in the binary format while the comparing code (with ground truth) is in source code format. Meanwhile, the source code is compiled to the comparing binary code with either a random or fixed configuration (e.g.,architecture, compiler family, compiler version, and optimization level), which significantly increases the difficulty of code similarity detection; and (2) the existence of different degrees of code similarity. Less similar code is known to be more, if not equally, important in various applications such as binary vulnerability study. To address these challenges, we design BugGraph, which performs source-binary code similarity detection in two steps. First, BugGraph identifies the compilation provenance of the target binary and compiles the comparing source code to a binary with the same provenance.Second, BugGraph utilizes a new graph triplet-loss network on the attributed control flow graph to produce a similarity ranking. The experiments on four real-world datasets show that BugGraph achieves 90% and 75% true positive rate for syntax equivalent and similar code, respectively, an improvement of 16% and 24% overstate-of-the-art methods. Moreover, BugGraph is able to identify 140 vulnerabilities in six commercial firmware. Yuede Ji, Lei Cui 0003, H. Howie Huang |
AsiaCCS | 1 |
| 2021 | DEFInit: An Analysis of Exposed Android Init Routines
Yuede Ji, Mohamed Elsabagh, Ryan Johnson 0002, Angelos Stavrou |
USENIX Security Symposium | 1 |
| 2021 | Discovering unknown advanced persistent threat using shared features mined by neural networks
Longkang Shang, Dong Guo 0002, Yuede Ji, Qiang Li 0008 |
Comput. Networks | 3 |
| 2020 | Aquila: Adaptive Parallel Computation of Graph Connectivity QueriesabstractGraph connectivity algorithms answer whether two nodes in a graph are connected under specific conditions, which are beneficial to a number of applications, such as pattern recognition and cybersecurity. Unfortunately, existing graph computing frameworks support only a small number of connectivity algorithms and achieve low computation parallelism. In this paper, we have designed an adaptive parallel computation framework, Aqila, that covers a wide range of different highly optimized graph connectivity algorithms. Given a graph, Aqila first transforms the query if it can be answered with partial computation. During the computation, Aqila is able to greatly reduce the workload by up to 98%. Furthermore, Aqila identifies the irregular tasks in the connectivity algorithms and applies different parallel strategies for different tasks. As a result, Aqila significantly outperforms existing systems such as Multistep, Galois, Ligra, GraphChi, X-Stream, DFS, and Boost, by average 13x, 53x, 264x, 364x, 1,369x, 45x, and 255x, respectively. Yuede Ji, H. Howie Huang |
HPDC | 1 |
| 2020 | Detecting Lateral Movement in Enterprise Computer Networks with Unsupervised Graph AI
Benjamin Bowman, Craig Laprade, Yuede Ji, H. Howie Huang |
RAID | 3 |
| 2020 | GRL: Knowledge graph completion with GAN-based reinforcement learning
Qi Wang 0044, Yuede Ji, Yongsheng Hao, Jie Cao 0011 |
Knowl. Based Syst. | 2 |
| 2019 | Stopping the Cyberattack in the Early Stage: Assessing the Security Risks of Social Network UsersabstractOnline social networks have become an essential part of our daily life. While we are enjoying the benefits from the social networks, we are inevitably exposed to the security threats, especially the serious Advanced Persistent Threat (APT) attack. The attackers can launch targeted cyberattacks on a user by analyzing its personal information and social behaviors. Due to the wide variety of social engineering techniques and undetectable zero-day exploits being used by attackers, the detection techniques of intrusion are increasingly difficult. Motivated by the fact that the attackers usually penetrate the social network to either propagate malwares or collect sensitive information, we propose a method to assess the security risk of the user being attacked so that we can take defensive measures such as security education, training, and awareness before users are attacked. In this paper, we propose a novel user analysis model to find potential victims by analyzing a large number of users’ personal information and social behaviors in social networks. For each user, we extract three kinds of features, i.e., statistical features, social-graph features, and semantic features. These features will become the input of our user analysis model, and the security risk score will be calculated. The users with high security risk score will be alarmed so that the risk of being attacked can be reduced. We have implemented an effective user analysis model and evaluated it on a real-world dataset collected from a social network, namely, Sina Weibo (Weibo). The results show that our model can effectively assess the risk of users’ activities in social networks with a high area under the ROC curve of 0.9607. Qiang Li 0008, Yuede Ji, Dong Guo 0002, Xiangyu Meng 0002 |
Secur. Commun. Networks | 3 |
| 2018 | iSpan: parallel identification of strongly connected components with spanning trees
Yuede Ji, Hang Liu 0001, H. Howie Huang |
SC | 1 |
| 2016 | Combating the evasion mechanisms of social bots
Yuede Ji, Xinyang Jiang, Qiang Li 0008 |
Comput. Secur. | 1 |
| 2015 | BotCatch: leveraging signature and behavior for bot detectionabstractAbstract The goal of bot detection is to discover malicious bot processes by signature comparison or behavior analysis. Existing approaches have several drawbacks, such as requiring a lot of prior knowledge, low detection accuracy, and high false alarm rate. In this paper, we propose a multi‐feedback approach, BotCatch, to detect bots effectively and efficiently on a host by leverage of a combination of signature and behavior. First, BotCatch assigns suspicious files to signature‐analysis and behavior‐analysis modules, which generate each detection result. Second, BotCatch correlates signature and behavior results to generate the final detection result through correlation engine. Third, BotCatch feeds back signature, behavior, and correlation results to dynamically adjust detecting modules through multi‐feedback engine. We evaluated the performance of BotCatch with 636 bot and 150 benign samples. Our results indicate that BotCatch achieves an accuracy of 97.1%and an F‐measure value of 0.982 simultaneously, which is better than existing approaches without feedbacks. BotCatch, due to the multi‐feedback mechanism, has the ability to gradually get more robust and accurate as the number of samples increases. The final stage even reaches an accuracy of 98.5%and F‐measure value of 0.991. Copyright © 2014 John Wiley & Sons, Ltd. Yuede Ji, Qiang Li 0008, Dong Guo 0002 |
Secur. Commun. Networks | 1 |
| 2014 | Towards social botnet behavior detecting in the end hostabstractSocial botnet utilizing online social network (OSN) as Command and Control channel (C&C) has caused enormous threats to Internet security. Server-side detection approaches mainly target on suspicious accounts, which cannot identify the specific bot hosts or processes. Host-side approaches target on suspicious process behaviors which are not robust enough to face the challenges of frequent variants and novel social bots. In this paper, we propose a novel social bot behavior detecting approach in the end host. Because social bot binaries or source codes are not easy to collect, we first design a novel social botnet, named wbbot, based on Sina Weibo. We analyze it from two aspects, wbbot architecture and wbbot behaviors. Second, we analyze the host behaviors of existing social botnets which come from public websites, other researchers, and our implementations. We identify six critical phases: infection, pre-defined host behaviors, establishment of C&C, receive the commands of botmaster, execution of social bot commands, and return the results. Third, we present our detection system which consists of three components: host behavior monitor, host behavior analyzer, and detection approach. We present behavior tree-based approach to detect social bot. After constructing the suspicious behavior tree, we match it with the template library to generate detection result. Finally, we collect real-world social botnet traces to evaluate the performance. We would like to share them for academic research. The results indicate that our system has an acceptable false positive rate of 29.6% and remarkable false negative rate of 4.5%. However, compared with other detection tools, our detection result is still remarkable. Yuede Ji, Xinyang Jiang, Qiang Li 0008 |
ICPADS | 1 |
| 2014 | A Mulitiprocess Mechanism of Evading Behavior-Based Bot Detection Approaches
Yuede Ji, Dewei Zhu, Qiang Li 0008, Dong Guo 0002 |
ISPEC | 1 |
| 2013 | BotInfer: A Bot Inference Approach by Correlating Host and Network Information
Qiang Li 0008, Yuede Ji, Dong Guo 0002 |
NPC | 3 |