Jiyong Jang

dblp:58/688 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 22 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 JBomAudit: Assessing the Landscape, Compliance, and Security Implications of Java SBOMs
Yue Xiao 0007, Dhilung Kirat, Douglas Lee Schales, Jiyong Jang, Luyi Xing, Xiaojing Liao
NDSS4
2024 Disentangled Knowledge Distillation for Unified Multi-Class Anomaly Detection
abstract
Anomaly detection is essential for image-based industrial inspection, yet class-specific models hinder its scalability and adaptability. This issue is exacerbated by ‘normality confusion’, where models struggle to distinguish normal from abnormal features across diverse classes. To address this, we introduce TwoStudents, a unified anomaly detection framework utilizing a knowledge distillation model with dual student decoders for efficient feature disentanglement. By separating class-specific abnormal features from normal ones, TwoStudents strengthens the unified normal representation. Our approach employs image transformations to simulate anomalies and uses cross-distillation training techniques on features derived from normal images and their corresponding simulated abnormal images. This effectively isolates abnormal features, enabling a unified model applicable to all classes. Evaluations on the MVTec AD, VisA, and BTAD datasets demonstrate that TwoStudents outperforms existing methods, achieving state-of-the-art performance. Notably, it significantly improves fine-grained localization, as measured by pixel AUPRO, with increases from 91.9% to 94.0% on MVTec AD and 87.6% to 90.5% on VisA.
Jiyong Jang, Hayeon Lee, Younkwan Lee
ICIP1
2024 Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization
abstract
We present a hybrid program analysis framework that automates the synthesis of stateful system call policies that describe admissible behaviors of containerized programs. Given a container image as input, the framework generates a reference policy that encodes a security automaton obtained by symbolically micro-executing the corresponding container's binary entrypoint under the constraints extracted from the container image metadata and environment.We demonstrate the utility and practicality of our approach by synthesizing security policies for 25 challenges in the DARPA Cyber Grand Challenge (CGC) corpus, 5 real-world containerized programs, including the widely used NGINX web server, and a complete microservice application from public benchmarks. We run each program or microservice using both benign and attack scenarios under the protection of a runtime policy monitor. Furthermore, we evaluate our approach by comparing our synthesized policies to those generated by four state-of-the-art system call specialization tools. Our results demonstrate that our techniques can scale to large programs and accurately extract concise reference application models for security monitoring.
William Blair, Frederico Araujo, Teryl Taylor, Jiyong Jang
SP4
2024 SyzGen++: Dependency Inference for Augmenting Kernel Driver Fuzzing
abstract
In recent years, kernel fuzzing research has experienced a significant surge. Among various kernel fuzzers, Syzkaller stands out as the state-of-the-art tool, having identified over 5,000 bugs in the Linux kernel. Syzkaller’s success can be attributed to its utilization of manually-curated syscall specifications provided by kernel experts. However, this process is time-consuming and not scalable due to complex input structures and unknown dependencies among syscalls. Consequently, a substantial portion of the kernel codebase, specifically kernel drivers, lacks specifications, posing a significant security risk.In this paper, we introduce SyzGen++, an innovative approach for automatically inferring dependencies between syscalls and generating specifications without relying on existing test suites. Specifically, we define two fundamental building blocks of insertion and lookup operations and their pairing to accurately identify dependencies. We evaluated SyzGen++ against existing state-of-the-art techniques on both Linux and macOS drivers. Our results demonstrate that SyzGen++ uncovered 245 more dependencies. Furthermore, SyzGen++ outperforms DIFUZE, KSG, and SyzDescribe in terms of code coverage, achieving 71%, 67%, and 39% improvement on average, respectively. Notably, our evaluation discovered 10 previously unknown bugs in Linux Kernel 6.2 using specifications generated by SyzGen++, resulting in 6 CVEs, which demonstrates its effectiveness in identifying vulnerabilities.
Weiteng Chen, Yu Hao 0006, Zheng Zhang 0058, Xiaochen Zou, Dhilung Kirat, Shachee Mishra, Douglas Lee Schales, Jiyong Jang, Zhiyun Qian
SP8
2023 EdgeTorrent: Real-time Temporal Graph Representations for Intrusion Detection
abstract
Anomaly-based intrusion detection aims to learn the normal behaviors of a system and detect activity that deviates from it. One of the best ways to represent the behavior of a computer network is through provenance graphs: dynamic networks of entity interactions over time. When provenance graphs deviate from their normal behaviors, it could be indicative of a malicious actor attempting to compromise the network. However, efficiently characterizing the normal behavior of large temporal graphs is challenging. To do this, we propose EdgeTorrent, an end-to-end anomaly-based intrusion detection system for provenance graph analysis. EdgeTorrent leverages a novel high-performance message passing neural network for graph embedding over a stream of edges to capture both temporal and topological changes in the system. These embeddings are then processed by a novel adversarially trained sequence analyzer that alerts when a series of graph embeddings changes in an unexpected way. EdgeTorrent preserves temporal ordering during message passing, and its streaming-focused design allows users to conduct out-of-core inference on billion-edge graphs, faster than real-time. We show that our method outperforms state-of-the-art graph-kernel approaches on several host monitoring data sets; notably, it is the first intrusion detection system to perfectly classify the StreamSpot data set. Additionally, we show it is the best-performing method on a real-world, billion-edge data set encompassing 11 days of benign and attack data.
Isaiah J. King, Xiaokui Shu, Jiyong Jang, Kevin Eykholt, Taesung Lee, H. Howie Huang
RAID3
2023 Fashion Faux Pas: Implicit Stylistic Fingerprints for Bypassing Browsers' Anti-Fingerprinting Defenses
abstract
Browser fingerprinting remains a topic of particular interest for both the research community and the browser ecosystem, and various anti-fingerprinting countermeasures have been proposed by prior work or deployed by browsers. While preventing fingerprinting presents a challenging task, modern fingerprinting techniques heavily rely on JavaScript APIs, which creates a choke point that can be targeted by countermeasures. In this paper, we explore how browser fingerprints can be generated without using any JavaScript APIs. To that end we develop StylisticFP, a novel fingerprinting system that relies exclusively on CSS features and implicitly infers system characteristics, including advanced fingerprinting attributes like the list of supported fonts, through carefully constructed and arranged HTML elements. We empirically demonstrate our system's effectiveness against privacy-focused browsers (e.g., Safari, Firefox, Brave, Tor) and popular privacy-preserving extensions. We also conduct a pilot study in a research organization and find that our system is comparable to a state-of-the-art JavaScript-based fingerprinting library at distinguishing devices, while outperforming it against browsers with anti-fingerprinting defenses. Our work highlights an additional dimension of the significant challenge posed by browser fingerprinting, and reaffirms the need for more robust detection systems and countermeasures.
Xu Lin 0003, Frederico Araujo, Teryl Taylor, Jiyong Jang, Iasonas Polakis
SP4
2023 URET: Universal Robustness Evaluation Toolkit (for Evasion)
Kevin Eykholt, Taesung Lee, Douglas Lee Schales, Jiyong Jang, Ian M. Molloy, Masha Zorin
USENIX Security Symposium4
2022 RAPID: Real-Time Alert Investigation with Context-aware Prioritization for Efficient Threat Discovery
abstract
Alerts reported by intrusion detection systems (IDSes) are often the starting points for attack campaign discovery and response procedures. However, the sheer number of alerts compared to the number of real attacks, along with the complexity of alert investigations, poses a challenge to achieving effective alert triage with limited computational resources. Automated procedures and human analysts could suffer from the burden of analyzing floods of alerts, and fail to respond to critical alerts promptly.
Yushan Liu 0004, Xiaokui Shu, Yixin Sun 0004, Jiyong Jang, Prateek Mittal
ACSAC4
2022 An In-Vehicle Warning Information Provision Strategy for V2V-Based Proactive Traffic Safety Management
abstract
The availability of vehicle interaction data, which is obtained by an in-vehicle forward collision warning system, including spacing between the leading and the following vehicle and time-to-collision, provides a valuable opportunity to predict crash risks in real time. When this opportunity is combined with connected vehicle technologies including vehicle-to-vehicle wireless communications, it is expected that more effective crash prevention would be achievable by providing predictive warning information as a part of proactive traffic safety management (PTSM). The purpose of this study is to develop a more reliable in-vehicle warning information provision strategy based on the prediction of crash risks using vehicle interaction data. A crash risk prediction model based on a long short-term memory was able to predict the crash risk after 3 seconds with a mean absolute percentage error of 8% using the data for the past 5 seconds. The predicted crash risk data were applied to derive the optimal threshold for triggering in-vehicle warning information, which is the essence of the proposed warning provision strategy. This study defined three indicators to evaluate the reliability of warning information: correct detection rate (CDR), detection failure rate (DFR), and information provision rate (IPR). An exemplar analysis result showed that the optimal threshold to minimize IPR in a situation where CDR and DFR are 100% and 0%, respectively, was identified as 0.69. The proposed methodology that predicts crash risks in real time and provides V2V-based warning information in a more proactive manner is expected to mitigate the crash risk significantly.
Young Jo, Jiyong Jang, Jieun Ko, Cheol Oh
IEEE Trans. Intell. Transp. Syst.2
2022 A Multi-Agent Driving Simulation Approach for Evaluating the Safety Benefits of Connected Vehicles
abstract
The in-vehicle warning information provided in the CVs environment allow the driver to respond rapidly to upcoming hazardous situations. The main purpose of this study is to evaluate the safety benefits due to the provision of warning information by analyzing vehicle interactions that are defined as the behavior change of the subject vehicle and the preceding vehicle. A notable feature of this study is the use of a multi-agent driving simulation (MADS) method to analyze the vehicle interaction with various vehicle pairs that are composed of the connected vehicle (CV) capable of receiving warning information and the regular vehicle (RV) that does not receive warning information. A total of four scenarios representing different vehicle pairs, which include CV-RV, RV-CV, CV-CV, and RV-RV, are evaluated in this study. The proposed analysis consists of four parts: the characteristics of subject vehicle maneuvering, variation of relative speed, evasive maneuvering by lane change, and overall crash potential. As an example, the result of analyzing crash potential index (CPI) showed that the greatest safety benefits were obtained with the CV-CV case among aforementioned four vehicle pairs. Approximately 45% of CPI reduction was achievable with the CV-CV case, compared to the RV-RV case. In addition to the CPI, useful findings obtained by investigating safety indicators for each vehicle pair are discussed in terms of safety benefits. The results of this study are expected to be used for both deriving valuable policies and developing more effective in-vehicle warning technologies to fully exploit the benefits of CVs.
Jieun Ko, Jiyong Jang, Cheol Oh
IEEE Trans. Intell. Transp. Syst.2
2021 Adaptive Verifiable Training Using Pairwise Class Similarity
Shiqi Wang 0002, Kevin Eykholt, Taesung Lee, Jiyong Jang, Ian M. Molloy
AAAI4
2020 Scarecrow: Deactivating Evasive Malware via Its Own Evasive Logic
abstract
Security analysts widely use dynamic malware analysis environments to exercise malware samples and derive virus signatures. Unfortunately, malware authors are becoming more aware of such analysis environments. Therefore, many have embedded evasive logic into malware to probe execution environments before exposing malicious behaviors. Consequently, such analysis environments become useless and evasive malware can damage victim systems with unforeseen malicious activities. However, adopting evasive techniques to bypass dynamic malware analysis is a double-edged sword. While evasive techniques can avoid early detection through sandbox analysis, it also significantly constrains the spectrum of execution environments where the malware activates. In this paper, we exploit this dilemma and seek to reverse the challenge by camouflaging end-user execution environments into analysis-like environments using a lightweight deception engine called SCARECROW. We thoroughly evaluate SCARECROW with real evasive malware samples and demonstrate that we can successfully deactivate 89.56% of evasive malware samples and the variants of ransomware (e.g., WannaCry and Locky) with little or no impact on the most commonly used benign software. Our evaluation also shows that SCARECROW is able to steer state-of-the-art analysis environment fingerprinting techniques so that end-user execution environments with SCARECROW and malware analysis environments with SCARECROW become indistinguishable.
Jialong Zhang 0001, Zhongshu Gu, Jiyong Jang, Dhilung Kirat, Marc Ph. Stoecklin, Xiaokui Shu, Heqing Huang 0001
DSN3
2019 Topology-Aware Hashing for Effective Control Flow Graph Similarity Analysis
Jiyong Jang, Xinming Ou
SecureComm (1)2
2018 Threat Intelligence Computing
abstract
Cyber threat hunting is the process of proactively and iteratively formulating and validating threat hypotheses based on security-relevant observations and domain knowledge. To facilitate threat hunting tasks, this paper introduces threat intelligence computing as a new methodology that models threat discovery as a graph computation problem. It enables efficient programming for solving threat discovery problems, equipping threat hunters with a suite of potent new tools for agile codifications of threat hypotheses, automated evidence mining, and interactive data inspection capabilities. A concrete realization of a threat intelligence computing platform is presented through the design and implementation of a domain-specific graph language with interactive visualization support and a distributed graph database. The platform was evaluated in a two-week DARPA competition for threat detection on a test bed comprising a wide variety of systems monitored in real time. During this period, sub-billion records were produced, streamed, and analyzed, dozens of threat hunting tasks were dynamically planned and programmed, and attack campaigns with diverse malicious intent were discovered. The platform exhibited strong detection and analytics capabilities coupled with high efficiency, resulting in a leadership position in the competition. Additional evaluations on comprehensive policy reasoning are outlined to demonstrate the versatility of the platform and the expressiveness of the language.
Xiaokui Shu, Frederico Araujo, Douglas Lee Schales, Marc Ph. Stoecklin, Jiyong Jang, Heqing Huang 0001, Josyula R. Rao
CCS5
2018 Protecting Intellectual Property of Deep Neural Networks with Watermarking
abstract
Deep learning technologies, which are the key components of state-of-the-art Artificial Intelligence (AI) services, have shown great success in providing human-level capabilities for a variety of tasks, such as visual analysis, speech recognition, and natural language processing and etc. Building a production-level deep learning model is a non-trivial task, which requires a large amount of training data, powerful computing resources, and human expertises. Therefore, illegitimate reproducing, distribution, and the derivation of proprietary deep learning models can lead to copyright infringement and economic harm to model creators. Therefore, it is essential to devise a technique to protect the intellectual property of deep learning models and enable external verification of the model ownership.
Jialong Zhang 0001, Zhongshu Gu, Jiyong Jang, Hui Wu 0006, Marc Ph. Stoecklin, Heqing Huang 0001, Ian M. Molloy
AsiaCCS3
2018 Error-Sensor: Mining Information from HTTP Error Traffic for Malware Intelligence
Jialong Zhang 0001, Jiyong Jang, Guofei Gu, Marc Ph. Stoecklin, Xin Hu 0001
RAID2
2017 Android Malware Clustering Through Malicious Payload Mining
Jiyong Jang, Xin Hu 0001, Xinming Ou
RAID2
2016 Detecting Malicious Exploit Kits using Tree-based Similarity Searches
Teryl Taylor, Xin Hu 0001, Ting Wang 0006, Jiyong Jang, Marc Ph. Stoecklin, Fabian Monrose, Reiner Sailer
CODASPY4
2016 BAYWATCH: Robust Beaconing Detection to Identify Infected Hosts in Large-Scale Enterprise Networks
abstract
Sophisticated cyber security threats, such as advanced persistent threats, rely on infecting end points within a targeted security domain and embedding malware. Typically, such malware periodically reaches out to the command and control infrastructures controlled by adversaries. Such callback behavior, called beaconing, is challenging to detect as (a) detection requires long-term temporal analysis of communication patterns at several levels of granularity, (b) malware authors employ various strategies to hide beaconing behavior, and (c) it is also employed by legitimate applications (such as updates checks). In this paper, we develop a comprehensive methodology to identify stealthy beaconing behavior from network traffic observations. We use an 8-step filtering approach to iteratively refine and eliminate legitimate beaconing traffic and pinpoint malicious beaconing cases for in-depth investigation and takedown. We provide a systematic evaluation of our core beaconing detection algorithm and conduct a large-scale evaluation of web proxy data (more than 30 billion events) collected over a 5-month period at a corporate network comprising over 130,000 end-user devices. Our findings indicate that our approach reliably exposes malicious beaconing behavior, which may be overlooked by traditional security mechanisms.
Xin Hu 0001, Jiyong Jang, Marc Ph. Stoecklin, Ting Wang 0006, Douglas Lee Schales, Dhilung Kirat, Josyula R. Rao
DSN2
2016 BotMeter: Charting DGA-Botnet Landscapes in Large Networks
abstract
Recent years have witnessed a rampant use of domain generation algorithms (DGAs) in major botnet crimewares, which tremendously strengthens a botnet's capability to evade detection or takedown. Despite a plethora of existing studies on detecting DGA-generated domains in DNS traffic, remediating such threats still relies on vetting the DNS behavior of each individual device. Yet, in large networks featuring complicated DNS infrastructures, we often lack the capability or the resource to exhaustively investigate every part of the networks to identify infected devices in a timely manner. It is therefore of great interest to first assess the population distribution of DGA-bots inside the networks and to prioritize the remediation efforts. In this paper, we present BotMeter, a novel tool that accurately charts the DGA-bot population landscapes in large networks. Specifically, we embrace the prevalent yet challenging setting of hierarchical DNS infrastructures with caching and forwarding mechanisms enabled, whereas DNS traffic is observable only at certain upper-level vantage points. We establish a new taxonomy of DGAs that captures their characteristic DNS dynamics. This allows us to develop a rich library of rigorous analytical models to describe the complex relationships between bot populations and DNS lookups observed at vantage points. We provide results from extensive empirical studies using both synthetic data and real DNS traces to validate the efficacy of BotMeter.
Ting Wang 0006, Xin Hu 0001, Jiyong Jang, Shouling Ji, Marc Ph. Stoecklin, Teryl Taylor
ICDCS3
2016 Hunting for invisibility: Characterizing and detecting malicious web infrastructures through server visibility analysis
abstract
Nowadays, cyber criminals often build web infrastructures rather than a single server to conduct their malicious activities. In order to continue their malevolent activities without being detected, cyber criminals make efforts to conceal the core servers (e.g., C&C servers, exploit servers, and drop-zone servers) in the malicious web infrastructure. Such deliberate invisibility of those concealed malicious servers, however, makes them particularly distinguishable from benign web servers that are usually promoted to be public. In this paper, we conduct the first large-scale measurement study to investigate the visibility of both malicious and benign servers. From our intensive analysis of over 100,000 benign servers, 45,000 malicious servers and 40,000 redirections, we identify a set of distinct features of malicious web infrastructures from their locations, structures, roles, and relationships perspectives, and propose a lightweight yet effective detection system called VisHunter. VisHunter identifies malicious redirections from visible servers to invisible servers at the entryway of malicious web infrastructures. We evaluate VisHunter on both online public data and large-scale enterprise network traffic, and demonstrate that VisHunter can achieve an average 96.2% detection rate with only 0.9% false positive rate on the real enterprise network traffic.
Jialong Zhang 0001, Xin Hu 0001, Jiyong Jang, Ting Wang 0006, Guofei Gu, Marc Ph. Stoecklin
INFOCOM3
2015 The Dropper Effect: Insights into Malware Distribution with Downloader Graph Analytics
abstract
Malware remains an important security threat, as miscreants continue to deliver a variety of malicious programs to hosts around the world. At the heart of all the malware delivery techniques are executable files (known as downloader trojans or droppers) that download other malware. Because the act of downloading software components from the Internet is not inherently malicious, benign and malicious downloaders are difficult to distinguish based only on their content and behavior. In this paper, we introduce the downloader-graph abstraction, which captures the download activity on end hosts, and we explore the growth patterns of benign and malicious graphs. Downloader graphs have the potential of exposing large parts of the malware download activity, which may otherwise remain undetected. By combining telemetry from anti-virus and intrusion-prevention systems, we reconstruct and analyze 19 million downloader graphs from 5 million real hosts. We identify several strong indicators of malicious activity, such as the growth rate, the diameter, and the Internet access patterns of downloader graphs. Building on these insights, we implement and evaluate a machine learning system for malware detection. Our system achieves a 96.0% true-positive rate, with a 1.0% false-positive rate, and detects malware an average of 9.24 days earlier than existing anti-virus products. We also perform an external validation by examining a sample of unlabeled files that our system detects as malicious, and we find that 41.41% are blocked by anti-virus products.
Bum Jun Kwon, Jayanta Mondal, Jiyong Jang, Leyla Bilge, Tudor Dumitras
CCS3
2015 FCCE: Highly scalable distributed Feature Collection and Correlation Engine for low latency big data analytics
abstract
In this paper, we present the design, architecture, and implementation of a novel analysis engine, called Feature Collection and Correlation Engine (FCCE), that finds correlations across a diverse set of data types spanning over large time windows with very small latency and with minimal access to raw data. FCCE scales well to collecting, extracting, and querying features from geographically distributed large data sets. FCCE has been deployed in a large production network with over 450,000 workstations for 3 years, ingesting more than 2 billion events per day and providing low latency query responses for various analytics. We explore two security analytics use cases to demonstrate how we utilize the deployment of FCCE on large diverse data sets in the cyber security domain: 1) detecting fluxing domain names of potential botnet activity and identifying all the devices in the production network querying these names, and 2) detecting advanced persistent threat infection. Both evaluation results and our experience with real-world applications show that FCCE yields superior performance over existing approaches, and excels in the challenging cyber security domain by correlating multiple features and deriving security intelligence.
Douglas Lee Schales, Xin Hu 0001, Jiyong Jang, Reiner Sailer, Marc Ph. Stoecklin, Ting Wang 0006
ICDE3
2015 Rateless and pollution-attack-resilient network coding
abstract
Consider the problem of reliable multicast over a network in the presence of adversarial errors. In contrast to traditional network error correction codes designed for a given network capacity and a given number of errors, we study an arguably more realistic setting that prior knowledge on the network and adversary parameters is not available. For this setting we propose efficient and throughput-optimal error correction schemes, provided that the source and terminals share randomness that is secret form the adversary. We discuss an application of cryptographic pseudorandom generators to efficiently produce the secret randomness, provided that a short key is shared between the source and terminals. Finally we present a secure key distribution scheme for our network setting.
Ting Wang 0006, Xin Hu 0001, Jiyong Jang, Theodoros Salonidis
ISIT4
2014 Lightweight authentication of freshness in outsourced key-value stores
abstract
Data outsourcing offers cost-effective computing power to manage massive data streams and reliable access to data. Data owners can forward their data to clouds, and the clouds provide data mirroring, backup, and online access services to end users. However, outsourcing data to untrusted clouds requires data authenticity and query integrity to remain in the control of the data owners and users.
Yuzhe Tang, Ting Wang 0006, Ling Liu 0001, Xin Hu 0001, Jiyong Jang
ACSAC5
2014 MUSE: asset risk scoring in enterprise network with mutually reinforced reputation propagation
abstract
Cyber security attacks are becoming ever more frequent and sophisticated. Enterprises often deploy several security protection mechanisms, such as anti-virus software, intrusion detection/prevention systems, and firewalls, to protect their critical assets against emerging threats. Unfortunately, these protection systems are typically ‘noisy’, e.g., regularly generating thousands of alerts every day. Plagued by false positives and irrelevant events, it is often neither practical nor cost-effective to analyze and respond to every single alert. The main challenges faced by enterprises are to extract important information from the plethora of alerts and to infer potential risks to their critical assets. A better understanding of risks will facilitate effective resource allocation and prioritization of further investigation. In this paper, we present MUSE, a system that analyzes a large number of alerts and derives risk scores by correlating diverse entities in an enterprise network. Instead of considering a risk as an isolated and static property pertaining only to individual users or devices, MUSE exploits a novel mutual reinforcement principle and models the dynamics of risk based on the interdependent relationship among multiple entities. We apply MUSE on real-world network traces and alerts from a large enterprise network consisting of more than 10,000 nodes and 100,000 edges. To scale up to such large graphical models, we formulate the algorithm using a distributed memory abstraction model that allows efficient in-memory parallel computations on large clusters. We implement MUSE on Apache Spark and demonstrate its efficacy in risk assessment and flexibility in incorporating a wide variety of datasets.
Xin Hu 0001, Ting Wang 0006, Marc Ph. Stoecklin, Douglas Lee Schales, Jiyong Jang, Reiner Sailer
EURASIP J. Inf. Secur.5
2013 Towards Automatic Software Lineage Inference
Jiyong Jang, Maverick Woo, David Brumley
USENIX Security Symposium1
2012 ReDeBug: Finding Unpatched Code Clones in Entire OS Distributions
abstract
Programmers should never fix the same bug twice. Unfortunately this often happens when patches to buggy code are not propagated to all code clones. Unpatched code clones represent latent bugs, and for security-critical problems, latent vulnerabilities, thus are important to detect quickly. In this paper we present ReDeBug, a system for quickly finding unpatched code clones in OS-distribution scale code bases. While there has been previous work on code clone detection, ReDeBug represents a unique design point that uses a quick, syntax-based approach that scales to OS distribution-sized code bases that include code written in many different languages. Compared to previous approaches, ReDeBug may find fewer code clones, but gains scale, speed, reduces the false detection rate, and is language agnostic. We evaluated ReDeBug by checking all code from all packages in the Debian Lenny/Squeeze, Ubuntu Maverick/Oneiric, all Source Forge C and C++ projects, and the Linux kernel for unpatched code clones. ReDeBug processed over 2.1 billion lines of code at 700,000 LoC/min to build a source code database, then found 15,546 unpatched copies of known vulnerable code in currently deployed code by checking 376 Debian/Ubuntu security-related patches in 8 minutes on a commodity desktop machine. We show the real world impact of ReDeBug by confirming 145 real bugs in the latest version of Debian Squeeze packages.
Jiyong Jang, Abeer Agrawal, David Brumley
IEEE Symposium on Security and Privacy1
2011 BitShred: feature hashing malware for scalable triage and semantic analysis
abstract
The sheer volume of new malware found each day is growing at an exponential pace. This growth has created a need for automatic malware triage techniques that determine what malware is similar, what malware is unique, and why. In this paper, we present BitShred, a system for large-scale malware similarity analysis and clustering, and for automatically uncovering semantic inter- and intra-family relationships within clusters. The key idea behind BitShred is using feature hashing to dramatically reduce the high-dimensional feature spaces that are common in malware analysis. Feature hashing also allows us to mine correlated features between malware families and samples using co-clustering techniques. Our evaluation shows that BitShred speeds up typical malware triage tasks by up to 2,365x and uses up to 82x less memory on a single CPU, all with comparable accuracy to previous approaches. We also develop a parallelized version of BitShred, and demonstrate scalability within the Hadoop framework.
Jiyong Jang, David Brumley, Shobha Venkataraman
CCS1
2010 SplitScreen: Enabling Efficient, Distributed Malware Detection
Sang Kil Cha, Iulian Moraru, Jiyong Jang, John Truelove, David Brumley, David G. Andersen
NSDI3
2007 A Time-Based Key Management Protocol for Wireless Sensor Networks
Jiyong Jang, Taekyoung Kwon 0002, JooSeok Song
ISPEC1
2006 Improving Resiliency Using Capacity-Aware Multicast Tree in P2P-Based Streaming Environments
Eunseok Kim, Jiyong Jang, Sungyoung Park, Alan Sussman, Jae Soo Yoo
HPCC2