Se Eun Oh

dblp:142/8371 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0000-7100-4613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 9 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TabCL: Continual Malware Classification with Tabular-Aware Generation
Haeseung Jeon, AHyun Ji, Aritran Piplai, Mohammad Saidur Rahman 0002, Se Eun Oh
PAKDD (3)6
2026 RoFiRe: Robust Website Fingerprinting on Real-World Tor Traffic via Improved Augmentation and Normalization
abstract
Website Fingerprinting (WF) attacks infer visited websites from encrypted traffic patterns, threatening the anonymity of Tor users. While deep learning has improved WF accuracy on synthetic datasets, its effectiveness on real Tor traffic remains unclear. We present RoFiRe, a novel WF model explicitly designed for real Tor exit traffic. It employs dynamic window-level augmentation and RMS-based pre-normalization tailored to irregular burst patterns and variable-length traces observed in the realistic GTT23 dataset, improving data efficiency and robustness. Across standard WF, concept-drift, and real-world open-world evaluations on GTT23, RoFiRe consistently outperforms state-of-the-art baselines by up to 9%. This study presents the first comprehensive analysis of WF on real Tor traffic, showing that WF models incorporating augmentation remain effective under realistic, data-limited conditions and can pose significant privacy risks.
Haeseung Jeon, Nate Mathews, Hosung Kang, Se Eun Oh
WWW5
2025 MalCL: Leveraging GAN-Based Generative Replay to Combat Catastrophic Forgetting in Malware Classification
abstract
Continual Learning (CL) for malware classification tackles the rapidly evolving nature of malware threats and the frequent emergence of new types. Generative Replay (GR)-based CL systems utilize a generative model to produce synthetic versions of past data, which are then combined with new data to retrain the primary model. Traditional machine learning techniques in this domain often struggle with catastrophic forgetting, where a model's performance on old data degrades over time. In this paper, we introduce a GR-based CL system that employs Generative Adversarial Networks (GANs) with feature matching loss to generate high-quality malware samples. Additionally, we implement innovative selection schemes for replay samples based on the model’s hidden representations. Our comprehensive evaluation across Windows and Android malware datasets in a class-incremental learning scenario -- where new classes are introduced continuously over multiple tasks -- demonstrates substantial performance improvements over previous methods. For example, our system achieves an average accuracy of 55% on Windows malware samples, significantly outperforming other GR-based models by 28%. This study provides practical insights for advancing GR-based malware classification systems.
AHyun Ji, Minji Park, Mohammad Saidur Rahman 0002, Se Eun Oh
AAAI5
2025 Enhancing Search Privacy on Tor: Advanced Deep Keyword Fingerprinting Attacks and BurstGuard Defense
Chaiwon Hwang, Haeseung Jeon, Jiwoo Hong, Hosung Kang, Nate Mathews, Goun Kim, Se Eun Oh
AsiaCCS7
2024 WhisperVoiceTrace: A Comprehensive Analysis of Voice Command Fingerprinting
abstract
Smart speakers such as Amazon Alexa or Google Home, have significantly enhanced the convenience and efficiency of our daily lives, leading to widespread adoption globally. Subsequently, the voice commands directed at these devices contain a wealth of private information, encompassing users' daily routines, details about clinic visits, and even shopping habits.
Minji Jo, Jiwoo Hong, Hosung Kang, Nate Mathews, Se Eun Oh
AsiaCCS6
2024 DeTorrent: An Adversarial Padding-only Traffic Analysis Defense
abstract
While anonymity networks like Tor aim to protect the privacy of their users, they are vulnerable to traffic analysis attacks such as Website Fingerprinting (WF) and Flow Correlation (FC). Recent implementations of WF and FC attacks, such as Tik-Tok and DeepCoFFEA, have shown that the attacks can be effectively carried out, threatening user privacy. Consequently, there is a need for effective traffic analysis defense. There are a variety of existing defenses, but most are either ineffective, incur high latency and bandwidth overhead, or require additional infrastructure. As a result, we aim to design a traffic analysis defense that is efficient and highly resistant to both WF and FC attacks. We propose DeTorrent, which uses competing neural networks to generate and evaluate traffic analysis defenses that insert 'dummy' traffic into real traffic flows. DeTorrent operates with moderate overhead and without delaying traffic. In a closed-world WF setting, it reduces an attacker's accuracy by 61.5%, a reduction 10.5% better than the next-best padding-only defense. Against the state-of-the-art FC attacker, DeTorrent reduces the true positive rate for a .00001 false positive rate to about .12, which is less than half that of the next-best defense. We also demonstrate DeTorrent's practicality by deploying it alongside the Tor network and find that it maintains its performance when applied to live traffic.
James K. Holland, Jason Carpenter, Se Eun Oh, Nicholas Hopper
Proc. Priv. Enhancing Technol.3
2023 SoK: A Critical Evaluation of Efficient Website Fingerprinting Defenses
abstract
Recent website fingerprinting attacks have been shown to achieve very high performance against traffic through Tor. These attacks allow an adversary to deduce the website a Tor user has visited by simply eavesdropping on the encrypted communication. This has consequently motivated the development of many defense strategies that obfuscate traffic through the addition of dummy packets and/or delays. The efficacy and practicality of many of these recent proposals have yet to be scrutinized in detail. In this study, we re-evaluate nine recent defense proposals that claim to provide adequate security with low-overheads using the latest Deep Learning-based attacks. Furthermore, we assess the feasibility of implementing these defenses within the current confines of Tor. To this end, we additionally provide the first on-network implementation of the DynaFlow defense to better assess its real-world utility.
Nate Mathews, James K. Holland, Se Eun Oh, Mohammad Saidur Rahman 0002, Nicholas Hopper, Matthew Wright 0001
SP3
2022 DeepCoFFEA: Improved Flow Correlation Attacks on Tor via Metric Learning and Amplification
abstract
End-to-end flow correlation attacks are among the oldest known attacks on low-latency anonymity networks, and are treated as a core primitive for traffic analysis of Tor. However, despite recent work showing that individual flows can be correlated with high accuracy, the impact of even these state-of-the-art attacks is questionable due to a central drawback: their pairwise nature, requiring comparison between N2pairs of flows to deanonymize N users. This results in a combinatorial explosion in computational requirements and an asymptotically declining base rate, leading to either high numbers of false positives or vanishingly small rates of successful correlation. In this paper, we introduce a novel flow correlation attack, DeepCoFFEA, that combines two ideas to overcome these drawbacks. First, DeepCoFFEA uses deep learning to train a pair of feature embedding networks that respectively map Tor and exit flows into a single low-dimensional space where correlated flows are similar; pairs of embedded flows can be compared at lower cost than pairs of full traces. Second, DeepCoFFEA uses amplification, dividing flows into short windows and using voting across these windows to significantly reduce false positives; the same embedding networks can be used with an increasing number of windows to independently lower the false positive rate. We conduct a comprehensive experimental analysis showing that DeepCoFFEA significantly outperforms state-of-the-art flow correlation attacks on Tor, e.g. 93% true positive rate versus at most 13% when tuned for high precision, with two orders of magnitude speedup over prior work. We also consider the effects of several potential countermeasures on DeepCoFFEA, finding that existing lightweight defenses are not sufficient to secure anonymity networks from this threat.
Se Eun Oh, Taiji Yang, Nate Mathews, James K. Holland, Mohammad Saidur Rahman 0002, Nicholas Hopper, Matthew Wright 0001
SP1
2021 GANDaLF: GAN for Data-Limited Fingerprinting
abstract
Abstract We introduce Generative Adversarial Networks for Data-Limited Fingerprinting (GANDaLF), a new deep-learning-based technique to perform Website Fingerprinting (WF) on Tor traffic. In contrast to most earlier work on deep-learning for WF, GANDaLF is intended to work with few training samples, and achieves this goal through the use of a Generative Adversarial Network to generate a large set of “fake” data that helps to train a deep neural network in distinguishing between classes of actual training data. We evaluate GANDaLF in low-data scenarios including as few as 10 training instances per site, and in multiple settings, including fingerprinting of website index pages and fingerprinting of non-index pages within a site. GANDaLF achieves closed-world accuracy of 87% with just 20 instances per site (and 100 sites) in standard WF settings. In particular, GANDaLF can outperform Var-CNN and Triplet Fingerprinting (TF) across all settings in subpage fingerprinting. For example, GANDaLF outperforms TF by a 29% margin and Var-CNN by 38% for training sets using 20 instances per site.
Se Eun Oh, Nate Mathews, Mohammad Saidur Rahman 0002, Matthew Wright 0001, Nicholas Hopper
Proc. Priv. Enhancing Technol.1
2019 p1-FP: Extraction, Classification, and Prediction of Website Fingerprints with Deep Learning
abstract
Abstract Recent advances in Deep Neural Network (DNN) architectures have received a great deal of attention due to their ability to outperform state-of-the-art machine learning techniques across a wide range of application, as well as automating the feature engineering process. In this paper, we broadly study the applicability of deep learning to website fingerprinting. First, we show that unsupervised DNNs can generate lowdimensional informative features that improve the performance of state-of-the-art website fingerprinting attacks. Second, when used as classifiers, we show that they can exceed performance of existing attacks across a range of application scenarios, including fingerprinting Tor website traces, fingerprinting search engine queries over Tor, defeating fingerprinting defenses, and fingerprinting TLS-encrypted websites. Finally, we investigate which site-level features of a website influence its fingerprintability by DNNs.
Se Eun Oh, Saikrishna Sunkam, Nicholas Hopper
Proc. Priv. Enhancing Technol.1
2017 Fingerprinting Keywords in Search Queries over Tor
abstract
Abstract Search engine queries contain a great deal of private and potentially compromising information about users. One technique to prevent search engines from identifying the source of a query, and Internet service providers (ISPs) from identifying the contents of queries is to query the search engine over an anonymous network such as Tor. In this paper, we study the extent to which Website Fingerprinting can be extended to fingerprint individual queries or keywords to web applications, a task we call Keyword Fingerprinting (KF). We show that by augmenting traffic analysis using a two-stage approach with new task-specific feature sets, a passive network adversary can in many cases defeat the use of Tor to protect search engine queries. We explore three popular search engines, Google, Bing, and Duckduckgo, and several machine learning techniques with various experimental scenarios. Our experimental results show that KF can identify Google queries containing one of 300 targeted keywords with recall of 80% and precision of 91%, while identifying the specific monitored keyword among 300 search keywords with accuracy 48%. We also further investigate the factors that contribute to keyword fingerprintability to understand how search engines and users might protect against KF.
Se Eun Oh, Nicholas Hopper
Proc. Priv. Enhancing Technol.1
2014 Privacy-preserving audit for broker-based health information exchange
abstract
Developments in health information technology have encouraged the establishment of distributed systems known as Health Information Exchanges (HIEs) to enable the sharing of patient records between institutions. In many cases, the parties running these exchanges wish to limit the amount of information they are responsible for holding because of sensitivities about patient information. Hence, there is an interest in broker-based HIEs that keep limited information in the exchange repositories. However, it is essential to audit these exchanges carefully due to risks of inappropriate data sharing. In this paper, we consider some of the requirements and present a design for auditing broker-based HIEs in a way that controls the information available in audit logs and regulates their release for investigations. Our approach is based on formal rules for audit and the use of Hierarchical Identity-Based Encryption (HIBE) to support staged release of data needed in audits and a balance between automated and manual reviews. We test our methodology via an extension of a standard for auditing HIEs called the Audit Trail and Node Authentication Profile (ATNA) protocol.
Se Eun Oh, Ji Young Chun, Limin Jia 0001, Deepak Garg 0001, Carl A. Gunter, Anupam Datta
CODASPY1