Euijin Choo

dblp:98/7002 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 MASER: Efficient Privacy-Preserving Cross-Silo Federated Learning with Multi-Key Homomorphic Encryption
Abdullah Al Omar, Euijin Choo, Omid Ardakanian
IEEE Big Data3
2025 3D-GSW: 3D Gaussian Splatting for Robust Watermarking
abstract
As 3D Gaussian Splatting (3D-GS) gains significant attention and its commercial usage increases, the need for watermarking technologies to prevent unauthorized use of the 3D-GS models and rendered images has become increasingly important. In this paper, we introduce a robust watermarking method for 3D-GS that secures copyright of both the model and its rendered images. Our proposed method remains robust against distortions in rendered images and model attacks while maintaining high rendering quality. To achieve these objectives, we present Frequency-Guided Densification (FGD), which removes 3D Gaussians based on their contribution to rendering quality, enhancing real-time rendering and the robustness of the message. FGD utilizes Discrete Fourier Transform to split 3D Gaussians in high-frequency areas, improving rendering quality. Furthermore, we employ a gradient mask for 3D Gaussians and design a wavelet-subband loss to enhance rendering quality. Our experiments show that our method embeds the message in the rendered images invisibly and robustly against various attacks, including model distortion. Our method achieves superior performance in both rendering quality and watermark robustness while improving real-time rendering efficiency. Project page: https: //kuai-lab.github.io/cvpr20253dgsw/
Youngdong Jang, Hyunje Park, Feng Yang 0008, Heeju Ko, Euijin Choo, Sangpil Kim
CVPR5
2025 Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection Models
Fujiao Ji, Kiho Lee, Hyungjoon Koo, Wenhao You, Euijin Choo, Hyoungshick Kim, Doowon Kim
USENIX Security Symposium5
2024 Unsupervised Parameter-free Outlier Detection using HDBSCAN* Outlier Profiles
abstract
In machine learning and data mining, outliers are data points that significantly differ from the dataset and often introduce irrelevant information that can induce bias in its statistics and models. Therefore, unsupervised methods are crucial to detect outliers if there is limited or no information about them. Global-Local Outlier Scores based on Hierarchies (GLOSH) is an unsupervised outlier detection method within HDBSCAN*, a state-of-the-art hierarchical clustering method. GLOSH estimates outlier scores for each data point by comparing its density to the highest density of the region they reside in the HDBSCAN* hierarchy. GLOSH may be sensitive to HDBSCAN*’s minptsparameter that influences density estimation. With limited knowledge about the data, choosing an appropriate minptsvalue beforehand is challenging as one or some minptsvalues may better represent the underlying cluster structure than others. Additionally, in the process of searching for "potential outliers", one has to define the number of outliers n a dataset has, which may be impractical and is often unknown. In this paper, we propose an unsupervised strategy to find the "best" minptsvalue, leveraging the range of GLOSH scores across minptsvalues to identify the value for which GLOSH scores can best identify outliers from the rest of the dataset. Moreover, we propose an unsupervised strategy to estimate a threshold for classifying points into inliers and (potential) outliers without the need to pre-define any value. Our experiments show that our strategies can automatically find the minptsvalue and threshold that yield the best or near best outlier detection results using GLOSH.
Kushankur Ghosh, Murilo Coelho Naldi, Jörg Sander 0001, Euijin Choo
IEEE Big Data4
2024 Poster: Advanced Features for Real-Time Website Fingerprinting Attacks on Tor
abstract
The Tor network has been identified as vulnerable to website fingerprinting (WF) attacks.Existing WF attacks have proven effective against the Tor network.However, prior research has mostly been limited to controlled experimental settings, leading to questions about the practicality of WF attacks in real-time environments.Recent advancements in feature engineering and machine learning aim to address this by exploring real-world scenarios, though they often overlook the preprocessing time required to design features from raw network traffic data.To tackle these issues, this research focuses on developing more efficient and high-performing feature vectors for WF attacks in real-time by analyzing previously successful feature vectors.The results indicate that advanced features, particularly those in a compact feature set, deliver competitive performance with reduced training times for real-time WF attacks.This study enhances our understanding of the feasibility of real-time WF attacks on Tor networks in practical settings and may inform future security improvements.
Andrew Booth, Euijin Choo, Doosung Hwang
CCS3
2023 DeviceWatch: A Data-Driven Network Analysis Approach to Identifying Compromised Mobile Devices with Graph-Inference
abstract
We propose to identify compromised mobile devices from a network administrator’s point of view. Intuitively, inadvertent users (and thus their devices) who download apps through untrustworthy markets are often lured to install malicious apps through in-app advertisements or phishing. We thus hypothesize that devices sharing similar apps would have a similar likelihood of being compromised, resulting in an association between a compromised device and its apps. We propose to leverage such associations to identify unknown compromised devices using the guilt-by-association principle. Admittedly, such associations could be relatively weak as it is hard, if not impossible, for an app to automatically download and install other apps without explicit user initiation. We describe how we can magnify such associations by carefully choosing parameters when applying graph-based inferences. We empirically evaluate the effectiveness of our approach on real datasets provided by a major mobile service provider. Specifically, we show that our approach achieves nearly 98% AUC (area under the ROC curve) and further detects as many as 6 ~ 7 times of new compromised devices not covered by the ground truth by expanding the limited knowledge on known devices. We show that the newly detected devices indeed present undesirable behavior in terms of leaking private information and accessing risky IPs and domains. We further conduct in-depth analysis of the effectiveness of graph inferences to understand the unique structure of the associations between mobile devices and their apps, and its impact on graph inferences, based on which we propose how to choose key parameters.
Euijin Choo, Mohamed Nabeel, Mashael Al Sabah, Issa M. Khalil, Ting Yu 0001, Wei Wang 0012
ACM Trans. Priv. Secur.1
2022 Content-Agnostic Detection of Phishing Domains using Certificate Transparency and Passive DNS
abstract
Existing phishing detection techniques mainly rely on blacklists or content-based analysis, which are not only evadable, but also exhibit considerable detection delays as they are reactive in nature. We observe through our deep dive analysis that artifacts of phishing are manifested in various sources of intelligence related to a domain even before its contents are online. In particular, we study various novel patterns and characteristics computed from viable sources of data including Certificate Transparency Logs, and passive DNS records. To compare benign and phishing domains, we construct thoroughly-verified realistic benign and phishing datasets. Our analysis shows clear differences between benign and phishing domains that can pave the way for content-agnostic approaches to predict phishing domains even before the contents of these webpages are up and running.
Mashael Al Sabah, Mohamed Nabeel, Yazan Boshmaf, Euijin Choo
RAID4
2022 SIRAJ: A Unified Framework for Aggregation of Malicious Entity Detectors
abstract
High-quality intelligence of Internet threat (e.g., malware files, malicious domains, phishing URLs and malicious IPs) are important for both security practitioners and the research community. Given the agility of attackers, the scale of the Internet, and the fast-evolving landscape of threats, one could not rely solely on a single source (such as an anti-malware engine or an IP blacklist) for obtaining accurate, up-to-date, and comprehensive threat analysis. Instead, we need to aggregate the analysis from multiple sources. However, it is non-trivial to do such aggregation effectively. A common practice is to label an indicator (malware, domains, URLs, etc.) as malicious if it is marked by a number of sources above an ad-hoc certain threshold. Often, this results in sub-optimal performance as it assumes that all sources are of similar quality/expertise, independent, and temporally stable, which unfortunately are often not true in practice. A natural alternative is to train a supervised machine learning model. However, this approach needs a sufficiently large amount of manually labeled ground truth, which is time-consuming to collect and has to be updated frequently, resulting in substantial recurring costs. In this paper, we propose SIRAJ, a novel framework for aggregating the detection output of various intelligence sources such as anti-malware engines. SIRAJ is based on the pretrain and fine-tune paradigm. Specifically, we use self-supervised learning-based approaches to learn a pre-trained embedding model that converts multi-source inputs into a high-dimensional embedding. The embeddings are learned through three carefully designed pretext tasks that imbue them with knowledge about dependencies between scanners and their temporal dynamics. The learned embeddings could be used for diverse downstream machine learning tasks. SIRAJ is designed to be general and can be used for diverse domains such as URLs, malware, and IPs. Further, SIRAJ works well even when there is limited to no labeled data available. Through extensive experiments, we show that our learned representations can produce results comparable to supervised methods while only requiring as little as 100 labeled samples. Importantly, the results show that SIRAJ accurately detects threat indicators much earlier than the baseline algorithms, a feat that is critical against short-lived indicators like Phishing URLs.
Saravanan Thirumuruganathan, Mohamed Nabeel, Euijin Choo, Issa M. Khalil, Ting Yu 0001
SP3
2021 Time-Window Based Group-Behavior Supported Method for Accurate Detection of Anomalous Users
abstract
Autoencoder-based anomaly detection methods have been used in identifying anomalous users from large-scale enterprise logs with the assumption that adversarial activities do not follow past habitual patterns. Most existing approaches typically build models by reconstructing single-day and individual-user behaviors. However, without capturing long-term signals and group-correlation signals, the models cannot identify low-signal yet long-lasting threats, and will wrongly report many normal users as anomalies on busy days, which, in turn, lead to high false positive rate. In this paper, we propose ACOBE, an Anomaly detection method based on COmpound BEhavior, which takes into consideration long-term patterns and group behaviors. ACOBE leverages a novel behavior representation and an ensemble of deep autoencoders and produces an ordered investigation list. Our evaluation shows that ACOBE outperforms prior work by a large margin in terms of precision and recall, and our case study demonstrates that ACOBE is applicable in practice for cyberattack detection.
Lun-Pin Yuan, Euijin Choo, Ting Yu 0001, Issa M. Khalil, Sencun Zhu
DSN2
2017 Detecting opinion spammer groups and spam targets through community discovery and sentiment analysis
abstract
In this paper we investigate on detecting opinion spammer groups through analyzing how users interact with each other. More specifically, our approaches are based on 1) discovering strong vs. weak implicit communities by mining user interaction patterns, and 2) revealing positive vs. negative communities through sentiment analysis on user interactions. Through extensive experiments over various datasets collected from Amazon, we found that the discovered strong, positive communities are significantly more likely to be opinion spammer groups than other communities. Interestingly, while our approach focused mainly on the characteristics of user interactions, it is comparable to the state of the art content-based classifier that mainly uses various content-based features extracted from user reviews. More importantly, we argue that our approach can be more robust than the latter in that if spammers superficially alter their review contents, our approach can still reliably identify them while the content-based approaches may fail.
Euijin Choo, Ting Yu 0001, Min Chi
J. Comput. Secur.1
2015 Detecting Opinion Spammer Groups Through Community Discovery and Sentiment Analysis
Euijin Choo, Ting Yu 0001, Min Chi
DBSec1
2014 COMPARS: toward an empirical approach for comparing the resilience of reputation systems
abstract
Reputation is a primary mechanism for trust management in decentralized systems. Many reputation-based trust functions have been proposed in the literature. However, picking the right trust function for a given decentralized system is a non-trivial task. One has to consider and balance a variety of factors, including computation and communication costs, scalability and resilience to manipulations by attackers. Although the former two are relatively easy to evaluate, the evaluation of resilience of trust functions is challenging. Most existing work bases evaluation on static attack models, which is unrealistic as it fails to reflect the adaptive nature of adversaries (who are often real human users rather than simple computing agents).
Euijin Choo, Jianchun Jiang, Ting Yu 0001
CODASPY1
2014 Revealing and incorporating implicit communities to improve recommender systems
abstract
Social connections often have a significant influence on personal decision making. Researchers have proposed novel recommender systems that take advantage of social relationship information to improve recommendations. These systems, while promising, are often hindered in practice. Existing social networks such as Facebook are not designed for recommendations and thus contain many irrelevant relationships. Many recommendation platforms such as Amazon often do not permit users to establish explicit social relationships. And direct integration of social and commercial systems raises privacy concerns.
Euijin Choo, Ting Yu 0001, Min Chi, Yan Lindsay Sun
EC1