EDBT 2026 Demo / reviewers in the wild / expert
Youngja Park
dblp:65/6400
· DBLP profile ↗
37ranked-venue papers
14as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-authorSecurity and privacy · 12 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-authorDatabases, data management, data science and information retrieval · 7 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Looking Beyond IoCs: Automatically Extracting Attack Patterns from External CTIabstractPublic and commercial organizations extensively share cyberthreat intelligence (CTI) to prepare systems to defend against existing and emerging cyberattacks. However, traditional CTI has primarily focused on tracking known threat indicators such as IP addresses and domain names, which may not provide long-term value in defending against evolving attacks. To address this challenge, we propose to use more robust threat intelligence signals called attack patterns. LADDER is a knowledge extraction framework that can extract text-based attack patterns from CTI reports at scale. The framework characterizes attack patterns by capturing the phases of an attack in Android and enterprise networks and systematically maps them to the MITRE ATT&CK pattern framework. LADDER can be used by security analysts to determine the presence of attack vectors related to existing and emerging threats, enabling them to prepare defenses proactively. We also present several use cases to demonstrate the application of LADDER in real-world scenarios. Finally, we provide a new, open-access benchmark malware dataset to train future cyberthreat intelligence models. Md Tanvirul Alam, Dipkamal Bhusal, Youngja Park, Nidhi Rastogi |
RAID | 3 |
| 2022 | Backdoor smoothing: Demystifying backdoor attacks on deep neural networks
Kathrin Grosse, Taesung Lee, Battista Biggio, Youngja Park, Michael Backes 0001, Ian M. Molloy |
Comput. Secur. | 4 |
| 2021 | An Ontology-driven Knowledge Graph for Android MalwareabstractWe present MalONT2.0 -- an ontology for malware threat intelligence [4]. New classes (attack patterns, infrastructural resources to enable attacks, malware analysis to incorporate static analysis, and dynamic analysis of binaries) and relations have been added following a broadened scope of core competency questions. MalONT2.0 allows researchers to extensively capture all requisite classes and relations that gather semantic and syntactic characteristics of an android malware attack. This ontology forms the basis for the malware threat intelligence knowledge graph, MalKG, which we exemplify using three different, non-overlapping demonstrations. Malware features have been extracted from openCTI reports on android threat intelligence shared on the Internet and written in the form of unstructured text. Some of these sources are blogs, threat intelligence reports, tweets, and news articles. The smallest unit of information that captures malware features is written as triples comprising head and tail entities, each connected with a relation. In the poster and demonstration, we discuss MalONT2.0 and MalKG. Christian Ryan, Sharmishtha Dutta, Youngja Park, Nidhi Rastogi |
CCS | 3 |
| 2019 | Supervising Unsupervised Open Information Extraction ModelsabstractArpita Roy, Youngja Park, Taesung Lee, Shimei Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Arpita Roy, Youngja Park, Taesung Lee, Shimei Pan |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Incorporating Domain Knowledge in Learning Word EmbeddingabstractWord embedding is a Natural Language Processing (NLP) technique that automatically maps words from a vocabulary to vectors of real numbers in an embedding space. It has been widely used in recent years to boost the performance of a variety of NLP tasks such as named entity recognition, syntactic parsing and sentiment analysis. Classic word embedding methods such as Word2Vec and GloVe work well when they are given a large text corpus. When the input texts are sparse as in many specialized domains (e.g., cybersecurity), these methods often fail to produce high-quality vectors. In this paper, we describe a novel method, called Annotation Word Embedding (AWE), to train domain-specific word embeddings from sparse texts. Our method is generic and can leverage diverse types of domain knowledge such as domain vocabulary, semantic relations and attribute specifications. Specifically, our method encodes diverse types of domain knowledge as text annotations and incorporates the annotations in word embedding. We have evaluated AWE in two cybersecurity applications: identifying malware aliases and identifying relevant Common Vulnerabilities and Exposures (CVEs). Our evaluation results have demonstrated the effectiveness of our method over state-of-the-art baselines. Arpita Roy, Youngja Park, Shimei Pan |
ICTAI | 2 |
| 2019 | Unsupervised Sentence Embedding Using Document Structure-Based Context
Taesung Lee, Youngja Park |
ECML/PKDD (2) | 2 |
| 2016 | DinTucker: Scaling Up Gaussian Process Models on Large Multidimensional ArraysabstractTensor decomposition methods are effective tools for modelling multidimensional array data (i.e., tensors). Among them, nonparametric Bayesian models, such as Infinite Tucker Decomposition (InfTucker), are more powerful than multilinear factorization approaches, including Tucker and PARAFAC, and usually achieve better predictive performance. However, they are difficult to handle massive data due to a prohibitively high training cost. To address this limitation, we propose Distributed infinite Tucker (DinTucker), a new hierarchical Bayesian model that enables local learning of InfTucker on subarrays and global information integration from local results. We further develop a distributed stochastic gradient descent algorithm, coupled with variational inference for model estimation. In addition, the connection between DinTucker and InfTucker is revealed in terms of model evidence. Experiments demonstrate that DinTucker maintains the predictive accuracy of InfTucker and is scalable on massive data: On multidimensional arrays with billions of elements from two real-world applications, DinTucker achieves significantly higher prediction accuracy with less training time, compared with the state-of-the-art large-scale tensor decomposition method, GigaTensor. Shandian Zhe, Yuan Qi 0001, Youngja Park, Zenglin Xu, Ian M. Molloy, Suresh Chari |
AAAI | 3 |
| 2016 | Comparing Password Ranking Algorithms on Real-World Password Datasets
Weining Yang, Ninghui Li 0001, Ian M. Molloy, Youngja Park, Suresh Chari |
ESORICS (1) | 4 |
| 2016 | Tri-Modularization of Firewall PoliciesabstractFirewall policies are notorious for having misconfiguration errors which can defeat its intended purpose of protecting hosts in the network from malicious users. We believe this is because today's firewall policies are mostly monolithic. Inspired by ideas from modular programming and code refactoring, in this work we introduce three kinds of modules: primary, auxiliary, and template, which facilitate the refactoring of a firewall policy into smaller, reusable, comprehensible, and more manageable components. We present algorithms for generating each of the three modules for a given legacy firewall policy. We also develop ModFP, an automated tool for converting legacy firewall policies represented in access control list to their modularized format. With the help of ModFP, when examining several real-world policies with sizes ranging from dozens to hundreds of rules, we were able to identify subtle errors. Haining Chen, Omar Chowdhury, Ninghui Li 0001, Warut Khern-am-nuai, Suresh Chari, Ian M. Molloy, Youngja Park |
SACMAT | 7 |
| 2015 | Scalable Nonparametric Multiway Data AnalysisabstractMultiway data analysis deals with multiway arrays, i.e., tensors, and the goal is twofold: predicting missing entries by modeling the interactions between array elements and discovering hidden patterns, such as clusters or communities in each mode. Despite the success of existing tensor factorization approaches, they are either unable to capture nonlinear interactions, or computationally expensive to handle massive data. In addition, most of the existing methods lack a principled way to discover latent clusters, which is important for better understanding of the data. To address these issues, we propose a scalable nonparametric tensor decomposition model. It employs Dirichlet process mixture (DPM) prior to model the latent clusters; it uses local Gaussian processes (GPs) to capture nonlinear relationships and to improve scalability. An efficient online variational Bayes Expectation-Maximization algorithm is proposed to learn the model. Experiments on both synthetic and real-world data show that the proposed model is able to discover latent clusters with higher prediction accuracy than competitive methods. Furthermore, the proposed model obtains significantly better predictive performance than the state-of-the-art large scale tensor decomposition algorithm, GigaTensor, on two large datasets with billions of entries. Shandian Zhe, Zenglin Xu, Xinqi Chu, Yuan Qi 0001, Youngja Park |
AISTATS | 5 |
| 2015 | Learning from Others: User Anomaly Detection Using Anomalous Samples from Other Users
Youngja Park, Ian M. Molloy, Suresh Chari, Zenglin Xu, Christopher Gates 0002, Ninghui Li 0001 |
ESORICS (2) | 1 |
| 2014 | Detecting Insider Information Theft Using Features from File Access Logs
Christopher Gates 0002, Ninghui Li 0001, Zenglin Xu, Suresh Chari, Ian M. Molloy, Youngja Park |
ESORICS (2) | 6 |
| 2014 | Hetero-Labeled LDA: A Partially Supervised Topic Model with Heterogeneous Labels
Dongyeop Kang, Youngja Park, Suresh Chari |
ECML/PKDD (1) | 2 |
| 2014 | PAKDD'12 best paper: generating balanced classifier-independent training samples from unlabeled data
Youngja Park, Zijie Qi, Suresh Chari, Ian M. Molloy |
Knowl. Inf. Syst. | 1 |
| 2013 | Estimating Asset Sensitivity by Profiling Users
Youngja Park, Christopher Gates 0002, Stephen C. Gates |
ESORICS | 1 |
| 2013 | A bigData platform for analytics on access control policies and logsabstractRelying on an access control security policy alone to protect valuable resources is a dangerous practice. Prudent security must engage in other risk management and mitigation techniques to rapidly detect and recover from breaches. In reality, many security policies are either wrong, containing errors, or are misused and abused by malicious employees or compromised accounts; not all granted access is desirable. A popular approach to mitigate against these and other residual threats is to monitor applications to detect misuse and abuse of credentials in near real-time. Suresh Chari, Ted Habeck, Ian M. Molloy, Youngja Park, Wilfried Teiken |
SACMAT | 4 |
| 2013 | Ensuring continuous compliance through reconciling policy with usageabstractOrganizations rarely define formal security properties or policies for their access control systems, often choosing to react to changing needs. This paper addresses the problem of reconciling entitlement usage with configured policies for multiple objectives: policy optimization and risk mitigation. Policies should remain up-to-date, maintaining least privilege, and using unambiguous constructs that reduce administrative stress. Suresh Chari, Ian M. Molloy, Youngja Park, Wilfried Teiken |
SACMAT | 3 |
| 2013 | Predicting Network Activity from High Throughput MetabolomicsabstractThe functional interpretation of high throughput metabolomics by mass spectrometry is hindered by the identification of metabolites, a tedious and challenging task. We present a set of computational algorithms which, by leveraging the collective power of metabolic pathways and networks, predict functional activity directly from spectral feature tables without a priori identification of metabolites. The algorithms were experimentally validated on the activation of innate immune cells. Shuzhao Li, Youngja Park, Sai Duraisingham, Frederick H. Strobel, Nooruddin Khan, Quinlyn A. Soltow, Dean P. Jones, Bali Pulendran |
PLoS Comput. Biol. | 2 |
| 2012 | Business email classification using incremental subspace learning
Youngja Park, He Yuan Huang |
ICPR | 2 |
| 2012 | Generating Balanced Classifier-Independent Training Samples from Unlabeled Data
Youngja Park, Zijie Qi, Suresh Chari, Ian M. Molloy |
PAKDD (1) | 1 |
| 2012 | Generative models for access control policies: applications to role mining over logs with attributionabstractWe consider a fundamentally new approach to role and policy mining: finding RBAC models which reflect the observed usage of entitlements and the attributes of users. Such policies are interpretable, i.e., there is a natural explanation of why a role is assigned to a user and are conservative from a security standpoint since they are based on actual usage. Further, such "generative" models provide many other benefits including reconciliation with policies based on entitlements, detection of provisioning errors, as well as the detection of anomalous behavior. Our contributions include defining the fundamental problem as extensions of the well-known role mining problem, as well as providing several new algorithms based on generative machine learning models. Our algorithms find models which are causally associated with actual usage of entitlements and any arbitrary combination of user attributes when such information is available. This is the most natural process to provision roles, thus addressing a key usability issue with existing role mining algorithms. Ian M. Molloy, Youngja Park, Suresh Chari |
SACMAT | 2 |
| 2011 | Automatic Call Quality Monitoring Using Cost-Sensitive Classification
Youngja Park |
INTERSPEECH | 1 |
| 2011 | System for automatic estimation of data sensitivity with applications to access control and other applicationsabstractThe Enterprise Information Security Management (EISM) system aims to semi-automatically estimate the sensitivity of enterprise data through advanced content analysis and business process mining. We demonstrate a proof-of-concept of EISM that crawls all the files in a personal computer and estimates the sensitivity of individual files and the overall sensitivity level of the computer. The system can identify 11 different personally identifiable information (PII) types and 11 sensitive data categories, and estimate data sensitivity based on the identified sensitive information in the data. Furthermore, the tool produces the evidences of the discovered sensitive information including the surrounding context in the document to help users understand what kinds of sensitive information are stored in their computer. The evidences allow users can easily redact the sensitive information or move it to a more secure location. Thus, this system can be used as a privacy enhancing tool as well as a security tool. Youngja Park, Stephen C. Gates, Wilfried Teiken, Suresh Chari |
SACMAT | 1 |
| 2009 | Towards real-time measurement of customer satisfaction using automatically generated call transcriptsabstractCustomer satisfaction is a very important indicator of how successful a contact center is at providing services to the customers. Contact centers typically conduct a manual survey with a randomly selected group of customers to measure customer satisfaction. Manual customer satisfaction surveys, however, provide limited values due to high cost and the time lapse between the service and the survey. Youngja Park, Stephen C. Gates |
CIKM | 1 |
| 2009 | Low-cost call type classification for contact center calls using partial transcriptsabstractCall type classification and topic classification for contact center calls using automatically generated transcripts is not yet widely available mainly due to the high cost and low accuracy of call-center grade automatic speech transcription. To address these challenges, we examine if using only partial conversations yields accuracy comparable to using the entire customer-agent conversations. We exploit two interesting characteristics of call center calls. First, contact center calls are highly scripted following prescribed steps, and the customers problem or request (i.e., the determinant of the call type) is typically stated in the beginning of a call. Thus, using only the beginning of calls may be sufficient to determine the call type. Second, agents often more clearly repeat or rephrase what customers said, thus it may be sufficient to process only agents’ speech. Our experiments with 1,677 customer calls show that two partial transcripts comprising only the agents utterances and the first 40 speaker turns actually produce slightly higher classification accuracy than a transcript set comprising the entire conversations. In addition, using partial conversations can significantly reduce the cost for speech transcription. Youngja Park, Wilfried Teiken, Stephen C. Gates |
INTERSPEECH | 1 |
| 2009 | apLCMS - adaptive processing of high-resolution LC/MS dataabstractMOTIVATION: Liquid chromatography-mass spectrometry (LC/MS) profiling is a promising approach for the quantification of metabolites from complex biological samples. Significant challenges exist in the analysis of LC/MS data, including noise reduction, feature identification/ quantification, feature alignment and computation efficiency. RESULT: Here we present a set of algorithms for the processing of high-resolution LC/MS data. The major technical improvements include the adaptive tolerance level searching rather than hard cutoff or binning, the use of non-parametric methods to fine-tune intensity grouping, the use of run filter to better preserve weak signals and the model-based estimation of peak intensities for absolute quantification. The algorithms are implemented in an R package apLCMS, which can efficiently process large LC/ MS datasets. AVAILABILITY: The R package apLCMS is available at www.sph.emory.edu/apLCMS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tianwei Yu, Youngja Park, Jennifer M. Johnson, Dean P. Jones |
Bioinform. | 2 |
| 2008 | Semi-automated logging of contact center telephone callsabstractModern businesses use contact centers as a communication channel with users of their products and services. The largest factor in the expense of running a telephone contact center is the labor cost of its agents. IBM Research has built a new system, Contact-Center Agent Buddies (CAB), which is designed to help reduce the average handle time (AHT) for customer calls, thereby also reducing their cost. In this paper, we focus on the call logging subsystem, which helps agents reduce the time they spend documenting those calls. We built a Template CAB and a Call Logging CAB, using a pipeline consisting of audio capture of a telephone conversation, automatic speech recognition, text analysis, and log generation. We developed techniques for ASR text cleansing, including normalization of expressions and acronyms, domain terms, capitalization, and boundaries for sentences, paragraphs, and call segments. We found that simple heuristics suffice to generate high-quality logs from the normalized sentences. The pipeline yields a candidate call log which the agents can edit in less time than it takes them to generate call logs manually. Evaluation of the Call Logging CAB in an industrial contact center environment shows that it reduces the amount of time agents spend logging calls by at least 50% without compromising the quality of the resulting call documentation. Roy J. Byrd, Mary S. Neff, Wilfried Teiken, Youngja Park, Keh-Shin F. Cheng, Stephen C. Gates, Karthik Visweswariah |
CIKM | 4 |
| 2008 | Automatically constructing blue pages for characters in instructional videosabstractThis paper presents our recent work on automatically constructing blue pages for main characters in instructional videos. Specifically, a blue page is a personal profile which contains various types of information such as name, affiliation, portrait and voice that are specific to each individual. To accomplish this, we first extract various types of personal identity information from the video by analyzing multiple media cues; then we carefully correlate them with each other w.r.t. individual video characters based on an advanced video context analysis. When additional information sources are available, the constructed blue pages could be further enriched with advanced data mining and information extraction techniques. To validate the proposed ideas, we have carried out some preliminary experiments on a set of instructional videos with acceptable results obtained. Ying Li 0121, Youngja Park |
ICME | 2 |
| 2008 | An empirical analysis of word error rate and keyword error rateabstractThis paper studies the relationship between word error rate (WER) and keyword error rate (KER) in speech transcripts and their effect on the performance of speech analytics applications. Automatic speech recognition (ASR) systems are increasingly used as input for speech analytics, which raises the question of whether WER or KER is the more suitable performance metric for calibrating the ASR system. ASR systems are typically evaluated in terms of WER. Many speech analytics applications, Youngja Park, Siddharth Patwardhan, Karthik Visweswariah, Stephen C. Gates |
INTERSPEECH | 1 |
| 2008 | Genetic algorithm-based feature selection in high-resolution NMR spectra
Hyun-Woo Cho, Seoung Bum Kim, Myong Kee Jeong, Youngja Park, Thomas R. Ziegler, Dean P. Jones |
Expert Syst. Appl. | 4 |
| 2007 | Automatic call section segmentation for contact-center callsabstractThis paper presents a SVM (Support Vector Machine) classification system which divides contact-center call transcripts into "Greeting", "Question", "Refine", "Research", "Resolution", "Closing" and "Out-of-topic" sections. This call section segmentation is useful to improve search and retrieval functions and to provide more detailed statistics on calls. We use an off-the-shelf automatic speech recognition (ASR) system to generate call transcripts from recorded calls between customers and service representatives. Youngja Park |
CIKM | 1 |
| 2006 | MAGICAL demonstration: system for automated metadata generation for instructional contentabstractThe "Tools for Automatic Generation of Learning Object Metadata" project addresses the requirement of developing advanced distributed learning delivery architecture and services for a large US government agency. We have developed a Webbased system called MAGIC (Metadata Automated Generation for Instructional Content) to assist content authors and course developers in generating metadata for learning objects and information assets to enable wider reuse of these objects across departments and organizations. Using the MAGIC system, content authors review and edit automatically-generated metadata sufficient to register and describe their assets for use and discovery in current and future distributed learning applications complying with the ADL SCORM standard. Course developers can use the system to assist in the conversion of existing courses to SCORM format or in developing new SCORM courses. The MAGIC system includes software tools to analyze and extract descriptive metadata from instructional videos, training documents, and other information assets. The tools generate some of the most critical SCORM metadata completely automatically. Benefits of MAGIC include easier reuse and repurposing, improved interoperability, and more timely registration of content for use by course developers. In this paper, we describe the system architecture, analysis tools developed, and services supported. A live demonstration of the system illustrating several use cases of the system will be presented at the conference, with a discussion of results from user studies and evaluation of the system. Chitra Dorai, Robert G. Farrell, Amy Katriel, Galina Kofman, Ying Li 0121, Youngja Park |
ACM Multimedia | 6 |
| 2006 | Atomic topical segments detection for instructional videosabstractThis paper presents our latest work on structuring instructional videos into units of atomic topical segment so as to facilitate topic-based video browsing and offer efficient video authoring. Specifically, we developed a comprehensive text analysis component to first extract informative text cues such as keyword synonym set and sentence boundary information, from a video's transcript. These text cues are then applied with various audiovisual cues such as silence/music break and speech similarity, to identify topical segments. Early experiments carried out on collections of real data from targeted user communities have yielded good results, and the user feedback on using the generated topical segment information is very encouraging. Ying Li 0121, Youngja Park, Chitra Dorai |
ACM Multimedia | 2 |
| 2006 | Extracting Salient Keywords from Instructional Videos Using Joint Text, Audio and Visual Cues
Youngja Park, Ying Li 0121 |
HLT-NAACL | 1 |
| 2004 | GlossOnt: A Concept-focused Ontology Building Tool
Youngja Park |
KR | 1 |
| 2002 | Automatic Glossary Extraction: Beyond Terminology Identification
Youngja Park, Roy J. Byrd, Branimir Boguraev |
COLING | 1 |
| 2001 | Hybrid Text Mining for Finding Abbreviations and their Definitions
Youngja Park, Roy J. Byrd |
EMNLP | 1 |