Taesung Lee

dblp:60/10355 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0003-1015-7004ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Batcher: Learning to Construct Cost-Efficient Batches of Small Queries in Big Data Processing Platforms
Yeonsu Park 0001, Taesung Lee, Byung-Chul Tak, Wook-Shin Han
ICDE2
2026 TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
Taesung Lee, Jaehyun Ha, Byung-Chul Tak, Wook-Shin Han
Proc. VLDB Endow.1
2023 Matching Pairs: Attributing Fine-Tuned Models to their Pre-Trained Large Language Models
abstract
Myles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou, Gabriele Picco, Giulio Zizzo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Myles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou 0001, Gabriele Picco, Giulio Zizzo
ACL (1)3
2023 EdgeTorrent: Real-time Temporal Graph Representations for Intrusion Detection
abstract
Anomaly-based intrusion detection aims to learn the normal behaviors of a system and detect activity that deviates from it. One of the best ways to represent the behavior of a computer network is through provenance graphs: dynamic networks of entity interactions over time. When provenance graphs deviate from their normal behaviors, it could be indicative of a malicious actor attempting to compromise the network. However, efficiently characterizing the normal behavior of large temporal graphs is challenging. To do this, we propose EdgeTorrent, an end-to-end anomaly-based intrusion detection system for provenance graph analysis. EdgeTorrent leverages a novel high-performance message passing neural network for graph embedding over a stream of edges to capture both temporal and topological changes in the system. These embeddings are then processed by a novel adversarially trained sequence analyzer that alerts when a series of graph embeddings changes in an unexpected way. EdgeTorrent preserves temporal ordering during message passing, and its streaming-focused design allows users to conduct out-of-core inference on billion-edge graphs, faster than real-time. We show that our method outperforms state-of-the-art graph-kernel approaches on several host monitoring data sets; notably, it is the first intrusion detection system to perfectly classify the StreamSpot data set. Additionally, we show it is the best-performing method on a real-world, billion-edge data set encompassing 11 days of benign and attack data.
Isaiah J. King, Xiaokui Shu, Jiyong Jang, Kevin Eykholt, Taesung Lee, H. Howie Huang
RAID5
2023 URET: Universal Robustness Evaluation Toolkit (for Evasion)
Kevin Eykholt, Taesung Lee, Douglas Lee Schales, Jiyong Jang, Ian M. Molloy, Masha Zorin
USENIX Security Symposium2
2022 Backdoor smoothing: Demystifying backdoor attacks on deep neural networks
Kathrin Grosse, Taesung Lee, Battista Biggio, Youngja Park, Michael Backes 0001, Ian M. Molloy
Comput. Secur.2
2021 Adaptive Verifiable Training Using Pairwise Class Similarity
Shiqi Wang 0002, Kevin Eykholt, Taesung Lee, Jiyong Jang, Ian M. Molloy
AAAI3
2021 iTurboGraph: Scaling and Automating Incremental Graph Analytics
abstract
With the rise of streaming data for dynamic graphs, large-scale graph analytics meets a new requirement of Incremental Computation because the larger the graph, the higher the cost for updating the analytics results by re-execution. A dynamic graph consists of an initial graph G and graph mutation updates Δ G$ of edge insertions or deletions. Given a query Q, its results $Q(G)$, and updates for Δ G$ to G, incremental graph analytics computes updates Δ Q$ such that Q($G \cup Δ G)$ = $Q(G)$ $\cup$ Δ Q$ where $\cup$ is a union operator. In this paper, we consider the problem of large-scale incremental neighbor-centric graph analytics (\NGA ). We solve the limitations of previous systems: lack of usability due to the difficulties in programming incremental algorithms for \NGA and limited scalability and efficiency due to the overheads in maintaining intermediate results for graph traversals in \NGA. First, we propose a domain-specific language, ŁNGA, and develop its compiler for intuitive programming of \NGA, automatic query incrementalization, and query optimizations. Second, we define Graph Streaming Algebra as a theoretical foundation for scalable processing of incremental \NGA. We introduce a concept of Nested Graph Windows and model graph traversals as the generation of walk streams. Lastly, we present a system \SystemName, which efficiently processes incremental \NGA for large graphs. Comprehensive experiments show that it effectively avoids costly re-executions and efficiently updates the analytics results with reduced IO and computations.
Seongyun Ko, Taesung Lee, Kijae Hong, In Seo, Jiwon Seo 0002, Wook-Shin Han
SIGMOD Conference2
2019 Supervising Unsupervised Open Information Extraction Models
abstract
Arpita Roy, Youngja Park, Taesung Lee, Shimei Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arpita Roy, Youngja Park, Taesung Lee, Shimei Pan
EMNLP/IJCNLP (1)3
2019 AdvIT: Adversarial Frames Identifier Based on Temporal Consistency in Videos
abstract
Deep neural networks (DNNs) have been widely applied in various applications, including autonomous driving and surveillance systems. However, DNNs are found to be vulnerable to adversarial examples, which are carefully crafted inputs aiming to mislead a learner to make incorrect predictions. While several defense and detection approaches are proposed for static image classification, many security-critical tasks use videos as their input and require efficient processing. In this paper, we propose an efficient and effective method advIT to detect adversarial frames within videos against different types of attacks based on temporal consistency property of videos. In particular, we apply optical flow estimation to the target and previous frames to generate pseudo frames and evaluate the consistency of the learner output between these pseudo frames and target. High inconsistency indicates that the target frame is adversarial. We conduct extensive experiments on various learning tasks including video semantic segmentation, human pose estimation, object detection, and action recognition, and demonstrate that we can achieve above 95% adversarial frame detection rate. To consider adaptive attackers, we show that even if an adversary has access to the detector and performs a strong adaptive attack based on the state of the art expectation of transformation method, the detection rate stays almost the same. We also tested the transferability among different optical flow estimators and show that it is hard for attackers to attack one and transfer the perturbation to others. In addition, as efficiency is important in video analysis, we show that advIT can achieve real-time detection in about 0.03--0.4 seconds.
Chaowei Xiao, Ruizhi Deng, Bo Li 0026, Taesung Lee, Benjamin Edwards, Jinfeng Yi, Dawn Song, Mingyan Liu, Ian M. Molloy
ICCV4
2019 Unsupervised Sentence Embedding Using Document Structure-Based Context
Taesung Lee, Youngja Park
ECML/PKDD (2)1
2018 List Intersection for Web Search: Algorithms, Cost Models, and Optimizations
abstract
This paper studies the optimization of list intersection, especially in the context of the matching phase of search engines. Given a user query, we intersect the postings lists corresponding to the query keywords to generate the list of documents matching all keywords. Since the speed of list intersection depends the algorithm, hardware, and list lengths and their correlations, none the existing intersection algorithms outperforms the others in every scenario. Therefore, we develop a cost-based approach in which we identify a search space, spanning existing algorithms and their combinations. We propose a cost model to estimate the cost of the algorithms with their combinations, and use the cost model to search for the lowest-cost algorithm. The resulting plan is usually a combination of 2-way algorithms, outperforming conventional 2-way and k -way algorithms. The proposed approach is more general than designing a specific algorithm, as the cost models can be adapted to different hardware. We validate the cost model experimentally on two different CPUs, and show that the cost model closely estimates the actual cost. Using both real and synthetic datasets, we show that the proposed cost-based optimizer outperforms the state-of-the-art alternatives.
Taesung Lee, Seung-won Hwang, Sameh Elnikety
Proc. VLDB Endow.2
2016 Trivia quiz mining using probabilistic knowledge
abstract
Recent work suggests that providing unexpected information is an important factor for drawing user traffic. Such examples can be easily found in the “Did you know” section of the Wikipedia main page, the ESPN quiz, the Google Doodles, and the Bing main page. Inspired by these applications, we propose a novel trivia quiz mining asking unexpected questions for a given entity. We solve this problem by linking different types of social media as input and output, and mine unexpected properties based on prototype theory to mediate the input and the output media.
Taesung Lee, Seung-won Hwang, Zhongyuan Wang 0006
ASONAM1
2016 Probabilistic Prototype Model for Serendipitous Property Mining
abstract
Besides providing the relevant information, amusing users has been an important role of the web. Many web sites provide serendipitous (unexpected but relevant) information to draw user traffic. In this paper, we study the representative scenario of mining an amusing quiz. An existing approach leverages a knowledge base to mine an unexpected property then find quiz questions on such property, based on prototype theory in cognitive science. However, existing deterministic model is vulnerable to noise in the knowledge base. Therefore, we instead propose to leverage probabilistic approach to build a prototype that can overcome noise. Our extensive empirical study shows that our approach not only significantly outperforms baselines by 0.06 in accuracy, and 0.11 in serendipity but also shows higher relevance than the traditional relevance-pursuing baseline using TF-IDF.
Taesung Lee, Seung-won Hwang, Zhongyuan Wang 0006
COLING1
2015 Processing and Optimizing Main Memory Spatial-Keyword Queries
abstract
Important cloud services rely on spatial-keyword queries, containing a spatial predicate and arbitrary boolean keyword queries. In particular, we study the processing of such queries in main memory to support short response times. In contrast, current state-of-the-art spatial-keyword indexes and relational engines are designed for different assumptions. Rather than building a new spatial-keyword index, we employ a cost-based optimizer to process these queries using a spatial index and a keyword index. We address several technical challenges to achieve this goal. We introduce three operators as the building blocks to construct plans for main memory query processing. We then develop a cost model for the operators and query plans. We introduce five optimization techniques that efficiently reduce the search space and produce a query plan with low cost. The optimization techniques are computationally efficient, and they identify a query plan with a formal approximation guarantee under the common independence assumption. Furthermore, we extend the framework to exploit interesting orders. We implement the query optimizer to empirically validate our proposed approach using real-life datasets. The evaluation shows that the optimizations provide significant reduction in the average and tail latency of query processing: 7- to 11-fold reduction over using a single index in terms of 99th percentile response time. In addition, this approach outperforms existing spatial-keyword indexes, and DBMS query optimizers for both average and high-percentile response times.
Taesung Lee, Seung-won Hwang, Sameh Elnikety, Yuxiong He
Proc. VLDB Endow.1
2014 Map Translation Using Geo-tagged Social Media
abstract
This paper discusses the problem of map translation, of servicing spatial entities in multiple languages.Existing work on entity translation harvests translation evidence from text resources, not considering spatial locality in translation.In contrast, we mine geo-tagged sources for multilingual tags to improve recall, and consider spatial properties of tags for translation to improve precision.Our approach empirically improves accuracy from 0.562 to 0.746 using Taiwanese spatial entities.
Sunyou Lee, Taesung Lee, Seung-won Hwang
EACL2
2014 Overcoming Asymmetry in Entity Graphs
abstract
This paper studies the problem of mining named entity translations by aligning comparable corpora. Current state-of-the-art approaches mine a translation pair by aligning an entity graph in one language to another based on node similarity or propagated similarity of related entities. However, they, building on the assumption of “symmetry”, quickly deteriorate on “weakly” comparable corpora with some asymmetry. In this paper, we pursue two directions for overcoming relation and entity asymmetry respectively. The first approach starts from weakly comparable corpora (for high recall) then ensures precision by selective propagation only to entities of symmetric relations. The second approach starts from parallel corpora (for high precision) then enhances recall by extending the translation matrix based on node similarity and contextual similarity. Our experimental results on English-Chinese corpora show that both approaches are effective and complementary. Our combined approach outperforms the best-performing baseline in terms of F1-score by up to 0.28.
Taesung Lee, Young-rok Cha, Seung-won Hwang
IEEE Trans. Knowl. Data Eng.1
2013 Bootstrapping Entity Translation on Weakly Comparable Corpora
Taesung Lee, Seung-won Hwang
ACL (1)1
2013 Attribute extraction and scoring: A probabilistic approach
abstract
Knowledge bases, which consist of concepts, entities, attributes and relations, are increasingly important in a wide range of applications. We argue that knowledge about attributes (of concepts or entities) plays a critical role in inferencing. In this paper, we propose methods to derive attributes for millions of concepts and we quantify the typicality of the attributes with regard to their corresponding concepts. We employ multiple data sources such as web documents, search logs, and existing knowledge bases, and we derive typicality scores for attributes by aggregating different distributions derived from different sources using different methods. To the best of our knowledge, ours is the first approach to integrate concept- and instance-based patterns into probabilistic typicality scores that scale to broad concept space. We have conducted extensive experiments to show the effectiveness of our approach.
Taesung Lee, Zhongyuan Wang 0006, Haixun Wang, Seung-won Hwang
ICDE1
2011 Web Scale Taxonomy Cleansing
Taesung Lee, Zhongyuan Wang 0006, Haixun Wang, Seung-won Hwang
Proc. VLDB Endow.1