EDBT 2026 Demo / reviewers in the wild / expert
Jun-Bum Shin
dblp:70/5871 · also Junbum Shin
· DBLP profile ↗
16ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0005-7985-6163ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IDFace: Face Template Protection for Efficient and Secure IdentificationabstractAs face recognition systems (FRS) become more widely used, user privacy becomes more important. A key privacy issue in FRS is protecting the user's face template, as the characteristics of the user's face image can be recovered from the template. Although recent advances in cryptographic tools such as homomorphic encryption (HE) have provided opportunities for securing the FRS, HE cannot be used directly with FRS in an efficient plug-and-play manner. In particular, although HE is functionally complete for arbitrary programs, it is basically designed for algebraic operations on encrypted data of predetermined shape, such as a polynomial ring. Thus, a non-tailored combination of HE and the system can yield very inefficient performance, and many previous HE-based face template protection methods are hundreds of times slower than plain systems without protection. In this study, we propose IDFace, a new HE-based secure and efficient face identification method with template protection. IDFace is designed on the basis of two novel techniques for efficient searching on a (homomorphically encrypted) biometric database with an angular metric. The first technique is a template representation transformation that sharply reduces the unit cost for the matching test. The second is a space-efficient encoding that reduces wasted space from the encryption algorithm, thus saving the number of operations on encrypted templates. Through experiments, we show that IDFace can identify a face template from among a database of 1M encrypted templates in 126ms, showing only 2X overhead compared to the identification over plaintexts. Sunpill Kim, Seunghun Paik, Chanwoo Hwang, Dongsoo Kim 0004, Jun-Bum Shin, Jae Hong Seo |
ICCV | 5 |
| 2023 | HETAL: Efficient Privacy-preserving Transfer Learning with Homomorphic EncryptionabstractTransfer learning is a de facto standard method for efficiently training machine learning models for data-scarce problems by adding and fine-tuning new classification layers to a model pre-trained on large datasets. Although numerous previous studies proposed to use homomorphic encryption to resolve the data privacy issue in transfer learning in the machine learning as a service setting, most of them only focused on encrypted inference. In this study, we present HETAL, an efficient Homomorphic Encryption based Transfer Learning algorithm, that protects the client's privacy in training tasks by encrypting the client data using the CKKS homomorphic encryption scheme. HETAL is the first practical scheme that strictly provides encrypted training, adopting validation-based early stopping and achieving the accuracy of nonencrypted training. We propose an efficient encrypted matrix multiplication algorithm, which is 1.8 to 323 times faster than prior methods, and a highly precise softmax approximation algorithm with increased coverage. The experimental results for five well-known benchmark datasets show total training times of 567--3442 seconds, which is less than an hour. Seewoo Lee, Garam Lee, Jun-Bum Shin, Mun-Kyu Lee |
ICML | 4 |
| 2023 | HEaaN.MLIR: An Optimizing Compiler for Fast Ring-Based Homomorphic EncryptionabstractHomomorphic encryption (HE) is an encryption scheme that provides arithmetic operations on the encrypted data without doing decryption. For Ring-based HE, an encryption scheme that uses arithmetic operations on a polynomial ring as building blocks, performance improvement of unit HE operations has been achieved by two kinds of efforts. The first one is through accelerating the building blocks, polynomial operations. However, it does not facilitate optimizations across polynomial operations such as fusing two polynomial operations. The second one is implementing highly optimized HE operations in an amalgamated manner. The written codes have superior performance, but they are hard to maintain. To resolve these challenges, we propose HEaaN.MLIR, a compiler that performs optimizations across polynomial operations. Also, we propose Poly and ModArith, compiler intermediate representations (IRs) for integer polynomial arithmetic and modulus arithmetic on integer arrays. HEaaN.MLIR has compiler optimizations that are motivated by manual optimizations that HE developers do. These include optimizing modular arithmetic operations, fusing loops, and vectorizing integer arithmetic instructions. HEaaN.MLIR can parse a program consisting of the Poly and ModArith instructions and generate a high-performance, multithreaded machine code for a CPU. Our experiment shows that the compiled operations outperform heavily optimized open-source and commercial HE libraries by up to 3.06x in a single thread and 4.55x in multiple threads. Sunjae Park, Woosung Song, Seunghyeon Nam, Hyeongyu Kim, Jun-Bum Shin, Juneyoung Lee |
Proc. ACM Program. Lang. | 5 |
| 2021 | Lattice-Based Secure Biometric Authentication for Hamming Distance
Jung Hee Cheon, Dongwoo Kim 0003, Duhyeong Kim, Joohee Lee, Jun-Bum Shin, Yongsoo Song |
ACISP | 5 |
| 2021 | Consistency Analysis of Data-Usage Purposes in Mobile AppsabstractWhile privacy laws and regulations require apps and services to disclose the purposes of their data collection to the users (i.e., why do they collect my data?), the data usage in an app's actual behavior does not always comply with the purposes stated in its privacy policy. Automated techniques have been proposed to analyze apps' privacy policies and their execution behavior, but they often overlooked the purposes of the apps' data collection, use and sharing. To mitigate this oversight, we propose PurPliance, an automated system that detects the inconsistencies between the data-usage purposes stated in a natural language privacy policy and those of the actual execution behavior of an Android app. PurPliance analyzes the predicate-argument structure of policy sentences and classifies the extracted purpose clauses into a taxonomy of data purposes. Purposes of actual data usage are inferred from network data traffic. We propose a formal model to represent and verify the data usage purposes in the extracted privacy statements and data flows to detect policy contradictions in a privacy policy and flow-to-policy inconsistencies between network data flows and privacy statements. Our evaluation results of end-to-end contradiction detection have shown PurPliance to improve detection precision from 19% to 95% and recall from 10% to 50% compared to a state-of-the-art method. Our analysis of 23.1k Android apps has also shown PurPliance to detect contradictions in 18.14% of privacy policies and flow-to-policy inconsistencies in 69.66% of apps, indicating the prevalence of inconsistencies of data practices in mobile apps. Duc Bui, Kang G. Shin, Jong-Min Choi, Jun-Bum Shin |
CCS | 5 |
| 2021 | Automated Extraction and Presentation of Data Practices in Privacy PoliciesabstractAbstract Privacy policies are documents required by law and regulations that notify users of the collection, use, and sharing of their personal information on services or applications. While the extraction of personal data objects and their usage thereon is one of the fundamental steps in their automated analysis, it remains challenging due to the complex policy statements written in legal (vague) language. Prior work is limited by small/generated datasets and manually created rules. We formulate the extraction of fine-grained personal data phrases and the corresponding data collection or sharing practices as a sequence-labeling problem that can be solved by an entity-recognition model. We create a large dataset with 4.1k sentences (97k tokens) and 2.6k annotated fine-grained data practices from 30 real-world privacy policies to train and evaluate neural networks. We present a fully automated system, called PI-Extract, which accurately extracts privacy practices by a neural model and outperforms, by a large margin, strong rule-based baselines. We conduct a user study on the effects of data practice annotation which highlights and describes the data practices extracted by PI-Extract to help users better understand privacy-policy documents. Our experimental evaluation results show that the annotation significantly improves the users’ reading comprehension of policy texts, as indicated by a 26.6% increase in the average total reading score. Duc Bui, Kang G. Shin, Jong-Min Choi, Jun-Bum Shin |
Proc. Priv. Enhancing Technol. | 4 |
| 2020 | Cryptanalysis of the obfuscated round boundary technique for whitebox cryptography
Yongjin Yeom, Dong-Chan Kim, Chung Hun Baek, Jun-Bum Shin |
Sci. China Inf. Sci. | 4 |
| 2020 | Learning New Words from Keystroke Data with Local Differential PrivacyabstractKeystroke data collected from smart devices includes various sensitive information about users. Collecting and analyzing such data raise serious privacy concerns. Google and Apple have recently applied local differential privacy (LDP) to address privacy issue on learning new words from users' keystroke data. However, these solutions require multiple LDP reports for a single word, which result in inefficient use of privacy budget and high computational cost. In this paper, we develop a novel algorithm for learning new words under LDP. Unlike the existing solutions, the proposed method generates only one LDP report for a single word. This enables the proposed method to use full privacy budget for generating a report and brings the benefit that the proposed method provides better utility at the same privacy degree than the existing methods. In our algorithm, each user appends a hash value to new word and sends only one LDP report of an n-gram selected randomly from the string packed by each new word and its hash value. The server then decodes frequent n-grams at each position of the string and discovers the candidate words by exploring graph-theoretic links between n-grams and checking integrity of candidates with hash values. Frequencies of frequent new words discovered are estimated from distribution estimates of n-grams by robust regression. We theoretically show that our algorithm can recover popular new words even though the server does not know the domain of the raw data. In addition, we theoretically and empirically demonstrate that our algorithm achieves higher accuracy compared to the existing solutions. Sungwook Kim 0001, Hyejin Shin, Chung Hun Baek, Soohyung Kim, Jun-Bum Shin |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Collecting and Analyzing Multidimensional Data with Local Differential PrivacyabstractLocal differential privacy (LDP) is a recently proposed privacy standard for collecting and analyzing data, which has been used, e.g., in the Chrome browser, iOS and macOS. In LDP, each user perturbs her information locally, and only sends the randomized version to an aggregator who performs analyses, which protects both the users and the aggregator against private information leaks. Although LDP has attracted much research attention in recent years, the majority of existing work focuses on applying LDP to complex data and/or analysis tasks. In this paper, we point out that the fundamental problem of collecting multidimensional data under LDP has not been addressed sufficiently, and there remains much room for improvement even for basic tasks such as computing the mean value over a single numeric attribute under LDP. Motivated by this, we first propose novel LDP mechanisms for collecting a numeric attribute, whose accuracy is at least no worse (and usually better) than existing solutions in terms of worst-case noise variance. Then, we extend these mechanisms to multidimensional data that can contain both numeric and categorical attributes, where our mechanisms always outperform existing solutions regarding worst-case noise variance. As a case study, we apply our solutions to build an LDP-compliant stochastic gradient descent algorithm (SGD), which powers many important machine learning tasks. Experiments using real datasets confirm the effectiveness of our methods, and their advantages over existing solutions. Ning Wang 0026, Xiaokui Xiao, Yin Yang 0001, Jun Zhao 0007, Siu Cheung Hui, Hyejin Shin, Jun-Bum Shin, Ge Yu 0001 |
ICDE | 7 |
| 2018 | PrivTrie: Effective Frequent Term Discovery under Local Differential PrivacyabstractA mobile operating system often needs to collect frequent new terms from users in order to build and maintain a comprehensive dictionary. Collecting keyboard usage data, however, raises privacy concerns. Local differential privacy (LDP) has been established as a strong privacy standard for collecting sensitive information from users. Currently, the best known solution for LDP-compliant frequent term discovery transforms the problem into collecting n-grams under LDP, and subsequently reconstructs terms from the collected n-grams by modelling the latter into a graph, and identifying cliques on this graph. Because the transformed problem (i.e., collecting n-grams) is very different from the original one (discovering frequent terms), the end result has poor utility. Further, this method is also rather expensive due to clique computation on a large graph. In this paper we tackle the problem head on: our proposal, PrivTrie, directly collects frequent terms from users by iteratively constructing a trie under LDP. While the methodology of building a trie is an obvious choice, obtaining an accurate trie under LDP is highly challenging. PrivTrie achieves this with a novel adaptive approach that conserves privacy budget by building internal nodes of the trie with the lowest level of accuracy necessary. Experiments using real datasets confirm that PrivTrie achieves high accuracy on common privacy levels, and consistently outperforms all previous methods. Ning Wang 0026, Xiaokui Xiao, Yin Yang 0001, Ta Duy Hoang, Hyejin Shin, Jun-Bum Shin, Ge Yu 0001 |
ICDE | 6 |
| 2018 | New Communication Technologies in Poster Design through the Religion Conflict in Global Social IssuesabstractWe are facing numerous social issues and problems now. As designer Paula Scher wisely said “Design matters.” Most people know that graphic designers have commercial clients-creating solutions for brands and corporations. The visual communication profession helps to drive the economy, provide information to the public, and promote competition. There is another side of graphic design that is less well known and vital to society: designers use their expertise to inform people about important social and political issues and promote good causes. Jun-Bum Shin |
IV | 1 |
| 2018 | Efficient Privacy-Preserving Matrix Factorization for Recommendation via Fully Homomorphic EncryptionabstractThere are recommendation systems everywhere in our daily life. The collection of personal data of users by a recommender in the system may cause serious privacy issues. In this article, we propose the first privacy-preserving matrix factorization for recommendation using fully homomorphic encryption. Our protocol performs matrix factorization over encrypted users’ rating data and returns encrypted outputs so that the recommendation system learns nothing on rating values and resulting user/item profiles. Furthermore, the protocol provides a privacy-preserving method to optimize the tuning parameters that can be a business benefit for the recommendation service providers. To overcome the performance degradation caused by the use of fully homomorphic encryption, we introduce a novel data structure to perform computations over encrypted vectors, which are essential for matrix factorization, through secure two-party computation in part. Our experiments demonstrate the efficiency of our protocol. Dongyoung Koo, Yuna Kim, Hyunsoo Yoon, Jun-Bum Shin, Sungwook Kim 0001 |
ACM Trans. Priv. Secur. | 5 |
| 2018 | Privacy Enhanced Matrix Factorization for Recommendation with Local Differential PrivacyabstractRecommender systems are collecting and analyzing user data to provide better user experience. However, several privacy concerns have been raised when a recommender knows user's set of items or their ratings. A number of solutions have been suggested to improve privacy of legacy recommender systems, but the existing solutions in the literature can protect either items or ratings only. In this paper, we propose a recommender system that protects both user's items and ratings. For this, we develop novel matrix factorization algorithms under local differential privacy (LDP). In a recommender system with LDP, individual users randomize their data themselves to satisfy differential privacy and send the perturbed data to the recommender. Then, the recommender computes aggregates of the perturbed data. This framework ensures that both user's items and ratings remain private from the recommender. However, applying LDP to matrix factorization typically raises utility issues with i) high dimensionality due to a large number of items and ii) iterative estimation algorithms. To tackle these technical challenges, we adopt dimensionality reduction technique and a novel binary mechanism based on sampling. We additionally introduce a factor that stabilizes the perturbed gradients. With MovieLens and LibimSeTi datasets, we evaluate recommendation accuracy of our recommender system and demonstrate that our algorithm performs better than the existing differentially private gradient descent algorithm for matrix factorization under stronger privacy requirements. Hyejin Shin, Sungwook Kim 0001, Jun-Bum Shin, Xiaokui Xiao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Efficient Privacy-Preserving Matrix Factorization via Fully Homomorphic Encryption: Extended AbstractabstractRecommendation systems become popular in our daily life. It is well known that the more the release of users' personal data, the better the quality of recommendation. However, such services raise serious privacy concerns for users. In this paper, focusing on matrix factorization-based recommendation systems, we propose the first privacy-preserving matrix factorization using fully homomorphic encryption. On inputs of encrypted users' ratings, our protocol performs matrix factorization over the encrypted data and returns encrypted outputs so that the recommendation system knows nothing on rating values and resulting user/item profiles. It provides a way to obfuscate the number and list of items a user rated without harming the accuracy of recommendation, and additionally protects recommender's tuning parameters for business benefit and allows the recommender to optimize the parameters for quality of service. To overcome performance degradation caused by the use of fully homomorphic encryption, we introduce a novel data structure to perform computations over encrypted vectors, which are essential operations for matrix factorization, through secure 2-party computation in part. With the data structure, the proposed protocol requires dozens of times less computation cost over those of previous works. Our experiments on a personal computer with 3.4 GHz 6-cores 64 GB RAM show that the proposed protocol runs in 1.5 minutes per iteration. It is more efficient than Nikolaenko et al.'s work proposed in CCS 2013, in which it took about 170 minutes on two servers with 1.9 GHz 16-cores 128 GB RAM. Sungwook Kim 0001, Dongyoung Koo, Yuna Kim, Hyunsoo Yoon, Jun-Bum Shin |
AsiaCCS | 6 |
| 2016 | New Paradigm of Social Poster Embedding Unexpected Graphic Pattern for the Ongoing Issue, Radiation Contamination in FukushimaabstractA poster must be sufficiently informative to convey a suitable message which reflects the genuine characteristics of the topic. A social poster design deals with social issues, which are not instantaneous events where existing in a dynamic society. Therefore, the social poster design has to represent the continuous consequences due to the social incident. On 11 March in 2011, a magnitude 9.0 earthquake hit off the east coast of Japan, caused 15,821 deaths, 3,962 missing, and 5,940 injuries in 20 Japanese prefectures as reported by The National Police Agency of Japan (Dunbar et al., 2011). Though the earthquake, aftershocks and tsunamis ended, it is an unfinished tragedy, due to the continuous exposure of the radioactive contaminated water by the Fukushima Daiichi plant. The highly contaminated water has accumulated on roofs and it keeps flowed into the Pacific Ocean when it rains (McCurry, 2015). This study aims to present the ongoing impact of radioactive materials to our society on a social poster through an animated graphic. The randomly created background patterns within the poster, symbolize the unfinished radioactive leaking and its circulation in our environment. Hence, this poster is a dynamically changing rather than a static graphic poster, much like the continuous changes occurring within the environment, due to the Fukushima event. In the middle, the red solid circle not only reflects the Japanese flag but also symbolizes the contaminated Earth. For the background patterns, I generated the random motion of oblique lines by computer programming using Processing, which is an open source programming language. The randomly generated graphic by Processing, referred to as animated graphic, exhibits the unexpected radioactive exposures due to the Fukushima Daiichi nuclear disaster. The red circle represented Japan and the Earth suffered by the ongoing contamination that it was depicted by transparent red that the symbol and the patterns are overlapped and coexist. It should be noted that the printed version of the poster is captured from the digital work, which can be found at https://dl.dropboxusercontent.com/s/ulcosh5dpm1xw85/index.html?dl=0. The animated graphic shown in the dynamic poster, informs an audience of the ongoing social issue that the radioactivity which has been leaking into the ocean since the Fukushima Daiichi nuclear disaster in 2011. Unlike a traditional, static print poster, it raises the audiences' awareness, empathy and understanding of the tragedy using dynamic data, affecting a different reaction. Jun-Bum Shin |
IV | 1 |
| 2016 | Hardware-Assisted On-Demand Hypervisor Activation for Efficient Security Critical Code Execution on Mobile Devices
Yeongpil Cho, Jun-Bum Shin, Donghyun Kwon, MyungJoo Ham, Yuna Kim, Yunheung Paek |
USENIX ATC | 2 |