Jong-Min Choi

dblp:290/5068 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Machine Learning Detection of Radio Occultation Electron Density Profiles Perturbed by the Equatorial Plasma Bubbles
abstract
The FORMOSAT-7/COSMIC-2 (F7C2) constellation consists of six small satellites that provide high temporal and spatial resolutions of ionosphere observations at mid- and low-latitudes using radio occultation (RO) technology. While having the advantage of such dense radio soundings, ensuring the quality of the derived electron density profiles (EDPs) is crucial for applications such as data assimilation forecasting models or monitoring of the ionosphere status. However, after the Hunga Tonga-Hunga Ha’apai volcano erupted on January 15, 2022, more than 70% of the F7C2 EDPs from the Pacific to the Indian Ocean exhibited significant fluctuations, indicating possible data quality degradation. In addition to the extreme event giving such a high proportion of EDPs with quality uncertainties, the fluctuated EDPs are also observed in daily RO soundings. More than 40% of EDPs fluctuate during premidnight hours from October to December within$90~^{\circ }$W–$0~^{\circ }$E, while 70% fluctuate during postmidnight hours from May to July. This study presents, for the first time, a comprehensive investigation into the fluctuating EDPs during usual and event days. The statistics indicate that these fluctuating or irregular EDPs primarily occur during nighttime (>70%). The high correlation (0.82) between the longitudinal and seasonal variations of irregular EDPs and the ion velocity meter (IVM) climatological occurrence of equatorial plasma bubbles (EPBs), as observed in previous studies, indicates that irregular EDPs during postmidnight hours primarily result from EPBs. The machine learning models utilizing the Bagged Trees classification are developed to classify the normal and irregular EDPs across varying times, locations, and solar activity levels showing the clear connection between the EPBs and the irregular EDPs.
Shih-Ping Chen, Charles Chien-Hung Lin, P. K. Rajesh, Pin-Hsuan Cheng, Ho-Fang Tsai, Richard Eastes, Jong-Min Choi, Jann-Yenq Liu, Alfred Bing-Chih Chen
IEEE Trans. Geosci. Remote. Sens.7
2021 Consistency Analysis of Data-Usage Purposes in Mobile Apps
abstract
While privacy laws and regulations require apps and services to disclose the purposes of their data collection to the users (i.e., why do they collect my data?), the data usage in an app's actual behavior does not always comply with the purposes stated in its privacy policy. Automated techniques have been proposed to analyze apps' privacy policies and their execution behavior, but they often overlooked the purposes of the apps' data collection, use and sharing. To mitigate this oversight, we propose PurPliance, an automated system that detects the inconsistencies between the data-usage purposes stated in a natural language privacy policy and those of the actual execution behavior of an Android app. PurPliance analyzes the predicate-argument structure of policy sentences and classifies the extracted purpose clauses into a taxonomy of data purposes. Purposes of actual data usage are inferred from network data traffic. We propose a formal model to represent and verify the data usage purposes in the extracted privacy statements and data flows to detect policy contradictions in a privacy policy and flow-to-policy inconsistencies between network data flows and privacy statements. Our evaluation results of end-to-end contradiction detection have shown PurPliance to improve detection precision from 19% to 95% and recall from 10% to 50% compared to a state-of-the-art method. Our analysis of 23.1k Android apps has also shown PurPliance to detect contradictions in 18.14% of privacy policies and flow-to-policy inconsistencies in 69.66% of apps, indicating the prevalence of inconsistencies of data practices in mobile apps.
Duc Bui, Kang G. Shin, Jong-Min Choi, Jun-Bum Shin
CCS4
2021 Automated Extraction and Presentation of Data Practices in Privacy Policies
abstract
Abstract Privacy policies are documents required by law and regulations that notify users of the collection, use, and sharing of their personal information on services or applications. While the extraction of personal data objects and their usage thereon is one of the fundamental steps in their automated analysis, it remains challenging due to the complex policy statements written in legal (vague) language. Prior work is limited by small/generated datasets and manually created rules. We formulate the extraction of fine-grained personal data phrases and the corresponding data collection or sharing practices as a sequence-labeling problem that can be solved by an entity-recognition model. We create a large dataset with 4.1k sentences (97k tokens) and 2.6k annotated fine-grained data practices from 30 real-world privacy policies to train and evaluate neural networks. We present a fully automated system, called PI-Extract, which accurately extracts privacy practices by a neural model and outperforms, by a large margin, strong rule-based baselines. We conduct a user study on the effects of data practice annotation which highlights and describes the data practices extracted by PI-Extract to help users better understand privacy-policy documents. Our experimental evaluation results show that the annotation significantly improves the users’ reading comprehension of policy texts, as indicated by a 26.6% increase in the average total reading score.
Duc Bui, Kang G. Shin, Jong-Min Choi, Jun-Bum Shin
Proc. Priv. Enhancing Technol.3