Jae Hyun Kim

dblp:15/4502 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author
YearPublicationVenuePosition
2024 CWAS-Plus: estimating category-wide association of rare noncoding variation from whole-genome sequencing data with cell-type-specific functional data
abstract
Variants in cis-regulatory elements link the noncoding genome to human pathology; however, detailed analytic tools for understanding the association between cell-level brain pathology and noncoding variants are lacking. CWAS-Plus, adapted from a Python package for category-wide association testing (CWAS), enhances noncoding variant analysis by integrating both whole-genome sequencing (WGS) and user-provided functional data. With simplified parameter settings and an efficient multiple testing correction method, CWAS-Plus conducts the CWAS workflow 50 times faster than CWAS, making it more accessible and user-friendly for researchers. Here, we used a single-nuclei assay for transposase-accessible chromatin with sequencing to facilitate CWAS-guided noncoding variant analysis at cell-type-specific enhancers and promoters. Examining autism spectrum disorder WGS data (n = 7280), CWAS-Plus identified noncoding de novo variant associations in transcription factor binding sites within conserved loci. Independently, in Alzheimer's disease WGS data (n = 1087), CWAS-Plus detected rare noncoding variant associations in microglia-specific regulatory elements. These findings highlight CWAS-Plus's utility in genomic disorders and scalability for processing large-scale WGS data and in multiple-testing corrections. CWAS-Plus and its user manual are available at https://github.com/joonan-lab/cwas/ and https://cwas-plus.readthedocs.io/en/latest/, respectively.
Minwoo Jeong, In Gyeong Koh, Chanhee Kim, Hyeji Lee, Jae Hyun Kim, Ronald Yurko, Il Bin Kim, Jeongbin Park, Donna M. Werling, Stephan J. Sanders, Joon-Yong An
Briefings Bioinform.6
2023 EvidenceMap: a three-level knowledge representation for medical evidence computation and comprehension
abstract
OBJECTIVE: To develop a computable representation for medical evidence and to contribute a gold standard dataset of annotated randomized controlled trial (RCT) abstracts, along with a natural language processing (NLP) pipeline for transforming free-text RCT evidence in PubMed into the structured representation. MATERIALS AND METHODS: Our representation, EvidenceMap, consists of 3 levels of abstraction: Medical Evidence Entity, Proposition and Map, to represent the hierarchical structure of medical evidence composition. Randomly selected RCT abstracts were annotated following EvidenceMap based on the consensus of 2 independent annotators to train an NLP pipeline. Via a user study, we measured how the EvidenceMap improved evidence comprehension and analyzed its representative capacity by comparing the evidence annotation with EvidenceMap representation and without following any specific guidelines. RESULTS: Two corpora including 229 disease-agnostic and 80 COVID-19 RCT abstracts were annotated, yielding 12 725 entities and 1602 propositions. EvidenceMap saves users 51.9% of the time compared to reading raw-text abstracts. Most evidence elements identified during the freeform annotation were successfully represented by EvidenceMap, and users gave the enrollment, study design, and study Results sections mean 5-scale Likert ratings of 4.85, 4.70, and 4.20, respectively. The end-to-end evaluations of the pipeline show that the evidence proposition formulation achieves F1 scores of 0.84 and 0.86 in the adjusted random index score. CONCLUSIONS: EvidenceMap extends the participant, intervention, comparator, and outcome framework into 3 levels of abstraction for transforming free-text evidence from the clinical literature into a computable structure. It can be used as an interoperable format for better evidence retrieval and synthesis and an interpretable representation to efficiently comprehend RCT findings.
Tian Kang, Yingcheng Sun, Jae Hyun Kim, Casey N. Ta, Adler J. Perotte, Kayla Schiffer, Mutong Wu, Nour Fahmy, Yifan Peng 0002, Chunhua Weng
J. Am. Medical Informatics Assoc.3
2021 Towards clinical data-driven eligibility criteria optimization for interventional COVID-19 clinical trials
abstract
OBJECTIVE: This research aims to evaluate the impact of eligibility criteria on recruitment and observable clinical outcomes of COVID-19 clinical trials using electronic health record (EHR) data. MATERIALS AND METHODS: On June 18, 2020, we identified frequently used eligibility criteria from all the interventional COVID-19 trials in ClinicalTrials.gov (n = 288), including age, pregnancy, oxygen saturation, alanine/aspartate aminotransferase, platelets, and estimated glomerular filtration rate. We applied the frequently used criteria to the EHR data of COVID-19 patients in Columbia University Irving Medical Center (CUIMC) (March 2020-June 2020) and evaluated their impact on patient accrual and the occurrence of a composite endpoint of mechanical ventilation, tracheostomy, and in-hospital death. RESULTS: There were 3251 patients diagnosed with COVID-19 from the CUIMC EHR included in the analysis. The median follow-up period was 10 days (interquartile range 4-28 days). The composite events occurred in 18.1% (n = 587) of the COVID-19 cohort during the follow-up. In a hypothetical trial with common eligibility criteria, 33.6% (690/2051) were eligible among patients with evaluable data and 22.2% (153/690) had the composite event. DISCUSSION: By adjusting the thresholds of common eligibility criteria based on the characteristics of COVID-19 patients, we could observe more composite events from fewer patients. CONCLUSIONS: This research demonstrated the potential of using the EHR data of COVID-19 patients to inform the selection of eligibility criteria and their thresholds, supporting data-driven optimization of participant selection towards improved statistical power of COVID-19 trials.
Jae Hyun Kim, Casey N. Ta, Cong Liu 0020, Cynthia Sung 0002, Alex M. Butler, Latoya A. Stewart, Lyudmila Ena, James R. Rogers, Anna Ostropolets, Patrick B. Ryan, Hao Liu 0054, Shing M. Lee, Mitchell S. V. Elkind, Chunhua Weng
J. Am. Medical Informatics Assoc.1
2021 The COVID-19 Trial Finder
abstract
Clinical trials are the gold standard for generating reliable medical evidence. The biggest bottleneck in clinical trials is recruitment. To facilitate recruitment, tools for patient search of relevant clinical trials have been developed, but users often suffer from information overload. With nearly 700 coronavirus disease 2019 (COVID-19) trials conducted in the United States as of August 2020, it is imperative to enable rapid recruitment to these studies. The COVID-19 Trial Finder was designed to facilitate patient-centered search of COVID-19 trials, first by location and radius distance from trial sites, and then by brief, dynamically generated medical questions to allow users to prescreen their eligibility for nearby COVID-19 trials with minimum human computer interaction. A simulation study using 20 publicly available patient case reports demonstrates its precision and effectiveness.
Yingcheng Sun, Alex M. Butler, Fengyang Lin, Hao Liu 0054, Latoya A. Stewart, Jae Hyun Kim, Betina Ross S. Idnay, Qingyin Ge, Xinyi Wei, Cong Liu 0020, Chi Yuan, Chunhua Weng
J. Am. Medical Informatics Assoc.6
2021 From clinical trials to clinical practice: How long are drugs tested and then used by patients?
abstract
OBJECTIVE: Evidence is scarce regarding the safety of long-term drug use, especially for drugs treating chronic diseases. To bridge this knowledge gap, this research investigated the differences in drug exposure between clinical trials and clinical practice. MATERIALS AND METHODS: We extracted drug follow-up times from clinical trials in ClinicalTrials.gov and compared the difference between clinical trials and real-world usage data for 914 drugs taken by 96 645 927 patients. RESULTS: A total of 17.5% of drugs had longer median exposure in practice than in trials, 6% of patients had extended exposure to at least 1 drug, and drugs treating nervous system disorders and cardiovascular diseases were the most common among drugs with high rates of extended exposure. CONCLUSIONS: For most of patients, the drug use length is shorter than the tested length in clinical trials. Still, a remarkable number of patients experienced extended drug exposure, particularly for drugs treating nervous system disorders or cardiovascular disorders.
Chi Yuan, Patrick B. Ryan, Casey N. Ta, Jae Hyun Kim, Ziran Li, Chunhua Weng
J. Am. Medical Informatics Assoc.4
2021 Building an OMOP common data model-compliant annotated corpus for COVID-19 clinical trials
Yingcheng Sun, Alex M. Butler, Latoya A. Stewart, Hao Liu 0054, Chi Yuan, Christopher T. Southard, Jae Hyun Kim, Chunhua Weng
J. Biomed. Informatics7
2020 CYPminer: an automated cytochrome P450 identification, classification, and data analysis tool for genome data sets across kingdoms
abstract
BACKGROUND: Cytochrome P450 monooxygenases (termed CYPs or P450s) are hemoproteins ubiquitously found across all kingdoms, playing a central role in intracellular metabolism, especially in metabolism of drugs and xenobiotics. The explosive growth of genome sequencing brings a new set of challenges and issues for researchers, such as a systematic investigation of CYPs across all kingdoms in terms of identification, classification, and pan-CYPome analyses. Such investigation requires an automated tool that can handle an enormous amount of sequencing data in a timely manner. RESULTS: CYPminer was developed in the Python language to facilitate rapid, comprehensive analysis of CYPs from genomes of all kingdoms. CYPminer consists of two procedures i) to generate the Genome-CYP Matrix (GCM) that lists all occurrences of CYPs across the genomes, and ii) to perform analyses and visualization of the GCM, including pan-CYPomes (pan- and core-CYPome), CYP co-occurrence networks, CYP clouds, and genome clustering data. The performance of CYPminer was evaluated with three datasets from fungal and bacterial genome sequences. CONCLUSIONS: CYPminer completes CYP analyses for large-scale genomes from all kingdoms, which allows systematic genome annotation and comparative insights for CYPs. CYPminer also can be extended and adapted easily for broader usage.
Ohgew Kweon, Seong-Jae Kim, Jae Hyun Kim, Seong Won Nho, Dongryeoul Bae, Jungwhan Chon, Mark Hart, Dong-Heon Baek, Young-Chang Kim, Sung-Kwan Kim, John B. Sutherland, Carl Cerniglia
BMC Bioinform.3
2012 An optimal motion vector regularization method using variance-distortion curve
abstract
This paper proposes a new motion vector (MV) regularization method. We present an optimization method of the tradeoff between motion accuracy and regularity. The proposed method adopts Lagrangian multiplier and greedy algorithm for optimization. By smoothing MVs while preserving motion accuracy, the proposed method can more efficiently regularize MVs than previous regularization methods. The experimental results show that the proposed method outperforms previous regularization methods by generating objectively and subjectively better results.
Hyungjun Lim, Dong Yoon Kim, Joonsung Choi, Seung-Ho Park, Se Hyeok Park, Jae Hyun Kim, Hyun Wook Park
ICIP6
2012 Motion estimation with adaptive block size for motion-compensated frame interpolation
abstract
This paper proposes a new frame rate up-conversion (FRUC) method. In the paper, we propose an overlap ratio between motion vectors, with which we determine adaptive block size for the motion estimation. In addition, the proposed method estimates the region of repeated patterns and treats them in a special manner. Conventional motion estimation methods usually produce wrong motion vectors for the repeated patterns. The experimental results show that the proposed method outperforms other FRUC methods in terms of generating better interpolated frames.
Hyungjun Lim, Dong Yoon Kim, Hyun Wook Park, Jun Ho Cho, Se Hyeok Park, Jae Hyun Kim
PCS6
2010 A Cooperative Scanning Mechanism for the Mobile Relay in the Moving Network Environment
abstract
This paper proposes a cooperative scanning mechanism to overcome the increased service disruption time due to the successive scanning of the mobile relay station(MRS) and its subordinated mobile stations(MSs) in the moving network environment. The proposed mechanism regulates that the scanning duration of the MSs attached to the MRS is overlapped with that of the MRS. It results in the decreased service disruption time of the MS. We also use the similarity of the signal quality of the MRS and the MSs that is caused by the feature of the moving network to reduce the scanning duration of the MRS. Compared with the conventional scanning mechanism, the proposed mechanism reduces the scanning time of the MRS about 33% and 17% at the two-tier cell structure environment(the number of the neighbor BSs is 18) and the one-tier cell structure environment(the number of the neighbor BSs is 6), respectively. The rough calibration mechanism improves the accuracy of the measurement result. Therefore, the proposed mechanism facilitates the deployment of the vehicular network using the MRS.
Sin-Hun Kang, Jae Hyun Kim
CCNC3
2010 The Active Buffer Management Scheme using Virtual Transmission Delay in the IEEE 802.11e Network
abstract
Due to the advance of WLAN technology, the use of the multimedia service such as the video streaming service has been increased in the home network or CCTV monitoring for cars or construction materials. However, we need to study the method which decreases the transmission delay and the frame loss rate to provide QoS of the video streaming service. Therefore, this paper proposes an active buffer management scheme to guarantee QoS of the video streaming service in the IEEE 802.11e EDCA. The proposed scheme discards the frame in the HoL(Head of Line) of the buffer based on the priority of each frame and the virtual transmission delay of frame newly arriving at the buffer. In the simulation results, the proposed scheme not only decreases the frame loss probability of important I and P frames but also stabilizes the transmission delay. It may increase the QoS of video streaming services.
Kyu-Hwan Lee, Jae Hyun Kim
CCNC3
2008 End mill design and machining via cutting simulation
Jae Hyun Kim, Jung Whan Park, Tae Jo Ko
Comput. Aided Des.1
1994 Two-layered DCT based coding scheme for recording digital HDTV signals
abstract
For the purpose of digital recording of HDTV signals we have designed a bit rate reduction coder that can reduce the input data rate by a compression ratio greater than 2, while maintaining excellent quality for studio applications. In this paper we have developed an adaptive quantization algorithm to decide scale factor that can optimize bit allocation in a cluster unit. Two-layered code data fixed cluster by cluster serves trick play. This algorithm has been designed to fulfil the constraints of professional studio recorders such as interframe editing, multiple copy, picture quality, trick play and robustness for burst and random errors.>
Jae Hyun Kim, Goo Man Park
ICASSP (5)1