EDBT 2026 Demo / reviewers in the wild / expert
Sunhee Kim
dblp:02/9234
· DBLP profile ↗
2ranked-venue papers in the field
2as first author
2since 2021 · last 2023
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A computational method of identifying false positives in the variant callingabstractWe studied the strain-specific effect in genetic variant calling. To this end, we used two major strains of the rice genome, Indica and Japonica, and called the variant with models that are different in the composition of samples from the two strains. We found that the more the samples differed in their strains from the reference, the more variants were predicted. We used confusion matrices in machine learning methods to compare the performance of different variant calling models. We found that a significant portion of predicted variants are potential false positive variants. We then proposed a method to identify the false positives. The proposed method involves calling true variants from the purebred samples and the reference of the same strain. We demonstrated the validity of the proposed method on the different variant calling models. Sunhee Kim, Sang-Ho Chu, Chang-Yong Lee |
IEEE Big Data | 1 |
| 2022 | Computational method of database construction for genetic variant callingabstractIn this study, we examined the impact of the variant database in recalibration and developed a database-generation model that gathers potential candidates directly from resequencing genome data. Based on human genome data, we optimize the hyper-parameters in the model and evaluate the performance improvements both in terms of recalibration and variant calling. To test whether our pseudo-database approach is applicable to species other than human, we constructed pseudo-databases for sheep, rice, and chickpea, and compared its performance with dbSNP. Consistently, we find that our pseudo-database provides improved recalibration and error rates. More importantly, the use of pseudo-databases led to the identification of additional genetic variants. Therefore, the reanalysis with our pseudo-databases approach effectively recalibrates the base quality scores and consequently uncovers hidden genetic variations in published resequencing data. Sunhee Kim, Chang-Yong Lee |
IEEE Big Data | 1 |