EDBT 2026 Demo / reviewers in the wild / expert
Hyunghoon Cho
dblp:160/6954
· DBLP profile ↗
16ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-2713-0150ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decor: Delegated Computation on Randomness for Secure Evaluation of Nonlinear Functions
Haris Smajlovic, Kyle Sheng, Timos Antonopoulos, Ruzica Piskac, Hyunghoon Cho |
SP | 5 |
| 2025 | TX-Phase: Secure Phasing of Private Genomes in a Trusted Execution Environment
Natnatee Dokmai, Kaiyuan Zhu, Süleyman Cenk Sahinalp, Hyunghoon Cho |
RECOMB | 4 |
| 2025 | Shechi: A Secure Distributed Computation Compiler Based on Multiparty Homomorphic Encryption
Haris Smajlovic, David Froelicher, Ariya Shajii, Bonnie Berger, Hyunghoon Cho, Ibrahim Numanagic |
USENIX Security Symposium | 5 |
| 2025 | Generating synthetic electronic health record data: a methodological scoping review with benchmarking on phenotype data and open-source softwareabstractOBJECTIVES: To conduct a scoping review (ScR) of existing approaches for synthetic Electronic Health Records (EHR) data generation, to benchmark major methods, and to provide an open-source software and offer recommendations for practitioners. MATERIALS AND METHODS: We search three academic databases for our scoping review. Methods are benchmarked on open-source EHR datasets, Medical Information Mart for Intensive Care III and IV (MIMIC-III/IV). Seven existing methods covering major categories and two baseline methods are implemented and compared. Evaluation metrics concern data fidelity, downstream utility, privacy protection, and computational cost. RESULTS: Forty-eight studies are identified and classified into five categories. Seven open-source methods covering all categories are selected, trained on MIMIC-III, and evaluated on MIMIC-III or MIMIC-IV for transportability considerations. Among them, Generative Adversarial Network (GAN)-based methods demonstrate competitive performance in fidelity and utility on MIMIC-III, rule-based methods excel in privacy protection. Similar findings are observed on MIMIC-IV, except that GAN-based methods further outperform the baseline methods in preserving fidelity. DISCUSSION: Method choice is governed by the relative importance of the evaluation metrics in downstream use cases. We provide a decision tree to guide the choice among the benchmarked methods. An extensible Python package, "SynthEHRella", is provided to facilitate streamlined evaluations. CONCLUSION: GAN-based methods excel when distributional shifts exist between the training and testing populations. Otherwise, CorGAN and MedGAN are most suitable for association modeling and predictive modeling, respectively. Future research should prioritize enhancing fidelity of the synthetic data while controlling privacy exposure, and comprehensive benchmarking of longitudinal or conditional generation methods. Xingran Chen, Zhenke Wu, Hyunghoon Cho, Bhramar Mukherjee |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Secure Discovery of Genetic Relatives Across Large-Scale and Distributed Genomic Datasets
Matthew M. Hong, David Froelicher, Ricky Magner, Victoria Popic, Bonnie Berger, Hyunghoon Cho |
RECOMB | 6 |
| 2023 | Scalable and Privacy-Preserving Federated Principal Component AnalysisabstractPrincipal component analysis (PCA) is an essential algorithm for dimensionality reduction in many data science domains. We address the problem of performing a federated PCA on private data distributed among multiple data providers while ensuring data confidentiality. Our solution, SF-PCA, is an end-to-end secure system that preserves the confidentiality of both the original data and all intermediate results in a passive-adversary model with up to all-but-one colluding parties. SF-PCA jointly leverages multiparty homomorphic encryption, interactive protocols, and edge computing to efficiently interleave computations on local cleartext data with operations on collectively encrypted data. SF-PCA obtains results as accurate as non-secure centralized solutions, independently of the data distribution among the parties. It scales linearly or better with the dataset dimensions and with the number of data providers. SF-PCA is more precise than existing approaches that approximate the solution by combining local analysis results, and between 3x and 250x faster than privacy-preserving alternatives based solely on secure multiparty computation or homomorphic encryption. Our work demonstrates the practical applicability of secure and federated PCA on private distributed datasets. David Froelicher, Hyunghoon Cho, Manaswitha Edupalli, João Sá Sousa, Jean-Philippe Bossuat, Apostolos Pyrgelis, Juan Ramón Troncoso-Pastoriza, Bonnie Berger, Jean-Pierre Hubaux |
SP | 2 |
| 2022 | k-SALSA: k-Anonymous Synthetic Averaging of Retinal Images via Local Style Alignmentabstract-SALSA leverages a new technique, called local style alignment, to generate a synthetic average that maximizes the retention of fine-grain visual patterns in the source images, thus improving the clinical utility of the generated images. On two benchmark datasets of diabetic retinopathy (EyePACS and APTOS), we demonstrate our improvement upon existing methods with respect to image fidelity, classification performance, and mitigation of membership inference attacks. Our work represents a step toward broader sharing of retinal images for scientific collaboration. Code is available at https://github.com/hcholab/k-salsa. Minkyu Jeon, Hyeon-Jin Park, Hyunwoo J. Kim, Michael G. Morley, Hyunghoon Cho |
ECCV (21) | 5 |
| 2022 | Mechanisms for Hiding Sensitive Genotypes With Information-Theoretic PrivacyabstractMotivated by the growing availability of personal genomics services, we study an information-theoretic privacy problem that arises when sharing genomic data: a user wants to share his or her genome sequence while keeping the genotypes at certain positions hidden, which could otherwise reveal critical health-related information. A straightforward solution of erasing (masking) the chosen genotypes does not ensure privacy, because the correlation between nearby positions can leak the masked genotypes. We introduce an erasure-based privacy mechanism with perfect information-theoretic privacy, whereby the released sequence is statistically independent of the sensitive genotypes. Our mechanism can be interpreted as a locally-optimal greedy algorithm for a given processing order of sequence positions, where utility is measured by the number of positions released without erasure. We show that finding an optimal order is NP-hard in general and provide an upper bound on the optimal utility. For sequences from hidden Markov models, a standard modeling approach in genetics, we propose an efficient algorithmic implementation of our mechanism with complexity polynomial in sequence length. Moreover, we illustrate the robustness of the mechanism by bounding the privacy leakage from erroneous prior distributions. Our work is a step towards more rigorous control of privacy in genomic data sharing. Fangwei Ye, Hyunghoon Cho, Salim El Rouayheb |
IEEE Trans. Inf. Theory | 2 |
| 2021 | Bayesian information sharing enhances detection of regulatory associations in rare cell typesabstractMOTIVATION: Recent advances in single-cell RNA-sequencing (scRNA-seq) technologies promise to enable the study of gene regulatory associations at unprecedented resolution in diverse cellular contexts. However, identifying unique regulatory associations observed only in specific cell types or conditions remains a key challenge; this is particularly so for rare transcriptional states whose sample sizes are too small for existing gene regulatory network inference methods to be effective. RESULTS: We present ShareNet, a Bayesian framework for boosting the accuracy of cell type-specific gene regulatory networks by propagating information across related cell types via an information sharing structure that is adaptively optimized for a given single-cell dataset. The techniques we introduce can be used with a range of general network inference algorithms to enhance the output for each cell type. We demonstrate the enhanced accuracy of our approach on three benchmark scRNA-seq datasets. We find that our inferred cell type-specific networks also uncover key changes in gene associations that underpin the complex rewiring of regulatory networks across cell types, tissues and dynamic biological processes. Our work presents a path toward extracting deeper insights about cell type-specific gene regulation in the rapidly growing compendium of scRNA-seq datasets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. AVAILABILITY AND IMPLEMENTATION: The code for ShareNet is available at http://sharenet.csail.mit.edu and https://github.com/alexw16/sharenet. Alexander P. Wu, Jian Peng 0001, Bonnie Berger, Hyunghoon Cho |
Bioinform. | 4 |
| 2020 | Mechanisms for Hiding Sensitive Genotypes with Information-Theoretic PrivacyabstractThe growing availability of personal genomics services comes with increasing concerns for genomic privacy. Individuals may wish to withhold sensitive genotypes that contain critical health-related information when sharing their data with such services. A straightforward solution that masks only the sensitive genotypes does not ensure privacy due to the correlation structure within the genome. Here, we develop an information-theoretic mechanism for masking sensitive genotypes, which ensures no information about the sensitive genotypes is leaked. We also propose an efficient algorithmic implementation of our mechanism for genomic data governed by hidden Markov models. Our work is a step towards more rigorous control of privacy in genomic data sharing. Fangwei Ye, Hyunghoon Cho, Salim El Rouayheb |
ISIT | 2 |
| 2020 | Privacy-Preserving Biomedical Database Queries with Optimal Privacy-Utility Trade-Offs
Hyunghoon Cho, Sean Simmons 0001, Ryan Kim, Bonnie Berger |
RECOMB | 1 |
| 2019 | Large-Margin Classification in Hyperbolic SpaceabstractRepresenting data in hyperbolic space can effectively capture latent hierarchical relationships. To enable accurate classification of points in hyperbolic space while respecting their hyperbolic geometry, we introduce hyperbolic SVM, a hyperbolic formulation of support vector machine classifiers, and describe its theoretical connection to the Euclidean counterpart. We also generalize Euclidean kernel SVM to hyperbolic space, allowing nonlinear hyperbolic decision boundaries and providing a geometric interpretation for a certain class of indefinite kernels. Hyperbolic SVM improves classification accuracy in simulation and in real-world problems involving complex networks and word embeddings. Our work enables end-to-end analyses based on the inherent hyperbolic geometry of the data without resorting to ill-fitting tools developed for Euclidean space. Hyunghoon Cho, Benjamin Demeo, Jian Peng 0001, Bonnie Berger |
AISTATS | 1 |
| 2019 | Geometric Sketching of Single-Cell Data Preserves Transcriptional Structure
Brian Hie, Hyunghoon Cho, Benjamin Demeo, Bryan Bryson, Bonnie Berger |
RECOMB | 2 |
| 2018 | Generalizable Visualization of Mega-Scale Single-Cell Data
Hyunghoon Cho, Bonnie Berger, Jian Peng 0001 |
RECOMB | 1 |
| 2015 | Diffusion Component Analysis: Unraveling Functional Topology in Biological Networks
Hyunghoon Cho, Bonnie Berger, Jian Peng 0001 |
RECOMB | 1 |
| 2015 | Exploiting ontology graph for predicting sparsely annotated gene functionabstractMOTIVATION: Systematically predicting gene (or protein) function based on molecular interaction networks has become an important tool in refining and enhancing the existing annotation catalogs, such as the Gene Ontology (GO) database. However, functional labels with only a few (<10) annotated genes, which constitute about half of the GO terms in yeast, mouse and human, pose a unique challenge in that any prediction algorithm that independently considers each label faces a paucity of information and thus is prone to capture non-generalizable patterns in the data, resulting in poor predictive performance. There exist a variety of algorithms for function prediction, but none properly address this 'overfitting' issue of sparsely annotated functions, or do so in a manner scalable to tens of thousands of functions in the human catalog. RESULTS: We propose a novel function prediction algorithm, clusDCA, which transfers information between similar functional labels to alleviate the overfitting problem for sparsely annotated functions. Our method is scalable to datasets with a large number of annotations. In a cross-validation experiment in yeast, mouse and human, our method greatly outperformed previous state-of-the-art function prediction algorithms in predicting sparsely annotated functions, without sacrificing the performance on labels with sufficient information. Furthermore, we show that our method can accurately predict genes that will be assigned a functional label that has no known annotations, based only on the ontology graph structure and genes associated with other labels, which further suggests that our method effectively utilizes the similarity between gene functions. AVAILABILITY AND IMPLEMENTATION: https://github.com/wangshenguiuc/clusDCA. Sheng Wang 0001, Hyunghoon Cho, ChengXiang Zhai, Bonnie Berger, Jian Peng 0001 |
Bioinform. | 2 |