EDBT 2026 Demo / reviewers in the wild / expert
Sanghyun Park 0003
dblp:45/2752-3 · also Sang-Hyun Park 0003
· DBLP profile ↗
81ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0002-5196-6193ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 31 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 6 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 8 since 2021Systems, architecture and hardware · 7 · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TheSelective: Dual Affinity-Guided Diffusion for Selective Molecular Generation
Hyoungjoon Park, Hwanhee Kim, Seungyeon Choi, Yoonju Kim, Sanghyun Park 0003 |
PAKDD (2) | 6 |
| 2026 | AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLMabstractVideo anomaly detection (VAD) is crucial for video analysis and surveillance in computer vision. However, existing VAD models rely on learned normal patterns, which makes them difficult to apply to diverse environments. Consequently, users should retrain models or develop separate AI models for new environments, which requires expertise in machine learning, high-performance hardware, and extensive data collection, limiting the practical usability of VAD. To address these challenges, this study proposes customizable video anomaly detection (C-VAD) technique and the AnyAnomaly model. C-VAD considers user-defined text as an abnormal event and detects frames containing a specified event in a video. We effectively implemented AnyAnomaly using a context-aware visual question answering without fine-tuning the large vision language model. To validate the effectiveness of the proposed model, we constructed C-VAD datasets and demonstrated the superiority of AnyAnomaly. Furthermore, our approach showed competitive results on VAD benchmarks, achieving state-of-the-art performance on UBnormal and UCF-Crime and surpassing other methods in generalization across all datasets. Our code is available online at github.com/SkiddieAhn/Paper-AnyAnomaly. Sunghyun Ahn, Youngwan Jo, Kijung Lee, Sein Kwon, Inpyo Hong, Sanghyun Park 0003 |
WACV | 6 |
| 2025 | JAM: Unlocking SAM2 Without Training and Prompts via Medical-Aware Mask SelectionabstractFoundation models allow zero-shot transfer, but SAM struggles on medical images where fine anatomy matters. We introduce JAM, a training-free one-shot prototype method that builds prototypes from a single support image and auto-generates optimal prompts at inference. We also propose PSC, a prototype similarity-coverage score that replaces confidence-based mask selection for more reliable results. JAM is plug-and-play and outperforms SAM-based and few-shot baselines on CHAOS-MRI- T2, Synapse-CT, and ETIS in both accuracy and scalability. Youngwan Jo, Sanghyun Park 0003 |
BIBM | 3 |
| 2025 | Efficient Approximate Nearest Neighbor Search via Data-Adaptive Parameter Adjustment in Hierarchical Navigable Small GraphsabstractHierarchical Navigable Small World (HNSW) graphs are a state-of-the-art solution for approximate nearest neighbor search, widely applied in areas like recommendation systems, computer vision, and natural language processing. However, the effectiveness of the HNSW algorithm is constrained by its reliance on static parameter settings, which do not account for variations in data density and dimensionality across different datasets. This paper introduces Dynamic HNSW, an adaptive method that dynamically adjusts key parameters - such as the$M$(number of connections per node) and ef (search depth) - based on both local data density and dimensionality of the dataset. The proposed approach improves flexibility and efficiency, allowing the graph to adapt to diverse data characteristics. Experimental results across multiple datasets demonstrate that Dynamic HNSW significantly reduces graph build time by up to 33.11% and memory usage by up to 32.44%, while maintaining comparable recall, thereby outperforming the conventional HNSW in both scalability and efficiency. Huijun Jin, Jieun Lee 0006, Shengmin Piao, Sangmin Seo 0001, Sein Kwon, Sanghyun Park 0003 |
DATE | 6 |
| 2025 | LatentTune: Efficient Tuning of High Dimensional Database Parameters via Latent Representation LearningabstractAs data volumes continue to grow, optimizing database performance has become increasingly critical, making the implementation of effective tuning methods essential. Among various approaches, database parameter tuning has proven to be a highly effective means of enhancing performance. Recent studies have shown that machine learning techniques can successfully optimize database parameters, leading to significant performance improvements. However, existing methods still face several limitations. First, they require substantial time to generate large training datasets. Second, to cope with the challenges of highdimensional optimization, they typically optimize only a subset of parameters rather than the full configuration space. Third, they often rely on information from similar workloads instead of directly leveraging information from the target workload. To address these limitations, we propose LatentTune, a novel approach that differs fundamentally from traditional methods. To reduce the time required for data generation, LatentTune incorporates a data augmentation strategy. Furthermore, it constructs a latent space that compresses information from all database parameters, enabling the optimization of the full configuration space. In addition, LatentTune integrates external metric information into the latent space, allowing for precise tuning tailored to the actual target workload. Experimental results demonstrate that LatentTune outperforms baseline models across four workloads on MySQL and RocksDB, achieving up to 1332 % improvement for RocksDB and 11.82 % throughput gain with 46.01 % latency reduction for MySQL. Sein Kwon, Youngwan Jo, Seungyeon Choi, Jieun Lee 0006, Huijun Jin, Sanghyun Park 0003 |
HiPC | 6 |
| 2025 | TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-ReflectionabstractShengmin Piao, Sanghyun Park. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shengmin Piao, Sanghyun Park 0003 |
NAACL (Long Papers) | 2 |
| 2025 | MDVAD: Multimodal Diffusion for Video Anomaly Detection
Kijung Lee, Youngwan Jo, Sunghyun Ahn, Sanghyun Park 0003 |
PAKDD (1) | 4 |
| 2024 | VideoPatchCore: An Effective Method to Memorize Normality for Video Anomaly Detection
Sunghyun Ahn, Youngwan Jo, Kijung Lee, Sanghyun Park 0003 |
ACCV (3) | 4 |
| 2024 | PretrainedBA: Enhancing Compound-Protein Binding Affinity Prediction Accuracy via Pre-training Large-Scale Interaction InformationabstractFinding potential drug candidates with high binding affinity for the specific target protein presents an important goal in early drug discovery. Although compound-protein complex structure-based affinity prediction methods have shown promising prediction accuracy, their dependency on high-resolution three-dimensional (3D) complex structure data considerably limits their practical application. Alternatively, many complex-free binding affinity prediction methods have been proposed; however, there is still room for improvement to compensate effectively for the lack of binding information. In particular, the interpretability of compound-protein interactions is a significant challenge that needs to be addressed. To alleviate the limitations of current complex-free models, we propose PretrainedBA, a predictive model that uses pre-training strategies on large-scale datasets, including interaction data. PretrainedBA pre-trains the interdependent relationships between compounds and proteins, rather than the independent pre-training of compounds and proteins utilized in existing studies. PretrainedBA consists of six modules and is designed to effectively model compound-protein interactions within identified binding pockets. Comparisons with state-of-the-art complex-free models on seven external benchmark datasets demonstrate that this pre-training strategy improves binding affinity prediction accuracy. In particular, the outstanding interpretive power of compound-protein interaction mechanisms compared with the previous method further emphasizes the value of PretrainedBA. Real-world application evaluation using the Database of Useful Decoys-Enhanced (DUD-E) dataset confirmed PretrainedBA’s practical applicability, demonstrating its utility in drug discovery. Sangmin Seo 0001, Seungyeon Choi, Hwanhee Kim, Sanghyun Park 0003 |
BIBM | 4 |
| 2024 | SPIN: SE(3)-Invariant Physics Informed Network for Binding Affinity PredictionabstractAccurate prediction of protein-ligand binding affinity is crucial for rapid and efficient drug development. Recently, the importance of predicting binding affinity has led to increased attention on research that models the three-dimensional structure of protein-ligand complexes using graph neural networks to predict binding affinity. However, traditional methods often fail to accurately model the complex’s spatial information or rely solely on geometric features, neglecting the principles of protein-ligand binding. This can lead to overfitting, resulting in models that perform poorly on independent datasets and ultimately reducing their usefulness in real drug development. To address this issue, we propose SPIN, a model designed to achieve superior generalization by incorporating various inductive biases applicable to this task, beyond merely training on empirical data from datasets. For prediction, we defined two types of inductive biases: a geometric perspective that maintains consistent binding affinity predictions regardless of the complex’s rotations and translations, and a physicochemical perspective that necessitates minimal binding free energy along their reaction coordinate for effective protein-ligand binding. These prior knowledge inputs enable the SPIN to outperform comparative models in benchmark sets such as CASF-2016 and CSAR HiQ. Furthermore, we demonstrated the practicality of our model through virtual screening experiments and validated the reliability and potential of our proposed model based on experiments assessing its interpretability. Seungyeon Choi, Sangmin Seo 0001, Sanghyun Park 0003 |
ECAI | 3 |
| 2024 | DrDiff: Drug Response Prediction Through Controllable Diffusion-GE and Graph Attention NetworkabstractThe accurate prediction of drug responses based on the genomic profile of a patient is essential to progress in the field of precision medicine. The advent of various deep-learning algorithms based on publicly available large-scale omics datasets is the driving force behind research in this field. The characteristics of biological datasets, characterized by high dimensions and low sample sizes, pose challenges of overfitting and limited generalization in prediction models. Additionally, constructing prediction models using biological data such as gene expression is further complicated by the need to account for the complex relationships among genes, which exacerbates the aforementioned challenges. To address these challenges, we propose a drug response prediction framework (DrDiff) that integrates a denoising diffusion probabilistic model (DDPM) based data augmentation module with a graph attention network based drug response prediction module. The proposed model showed a 10% higher AUC than the state-of-the-art models for drug response prediction for the six drugs considered in the study, suggesting the superior generalization performance of DrDiff over other baseline models. Furthermore, we demonstrated the feasibility of generative models, which form one of the modules of the proposed framework, in overcoming the fundamental limitations of omics datasets. Further experiments bear out the feasibility of generative models, which form one of the modules of the proposed framework, in augmenting gene expression data. Seungyeon Choi, Sangmin Seo 0001, Jonghwan Choi, Chihyun Park, Sanghyun Park 0003 |
ECAI | 5 |
| 2024 | Towards Workload-Specific Configuration Tuning via Meta-Learning for RocksDBabstractA persistent key-value store, RocksDB, is adapt-able to various workloads and provides fast and low-latency storage for devices that are utilized by numerous applications. RocksDB has been introduced with numerous configuration options for customization and performance optimization. Unfor-tunately, determining an optimal configuration for each given workload remains challenging due to the overwhelming number of options. This complexity is compounded by different types of workloads, thereby requiring efficient configuration tuning. Recent studies have approached automatic tuning techniques to solve this problem by applying reinforcement learning approaches or transferring prior knowledge to predictive models in order to tune unobserved target workloads. However, the former method is time-consuming, and the latter results in unstable optimal performance according to the accuracy of the predictive models. The models trained with prior knowledge, estimate RocksDB performance of given configurations on the target workload, where those workload mismatches degrade tuning performance. To address these challenges, we propose MetaTune, which introduces a meta learner, which is a meta-learning technique, to train a workload -specific predictive model. MetaTune effectively transfers prior knowledge and effi-ciently fine-tunes the model for new workloads. We conducted a comparative analysis of MetaTune with the state-of-the-art baselines across a heterogeneous set of workloads. MetaTune achieved 3.78% to 53.25% improvement in tuning performance compared to the most recent baseline. Chanho Yearn, Jieun Lee 0006, Sangmin Seo 0001, Sanghyun Park 0003 |
SMC | 4 |
| 2024 | Pseq2Sites: Enhancing protein sequence-based ligand binding-site prediction accuracy via the deep convolutional network and attention mechanismabstractProtein-ligand interactions play an essential role in many biological processes, and prior knowledge of ligand binding sites is necessary for successful drug design. Many 3D structure- and sequence-based methods have been proposed for identifying ligand binding sites. The 3D structure-based methods typically achieve better binding site prediction than the sequence-based methods. However, as deep-learning techniques that can extract structural information from large-scale sequence data have been developed, the performance gap between 3D structure- and sequence-based methods is narrowing. Nonetheless, there remains room for improvement in sequence-based prediction. We propose Pseq2Sites, a sequence-based deep-learning model for predicting ligand binding sites. Pseq2Sites comprises a 1D convolutional neural network that extracts local features from the protein sequence, and a position-based attention mechanism that captures long-distance dependencies between binding residues. To verify the effectiveness of the proposed method, we compared it with other state-of-the-art methods using three public datasets: COACH420, HOLO4K, and CSAR-NRC HiQ. Utilizing solely protein sequence information, Pseq2Sites outperformed 3D structure-based state-of-the-art methods on external test datasets; within the COACH420 dataset, Pseq2Sites remarkably identified 97% of the binding pockets (at a significance level δ = 0.5), which was 27% higher than the second highest-performing model. Pseq2Sites also achieved outstanding binding site prediction, even for proteins with low similarity to the training dataset. Our code is available at https://github.com/Blue1993/Pseq2Sites. Sangmin Seo 0001, Jonghwan Choi, Seungyeon Choi, Jieun Lee 0006, Chihyun Park, Sanghyun Park 0003 |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | K2vTune: A workload-aware configuration tuning for RocksDB
Jieun Lee 0006, Sangmin Seo 0001, Jonghwan Choi, Sanghyun Park 0003 |
Inf. Process. Manag. | 4 |
| 2023 | SELF-EdiT: Structure-constrained molecular optimisation using SELFIES editing transformerabstractAbstract Structure-constrained molecular optimisation aims to improve the target pharmacological properties of input molecules through small perturbations of the molecular structures. Previous studies have exploited various optimisation techniques to satisfy the requirements of structure-constrained molecular optimisation tasks. However, several studies have encountered difficulties in producing property-improved and synthetically feasible molecules. To achieve both property improvement and synthetic feasibility of molecules, we proposed a molecular structure editing model called SELF-EdiT that uses self-referencing embedded strings (SELFIES) and Levenshtein transformer models. The SELF-EdiT generates new molecules that resemble the seed molecule by iteratively applying fragment-based deletion-and-insertion operations to SELFIES. The SELF-EdiT exploits a grammar-based SELFIES tokenization method and the Levenshtein transformer model to efficiently learn deletion-and-insertion operations for editing SELFIES. Our results demonstrated that SELF-EdiT outperformed existing structure-constrained molecular optimisation models by a considerable margin of success and total scores on the two benchmark datasets. Furthermore, we confirmed that the proposed model could improve the pharmacological properties without large perturbations of the molecular structures through edit-path analysis. Moreover, our fragment-based approach significantly relieved the SELFIES collapse problem compared to the existing SELFIES-based model. SELF-EdiT is the first attempt to apply editing operations to the SELFIES to design an effective editing-based optimisation, which can be helpful for fellow researchers planning to utilise the SELFIES. Shengmin Piao, Jonghwan Choi, Sangmin Seo 0001, Sanghyun Park 0003 |
Appl. Intell. | 4 |
| 2023 | DARK: Deep automatic Redis knobs tuning system depending on the persistence method
Ju Yeon Seo, Kyeonghun Kim, Sangmin Seo 0001, Sanghyun Park 0003 |
Expert Syst. Appl. | 4 |
| 2023 | NCMD: Node2vec-Based Neural Collaborative Filtering for Predicting MiRNA-Disease AssociationabstractNumerous studies have reported that micro RNAs (miRNAs) play pivotal roles in disease pathogenesis based on the deregulation of the expressions of target messenger RNAs. Therefore, the identification of disease-related miRNAs is of great significance in understanding human complex diseases, which can also provide insight into the design of novel prognostic markers and disease therapies. Considering the time and cost involved in wet experiments, most recent works have focused on the effective and feasible modeling of computational frameworks to uncover miRNA-disease associations. In this study, we propose a novel framework called node2vec-based neural collaborative filtering for predicting miRNA-disease association (NCMD) based on deep neural networks. Initially, NCMD exploits Node2vec to learn low-dimensional vector representations of miRNAs and diseases. Next, it utilizes a deep learning framework that combines the linear ability of generalized matrix factorization and nonlinear ability of a multilayer perceptron. Experimental results clearly demonstrate the comparable performance of NCMD relative to the state-of-the-art methods according to statistical measures. In addition, case studies on breast cancer, lung cancer and pancreatic cancer validate the effectiveness of NCMD. Extensive experiments demonstrate the benefits of modeling a neural collaborative-filtering-based approach for discovering novel miRNA-disease associations. Jihwan Ha, Sanghyun Park 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | RTune: a RocksDB tuning system with deep genetic algorithmabstractDatabase systems typically have many knobs that must be configured by database administrators to achieve high performance. RocksDB achieves fast data writing performance using a log-structured merge-tree. This database contains many knobs related to write and space amplification, which are important performance indicators in RocksDB. Previously, it was proved that significant performance improvements could be achieved by tuning database knobs. However, tuning multiple knobs simultaneously is a laborious task owing to the large number of potential configuration combinations and trade-offs. Huijun Jin, Jieun Lee 0006, Sanghyun Park 0003 |
GECCO | 3 |
| 2022 | DeepGate: Global-local decomposition for multivariate time series modeling
Jinuk Park, Chanhee Park, Jonghwan Choi, Sanghyun Park 0003 |
Inf. Sci. | 4 |
| 2021 | MolBit: De novo Drug Design via Binary Representations of SMILES for avoiding the Posterior Collapse ProblemabstractDeep generative models for molecular generation have accelerated the development of de novo drug design by introducing how to generate novel molecular structures expressed in simplified molecular-input line-entry system (SMILES) or molecular graph formats. Numerous drug design studies have proposed combinations of variational autoencoder (VAE) and autoregressive generators such as recurrent neural networks (RNNs) to generate SMILES strings. However, RNN-VAE has one notorious issue, called posterior collapse, in which different latent vectors produce indistinguishable molecular distributions. In this study, we proposed a Gumbel-Softmax-based generative model, MolBit, and a genetic algorithm-based molecular property optimization method. We confirmed that the proposed model avoided the posterior collapse problem and outperformed the existing drug design models with SMILES. Jonghwan Choi, Sangmin Seo 0001, Jinuk Park, Sanghyun Park 0003 |
BIBM | 4 |
| 2021 | Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactionsabstractBACKGROUND: Accurate prediction of protein-ligand binding affinity is important for lowering the overall cost of drug discovery in structure-based drug design. For accurate predictions, many classical scoring functions and machine learning-based methods have been developed. However, these techniques tend to have limitations, mainly resulting from a lack of sufficient energy terms to describe the complex interactions between proteins and ligands. Recent deep-learning techniques can potentially solve this problem. However, the search for more efficient and appropriate deep-learning architectures and methods to represent protein-ligand complex is ongoing. RESULTS: In this study, we proposed a deep-neural network model to improve the prediction accuracy of protein-ligand complex binding affinity. The proposed model has two important features, descriptor embeddings with information on the local structures of a protein-ligand complex and an attention mechanism to highlight important descriptors for binding affinity prediction. The proposed model performed better than existing binding affinity prediction models on most benchmark datasets. CONCLUSIONS: We confirmed that an attention mechanism can capture the binding sites in a protein-ligand complex to improve prediction performance. Our code is available at https://github.com/Blue1993/BAPA . Sangmin Seo 0001, Jonghwan Choi, Sanghyun Park 0003, Jaegyoon Ahn |
BMC Bioinform. | 3 |
| 2021 | OurRocks: Offloading Disk Scan Directly to GPU in Write-Optimized Database SystemabstractThe log structured merge (LSM) tree has been widely adopted by database systems owing to its superior write performance. However, LSM-tree based databases face vulnerabilities when processing analytical queries due to the read amplification caused by its architecture and the limited use of storage devices with high bandwidth. To flexibly handle transactional and analytical workloads, we proposed and implemented OurRocks taking full advantage of NVMe SSD and GPU devices, which improves scan performance. Although the NVMe SSD serves multi GB/s I/O rates, it is necessary to solve the data transfer overhead which limits the benefits of the GPU processing. The primary idea is to offload the scan operation to the GPU with filtering predicate pushdown and resolve the bottleneck from the data transfer between devices with direct memory access (DMA). OurRocks benefits from all the features of write-optimized database systems, in addition to accelerating the analytic queries using the aforementioned idea. Experimental results indicate that OurRocks effectively leverages resources of the NVMe SSD and GPU and significantly improves the execution of queries in the YCSB and TPC-H benchmarks, compared to the conventional write-optimized database. Our research demonstrates that the proposed approach can speed up the handling of the data-intensive workloads. Won Gi Choi, Hongchan Roh, Sanghyun Park 0003 |
IEEE Trans. Computers | 4 |
| 2020 | Prediction of Alzheimer's disease based on deep neural network by integrating gene expression and DNA methylation dataset
Chihyun Park, Jihwan Ha, Sanghyun Park 0003 |
Expert Syst. Appl. | 3 |
| 2020 | Combinatorial feature embedding based on CNN and LSTM for biomedical named entity recognition
Minsoo Cho, Jihwan Ha, Chihyun Park, Sanghyun Park 0003 |
J. Biomed. Informatics | 4 |
| 2020 | IMIPMF: Inferring miRNA-disease interactions using probabilistic matrix factorization
Jihwan Ha, Chihyun Park, Chanyoung Park 0001, Sanghyun Park 0003 |
J. Biomed. Informatics | 4 |
| 2019 | ADC: Advanced document clustering using contextualized representations
Jinuk Park, Chanhee Park, Jeongwoo Kim 0002, Minsoo Cho, Sanghyun Park 0003 |
Expert Syst. Appl. | 5 |
| 2019 | A write-friendly approach to manage namespace of Hadoop distributed file system by utilizing nonvolatile memory
Won Gi Choi, Sanghyun Park 0003 |
J. Supercomput. | 2 |
| 2018 | Classifying malwares for identification of author groupsabstractSummary Malwares are growing exponentially in number, and authors of malwares are continuously releasing new ones. Malwares developed by the same author group might have similar signatures. For a number of applications including digital forensic and law enforcement, such characteristics can be used to determine which author group is likely to have released a given malware. In this paper, we describe a new type of classification that identifies which group of authors is most likely to have developed a given malware. We identify and verify a set of various features obtained through static and dynamic analyses of malwares and exploit them for classification. We evaluate our approach through extensive experiments with a real‐world dataset labeled by a group of domain experts. The results show that our approach is effective and provides good accuracy in malware classification. Jiwon Hong, Sanghyun Park 0003, Sang-Wook Kim, Dongphil Kim, Wonho Kim |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | Selective I/O Bypass and Load Balancing Method for Write-Through SSD Caching in Big Data AnalyticsabstractFast network quality analysis in the telecom industry is an important method used to provide quality service. SK Telecom, based in South Korea, built a Hadoop-based analytical system consisting of a hundred nodes, each of which only contains hard disk drives (HDDs). Because the analysis process is a set of parallel I/O intensive jobs, adding solid state drives (SSDs) with appropriate settings is the most cost-efficient way to improve the performance, as shown in previous studies. Therefore, we decided to configure SSDs as a write-through cache instead of increasing the number of HDDs. To improve the cost-per-performance of the SSD cache, we introduced a selective I/O bypass (SIB) method, redirecting the automatically calculated number of read I/O requests from the SSD cache to idle HDDs when the SSDs are I/O over-saturated, which means the disk utilization is greater than 100 percent. To precisely calculate the disk utilization, we also introduced a combinational approach for SSDs because the current method used for HDDs cannot be applied to SSDs because of their internal parallelism. In our experiments, the proposed approach achieved a maximum 2x faster performance than other approaches. Hongchan Roh, Sanghyun Park 0003 |
IEEE Trans. Computers | 3 |
| 2018 | GSEH: A Novel Approach to Select Prostate Cancer-Associated Genes Using Gene Expression HeterogeneityabstractWhen a gene shows varying levels of expression among normal people but similar levels in disease patients or shows similar levels of expression among normal people but different levels in disease patients, we can assume that the gene is associated with the disease. By utilizing this gene expression heterogeneity, we can obtain additional information that abets discovery of disease-associated genes. In this study, we used collaborative filtering to calculate the degree of gene expression heterogeneity between classes and then scored the genes on the basis of the degree of gene expression heterogeneity to find "differentially predicted" genes. Through the proposed method, we discovered more prostate cancer-associated genes than 10 comparable methods. The genes prioritized by the proposed method are potentially significant to biological processes of a disease and can provide insight into them. Sang-Min Choi, Sanghyun Park 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | MV-FTL: An FTL That Provides Page-Level Multi-Version ManagementabstractIn this paper, we propose MV-FTL, a multi-version flash transition layer (FTL) that provides page-level multi-version management. By extending a unique characteristic of solid-state drives (SSDs), the out-of-place (OoP) update to multi-version management, MV-FTL can both guarantee atomic page updates from each transaction and provide concurrency without requiring redundant log data writes as well. For evaluation, we first modified SQLite, a lightweight database management system (DBMS), to cooperate with MV-FTL. Owing to the architectural simplicity of SQLite, we clearly show that MV-FTL improves both the performance and the concurrency aspects of the system. In addition, to prove the effectiveness in a full-fledged enterprise-level DBMS, we modified MyRocks, a MySQL variant by Facebook, to use our new Patch Compaction algorithm, which deeply relies on MV-FTL. The TPC-C and LinkBench benchmark tests demonstrated that MV-FTL reduces the overall amount of writes, implying that MV-FTL can be effective in such DBMSs. Doogie Lee, Won Gi Choi, Hongchan Roh, Sanghyun Park 0003 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2017 | Machine Learning based Performance Modeling of Flash SSDsabstractFlash memory based solid state drives(SSDs) have alleviated the I/O bottleneck by exploiting its data parallel design. In an enterprise environment, Flash SSD used in the form of a hybrid storage architecture to achieve the better performance with lower cost. In this architecture, I/O load balancing is one of the important factors. However, the internal parallelism distorts the performance measures of the flash SSDs. Despite the criticality of load balancing on I/O intensive environments, these studies have rarely been addressed. In this paper, we examine the effectiveness of applying classification method using machine learning techniques to the I/O saturation estimation by using Linux kernel I/O statistics instead of the utilization measure that is currently used for HDDs. We conclude that machine learning techniques that we employed (Support Vector Machine and LASSO Generalized Linear Model) performs well compared to the existing utilization measure even we cannot collect the internal information of the flash SSDs. Jinuk Park, Sanghyun Park 0003 |
CIKM | 3 |
| 2017 | Improved prediction of breast cancer outcome by identifying heterogeneous biomarkersabstractMOTIVATION: Identification of genes that can be used to predict prognosis in patients with cancer is important in that it can lead to improved therapy, and can also promote our understanding of tumor progression on the molecular level. One of the common but fundamental problems that render identification of prognostic genes and prediction of cancer outcomes difficult is the heterogeneity of patient samples. RESULTS: To reduce the effect of sample heterogeneity, we clustered data samples using K-means algorithm and applied modified PageRank to functional interaction (FI) networks weighted using gene expression values of samples in each cluster. Hub genes among resulting prioritized genes were selected as biomarkers to predict the prognosis of samples. This process outperformed traditional feature selection methods as well as several network-based prognostic gene selection methods when applied to Random Forest. We were able to find many cluster-specific prognostic genes for each dataset. Functional study showed that distinct biological processes were enriched in each cluster, which seems to reflect different aspect of tumor progression or oncogenesis among distinct patient groups. Taken together, these results provide support for the hypothesis that our approach can effectively identify heterogeneous prognostic genes, and these are complementary to each other, improving prediction accuracy. AVAILABILITY AND IMPLEMENTATION: https://github.com/mathcom/CPR. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonghwan Choi, Sanghyun Park 0003, Youngmi Yoon, Jaegyoon Ahn |
Bioinform. | 2 |
| 2017 | IMA: Identifying disease-related genes using MeSH terms and association rules
Jeongwoo Kim 0002, Changbae Bang, Hyeonseo Hwang, Chihyun Park, Sanghyun Park 0003 |
J. Biomed. Informatics | 6 |
| 2017 | Advanced Block Nested Loop Join for Extending SSD LifetimeabstractFlash technology trends have shown that greater densities between flash memory cells increase read/write error rates and shorten solid-state drive (SSD) device lifetimes. This is critical for enterprise systems, causing such problems as service instability and increased total cost of ownership (TCO) because of SSD replacement. Therefore, numerous studies have focused on decreasing the amount of the DBMS writes. However, there has been no research that focused on decreasing the amount of temporary writes, which are primarily created by join processing. In DBMSs, there are two major join-processing algorithms, i.e., hybrid hash join (HHJ) and sort merge join (SMJ), proven to be the best according to DBMS workload; however, the two algorithms produce temporary writes of intermediate results. Therefore, we instead look to the block-nested loop join (BNLJ); it is well-known that the two algorithms are better than BNLJ, but BNLJ creates no intermediate result writes. It is reasonable to use BNLJ for a major join algorithm if its performance can be enhanced similar to those of HHJ and SMJ, considering BNLJ's advantage of extending SSD lifetimes. Therefore, in this paper, we propose an advanced BNLJ (ANLJ) algorithm that can match the performance of the two main join algorithms. Hongchan Roh, Wonmook Jung, Sanghyun Park 0003 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | Optimization of a multiversion index on SSDs to improve system performanceabstractIn this paper, we propose a multiversion index utilizing key features of SSDs (solid state drives). SSDs have many advantages, e.g., fast read/write performance, high energy efficiency, and non-volatility. Thus, SSDs have been considered for several years as a promising alternative to HDDs (hard disk drives). Many studies have made progress in optimizing and modifying HDD-based database management system (DBMS) to suit SSDs. In the case of multiversion databases, which manage not only keys and but also versions, research optimizing SSD query processing has been ignored in comparison with single versioned databases. Generally, the multiversion databases manage an evolving data which is processed in a cyber physical system or an accounting system. Therefore, the data is large and the index structure requires frequent rearrangement of its structure to maximize efficiency, which is called structure modification operation. The multiversion index based on HDDs utilizes random writes to conduct the structure modification operation. This feature can introduce crucial performance problems on SSDs, because the speed of random writes on SSDs is much slower than the speed of sequential writes. We propose a Bulk Split multiversion tree (BSMVBT) index that utilizes sequential pattern I/O and out-of-place updates of SSDs. Experimental results showed that it is 10% – 30% faster than the compared version. Won Gi Choi, Doogie Lee, Sanghyun Park 0003 |
SMC | 5 |
| 2016 | DSS: A biclustering method to identify diverse and state specific gene modules in gene expression dataabstractThe biclustering method is a useful co-clustering technique to identify biologically relevant gene modules. In this paper, we propose a novel method to find not only functionally-related gene modules but also state specific gene modules by applying a genetic algorithm to gene expression data. To identify these gene modules, the proposed method finds biclusters in which genes are statistically overexpressed or under expressed, and are differentially-expressed in the samples in the bicluster compared to the samples not in the bicluster. In addition, we improve the genetic algorithm by adding a selection pool for preserving the diversity of the population. The resulting gene modules exhibit better performances than comparative methods in the GO (Gene Ontology) term enrichment test and an analysis connection between gene modules and disease. This is especially the case with gene modules that receive the highest score in the breast cancer dataset; they are closely linked to the ribosome pathway. Recent studies show that dysregulation of ribosome biogenesis is associated with breast tumor progression. Jungrim Kim, Yunku Yeu, Jeongwoo Kim 0002, Youngmi Yoon, Sanghyun Park 0003 |
SMC | 5 |
| 2016 | External Mergesort for Flash-Based Solid State DrivesabstractMergesort is the most widely-known external sorting algorithm, which is used when the data being sorted do not fit into the available main memory. There have been several attempts to improve mergesort by reducing I/O time, since mergesort is I/O intensive. However, these methods assumed that mergesort runs on hard disk drives (HDDs). Flash-based solid state drives (SSDs) are emerging as next generation storage devices and becoming alternatives to HDDs. SSDs outperform HDDs in access latency, because they have no physical arms to move. In addition, SSDs benefit from their inner structure by exploiting internal parallelism, resulting in high I/O bandwidth. Previous methods for improving mergesort focused on reducing random access cost, which is insignificant on SSDs. In this paper we propose an external mergesort algorithm for SSDs called FMsort. FMsort calculates a block read order which is the order of blocks needed in the merge phase. With a block read order, a number of blocks required during the merge phase are read into main memory via multiple asynchronous I/Os. Our experiments show that FMsort outperforms other mergesort algorithms, at an invisible cost of calculating a block read order. Hongchan Roh, Sanghyun Park 0003 |
IEEE Trans. Computers | 3 |
| 2015 | Inverted index maintenance strategy for flashSSDs: Revitalization of in-place index update strategy
Wonmook Jung, Hongchan Roh, Sanghyun Park 0003 |
Inf. Syst. | 4 |
| 2015 | BulkAligner: A novel sequence alignment algorithm based on graph theory and Trinity
Junsu Lee, Yunku Yeu, Hongchan Roh, Youngmi Yoon, Sanghyun Park 0003 |
Inf. Sci. | 5 |
| 2015 | LGscore: A method to identify disease-related genes using biological literature and Google data
Jeongwoo Kim 0002, Youngmi Yoon, Sanghyun Park 0003 |
J. Biomed. Informatics | 4 |
| 2014 | Discovering phenotype specific gene module using a novel biclustering algorithm in colorectal cancerabstractGene clustering is a method for finding gene sets which are related to the same biological processes or molecular function. In order to find these gene sets, previous studies have clustered genes which showed similar mRNA expression or a specific expression pattern in a (sub) sample set. However, for two contrasting groups of samples, it is not easy to identify gene sets which show significant expression pattern in only one group using current gene clustering methods. Existing biclustering methods use only one group (disease) of samples. It is hard to identify disease specific biclusters which are differentially expressed in the disease although those methods can find biclusters which have specific expression pattern. Here, we proposed a novel method using a genetic algorithm in gene expression data, in order to find gene sets which can represent specific subtype of cancer. Proposed method finds gene sets which have statistically differential mRNA expression on two contrasting samples and fraction of cancer samples. The resulting gene modules share higher number of GO (Gene Ontology) terms related to a specific disease than gene modules identified by current algorithms. We also identify that when we integrate protein-protein interaction data with gene expression data of colorectal cancer samples, proposed method can find more functionally related gene sets. Jungrim Kim, Youngmi Yoon, Sanghyun Park 0003, Jaegyoon Ahn, Yunku Yeu |
BIBM | 3 |
| 2013 | Preface
Sang-Wook Kim, Sanghyun Park 0003, Haixun Wang |
J. Comput. Sci. Technol. | 3 |
| 2013 | CORE: Common Region Extension Based Multiple Protein Structure Alignment for Producing Multiple Solution
Woo-Cheol Kim, Sanghyun Park 0003, Jung-Im Won |
J. Comput. Sci. Technol. | 2 |
| 2012 | Identification of functional CNV region networks using a CNV-gene mapping algorithm in a genome-wide scaleabstractMOTIVATION: Identifying functional relation of copy number variation regions (CNVRs) and gene is an essential process in understanding the impact of genotypic variations on phenotype. There have been many related works, but only a few attempts were made to normal populations. RESULTS: To analyze the functions of genome-wide CNVRs, we applied a novel correlation measure called Correlation based on Sample Set (CSS) to paired Whole Genome TilePath array and messenger RNA (mRNA) microarray data from 210 HapMap individuals with normal phenotypes and calculated the confident CNVR-gene relationships. Two CNVR nodes form an edge if they regulate a common set of genes, allowing the construction of a global CNVR network. We performed functional enrichment on the common genes that were trans-regulated from CNVRs clustered together in our CNVR network. As a result, we observed that most of CNVR clusters in our CNVR network were reported to be involved in some biological processes or cellular functions, while most CNVR clusters from randomly constructed CNVR networks showed no evidence of functional enrichment. Those results imply that CSS is capable of finding related CNVR-gene pairs and CNVR networks that have functional significance. AVAILABILITY: http://embio.yonsei.ac.kr/~ Park/cnv_net.php. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chihyun Park, Jaegyoon Ahn, Youngmi Yoon, Sanghyun Park 0003 |
Bioinform. | 4 |
| 2011 | Integrative gene network construction for predicting a set of complementary prostate cancer genesabstractMOTIVATION: Diagnosis and prognosis of cancer and understanding oncogenesis within the context of biological pathways is one of the most important research areas in bioinformatics. Recently, there have been several attempts to integrate interactome and transcriptome data to identify subnetworks that provide limited interpretations of known and candidate cancer genes, as well as increase classification accuracy. However, these studies provide little information about the detailed roles of identified cancer genes. RESULTS: To provide more information to the network, we constructed the network by incorporating genetic interactions and manually curated gene regulations to the protein interaction network. To make our newly constructed network cancer specific, we identified edges where two genes show different expression patterns between cancer and normal phenotypes. We showed that the integration of various datasets increased classification accuracy, which suggests that our network is more complete than a network based solely on protein interactions. We also showed that our network contains significantly more known cancer-related genes than other feature selection algorithms. Through observations of some examples of cancer-specific subnetworks, we were able to predict more detailed and interpretable roles of oncogenes and other cancer candidate genes in the prostate cancer cells. AVAILABILITY: http://embio.yonsei.ac.kr/~Ahn/tc.php. CONTACT: [email protected] Jaegyoon Ahn, Youngmi Yoon, Chihyun Park, Eunji Shin, Sanghyun Park 0003 |
Bioinform. | 5 |
| 2011 | Noise-robust algorithm for identifying functionally associated biclusters from gene expression data
Jaegyoon Ahn, Youngmi Yoon, Sanghyun Park 0003 |
Inf. Sci. | 3 |
| 2011 | B+-tree Index Optimization by Exploiting Internal Parallelism of Flash-based Solid State DrivesabstractPrevious research addressed the potential problems of the hard-disk oriented design of DBMSs of flashSSDs. In this paper, we focus on exploiting potential benefits of flashSSDs. First, we examine the internal parallelism issues of flashSSDs by conducting benchmarks to various flashSSDs. Then, we suggest algorithm-design principles in order to best benefit from the internal parallelism. We present a new I/O request concept, called psync I/O that can exploit the internal parallelism of flashSSDs in a single process. Based on these ideas, we introduce B+-tree optimization methods in order to utilize internal parallelism. By integrating the results of these methods, we present a B+-tree variant, PIO B-tree. We confirmed that each optimization method substantially enhances the index performance. Consequently, PIO B-tree enhanced B+-tree's insert performance by a factor of up to 16.3, while improving point-search performance by a factor of 1.2. The range search of PIO B-tree was up to 5 times faster than that of the B+-tree. Moreover, PIO B-tree outperformed other flash-aware indexes in various synthetic workloads. We also confirmed that PIO B-tree outperforms B+-tree in index traces collected inside the Postgresql DBMS with TPC-C benchmark. Hongchan Roh, Sanghyun Park 0003, Sang-Won Lee 0001 |
Proc. VLDB Endow. | 2 |
| 2010 | Yet another write-optimized DBMS layer for flash-based solid state storageabstractFlash-based Solid State Storage (flashSSS) has write-oriented problems such as low write throughput, and limited life-time. Especially, flashSSDs have a characteristic vulnerable to random-writes, due to its control logic utilizing parallelism between the flash memory chips. In this paper, we present a write-optimized layer of DBMSs to address the write-oriented problems of flashSSS in on-line transaction processing environments. The layer consists of a write-optimized buffer, a corresponding log space, and an in-memory mapping table, closely associated with a novel logging scheme called InCremental Logging (ICL). The ICL scheme enables DBMSs to reduce page-writes at the least expense of additional page-reads, while replacing random-writes into sequential-writes. Through experiments, our approach demonstrated up-to an order of magnitude performance enhancement in I/O processing time compared to the original DBMS, increasing the longevity of flashSSS by approximately a factor of two. Hongchan Roh, Daewook Lee, Sanghyun Park 0003 |
CIKM | 3 |
| 2010 | Efficient processing of spatial joins with DOT-based indexing
Hyun Back, Jung-Im Won, Jeehee Yoon, Sanghyun Park 0003, Sang-Wook Kim |
Inf. Sci. | 4 |
| 2010 | Microarray Data Classifier Consisting of k-Top-Scoring Rank-Comparison Decision Rules With a Variable Number of GenesabstractMicroarray experiments generate quantitative expression measurements for thousands of genes simultaneously, which is useful for phenotype classification of many diseases. Our proposed phenotype classifier is an ensemble method with k-top-scoring decision rules. Each rule involves a number of genes, a rank comparison relation among them, and a class label. Current classifiers, which are also ensemble methods, consist of k-top-scoring decision rules. Some of these classifiers fix the number of genes in each rule as a triple or a pair. In this paper, we generalize the number of genes involved in each rule. The number of genes in each rule ranges from 2 to N, respectively. Generalizing the number of genes increases the robustness and the reliability of the classifier for the class prediction of an independent sample. Our algorithm saves resources by combining shorter rules in order to build a longer rule. It converges rapidly toward its high-scoring rule list by implementing several heuristics. The parameter k is determined by applying leave-one-out cross validation to the training dataset. Youngmi Yoon, Sangjay Bien, Sanghyun Park 0003 |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2009 | A Computational Approach to Detect CNVs Using High-throughput SequencingabstractCopy-Number Variations (CNVs) can be defined as gains or losses that are greater than 1kbs of genomic DNA among phenotypically normal individuals. CNVs detected by microarray based approach are limited to medium or large sized ones because of its low resolution. Here we propose a novel approach to detect CNVs by aligning the short reads obtained by high-throughput sequencer to the previously assembled human genome sequence, and analyzing the distribution of the aligned reads. Application of our algorithm demonstrates the feasibility of detecting CNVs of arbitrary length, which include short ones that microarray based algorithms cannot detect. Also, false positive and false negative rates of the results were relatively low compared to those of microarray based algorithms. Myungjin Moon, Jaegyoon Ahn, Youngmi Yoon, Chihyun Park, Sanghyun Park 0003, Jeehee Yoon |
BIBE | 5 |
| 2009 | HAPLOWSER: a whole-genome haplotype browser for personal genome and metagenomeabstractSUMMARY: Haplotype assembly is becoming a very important tool in genome sequencing of human and other organisms. Although haplotypes were previously inferred from genome assemblies, there has never been a comparative haplotype browser that depicts a global picture of whole-genome alignments among haplotypes of different organisms. We introduce a whole-genome HAPLotype brOWSER (HAPLOWSER), providing evolutionary perspectives from multiple aligned haplotypes and functional annotations. Haplowser enables the comparison of haplotypes from metagenomes, and associates conserved regions or the bases at the conserved regions with functional annotations and custom tracks. The associations are quantified for further analysis and presented as pie charts. Functional annotations and custom tracks that are projected onto haplotypes are saved as multiple files in FASTA format. Haplowser provides a user-friendly interface, and can display alignments of haplotypes with functional annotations at any resolution. AVAILABILITY: Haplowser, written in Java, supports multiple platforms including Windows and Linux. Haplowser is publicly available at http://embio.yonsei.ac.kr/haplowser . Woo-Cheol Kim, Michael S. Waterman, Sanghyun Park 0003, Lei M. Li |
Bioinform. | 4 |
| 2009 | A stock recommendation system exploiting rule discovery in stock databases
You-min Ha, Sanghyun Park 0003, Sang-Wook Kim, Jung-Im Won, Jeehee Yoon |
Inf. Softw. Technol. | 2 |
| 2009 | A B-Tree index extension to enhance response time and the life cycle of flash memory
Hongchan Roh, Woo-Cheol Kim, Seung-Woo Kim, Sanghyun Park 0003 |
Inf. Sci. | 4 |
| 2008 | A novel evolutionary algorithm for bi-clustering of gene expression data based on the Order Preserving Sub-Matrix (OPSM) constraintabstractBiclustering is a popular method which can reveal unknown genetic pathways. However, even though many algorithms have been suggested, no overwhelming algorithm has been suggested, due to its significant search space, until now. In this respect, several evolutionary algorithms tried to address this problem utilizing the powerful search capability of Evolutionary Computation (EC). However, most algorithms focused on exploiting the Mean Square Residue (MSR) measure which was proposed by Cheng and Church. The Order Preserving Sub-Matrix (OPSM) constraint was rarely considered even though it promises more biologically relevant biclusters than the MSR measure. The goal of this paper is to design an EC algorithm which ensures biologically significant biclusters by using the OPSM constraint and better biclusters than the original OPSM algorithm. We designed a novel encoding method and evolutionary operators suitable for the OPSM constraint. To efficiently explore the search space, we modulized our evolutionary algorithm and applied the co-evolution concept. Through a set of experiments, it was confirmed that our algorithm outperformed a representative EC biclustering algorithm based on CC and the original OPSM algorithm. Hongchan Roh, Sanghyun Park 0003 |
BIBE | 2 |
| 2008 | Rule Discovery and Matching in Stock DatabasesabstractThis paper addresses an approach that recommends investment types to stock investors by discovering useful rules from past changing patterns of stock prices in databases. First, we define a new rule model for recommending stock investment types. For a frequent pattern of stock prices, if its subsequent stock prices are matched to a condition of an investor, the model recommends a corresponding investment type for this stock. The frequent pattern is regarded as a rule head, and the subsequent part a rule body. We observed that the conditions on rule bodies are quite different depending on dispositions of investors while rule heads are independent of characteristics of investors in most cases. With this observation, we propose a new method that discovers and stores only the rule heads rather than the whole rules in a rule discovery process. This allows investors to impose various conditions on rule bodies flexibly, and also improves the performance of a rule discovery process by reducing the number of rules to be discovered. For efficient discovery and matching of rules, we propose methods for discovering frequent patterns, constructing a frequent pattern base, and its indexing. We also suggest a method that finds the rules matched to a query from a frequent pattern base, and a method that recommends an investment type by using the rules. Finally, we verify the effectiveness and the efficiency of our approach through extensive experiments with real-life stock data. You-min Ha, Sanghyun Park 0003, Sang-Wook Kim, Jung-Im Won, Jeehee Yoon |
COMPSAC | 2 |
| 2008 | Extraction of Informative Genes from Integrated Microarray Data
Dongwan Hong, Jongkeun Lee, Sang-Kyoon Hong, Jeehee Yoon, Sanghyun Park 0003 |
ISMIS | 5 |
| 2008 | Privacy preserving data mining of sequential patterns for network traffic data
Seung-Woo Kim, Sanghyun Park 0003, Jung-Im Won, Sang-Wook Kim |
Inf. Sci. | 2 |
| 2008 | Image retrieval model based on weighted visual features determined by relevance feedback
Woo-Cheol Kim, Ji-Young Song, Seung-Woo Kim, Sanghyun Park 0003 |
Inf. Sci. | 4 |
| 2008 | Direct integration of microarrays for selecting informative genes and phenotype classification
Youngmi Yoon, Jongchan Lee, Sanghyun Park 0003, Sangjay Bien, Hyun Cheol Chung, Sun Young Rha |
Inf. Sci. | 3 |
| 2007 | Privacy Preserving Data Mining of Sequential Patterns for Network Traffic Data
Seung-Woo Kim, Sanghyun Park 0003, Jung-Im Won, Sang-Wook Kim |
DASFAA | 2 |
| 2007 | A Practical Method for Approximate Subsequence Search in DNA Databases
Jung-Im Won, Sang-Kyoon Hong, Jeehee Yoon, Sanghyun Park 0003, Sang-Wook Kim |
PAKDD | 4 |
| 2007 | Towards Efficient Searching on the Secondary Structure of Protein Sequences
Minkoo Seo, Sanghyun Park 0003, Jung-Im Won |
Fundam. Informaticae | 2 |
| 2007 | An efficient location encoding method for moving objects using hierarchical administrative district and road network
Sanghyun Park 0003, Woo-Cheol Kim, Dongwon Lee 0001 |
Inf. Sci. | 2 |
| 2007 | A multi-dimensional indexing approach for timestamped event sequence matching
Sanghyun Park 0003, Jung-Im Won, Jeehee Yoon, Sang-Wook Kim |
Inf. Sci. | 1 |
| 2006 | Building a Classifier for Integrated Microarray Datasets through Two-Stage ApproachabstractSince microarray data acquire tens of thousands of gene expression values simultaneously, they could be very useful in identifying the phenotypes of diseases. However, the results of analyzing several microarray datasets which were independently carried out with the same biological objectives, could turn out to be different. One of the main reasons is attributable to the limited number of samples involved in one microarray experiment. In order to increase the classification accuracy, it is desirable to augment the sample size by integrating and maximizing the use of independently-conducted microarray datasets. In this paper, we propose a two-stage approach which firstly integrates individual microarray datasets to overcome the problem caused by limited number of samples, and identifies informative genes, secondly builds a classifier using only the informative genes. The classifier from large samples by integrating independent microarray datasets achieves high accuracy, sensitivity, and specificity on independent test sample dataset Youngmi Yoon, Jongchan Lee, Sanghyun Park 0003 |
BIBE | 3 |
| 2006 | Shape-based retrieval in time-series databases
Sang-Wook Kim, Jeehee Yoon, Sanghyun Park 0003, Jung-Im Won |
J. Syst. Softw. | 3 |
| 2005 | An Efficient Location Encoding Method Based on Hierarchical Administrative District
Sanghyun Park 0003, Woo-Cheol Kim, Dongwon Lee 0001 |
DEXA | 2 |
| 2005 | An Index-Based Method for Timestamped Event Sequence Matching
Sanghyun Park 0003, Jung-Im Won, Jeehee Yoon, Sang-Wook Kim |
DEXA | 1 |
| 2005 | CSI: Clustered Segment Indexing for Efficient Approximate Searching on the Secondary Structure of Protein Sequences
Minkoo Seo, Sanghyun Park 0003, Jung-Im Won |
ISMIS | 2 |
| 2005 | A DNA Index Structure Using Frequency and Position Information of Genetic Alphabet
Woo-Cheol Kim, Sanghyun Park 0003, Jung-Im Won, Sang-Wook Kim, Jeehee Yoon |
PAKDD | 2 |
| 2005 | A Novel Indexing Method for Efficient Sequence Matching in Large DNA Database Environment
Jung-Im Won, Jeehee Yoon, Sanghyun Park 0003, Sang-Wook Kim |
PAKDD | 3 |
| 2004 | Efficient processing of similarity search under time warping in sequence databases: an index-based approach
Sang-Wook Kim, Sanghyun Park 0003, Wesley W. Chu |
Inf. Syst. | 2 |
| 2003 | Indexing Weighted-Sequences in Large DatabasesabstractWe present an index structure for managing weighted-sequences in large databases. A weighted-sequence is defined as a two-dimensional structure where each element in the sequence is associated with a weight. A series of network events, for instance, is a weighted-sequence in that each event has a timestamp. Querying a large sequence database by events' occurrence patterns is a first step towards understanding the temporal causal relationships among the events. The index structure proposed enables us to efficiently retrieve from the database all subsequences, possibly noncontiguous, that match a given query sequence both by events and by weights. The index method also takes into consideration the nonuniformfrequency distribution of events in the sequence data. In addition, our method finds a broad range of applications in indexing scientific data consisting of multiple numerical columns for discovery of correlations among these columns. For instance, indexing a DNA microarray that records expression levels of genes under different conditions enables us to search for genes whose responses to various experimental perturbations follow a given pattern. We demonstrate, using real-world data sets, that our method is effective and efficient. Haixun Wang, Chang-Shing Perng, Wei Fan 0001, Sanghyun Park 0003, Philip S. Yu |
ICDE | 4 |
| 2003 | ViST: A Dynamic Index Method for Querying XML Data by Tree StructuresabstractWith the growing importance of XML in data exchange, much research has been done in providing flexible query facilities to extract data from structured XML documents. In this paper, we propose ViST, a novel index structure for searching XML documents. By representing both XML documents and XML queries in structure-encoded sequences, we show that querying XML data is equivalent to finding subsequence matches. Unlike index methods that disassemble a query into multiple sub-queries, and then join the results of these sub-queries to provide the final answers, ViST uses tree structures as the basic unit of query to avoid expensive join operations. Furthermore, ViST provides a unified index on both content and structure of the XML documents, hence it has a performance advantage over methods indexing either just content or structure. ViST supports dynamic index update, and it relies solely on B+ Trees without using any specialized data structures that are not well supported by DBMSs. Our experiments show that ViST is effective, scalable, and efficient in supporting structural queries. Haixun Wang, Sanghyun Park 0003, Wei Fan 0001, Philip S. Yu |
SIGMOD Conference | 2 |
| 2003 | Similarity search of time-warped subsequences via a suffix tree
Sanghyun Park 0003, Wesley W. Chu, Jeehee Yoon, Jung-Im Won |
Inf. Syst. | 1 |
| 2001 | An Index-Based Approach for Similarity Search Supporting Time Warping in Large Sequence DatabasesabstractThis paper proposes a new novel method for similarity search that supports time warping in large sequence databases. Time warping enables finding sequences with similar patterns even when they are of different lengths. Previous methods for processing similarity search that supports time warping fail to employ multi-dimensional indexes without false dismissal since the time warping distance does not satisfy the triangular inequality. Our primary goal is to innovate on search performance without permitting any false dismissal. To attain this goal, we devise a new distance function D/sub tw-lb/ that consistently underestimates the time warping distance and also satisfies the triangular inequality D/sub tw-lb/ uses a 4-tuple feature vector that is extracted from each sequence and is invariant to time warping. For efficient processing of similarity search, we employ a multi-dimensional index that uses the 4-tuple feature vector as indexing attributes and D/sub tw-lb/ as a distance function. The extensive experimental results reveal that our method achieves significant speedup up to 43 times with real-world S&P 500 stock data and up to 720 times with very large synthetic data. Sang-Wook Kim, Sanghyun Park 0003, Wesley W. Chu |
ICDE | 2 |
| 2001 | Discovering and Matching Elastic Rules from Sequence Databases
Sanghyun Park 0003, Wesley W. Chu |
Fundam. Informaticae | 1 |
| 2000 | Efficient Searches for Similar Subsequences of Different Lengths in Sequence DatabasesabstractWe propose an indexing technique for fast retrieval of similar subsequences using time warping distances. A time warping distance is a more suitable similarity measure than the Euclidean distance in many applications, where sequences may be of different lengths or different sampling rates. Our indexing technique uses a disk-based suffix tree as an index structure and employs lower-bound distance functions to filter out dissimilar subsequences without false dismissals. To make the index structure compact and thus accelerate the query processing, we convert sequences of continuous values to sequences of discrete values via a categorization method and store only a subset of suffixes whose first values are different from their preceding values. The experimental results reveal that our proposed technique can be a few orders of magnitude faster than sequential scanning. Sanghyun Park 0003, Wesley W. Chu, Jeehee Yoon, Chih-Cheng Hsu |
ICDE | 1 |
| 2000 | Discovering and Matching Elastic Rules from Sequence Databases
Sanghyun Park 0003, Wesley W. Chu |
ISMIS | 1 |