VLDB 2026 Research / reviewers in the wild / expert
Seung-Hee Bae
dblp:79/379
· DBLP profile ↗
18ranked-venue papers
7as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Data mining · 95% Information retrieval · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 64% Distributed systems · 32% High-performance computing · 5% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › structured data mining › graph mining
community detection |
0.2 | 1 | 2015 | GossipMap: a distributed community detection algorithm for billion-edge directed graphs · SC 2015 |
Data mining › structured data mining
graph mining |
0.2 | 1 | 2015 | GossipMap: a distributed community detection algorithm for billion-edge directed graphs · SC 2015 |
Distributed systems
distributed graph processing |
0.2 | 1 | 2015 | GossipMap: a distributed community detection algorithm for billion-edge directed graphs · SC 2015 |
Visualization and visual analytics
high-dimensional data visualization |
0.1 | 2 | 2010 | Browsing large scale cheminformatics data with dimension reduction · HPDC 2010 Dimension reduction and visualization of large high-dimensional data via interpolation · HPDC 2010 |
Data mining
dimensionality reduction |
0.1 | 1 | 2010 | Dimension reduction and visualization of large high-dimensional data via interpolation · HPDC 2010 |
Data mining › dimensionality reduction
multidimensional scaling |
0.1 | 1 | 2010 | Dimension reduction and visualization of large high-dimensional data via interpolation · HPDC 2010 |
Visualization and visual analytics
dimensionality reduction |
0.1 | 1 | 2010 | Browsing large scale cheminformatics data with dimension reduction · HPDC 2010 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Parallel and multicore computing › parallel programming runtimes
mapreduce runtime |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Bioinformatics and computational biology
gene regulation |
0.1 | 1 | 2007 | dPattern: transcription factor binding site (TFBS) discovery in human genome using a discriminative pattern analysis · Bioinform. 2007 |
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction |
0.1 | 1 | 2007 | dPattern: transcription factor binding site (TFBS) discovery in human genome using a discriminative pattern analysis · Bioinform. 2007 |
Bioinformatics and computational biology › molecular informatics
cheminformatics |
0.0 | 1 | 2010 | Browsing large scale cheminformatics data with dimension reduction · HPDC 2010 |
Information retrieval › interactive information retrieval
browsing |
0.0 | 1 | 2010 | Browsing large scale cheminformatics data with dimension reduction · HPDC 2010 |
High-performance computing › data-intensive computing
parallel data analysis |
0.0 | 1 | 2010 | Dimension reduction and visualization of large high-dimensional data via interpolation · HPDC 2010 |
Bioinformatics and computational biology › gene expression analysis
differential expression analysis |
0.0 | 1 | 2007 | dPattern: transcription factor binding site (TFBS) discovery in human genome using a discriminative pattern analysis · Bioinform. 2007 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2007 | dPattern: transcription factor binding site (TFBS) discovery in human genome using a discriminative pattern analysis · Bioinform. 2007 |
Methods — techniques the papers use, named apart from their topics
generative topographic mapping · 0.7gossip-based distributed computation · 0.4approximation algorithm · 0.4parallelization · 0.3multidimensional scaling · 0.3interpolation · 0.3iterative mapreduce · 0.1discriminative pattern analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Scalable and Efficient Flow-Based Community Detection for Large-Scale Graph AnalysisabstractCommunity detection is an increasingly popular approach to uncover important structures in large networks. Flow-based community detection methods rely on communication patterns of the network rather than structural properties to determine communities. The Infomap algorithm in particular optimizes a novel objective function called the map equation and has been shown to outperform other approaches in third-party benchmarks. However, Infomap and its variants are inherently sequential, limiting their use for large-scale graphs. In this article, we propose a novel algorithm to optimize the map equation called RelaxMap. RelaxMap provides two important improvements over Infomap: parallelization, so that the map equation can be optimized over much larger graphs, and prioritization, so that the most important work occurs first, iterations take less time, and the algorithm converges faster. We implement these techniques using OpenMP on shared-memory multicore systems, and evaluate our approach on a variety of graphs from standard graph clustering benchmarks as well as real graph datasets. Our evaluation shows that both techniques are effective: RelaxMap achieves 70% parallel efficiency on eight cores, and prioritization improves algorithm performance by an additional 20--50% on average, depending on the graph properties. Additionally, RelaxMap converges in the similar number of iterations and provides solutions of equivalent quality as the serial Infomap implementation. Seung-Hee Bae, Daniel Halperin, Jevin D. West, Martin Rosvall, Bill Howe |
ACM Trans. Knowl. Discov. Data | 1 |
| 2015 | GossipMap: a distributed community detection algorithm for billion-edge directed graphsabstractIn this paper, we describe a new distributed community detection algorithm for billion-edge directed graphs that, unlike modularity-based methods, achieves cluster quality on par with the best-known algorithms in the literature. We show that a simple approximation to the best-known serial algorithm dramatically reduces computation and enables distributed evaluation yet incurs only a very small impact on cluster quality. Seung-Hee Bae, Bill Howe |
SC | 1 |
| 2012 | Interpolative multidimensional scaling techniques for the identification of clusters in very large sequence setsabstractBACKGROUND: Modern pyrosequencing techniques make it possible to study complex bacterial populations, such as 16S rRNA, directly from environmental or clinical samples without the need for laboratory purification. Alignment of sequences across the resultant large data sets (100,000+ sequences) is of particular interest for the purpose of identifying potential gene clusters and families, but such analysis represents a daunting computational task. The aim of this work is the development of an efficient pipeline for the clustering of large sequence read sets. METHODS: Pairwise alignment techniques are used here to calculate genetic distances between sequence pairs. These methods are pleasingly parallel and have been shown to more accurately reflect accurate genetic distances in highly variable regions of rRNA genes than do traditional multiple sequence alignment (MSA) approaches. By utilizing Needleman-Wunsch (NW) pairwise alignment in conjunction with novel implementations of interpolative multidimensional scaling (MDS), we have developed an effective method for visualizing massive biosequence data sets and quickly identifying potential gene clusters. RESULTS: This study demonstrates the use of interpolative MDS to obtain clustering results that are qualitatively similar to those obtained through full MDS, but with substantial cost savings. In particular, the wall clock time required to cluster a set of 100,000 sequences has been reduced from seven hours to less than one hour through the use of interpolative MDS. CONCLUSIONS: Although work remains to be done in selecting the optimal training set size for interpolative MDS, substantial computational cost savings will allow us to cluster much larger sequence sets in the future. Adam Hughes, Yang Ruan 0001, Saliya Ekanayake, Seung-Hee Bae, Qunfeng Dong, Mina Rho, Judy Qiu, Geoffrey C. Fox |
BMC Bioinform. | 4 |
| 2012 | Performance of windows multicore systems on threading and MPIabstractSUMMARY We present performance results on a Windows cluster with up to 768 cores using Message Passing Interface (MPI) and two variants of threading—Concurrency and Coordination Runtime (CCR) and Task Parallel Library (TPL). CCR presents a message‐based interface, while TPL allows for loops to be automatically parallelized. MPI is used between the cluster nodes (up to 32) and either threading or MPI for parallelism on the 24 cores of each node. We look at the performance of two significant bioinformatics applications; gene clustering and dimension reduction. We find that the two threading runtimes offer similar performance with MPI outperforming both at low levels of parallelism but threading much better when the grain size (problem size per process/thread) is small. We develop simple models for the performance of the clustering code. Copyright © 2011 John Wiley & Sons, Ltd. Judy Qiu, Seung-Hee Bae |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Browsing large-scale cheminformatics data with dimension reductionabstractSUMMARY Visualization of large‐scale high dimensional data is highly valuable for data analysis facilitating scientific discovery in many fields. We present PubChemBrowse, a customized visualization tool for cheminformatics research. It provides a novel 3D data point browser that displays complex properties of massive data on commodity clients. As in Geographic Information System browsers for Earth and Environment data, chemical compounds with similar properties are nearby in the browser. PubChemBrowse is built around in‐house high performance parallel Multi‐dimensional scaling and Generative topographic mapping services and supports fast interaction with an external property database. These properties can be overlaid on 3D mapped compound space or queried for individual points. We prototype the integration with Chem2Bio2RDF system using SPARQL endpoint to access over 20 publicly accessible bioinformatics databases. We describe our design and implementation of the integrated PubChemBrowse application and outline its use in drug discovery. The same core technologies are generally applicable to develop high performance scientific data browsing systems for other applications. Copyright © 2011 John Wiley & Sons, Ltd. Jong Choi 0001, Seung-Hee Bae, Judy Qiu, Bin Chen 0002, David J. Wild 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Cloud computing paradigms for pleasingly parallel biomedical applicationsabstractSUMMARY Cloud computing offers exciting new approaches for scientific computing that leverage major commercial players’ hardware and software investments in large‐scale data centers. Loosely coupled problems are very important in many scientific fields, and with the ongoing move towards data‐intensive computing, they are on the rise. There exist several different approaches to leveraging clouds and cloud‐oriented data processing frameworks to perform pleasingly parallel (also called embarrassingly parallel) computations. In this paper, we present three pleasingly parallel biomedical applications: (i) assembly of genome fragments; (ii) sequence alignment and similarity search; and (iii) dimension reduction in the analysis of chemical structures, which are implemented utilizing a cloud infrastructure service‐based utility computing models of Amazon Web Services ( http://Amazon.com Inc., Seattle, WA, USA) and Microsoft Windows Azure (Microsoft Corp., Redmond, WA, USA) as well as utilizing MapReduce‐based data processing frameworks Apache Hadoop (Apache Software Foundation, Los Angeles, CA, USA) and Microsoft DryadLINQ. We review and compare each of these frameworks, performing a comparative study among them based on performance, cost, and usability. High latency, eventually consistent cloud infrastructure service‐based frameworks that rely on off‐the‐node cloud storage were able to exhibit performance efficiencies and scalability comparable to the MapReduce‐based frameworks with local disk‐based storage for the applications considered. In this paper, we also analyze variations in cost among the different platform choices (e.g., Elastic Compute Cloud instance types), highlighting the importance of selecting an appropriate platform based on the nature of the computation. Copyright © 2011 John Wiley & Sons, Ltd. Thilina Gunarathne, Tak-Lon Wu, Jong Choi 0001, Seung-Hee Bae, Judy Qiu |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | High Performance Dimension Reduction and Visualization for Large High-Dimensional Data AnalysisabstractLarge high dimension datasets are of growing importance in many fields and it is important to be able to visualize them for understanding the results of data mining approaches or just for browsing them in a way that distance between points in visualization (2D or 3D) space tracks that in original high dimensional space. Dimension reduction is a well understood approach but can be very time and memory intensive for large problems. Here we report on parallel algorithms for Scaling by MAjorizing a Complicated Function (SMACOF) to solve Multidimensional Scaling problem and Generative Topographic Mapping (GTM). The former is particularly time consuming with complexity that grows as square of data set size but has advantage that it does not require explicit vectors for dataset points but just measurement of inter-point dissimilarities. We compare SMACOF and GTM on a subset of the NIH PubChem database which has binary vectors of length 166 bits. We find good parallel performance for both GTM and SMACOF and strong correlation between the dimension-reduced PubChem data from these two methods. Jong Choi 0001, Seung-Hee Bae, Xiaohong Qiu, Geoffrey C. Fox |
CCGRID | 2 |
| 2010 | Performance of Windows Multicore Systems on Threading and MPIabstractWe present performance results on a Windows cluster with up to 768 cores using MPI and two variants of threading - CCR and TPL. CCR (Concurrency and Coordination Runtime) presents a message based interface while TPL (Task Parallel Library) allows for loops to be automatically parallelized. MPI is used between the cluster nodes (up to 32) and either threading or MPI for parallelism on the 24 cores of each node. We use a simple matrix multiplication kernel as well as a significant bioinformatics gene clustering application. We find that the two threading models offer similar performance with MPI outperforming both at low levels of parallelism but threading much better when the grain size (problem size per process) is small. We find better performance on Intel compared to AMD on comparable 24 core systems. We develop simple models for the performance of the clustering code. Judy Qiu, Scott Beason, Seung-Hee Bae, Saliya Ekanayake, Geoffrey C. Fox |
CCGRID | 3 |
| 2010 | Multidimensional Scaling by Deterministic Annealing with Iterative Majorization AlgorithmabstractMultidimensional Scaling (MDS) is a dimension reduction method for information visualization, which is set up as a non-linear optimization problem. It is applicable to many data intensive scientific problems including studies of DNA sequences but tends to get trapped in local minima. Deterministic Annealing (DA) has been applied to many optimization problems to avoid local minima. We apply DA approach to MDS problem in this paper and show that our proposed DA approach improves the mapping quality and shows high reliability in a variety of experimental results. Further its execution time is similar to that of the un-annealed approach. We use different data sets for comparing the proposed DA approach with both a well known algorithm called SMACOF and a MDS with distance smoothing method which aims to avoid local optima. Our proposed DA method outperforms SMACOF algorithm and the distance smoothing MDS algorithm in terms of the mapping quality and shows much less sensitivity with respect to initial configurations and stopping condition. We also investigate various temperature cooling parameters for our deterministic annealing method within an exponential cooling scheme. Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox |
eScience | 1 |
| 2010 | Dimension reduction and visualization of large high-dimensional data via interpolationabstractThe recent explosion of publicly available biology gene sequences and chemical compounds offers an unprecedented opportunity for data mining. To make data analysis feasible for such vast volume and high-dimensional scientific data, we apply high performance dimension reduction algorithms. It facilitates the investigation of unknown structures in a three dimensional visualization. Among the known dimension reduction algorithms, we utilize the multidimensional scaling and generative topographic mapping algorithms to configure the given high-dimensional data into the target dimension. However, both algorithms require large physical memory as well as computational resources. Thus, the authors propose an interpolated approach to utilizing the mapping of only a subset of the given data. This approach effectively reduces computational complexity. With minor trade-off of approximation, interpolation method makes it possible to process millions of data points with modest amounts of computation and memory requirement. Since huge amount of data are dealt, we represent how to parallelize proposed interpolation algorithms, as well. For the evaluation of the interpolated MDS by STRESS criteria, it is necessary to compute symmetric all pairwise computation with only subset of required data per process, so we also propose a simple but efficient parallel mechanism for the symmetric all pairwise computation when only a subset of data is available to each process. Our experimental results illustrate that the quality of interpolated mapping results are comparable to the mapping results of original algorithm only. In parallel performance aspect, those interpolation methods are well parallelized with high efficiency. With the proposed interpolation method, we construct a configuration of two-million out-of-sample data into the target dimension, and the number of out-of-sample data can be increased further. Seung-Hee Bae, Jong Choi 0001, Judy Qiu, Geoffrey C. Fox |
HPDC | 1 |
| 2010 | Browsing large scale cheminformatics data with dimension reductionabstractVisualization of large-scale high dimensional data tool is highly valuable for scientific discovery in many fields. We presentPubChemBrowse, acustomizedvisualizationtoolfor cheminformatics research. It provides a novel 3D data point browser that displays complex properties of massive data on commodity clients. As in GIS browsers for Earth and Environment data, chemical compounds with similar properties are nearby in the browser. PubChemBrowse is built around in-househighperformanceparallel MDS(Multi-Dimensional Scaling) and GTM (Generative Topographic Mapping) services andsupports fast interaction with anexternalproperty database. These properties can be overlaid on 3D mapped compound space or queried for individual points. We prototype use with Chem2Bio2RDF system using SPARQLquery language to access over 20 publicly accessible bioinformatics databases. We describe our design and implementation of the integrated PubChemBrowse application and outline its use in drug discovery. The same core technologies can be used to develop similar high dimensional browsers in other scientific areas. Jong Choi 0001, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox, Bin Chen 0002, David J. Wild 0001 |
HPDC | 2 |
| 2010 | Twister: a runtime for iterative MapReduceabstractMapReduce programming model has simplified the implementation of many data parallel applications. The simplicity of the programming model and the quality of services provided by many implementations of MapReduce attract a lot of enthusiasm among distributed computing communities. From the years of experience in applying MapReduce to various scientific applications we identified a set of extensions to the programming model and improvements to its architecture that will expand the applicability of MapReduce to more classes of applications. In this paper, we present the programming model and the architecture of Twister an enhanced MapReduce runtime that supports iterative MapReduce computations efficiently. We also show performance comparisons of Twister with other similar runtimes such as Hadoop and DryadLINQ for large scale data parallel applications. Jaliya Ekanayake, Bingjing Zhang, Thilina Gunarathne, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox |
HPDC | 5 |
| 2010 | Hybrid cloud and cluster computing paradigms for life science applicationsabstractBACKGROUND: Clouds and MapReduce have shown themselves to be a broadly useful approach to scientific computing especially for parallel data intensive applications. However they have limited applicability to some areas such as data mining because MapReduce has poor performance on problems with an iterative structure present in the linear algebra that underlies much data analysis. Such problems can be run efficiently on clusters using MPI leading to a hybrid cloud and cluster environment. This motivates the design and implementation of an open source Iterative MapReduce system Twister. RESULTS: Comparisons of Amazon, Azure, and traditional Linux and Windows environments on common applications have shown encouraging performance and usability comparisons in several important non iterative cases. These are linked to MPI applications for final stages of the data analysis. Further we have released the open source Twister Iterative MapReduce and benchmarked it against basic MapReduce (Hadoop) and MPI in information retrieval and life sciences applications. CONCLUSIONS: The hybrid cloud (MapReduce) and cluster (MPI) approach offers an attractive production environment while Twister promises a uniform programming environment for many Life Sciences applications. METHODS: We used commercial clouds Amazon and Azure and the NSF resource FutureGrid to perform detailed comparisons and evaluations of different approaches to data intensive computing. Several applications were developed in MPI, MapReduce and Twister in these different environments. Judy Qiu, Jaliya Ekanayake, Thilina Gunarathne, Jong Choi 0001, Seung-Hee Bae, Bingjing Zhang, Tak-Lon Wu, Yang Ruan 0001, Saliya Ekanayake, Adam Hughes, Geoffrey C. Fox |
BMC Bioinform. | 5 |
| 2008 | Parallel Multidimensional Scaling Performance on Multicore SystemsabstractMultidimensional scaling constructs a configuration points into the target low-dimensional space, while the interpoint distances are approximated to the corresponding known dissimilarity values as much as possible. SMACOF algorithm is an elegant gradient descent approach to solve Multidimensional scaling problem. We design parallel SMACOF program using parallel matrix multiplication to run on a multicore machine. Also, we propose a block decomposition algorithm based on the number of threads for the purpose of keeping good load balance. The proposed block decomposition algorithm works very well if the number of block columns is at least a half of the number of threads. In this paper, we investigate performance results of the implemented parallel SMACOF in terms of the block size, data size, and the number of threads. The speedup factor is almost 7.7 with 2048 points data over 8 running threads. In addition, performance comparison between jagged array and two-dimensional array in C# language is carried out. The jagged array data structure performs at least 40% better than the two-dimensional array structure. Seung-Hee Bae |
eScience | 1 |
| 2008 | SALSA Project: Parallel Data Mining of GIS, Web, Medical, Physics, Chemical, and Biology DataabstractThe multicore revolution promises potentially hundreds of cores in desktop computers. The ever increasing number of cores per chip will be accompanied by a pervasive data deluge whose size will probably increase even faster than CPU core count over the next few years. This suggests the importance of parallel data analysis and data mining applications with good multicore, cluster and grid performance. The SALSA project at Community Grid Lab of Indiana University is looking to revolutionize the way software is written in parallel for real applications that advance scientific discovery and improve the quality of people's life. Xiaohong Qiu, Geoffrey C. Fox, Seung-Hee Bae, Jong Choi 0001, Jaliya Ekanayake, Yang Ruan 0001 |
eScience | 3 |
| 2007 | High Performance Multi-paradigm Messaging Runtime Integrating Grids and Multicore SystemsabstracteScience applications need to use distributed Grid environments where each component is an individual or cluster of multicore machines. These are expected to have 64-128 cores 5 years from now and need to support scalable parallelism. Users will want to compose heterogeneous components into single jobs and run seamlessly in both distributed fashion and on a future "Grid on a chip" with different subsets of cores supporting individual components. We support this with a simple programming model made up of two layers supporting traditional parallel and Grid programming paradigms (workflow) respectively. We examine for a parallel clustering application, the Concurrency and Coordination Runtime CCR from Microsoft as a multi-paradigm runtime that integrates the two layers. Our work uses managed code (C#) and for AMD and Intel processors shows around a factor of 5 better performance than Java. CCR has MPI pattern and dynamic threading latencies of a few microseconds that are competitive with the performance of standard MPI for C. Xiaohong Qiu, Geoffrey C. Fox, Huapeng Yuan, Seung-Hee Bae, George Chrysanthakopoulos, Henrik Frystyk Nielsen |
eScience | 4 |
| 2007 | dPattern: transcription factor binding site (TFBS) discovery in human genome using a discriminative pattern analysisabstractAbstract Motivation: Transcription factor binding sites (TFBSs) are typically short in length, thus search with a profile model from known TFBSs produces many false positives. When combined with additional information, gene expression data in this article, sensitivity and specificity of TFBS search can be improved significantly. Results: By modifying our previous REFINEMENT approach, we developed dPattern that searches for occurrences of TFBSs in the promotor regions of up/down regulated or random genes. Availability: http://platcom.org/projects/dpattern Contact: [email protected] or [email protected] Seung-Hee Bae, Haixu Tang, Sun Kim |
Bioinform. | 1 |
| 2004 | Mutation Rates in the Context of Hybrid Genetic Algorithms
Seung-Hee Bae, Byung Ro Moon |
GECCO (2) | 1 |