Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

You Jung Kim

dblp:28/4685 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 46% Query processing and optimization · 30% Indexing and storage engines · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Data mining › data reduction
data pruning
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Query processing and optimization › runtime optimization › data skipping
partition pruning
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Indexing and storage engines › synopsis structure
zone maps
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Bioinformatics and computational biology › sequence analysis
read mapping
0.112009
ProbeMatch: rapid alignment of oligonucleotides to genome allowing both gaps and mismatches · Bioinform. 2009
Bioinformatics and computational biology
sequence alignment
0.112009
ProbeMatch: rapid alignment of oligonucleotides to genome allowing both gaps and mismatches · Bioinform. 2009
Bioinformatics and computational biology
sequence analysis
0.112009
ProbeMatch: rapid alignment of oligonucleotides to genome allowing both gaps and mismatches · Bioinform. 2009
Bioinformatics and computational biology › sequence analysis
sequencing data analysis
0.112009
ProbeMatch: rapid alignment of oligonucleotides to genome allowing both gaps and mismatches · Bioinform. 2009

Methods — techniques the papers use, named apart from their topics

gapped q-grams · 0.1gapped alignment · 0.1
YearPublicationVenuePosition
2017 Dimensions Based Data Clustering and Zone Maps
abstract
In recent years, the data warehouse industry has witnessed decreased use of indexing but increased use of compression and clustering of data facilitating efficient data access and data pruning in the query processing area. A classic example of data pruning is the partition pruning, which is used when table data is range or list partitioned. But lately, techniques have been developed to prune data at a lower granularity than a table partition or sub-partition. A good example is the use of data pruning structure called zone map. A zone map prunes zones of data from a table on which it is defined. Data pruning via zone map is very effective when the table data is clustered by the filtering columns. The database industry has offered support to cluster data in tables by its local columns, and to define zone maps on clustering columns of such tables. This has helped improve the performance of queries that contain filter predicates on local columns. However, queries in data warehouses are typically based on star/snowflake schema with filter predicates usually on columns of the dimension tables joined to a fact table. Given this, the performance of data warehouse queries can be significantly improved if the fact table data is clustered by columns of dimension tables together with zone maps that maintain min/max value ranges of these clustering columns over zones of fact table data. In recognition of this opportunity of significantly improving the performance of data warehouse queries, Oracle 12c release 1 has introduced the support for dimension based clustering of fact tables together with data pruning of the fact tables via dimension based zone maps.
Mohamed Ziauddin, Andrew Witkowski, You Jung Kim, Janaki Lahorani, Dmitry Potapov, Murali Krishna
Proc. VLDB Endow.3
2010 Performance Comparison of the R*-Tree and the Quadtree for kNN and Distance Join Queries
abstract
Multidimensional point indexing plays a critical role in a variety of data-centric applications, including image retrieval, sequence matching, and moving object database search. A common choice of indexing method for these applications is often the "ubiquitous” R*-tree. Choosing the right indexing method requires careful consideration of various factors such as query operations and index construction methods. In this work, we present an experimental study comparing the R*-tree and Quadtree using various criteria including the query operations and index construction methods. Although a variety of query operations can be performed using these index structures, previous work has largely focused only on the range search operation. We go beyond this previous work and compare the performance of these index structures using k-nearest neighbor (kNN) and distance join queries. In addition, we also consider the impact of index construction methods in evaluating these index structures. Our study sheds light on how the choice of the underlying index structure affects the performance of different query operations, and shows that the method used for constructing the index and the dynamic nature of the data set has a dramatic impact on the performance of these index structures.
You Jung Kim, Jignesh M. Patel
IEEE Trans. Knowl. Data Eng.1
2009 ProbeMatch: rapid alignment of oligonucleotides to genome allowing both gaps and mismatches
abstract
SUMMARY: We have developed a tool, called ProbeMatch, for matching a large set of oligonucleotide sequences against a genome database using gapped alignments. Unlike most of the existing tools such as ELAND which only perform ungapped alignments allowing at most two mismatches, ProbeMatch generates both ungapped and gapped alignments allowing up to three errors including insertion, deletion and mismatch. To speedup sequence alignment, ProbeMatch uses gapped q-grams and q-grams of various patterns to identify target hits to a query sequence. This approach results in fewer initial sequences to examine with no loss in sensitivity. ProbeMatch has been used to align 169,095 Illumina GAII reads against the human genome, which could not be mapped by ELAND, and found alignments for 28,625 reads of the 169,095 reads in less than 3 h. AVAILABILITY: Source code is freely available at (http://www.cs.wisc.edu/~jignesh/probematch/).
You Jung Kim, Nikhil Teletia, Victor Ruotti, Christopher A. Maher, Arul M. Chinnaiyan, Ron M. Stewart, James A. Thomson, Jignesh M. Patel
Bioinform.1
2007 Rethinking Choices for Multi-dimensional Point Indexing: Making the Case for the Often Ignored Quadtree
You Jung Kim, Jignesh M. Patel
CIDR1
2006 A framework for protein structure classification and identification of novel protein structures
abstract
BACKGROUND: Protein structure classification plays a central role in understanding the function of a protein molecule with respect to all known proteins in a structure database. With the rapid increase in the number of new protein structures, the need for automated and accurate methods for protein classification is increasingly important. RESULTS: In this paper we present a unified framework for protein structure classification and identification of novel protein structures. The framework consists of a set of components for comparing, classifying, and clustering protein structures. These components allow us to accurately classify proteins into known folds, to detect new protein folds, and to provide a way of clustering the new folds. In our evaluation with SCOP 1.69, our method correctly classifies 86.0%, 87.7%, and 90.5% of new domains at family, superfamily, and fold levels. Furthermore, for protein domains that belong to new domain families, our method is able to produce clusters that closely correspond to the new families in SCOP 1.69. As a result, our method can also be used to suggest new classification groups that contain novel folds. CONCLUSION: We have developed a method called proCC for automatically classifying and clustering domains. The method is effective in classifying new domains and suggesting new domain families, and it is also very efficient. A web site offering access to proCC is freely available at http://www.eecs.umich.edu/periscope/procc.
You Jung Kim, Jignesh M. Patel
BMC Bioinform.1