Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Byoung-Kee Yi

dblp:12/4952 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
0since 2021 · last 2012
0000-0002-7699-9629ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 4 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Spatial and temporal data management · 40% Data mining · 21% Query processing and optimization · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
pattern matching query
0.112011
Supporting Pattern-Matching Queries over Trajectories on Road Networks · IEEE Trans. Knowl. Data Eng. 2011
Spatial and temporal data management › trajectory data management
road network trajectory
0.112011
Supporting Pattern-Matching Queries over Trajectories on Road Networks · IEEE Trans. Knowl. Data Eng. 2011
Spatial and temporal data management
trajectory data
0.112011
Supporting Pattern-Matching Queries over Trajectories on Road Networks · IEEE Trans. Knowl. Data Eng. 2011
Bioinformatics and computational biology › biomedical text mining
biomedical named entity recognition
0.112005
POSBIOTM-NER: a trainable biomedical named-entity recognition system · Bioinform. 2005
Bioinformatics and computational biology
biomedical text mining
0.112005
POSBIOTM-NER: a trainable biomedical named-entity recognition system · Bioinform. 2005
Data mining › predictive modeling
forecasting
0.012000
Online Data Mining for Co-Evolving Time Sequences · ICDE 2000
Data integration and cleaning › missing data
missing value imputation
0.012000
Online Data Mining for Co-Evolving Time Sequences · ICDE 2000
Indexing and storage engines
multidimensional indexing
0.012000
Fast Time Sequence Indexing for Arbitrary Lp Norms · VLDB 2000
Data mining › anomaly detection
outlier detection
0.012000
Online Data Mining for Co-Evolving Time Sequences · ICDE 2000
Data mining › anomaly detection
time series anomaly detection
0.012000
Online Data Mining for Co-Evolving Time Sequences · ICDE 2000
Indexing and storage engines › temporal indexing
time series indexing
0.012000
Fast Time Sequence Indexing for Arbitrary Lp Norms · VLDB 2000
Data mining › temporal data mining
time series mining
0.012000
Online Data Mining for Co-Evolving Time Sequences · ICDE 2000
Information retrieval › similarity search › sequence similarity search
time series similarity search
0.011998
Efficient Retrieval of Similar Time Sequences Under Time Warping · ICDE 1998
Data mining › time series analysis
time warping
0.011998
Efficient Retrieval of Similar Time Sequences Under Time Warping · ICDE 1998
Indexing and storage engines
vector index
0.011998
Efficient Retrieval of Similar Time Sequences Under Time Warping · ICDE 1998

Methods — techniques the papers use, named apart from their topics

linguistic feature analysis · 0.1conditional random field · 0.1incremental algorithm · 0.0Selective MUSCLES · 0.0MUSCLES · 0.0time warping · 0.0linear filter · 0.0fastmap · 0.0
YearPublicationVenuePosition
2012 Efficient bitmap-based indexing of time-based interval sequences
Jong-Won Roh, Seung-won Hwang, Byoung-Kee Yi
Inf. Sci.3
2011 Supporting Pattern-Matching Queries over Trajectories on Road Networks
abstract
With the advent of ubiquitous computing, we can easily collect large-scale trajectory data, say, from moving vehicles. This paper studies pattern-matching problems for trajectory data over road networks, which complements existing efforts focusing on (1) a spatiotemporal window query for location-based service or (2) euclidean space with no restriction. In contrast, we first identify some desirable properties for pattern-matching queries to the road network trajectories. As the existing work does not fully satisfy these properties, we develop (1) trajectory representation and (2) distance metric that satisfy all the desirable properties we identified. Based on this representation and metric, we develop efficient algorithms for three types of pattern-matching queries-whole, subpattern, and reverse subpattern matching. We analytically validate the correctness of our algorithms and also empirically validate their scalability over large-scale, real-life, and synthetic trajectory data sets.
Gook-Pil Roh, Jong-Won Roh, Seung-won Hwang, Byoung-Kee Yi
IEEE Trans. Knowl. Data Eng.4
2008 Efficient indexing of interval time sequences
Jong-Won Roh, Byoung-Kee Yi
Inf. Process. Lett.2
2006 Two-phase learning for biological event extraction and verification
abstract
Many previous biological event-extraction systems were based on hand-crafted rules which were specifically tuned to a specific biological application domain. But manually constructing and tuning the rules are time-consuming processes and make the systems less portable. So supervised machine-learning methods were developed to generate the extraction rules automatically, but accepting the trade-off between precision and recall (high recall with low precision, and vice versa) is a barrier to improving performance. To make matters worse, a text in the biological domain is more complex because it often contains more than two biological events in a sentence, and one event in a noun chunk can be an entity for the other event. As a result, there are as yet no systems that give a good performance in extracting events in biological domains by using supervised machine learning.To overcome the limitations of previous systems and the complexity of biological texts, we present the following new ideas. First, we adopted a supervised machine-learning method to reduce the human effort in making extraction rules in order to obtain a highly domain-portable system. Second, we overcame the classical trade-off between precision and recall by using an event component verification method. Thus, machine learning occurs in two phases in our architecture. In the first phase, the system focuses on improving recall in extracting events between biological entities during a supervised machine-learning period. After extracting the biological events with automatically learned rules, in the second phase the system removes incorrect biological events by verifying the extracted event components with a maximum entropy (ME) classification method. In other words, the system targets for high recall in the first phase and tries to achieve high precision with a classifier in the second phase. Finally, we improved a supervised machine-learning algorithm so that it could learn a rule in a noun chunk and a rule extending throughout a sentence at two different levels, separately, for nested biological events.
Eunju Kim, Cheongjae Lee, Kyungduk Kim, Gary Geunbae Lee, Byoung-Kee Yi, Jeongwon Cha
ACM Trans. Asian Lang. Inf. Process.6
2005 POSBIOTM-NER: a trainable biomedical named-entity recognition system
abstract
SUMMARY: POSBIOTM-NER is a trainable biomedical named-entity recognition system. POSBIOTM-NER can be automatically trained and adapted to new datasets without performance degradation, using CRF (conditional random field) machine learning techniques and automatic linguistic feature analysis. Currently, we have trained our system on three different datasets. GENIA-NER was trained based on GENIA Corpus, GENE-NER based on BioCreative data and GPCR-NER based on our own POSBIOTM/NE corpus, respectively, which would be used in GPCR-related pathway extraction.
Eunju Kim, Gary Geunbae Lee, Byoung-Kee Yi
Bioinform.4
2004 Similarity Search for Interval Time Sequences
Byoung-Kee Yi, Jong-Won Roh
DASFAA1
2000 Online Data Mining for Co-Evolving Time Sequences
abstract
In many applications, the data of interest comprises multiple sequences that evolve over time. Examples include currency exchange rates and network traffic data. We develop a fast method to analyze such co-evolving time sequences jointly to allow (a) estimation/forecasting of missing/delayed/future values, (b) quantitative data mining, and (c) outlier detection. Our method, MUSCLES, adapts to changing correlations among time sequences. It can handle indefinitely long sequences efficiently using an incremental algorithm and requires only a small amount of storage and less I/O operations. To make it scale for a large number of sequences, we present a variation, the Selective MUSCLES method and propose an efficient algorithm to reduce the problem size. Experiments on real datasets show that MUSCLES outperforms popular competitors in prediction accuracy up to 10 times, and discovers interesting correlations. Moreover, Selective MUSCLES scales up very well for large numbers of sequences, reducing response time up to 110 times over MUSCLES, and sometimes even improves the prediction quality.
Byoung-Kee Yi, Nicholas D. Sidiropoulos, Theodore Johnson, H. V. Jagadish, Christos Faloutsos, Alexandros Biliris
ICDE1
2000 Fast Time Sequence Indexing for Arbitrary Lp Norms
Byoung-Kee Yi, Christos Faloutsos
VLDB1
1999 Interactive Authoring of Multimedia Documents in a Constraint-Based Authoring System
Junehwa Song, G. Ramalingam, Raymond E. Miller, Byoung-Kee Yi
Multim. Syst.4
1998 Efficient Retrieval of Similar Time Sequences Under Time Warping
abstract
Fast similarity searching in large time sequence databases has typically used Euclidean distance as a dissimilarity metric. However, for several applications, including matching of voice, audio and medical signals (e.g., electrocardiograms), one is required to permit local accelerations and decelerations in the rate of sequences, leading to a popular, field tested dissimilarity metric called the "time warping" distance. From the indexing viewpoint, this metric presents two major challenges: (a) it does not lead to any natural indexable "features", and (b) comparing two sequences requires time quadratic in the sequence length. To address each problem, we propose to use: (a) a modification of the so called "FastMap", to map sequences into points, with little compromise of "recall" (typically zero); and (b) a fast linear test, to help us discard quickly many of the false alarms that FastMap will typically introduce. Using both ideas in cascade, our proposed method achieved up to an order of magnitude speed-up over sequential scanning on both real and synthetic datasets.
Byoung-Kee Yi, H. V. Jagadish, Christos Faloutsos
ICDE1