EDBT 2026 Demo / reviewers in the wild / expert
Jong Wook Kim
dblp:49/3726
· DBLP profile ↗
36ranked-venue papers
15as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 7 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Computer networks · 4 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 42% Transfer learning and domain adaptation · 19% Trustworthy machine learning · 13% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 79% Query processing and optimization · 14% Web and social media mining · 7% | |
| Human-computer interaction and pervasive computing
3 papers |
Accessibility and assistive technology · 46% Usability and user experience research · 41% Learning and educational technologies · 12% |
Topics — the 28 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
1.2 | 2 | 2023 | Robust Speech Recognition via Large-Scale Weak Supervision · ICML 2023 Learning Transferable Visual Models From Natural Language Supervision · ICML 2021 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.9 | 1 | 2025 | CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image · ICCV 2025 |
Computer vision › 3D vision
3d scene reconstruction |
0.9 | 1 | 2025 | CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image · ICCV 2025 |
Computer vision › 3D vision
novel view synthesis |
0.9 | 1 | 2025 | CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image · ICCV 2025 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.9 | 1 | 2025 | CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image · ICCV 2025 |
Digital forensics and information hiding › watermarking
neural radiance field watermarking |
0.8 | 1 | 2024 | WateRF: Robust Watermarks in Radiance Fields for Protection of Copyrights · CVPR 2024 |
Digital forensics and information hiding
watermarking |
0.8 | 1 | 2024 | WateRF: Robust Watermarks in Radiance Fields for Protection of Copyrights · CVPR 2024 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition |
0.7 | 1 | 2023 | Robust Speech Recognition via Large-Scale Weak Supervision · ICML 2023 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
robust fine-tuning |
0.6 | 1 | 2022 | Robust fine-tuning of zero-shot models · CVPR 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | Robust fine-tuning of zero-shot models · CVPR 2022 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.6 | 1 | 2022 | Robust fine-tuning of zero-shot models · CVPR 2022 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
weight-space ensembling |
0.6 | 1 | 2022 | Robust fine-tuning of zero-shot models · CVPR 2022 |
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining |
0.5 | 1 | 2021 | Learning Transferable Visual Models From Natural Language Supervision · ICML 2021 |
Computer vision › Vision and language › vision-language model
vision-language model guidance |
0.3 | 1 | 2025 | CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image · ICCV 2025 |
Computer vision › 3D vision
neural radiance field |
0.2 | 1 | 2024 | WateRF: Robust Watermarks in Radiance Fields for Protection of Copyrights · CVPR 2024 |
Usability and user experience research
user modeling |
0.2 | 1 | 2015 | Predicting User Performance and Learning in Human-Computer Interaction with the Herbal Compiler · ACM Trans. Comput. Hum. Interact. 2015 |
Machine learning › Learning paradigms
weakly supervised learning |
0.2 | 1 | 2023 | Robust Speech Recognition via Large-Scale Weak Supervision · ICML 2023 |
Computer vision › Vision and language › cross-modal supervision
natural language supervision |
0.1 | 1 | 2021 | Learning Transferable Visual Models From Natural Language Supervision · ICML 2021 |
Accessibility and assistive technology › screen reader accessibility
screen reader navigation |
0.1 | 2 | 2009 | SEA: Segment-enrich-annotate paradigm for adapting dialog-based content for improved accessibility · ACM Trans. Inf. Syst. 2009 Topic segmentation of message hierarchies for indexing and navigation support · WWW 2005 |
Information retrieval › document retrieval
context-sensitive retrieval |
0.1 | 1 | 2009 | Skip-and-prune: cosine-based top-k query processing for efficient context-sensitive document retrieval · SIGMOD Conference 2009 |
Information retrieval › similarity search
near-duplicate detection |
0.1 | 1 | 2009 | Efficient overlap and content reuse detection in blogs and online news articles · WWW 2009 |
Information retrieval › retrieval models
ranked retrieval |
0.1 | 1 | 2009 | Skip-and-prune: cosine-based top-k query processing for efficient context-sensitive document retrieval · SIGMOD Conference 2009 |
Information retrieval
retrieval models |
0.1 | 1 | 2009 | Skip-and-prune: cosine-based top-k query processing for efficient context-sensitive document retrieval · SIGMOD Conference 2009 |
Information retrieval › document processing › document analysis
text reuse detection |
0.1 | 1 | 2009 | Efficient overlap and content reuse detection in blogs and online news articles · WWW 2009 |
Query processing and optimization
top-k query processing |
0.1 | 1 | 2009 | Skip-and-prune: cosine-based top-k query processing for efficient context-sensitive document retrieval · SIGMOD Conference 2009 |
Learning and educational technologies
skill acquisition |
0.1 | 1 | 2015 | Predicting User Performance and Learning in Human-Computer Interaction with the Herbal Compiler · ACM Trans. Comput. Hum. Interact. 2015 |
Information retrieval › text analysis › text segmentation
topic segmentation |
0.1 | 1 | 2005 | Topic segmentation of message hierarchies for indexing and navigation support · WWW 2005 |
Web and social media mining › social media analysis
blog analysis |
0.0 | 1 | 2009 | Efficient overlap and content reuse detection in blogs and online news articles · WWW 2009 |
Methods — techniques the papers use, named apart from their topics
patch-wise loss · 1.5discrete wavelet transform · 1.5deferred backpropagation · 1.5visual language model · 0.9transformer · 0.9cross-attention · 0.9multi-task learning · 0.7large-scale weak supervision · 0.7zero-shot inference · 0.6weight ensembling · 0.6cognitive modeling · 0.2GOMS · 0.2ACT-R · 0.2topic segmentation · 0.1text analysis · 0.1user study · 0.1qsign algorithm · 0.1incremental processing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing differentially private distributed k-means clustering through data-driven sensitivity calibration and optimized dropout handling
Jong Wook Kim, Sae-Hong Cho |
Inf. Sci. | 1 |
| 2026 | Privacy-preserving trajectory data publication: A distributed approach without trusted servers
Jong Wook Kim, Beakcheol Jang |
J. Netw. Comput. Appl. | 1 |
| 2025 | CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View ImageabstractRecently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. However, unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from a single image. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under single-view settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis. Wonseok Roh, Hwanhee Jung, Jong Wook Kim, Seunggwan Lee, Innfarn Yoo, Andreas Lugmayr, Seunggeun Chi, Karthik Ramani, Sangpil Kim |
ICCV | 3 |
| 2025 | Spatial-Channel Mixing Block for Neural Network-based Video Coding (NNVC) ToolsabstractIn this paper, we propose an efficient building block for neural network-based video coding (NNVC) tools, which are part of the post-VVC research. Our proposed Spatial-Channel Mixing (SCM) block explicitly separates spatial and channel mixing operations and applies them sequentially. The SCM block is integrated into existing NNLF (NN-based in-Loop Filter) and NNSR (NN-based Super Resolution) tools in NNVC by replacing their backbone blocks. Under the Common Test Conditions (CTC) of the Joint Video Experts Team (JVET) NNVC, the proposed NNLF and NNSR showed consistent BD-Rate improvements across Y, U, and V channels in various configurations, along with reductions in parameter count and MACs per pixel, demonstrating the effectiveness and efficiency of the proposed SCM block for NNVC tools. HyunDong Cho, Suyong Bahk, Jong Wook Kim, Donghyun Kim 0017, Sung-Chang Lim, Hui Yong Kim |
VCIP | 3 |
| 2025 | High-quality three-dimensional cartoon avatar reconstruction with Gaussian splatting
MinHyuk Jang, Jong Wook Kim, Youngdong Jang, Donghyun Kim 0006, Wonseok Roh, Inyong Hwang, Guang Lin 0001, Sangpil Kim |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | WateRF: Robust Watermarks in Radiance Fields for Protection of CopyrightsabstractThe advances in the Neural Radiance Fields (NeRF) research offer extensive applications in diverse domains, but protecting their copyrights has not yet been researched in depth. Recently, NeRF watermarking has been considered one of the pivotal solutions for safely deploying NeRF-based 3D representations. However, existing methods are designed to apply only to implicit or explicit NeRF representations. In this work, we introduce an innovative watermarking method that can be employed in both representations of NeRF. This is achieved by fine-tuning NeRF to embed binary messages in the rendering process. In detail, we propose utilizing the discrete wavelet transform in the NeRF space for watermarking. Furthermore, we adopt a deferred back-propagation technique and introduce a combination with the patch-wise loss to improve rendering quality and bit accuracy with minimum trade-offs. We evaluate our method in three different aspects: capacity, invisibility, and robustness of the embedded watermarks in the 2D-rendered images. Our method achieves state-of-the-art performance with faster training speed over the compared state-of-the-art methods. Project page: https://kuai-lab.github.io/cvpr2024waterf/ Youngdong Jang, Dong In Lee, MinHyuk Jang, Jong Wook Kim, Feng Yang 0008, Sangpil Kim |
CVPR | 4 |
| 2024 | Privacy-preserving generation and publication of synthetic trajectory microdata: A comprehensive survey
Jong Wook Kim, Beakcheol Jang |
J. Netw. Comput. Appl. | 1 |
| 2023 | Robust Speech Recognition via Large-Scale Weak SupervisionabstractWe study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results without the need for any dataset specific fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing. Alec Radford, Jong Wook Kim, Greg Brockman, Christine McLeavey, Ilya Sutskever |
ICML | 2 |
| 2023 | Long-Term Influenza Outbreak Forecast Using Time-Precedence Correlation of Web DataabstractInfluenza leads to many deaths every year and is a threat to human health. For effective prevention, traditional national-scale statistical surveillance systems have been developed, and numerous studies have been conducted to predict influenza outbreaks using web data. Most studies have captured the short-term signs of influenza outbreaks, such as one-week prediction using the characteristics of web data uploaded in real time; however, long-term predictions of more than 2-10 weeks are required to effectively cope with influenza outbreaks. In this study, we determined that web data uploaded in real time have a time-precedence relationship with influenza outbreaks. For example, a few weeks before an influenza pandemic, the word "colds" appears frequently in web data. The web data after the appearance of the word "colds" can be used as information for forecasting future influenza outbreaks, which can improve long-term influenza prediction accuracy. In this study, we propose a novel long-term influenza outbreak forecast model utilizing the time precedence between the emergence of web data and an influenza outbreak. Based on the proposed model, we conducted experiments on: 1) selecting suitable web data for long-term influenza prediction; 2) determining whether the proposed model is regionally dependent; and 3) evaluating the accuracy according to the prediction timeframe. The proposed model showed a correlation of 0.87 in the long-term prediction of ten weeks while significantly outperforming other state-of-the-art methods. Beakcheol Jang, Inhwan Kim, Jong Wook Kim |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Robust fine-tuning of zero-shot modelsabstractLarge pre-trained models such as CLIP or ALIGN offer consistent accuracy across a range of data distributions when performing zero-shot inference (i.e., without fine-tuning on a specific dataset). Although existing fine-tuning methods substantially improve accuracy on a given target distribution, they often reduce robustness to distribution shifts. We address this tension by introducing a simple and effective method for improving robustness while fine-tuning: ensembling the weights of the zero-shot and fine-tuned models (WiSE-FT). Compared to standard fine-tuning, WiSE-FT provides large accuracy improvements under distribution shift, while preserving high accuracy on the target distribution. On ImageNet and five derived distribution shifts, WiSE-FT improves accuracy under distribution shift by 4 to 6 percentage points (pp) over prior work while increasing ImageNet accuracy by 1.6 pp. WiSE-FT achieves similarly large robustness gains (2 to 23 pp) on a diverse set of six further distribution shifts, and accuracy gains of 0.8 to 3.3 pp compared to standard fine-tuning on commonly used transfer learning datasets. These improvements come at no additional computational cost during fine-tuning or inference. Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, Ludwig Schmidt |
CVPR | 3 |
| 2022 | Deep similarity analysis and forecasting of actual outbreak of major infectious diseases using Internet-Sourced data
Beakcheol Jang, Yeongha Kim, Gun Il Kim, Jong Wook Kim |
J. Biomed. Informatics | 4 |
| 2022 | Privacy-preserving mechanisms for location privacy in mobile crowdsensing: A survey
Jong Wook Kim, Kennedy Edemacu, Beakcheol Jang |
J. Netw. Comput. Appl. | 1 |
| 2022 | Deep learning-based privacy-preserving framework for synthetic trajectory generation
Jong Wook Kim, Beakcheol Jang |
J. Netw. Comput. Appl. | 1 |
| 2021 | Learning Transferable Visual Models From Natural Language SupervisionabstractState-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever |
ICML | 2 |
| 2021 | A Survey Of differential privacy-based techniques and their applicability to location-Based services
Jong Wook Kim, Kennedy Edemacu, Jong Seon Kim, Yon Dohn Chung, Beakcheol Jang |
Comput. Secur. | 1 |
| 2021 | Reliability check via weight similarity in privacy-preserving multi-party machine learning
Kennedy Edemacu, Beakcheol Jang, Jong Wook Kim |
Inf. Sci. | 3 |
| 2021 | A deep attention model to forecast the Length Of Stay and the in-hospital mortality right on admission from ICD codes and demographic data
Gaspard Harerimana, Jong Wook Kim, Beakcheol Jang |
J. Biomed. Informatics | 2 |
| 2020 | Practical Advice on How to Run Human Behavioral Studies
Frank E. Ritter, Jonathan H. Morgan, Jong Wook Kim |
CogSci | 3 |
| 2020 | Collaborative Ehealth Privacy and Security: An Access Control With Attribute Revocation Based on OBDD Access StructureabstractThe digitization of health records due to technological developments has paved the way for patients to be collaboratively treated by different healthcare institutions. In collaborative ehealth systems, a patient's health data is stored remotely in the cloud for sharing with different healthcare service providers. However, the use of third parties for storage exposes the data to several privacy and security violation threats. Ciphertext policy attribute-based encryption (CP-ABE) which provides a fine-grained access control is a promising solution to privacy and security issues in the cloud environment and as a result, it has been widely studied for secure sharing of health data in cloud-based ehealth systems. Addressing the aspects of expressiveness, efficiency, user collusion resistance and attribute/user revocation in CP-ABE have been at the forefront of these studies. Thus, in this article, we proposed a novel expressive, efficient and collusion-resistant access control scheme with immediate attribute/user revocation for secure sharing of health data in collaborative ehealth systems. The proposed scheme additionally achieves forward and backward security. To realize these features, our access control is based on the ordered binary decision diagram (OBDD) access structure and it binds the user keys to the user identities. Security and performance analysis show that our proposed scheme is secure, expressive and efficient. Kennedy Edemacu, Beakcheol Jang, Jong Wook Kim |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | Neural Music Synthesis for Flexible Timbre ControlabstractThe recent success of raw audio waveform synthesis models like WaveNet motivates a new approach for music synthesis, in which the entire process - creating audio samples from a score and instrument information - is modeled using generative neural networks. This paper describes a neural music synthesis model with flexible timbre controls, which consists of a recurrent neural network conditioned on a learned instrument embedding followed by a WaveNet vocoder. The learned embedding space successfully captures the diverse variations in timbres within a large dataset and enables timbre control and morphing by interpolating between instruments in the embedding space. The synthesis quality is evaluated both numerically and perceptually, and an interactive web demo is presented. Jong Wook Kim, Rachel M. Bittner, Aparna Kumar, Juan Pablo Bello |
ICASSP | 1 |
| 2018 | Crepe: A Convolutional Representation for Pitch EstimationabstractThe task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are based on a combination of DSP pipelines and heuristics. While such techniques perform very well on average, there remain many cases in which they fail to correctly estimate the pitch. In this paper, we propose a data-driven pitch tracking algorithm, CREPE, which is based on a deep convolutional neural network that operates directly on the time-domain waveform. We show that the proposed model produces state-of-the-art results, performing equally or better than pYIN. Furthermore, we evaluate the model's generalizability in terms of noise robustness. A pre-trained version of CREPE is made freely available as an open-source Python module for easy application. Jong Wook Kim, Justin Salamon, Peter Li, Juan Pablo Bello |
ICASSP | 1 |
| 2015 | Learning, Forgetting, and Relearning for Keystroke- and Mouse-Driven Tasks: Relearning Is ImportantabstractWe investigate performance change arising through learning, forgetting, and relearning. Participants learned a spreadsheet task with either keystroke-driven (keyboard, n = 30) or mouse-based menu-driven (mouse, n = 30) commands. Their performance confirmed the power law of practice. The keyboard users learned to complete the task faster than the mouse users on the last learning session (Day 4). At a 6-day retention interval, the mouse users were observed to forget more—they took more time to complete the task than the keyboard users. Of interest, the participants in the two modality groups showed no reliable differences in their forgetting under the retention of 12 and 18 days. With additional practice, the mouse group users with the 6-day retention relearned more—they reliably reduced the time to complete the task in comparison to the paired keyboard group. These results help understand why people may choose to use a mouse-driven graphical user interface rather than a keystroke-driven interface: People choosing to use a mouse-based menu-driven interface may not need to use a knowledge-in-the-head strategy but knowledge-in-the-world, and may be doing so because this strategy provides better relearning, rather than because it is faster or easier initially or because it is better for learning or forgetting. These results provide a richer explanation of why menu-driven interfaces (knowledge-in-the-world) are more ubiquitous and suggests when they can be replaced, for example, where use is infrequent but often enough that forgetting does not substantially occur. Our results provide preliminary suggestions for choosing optimal training strategies and supporting these strategies in terms of the three stages of learning and forgetting. Jong Wook Kim, Frank E. Ritter |
Hum. Comput. Interact. | 1 |
| 2015 | Predicting User Performance and Learning in Human-Computer Interaction with the Herbal CompilerabstractWe report a way to build a series of GOMS-like cognitive user models representing a range of performance at different stages of learning. We use a spreadsheet task across multiple sessions as an example task; it takes about 20--30 min. to perform. The models were created in ACT-R using a compiler. The novice model has 29 rules and 1,152 declarative memory task elements (chunks)—it learns to create procedural knowledge to perform the task. The expert model has 617 rules and 614 task chunks (that it does not use) and 538 command string chunks—it gets slightly faster through limited declarative learning of the command strings and some further production compilation; there are a range of intermediate models. These models were tested against aggregate and individual human learning data, confirming the models’ predictions. This work suggests that user models can be created that learn like users while doing the task. Jaehyon Paik, Jong Wook Kim, Frank E. Ritter, David Reitter |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2014 | Practical Advice on How to Run Human Behavioral Studies
Frank E. Ritter, Jong Wook Kim |
CogSci | 2 |
| 2012 | Practical Advice on How to Run Human Behavioral Studies
Frank E. Ritter, Jong Wook Kim |
CogSci | 2 |
| 2010 | Efficient wikipedia-based semantic interpreter by exploiting top-k processingabstractProper representation of the meaning of texts is crucial to enhancing many data mining and information retrieval tasks, including clustering, computing semantic relatedness between texts, and searching. Representing of texts in the concept space derived from Wikipedia has received growing attention recently, due to its comprehensiveness and expertise, This concept-based representation is capable of extracting semantic relatedness between texts that cannot be deduced with the bag of words model. A key obstacle, however, for using Wikipedia as a semantic interpreter is that the sheer size of the concepts derived from Wikipedia makes it hard to efficiently map texts into concept-space. In this paper, we develop an efficient algorithm which is able to represent the meaning of a text by using the concepts that best match it. In particular, our approach first computes the approximate top-k concepts that are most relevant to the given text. We then leverage these concepts for representing the meaning of the given text. The experimental results show that the proposed technique provides significant gains in execution time over current solutions to the problem. Jong Wook Kim, Ashwin Kashyap, Sandilya Bhamidipati |
CIKM | 1 |
| 2009 | Enabling accessible interfaces to digital library contentabstractMost of the Web interfaces are primarily designed for people with sight, with visually rich features that makes effective use of the tools to enhance visual usability but in process making it impossible for users who are blind or visually impaired to use them. In this work, our goal is to improve participation to NSF's National Science Digital Library (NSDL) by teachers, librarians, and learners who are blind. The middleware for accessible information spaces on NSDL (MAISON) is enhancing the accessibility of NSDL, its internal and external resources and existing services (such as strand maps of educational benchmarks). Relying on cutting-edge, context-aware graph segmentation, filtering and summarization, and concept propagation techniques, the middleware provides information space adaptation, reduction, and preview services through open Web-based service APIs to enable implementation of informative navigation interfaces that are able to reduce the complexity of the information space and provide previews to prevent user disorientation. Syed Toufeeq Ahmed, K. Selçuk Candan, Suganthi Cidambaram, Shruti Gaur, Jong Wook Kim, Mijung Kim, Hari Sundaram, Renwei Yu |
ICME | 5 |
| 2009 | PICC Counting: Who Needs Joins When You Can Propagate Efficiently?abstractCounting is a common task in many data mining applications, including market basket data analysis, scientific inquiry, and other high dimensional data management applications. Given a single table, obtaining the instance counts of the entries in the table is relatively cheap. In situations where the attributes of interest are distributed across different tables, however, the problem of computing instance counts can be very expensive. The naive solution, joining all the relevant relations to obtain a single table suitable for counting, is rarely practical. In this paper, we propose PICC (Propagation-based Instance Counts on Concise Graphs), a novel counting technique for discovering instance counts in databases. We first propose a propagation-based instance counting scheme which avoids joins to obtain a single table. We then present a method for summarizing a database into a concise synopsis and describe how to use this along with the propagation scheme to estimate the required counts efficiently. The experiment results show that the proposed technique, PICC, provides significant execution time and accuracy gains over the existing solutions to this problem. Jong Wook Kim, K. Selçuk Candan |
SDM | 1 |
| 2009 | Skip-and-prune: cosine-based top-k query processing for efficient context-sensitive document retrievalabstractKeyword search and ranked retrieval together emerged as popular data access paradigms for various kinds of data, from web pages to XML and relational databases. A user can submit keywords without knowing much (sometimes nothing) about the complex structure underlying a data collection, yet the system can identify, rank, and return a set of relevant matches by exploiting statistics about the distribution and structure of the data. Keyword-based data models are also suitable for capturing user's search context in terms of weights associated to the keywords in the query. Given a search context, the data in the database can also be re-interpreted for semantically correct retrieval. This option, however, is often ignored as the cost of re-assessing the content in the database naively tends to be prohibitive. In this paper, we first argue that top-k query processing can help tackle this challenge by re-assessing only the relevant parts of the database, efficiently. A road-block in this process, however, is that most efficient implementations of top-k query processing assume that the scoring function is monotonic, whereas the cosine-based scoring function needed for re-interpretation of content based on user context is not. In this paper, we develop an efficient top-k query processing algorithm, skip-and-prune (SnP), which is able to process top-k queries under cosine-based non-monotonic scoring functions. We compare the use of proposed algorithm against the alternative implementations of the context-aware retrieval, including naive top-k, accumulator-based inverted files, and full-scan. The experiment results show that while being fast, naive top-k is not an effective solution due to the non-monotonicity of underlying scoring function. The proposed technique, SnP, however, matches the precision of accumulator-based inverted files and full-scan, yet it is orders of magnitude faster than these. Jong Wook Kim, K. Selçuk Candan |
SIGMOD Conference | 1 |
| 2009 | Efficient overlap and content reuse detection in blogs and online news articlesabstractThe use of blogs to track and comment on real world (political, news, entertainment) events is growing. Similarly, as more individuals start relying on the Web as their primary information source and as more traditional media outlets try reaching consumers through alternative venues, the number of news sites on the Web is also continuously increasing. Content-reuse, whether in the form of extensive quotations or content borrowing across media outlets, is very common in blogs and news entries outlets tracking the same real-world event. Knowledge about which web entries re-use content from which others can be an effective asset when organizing these entries for presentation. On the other hand, this knowledge is not cheap to acquire: considering the size of the related space web entries, it is essential that the techniques developed for identifying re-use are fast and scalable. Furthermore, the dynamic nature of blog and news entries necessitates incremental processing for reuse detection. In this paper, we develop a novel qSign algorithm that efficiently and effectively analyze the blogosphere for quotation and reuse identification. Experiment results show that with qSign processing time gains from 10X to 100X are possible while maintaining reuse detection rates of upto 90%. Furthermore, processing time gains can be pushed multiple orders of magnitude (from 100X to 1000X) for 70% recall. Jong Wook Kim, K. Selçuk Candan, Jun'ichi Tatemura |
WWW | 1 |
| 2009 | SEA: Segment-enrich-annotate paradigm for adapting dialog-based content for improved accessibilityabstractWhile navigation within complex information spaces is a problem for all users, the problem is most evident with individuals who are blind who cannot simply locate, point, and click on a link in hypertext documents with a mouse. Users who are blind have to listen searching for the link in the document using only the keyboard and a screen reader program, which may be particularly inefficient in large documents with many links or deep hierarchies that are hard to navigate. Consequently, they are especially penalized when the information being searched is hidden under multiple layers of indirections. In this article, we introduce a segment-enrich-annotate (SEA) paradigm for adapting digital content with deep structures for improved accessibility. In particular, we instantiate and evaluate this paradigm through the iCare-Assistant, an assistive system for helping students who are blind in accessing Web and electronic course materials. Our evaluations, involving the participation of students who are blind, showed that the iCare-Assistant system, built based on the SEA paradigm, reduces the navigational overhead significantly and enables user who are blind access complex online course servers effectively. K. Selçuk Candan, Mehmet Emin Dönderler, Terri Hedgpeth, Jong Wook Kim, Maria Luisa Sapino |
ACM Trans. Inf. Syst. | 4 |
| 2006 | CP/CV: concept similarity mining without frequency information from domain describing taxonomiesabstractDomain specific ontologies are heavily used in many applications. For instance, these form the bases on which similarity/dissimilarity between keywords are extracted for various knowledge discovery and retrieval tasks. Existing similarity computation schemes can be categorized as (a) structure- or (b) information-based approaches. Structure based approaches compute dissimilarity between keywords using a (weighted) count of edges between two keywords. Information-base approaches, on the other hand, leverage available corpora to extract additional information, such as keyword frequency, to achieve better performance in similarity computation than structure-based approaches. Unfortunately, in many application domains (such as applications that rely on unique-keys in a relational database), frequency information required by information-based approaches does not exist. In this paper, we note that there is a third way of computing similarity: if each node in a given hierarchy can be represented as a vector of related concepts, these vectors could be compared to compute similarities. This requires mapping concept-nodes in a given hierarchy onto a concept space. In this paper, we propose a concept propagation (CP) scheme, which relies on the semantical relationships between concepts implied by the structure of the hierarchy to annotate each concept-node with a concept-vector (CV). We refer to this approach as CP/CV. Comparison of keyword similarity results shows that CP/CV provides significantly better (upto 33%) results than existing structure-based schemes. Also, even if CP/CV does not assume the availability of an appropriate corpus to extract keyword frequency information, our approach matches (and slightly improves on) the performance of information-based approaches. Jong Wook Kim, K. Selçuk Candan |
CIKM | 1 |
| 2006 | FMware: Middleware for Efficient Filtering and Matching of XML Messages with Local Data
K. Selçuk Candan, Mehmet Emin Dönderler, Yan Qi 0002, Jaikannan Ramamoorthy, Jong Wook Kim |
Middleware | 5 |
| 2006 | Discovering mappings in hierarchical data from multiple sources using the inherent structure
K. Selçuk Candan, Jong Wook Kim, Huan Liu 0001, Reshma Suvarna |
Knowl. Inf. Syst. | 2 |
| 2005 | Topic segmentation of message hierarchies for indexing and navigation supportabstractMessage hierarchies in web discussion boards grow with new postings. Threads of messages evolve as new postings focus within or diverge from the original themes of the threads. Thus, just by investigating the subject headings or contents of earlier postings in a message thread, one may not be able to guess the contents of the later postings. The resulting navigation problem is further compounded for blind users who need the help of a screen reader program that can provide only a linear representation of the content. We see that, in order to overcome the navigation obstacle for blind as well as sighted users, it is essential to develop techniques that help identify how the content of a discussion board grows through generalizations and specializations of topics. This knowledge can be used in segmenting the content in coherent units and guiding the users through segments relevant to their navigational goals. Our experimental results showed that the segmentation algorithm described in this paper provides up to 80-85% success rate in labeling messages. The algorithm is being deployed in a software system to reduce the navigational load of blind students in accessing web-based electronic course materials; however, we note that the techniques are equally applicable for developing web indexing and summarization tools for users with sight. Jong Wook Kim, K. Selçuk Candan, Mehmet Emin Dönderler |
WWW | 1 |
| 2003 | Efficient preprocessing of XML queries using structured signatures
Yon Dohn Chung, Jong Wook Kim, Myoung-Ho Kim |
Inf. Process. Lett. | 2 |