Jiwoon Jeon

dblp:69/6160 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
1since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 4 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Information retrieval · 100%
Artificial intelligence
2 papers
Representation and self-supervised learning · 64% Image recognition and object detection · 36%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
question answering
0.232008
Retrieval models for question and answer archives · SIGIR 2008
A framework to predict the quality of answers with non-textual features · SIGIR 2006
Finding semantically similar questions based on their answers · SIGIR 2005
Information retrieval
retrieval models
0.132008
Retrieval models for question and answer archives · SIGIR 2008
A framework to predict the quality of answers with non-textual features · SIGIR 2006
Automatic image annotation and retrieval using cross-media relevance models · SIGIR 2003
Information retrieval › document processing
document analysis
0.112009
High precision retrieval using relevance-flow graph · SIGIR 2009
Information retrieval
precision-oriented retrieval
0.112009
High precision retrieval using relevance-flow graph · SIGIR 2009
Information retrieval › cross-language information retrieval
translation-based language model
0.112008
Retrieval models for question and answer archives · SIGIR 2008
Machine learning › Representation and self-supervised learning › multimodal representation learning
multi-modal clustering
0.112007
Multi-modal Clustering for Multimedia Collections · CVPR 2007
Information retrieval › question answering › community question answering
answer quality prediction
0.112006
A framework to predict the quality of answers with non-textual features · SIGIR 2006
Computer vision › Image recognition and object detection
image annotation
0.012003
A Model for Learning the Semantics of Pictures · NIPS 2003
Information retrieval › image retrieval
image annotation and retrieval
0.012003
Automatic image annotation and retrieval using cross-media relevance models · SIGIR 2003
Information retrieval
image retrieval
0.012003
A Model for Learning the Semantics of Pictures · NIPS 2003
Information retrieval
multimedia analysis and retrieval
0.012003
Automatic image annotation and retrieval using cross-media relevance models · SIGIR 2003
Information retrieval › image retrieval
text-based image retrieval
0.012003
A Model for Learning the Semantics of Pictures · NIPS 2003
Information retrieval › retrieval models › language model
query likelihood model
0.012008
Retrieval models for question and answer archives · SIGIR 2008
Information retrieval › retrieval models › probabilistic retrieval model
binary independence model
0.012003
Automatic image annotation and retrieval using cross-media relevance models · SIGIR 2003

Methods — techniques the papers use, named apart from their topics

k-means · 0.1combinatorial markov random field · 0.1translation-based language model · 0.1query likelihood · 0.1probabilistic model · 0.1joint probabilistic model · 0.1non-textual feature analysis · 0.1similarity measure · 0.1language modeling · 0.1probabilistic modeling · 0.0clustering · 0.0
YearPublicationVenuePosition
2024 Real-Time Polyp Detection in Colonoscopy using Lightweight Transformer
abstract
Colorectal cancer (CRC) represents a major global health challenge, and early detection of polyps is crucial in preventing its progression. Although colonoscopy is the gold standard for polyp detection, it has limitations, such as human error and missed detection rates. In response, computer-aided detection (CADe) systems have been developed to enhance the efficiency and accuracy of polyp detection. As deep learning gained prominence, the incorporation of Convolutional Neural Networks (CNNs) into CADe systems emerged as a breakthrough approach. However, CADe systems based on CNNs often demand significant computational resources, making them unsuitable for deployment in resource-constrained environments. To mitigate this, we propose a novel and lightweight polyp detection model that integrates a Transformer layer into the You Only Look Once (YOLO) architecture, focusing on optimizing the neck part responsible for feature fusion and rescaling. Our model demonstrates a substantial reduction in computational complexity and the number of parameters, without compromising detection performances. The lightweight model makes it accessible and feasibly deployable in medically underserved regions, serving a significant public interest by potentially expanding the reach of critical diagnostic tools for CRC prevention. By optimizing the architecture to reduce resource requirements while maintaining performance, our model becomes a practical solution to assist healthcare professionals in the real-time identification of polyps, even with resource-constraint devices.
Youngbeom Yoo, Jae Young Lee 0002, Jiwoon Jeon, Junmo Kim 0002
WACV4
2011 Sentence-based relevance flow analysis for high accuracy retrieval
abstract
Traditional ranking models for information retrieval lack the ability to make a clear distinction between relevant and nonrelevant documents at top ranks if both have similar bag-of-words representations with regard to a user query. We aim to go beyond the bag-of-words approach to document ranking in a new perspective, by representing each document as a sequence of sentences. We begin with an assumption that relevant documents are distinguishable from nonrelevant ones by sequential patterns of relevance degrees of sentences to a query. We introduce the notion of relevance flow, which refers to a stream of sentence-query relevance within a document. We then present a framework to learn a function for ranking documents effectively based on various features extracted from their relevance flows and leverage the output to enhance existing retrieval models. We validate the effectiveness of our approach by performing a number of retrieval experiments on three standard test collections, each comprising a different type of document: news articles, medical references, and blog posts. Experimental results demonstrate that the proposed approach can improve the retrieval performance at the top ranks significantly as compared with the state-of-the-art retrieval models regardless of document type.
Jung-Tae Lee, Jangwon Seo, Jiwoon Jeon, Hae-Chang Rim
J. Assoc. Inf. Sci. Technol.3
2009 High precision retrieval using relevance-flow graph
abstract
Traditional bag-of-words information retrieval models use aggregated term statistics to measure the relevance of documents, making it difficult to detect non-relevant documents that contain many query terms by chance or in the wrong context. In-depth document analysis is needed to filter out these deceptive documents. In this paper, we hypothesize that truly relevant documents have relevant sentences in predictable patterns. Our experimental results show that we can successfully identify and exploit these patterns to significantly improve retrieval precision at top ranks.
Jangwon Seo, Jiwoon Jeon
SIGIR2
2008 Retrieval models for question and answer archives
abstract
Retrieval in a question and answer archive involves finding good answers for a user's question. In contrast to typical document retrieval, a retrieval model for this task can exploit question similarity as well as ranking the associated answers. In this paper, we propose a retrieval model that combines a translation-based language model for the question part with a query likelihood approach for the answer part. The proposed model incorporates word-to-word translation probabilities learned through exploiting different sources of information. Experiments show that the proposed translation based language model for the question part outperforms baseline methods significantly. By combining with the query likelihood language model for the answer part, substantial additional effectiveness improvements are obtained.
Xiaobing Xue, Jiwoon Jeon, W. Bruce Croft
SIGIR2
2007 Multi-modal Clustering for Multimedia Collections
abstract
Most of the online multimedia collections, such as picture galleries or video archives, are categorized in a fully manual process, which is very expensive and may soon be infeasible with the rapid growth of multimedia repositories. In this paper, we present an effective method for automating this process within the unsupervised learning framework. We exploit the truly multi-modal nature of multimedia collections - they have multiple views, or modalities, each of which contributes its own perspective to the collection's organization. For example, in picture galleries, image captions are often provided that form a separate view on the collection. Color histograms (or any other set of global features) form another view. Additional views are blobs, interest points and other sets of local features. Our model, called Comraf* (pronounced Comraf-Star), efficiently incorporates various views in multi-modal clustering, by which it allows great modeling flexibility. Comraf* is a light-weight version of the recently introduced combinatorial Markov random field (Comraf). We show how to translate an arbitrary Comraf into a series of Comraf* models, and give an empirical evidence for comparable effectiveness of the two. Comraf* demonstrates excellent results on two real-world image galleries: it obtains 2.5-3 times higher accuracy compared with a uni-modal k-means.
Ron Bekkerman, Jiwoon Jeon
CVPR2
2006 A framework to predict the quality of answers with non-textual features
abstract
New types of document collections are being developed by various web services. The service providers keep track of non-textual features such as click counts. In this paper, we present a framework to use non-textual features to predict the quality of documents. We also show our quality measure can be successfully incorporated into the language modeling-based retrieval model. We test our approach on a collection of question and answer pairs gathered from a community based question answering service where people ask and answer questions. Experimental results using our quality measure show a significant improvement over our baseline.
Jiwoon Jeon, W. Bruce Croft, Joon Ho Lee
SIGIR1
2005 Finding similar questions in large question and answer archives
abstract
There has recently been a significant increase in the number of community-based question and answer services on the Web where people answer other peoples' questions. These services rapidly build up large archives of questions and answers, and these archives are a valuable linguistic resource. One of the major tasks in a question and answer service is to find questions in the archive that a semantically similar to a user's question. This enables high quality answers from the archive to be retrieved and removes the time lag associated with a community-based system. In this paper, we discuss methods for question retrieval that are based on using the similarity between answers in the archive to estimate probabilities for a translation-based retrieval model. We show that with this model it is possible to find semantically similar questions with relatively little word overlap.
Jiwoon Jeon, W. Bruce Croft, Joon Ho Lee
CIKM1
2005 Finding semantically similar questions based on their answers
abstract
A large number of question and answer pairs can be collected from question and answer boards and FAQ pages on the Web. This paper proposes an automatic method of finding the questions that have the same meaning. The method can detect semantically similar questions that have little word overlap because it calculates question-question similarities by using the corresponding answers as well as the questions. We develop two different similarity measures based on language modeling and compare them with the traditional similarity measures. Experimental results show that semantically similar questions pairs can be effectively found with the proposed similarity measures.
Jiwoon Jeon, W. Bruce Croft, Joon Ho Lee
SIGIR1
2003 A Model for Learning the Semantics of Pictures
abstract
We propose an approach to learning the semantics of images which al- lows us to automatically annotate an image with keywords and to retrieve images based on text queries. We do this using a formalism that models the generation of annotated images. We assume that every image is di- vided into regions, each described by a continuous-valued feature vector. Given a training set of images with annotations, we compute a joint prob- abilistic model of image features and words which allow us to predict the probability of generating a word given the image regions. This may be used to automatically annotate and retrieve images given a word as a query. Experiments show that our model significantly outperforms the best of the previously reported results on the tasks of automatic image annotation and retrieval.
Victor Lavrenko, R. Manmatha, Jiwoon Jeon
NIPS3
2003 Automatic image annotation and retrieval using cross-media relevance models
abstract
Libraries have traditionally used manual image annotation for indexing and then later retrieving their image collections. However, manual image annotation is an expensive and labor intensive procedure and hence there has been great interest in coming up with automatic ways to retrieve images based on content. Here, we propose an automatic approach to annotating and retrieving images based on a training set of images. We assume that regions in an image can be described using a small vocabulary of blobs. Blobs are generated from image features using clustering. Given a training set of images with annotations, we show that probabilistic models allow us to predict the probability of generating a word given the blobs in an image. This may be used to automatically annotate and retrieve images given a word as a query. We show that relevance models allow us to derive these probabilities in a natural way. Experiments show that the annotation performance of this cross-media relevance model is almost six times as good (in terms of mean precision) than a model based on word-blob co-occurrence model and twice as good as a state of the art model derived from machine translation. Our approach shows the usefulness of using formal information retrieval models for the task of image annotation and retrieval.
Jiwoon Jeon, Victor Lavrenko, R. Manmatha
SIGIR1