Kyusong Lee

dblp:44/9770 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
5since 2021 · last 2024
0009-0003-8113-4667ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Image recognition and object detection · 42% Video understanding and tracking · 21% Multi-agent systems · 21%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 7 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
long video understanding
0.812024
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer · EMNLP 2024
Knowledge, reasoning and agents › Multi-agent systems
multimodal agent
0.812024
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer · EMNLP 2024
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.812024
How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection · AAAI 2024
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.312018
Unsupervised Discrete Sentence Representation Learning for Interpretable Neural Dialog Generation · ACL (1) 2018
Information retrieval
cross-modal retrieval
0.112021
VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › error detection
grammatical error detection
0.112011
Grammatical Error Detection for Corrective Feedback Provision in Oral Conversations · AAAI 2011
Machine learning › Generative modeling
variational autoencoder
0.112018
Unsupervised Discrete Sentence Representation Learning for Interpretable Neural Dialog Generation · ACL (1) 2018

Methods — techniques the papers use, named apart from their topics

vision-language pretraining · 0.8tool calling · 0.8frame retrieval · 0.8divide-and-conquer · 0.8weighted bag-of-words · 0.5variational autoencoder · 0.3discrete sentence representation learning · 0.3error pattern matching · 0.2confidence score classification · 0.2
YearPublicationVenuePosition
2024 How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection
abstract
Object detection (OD) in computer vision has made significant progress in recent years, transitioning from closed-set labels to open-vocabulary detection (OVD) based on large-scale vision-language pre-training (VLP). However, current evaluation methods and datasets are limited to testing generalization over object types and referral expressions, which do not provide a systematic, fine-grained, and accurate benchmark of OVD models' abilities. In this paper, we propose a new benchmark named OVDEval, which includes 9 sub-tasks and introduces evaluations on commonsense knowledge, attribute understanding, position understanding, object relation comprehension, and more. The dataset is meticulously created to provide hard negatives that challenge models' true understanding of visual and linguistic input. Additionally, we identify a problem with the popular Average Precision (AP) metric when benchmarking models on these fine-grained label datasets and propose a new metric called Non-Maximum Suppression Average Precision (NMS-AP) to address this issue. Extensive experimental results show that existing top OVD models all fail on the new tasks except for simple object types, demonstrating the value of the proposed dataset in pinpointing the weakness of current OVD models and guiding future research. Furthermore, the proposed NMS-AP metric is verified by experiments to provide a much more truthful evaluation of OVD models, whereas traditional AP metrics yield deceptive results. Data is available at https://github.com/om-ai-lab/OVDEval
Yiyang Yao, Jiajia Liao, Chunxin Fang, Kyusong Lee
AAAI7
2024 OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
abstract
Recent advancements in Large Language Models (LLMs) have expanded their capabilities to multimodal contexts, including comprehensive video understanding.However, processing extensive videos such as 24-hour CCTV footage or full-length films presents significant challenges due to the vast data and processing demands.Traditional methods, like extracting key frames or converting frames to text, often result in substantial information loss.To address these shortcomings, we develop OmAgent, efficiently stores and retrieves relevant video frames for specific queries, preserving the detailed content of videos.Additionally, it features an Divide-and-Conquer Loop capable of autonomous reasoning, dynamically invoking APIs and tools to enhance query processing and accuracy.This approach ensures robust video understanding, significantly reducing information loss.Experimental results affirm OmAgent's efficacy in handling various types of videos and complex tasks.Moreover, we have endowed it with greater autonomy and a robust tool-calling system, enabling it to accomplish even more intricate tasks.Code:
Heting Ying, Yibo Ma, Kyusong Lee
EMNLP5
2024 OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
abstract
Abstract The advancement of object detection (OD) in open‐vocabulary and open‐world scenarios is a critical challenge in computer vision. OmDet, a novel language‐aware object detection architecture and an innovative training mechanism that harnesses continual learning and multi‐dataset vision‐language pre‐training is introduced. Leveraging natural language as a universal knowledge representation, OmDet accumulates “visual vocabularies” from diverse datasets, unifying the task as a language‐conditioned detection framework. The multimodal detection network (MDN) overcomes the challenges of multi‐dataset joint training and generalizes to numerous training datasets without manual label taxonomy merging. The authors demonstrate superior performance of OmDet over strong baselines in object detection in the wild, open‐vocabulary detection, and phrase grounding, achieving state‐of‐the‐art results. Ablation studies reveal the impact of scaling the pre‐training visual vocabulary, indicating a promising direction for further expansion to larger datasets. The effectiveness of our deep fusion approach is underscored by its ability to learn jointly from multiple datasets, enhancing performance through knowledge sharing.
Kyusong Lee
IET Comput. Vis.3
2021 VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words
abstract
Xiaopeng Lu, Tiancheng Zhao, Kyusong Lee. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xiaopeng Lu, Kyusong Lee
ACL/IJCNLP (1)3
2021 SPARTA: Efficient Open-Domain Question Answering via Sparse Transformer Matching Retrieval
abstract
We introduce SPARTA, a novel neural retrieval method that shows great promise in performance, generalization, and interpretability for open-domain question answering.Unlike many neural ranking methods that use dense vector nearest neighbor search, SPARTA learns a sparse representation that can be efficiently implemented as an Inverted Index.The resulting representation enables scalable neural retrieval that does not require expensive approximate vector search and leads to better performance than its dense counterpart.We validated our approaches on 4 opendomain question answering (OpenQA) tasks and 11 retrieval question answering (ReQA) tasks.SPARTA achieves new state-of-the-art results across a variety of open-domain question answering tasks in both English and Chinese datasets, including open SQuAD, CMRC and etc. Analysis also confirms that the proposed method creates human interpretable representation and allows flexible control over the trade-off between performance and efficiency.
Xiaopeng Lu, Kyusong Lee
NAACL-HLT3
2018 Unsupervised Discrete Sentence Representation Learning for Interpretable Neural Dialog Generation
abstract
The encoder-decoder dialog model is one of the most prominent methods used to build dialog systems in complex domains.Yet it is limited because it cannot output interpretable actions as in traditional systems, which hinders humans from understanding its generation process.We present an unsupervised discrete sentence representation learning method that can integrate with any existing encoderdecoder dialog models for interpretable response generation.Building upon variational autoencoders (VAEs), we present two novel models, DI-VAE and DI-VST that improve VAEs and can discover interpretable semantics via either auto encoding or context predicting.Our methods have been validated on real-world dialog datasets to discover semantic representations and enhance encoder-decoder models with interpretable generation. 1
Kyusong Lee, Maxine Eskénazi
ACL (1)2
2018 DialCrowd: A toolkit for easy dialog system assessment
abstract
When creating a dialog system, developers need to test each version to ensure that it is performing correctly.Recently the trend has been to test on large datasets or to ask many users to try out a system.Crowdsourcing has solved the issue of finding users, but it presents new challenges such as how to use a crowdsourcing platform and what type of test is appropriate.Di-alCrowd makes system assessment using crowdsourcing easier by providing tools, templates and analytics.This paper describes the services that DialCrowd provides and how it works.It also describes a test of DialCrowd by a group of dialog system developers.
Kyusong Lee, Alan W. Black, Maxine Eskénazi
SIGDIAL Conference1
2017 DialPort, Gone Live: An Update After A Year of Development
abstract
Kyusong Lee, Tiancheng Zhao, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David Traum, Stefan Ultes, Lina M. Rojas-Barahona, Milica Gasic, Steve Young, Maxine Eskenazi. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017.
Kyusong Lee, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David R. Traum, Stefan Ultes, Lina Maria Rojas-Barahona, Milica Gasic, Steve J. Young, Maxine Eskénazi
SIGDIAL Conference1
2017 Generative Encoder-Decoder Models for Task-Oriented Spoken Dialog Systems with Chatting Capability
abstract
Generative encoder-decoder models offer great promise in developing domaingeneral dialog systems.However, they have mainly been applied to open-domain conversations.This paper presents a practical and novel framework for building task-oriented dialog systems based on encoder-decoder models.This framework enables encoder-decoder models to accomplish slot-value independent decision-making and interact with external databases.Moreover, this paper shows the flexibility of the proposed method by interleaving chatting capability with a slotfilling system for better out-of-domain recovery.The models were trained on both real-user data from a bus information system and human-human chat data.Results show that the proposed framework achieves good performance in both offline evaluation metrics and in task success rate with human users.
Allen Lu, Kyusong Lee, Maxine Eskénazi
SIGDIAL Conference3
2016 DialPort: Connecting the spoken dialog research community to real user data
abstract
This paper describes a new spoken dialog portal that connects systems produced by the spoken dialog academic research community and gives them access to real users. We introduce a distributed, multi-modal, multi-agent prototype dialog framework that affords easy integration with various remote resources, ranging from end-to-end dialog systems to external knowledge APIs. The portal provides seamless passage from one spoken dialog system to another. To date, the DialPort portal has successfully connected to the multi-domain spoken dialog system at Cambridge University, the NOAA (National Oceanic and Atmospheric Administration) weather API and the Yelp API. We present statistics derived from log data gathered during preliminary tests of the portal on the performance of the portal and on the quality (seamlessness) of the transition from one system to another.
Kyusong Lee, Maxine Eskénazi
SLT2
2015 Open-domain personalized dialog system using user-interested topics in system responses
abstract
We built a personalized example-based dialog system that constructs its responses by considering entities that the user has uttered, and topics in which the user has expressed interest. The system analyzes user input utterances, then uses DBpedia and Freebase to extract relevant entities and topics. The extracted entities and topics are stored in personal knowledge memory and are used when the system selects responses from the example database and generates responses. We conducted a human experiment in which evaluators rated dialog systems based on subjective metrics. The proposed dialog system that uses topics that are of interest to the user achieved higher evaluation scores for both personalization and satisfaction than the baseline systems. These results demonstrate that the use of topics in the system response provides a sense that the system pays attention to the user's utterances; as a consequence the user has a satisfactory dialog experience.
Jeesoo Bang, Sangdo Han, Kyusong Lee, Gary Geunbae Lee
ASRU3
2015 Conversational Knowledge Teaching Agent that uses a Knowledge Base
abstract
When implementing a conversational educational teaching agent, user-intent understanding and dialog management in a dialog system are not sufficient to give users educational information.In this paper, we propose a conversational educational teaching agent that gives users some educational information or triggers interests on educational contents.The proposed system not only converses with a user but also answer questions that the user asked or asks some educational questions by integrating a dialog system with a knowledge base.We used the Wikipedia corpus to learn the weights between two entities and embedding of properties to calculate similarities for the selection of system questions and answers.
Kyusong Lee, Hongsuck Seo, Junhwi Choi, Sangjun Koo, Gary Geunbae Lee
SIGDIAL Conference1
2014 Vowel-reduction feedback system for non-native learners of English
abstract
In spoken English, vowels in non-stressed syllables are often reduced to a brief neutral vowel (e.g, e or ι). Non-native speakers of English may not use this `vowel reduction' correctly, so their utterances may sound unnatural. We propose an automatic system to provide feedback about vowel-reduction to non-native speakers of English. The system has three parts: it predicts vowel reduction, detects vowel reduction in speech, compares the prediction to the detected sound to generate a score then uses this score to provide corrective feedback to the speaker. The system had good accuracy and provided positive learning results for the user. The proposed system can be used as a part of a computer-assisted language learning system.
Jeesoo Bang, Kyusong Lee, Seonghan Ryu, Gary Geunbae Lee
ICASSP2
2014 Grammatical error correction based on learner comprehension model in oral conversation
abstract
We aim to provide grammar error feedback to learners. It is known that grammar error detection and feedback are challenging problems in written language, however, they become much more difficult tasks in oral conversation because it is difficult for a system to judge whether an error is due to grammar or automatic speech recognition (ASR). False alarms occur when a learner correctly utters a remark, but the system gives feedback implying an error. Minimizing the false alarm rate is especially critical in education applications because it is imperative that the tutor give correct instruction to learners. Thus, to reduce the false alarm rate in grammar error detection and feedback, we apply a partially observable Markov decision process (POMDP) when the system provides feedback about a learner's mistake. The POMDP models uncertainty between grammar errors and ASR errors. An additional advantage of our method is that “belief states” in POMDP can be used for learner models which indicate each individual learner's grammar comprehension level.
Kyusong Lee, Seonghan Ryu, Hongsuck Seo, Seokhwan Kim, Gary Geunbae Lee
SLT1
2013 Counseling Dialog System with 5W1H Extraction
Sangdo Han, Kyusong Lee, Gary Geunbae Lee
SIGDIAL Conference2
2012 Grammatical Error Annotation for Korean Learners of Spoken English
Hongsuck Seo, Kyusong Lee, Gary Geunbae Lee, Soo-Ok Kweon, Hae-Ri Kim
LREC2
2012 Generating grammar questions using corpus data in L2 learning
abstract
This paper examines how grammar questions are automatically generated for L2 learning by applying a sequential labeling technique to learner corpora. We developed a model that helps detect possible error positions and select the most appropriate form among choices. Discriminant models such as conditional random field and maximum entropy are used to generate the error identification question. Questions generated by the proposed method corresponded highly to questions that experts made. Our data-driven approach lends itself to any language without costing expensive expertise.
Kyusong Lee, Soo-Ok Kweon, Hongsuck Seo, Gary Geunbae Lee
SLT1
2011 Grammatical Error Detection for Corrective Feedback Provision in Oral Conversations
abstract
The demand for computer-assisted language learning systems that can provide corrective feedback on language learners’ speaking has increased. However, it is not a trivial task to detect grammatical errors in oral conversations because of the unavoidable errors of automatic speech recognition systems. To provide corrective feedback, a novel method to detect grammatical errors in speaking performance is proposed. The proposed method consists of two sub-models: the grammaticality-checking model and the error-type classification model. We automatically generate grammatical errors that learners are likely to commit and construct error patterns based on the articulated errors. When a particular speech pattern is recognized, the grammaticality-checking model performs a binary classification based on the similarity between the error patterns and the recognition result using the confidence score. The error-type classification model chooses the error type based on the most similar error pattern and the error frequency extracted from a learner corpus. The grammaticality checking method largely outperformed the two comparative models by 56.36% and 42.61% in F-score while keeping the false positive rate very low. The error-type classification model exhibited very high performance with a 99.6% accuracy rate. Because high precision and a low false positive rate are important criteria for the language-tutoring setting, the proposed method will be helpful for intelligent computer-assisted language learning systems.
Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee
AAAI3
2011 POMY: A Conversational Virtual Environment for Language Learning in POSTECH
Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee
SIGDIAL Conference2
2011 Grammatical error simulation for computer-assisted language learning
Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee
Knowl. Based Syst.4
2011 Iteratively constrained selection of word alignment links using knowledge and statistics
Hyeongjong Noh, Kyusong Lee, Gary Geunbae Lee
Knowl. Based Syst.4
2010 Affective effects of speech-enabled robots for language learning
abstract
This study introduces the speech and language technologies used in the educational assistant robots that we developed for language learning and exploring the affective effects of robot-assisted language learning (RALL). To achieve this purpose, a course was designed in which students have meaningful interaction with intelligent robots in an immersive environment. A total of 24 elementary students, ranging in age over 9-13, were enrolled in English lessons. Descriptive statistics and pre-test/post-test design were used to investigate the affective effects of RALL approach. The result showed that RALL is promoting and improving students' satisfaction, interest, confidence, and motivation at the significance level of 0.01.
Changgu Kim, Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee
SLT5