Sungzoon Cho

dblp:60/2556 · DBLP profile ↗
← Back
116ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 85 · 8 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 since 2021Databases, data management, data science and information retrieval · 11 · 3 since 2021Security and privacy · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 4
YearPublicationVenuePosition
2026 HarmonyRouting with traffic impact prediction based on graph neural network for large-scale semiconductor fabrication
Younkook Kang, Kwangyoung Im, Sungzoon Cho
Eng. Appl. Artif. Intell.4
2026 Predicting the difference between intent and perception with transformer using athlete self-report measures in sports training
Sanggi Lee, Sungzoon Cho
Eng. Appl. Artif. Intell.3
2026 Human-guided collective LLM intelligence for strategic planning via two-stage information retrieval
Sangyeop Kim 0001, Junguk Ha, Hangyeul Lee, Sohhyung Park, Sungzoon Cho
Inf. Process. Manag.5
2026 Sentiment lexicon integrated meta-training for sentiment analysis in data-scarce settings
Hyunjong Kim, Sungzoon Cho
J. Intell. Inf. Syst.2
2025 Upcycling Candidate Tokens of Large Language Models for Query Expansion
Sukmin Cho, Soyeong Jeong, Sangyeop Kim 0001, Sungzoon Cho
CIKM5
2025 Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
abstract
Recent advancements in Large Vision-Language Models (LVLMs) have significantly expanded their utility in tasks like image captioning and visual question answering. However, they still struggle with object hallucination, where models generate descriptions that inaccurately reflect the visual content by including nonexistent objects or misrepresenting existing ones. While previous methods, such as data augmentation and training-free approaches, strive to tackle this issue, they still encounter scalability challenges and often depend on additional external modules. In this work, we propose Ensemble Decoding (ED), a novel strategy that splits the input image into sub-images and combines logit distributions by assigning weights through the attention map. Furthermore, we introduce ED adaptive plausibility constraint to calibrate logit distribution and FastED, a variant designed for speed-critical applications. Extensive experiments across hallucination benchmarks demonstrate that our proposed method achieves state-of-the-art performance, validating the effectiveness of our approach.
Yeongjae Cho, Keonwoo Kim 0004, Taebaek Hwang, Sungzoon Cho
ICLR4
2025 InstructX2X: An Interpretable Local Editing Model for Counterfactual Medical Image Generation
Hyungi Min, Taeseung You, Hangyeul Lee, Yeongjae Cho, Sungzoon Cho
MICCAI (16)5
2025 KMI: A Dataset of Korean Motivational Interviewing Dialogues for Psychotherapy
abstract
Hyunjong Kim, Suyeon Lee, Yeongjae Cho, Eunseo Ryu, Yohan Jo, Suran Seong, Sungzoon Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hyunjong Kim, Suyeon Lee, Yeongjae Cho, Eunseo Ryu, Yohan Jo, Suran Seong, Sungzoon Cho
NAACL (Long Papers)7
2025 LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue
abstract
Sangyeop Kim, Sohhyung Park, Jaewon Jung, Jinseok Kim, Sungzoon Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sangyeop Kim 0001, Sohhyung Park, Jaewon Jung 0002, Sungzoon Cho
NAACL (Long Papers)5
2024 A Revisit to the Decoder for Camouflaged Object Detection
Seung Woo Ko 0002, Joopyo Hong, Suyoung Kim, Seungjai Bang, Sungzoon Cho, Nojun Kwak, Hyung-Sin Kim, Joonseok Lee
BMVC5
2024 Stock price momentum modeling using social media data
Hye Jin Lee, Soh Hyung Park, Sung Whan Jeon, Sungzoon Cho
Expert Syst. Appl.5
2024 Building knowledge graphs from technical documents using named entity recognition and edge weight updating neural network with triplet loss for entity normalization
abstract
Attempts to express information from various documents in graph form are rapidly increasing. The speed and volume in which these documents are being generated call for an automated process, based on machine learning techniques, for cost-effective and timely analysis. Past studies responded to such needs by building knowledge graphs or technology trees from the bibliographic information of documents, or by relying on text mining techniques in order to extract keywords and/or phrases. While these approaches provide an intuitive glance into the technological hotspots or the key features of the select field, there still is room for improvement, especially in terms of recognizing the same entities appearing in different forms so as to interconnect closely related technological concepts properly. In this paper, we propose to build a patent knowledge network using the United States Patent and Trademark Office (USPTO) patent filings for the semiconductor device sector by fine-tuning Huggingface’s named entity recognition (NER) model with our novel edge weight updating neural network. For the named entity normalization, we employ edge weight updating neural network with positive and negative candidates that are chosen by substring matching techniques. Experiment results show that our proposed approach performs very competitively against the conventional keyword extraction models frequently employed in patent analysis, especially for the named entity normalization (NEN) and document retrieval tasks. By grouping entities with named entity normalization model, the resulting knowledge graph achieves higher scores in retrieval tasks. We also show that our model is robust to the out-of-vocabulary problem by employing the fine-tuned BERT NER model.
Sung Hwan Jeon, Hye Jin Lee, Sungzoon Cho
Intell. Data Anal.4
2023 MEMTO: Memory-guided Transformer for Multivariate Time Series Anomaly Detection
abstract
Detecting anomalies in real-world multivariate time series data is challenging due to complex temporal dependencies and inter-variable correlations. Recently, reconstruction-based deep models have been widely used to solve the problem. However, these methods still suffer from an over-generalization issue and fail to deliver consistently high performance. To address this issue, we propose the MEMTO, a memory-guided Transformer using a reconstruction-based approach. It is designed to incorporate a novel memory module that can learn the degree to which each memory item should be updated in response to the input data. To stabilize the training procedure, we use a two-phase training paradigm which involves using K-means clustering for initializing memory items. Additionally, we introduce a bi-dimensional deviation-based detection criterion that calculates anomaly scores considering both input space and latent space. We evaluate our proposed method on five real-world datasets from diverse domains, and it achieves an average anomaly detection F1-score of 95.74%, significantly outperforming the previous state-of-the-art methods. We also conduct extensive experiments to empirically validate the effectiveness of our proposed model's key components.
Keonwoo Kim 0004, Jeonglyul Oh, Sungzoon Cho
NeurIPS4
2023 Incorporation of company-related factual knowledge into pre-trained language models for stock-related spam tweet filtering
Sungzoon Cho
Expert Syst. Appl.2
2023 Hot topic detection in central bankers' speeches
Hye Jin Lee, Sungzoon Cho
Expert Syst. Appl.3
2023 General-use unsupervised keyword extraction model for keyword analysis
Hunsik Shin, Hye Jin Lee, Sungzoon Cho
Expert Syst. Appl.3
2023 Edge Weight Updating Neural Network for Named Entity Normalization
Sung Hwan Jeon, Sungzoon Cho
Neural Process. Lett.2
2022 Domain-Slot Relationship Modeling Using a Pre-Trained Language Encoder for Multi-Domain Dialogue State Tracking
abstract
Dialogue state tracking for multi-domain dialogues is challenging because the model should be able to track dialogue states across multiple domains and slots. As using pre-trained language models is the de facto standard for natural language processing tasks, many recent studies use them to encode the dialogue context for predicting the dialogue states. Model architectures that have certain inductive biases for modeling the relationship among different domain-slot pairs are also emerging. Our work is based on these research approaches on multi-domain dialogue state tracking. We propose a model architecture that effectively models the relationship among domain-slot pairs using a pre-trained language encoder. Inspired by the way the special [CLS] token in BERT is used to aggregate the information of the whole sequence, we use multiple special tokens for each domain-slot pair that encodes information corresponding to its domain and slot. The special tokens are run together with the dialogue context through the pre-trained language encoder, which effectively models the relationship among different domain-slot pairs. Our experimental results on the datasets MultiWOZ-2.0 and MultiWOZ-2.1 show that our model outperforms other models with the same setting. Our ablation studies incorporate three main parts. The first component shows the effectiveness of our approach exploiting the relationship modeling. The second component compares the effect of using different pre-trained language encoders. The final component involves comparing different initialization methods that could be used for the special tokens. Qualitative analysis of the attention map of the pre-trained language encoder shows that our special tokens encode relevant information through the encoding process by attending to each other.
Jinwon An, Sungzoon Cho, Junseong Bang, Misuk Kim
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Active cluster annotation for wafer map pattern classification in semiconductor manufacturing
Jaewoong Shim, Seokho Kang 0001, Sungzoon Cho
Expert Syst. Appl.3
2021 Abuse detection in healthcare insurance with disease-treatment network embedding
Je Hyuk Lee, Sungzoon Cho
J. Biomed. Informatics2
2021 A scoring model to detect abusive medical institutions based on patient classification system: Diagnosis-related group and ambulatory patient group
Hunsik Shin, Je Hyuk Lee, Yongdae An, Sungzoon Cho
J. Biomed. Informatics4
2021 User segmentation via interpretable user representation and relative similarity-based segmentation method
Younghoon Lee, Sungzoon Cho
Multim. Syst.2
2020 Expected margin-based pattern selection for support vector machines
Dongil Kim, Seokho Kang 0001, Sungzoon Cho
Expert Syst. Appl.3
2020 Improving spherical k-means for document clustering: Fast initialization, sparse centroid projection, and efficient cluster labeling
Hyunjoong Kim, Han Kyul Kim, Sungzoon Cho
Expert Syst. Appl.3
2020 Abnormal Usage Sequence Detection for Identification of User Needs via Recurrent Neural Network Semantic Variational Autoencoder
abstract
In this paper, we propose an advanced method to detect abnormal usage patterns for identifying the fine-grained levels of user needs. Most previous studies investigated user need identification based on users textual reviews. Thus, they focused only on the explicit needs of the product levels, and not on the implied needs of the fine-grained levels. Although in a few recent studies the authors attempted to identify user needs based on abnormality detection, they considered only limited elements of the usage sequence, such as touching buttons, and did not consider the important elements, such as the dragging interaction and the pop-up and notification components. Thus, in this study, we considered all the elements of the usage sequence to identify abnormal usage sequences for recognizing user needs at the fine-grained level. Moreover, we utilized the recurrent neural network semantic variational autoencoder (RNN-SVAE) architecture, which is a state-of-the-art architecture for sentence embedding, to represent the usage sequences effectively. In detail, we calculate the vector representation of the entire usage sequence utilizing the RNN-SVAE architecture based on heterogeneous embedding to apply the abnormality detection method for determining abnormal sequences corresponding to user needs. The experimental results verify that our proposed method extracts meaningful abnormal usage patterns that previous approaches do not identify. Additionally, our proposed method shows a higher correlation of the coefficient score between the abnormality score and the importance score of the extracted sequences than do previous approaches.
Younghoon Lee, Sungzoon Cho
Int. J. Hum. Comput. Interact.2
2020 Design of Semantic-Based Colorization of Graphical User Interface Through Conditional Generative Adversarial Nets
abstract
There has recently been a significant movement toward aiding graphic design tasks based on artificial intelligence or machine learning. In addition, colorization plays an important role within the topic of GUI design. Previous studies regarding automatic colorization have focused on a consideration of the realistic aspects of an image without consideration of the design semantics or usability, which are critical aspects for a practical GUI design. We, therefore, propose an end-to-end network for a generative combination of color sets for a GUI design based on the design semantics, while utilizing thousands of actual GUI design datasets acquired from LG Electronics to train the network. By utilizing the GUI design dataset, our network effectively generates color sets for a GUI design by considering various design aspects, such as the usability factors. In detail, we concatenate the textual design concept, characteristics of the application, and usage frequency for the elements of the design semantics. We then construct a conditional generative adversarial net processing of the design semantics as a condition to generate suitable color sets and construct the GUI design based on these sets. The experiments indicate that our proposed method effectively generates color sets for a GUI design based on the design semantics. In addition, our proposed method shows a better score than other methods on a user test conducted to verify the practicality, perception, recognition, diversity, and esthetic features. Moreover, experimental results prove that users can effectively grasp the intended design concept of our generated GUI design with higher top-1, top-2, and top-3 levels of accuracy.
Younghoon Lee, Sungzoon Cho
Int. J. Hum. Comput. Interact.2
2020 A medical treatment based scoring model to detect abusive institutions
Je Hyuk Lee, Hunsik Shin, Sungzoon Cho
J. Biomed. Informatics3
2020 Extraction and prioritization of product attributes using an explainable neural network
Younghoon Lee, Jungmin Park, Sungzoon Cho
Pattern Anal. Appl.3
2019 Dynamic Vehicle Traffic Control Using Deep Reinforcement Learning in Automated Material Handling System
abstract
In automated material handling systems (AMHS), delivery time is an important issue directly associated with the production cost and the quality of the product. In this paper, we propose a dynamic routing strategy to shorten delivery time and delay. We set the target of control by analyzing traffic flows and selecting the region with the highest flow rate and congestion frequency. Then, we impose a routing cost in order to dynamically reflect the real-time changes of traffic states. Our deep reinforcement learning model consists of a Q-learning step and a recurrent neural network, through which traffic states and action values are predicted. Experiment results show that the proposed method decreases manufacturing costs while increasing productivity. Additionally, we find evidence the reinforcement learning structure proposed in this study can autonomously and dynamically adjust to the changes in traffic patterns.
Younkook Kang, Sungwon Lyu, Jeeyung Kim, Bongjoon Park, Sungzoon Cho
AAAI5
2019 Champion-challenger analysis for credit card fraud detection: Hybrid ensemble and deep learning
Eunji Kim 0001, Je Hyuk Lee, Hunsik Shin, Hoseong Yang, Sungzoon Cho, Seung-kwan Nam, Youngmi Song, Jeong-a Yoon, Jong-il Kim
Expert Syst. Appl.5
2019 Smartphone help contents re-organization considering user specification via conditional GAN
Younghoon Lee, Sungzoon Cho, Jinhae Choi
Int. J. Hum. Comput. Stud.2
2019 Extraction of Product Evaluation Factors with a Convolutional Neural Network and Transfer Learning
Younghoon Lee, Minki Chung, Sungzoon Cho, Jinhae Choi
Neural Process. Lett.3
2019 Document representation based on probabilistic word clustering in customer-voice classification
Younghoon Lee, Seokmin Song, Sungzoon Cho, Jinhae Choi
Pattern Anal. Appl.3
2019 App usage prediction for dual display device via two-phase sequence modeling
Younghoon Lee, Sungzoon Cho, Jinhae Choi
Pervasive Mob. Comput.2
2018 Knowledge extraction and visualization of digital design process
Jiwon Yang, Eunji Kim 0001, Minhoe Hur, Sungzoon Cho, Myungbin Han, Iksang Seo
Expert Syst. Appl.4
2018 Stock price prediction through sentiment analysis of corporate disclosures using distributed representation
abstract
Many researches have exploited textual data, such as news, online blogs, and financial reports, in order to predict stock price movements effectively. Previous studies formed the task as a classification problem predicting upward or downward movement of stock prices from text documents. Such an app roach, however, may be deemed inappropriate when combined with sentiment analysis. In financial documents, same words may convey different sentiments across different sectors; if documents from multiple sectors are learned simultaneously, performance can deteriorate. Therefore, we conducted sentiment analysis of 8-K financial reports of firms sector by sector. In particular, we also employed distributed representation for predicting stock price movements. Experiment results show that our approach improves prediction performance by 25.4% over the baseline model, and that the direction of post-announcement stock price movements shifts accordingly with the polarity of the sentiment of reports. Not only does our model improve predictability, but also provides visualizations, which may assist agents actively trading in the field with understanding the drivers for the observed stock movements. The two main aspects of our model, predictability and interpretability, will provide meaningful insights to help decision-makers in the industry with time-split trading decisions or data-driven detection of promising companies.
Misuk Kim, Eunjeong L. Park, Sungzoon Cho
Intell. Data Anal.3
2018 De-noising documents with a novelty detection method utilizing class vectors
abstract
The classification of customer-voice data is an important matter in real business since it is necessary for customer-voice data to be delivered to relevant departments and responsible individuals. Additionally, customer-voice data typically includes several novel words, such as typo’s, informal ter ms, or exceedingly general words to discriminate between categories of customer-voice data. Furthermore, noisy data often has a negative effect on the classification task. In this study, advanced novelty detection method is proposed to utilize class vector that possessed high cosine similarity with words to effectively discriminate between classes. The class vector is considered as the centroid or the mean of each word vector distribution as derived from the neural embedding model, and the novelty score of each word is calculated and novel words are effectively detected. Each novelty score is calculated by improvements of GMM and KMC methods utilizing a class vector. The experiments verify the propriety of the proposed method with qualitative observations, and the application of the proposed method with quantitative experiments verifies the representational effectiveness and classification performance of customer-voice data. The experiment results indicate that the performance of a classification of customer-voice data improved with the application of the newly proposed novelty detection method in this study.
Younghoon Lee, Sungzoon Cho, Jinhae Choi
Intell. Data Anal.2
2018 Applying convolution filter to matrix of word-clustering based document representation
Younghoon Lee, Jinbae Im, Sungzoon Cho, Jinhae Choi
Neurocomputing3
2018 Obtaining calibrated probability using ROC Binning
Meesun Sun, Sungzoon Cho
Pattern Anal. Appl.2
2017 Building industry network based on business text: Corporate disclosures and news
abstract
Industry classification has served as an important tool for sector analysis in capital market research. Most of the existing classification schemes that are commonly used in academia or industry research require human input in the process of its design and construction. Because compilation and maintenance of such schemes demand comprehensive domain knowledge, it is highly costly and exhaustively time-consuming to update them to correctly reflect the fast-changing trends of the market. In addition, due to its subjective nature, interpretability is limited. As a remedy to such shortcomings, this paper adds to our earlier work, Business Text Industry Classification (BTIC) and proposes a new classification scheme, Business Text Industry Network (BTIN), namely, that automatically produces industry groupings based on the textual information from the corporate disclosures and news articles. BTIN exploits the business section of the Form 10-Ks, in which firms describe their business identities in a rich context, as well as the contents of news articles, which represent market perception of the subject firm. We employ Doc2vec for document embedding and a novel news curation algorithm to represent the relationship among firms in a graph structure, and then apply Louvain method to detect communities ??? or, industries ??? from the network and categorize firms into BTIN groups. Evaluation results using market ratios show that BTIN performs quite competitively against SIC, GICS, and its predecessor, BTIC, in terms of clustering fitness, inter-industry heterogeneity, and intra-industry homogeneity.
Sung Whan Jeon, Hye Jin Lee, Sungzoon Cho
IEEE BigData3
2017 Bag-of-concepts: Comprehending document representation through clustering words in distributed representation
Han Kyul Kim, Hyunjoong Kim, Sungzoon Cho
Neurocomputing3
2017 Reliable prediction of anti-diabetic drug failure using a reject option
Seokho Kang 0001, Sungzoon Cho, Su-jin Rhee, Kyung-Sang Yu
Pattern Anal. Appl.2
2016 Automatic classification of securities using hierarchical clustering of the 10-Ks
abstract
Industry classification has been rigorously utilized in academic research and business analytics. The existing classification schemes, however, have been constructed and maintained manually by domain experts, which require exhaustive time and human effort while vulnerable to subjectivity. Hence, the existing classification systems do not properly reflect the fast-changing trends of the firms and the capital market. As a remedy to such shortcomings, this paper proposes a new classification scheme, Business Text Industry Classification (BTIC), namely, that automatically clusters securities based on the textual information from the corporate disclosures. BTIC exploits the business section of the Form 10-Ks, in which firms provide their self-identities in a rich context. We employ doc2vec for document embedding and apply Ward's hierarchical clustering method to categorize securities into BTIC groups. Evaluation results using 12 financial ratios commonly found in financial research show that BTIC performs just as good as SIC and GICS in terms of inter- and intra-industry homogeneity, especially for the higher level of clustering. Given that, we claim that BTIC outperforms SIC and GICS in four aspects: process automation, objectivity, clustering flexibility, and result interpretability.
Hoseong Yang, Hye Jin Lee, Sungzoon Cho, Eugene C. Snyder
IEEE BigData3
2016 Semi-supervised support vector regression based on self-training with label uncertainty: An application to virtual metrology in semiconductor manufacturing
Pilsung Kang 0001, Dongil Kim, Sungzoon Cho
Expert Syst. Appl.3
2016 Detecting financial misstatements with fraud intention using multi-class cost-sensitive learning
Yeonkook J. Kim, Bok Baik, Sungzoon Cho
Expert Syst. Appl.3
2016 Box-office forecasting based on sentiments of movie reviews and Independent subspace method
Minhoe Hur, Pilsung Kang 0001, Sungzoon Cho
Inf. Sci.3
2015 Multi-class classification via heterogeneous ensemble of one-class classifiers
Seokho Kang 0001, Sungzoon Cho, Pilsung Kang 0001
Eng. Appl. Artif. Intell.2
2015 An efficient and effective ensemble of support vector machines for anti-diabetic drug failure prediction
Seokho Kang 0001, Pilsung Kang 0001, Taehoon Ko, Sungzoon Cho, Su-jin Rhee, Kyung-Sang Yu
Expert Syst. Appl.4
2015 A novel multi-class classification algorithm based on one-class support vector machine
abstract
The existing multi-class classification algorithms based on support vector machine (SVM) generally decompose the original problem into smaller subproblems. However, the decomposition approach raises the problems of unreliable and unbalanced training of individual classifiers. The aim of this work is to alleviate these problems. In this paper, a novel multi-class classification algorithm based on one-class SVM is presented. The distinguishing point of our proposed algorithm is that the algorithm solves a one-class classification problem rather than decomposing the problem into several smaller subproblems. The proposed algorithm solves a multi-class classification problem by expanding the training input vector with the class label and constructing a one-class SVM with the expanded input vectors. In the test phase, test input vectors are assessed by fit to the decision boundary under the assumption of belonging to a class. We conducted experiments on several real-world benchmark datasets. Experimental results showed that the proposed algorithm outperforms other SVM-based multi-class classification algorithms.
Seokho Kang 0001, Sungzoon Cho
Intell. Data Anal.2
2015 Optimal construction of one-against-one classifier based on meta-learning
Seokho Kang 0001, Sungzoon Cho
Neurocomputing2
2015 Constructing a multi-class classifier using one-against-one approach with different binary classifiers
Seokho Kang 0001, Sungzoon Cho, Pilsung Kang 0001
Neurocomputing2
2015 Keystroke dynamics-based user authentication using long and free text strings from various input devices
Pilsung Kang 0001, Sungzoon Cho
Inf. Sci.2
2015 Improvement of virtual metrology performance by removing metrology noises in a training dataset
Dongil Kim, Pilsung Kang 0001, Seung-kyung Lee, Seokho Kang 0001, Seungyong Doh, Sungzoon Cho
Pattern Anal. Appl.6
2014 Approximating support vector machine with artificial neural network for fast prediction
Seokho Kang 0001, Sungzoon Cho
Expert Syst. Appl.2
2014 Knowledge discovery in inspection reports of marine structures
Seung-kyung Lee, Bongseok Kim, Minhoe Huh, Jooseoung Park, Seokho Kang 0001, Sungzoon Cho, Dongha Lee 0009, Daehyung Lee
Expert Syst. Appl.6
2014 Data based segmentation and summarization for sensor data in semiconductor manufacturing
Eunjeong L. Park, Jooseoung Park, Jiwon Yang, Sungzoon Cho, Young-Hak Lee, Hae-Sang Park
Expert Syst. Appl.4
2014 Probabilistic local reconstruction for k-NN regression and its application to virtual metrology in semiconductor manufacturing
Seung-kyung Lee, Pilsung Kang 0001, Sungzoon Cho
Neurocomputing3
2014 Evaluating the reliability level of virtual metrology results for flexible process control: a novelty detection-based approach
Pilsung Kang 0001, Dongil Kim, Sungzoon Cho
Pattern Anal. Appl.3
2013 Mining transportation logs for understanding the after-assembly block manufacturing process in the shipbuilding industry
Seung-kyung Lee, Bongseok Kim, Minhoe Huh, Sungzoon Cho, Sungkyu Park, Daehyung Lee
Expert Syst. Appl.4
2012 Improved response modeling based on clustering, under-sampling, and ensemble
Pilsung Kang 0001, Sungzoon Cho, Douglas L. MacLachlan
Expert Syst. Appl.2
2012 Pattern selection for support vector regression based response modeling
Dongil Kim, Sungzoon Cho
Expert Syst. Appl.2
2012 Machine learning-based novelty detection for faulty wafer detection in semiconductor manufacturing
Dongil Kim, Pilsung Kang 0001, Sungzoon Cho, Hyoungjoo Lee, Seungyong Doh
Expert Syst. Appl.3
2012 Support vector class description (SVCD): Classification in kernel space
abstract
We proposed a kernel-based binary classification algorithm, named support vector class description (SVCD), which is an extended version of support vector domain description (SVDD) for one-class classification. SVCD constructs two compact hyperspheres
Pilsung Kang 0001, Sungzoon Cho
Intell. Data Anal.2
2011 Virtual metrology for run-to-run control in semiconductor manufacturing
Pilsung Kang 0001, Dongil Kim, Hyoungjoo Lee, Seungyong Doh, Sungzoon Cho
Expert Syst. Appl.5
2009 K-Means Clustering Seeds Initialization Based on Centrality, Sparsity, and Isotropy
Pilsung Kang 0001, Sungzoon Cho
IDEAL2
2009 Keystroke dynamics-based authentication for mobile devices
Seongseob Hwang, Sungzoon Cho, Sunghoon Park
Comput. Secur.2
2009 Improving authentication accuracy using artificial rhythms and cues for keystroke dynamics-based authentication
Seongseob Hwang, Hyoungjoo Lee, Sungzoon Cho
Expert Syst. Appl.3
2009 A virtual metrology system for semiconductor manufacturing
Pilsung Kang 0001, Hyoungjoo Lee, Sungzoon Cho, Dongil Kim, Chan-Kyoo Park, Seungyong Doh
Expert Syst. Appl.3
2009 A hybrid novelty score and its use in keystroke dynamics-based user authentication
Pilsung Kang 0001, Sungzoon Cho
Pattern Recognit.2
2008 Bootstrap Based Pattern Selection for Support Vector Regression
Dongil Kim, Sungzoon Cho
PAKDD2
2008 Supporting diagnosis of attention-deficit hyperactive disorder with novelty detection
Hyoungjoo Lee, Sungzoon Cho, Min Sup Shin
Artif. Intell. Medicine2
2008 Improvement of keystroke data quality through artificial rhythms and cues
Pilsung Kang 0001, Sunghoon Park, Seongseob Hwang, Hyoungjoo Lee, Sungzoon Cho
Comput. Secur.5
2008 Response modeling with support vector regression
Dongil Kim, Hyoungjoo Lee, Sungzoon Cho
Expert Syst. Appl.3
2008 Locally linear reconstruction for instance-based learning
Pilsung Kang 0001, Sungzoon Cho
Pattern Recognit.2
2007 Clustering-Based Reference Set Reduction for k-Nearest Neighbor
Seongseob Hwang, Sungzoon Cho
ISNN (2)2
2007 Retraining a keystroke dynamics-based authenticator with impostor patterns
Hyoungjoo Lee, Sungzoon Cho
Comput. Secur.2
2007 Focusing on non-respondents: Response modeling with novelty detectors
Hyoungjoo Lee, Sungzoon Cho
Expert Syst. Appl.2
2007 Neighborhood Property-Based Pattern Selection for Support Vector Machines
abstract
The support vector machine (SVM) has been spotlighted in the machine learning community because of its theoretical soundness and practical performance. When applied to a large data set, however, it requires a large memory and a long time for training. To cope with the practical difficulty, we propose a pattern selection algorithm based on neighborhood properties. The idea is to select only the patterns that are likely to be located near the decision boundary. Those patterns are expected to be more informative than the randomly selected patterns. The experimental results provide promising evidence that it is possible to successfully employ the proposed algorithm ahead of SVM training.
Hyunjung Shin, Sungzoon Cho
Neural Comput.2
2006 EUS SVMs: Ensemble of Under-Sampled SVMs for Data Imbalance Problems
Pilsung Kang 0001, Sungzoon Cho
ICONIP (1)2
2006 The Novelty Detection Approach for Different Degrees of Class Imbalance
Hyoungjoo Lee, Sungzoon Cho
ICONIP (2)2
2006 Prototype based outlier detection
abstract
Outliers refer to "minority" data that are different from most other data. They usually disturb data mining process. But, sometimes they provide valuable information. Thus, it is important to identify outliers in a given data set. In this paper, we propose a novel approach which scores "outlierness" based on the distance from majority data. First, prototype data are identified. Second, those prototypes that are distant from others are eliminated. Finally, the outlierness of each data point is computed as the distance from the remaining prototypes. Experiments involving various data sets show that the proposed approach performs well in terms of accuracy, robustness and versatility.
Seungtaek Kim, Sungzoon Cho
IJCNN2
2006 Pattern Selection for Support Vector Regression based on Sparseness and Variability
abstract
Support Vector Machine has been well received in machine learning community with its theoretical as well as practical value. However, since its training time complexity is cubic, its use is limited in data mining involving problems with a huge pattern set with a cubic time complexity of its training time. In this paper, we propose a pattern selection method for Support Vector Regression (SVR), using notions of sparseness, variability and uniqueness. Two versions of algorithms, deterministic and stochastic, are presented, which are then applied to an artificial data set and two well known real world data sets. Preliminary results justify further investigation. The proposed method should work well with non-SVM function approximators such as neural networks.
Jiyoung Sun, Sungzoon Cho
IJCNN2
2006 epsilon-Tube Based Pattern Selection for Support Vector Machines
Dongil Kim, Sungzoon Cho
PAKDD2
2006 Response modeling with support vector machines
Hyunjung Shin, Sungzoon Cho
Expert Syst. Appl.2
2006 Constructing response model using ensemble based on feature subset selection
Enzhe Yu, Sungzoon Cho
Expert Syst. Appl.2
2006 Application of LVQ to novelty detection using outlier training data
Hyoungjoo Lee, Sungzoon Cho
Pattern Recognit. Lett.2
2005 SOM-Based Novelty Detection Using Novel Data
Hyoungjoo Lee, Sungzoon Cho
IDEAL2
2005 Invariance of neighborhood relation under input space to feature space mapping
Hyunjung Shin, Sungzoon Cho
Pattern Recognit. Lett.2
2004 Combining Gaussian Mixture Models
Hyoungjoo Lee, Sungzoon Cho
IDEAL2
2004 Keystroke dynamics identity verification - its problems and practical solutions
Enzhe Yu, Sungzoon Cho
Comput. Secur.2
2003 Bankruptcy prediction for credit risk using an auto-associative neural network in Korean firms
abstract
Empirical bankruptcy prediction models have been proposed and widely used in the last decades or so. Historic solvent and default firm data are collected and labeled appropriately. Statistical and neural network models are then "trained" to fit these data. A major problem is the imbalance of data, i.e. much more solvent data than default data. We propose an auto-associative neural network (AANN) that learns the identity mapping of input. By training the network with only solvent data, we built a bankruptcy predictor with better accuracies than conventional 2-class neural network based predictor.
Jinwoo Baek, Sungzoon Cho
CIFEr2
2003 Trend detection using auto-associative neural networks: Intraday KOSPI 200 futures
abstract
This paper reports the results of a new neural network based trend detector. An auto-associative neural network was trained with the "trend" data obtained from the intra-day KOSPI 200 future price. It was then used to predict a trend. Simple investment strategies based on the detector achieved a one-year return of 31.2 points with no leverage.
Jun-Myung Lee, Sungzoon Cho, Jinwoo Baek
CIFEr2
2003 Fast Pattern Selection Algorithm for Support Vector Classifiers: Time Complexity Analysis
Hyunjung Shin, Sungzoon Cho
IDEAL2
2003 Novelty Detection Approach for Keystroke Dynamics Identity Verification
Enzhe Yu, Sungzoon Cho
IDEAL2
2003 How many neighbors to consider in pattern pre-selection for support vector classifiers?
abstract
Training support vector classifiers (SVC) requires large memory and long cpu time when the pattern set is large. To alleviate the computational burden in SVC training, we previously proposed a preprocessing algorithm which selects only the patterns in the overlap region around the decision boundary, based on neighborhood properties. The k-nearest neighbors' class label entropy for each pattern was used to estimate the pattern's proximity to the decision boundary. The value of parameter k is critical, yet has been determined by a rather ad-hoc fashion. We propose in this paper a systematic procedure to determine k and show its effectiveness through experiments.
Hyunjung Shin, Sungzoon Cho
IJCNN2
2003 GA-SVM wrapper approach for feature subset selection in keystroke dynamics identity verification
abstract
Password is the most widely used identity verification method in computer security domain. However, due to its simplicity, it is vulnerable to imposter attacks. Keystroke dynamics adds a shield to password. Password typing patterns or timing vectors of a user are measured and used to train a novelty detector model. However, without manual pre-processing to remove noises and outliers resulting from typing inconsistencies, a poor detection accuracy results. Thus, in this paper, we propose an automatic feature subset selection process that can automatically selects a relevant subset of features and ignores the rest, thus producing a better accuracy. Genetic algorithm is employed to implement a randomized search and SVM, an excellent novelty detector with fast learning speed, is employed as a base learner. Preliminary experiments show a promising result.
Enzhe Yu, Sungzoon Cho
IJCNN2
2003 Fast Pattern Selection for Support Vector Classifiers
Hyunjung Shin, Sungzoon Cho
PAKDD2
2002 An Up-Trend Detection Using an Auto-Associative Neural Network: KOSPI 200 Futures
Jinwoo Baek, Sungzoon Cho
IDEAL2
2002 Pattern Selection for Support Vector Classifiers
Hyunjung Shin, Sungzoon Cho
IDEAL2
2002 Observational Learning Algorithm for an Ensemble of Neural Networks
Min Jang, Sungzoon Cho
Pattern Anal. Appl.2
2001 Neural Network Based Automatic Diagnosis of Children with Brain Dysfunction
abstract
This paper proposes the use of multilayer perceptron for brain dysfunction diagnosis. The performance of MLP was better than that of Discriminant Analysis and Decision Tree classifiers, with an 85% accuracy rate in an experimental test involving 332 subjects. In addition, the neural network employing Bayesian learning was able to identify the most important input variable. These two results demonstrate that the neural network can be effectively used in the diagnosis of children with brain dysfunction.
Sungzoon Cho, Min Sup Shin
Int. J. Neural Syst.1
2001 Smoothed Bagging with Kernel Bandwidth Selectors
Shinjae Lee, Sungzoon Cho
Neural Process. Lett.2
2000 "Left Shoulder" Detection in Korea Composite Stock Price Index Using an Auto-Associative Neural Network
Jinwoo Baek, Sungzoon Cho
IDEAL2
2000 Observational Learning with Modular Networks
Hyunjung Shin, Hyoungjoo Lee, Sungzoon Cho
IDEAL3
1999 Characteristics of auto-associative MLP as a novelty detector
abstract
In novelty detection, one tries to discriminate abnormal patterns from normal patterns. As a two class pattern classification problem, novelty detection is quite difficult since in practice only normal patterns are available for training. Novel or abnormal patterns are very few or not available at all. Recently, an auto-associative MLP (AaMLP) has been shown to give a good performance. In this paper, we analyze the output characteristics of trained AaMLPs and show that the AaMLP is indeed a reliable solution for novelty detection. In particular, we prove why nonlinearity in the hidden layer is necessary for novelty detection.
Byungho Hwang, Sungzoon Cho
IJCNN2
1999 Ensemble learning using observational learning theory
abstract
In this paper, we propose an ensemble learning algorithm, which we call the observational learning algorithm (OLA), motivated from the observational learning theory suggested by Bandura (1971). According to the theory, in a group of children who have different and insufficient knowledge for a task, each child can lean how to do the task by observing other children. In the OLA, a neural network ensemble is regarded as a group of children. Each network is trained with the virtual data that are generated from observing other networks as well as the bootstrapping data from the original data set. The virtual data function as both temporal hints having the auxiliary information about the target function and a regularization penalty for making networks smooth. From numerical experiments involving both regression and classification problems, the OLA is shown to give better generalization performance than simple committee and bagging approaches when insufficient and noisy data are given.
Min Jang, Sungzoon Cho
IJCNN2
1999 Constructing belief networks from realistic data
abstract
In this paper, we report our investigation on the behavior of some Bayesian belief network learning methods, when applied to realistic case data that do not accurately reflect the probability distribution of the underlying domain. The investigation focuses on two types of such data of practical interest, namely data excluding null or normal cases and data contaminated by noise, and their effects on the performance of two learning methods: the K2 algorithm, a representative method of Bayesian approach, and the extended Hebbian learning (EHL) method, representing those of neural network approach. Both analytical and experimental results show that when null cases are removed from the case database, the EHL method is able to accurately construct the underlying causal structure; whereas K2 algorithm may fail to properly handle root nodes. On the other hand, the EHL method is noise sensitive while K2 has certain inherent noise-resistance capability. Suggestions to extend both K2 and EHL to better cope with these problems are also made, and their effectiveness tested through a systematic computer simulation experiment. ©1999 John Wiley & Sons, Inc.
Yun Peng 0001, Zonglin Zhou, Sungzoon Cho
Int. J. Intell. Syst.3
1998 Rock Porosity Prediction Using Multilayer Perceptrons
Min Jang, Sungzoon Cho, Patrick M. Wong, Lance Chun Che Fung, Kevin Kok Wai Wong
ICONIP2
1997 Does Rotation of Neuronal Population Vectors Equal Mental Rotation?
abstract
It has been claimed that rotation of neuronal population vectors in the primary motor cortex during tasks involving transformation of movement directions provides direct evidence that a mental rotation algorithm is being used. An alternative explanation offered here, the 'asynchronous superposition hypothesis', asserts that the apparent rotation can result from the summation of two stationary vectors, one in the stimulus direction and one in the target movement direction, which vary in length over time. Computer simulations demonstrate the surprising result that the asynchronous superposition hypothesis can produce activation of motor corex cells with intermediate preferred directions, something previously assumed to confirm a mental rotation algorithm. Simulations also demonstrate that the asynchronous superposition hypothesis accounts for some aspects of the data more naturally than does an explanation based on a mental rotation algorithm, and suggest experimental tests that can distinguish between these two hypotheses.
Carol S. Whitney, James A. Reggia, Sungzoon Cho
Connect. Sci.3
1997 Virtual sample generation using a population of networks
Sungzoon Cho, Min Jang, Sungjune Chang
Neural Process. Lett.1
1997 Reliable roll force prediction in cold mill using multiple neural networks
abstract
The cold rolling mill process in steel works uses stands of rolls to flatten a strip to a desired thickness. The accurate prediction of roll force is essential for product quality. Currently, a suboptimal mathematical model is used. We trained two multilayer perceptrons, one to directly predict the roll force and the other to compute a corrective coefficient to be multiplied to the prediction made by the mathematical model. Both networks were shown to improve the accuracy by 30-50%. Combining the two networks and the mathematical model results in systems with an improved reliability.
Sungzoon Cho, Yongjung Cho, Sungchul Yoon
IEEE Trans. Neural Networks1
1996 Effects of Varying Parameters on Properties of Self-Organizing Feature Maps
Sungzoon Cho, Min Jang, James A. Reggia
Neural Process. Lett.1
1994 Map Formation in Proprioceptive Cortex
abstract
Current understanding of feature maps in proprioceptive cortex is quite limited. To complement experimental studies, we developed a computational model of map formation in proprioceptive cortex. Muscle length and tension from six muscle groups controlling the position of a model arm in three-dimensional space served as input to the simulated cortex. The resultant feature map consisted of regularly spaced clusters of cortical columns representing individual muscle lengths and tensions. Cortical units became tuned to plausible combinations of tension and length, and multiple representations of each muscle group were present. The map was organized such that compact regions within which all muscle group lengths and tensions are represented could be identified. Most striking was the observation that, although not explicitly present in the input, the cortical map developed a representation of the three-dimensional space in which the arm moved. These findings represent testable predictions about proprioceptive cortex, and may also help clarify some organizational issues concerning primary motor cortex.
Sungzoon Cho, James A. Reggia
Int. J. Neural Syst.1
1993 Multiple disorder diagnosis with adaptive competitive neural networks
Sungzoon Cho, James A. Reggia
Artif. Intell. Medicine1
1993 Learning Competition and Cooperation
abstract
Competitive activation mechanisms introduce competitive or inhibitory interactions between units through functional mechanisms instead of inhibitory connections. A unit receives input from another unit proportional to its own activation as well as to that of the sending unit and the connection strength between the two. This, plus the finite output from any unit, induces competition among units that receive activation from the same unit. Here we present a backpropagation learning rule for use with competitive activation mechanisms and show empirically how this learning rule successfully trains networks to perform an exclusive-OR task and a diagnosis task. In particular, networks trained by this learning rule are found to outperform standard backpropagation networks with novel patterns in the diagnosis problem. The ability of competitive networks to bring about context-sensitive competition and cooperation among a set of units proved to be crucial in diagnosing multiple disorders.
Sungzoon Cho, James A. Reggia
Neural Comput.1
1989 Improvement of kittler and illingworth's minimum error thresholding
Sungzoon Cho, Robert M. Haralick, Seungku Yi
Pattern Recognit.1