Heeyoul Choi

dblp:78/2514 · DBLP profile ↗
← Back
43ranked-venue papers
20as first author
11since 2021 · last 2026
0000-0002-0855-8725ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 16 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 3 since 2021Computer networks · 5 · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Joint Models for Sentence Segmentation and Named Entity Recognition in Literary Sinitic Text
abstract
It is challenging to understand Literary Sinitic text from the Joseon dynasty, since there is a lack of explicit word separators, which creates significant semantic ambiguity. To address this, both sentence segmentation and named entity recognition (NER) are essential. We propose a Transformer-based analyzer that performs these two tasks simultaneously. Trained on a labeled corpus from the Seungjeongwon Ilgi, our model effectively segments sentences and identifies named entities, thereby significantly improving the understanding of sentence structure and overall context.
DongNyeong Heo, Yunhee Kang, Chul Heo, Heeyoul Choi, Kyounghun Jung
J. Web Eng.4
2026 Generalized Probabilistic Attention Mechanism in Transformers
DongNyeong Heo, Heeyoul Choi
Mach. Learn.2
2025 Pre-trained Models for Bytecode Instructions
abstract
Recent advancements in pre-trained models have rapidly expanded their applicability to various software engineering challenges. Despite this progress, current research predominantly focuses on source code and natural language processing, largely overlooking Java bytecode. Java bytecode, with its well-defined structure and high availability, presents a promising yet under-explored domain for leveraging pre-trained models. Its inherent properties, such as platform independence and optimized performance, make Java bytecode an ideal candidate for developing robust and efficient software engineering solutions. Addressing this gap could unlock new opportunities for enhancing automated program analysis, bug detection, and code generation tasks. In this study, we propose byteT5 and byteBERT, which are pre-trained models with hexadecimal bytecode. To build our models, we developed a bytecode tokenizer, ByteTok, to generate hexadecimal input representations for our pre-trained models. We conduct an empirical study comparing our models and GPT-40. The results indicate that byteT5 and byteBERT outperform GPT-40 in the span masking task. We anticipate these findings will pave the way for novel approaches to addressing various software engineering challenges, particularly live patching.
Heeyoul Choi, Jaechang Nam
ICST5
2025 Shared Latent Space by Both Languages for Non-Autoregressive Neural Machine Translation
abstract
Non-autoregressive neural machine translation (NAT) offers substantial translation speed up compared to autoregressive neural machine translation (AT) at the cost of translation quality. Latent variable modeling has emerged as a promising approach to bridge this quality gap, particularly for addressing the chronic multimodality problem in NAT. In the previous works that used latent variable modeling, they added an auxiliary model to estimate the posterior distribution of the latent variable conditioned on the source and target sentences. However, it causes several disadvantages, such as redundant information extraction in the latent variable, increasing the number of parameters, and a tendency to ignore some information from the inputs. In this paper, we propose a novel latent variable modeling that integrates a dual reconstruction perspective and an advanced hierarchical latent modeling with a shared intermediate latent space across languages. This latent variable modeling hypothetically alleviates or prevents the above disadvantages. In our experiment results, we present comprehensive demonstrations that our proposed approach infers superior latent variables which lead better translation quality. Finally, in the benchmark translation tasks, such as WMT, we demonstrate that our proposed method significantly improves translation quality compared to previous strong NAT baselines.
DongNyeong Heo, Heeyoul Choi
IJCNN2
2025 Harnessing EHRs for Diffusion-Based Anomaly Detection on Chest X-Rays
Harim Kim, Yuhan Wang 0001, Minkyu Ahn, Heeyoul Choi, Yuyin Zhou, Charmgil Hong
MICCAI (3)4
2025 Diverse Target Representations for Language Models with Word Difference Representations
DongNyeong Heo, Daniela N. Rim, Heeyoul Choi
PACLIC3
2024 End-To-End Training of Back-Translation Framework with Categorical Reparameterization Trick
DongNyeong Heo, Heeyoul Choi
ICANN (7)2
2024 Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
Tsun-An Hsieh, Heeyoul Choi
INTERSPEECH2
2022 Reinforcement Learning of Graph Neural Networks for Service Function Chaining in Computer Network Management
abstract
In management of computer network systems, a service function chaining (SFC) module plays a vital role in generating an efficient network traffic path that connects virtualized network functions (VNF) on network topology to serve a user request. The SFC module needs to generate a complete path quickly even in various network situations, including dynamic VNF resources, various types of requests, and various network topologies to provide the best quality of service. The previous supervised learning method demonstrated that graph neural networks (GNN) could represent network features for the SFC task. However, the supervised learning method works properly only in the network situation in which the model was trained with labels. Due to the limitation, it showed poor performance on new unseen network situations. In this paper, we apply a reinforcement learning algorithm to train GNN based models in various network situations even without label information. In the experiments, compared to the previous supervised learning method, the proposed methods demonstrate remarkable generalization effects, showing that the proposed methods could work successfully on unseen network situations without re-designing and re-training.
DongNyeong Heo, Doyoung Lee, Heegon Kim, Heeyoul Choi
APNOMS5
2022 Updating VNF deployment with Scaling Actions using Reinforcement Algorithms
abstract
Softwarization of the internet network is promising for network service providers (NSPs) to satisfy various types of user requests while dynamically operating the networking system. However, it is challenging to provide optimal service quality as the complexity of the network increases. In particular, deployment of the VNF instances is a critical issue in providing better QoS while maintaining resources optimally. For the management of the VNF deployment task, dynamic programming algorithms are only feasible in small networks or rely on heuristics. In this paper, we propose a VNF deployment method based on reinforcement learning (RL) that can effectively satisfy the QoS with minimized resource consumption in the network. Our approach makes adjustment decisions, which are scale-in/keep/out for target nodes and target VNF types given ILP-based deployments. We formulate this VNF deployment task as RL by setting the reward as QoS and resources. Moreover, we propose a model architecture for our RL agent's policy, based on Graph Neural Network. In the experiment, our approach optimizes VNF deployment with improved QoS while keeping the similar or slightly less amount of resources, compared to ILP-based deployment.
Namjin Seo, DongNyeong Heo, Jibum Hong, Heegon Kim, Jae-Hyoung Yoo, James Won-Ki Hong, Heeyoul Choi
APNOMS7
2022 Partitioning Image Representation in Contrastive Learning
abstract
In contrastive learning in the image domain, the anchor and positive samples are forced to have as close representations as possible. However, forcing the two samples to have the same representation could be misleading because the data augmentation techniques make the two samples different. In this paper, we introduce a new representation, partitioned representation, which can learn both common and unique features of the anchor and positive samples in contrastive learning. The partitioned representation consists of two parts: content part and style part. The content part represents common features of the class, and the style part represents own features of each sample, which can lead to the representation of the data augmentation method. We can achieve the partitioned representation simply by decomposing a loss function of contrastive learning into two terms on the two separate representations, respectively. To evaluate our representation with two parts, we take two framework models: Variational AutoEncoder (VAE) and Bootstrap Your Own Latent (BYOL), to show the content and style’s separability and confirm the generalization ability in classification, respectively. Based on the experiments, we show that our approach can separate two types of information in the VAE framework and outperforms the conventional BYOL in the classification and a few-shot learning task as downstream tasks.
Hyunsub Lee, Heeyoul Choi
ICPR2
2020 Graph Neural Network based Service Function Chaining for Automatic Network Control
abstract
Software-defined networking (SDN) and the network function virtualization (NFV) led to great developments in software based control technology by decreasing expenditures. Service function chaining (SFC) is an important technology to find efficient paths in network servers to process all of the requested virtualized network functions (VNF). However, SFC is challenging since it has to maintain high Quality of Service (QoS) even for complicated situations. Although some works have been conducted for such tasks with high-level intelligent models like deep neural networks (DNNs), those approaches are not efficient in utilizing the topology information of networks and cannot be applied to networks with dynamically changing topology since their models assume that the topology is fixed. In this paper, we propose a new neural network architecture for SFC, which is based on graph neural network (GNN) considering the graph-structured properties of network topology. The proposed SFC model consists of an encoder and a decoder, where the encoder finds the representation of the network topology, and then the decoder estimates probabilities of neighborhood nodes and their probabilities to process a VNF. In the experiments, our proposed architecture outperformed previous performances of DNN based baseline model. Moreover, the GNN based model can be applied to a new network topology without re-designing and re-training.
DongNyeong Heo, Stanislav Lange, Heegon Kim, Heeyoul Choi
APNOMS4
2020 Graph Neural Network-based Virtual Network Function Management
abstract
Software-Defined Networking (SDN) and Network Function Virtualization (NFV) help reduce OPEX and CAPEX as well as increase network flexibility and agility. But at the same time, operators have to cope with the increased complexity of managing virtual networks and machines, which are more dynamic and heterogeneous than before. Since this complexity is paired with strict time requirements for making management decisions, traditional mechanisms that rely on, e.g., Integer Linear Programming (ILP) models are no longer feasible. Machine learning has emerged as a possible solution to address network management problems to get near-optimal solutions in a short time. In this paper, we propose a Graph Neural Network (GNN) based algorithm to manage VNFs. The proposed model solves the complex VNF management problem in a short time and gets near-optimal solutions.
Heegon Kim, Stanislav Lange, Doyoung Lee, DongNyeong Heo, Heeyoul Choi, Jae-Hyoung Yoo, James Won-Ki Hong
APNOMS6
2020 Graph Neural Network-based Virtual Network Function Deployment Prediction
abstract
Software-Defined Networking (SDN) and Network Function Virtualization (NFV) help reduce OPEX and CAPEX as well as increase network flexibility and agility. But at the same time, operators have to cope with the increased complexity of managing virtual networks and machines, which are more dynamic and heterogeneous than before. Since this complexity is paired with strict time requirements for making management decisions, traditional mechanisms that rely on, e.g., Integer Linear Programming (ILP) models are no longer feasible. Machine learning has emerged as a possible solution to address network management problems to get near-optimal solutions in a short time. In this paper, we propose a Graph Neural Network (GNN) based algorithm to manage Virtual Network Functions (VNFs). The proposed model solves the complex VNF management prob-lem in a short time and gets near-optimal solutions.
Heegon Kim, DongNyeong Heo, Stanislav Lange, Heeyoul Choi, Jae-Hyoung Yoo, James Won-Ki Hong
CNSM5
2020 Understanding dropout as an optimization trick
Sangchul Hahn, Heeyoul Choi
Neurocomputing2
2019 Machine Learning-based Prediction of VNF Deployment Decisions in Dynamic Networks
abstract
In addition to providing network operators with benefits in terms of flexibility and cost efficiency, softwarization paradigms like SDN and NFV are key enablers for the concept of Service Function Chaining (SFC). The corresponding networks need to support a wide range of services and applications with highly dynamic temporal profiles and heterogeneous demands. Hence, efficient management and operation of such networks requires a high degree of automation that is paired with fast and proactive decisions in order to cope with these phenomena. In particular, determining the optimal number of VNF instances that is required for accommodating current and upcoming demands is a crucial task that also affects subsequent management decisions. To enable fast and proactive decisions in this context, we propose a machine learning-based approach that uses recent monitoring data to predict whether to adapt the current number of VNF instances of a given type. Furthermore, we present a work flow for generating labeled training data that reflects temporal dynamics and heterogeneous demands of real world networks. In addition to demonstrating the feasibility of the approach in a case study, we provide guidelines regarding the choice of monitoring data that should be collected for reliable prediction as well as the amount of data that is required to train such a predictor.
Stanislav Lange, Heegon Kim, Seyeon Jeong, Heeyoul Choi, Jae-Hyoung Yoo, James Won-Ki Hong
APNOMS4
2019 Predicting VNF Deployment Decisions under Dynamically Changing Network Conditions
abstract
In addition to providing network operators with benefits in terms of flexibility and cost efficiency, softwarization paradigms like SDN and NFV are key enablers for the concept of Service Function Chaining (SFC). The corresponding networks need to support a wide range of services and applications with highly heterogeneous requirements that change dynamically during the network's lifetime. Hence, efficient management and operation of such networks requires a high degree of automation that is paired with fast and proactive decisions in order to cope with these phenomena. In particular, determining the optimal number of VNF instances that is required for accommodating current and upcoming demands is a crucial task that also affects subsequent management decisions. To enable fast and proactive decisions in this context, we propose a machine learning-based approach that uses recent monitoring data to predict whether to adapt the current number of VNF instances of a given type. Furthermore, we present a methodology for generating labeled training data that reflects temporal dynamics and heterogeneous demands of real world networks. We demonstrate the feasibility of the approach using two different network topologies that represent WAN and mobile edge computing use cases, respectively. Additionally, we investigate how well the models generalize among networks and provide guidelines regarding the prediction horizon, i.e., how far ahead predictions can be performed in a reliable manner.
Stanislav Lange, Heegon Kim, Seyeon Jeong, Heeyoul Choi, Jae-Hyoung Yoo, James Won-Ki Hong
CNSM4
2019 Disentangling Latent Factors of Variational Auto-encoder with Whitening
Sangchul Hahn, Heeyoul Choi
ICANN (3)2
2019 A Deep Learning Approach to VNF Resource Prediction using Correlation between VNFs
abstract
Software-Defined Networking (SDN) and Network Function Virtualization (NFV) greatly facilitate network service management. Specifically, these new network paradigms help manage the network environment dynamically and cost-efficiently. Virtual Network Function (VNF) and Service Function Chaining (SFC) are important aspects of the NFV environment. In terms of NFV management, resource demand of VNFs can be predicted at a future time to handle Quality of Service (QoS) and resource allocation problems efficiently. Hence, researchers study and build a management system where machine-learning-based predictions of VNF information are used to handle auto-scaling, deployment and migration of VNFs. In addition, in recent studies, these systems have involved SFC to obtain useful information, not just a lone VNF. However, not many of studies explain clearly how chaining dependency among VNFs in a SFC can be used to predict future resource demand of a VNF. In this paper, we introduce VNF resource prediction machine learning model that maximizes the benefits of using SFC. Then, we compare several machine learning models and analyze how SFC data can help predict resource usage patterns of VNFs. We also show benefits of Attention model to improve prediction accuracy and convergence time through experiments.
Heegon Kim, Seyeon Jeong, Doyoung Lee, Heeyoul Choi, Jae-Hyoung Yoo, James Won-Ki Hong
NetSoft4
2019 Machine Learning-Based Method for Prediction of Virtual Network Function Resource Demands
abstract
Software-Defined Networking (SDN) and Network Function Virtualization (NFV) are paradigms that help administrators to manage dynamic networks. While SDN allows centralized network control, NFV provides flexible and scalable Virtual Network Functions (VNFs). These paradigms are also enablers for concepts such as Service Function Chaining (SFC) where chains are composed of several VNFs to provide a specific service. However, in order to maximize the benefits from the above-mentioned flexibility, new research questions need to be addressed, e.g., regarding effective management processes for dynamic networks. We proposed a novel learning model based on the flexibility of softwarization and abundant volume of monitoring data in NFV environments to predict VNF resource demands using SFC data. Our model is based on Context and Aspect Embedded Attentive Target Dependent Long Short Term Memory (CAT-LSTM) that consists of Target-Dependent LSTM (TD-LSTM), context embedding, aspect embedding, and attention. We developed this model to obtain high accuracy for the prediction of VNF resources such as the CPU. Our model uses two labeling systems: the qualitative resource state and the quantitative resource usage, both of which are used to evaluate its performance. This assists the administrator in understanding the network conditions, improves prediction performance, and provides practically useful information. Our learning model for predicting VNF resource demands can be utilized to solve essential SFC problems such as auto-scaling and optimal placement, which in turn prevent service interruption and provide high reliability.
Heegon Kim, Doyoung Lee, Seyeon Jeong, Heeyoul Choi, Jae-Hyoung Yoo, James Won-Ki Hong
NetSoft4
2019 Persistent hidden states and nonlinear transformation for long short-term memory
Heeyoul Choi
Neurocomputing1
2018 Fast Depthwise Separable Convolution for Embedded Systems
Byeongheon Yoo, Yongjun Choi, Heeyoul Choi
ICONIP (7)3
2018 Fine-grained attention mechanism for neural machine translation
Heeyoul Choi, Kyunghyun Cho, Yoshua Bengio
Neurocomputing1
2017 Context-dependent word representation for neural machine translation
Heeyoul Choi, Kyunghyun Cho, Yoshua Bengio
Comput. Speech Lang.1
2015 RNNDROP: A novel dropout for RNNS in ASR
abstract
Recently, recurrent neural networks (RNN) have achieved the state-of-the-art performance in several applications that deal with temporal data, e.g., speech recognition, handwriting recognition and machine translation. While the ability of handling long-term dependency in data is the key for the success of RNN, combating over-fitting in training the models is a critical issue for achieving the cutting-edge performance particularly when the depth and size of the network increase. To that end, there have been some attempts to apply the dropout, a popular regularization scheme for the feed-forward neural networks, to RNNs, but they do not perform as well as other regularization scheme such as weight noise injection. In this paper, we propose rnnDrop, a novel variant of the dropout tailored for RNNs. Unlike the existing methods where dropout is applied only to the non-recurrent connections, the proposed method applies dropout to the recurrent connections as well in such a way that RNNs generalize well. Our experiments show that rnnDrop is a better regularization method than others including weight noise injection. Namely, when deep bidirectional long short-term memory (LSTM) RNNs were trained with rnnDrop as acoustic models for phoneme and speech recognition, they significantly outperformed the current state-of-the-arts; we achieved the phoneme error rate of 16.29% on the TIMIT core test set for phoneme recognition and the word error rate of 5.53% on the Wall Street Journal (WSJ) dataset, dev93, for speech recognition, which are the best reported results on both of the datasets.
Taesup Moon, Heeyoul Choi, Hoshik Lee, Inchul Song
ASRU2
2014 Data visualization for asymmetric relations
Heeyoul Choi
Neurocomputing1
2014 Localization and regularization of normalized transfer entropy
Heeyoul Choi
Neurocomputing1
2013 Parameter Learning for Alpha Integration
abstract
In pattern recognition, data integration is an important issue, and when properly done, it can lead to improved performance. Also, data integration can be used to help model and understand multimodal processing in the brain. Amari proposed α-integration as a principled way of blending multiple positive measures (e.g., stochastic models in the form of probability distributions), enabling an optimal integration in the sense of minimizing the α-divergence. It also encompasses existing integration methods as its special case, for example, a weighted average and an exponential mixture. The parameter α determines integration characteristics, and the weight vector w assigns the degree of importance to each measure. In most work, however, α and w are given in advance rather than learned. In this letter, we present a parameter learning algorithm for learning α and ω from data when multiple integrated target values are available. Numerical experiments on synthetic as well as real-world data demonstrate the effectiveness of the proposed method.
Heeyoul Choi, Seungjin Choi 0001, Yoonsuck Choe
Neural Comput.1
2012 Dynamic learning for visual representation of asymmetric proximity
abstract
While many methods like multidimensional scaling (MDS) are exploited to represent and visualize symmetric distance matrices on 2-dimensional spaces, asymmetric proximity matrices such as an importing/exporting matrix from/to countries cannot be perfectly represented on metric spaces, since the methods assume a symmetric distance matrix. To overcome such an intrinsic limitation, in this paper, we propose a dynamic learning for metric representations of asymmetric proximity data to better understand the data. The proposed learning generates two representations (maps) with the column vectors (importing) and row vectors (exporting) of the matrix, respectively. To better present the patterns, we supplement the maps with two analysis tools: cluster analysis and flow analysis, which connect and compare the different patterns from the different maps. Experimental results using cola-brand-switching data and world-trade data confirm that the proposed learning method is useful to understand asymmetric proximity data.
Heeyoul Choi
SMC1
2011 From Data Streams to Information Flow: Information Exchange in Child-Parent Interaction
Heeyoul Choi, Chen Yu 0001, Linda B. Smith, Olaf Sporns
CogSci1
2010 Learning alpha-integration with partially-labeled data
abstract
Sensory data integration is an important task in human brain for multimodal processing as well as in machine learning for multisensor processing. α-integration was proposed by Amari as a principled way of blending multiple positive measures (e.g., stochastic models in the form of probability distributions), providing an optimal integration in the sense of minimizing the α-divergence. It also encompasses existing integration methods as its special case, e.g., weighted average and exponential mixture. In α-integration, the value of α determines the characteristics of the integration and the weight vector w assigns the degree of importance to each measure. In most of the existing work, however, α and w are given in advance rather than learned. In this paper we present two algorithms, for learning α and w from data when only a few integrated target values are available. Numerical experiments on synthetic as well as real-world data confirm the proposed method's effectiveness.
Heeyoul Choi, Seungjin Choi 0001, Anup Katake, Yoonsuck Choe
ICASSP1
2010 Alpha-integration of multiple evidence
abstract
In pattern recognition, data integration is a processing method to combine multiple sources so that the combined result can be more accurate than a single source. Evidence theory is one of the methods that have been successfully applied to the data integration task. Since Dempster-Shafer theory as the first evidence theory can be against our intuitive reasoning with some data sets, many researchers have proposed different rules for evidence theory. Among all these rules, the averaging rule is known to be better than others. On the other hand, a-integration was proposed by Amari as a principled way of blending multiple positive measures. It is a generalized averaging algorithm including arithmetic, geometric and harmonic means as its special case. In this paper, we generalize evidence theory with α-integration. Our experimental results show how our proposed methods work.
Heeyoul Choi, Anup Katake, Seungjin Choi 0001, Yoonsuck Choe
ICASSP1
2010 Manifold Alpha-Integration
Heeyoul Choi, Seungjin Choi 0001, Anup Katake, Yoonseop Kang, Yoonsuck Choe
PRICAI1
2009 Probabilistic Combination of Multiple Evidence
Heeyoul Choi, Anup Katake, Seungjin Choi 0001, Yoonseop Kang, Yoonsuck Choe
ICONIP (1)1
2009 Fast and accurate retinal vasculature tracing and kernel-Isomap-based feature selection
abstract
The blood vessels in the retina have a characteristic radiating pattern, while there exists a significant variation dependent on the individual and/or medical condition. Extracting the geometric properties of these blood vessels have several important applications, such as biometrics (for identification) and medical diagnosis. In this paper, we will focus on biometric applications. For this, we propose a fast and accurate algorithm for tracing the blood vessels, and compare several candidate summary features based on the tracing results. Existing tracing algorithms based on a detailed analysis of the image can be too slow to quickly process a large volume of retinal images in real time (e.g., at a security check point). In order to select good features that can be extracted from the traces, we used kernel Isomap to test the distance between different retinal images as projected onto their respective feature spaces. We tested the following feature set: (1) angle among branches, (2) the number of fiber based on distance, (3) distance between branches, and (4) inner product among branches. Our results indicate that features 3 and 4 are prime candidates for use in fast, realtime biometric tasks. We expect our method to lead to fast and accurate biometric systems based on retinal images.
Donghyeop Han, Heeyoul Choi, Choonseog Park, Yoonsuck Choe
IJCNN2
2008 Manifold Integration with Markov Random Walks
Heeyoul Choi, Seungjin Choi 0001, Yoonsuck Choe
AAAI1
2008 Sketch Recognition Based on Manifold Learning
Heeyoul Choi, Tracy Anne Hammond
AAAI1
2008 Kernel oriented discriminant analysis for speaker-independent phoneme spaces
abstract
Speaker independent feature extraction is a critical problem in speech recognition. Oriented principal component analysis (OPCA) is a potential solution that can find a subspace robust against noise of the data set. The objective of this paper is to find a speaker-independent subspace by generalizing OPCA in two steps: First, we find a nonlinear subspace with the help of a kernel trick, which we refer to as kernel OPCA. Second, we generalize OPCA to problems with more than two phonemes, which leads to oriented discriminant analysis (ODA). In addition, we equip ODA with the kernel trick again, which we refer to as kernel ODA. The models are tested on the CMU ARCTIC speech database. Our results indicate that our proposed kernel methods can outperform linear OPCA and linear ODA at finding a speaker-independent phoneme space.
Heeyoul Choi, Ricardo Gutierrez-Osuna, Seungjin Choi 0001, Yoonsuck Choe
ICPR1
2007 A relative trust-region algorithm for independent component analysis
Heeyoul Choi, Seungjin Choi 0001
Neurocomputing1
2007 Robust kernel Isomap
Heeyoul Choi, Seungjin Choi 0001
Pattern Recognit.1
2006 Relative Gradient Learning for Independent Subspace Analysis
abstract
Independent subspace analysis (ISA) is a generalization of independent component analysis (ICA), where multidimensional ICA is incorporated with the idea of invariant feature subspaces, allowing components in the same subspace to be dependent, but requiring independence between feature subspaces. In this paper we present a relative gradient algorithm for ISA, derived in the framework of the relative optimization as well as in a direct manner. Empirical comparison with the gradient ISA algorithm, shows that the relative gradient ISA algorithm achieves faster convergence, compared to the conventional gradient algorithm.
Heeyoul Choi, Seungjin Choi 0001
IJCNN1
2005 Relative trust region learning for ICA
abstract
We present a new learning method, relative trust-region learning, where we incorporate the relative optimization technique (M. Zibulevsky, Proc. ICA, pp. 897-902, 2003) into the trust-region method. We apply this relative trust-region learning method to the problem of independent component analysis (ICA), which leads to the relative TR-ICA algorithm which turns out to be faster than Newton-type ICA algorithms as well as gradient-based ICA algorithms and to possess the equivariant property. Empirical comparisons with several existing ICA algorithms confirm the fast convergence of the relative TR-ICA algorithm.
Heeyoul Choi, Seungjin Choi 0001
ICASSP (5)1
2004 Trust-region learning for ICA
abstract
A trust-region method is a quite attractive optimization technique, which finds a direction and a step size in an efficient and reliable manner with the help of a quadratic model of the objective function. It is, in general, faster than the steepest descent method and is free of a pre-selected constant learning rate. In addition to its convergence property (between linear and quadratic convergence), its stability is always guaranteed, in contrast to the Newton's method. We present an efficient implementation of the maximum likelihood independent component analysis (ICA) using the trust-region method, which leads to trust-region-based ICA (TR-ICA) algorithms. The useful behavior of our TR-ICA algorithms is confirmed through numerical experimental results.
Heeyoul Choi, Sookjeong Kim, Seungjin Choi 0001
IJCNN1