EDBT 2026 Demo / reviewers in the wild / expert
Byoung-Tak Zhang
dblp:09/5682
· DBLP profile ↗
207ranked-venue papers
16as first author
50since 2021 · last 2026
0000-0001-9890-0389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 166 · 14 first-author · 46 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 18 since 2021Systems, architecture and hardware · 20 · 1 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 since 2021Databases, data management, data science and information retrieval · 12
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion ModelsabstractImage generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work, we leverage intrinsic scene properties (e.g., depth, segmentation maps) that provide rich information about the underlying scene, unlike prior approaches that solely rely on image-text pairs or use intrinsics as conditional inputs. Our approach aims to co-generate both images and their corresponding intrinsics, enabling the model to implicitly capture the underlying scene structure and generate more spatially consistent and realistic images. Specifically, we first extract rich intrinsic scene properties from a large image dataset with pre-trained estimators, eliminating the need for additional scene information or explicit 3D representations. We then aggregate various intrinsic scene properties into a single latent variable using an autoencoder. Building upon pre-trained large-scale Latent Diffusion Models (LDMs), our method simultaneously denoises the image and intrinsic domains by carefully sharing mutual information so that the image and intrinsic reflect each other without degrading image quality. Experimental results demonstrate that our method corrects spatial inconsistencies and produces a more natural layout of scenes while maintaining the fidelity and textual alignment of the base model (e.g., Stable Diffusion). Hyundo Lee, Suhyung Choi, Inwoo Hwang, Byoung-Tak Zhang |
AAAI | 4 |
| 2026 | Neural Collapse-Informed Initialization with Perturbation Injection in Classification-based Metric LearningabstractRecent studies have revealed Neural Collapse (NC) in deep classifiers, where last-layer weights and features align into an equiangular tight frame (ETF), concentrating class information along specific embedding directions. However, conventional fine-tuning typically disregards this structure, initializing task-specific classifier heads randomly. To explicitly leverage this phenomenon, we propose a simple yet effective method for metric learning: (1) initializing the classifier head along each class’s NC direction from a pretrained model to preserve the emergent structure, and (2) injecting small isotropic Gaussian noise during finetuning to boost generalization. In addition, we provide a theoretical bound proving that our method explicitly reduces cumulative weight drift from the NC-initialization, compared to standard finetuning. This suggests that our method better preserves the pretrained model’s class-specific structure. Empirically, this structural preservation yields Recall@K gains: reduced weight drift correlates with better performance. Concurrent decreases in the Neural Collapse 1 (NC1) measure confirm that stronger intra‐class cohesion underlies these improvements. Furthermore, we validate the effectiveness of our method on class‐imbalanced benchmarks. Jinhee Park 0001, Hee Bin Yoo, Byoung-Tak Zhang, Junseok Kwon |
AAAI | 4 |
| 2026 | PeriUn: Enhancing Unlearning by Selectively Forgetting Peripheral SamplesabstractOnce trained, neural networks memorize information in diffusely encoded parameters, making it difficult to forget in support of the right to be forgotten. Unlearning aims to remove the influence of data, with performance measured against a retrained model that excludes the data. However, understanding the behavior of gold-standard retraining remains underexplored. We compare original and retrained models and observe that most prediction changes occur in peripheral samples near decision boundaries. Consequently, we propose PeriUn, a selective strategy that unlearns only peripheral samples to mimic retrained model behavior with minimal disruption, unlike prior works that remove the entire request. Combined with the Random Label based method, PeriUn significantly improves both generalization and privacy metrics. Specifically, on TinyImageNet with VGG16, PeriUn increases the Tug-of-War score by 22 points compared to the strongest. Besides, the MIA gap score surpasses the state-of-the-art method, improving by 8.7 points after applying PeriUn. Further analyses confirm that PeriUn better preserves the feature space and aligns closely with the retrained model. Hee Bin Yoo, Dong-Sig Han, Jaein Kim 0004, Byoung-Tak Zhang |
AAAI | 4 |
| 2026 | Hybrid State Representation for Video Procedure PlanningabstractAccurate state representation is critical for effective procedure planning from visual inputs. Existing methods typically rely on sentence-level natural language descriptions to represent states. However, such representations are often ambiguous and fail to capture fine-grained object interactions, leading to errors in complex scenarios. To overcome these limitations, we propose a Hybrid State Representation (HSR), inspired by human-like reasoning, that models procedural states with both structural precision and contextual clarity. HSR integrates two complementary modalities: (1) Semantic State Graphs (SSGs), which explicitly encode objects, attributes, and their relations, and (2) contextual Question-Answer (QA) pairs, which act as semantic probes to disambiguate critical state transitions. We further design a heterogeneous encoder to fuse these components and intro-duce a visual-state alignment objective to ground the hybrid representation in the visual context. Extensive experiments on the CrossTask, COIN, and NIV benchmarks demonstrate that our method establishes a new state-of-the-art, achieving significant gains on the strict Success Rate (SR) metric. Ablation studies confirm that both the structural (SSG) and contextual (QA) components of HSR are essential for the observed performance gains. Woosuk Choi, Youwon Jang, Minsu Lee 0001, Byoung-Tak Zhang |
WACV | 4 |
| 2026 | Toward object-centric spatial abstraction for probabilistic environment understandingabstractRecent advances in neural networks have enabled agents to perform high-level reasoning in complex environments, where robust environmental understanding must often be obtained from accumulated observations rather than from fixed scene labels. Existing spatial abstractions commonly rely on scene-level representations, object-presence cues, scene graphs, or language-grounded category prompts; depending on the task, these abstractions can either discard distributional object evidence or introduce additional structure that is not directly grounded in the agent’s own experience. To address this information abstraction imbalance, we design an object-centric spatial abstraction (OSA) that models environments as empirical distributions over object embeddings. Building on this abstraction, we present PEU-OSA, a Probabilistic Environment Understanding framework via Object-centric Spatial Abstraction. PEU-OSA comprises three conceptual modules: an object perception module, the OSA module, and a memory component. Scene observations are transformed into object representations, abstracted into empirical distributions through OSA, and accumulated in memory for comparison with new observations. The OSA module captures three fundamental relationships in the latent space via kernel density estimation: object-object, object-environment, and environment-environment relationships. To characterize when these measurements are reliable, we propose the ( )-Statistically Separable (EDS) function, which quantifies the separability and concentration of object representations and provides theoretical conditions and empirical evidence for OSA estimates. We evaluate PEU-OSA through EDS-based analysis on ImageNet, scene classification on Places, chained inference in Replica, and applicability studies for embodied agents in Replica and Minecraft. In a Replica chained-inference setting designed to test generalization to unseen environments, PEU-OSA improves top-1 retrieval accuracy from 0.36 with CLIP-based vision-language retrieval to 0.78, suggesting that object-distribution memory can provide a useful grounding source when scene-level language retrieval is insufficient. These findings support PEU-OSA as an annotation-free, experience-grounded framework for probabilistic environment understanding across recognition, reasoning, and embodied memory scenarios. Won-Seok Choi 0006, Dong-Sig Han, Suhyung Choi, Hyeonseo Yang, Byoung-Tak Zhang |
Neurocomputing | 5 |
| 2026 | Confidence Controls Deep Metric Learning
Jinhee Park 0001, Hee Bin Yoo, Byoung-Tak Zhang, Junseok Kwon |
Mach. Learn. | 3 |
| 2025 | Truncated Gaussian Policy for Debiased Continuous ControlabstractIn continuous domains, reinforcement learning policies are often based on Gaussian distributions for their generality. However, the unbounded support of Gaussian policy can cause a bias toward sampling boundary actions in many continuous control tasks that impose action limits due to physical constraints. This "boundary action bias'' can negatively impact training in algorithms like Proximal Policy Optimization. Despite this, it has been overlooked in many existing research and applications. In this paper, we revisit this issue by presenting illustrative explanations and analysis from the sampling point of view. Then, we introduce a truncated Gaussian policy with inherent bounds as a minimal alternative to mitigate the bias. However, we find that the plain truncated Gaussian policy may lay the counter-bias, preferring interior actions: to balance the bias, we ultimately propose a scale-adjusted truncated Gaussian policy, where the distribution scale shrinks if the location is near the boundaries. This property makes boundary actions deterministic more than in plain truncated Gaussian, but still less than in original Gaussian. Extensive empirical studies and comparisons on various continuous control tasks demonstrate that the truncated Gaussian policies significantly reduce the rate of boundary action usage, while scale-adjusted ones successfully balance the bias and counter-bias. It generally outperforms the Gaussian policy and shows competitive results compared to other approaches designed to counteract the bias. Ganghun Lee, Minji Kim 0005, Minsu Lee 0001, Byoung-Tak Zhang |
AAAI | 4 |
| 2025 | On the Consistency of Video Large Language Models in Temporal ComprehensionabstractVideo large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency –a key indicator for robustness and trustworthiness of temporal grounding. After the model identifies an initial moment within the video content, we apply a series of probes to check if the model’s responses align with this initial grounding as an indicator of reliable comprehension. Our results reveal that current Video-LLMs are sensitive to variations in video contents, language queries, and task settings, unveiling severe deficiencies in maintaining consistency. We further explore common prompting and instruction-tuning methods as potential solutions, but find that their improvements are often unstable. To that end, we propose event temporal verification tuning that explicitly accounts for consistency, and demonstrate significant improvements for both grounding and consistency. Our data and code are open-sourced at https://github.com/minjoong507/Consistency-of-Video-LLM. Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang, Angela Yao |
CVPR | 3 |
| 2025 | Confidence-guided Refinement Reasoning for Zero-shot Question AnsweringabstractWe propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains.C2R strategically constructs and refines subquestions and their answers (sub-QAs), deriving a better confidence score for the target answer.C2R first curates a subset of sub-QAs to explore diverse reasoning paths, then compares the confidence scores of the resulting answer candidates to select the most reliable final answer.Since C2R relies solely on confidence scores derived from the model itself, it can be seamlessly integrated with various existing QA models, demonstrating consistent performance improvements across diverse models and benchmarks.Furthermore, we provide essential yet underexplored insights into how leveraging sub-QAs affects model behavior, specifically analyzing the impact of both the quantity and quality of sub-QAs on achieving robust and reliable reasoning.Q: What data structure is pictured in the image?Answer: Array Q: What data structure is pictured in the image?𝒒 𝟏 : How many elements are in the data structure?𝒂 𝟏 : 5. 𝒒 𝟐 : What do the arrows between the boxes indicate?𝒂 𝟐 : The arrows point from one box to the next, suggesting pointers or references to the next element in the sequence.𝒒 𝟑 : Youwon Jang, Woosuk Choi, Minjoon Jung, Minsu Lee 0001, Byoung-Tak Zhang |
EMNLP | 5 |
| 2025 | Ock: Unsupervised Dynamic Video Prediction With Object-Centric Kinematics
Yeon-Ji Song, Suhyung Choi, Jin-Hwa Kim, Byoung-Tak Zhang |
ICCV | 5 |
| 2025 | DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance SegmentationabstractIn logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelfpicking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB-based methods tend to over-segment objects due to their reliance on texture, while depth-based methods often under-segment by focusing primarily on geometric features. To address these limitations, we propose DA-Fusion, a deformable attention-based RGB-D fusion Transformer designed for unseen object instance segmentation. DA-Fusion effectively combines the strengths of both RGB and depth data, enhancing segmentation accuracy in cluttered and multi-layered object environments. We also introduce the Object Clutter Bin Dataset (OCBD), a benchmark dataset specifically tailored for evaluating bin-picking scenarios in top-down views. Extensive evaluations demonstrate that DA-Fusion outperforms state-of-the-art methods across diverse environments, making it particularly suited for real-world logistics tasks. Yesol Park, Hye Jung Yoon, Juno Kim, Byoung-Tak Zhang |
ICRA | 4 |
| 2025 | Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction FollowingabstractEmbodied Instruction Following (EIF) is the task of executing natural language instructions by navigating and interacting with objects in interactive environments. A key challenge in EIF is compositional task planning, typically addressed through supervised learning or few-shot in-context learning with labeled data. To this end, we introduce the Socratic Planner, a self-QA-based zero-shot planning method that infers an appropriate plan without any further training. The Socratic Planner first facilitates self-questioning and answering by the Large Language Model (LLM), which in turn helps generate a sequence of subgoals. While executing the subgoals, an embodied agent may encounter unexpected situations, such as unforeseen obstacles. The Socratic Planner then adjusts plans based on dense visual feedback through a visuallygrounded re-planning mechanism. Experiments demonstrate the effectiveness of the Socratic Planner, outperforming current state-of-the-art planning models on the ALFRED benchmark across all metrics, particularly excelling in long-horizon tasks that demand complex inference. We further demonstrate its real-world applicability through deployment on a physical robot for long-horizon tasks. Suyeon Shin, Sujin Jeon, Junghyun Kim 0009, Gi-Cheon Kang, Byoung-Tak Zhang |
ICRA | 5 |
| 2025 | CDIS : Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection MergingabstractClass-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation. Existing approaches typically project per-frame 2D instance masks into 3D and merge them, which often breaks object identities across time and yields fragmented 3D instances. We introduce Cross-Dimensional Class-Agnostic 3D Instance Segmentation (CDIS), a zero-shot framework that explicitly tracks 2D instance masks across frames and associates them with 3D superpoints, creating a feedback loop between 2D and 3D. This cross-dimensional reasoning links temporally stable 2D tracks with spatially coherent 3D regions, producing globally consistent 3D instance labels without any 3D-specific training. Experiments on benchmark datasets demonstrate that CDIS achieves higher accuracy and consistency than state-of-the-art zero-shot methods, while remaining efficient and scalable to diverse real-world environments. Juno Kim, Hye Jung Yoon, Yesol Park, Byoung-Tak Zhang |
IROS | 4 |
| 2025 | How Classifier Features Transfer to Downstream: An Asymptotic Analysis in a Two-Layer ModelabstractNeural networks learn effective feature representations, which can be transferred to new tasks without additional training.
While larger datasets are known to improve feature transfer, the theoretical conditions for the success of such transfer remain unclear.
This work investigates feature transfer in networks trained for classification to identify the conditions that enable effective clustering in unseen classes.
We first reveal that higher similarity between training and unseen distributions leads to improved Cohesion and Separability.
We then show that feature expressiveness is enhanced when inputs are similar to the training classes, while the features of irrelevant inputs remain indistinguishable.
We validate our analysis on synthetic and benchmark datasets, including CAR, CUB, SOP, ISC, and ImageNet.
Our analysis highlights the importance of the similarity between training classes and the input distribution for successful feature transfer. Hee Bin Yoo, Sungyoon Lee, Cheongjae Jang, Dong-Sig Han, Jaein Kim 0004, Seunghyeon Lim, Byoung-Tak Zhang |
NeurIPS | 7 |
| 2025 | Background-Aware Moment Detection for Video Moment RetrievalabstractVideo moment retrieval (VMR) identifies a specific moment in an untrimmed video for a given natural language query. This task is prone to suffer the weak alignment problem innate in video datasets. Due to the ambiguity, a query does not fully cover the relevant details of the corresponding moment, or the moment may contain misaligned and irrelevant frames, potentially limiting further performance gains. To tackle this problem, we propose a background-aware moment detection transformer (BM-DETR). Our model adopts a contrastive approach, carefully utilizing the negative queries matched to other moments in the video. Specifically, our model learns to predict the target moment from the joint probability of each frame given the positive query and the complement of negative queries. This leads to effective use of the surrounding background, improving moment sensitivity and enhancing overall alignments in videos. Extensive experiments on four benchmarks demonstrate the effectiveness of our approach. Our code is available at: https://github.com/minjoong507/BM-DETR Minjoon Jung, Youwon Jang, Seongho Choi 0001, Joochan Kim, Jin-Hwa Kim, Byoung-Tak Zhang |
WACV | 6 |
| 2025 | INQUIRER: Harnessing internal knowledge graphs for video question generationabstractVideo question generation (VideoQG) aims to generate questions about video content to facilitate and assess video understanding. Existing works which primarily condition question generation on answer-related information such as the answer itself or its attributes. However, these methods are primarily designed as data augmentation techniques and thus struggle to produce semantically diverse questions. We propose INQUIRER, a novel VideoQG framework that leverages internal knowledge graphs derived from video information to generate meaningful and diverse questions. INQUIRER consists of three key steps: KCon, which constructs an internal knowledge graph to represent a video similarly to human knowledge structures, QGen which generates questions based on the video and the knowledge graph, and QCur which refines the generated questions to ensure quality and contextual relevance. Each generated question is accompanied by a correct answer and plausible distractors to support downstream QA evaluation. To comprehensively evaluate the generated QAs and utility of INQUIRER from multiple perspectives, we utilize widely used video question answering (VideoQA) benchmarks, including DramaQA, TVQA, How2QA, and STAR. Experiment results demonstrate that INQUIRER not only generates high-quality question-answer pairs but also significantly enhances VideoQA performance, validating its effectiveness as a robust framework for video question generation. Woosuk Choi, Youwon Jang, Minsu Lee 0001, Byoung-Tak Zhang |
Knowl. Based Syst. | 4 |
| 2024 | DUEL: Duplicate Elimination on Active Memory for Self-Supervised Class-Imbalanced LearningabstractRecent machine learning algorithms have been developed using well-curated datasets, which often require substantial cost and resources. On the other hand, the direct use of raw data often leads to overfitting towards frequently occurring class information. To address class imbalances cost-efficiently, we propose an active data filtering process during self-supervised pre-training in our novel framework, Duplicate Elimination (DUEL). This framework integrates an active memory inspired by human working memory and introduces distinctiveness information, which measures the diversity of the data in the memory, to optimize both the feature extractor and the memory. The DUEL policy, which replaces the most duplicated data with new samples, aims to enhance the distinctiveness information in the memory and thereby mitigate class imbalances. We validate the effectiveness of the DUEL framework in class-imbalanced environments, demonstrating its robustness and providing reliable results in downstream tasks. We also analyze the role of the DUEL policy in the training process through various metrics and visualizations. Won-Seok Choi 0006, Hyundo Lee, Dong-Sig Han, Heeyeon Koo, Byoung-Tak Zhang |
AAAI | 6 |
| 2024 | Unveiling the Significance of Toddler-Inspired Reward Transition in Goal-Oriented Reinforcement LearningabstractToddlers evolve from free exploration with sparse feedback to exploiting prior experiences for goal-directed learning with denser rewards. Drawing inspiration from this Toddler-Inspired Reward Transition, we set out to explore the implications of varying reward transitions when incorporated into Reinforcement Learning (RL) tasks. Central to our inquiry is the transition from sparse to potential-based dense rewards, which share optimal strategies regardless of reward changes. Through various experiments, including those in egocentric navigation and robotic arm manipulation tasks, we found that proper reward transitions significantly influence sample efficiency and success rates. Of particular note is the efficacy of the toddler-inspired Sparse-to-Dense (S2D) transition. Beyond these performance metrics, using Cross-Density Visualizer technique, we observed that transitions, especially the S2D, smooth the policy loss landscape, promoting wide minima that enhance generalization in RL models. Yoonsung Kim, Hee Bin Yoo, Min Whoo Lee, Kibeom Kim, Won-Seok Choi 0006, Minsu Lee 0001, Byoung-Tak Zhang |
AAAI | 8 |
| 2024 | CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
Minjung Shin, Seongho Choi 0001, Yu-Jung Heo, Minsu Lee 0001, Byoung-Tak Zhang, Jeh-Kwang Ryu |
CogSci | 5 |
| 2024 | Continuous SO(3) Equivariant Convolution for 3D Point Cloud Analysis
Jaein Kim 0004, Hee Bin Yoo, Dong-Sig Han, Yeon-Ji Song, Byoung-Tak Zhang |
ECCV (52) | 5 |
| 2024 | Fine-Grained Causal Dynamics Learning with Quantization for Improving Robustness in Reinforcement LearningabstractCausal dynamics learning has recently emerged as a promising approach to enhancing robustness in reinforcement learning (RL). Typically, the goal is to build a dynamics model that makes predictions based on the causal relationships among the entities. Despite the fact that causal connections often manifest only under certain contexts, existing approaches overlook such fine-grained relationships and lack a detailed understanding of the dynamics. In this work, we propose a novel dynamics model that infers fine-grained causal structures and employs them for prediction, leading to improved robustness in RL. The key idea is to jointly learn the dynamics model with a discrete latent variable that quantizes the state-action space into subgroups. This leads to recognizing meaningful context that displays sparse dependencies, where causal structures are learned for each subgroup throughout the training. Experimental results demonstrate the robustness of our method to unseen states and locally spurious correlations in downstream tasks where fine-grained causal reasoning is crucial. We further illustrate the effectiveness of our subgroup-based approach with quantization in discovering fine-grained causal relationships compared to prior methods. Inwoo Hwang, Yunhyeok Kwak, Suhyung Choi, Byoung-Tak Zhang, Sanghack Lee |
ICML | 4 |
| 2024 | HAPFI: History-Aware Planning based on Fused InformationabstractEmbodied Instruction Following (EIF) is a task of planning a long sequence of sub-goals given high-level natural language instructions, such as "Rinse a slice of lettuce and place on the white table next to the fork". To successfully execute these long-term horizon tasks, we argue that an agent must consider its past, i.e., historical data, when making decisions in each step. Nevertheless, recent approaches in EIF often neglects the knowledge from historical data and also do not effectively utilize information across the modalities. To this end, we propose History-Aware Planning based on Fused Information(HAPFI), effectively leveraging the historical data from diverse modalities that agents collect while interacting with the environment. Specifically, HAPFI integrates multiple modalities, including historical RGB observations, bounding boxes, sub-goals, and high-level instructions, by effectively fusing modalities via our Mutually Attentive Fusion method. Through experiments with diverse comparisons, we show that an agent utilizing historical multi-modal information surpasses all the compared methods that neglect the historical data in terms of action planning capability, enabling the generation of well-informed action plans for the next step. Moreover, we provided qualitative evidence highlighting the significance of leveraging historical multi-modal data, particularly in scenarios where the agent encounters intermediate failures, showcasing its robust re-planning capabilities. Sujin Jeon, Suyeon Shin, Byoung-Tak Zhang |
ICRA | 3 |
| 2024 | PROGrasp: Pragmatic Human-Robot Communication for Object GraspingabstractInteractive Object Grasping (IOG) is the task of identifying and grasping the desired object via human-robot natural language interaction. Current IOG systems assume that a human user initially specifies the target object’s category (e.g., bottle). Inspired by pragmatics, where humans often convey their intentions by relying on context to achieve goals, we introduce a new IOG task, Pragmatic-IOG, and the corresponding dataset, Intention-oriented Multi-modal Dialogue (IM-Dial). In our proposed task scenario, an intention-oriented utterance (e.g., "I am thirsty") is initially given to the robot. The robot should then identify the target object by interacting with a human user. Based on the task setup, we propose a new robotic system that can interpret the user’s intention and pick up the target object, Pragmatic Object Grasping (PROGrasp). PROGrasp performs Pragmatic-IOG by incorporating modules for visual grounding, question asking, object grasping, and most importantly, answer interpretation for pragmatic inference. Experimental results show that PROGrasp is effective in offline (i.e., target object discovery) and online (i.e., IOG with a physical robot arm) settings. Code and data are available at https://github.com/gicheonkang/prograsp. Gi-Cheon Kang, Junghyun Kim 0009, Jaein Kim 0004, Byoung-Tak Zhang |
ICRA | 4 |
| 2024 | Multi-Object RANSAC: Efficient Plane Clustering Method in a ClutterabstractIn this paper, we propose a novel method for plane clustering specialized in cluttered scenes using an RGB-D camera and validate its effectiveness through robot grasping experiments. Unlike existing methods, which focus on large- scale indoor structures, our approach—Multi-Object RANSAC emphasizes cluttered environments that contain a wide range of objects with different scales. It enhances plane segmentation by generating subplanes in Deep Plane Clustering (DPC) module, which are then merged with the final planes by postprocessing. DPC rearranges the point cloud by voting layers to make subplane clusters, trained in a self-supervised manner using pseudo-labels generated from RANSAC. Multi-Object RANSAC demonstrates superior plane instance segmentation performances over other recent RANSAC applications. We conducted an experiment on robot suction-based grasping, comparing our method with vision-based grasping network and RANSAC applications. The results from this real-world scenario showed its remarkable performance surpassing the baseline methods, highlighting its potential for advanced scene understanding and manipulation. Seunghyeon Lim, Youngjae Yoo, Jun Ki Lee, Byoung-Tak Zhang |
ICRA | 4 |
| 2024 | PGA: Personalizing Grasping Agents with Single Human-Robot InteractionabstractLanguage-Conditioned Robotic Grasping (LCRG) aims to develop robots that comprehend and grasp objects based on natural language instructions. While the ability to understand personal objects like my wallet facilitates more natural interaction with human users, current LCRG systems only allow generic language instructions, e.g., the black-colored wallet next to the laptop. To this end, we introduce a task scenario GraspMine alongside a novel dataset aimed at pinpointing and grasping personal objects given personal indicators via learning from a single human-robot interaction, rather than a large labeled dataset. Our proposed method, Personalized Grasping Agent (PGA), addresses GraspMine by leveraging the unlabeled image data of the user’s environment, called Reminiscence. Specifically, PGA acquires personal object information by a user presenting a personal object with its associated indicator, followed by PGA inspecting the object by rotating it. Based on the acquired information, PGA pseudo-labels objects in the Reminiscence by our proposed label propagation algorithm. Harnessing the information acquired from the interactions and the pseudo-labeled objects in the Reminiscence, PGA adapts the object grounding model to grasp personal objects. This results in significant efficiency while previous LCRG systems rely on resource-intensive human annotations—necessitating hundreds of labeled data to learn my wallet. Moreover, PGA outperforms baseline methods across all metrics and even shows comparable performance compared to the fully-supervised method, which learns from 9k annotated data samples. We further validate PGA’s real-world applicability by employing a physical robot to execute GrsapMine. Code and data are publicly available at https://github.com/JHKim-snu/PGA. Junghyun Kim 0009, Gi-Cheon Kang, Jaein Kim 0004, Seoyun Yang, Minjoon Jung, Byoung-Tak Zhang |
IROS | 6 |
| 2024 | OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for RobotsabstractWe introduce OV-MAP, a novel approach to open-world 3D mapping for mobile robots by integrating open-features into 3D maps to enhance object recognition capabilities. A significant challenge arises when overlapping features from adjacent voxels reduce instance-level precision, as features spill over voxel boundaries, blending neighboring regions together. Our method overcomes this by employing a class-agnostic segmentation model to project 2D masks into 3D space, combined with a supplemented depth image created by merging raw and synthetic depth from point clouds. This approach, along with a 3D mask voting mechanism, enables accurate zero-shot 3D instance segmentation without relying on 3D supervised segmentation models. We assess the effectiveness of our method through comprehensive experiments on public datasets such as ScanNet200 and Replica, demonstrating superior zero-shot performance, robustness, and adaptability across diverse environments. Additionally, we conducted real-world experiments to demonstrate our method’s adaptability and robustness when applied to diverse real-world environments. Juno Kim, Yesol Park, Hye Jung Yoon, Byoung-Tak Zhang |
IROS | 4 |
| 2024 | Seg2Grasp: A Robust Modular Suction Grasping in Bin PickingabstractCurrent bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in unstructured environments. To overcome these limitations, we introduce Seg2Grasp, a modular pipeline designed for robust suction grasping in dynamic and cluttered bin scenarios. Seg2Grasp is built on a three-step process: Segmentation, Grasping, and Classification. The Segmentation module employs a Transformer-based model to generate class-agnostic object masks from RGB-D images, ensuring accurate detection across various conditions. The Grasping module uses surface normals and mask proposals to determine the optimal suction points, enhancing grasp success. Finally, the Classification module leverages fine-tuned open-vocabulary Mask-CLIP for precise object identification, enabling versatile handling of diverse objects. Real-world robotic experiments demonstrate that Seg2Grasp outperforms existing methods in success rates and adaptability, establishing it as a powerful tool for automated bin picking in industrial settings. Hye Jung Yoon, Juno Kim, Yesol Park, Jun-Ki Lee, Byoung-Tak Zhang |
IROS | 5 |
| 2024 | Efficient Monte Carlo Tree Search via On-the-Fly State-Conditioned Action AbstractionabstractMonte Carlo Tree Search (MCTS) has showcased its efficacy across a broad spectrum of decision-making problems. However, its performance often degrades under vast combinatorial action space, especially where an action is composed of multiple sub-actions. In this work, we propose an action abstraction based on the compositional structure between a state and sub-actions for improving the efficiency of MCTS under a factored action space. Our method learns a latent dynamics model with an auxiliary network that captures sub-actions relevant to the transition on the current state, which we call state-conditioned action abstraction. Notably, it infers such compositional relationships from high-dimensional observations without the known environment model. During the tree traversal, our method constructs the state-conditioned action abstraction for each node on-the-fly, reducing the search space by discarding the exploration of redundant sub-actions. Experimental results demonstrate the superior sample efficiency of our method compared to vanilla MuZero, which suffers from expansive action space. Yunhyeok Kwak, Inwoo Hwang, Sanghack Lee, Byoung-Tak Zhang |
UAI | 5 |
| 2023 | The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingabstractVisual dialog (VisDial) is a task of answering a sequence of questions grounded in an image, using the dialog history as context. Prior work has trained the dialog agents solely on VisDial data via supervised learning or leveraged pretraining on related vision-and-language datasets. This paper presents a semi-supervised learning approach for visually-grounded dialog, called Generative Self-Training (GST), to leverage unlabeled images on the Web. Specifically, GST first retrieves in-domain images through out-of-distribution detection and generates synthetic dialogs regarding the images via multimodal conditional text generation. GST then trains a dialog agent on the synthetic and the original Vis-Dial data. As a result, GST scales the amount of training data up to an order of magnitude that of VisDial (1.2M → 12.9M QA data). For robust training of the synthetic dialogs, we also propose perplexity-based data selection and multimodal consistency regularization. Evaluation on Vis-Dial v1.0 and v0.9 datasets shows that GST achieves new state-of-the-art results on both datasets. We further observe the robustness of GST against both visual and textual adversarial attacks. Finally, GST yields strong performance gains in the low-data regime. Code is available at https:/github.com/gicheonkang/gst-visdial. Gi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak, Byoung-Tak Zhang |
CVPR | 5 |
| 2023 | Learning Geometry-aware Representations by SketchingabstractUnderstanding geometric concepts, such as distance and shape, is essential for understanding the real world and also for many vision tasks. To incorporate such information into a visual representation of a scene, we propose learning to represent the scene by sketching, inspired by human behavior. Our method, coined Learning by Sketching (LBS), learns to convert an image into a set of colored strokes that explicitly incorporate the geometric information of the scene in a single inference step without requiring a sketch dataset. A sketch is then generated from the strokes where CLIP-based perceptual loss maintains a semantic similarity between the sketch and the image. We show theoretically that sketching is equivariant with respect to arbitrary affine transformations and thus provably preserves geometric information. Experimental results show that LBS substantially improves the performance of object attribute classification on the unlabeled CLEVR dataset, domain transfer between CLEVR and STL-10 datasets, and for diverse downstream tasks, confirming that LBS provides rich geometric information. Hyundo Lee, Inwoo Hwang, Hyunsung Go, Won-Seok Choi 0006, Kibeom Kim, Byoung-Tak Zhang |
CVPR | 6 |
| 2023 | Neural Collage Transfer: Artistic Reconstruction via Material ManipulationabstractCollage is a creative art form that uses diverse material scraps as a base unit to compose a single image. Although pixel-wise generation techniques can reproduce a target image in collage style, it is not a suitable method due to the solid stroke-by-stroke nature of the collage form. While some previous works for stroke-based rendering produced decent sketches and paintings, collages have received much less attention in research despite their popularity as a style. In this paper, we propose a method for learning to make collages via reinforcement learning without the need for demonstrations or collage artwork data. We design the collage Markov Decision Process (MDP), which allows the agent to handle various materials and propose a model-based soft actor-critic to mitigate the agent’s training burden derived from the sophisticated dynamics of collage. Moreover, we devise additional techniques such as active material selection and complexity-based multi-scale collage to handle target images at any size and enhance the results’ aesthetics by placing relatively more scraps in areas of high complexity. Experimental results show that the trained agent appropriately selected and pasted materials to regenerate the target image into a collage and obtained a higher evaluation score on content and style than pixel-wise generation methods. Code is available at https://github.com/northadventure/CollageRL. Ganghun Lee, Minji Kim 0005, Yunsu Lee, Minsu Lee 0001, Byoung-Tak Zhang |
ICCV | 5 |
| 2023 | Robust Map Fusion with Visual Attention Utilizing Multi-agent RendezvousabstractThe map fusion for multi-robot simultaneous localization and mapping (SLAM) consistently combines robot maps built independently into the global map. An established approach to map fusion is utilizing rendezvous, which refers to an encounter between multiple agents, to calculate the transformation into the global map. However, previous works using rendezvous have a limitation in that they are unreliable for certain circumstances, where the amount of agent observations or overlapping landmarks is limited. This work proposes a novel map fusion system which robustly fuses local maps in challenging rendezvous that lack shared information. Our system utilizes the single visual perception from rendezvous and estimates the relative pose between agents with the DOPE. Then our scheme transforms local maps with an estimated relative pose and predicts the misalignment from approximated maps by utilizing the attention mechanism of the vision transformer. Comparisons with the Hough transform-based method show that ours is significantly better when the overlap between local maps is insufficient. We also verify the robustness of our system against a similar real-world scenario. Jaein Kim 0004, Dong-Sig Han, Byoung-Tak Zhang |
ICRA | 3 |
| 2023 | EXOT: Exit-aware Object Tracker for Safe Robotic Manipulation of Moving ObjectabstractCurrent robotic hand manipulation narrowly operates with objects in predictable positions in limited environments. Thus, when the location of the target object deviates severely from the expected location, a robot sometimes responds in an unexpected way, especially when it operates with a human. For safe robot operation, we propose the EXit-aware Object Tracker (EXOT) on a robot hand camera that recognizes an object's absence during manipulation. The robot decides whether to proceed by examining the tracker's bounding box output containing the target object. We adopt an out-of-distribution classifier for more accurate object recognition since trackers can mistrack a background as a target object. To the best of our knowledge, our method is the first approach of applying an out-of-distribution classification technique to a tracker output. We evaluate our method on the first-person video benchmark dataset, TREK-150, and on the custom dataset, RMOT-223, that we collect from the UR5e robot. Then we test our tracker on the UR5e robot in real-time with a conveyor-belt sushi task, to examine the tracker's ability to track target dishes and to determine the exit status. Our tracker shows 38% higher exit-aware performance than a baseline method. The dataset and the code will be released at https://github.com/hskAlena/EXOT. Hyunseo Kim 0001, Hye Jung Yoon, Minji Kim 0005, Dong-Sig Han, Byoung-Tak Zhang |
ICRA | 5 |
| 2023 | GVCCI: Lifelong Learning of Visual Grounding for Language-Guided Robotic ManipulationabstractLanguage-Guided Robotic Manipulation (LGRM) is a challenging task as it requires a robot to understand human instructions to manipulate everyday objects. Recent approaches in LGRM rely on pre-trained Visual Grounding (VG) models to detect objects without adapting to manipulation environments. This results in a performance drop due to a substantial domain gap between the pre-training and real-world data. A straight-forward solution is to collect additional training data, but the cost of human-annotation is extortionate. In this paper, we propose Grounding Vision to Ceaselessly Created Instructions (GVCCI), a lifelong learning framework for LGRM, which continuously learns VG without human supervision. GVCCI iteratively generates synthetic instruction via object detection and trains the VG model with the generated data. We validate our framework in offline and online settings across diverse environments on different VG models. Experimental results show that accumulating synthetic data from GVCCI leads to a steady improvement in VG by up to 56.7% and improves resultant LGRM by up to 29.4%. Furthermore, the qualitative analysis shows that the unadapted VG model often fails to find correct objects due to a strong bias learned from the pre-training data. Finally, we introduce a novel VG dataset for LGRM, consisting of nearly 252k triplets of image-object-instruction from diverse manipulation environments. Junghyun Kim 0009, Gi-Cheon Kang, Jaein Kim 0004, Suyeon Shin, Byoung-Tak Zhang |
IROS | 5 |
| 2022 | Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question AnsweringabstractKnowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself.Answering complex questions that require multi-hop reasoning under weak supervision is considered as a challenging problem since i) no supervision is given to the reasoning process and ii) highorder semantics of multi-hop knowledge facts need to be captured.In this paper, we introduce a concept of hypergraph to encode highlevel semantics of a question and a knowledge base, and to learn high-order associations between them.The proposed model, Hypergraph Transformer, constructs a question hypergraph and a query-aware knowledge hypergraph, and infers an answer by encoding inter-associations between two hypergraphs and intra-associations in both hypergraph itself.Extensive experiments on two knowledgebased visual QA and two knowledge-based textual QA demonstrate the effectiveness of our method, especially for multi-hop reasoning problem.Our source code is available at https://github.com/yujungheo/ kbvqa-public. Yu-Jung Heo, Eun-Sol Kim, Woosuk Choi, Byoung-Tak Zhang |
ACL (1) | 4 |
| 2022 | Smooth-Swap: A Simple Enhancement for Face-Swapping with SmoothnessabstractFace-swapping models have been drawing attention for their compelling generation quality, but their complex architectures and loss functions often require careful tuning for successful training. We propose a new face-swapping model called ‘Smooth-Swap’, which excludes complex handcrafted designs and allows fast and stable training. The main idea of Smooth-Swap is to build smooth identity embedding that can provide stable gradients for identity change. Unlike the one used in previous models trained for a purely discriminative task, the proposed embedding is trained with a supervised contrastive loss promoting a smoother space. With improved smoothness, Smooth-Swap suffices to be composed of a generic U-Net-based generator and three basic loss functions, a far simpler design compared with the previous models. Extensive experiments on face-swapping benchmarks (FFHQ,$Face-Forensics++$) and face images in the wild show that our model is also quantitatively and qualitatively comparable or even superior to the existing methods. Jiseob Kim, Byoung-Tak Zhang |
CVPR | 3 |
| 2022 | Modal-specific Pseudo Query Generation for Video Corpus Moment RetrievalabstractVideo corpus moment retrieval (VCMR) is the task to retrieve the most relevant video moment from a large video corpus using a natural language query.For narrative videos, e.g., dramas or movies, the holistic understanding of temporal dynamics and multimodal reasoning is crucial.Previous works have shown promising results; however, they relied on the expensive query annotations for VCMR, i.e., the corresponding moment intervals.To overcome this problem, we propose a self-supervised learning framework: Modal-specific Pseudo Query Generation Network (MPGN).First, MPGN selects candidate temporal moments via subtitle-based moment sampling.Then, it generates pseudo queries exploiting both visual and textual information from the selected temporal moments.Through the multimodal information in the pseudo queries, we show that MPGN successfully learns to localize the video corpus moment without any explicit annotation.We validate the effectiveness of MPGN on the TVR dataset, showing competitive results compared with both supervised models and unsupervised setting models. Minjoon Jung, Seongho Choi 0001, Joochan Kim, Jin-Hwa Kim, Byoung-Tak Zhang |
EMNLP | 5 |
| 2022 | From Scratch to Sketch: Deep Decoupled Hierarchical Reinforcement Learning for Robotic Sketching AgentabstractWe present an automated learning framework for a robotic sketching agent that is capable of learning stroke-based rendering and motor control simultaneously. We formulate the robotic sketching problem as a deep decoupled hierarchical reinforcement learning; two policies for stroke-based rendering and motor control are learned independently to achieve sub-tasks for drawing, and form a hierarchy when cooperating for real-world drawing. Without hand-crafted features, drawing sequences or trajectories, and inverse kinematics, the proposed method trains the robotic sketching agent from scratch. We performed experiments with a 6-DoF robot arm with 2F gripper to sketch doodles. Our experimental results show that the two policies successfully learned the sub-tasks and collaborated to sketch the target images. Also, the robustness and flexibility were examined by varying drawing tools and surfaces. Ganghun Lee, Minji Kim 0005, Minsu Lee 0001, Byoung-Tak Zhang |
ICRA | 4 |
| 2022 | PlaceNet: Neural Spatial Representation Learning with Multimodal AttentionabstractSpatial representation capable of learning a myriad of environmental features is a significant challenge for natural spatial understanding of mobile AI agents. Deep generative models have the potential of discovering rich representations of observed 3D scenes. However, previous approaches have been mainly evaluated on simple environments, or focused only on high-resolution rendering of small-scale scenes, hampering generalization of the representations to various spatial variability. To address this, we present PlaceNet, a neural representation that learns through random observations in a self-supervised manner, and represents observed scenes with triplet attention using visual, topographic, and semantic cues. We evaluate the proposed method on a large-scale multimodal scene dataset consisting of 120 million indoor scenes, and show that PlaceNet successfully generalizes to various environments with lower training loss, higher image quality and structural similarity of predicted scenes, compared to a competitive baseline model. Additionally, analyses of the representations demonstrate that PlaceNet activates more specialized and larger numbers of kernels in the spatial representation, capturing multimodal spatial properties in complex environments. Chung-Yeon Lee, Youngjae Yoo, Byoung-Tak Zhang |
IJCAI | 3 |
| 2022 | Robust Imitation via Mirror Descent Inverse Reinforcement LearningabstractRecently, adversarial imitation learning has shown a scalable reward acquisition method for inverse reinforcement learning (IRL) problems. However, estimated reward signals often become uncertain and fail to train a reliable statistical model since the existing methods tend to solve hard optimization problems directly. Inspired by a first-order optimization method called mirror descent, this paper proposes to predict a sequence of reward functions, which are iterative solutions for a constrained convex problem. IRL solutions derived by mirror descent are tolerant to the uncertainty incurred by target density estimation since the amount of reward learning is regulated with respect to local geometric constraints. We prove that the proposed mirror descent update rule ensures robust minimization of a Bregman divergence in terms of a rigorous regret bound of $\mathcal{O}(1/T)$ for step sizes $\{\eta_t\}_{t=1}^{T}$. Our IRL method was applied on top of an adversarial framework, and it outperformed existing adversarial methods in an extensive suite of benchmarks. Dong-Sig Han, Hyunseo Kim 0001, Hyundo Lee, Je-Hwan Ryu, Byoung-Tak Zhang |
NeurIPS | 5 |
| 2022 | SelecMix: Debiased Learning by Contradicting-pair SamplingabstractNeural networks trained with ERM (empirical risk minimization) sometimes learn unintended decision rules, in particular when their training data is biased, i.e., when training labels are strongly correlated with undesirable features. To prevent a network from learning such features, recent methods augment training data such that examples displaying spurious correlations (i.e., bias-aligned examples) become a minority, whereas the other, bias-conflicting examples become prevalent. However, these approaches are sometimes difficult to train and scale to real-world data because they rely on generative models or disentangled representations. We propose an alternative based on mixup, a popular augmentation that creates convex combinations of training examples. Our method, coined SelecMix, applies mixup to contradicting pairs of examples, defined as showing either (i) the same label but dissimilar biased features, or (ii) different labels but similar biased features. Identifying such pairs requires comparing examples with respect to unknown biased features. For this, we utilize an auxiliary contrastive model with the popular heuristic that biased features are learned preferentially during training. Experiments on standard benchmarks demonstrate the effectiveness of the method, in particular when label noise complicates the identification of bias-conflicting examples. Inwoo Hwang, Yunhyeok Kwak, Seong Joon Oh, Damien Teney, Jin-Hwa Kim, Byoung-Tak Zhang |
NeurIPS | 7 |
| 2021 | DramaQA: Character-Centered Video Story Understanding with Hierarchical QAabstractDespite recent progress on computer vision and natural language processing, developing a machine that can understand video story is still hard to achieve due to the intrinsic difficulty of video story. Moreover, researches on how to evaluate the degree of video understanding based on human cognitive process have not progressed as yet. In this paper, we propose a novel video question answering (Video QA) task, DramaQA, for a comprehensive understanding of the video story. The DramaQA focuses on two perspectives: 1) Hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence. 2) Character-centered video annotations to model local coherence of the story. Our dataset is built upon the TV drama "Another Miss Oh" and it contains 17,983 QA pairs from 23,928 various length video clips, with each QA pair belonging to one of four difficulty levels. We provide 217,308 annotated images with rich character-centered annotations, including visual bounding boxes, behaviors and emotions of main characters, and coreference resolved scripts. Additionally, we suggest Multi-level Context Matching model which hierarchically understands character-centered representations of video to answer questions. We release our dataset and model publicly for research purposes, and we expect our work to provide a new perspective on video story understanding research. Seongho Choi 0001, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Minsu Lee 0001, Byoung-Tak Zhang |
AAAI | 7 |
| 2021 | Attend What You Need: Motion-Appearance Synergistic Networks for Video Question AnsweringabstractAhjeong Seo, Gi-Cheon Kang, Joonhan Park, Byoung-Tak Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ahjeong Seo, Gi-Cheon Kang, Joonhan Park, Byoung-Tak Zhang |
ACL/IJCNLP (1) | 4 |
| 2021 | Passive Versus Active: Frameworks of Active Learning for Linking Humans to Machines
Jaeseo Lim, Hwiyeol Jo, Byoung-Tak Zhang, Jooyong Park |
CogSci | 3 |
| 2021 | Co-Attentional Transformers for Story-Based Video UnderstandingabstractInspired by recent trends in vision and language learning, we explore applications of attention mechanisms for visio-lingual fusion within an application to story-based video understanding. Like other video-based QA tasks, video story understanding requires agents to grasp complex temporal dependencies. However, as it focuses on the narrative aspect of video it also requires understanding of the interactions between different characters, as well as their actions and their motivations. We propose a novel co-attentional transformer model to better capture long-term dependencies seen in visual stories such as dramas and measure its performance on the video question answering task. We evaluate our approach on the recently introduced DramaQA dataset which features character-centered video story understanding questions. Our model outperforms the baseline model by 8 percentage points overall, at least 4.95 and up to 12.8 percentage points on all difficulty levels and manages to beat the winner of the DramaQA challenge. Björn Bebensee, Byoung-Tak Zhang |
ICASSP | 2 |
| 2021 | Toddler-Guidance Learning: Impacts of Critical Period on Multimodal AI AgentsabstractCritical periods are phases during which a toddler’s brain develops in spurts. To promote children’s cognitive development, proper guidance is critical in this stage. However, it is not clear whether such a critical period also exists for the training of AI agents. Similar to human toddlers, well-timed guidance and multimodal interactions might significantly enhance the training efficiency of AI agents as well. To validate this hypothesis, we adapt this notion of critical periods to learning in AI agents and investigate the critical period in the virtual environment for AI agents. We formalize the critical period and Toddler-guidance learning in the reinforcement learning (RL) framework. Then, we built up a toddler-like environment with VECA toolkit to mimic human toddlers’ learning characteristics. We study three discrete levels of mutual interaction: weak-mentor guidance (sparse reward), moderate mentor guidance (helper-reward), and mentor demonstration (behavioral cloning). We also introduce the EAVE dataset consisting of 30,000 real-world images to fully reflect the toddler’s viewpoint. We evaluate the impact of critical periods on AI agents from two perspectives: how and when they are guided best in both uni- and multimodal learning. Our experimental results show that both uni- and multimodal agents with moderate mentor guidance and critical period on 1 million and 2 million training steps show a noticeable improvement. We validate these results with transfer learning on the EAVE dataset and find the performance advancement on the same critical period and the guidance. Kwanyoung Park, Hyunseok Oh, Ganghun Lee, Minsu Lee 0001, Youngki Lee 0001, Byoung-Tak Zhang |
ICMI | 7 |
| 2021 | Message Passing Adaptive Resonance Theory for Online Active Semi-supervised LearningabstractActive learning is widely used to reduce labeling effort and training time by repeatedly querying only the most beneficial samples from unlabeled data. In real-world problems where data cannot be stored indefinitely due to limited storage or privacy issues, the query selection and the model update should be performed as soon as a new data sample is observed. Various online active learning methods have been studied to deal with these challenges; however, there are difficulties in selecting representative query samples and updating the model efficiently without forgetting. In this study, we propose Message Passing Adaptive Resonance Theory (MPART) that learns the distribution and topology of input data online. Through message passing on the topological graph, MPART actively queries informative and representative samples, and continuously improves the classification performance using both labeled and unlabeled data. We evaluate our model in stream-based selective sampling scenarios with comparable query selection strategies, showing that MPART significantly outperforms competitive models. Injune Hwang, Hyundo Lee, Hyunseo Kim 0001, Won-Seok Choi 0006, Joseph J. Lim, Byoung-Tak Zhang |
ICML | 7 |
| 2021 | Multimodal Anomaly Detection based on Deep Auto-Encoder for Object Slip Perception of Mobile Manipulation RobotsabstractObject slip perception is essential for mobile manipulation robots to perform manipulation tasks reliably in the dynamic real-world. Traditional approaches to robot arms’ slip perception use tactile or vision sensors. However, mobile robots still have to deal with noise in their sensor signals caused by the robot’s movement in a changing environment. To solve this problem, we present an anomaly detection method that utilizes multisensory data based on a deep autoencoder model. The proposed framework integrates heterogeneous data streams collected from various robot sensors, including RGB and depth cameras, a microphone, and a force-torque sensor. The integrated data is used to train a deep autoencoder to construct latent representations of the multisensory data that indicate the normal status. Anomalies can then be identified by error scores measured by the difference between the trained encoder’s latent values and the latent values of reconstructed input data. In order to evaluate the proposed framework, we conducted an experiment that mimics an object slip by a mobile service robot operating in a real-world environment with diverse household objects and different moving patterns. The experimental results verified that the proposed framework reliably detects anomalies in object slip situations despite various object types and robot behaviors, and visual and auditory noise in the environment. Youngjae Yoo, Chung-Yeon Lee, Byoung-Tak Zhang |
ICRA | 3 |
| 2021 | Goal-Aware Cross-Entropy for Multi-Target Reinforcement LearningabstractLearning in a multi-target environment without prior knowledge about the targets requires a large amount of samples and makes generalization difficult. To solve this problem, it is important to be able to discriminate targets through semantic understanding. In this paper, we propose goal-aware cross-entropy (GACE) loss, that can be utilized in a self-supervised way using auto-labeled goal states alongside reinforcement learning. Based on the loss, we then devise goal-discriminative attention networks (GDAN) which utilize the goal-relevant information to focus on the given instruction. We evaluate the proposed methods on visual navigation and robot arm manipulation tasks with multi-target environments and show that GDAN outperforms the state-of-the-art methods in terms of task success ratio, sample efficiency, and generalization. Additionally, qualitative analyses demonstrate that our proposed method can help the agent become aware of and focus on the given instruction clearly, promoting goal-directed behavior. Kibeom Kim, Min Whoo Lee, Yoonsung Kim, Je-Hwan Ryu, Minsu Lee 0001, Byoung-Tak Zhang |
NeurIPS | 6 |
| 2021 | RoboCup@Home 2021 Domestic Standard Platform League Winner
Dongwoon Song, Taewoong Kang, Jae-Bong Yi, Joonyoung Kim 0004, Taeyang Kim, Chung-Yeon Lee, Je-Hwan Ryu, Minji Kim 0005, Hyun-Jun Jo, Byoung-Tak Zhang, Jae-Bok Song, Seung-Joon Yi |
RoboCup | 10 |
| 2020 | Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video DataabstractConventional sequential learning methods such as Recurrent Neural Networks (RNNs) focus on interactions between consecutive inputs, i.e. first-order Markovian dependency. However, most of sequential data, as seen with videos, have complex dependency structures that imply variable-length semantic flows and their compositions, and those are hard to be captured by conventional methods. Here, we propose Cut-Based Graph Learning Networks (CB-GLNs) for learning video data by discovering these complex structures of the video. The CB-GLNs represent video data as a graph, with nodes and edges corresponding to frames of the video and their dependencies respectively. The CB-GLNs find compositional dependencies of the data in multilevel graph forms via a parameterized kernel with graph-cut and a message passing framework. We evaluate the proposed method on the two different tasks for video understanding: Video theme classification (Youtube-8M dataset (Abu-El-Haija et al. 2016)) and Video Question and Answering (TVQA dataset(Lei et al. 2018)). The experimental results show that our model efficiently learns the semantic compositional structure of video data. Furthermore, our model achieves the highest performance in comparison to other baseline methods. Kyoung-Woon On, Eun-Sol Kim, Yu-Jung Heo, Byoung-Tak Zhang |
AAAI | 4 |
| 2020 | Effect of Active Pre-Learning Activities on Humans and Machines
Jaeseo Lim, Hwiyeol Jo, Byoung-Tak Zhang, Jooyong Park |
CogSci | 3 |
| 2020 | Hypergraph Attention Networks for Multimodal LearningabstractOne of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and extract a joint representation of the modalities based on a co-attention map constructed in the semantic space. HANs follow the process: constructing the common semantic space with symbolic graphs of each modality, matching the semantics between sub-structures of the symbolic graphs, constructing co-attention maps between the graphs in the semantic space, and integrating the multimodal inputs using the co-attention maps to get the final joint representation. From the qualitative analysis with two Visual Question and Answering datasets, we discover that 1) the alignment of the information levels between the modalities is important, and 2) the symbolic graphs are very powerful ways to represent the information of the low-level signals in alignment. Moreover, HANs dramatically improve the state-of-the-art accuracy on the GQA dataset from 54.6\% to 61.88\% only using the symbolic information in quantitatively. Eun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo, Byoung-Tak Zhang |
CVPR | 5 |
| 2020 | Label Propagation Adaptive Resonance Theory for Semi-Supervised Continuous LearningabstractSemi-supervised learning and continuous learning are fundamental paradigms for human-level intelligence. To deal with real-world problems where labels are rarely given and the opportunity to access the same data is limited, it is necessary to apply these two paradigms in a joined fashion. In this paper, we propose Label Propagation Adaptive Resonance Theory (LPART) for semi-supervised continuous learning. LPART uses an online label propagation mechanism to perform classification and gradually improves its accuracy as the observed data accumulates. We evaluated the proposed model on visual (MNIST, SVHN, CIFAR-10) and audio (NSynth) datasets by adjusting the ratio of the labeled and unlabeled data. The accuracies are much higher when both labeled and unlabeled data are used, demonstrating the significant advantage of LPART in environments where the data labels are scarce. Injune Hwang, Gi-Cheon Kang, Won-Seok Choi 0006, Hyunseo Kim 0001, Byoung-Tak Zhang |
ICASSP | 6 |
| 2019 | CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven CommunicationabstractJin-Hwa Kim, Nikita Kitaev, Xinlei Chen, Marcus Rohrbach, Byoung-Tak Zhang, Yuandong Tian, Dhruv Batra, Devi Parikh. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Jin-Hwa Kim, Nikita Kitaev, Xinlei Chen, Marcus Rohrbach, Byoung-Tak Zhang, Yuandong Tian, Dhruv Batra, Devi Parikh |
ACL (1) | 5 |
| 2019 | Modeling Delay Discounting using Gaussian Process with Active Learning
Jorge Chang, Jiseob Kim, Byoung-Tak Zhang, Mark A. Pitt, Jay I. Myung |
CogSci | 3 |
| 2019 | Problem Difficulty in Arithmetic Cognition: Humans and Connectionist Models
Sungjae Cho, Jaeseo Lim, Chris Hickey, Byoung-Tak Zhang |
CogSci | 4 |
| 2019 | Dual Attention Networks for Visual Reference Resolution in Visual DialogabstractGi-Cheon Kang, Jaeseo Lim, Byoung-Tak Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Gi-Cheon Kang, Jaeseo Lim, Byoung-Tak Zhang |
EMNLP/IJCNLP (1) | 3 |
| 2019 | WithDorm: Dormitory Solution for Linking RoommatesabstractExperiences in universities are important for emotional maturation and offer an opportunity to develop individual characteristics and skills needed for social life. There are diverse issues affecting the quality of dormitory life and roommate relationships, which can influence one's psychosocial development. In this paper, we propose WithDorm, a mobile application to help communication with roommates and tighten their connections, and thereby assisting the users' emotional health and psychosocial development. We analyzed dormitory roommate issues from a human-centered perspective and narrowed down to three design implications after dormitory life modeling. Furthermore, we implemented the design implications in a prototype and performed a usability test to evaluate and improve the design. The final design, WithDorm, is aware of dormitory-specific concerns, collects and adapts to users' lifestyles, and initiates humanhuman interaction among roommates. Minji Kwak, Seung-Hee Yang, Jaeseo Lim, Byoung-Tak Zhang |
MobileHCI | 5 |
| 2018 | Perception-Action-Learning System for Mobile Social-Service Robots Using Deep LearningabstractWe introduce a novel perception-action-learning system for mobile social-service robots. The state-of-the-art deep learning techniques were incorporated into each module which significantly improves the performance in solving social service tasks. The system not only demonstrated fast and robust performance in a homelike environment but also achieved the highest score in the RoboCup2017@Home Social Standard Platform League (SSPL) held in Nagoya, Japan. Beom-Jin Lee, Chung-Yeon Lee, Kyung-Wha Park, Sungjun Choi, Cheolho Han, Dong-Sig Han, Christina Baek, Patrick Mokodir Emaase, Byoung-Tak Zhang |
AAAI | 10 |
| 2018 | Multimodal Dual Attention Memory for Video Story Question Answering
Seongho Choi 0001, Jin-Hwa Kim, Byoung-Tak Zhang |
ECCV (15) | 4 |
| 2018 | Robust Human Following by Deep Bayesian Trajectory Prediction for Home Service RobotsabstractThe capability of following a person is crucial in service-oriented robots for human assistance and cooperation. Though a vast variety of following systems exist, they lack robustness against dynamic changes of the environment and relocating to continue following a lost target. Here we present a robust human following system that has the extendability to commercial service robot platforms having a RGB-D camera. The proposed framework integrates deep learning methods for perception and variational Bayesian techniques for trajectory prediction. Deep learning modules enable robots to accompany a person by detecting the target, learning the target and following while avoiding collision within the dynamic home environment. The variational Bayesian techniques robustly predict the trajectory of the target by empowering the following ability of the robot when target is lost. We experimentally demonstrate the capability of the deep Bayesian trajectory prediction method on real-time usage, following abilities, collision avoidance and trajectory prediction of the system. The proposed system was deployed at the RoboCup@Home 2017 Social Standard Platform League and successfully demonstrated its robust functions and smooth person following capability resulting in winning the 1st place. Beom-Jin Lee, Christina Baek, Byoung-Tak Zhang |
ICRA | 4 |
| 2018 | Bilinear Attention NetworksabstractAttention networks in multimodal learning provide an efficient way to utilize given visual information selectively. However, the computational cost to learn attention distributions for every pair of multimodal input channels is prohibitively expensive. To solve this problem, co-attention builds two separate attention distributions for each modality neglecting the interaction between multimodal inputs. In this paper, we propose bilinear attention networks (BAN) that find bilinear attention distributions to utilize given vision-language information seamlessly. BAN considers bilinear interactions among two groups of input channels, while low-rank bilinear pooling extracts the joint representations for each pair of channels. Furthermore, we propose a variant of multimodal residual networks to exploit eight-attention maps of the BAN efficiently. We quantitatively and qualitatively evaluate our model on visual question answering (VQA 2.0) and Flickr30k Entities datasets, showing that BAN significantly outperforms previous methods and achieves new state-of-the-arts on both datasets. Jin-Hwa Kim, Jaehyun Jun, Byoung-Tak Zhang |
NeurIPS | 3 |
| 2018 | Answerer in Questioner's Mind: Information Theoretic Approach to Goal-Oriented Visual DialogabstractGoal-oriented dialog has been given attention due to its numerous applications in artificial intelligence. Goal-oriented dialogue tasks occur when a questioner asks an action-oriented question and an answerer responds with the intent of letting the questioner know a correct action to take. To ask the adequate question, deep learning and reinforcement learning have been recently applied. However, these approaches struggle to find a competent recurrent neural questioner, owing to the complexity of learning a series of sentences. Motivated by theory of mind, we propose "Answerer in Questioner's Mind" (AQM), a novel information theoretic algorithm for goal-oriented dialog. With AQM, a questioner asks and infers based on an approximated probabilistic model of the answerer. The questioner figures out the answerer’s intention via selecting a plausible question by explicitly calculating the information gain of the candidate intentions and possible answers to each question. We test our framework on two goal-oriented visual dialog tasks: "MNIST Counting Dialog" and "GuessWhat?!". In our experiments, AQM outperforms comparative algorithms by a large margin. Sang-Woo Lee 0001, Yu-Jung Heo, Byoung-Tak Zhang |
NeurIPS | 3 |
| 2018 | Fluid Dynamic Models for Bhattacharyya-Based Discriminant AnalysisabstractClassical discriminant analysis attempts to discover a low-dimensional subspace where class label information is maximally preserved under projection. Canonical methods for estimating the subspace optimize an information-theoretic criterion that measures the separation between the class-conditional distributions. Unfortunately, direct optimization of the information-theoretic criteria is generally non-convex and intractable in high-dimensional spaces. In this work, we propose a novel, tractable algorithm for discriminant analysis that considers the class-conditional densities as interacting fluids in the high-dimensional embedding space. We use the Bhattacharyya criterion as a potential function that generates forces between the interacting fluids, and derive a computationally tractable method for finding the low-dimensional subspace that optimally constrains the resulting fluid flow. We show that this model properly reduces to the optimal solution for homoscedastic data as well as for heteroscedastic Gaussian distributions with equal means. We also extend this model to discover optimal filters for discriminating Gaussian processes and provide experimental results and comparisons on a number of datasets. Yung-Kyun Noh, Jihun Hamm, Frank C. Park 0001, Byoung-Tak Zhang, Daniel D. Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Generative Local Metric Learning for Nearest Neighbor ClassificationabstractWe consider the problem of learning a local metric in order to enhance the performance of nearest neighbor classification. Conventional metric learning methods attempt to separate data distributions in a purely discriminative manner; here we show how to take advantage of information from parametric generative models. We focus on the bias in the information-theoretic error arising from finite sampling effects, and find an appropriate local metric that maximally reduces the bias based upon knowledge from generative models. As a byproduct, the asymptotic theoretical analysis in this work relates metric learning to dimensionality reduction from a novel perspective, which was not understood from previous discriminative approaches. Empirical experiments show that this learned local metric enhances the discriminative nearest neighbor performance on various datasets using simple class conditional generative models such as a Gaussian. Yung-Kyun Noh, Byoung-Tak Zhang, Daniel D. Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Extremely Sparse Deep Learning Using Inception Modules with DropfiltersabstractThis paper reports a successful application of highly sparse convolutional network model for offline handwritten character recognition. The model makes use of spatial dropout techniques named dropfilters for sparsifying the inception modules in GoogLeNet, resulting in extremely sparse deep networks. The model is industry-deployable regarding model size and performance, which trained by a handwritten dataset of 520 classes and 260,000 Hangul(Korean) characters for tablet PCs and smartphones. The proposed model obtained significant improvement in recognition performance while the number of parameters is much smaller than that of the LeNet, a classical sparse convolutional network. We also evaluated the dropfiltered inception networks on the handwritten Hangul dataset and achieved 3.275% higher recognition accuracy with approximately three times fewer parameters than a deep network based on LeNet structure without dropfilters. Woo-Young Kang, Kyung-Wha Park, Byoung-Tak Zhang |
ICDAR | 3 |
| 2017 | Hadamard Product for Low-rank Bilinear Pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha 0001, Byoung-Tak Zhang |
ICLR (Poster) | 6 |
| 2017 | DeepStory: Video Story QA by Deep Embedded Memory NetworksabstractQuestion-answering (QA) on video contents is a significant challenge for achieving human-level intelligence as it involves both vision and language in real-world settings. Here we demonstrate the possibility of an AI agent performing video story QA by learning from a large amount of cartoon videos. We develop a video-story learning model, i.e. Deep Embedded Memory Networks (DEMN), to reconstruct stories from a joint scene-dialogue video stream using a latent embedding space of observed data. The video stories are stored in a long-term memory component. For a given question, an LSTM-based attention model uses the long-term memory to recall the best question-story-answer triplet by focusing on specific words containing key information. We trained the DEMN on a novel QA dataset of children’s cartoon video series, Pororo. The dataset contains 16,066 scene-dialogue pairs of 20.5-hour videos, 27,328 fine-grained sentences for scene description, and 8,913 story-related QA pairs. Our experimental results show that the DEMN outperforms other QA models. This is mainly due to 1) the reconstruction of video stories in a scene-dialogue combined form that utilize the latent embedding and 2) attention. DEMN also achieved state-of-the-art results on the MovieQA benchmark. Min-Oh Heo, Seongho Choi 0001, Byoung-Tak Zhang |
IJCAI | 4 |
| 2017 | Overcoming Catastrophic Forgetting by Incremental Moment MatchingabstractCatastrophic forgetting is a problem of neural networks that loses the information of the first task after training the second task. Here, we propose a method, i.e. incremental moment matching (IMM), to resolve this problem. IMM incrementally matches the moment of the posterior distribution of the neural network which is trained on the first and the second task, respectively. To make the search space of posterior parameter smooth, the IMM procedure is complemented by various transfer learning techniques including weight transfer, L2-norm of the old and the new parameter, and a variant of dropout with the old parameter. We analyze our approach on a variety of datasets including the MNIST, CIFAR-10, Caltech-UCSD-Birds, and Lifelog datasets. The experimental results show that IMM achieves state-of-the-art performance by balancing the information between an old and a new network. Sang-Woo Lee 0001, Jin-Hwa Kim, Jaehyun Jun, Jung-Woo Ha 0001, Byoung-Tak Zhang |
NIPS | 5 |
| 2017 | Dual-memory neural networks for modeling cognitive activities of humans via wearable sensors
Sang-Woo Lee 0001, Chung-Yeon Lee, Donghyun Kwak, Jung-Woo Ha 0001, Jeonghee Kim, Byoung-Tak Zhang |
Neural Networks | 6 |
| 2016 | DeepSchema: Automatic Schema Acquisition from Wearable Sensor Data in Restaurant Situations
Eun-Sol Kim, Kyoung-Woon On, Byoung-Tak Zhang |
IJCAI | 3 |
| 2016 | Dual-Memory Deep Learning Architectures for Lifelong Learning of Everyday Human Behaviors
Sang-Woo Lee 0001, Chung-Yeon Lee, Donghyun Kwak, Jeonghee Kim, Byoung-Tak Zhang |
IJCAI | 6 |
| 2016 | Multimodal Residual Learning for Visual QAabstractDeep neural networks continue to advance the state-of-the-art of image recognition tasks with various methods. However, applications of these methods to multimodality remain limited. We present Multimodal Residual Networks (MRN) for the multimodal residual learning of visual question-answering, which extends the idea of the deep residual learning. Unlike the deep residual learning, MRN effectively learns the joint representation from visual and language information. The main idea is to use element-wise multiplication for the joint residual mappings exploiting the residual learning of the attentional models in recent studies. Various alternative models introduced by multimodality are explored based on our study. We achieve the state-of-the-art results on the Visual QA dataset for both Open-Ended and Multiple-Choice tasks. Moreover, we introduce a novel method to visualize the attention effect of the joint representations for each learning block using back-propagation algorithm, even though the visual features are collapsed without spatial information. Jin-Hwa Kim, Sang-Woo Lee 0001, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha 0001, Byoung-Tak Zhang |
NIPS | 7 |
| 2016 | Special issue: First International Conference on Big Data and Smart Computing (BigComp2014)
James T. Kwok, Kazutoshi Sumiya, Byoung-Tak Zhang |
Data Knowl. Eng. | 4 |
| 2015 | Automated Construction of Visual-Linguistic Knowledge via Concept Learning from Cartoon VideosabstractLearning mutually-grounded vision-language knowledge is a foundational task for cognitive systems and human-level artificial intelligence. Most of knowledge-learning techniques are focused on single modal representations in a static environment with a fixed set of data. Here, we explore an ecologically more-plausible setting by using a stream of cartoon videos to build vision-language concept hierarchies continuously. This approach is motivated by the literature on cognitive development in early childhood. We present the model of deep concept hierarchy (DCH) that enables the progressive abstraction of concept knowledge in multiple levels. We develop a stochastic method for graph construction, i.e. a graph Monte Carlo algorithm, to search efficiently the huge compositional space of the vision-language concepts. The concept hierarchies are built incrementally and can handle concept drift, allowing for being deployed in lifelong learning environments. Using a series of approximately 200 episodes of educational cartoon videos we demonstrate the emergence and evolution of the concept hierarchies as the video stories unfold. We also present the application of the deep concept hierarchies for context-dependent translation between vision and language, i.e. the transcription of a visual scene into text and the generation of visual imagery from text. Jung-Woo Ha 0001, Byoung-Tak Zhang |
AAAI | 3 |
| 2015 | Social Network Analysis of TV Drama Characters via Deep Concept HierarchiesabstractTV drama is a kind of big data, containing enormous knowledge of modern human society. As the character-centered stories unfold, diverse knowledge, such as economics, politics and the culture, is displayed. However, unless we have efficient dynamic multi-modal data processing and picture processing methods, we cannot analyze drama data effectively. Here, we adopt the recently proposed deep concept hierarchies (DCH) and convolutional-recursive neural network (C-RNN) models to analyze the social network between the drama characters. DCH uses multi hierarchies structure to translate the vision-language concepts of drama characters into diversified abstract concepts, and utilizes Markov Chain Monte Carlo algorithm to improve the retrieval efficiency of organizing conceptual spaces. Adopting approximately 4400-minute data of TV drama - Friends, we process face recognition on the characters by using convolutional-recursive deep learning model. Then we establish the social network between the characters by deep concept hierarchies model and analyze their affinity and the change of social network while the stories unfold. Chang-Jun Nan, Byoung-Tak Zhang |
ASONAM | 3 |
| 2015 | Analyzing Human Behavioral Data to Interact with Restaurant Server AgentsabstractIn this paper, we consider a problem of analyzing human behavioral data to predict the human cognitive states and generate corresponding actions of sever-agent. Specifically, we aim at predicting human cognitive states during meal time and generating relevant dining services for the human. For this study, we collect behavioral data using 2 kinds of wearable devices, which are an eye tracker and a watch type EDA device, during meal time. We focus on the characteristics of the behavioral data, which are heterogeneous, noisy and temporal, and suggest a novel machine learning algorithm which can analyze the data integrally. Suggested model has hierarchical structure: the bottom layer combines the multi-modal behavioral data based on causal structure of the data and extracts the feature vector. Using the extracted feature vectors, the upper layer predicts the cognitive states based on temporal correlation between feature vectors. Experimental results show that the suggested model can analyze the behavioral data efficiently and predict the human cognitive states correctly. Eun-Sol Kim, Kyoung-Woon On, Byoung-Tak Zhang |
HAI | 3 |
| 2015 | Preface
Jin-Woo Jung, Hiroshi Wakuya, Byoung-Tak Zhang |
Soft Comput. | 3 |
| 2015 | Consensus Analysis and Modeling of Visual Aesthetic PerceptionabstractThis paper reports a characteristic relation between skewness and kurtosis of aesthetic score distributions in a massive photo aesthetics dataset generated from online voting. Analysis results reveal an unexpectedly wide range of kurtosis in the mediocre photo group, asymmetric consensus, the 4/3 power-law regime in both extremes, and tag-specific relation in the skewness-kurtosis plane. From the human cognition perspective on affective content analysis, these patterns are interpreted as supporting the necessity of a consensus property in addition to the preference used so far for accurate modeling of aesthetic evaluation process in human mind. For explaining the observed patterns, we propose a new computational model of a dynamic system based on the interaction between multiple attractors. Characteristic patterns in response time and consensus are predicted from the proposed model and observed in the experiments with human subjects for model validation. Tae-Suh Park, Byoung-Tak Zhang |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | Effective EEG Connectivity Analysis of Episodic Memory Retrieval
Chung-Yeon Lee, Byoung-Tak Zhang |
CogSci | 2 |
| 2014 | Bayesian evolutionary hypergraph learning for predicting cancer clinical outcomes
Soo-Jin Kim, Jung-Woo Ha 0001, Byoung-Tak Zhang |
J. Biomed. Informatics | 3 |
| 2013 | Evolutionary concept learning from cartoon videos by multimodal hypernetworksabstractConcepts have been widely used for categorizing and representing knowledge in artificial intelligence. Previous researches on concept learning have focused on unimodal data, usually on linguistic domains in a static environment. Concept learning from multimodal stream data, such as videos, remains a challenge due to their dynamic change and high-dimensionality. Here we propose an evolutionary method that simulates the process of human concept learning from multimodal video streams. Two key ideas on evolutionary concept learning are representing concepts in a large collection (population) of hyperedges or a hypergraph and to incrementally learning from video streams based on an evolutionary approach. The hypergraph is learned "evolutionarily" by repeating the generation and selection process of hyperedge concepts from the video data. The advantage of this evolutionary learning process is that the population-based distributed coding allows flexible and robust trace of the change of concept relations as the video story unfolds. We evaluate the proposed method on a suite of children's cartoon videos for 517 minutes of total playing time. Experimental results show that the proposed method effectively represents visual-textual concept relations and our evolutionary concept learning method effectively models the conceptual change as an evolutionary process. We also investigate the structure properties of the constructed concept networks. Beom-Jin Lee, Jung-Woo Ha 0001, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 4 |
| 2013 | Estimating Multiple Evoked Emotions from Videos
Wonhee Choe, Hyo-Sun Chun, Junhyug Noh, Seong-Deok Lee, Byoung-Tak Zhang |
CogSci | 5 |
| 2013 | Online Incremental Structure Learning of Sum-Product Networks
Sang-Woo Lee 0001, Min-Oh Heo, Byoung-Tak Zhang |
ICONIP (2) | 3 |
| 2013 | Online learning of low dimensional strategies for high-level push recovery in bipedal humanoid robotsabstractBipedal humanoid robots will fall under unforeseen perturbations without active stabilization. Humans use dynamic full body behaviors in response to perturbations, and recent bipedal robot controllers for balancing are based upon human biomechanical responses. However these controllers rely on simplified physical models and accurate state information, making them less effective on physical robots in uncertain environments. In our previous work, we have proposed a hierarchical control architecture that learns from repeated trials to switch between low-level biomechanically-motivated strategies in response to perturbations. However in practice, it is hard to learn a complex strategy from limited number of trials available with physical robots. In this work, we focus on the very problem of efficiently learning the high-level push recovery strategy, using simulated models of the robot with different levels of abstraction, and finally the physical robot. From the state trajectory information generated using different models and a physical robot, we find a common low dimensional strategy for high level push recovery, which can be effectively learned in an online fashion from a small number of experimental trials on a physical robot. This learning approach is evaluated in physics-based simulations as well as on a small humanoid robot. Our results demonstrate how well this method stabilizes the robot during walking and whole body manipulation tasks. Seung-Joon Yi, Byoung-Tak Zhang, Dennis W. Hong, Daniel D. Lee |
ICRA | 2 |
| 2012 | Evolutionary particle filtering for sequential dependency learning from video dataabstractWe describe a novel learning scheme for hidden dependencies in video streams. The proposed scheme aims to transform a given sequential stream into a dependency structure of particle populations. Each particle population summarizes an associated segment. The novel point of the proposed scheme is that both of dependency learning and segment summarization are performed in an unsupervised online manner without assuming priors. The proposed scheme is executed in two-stage learning. At the first stage, a segment corresponding to a common dominant image is estimated using evolutionary particle filtering. Each dominant image is depicted based on combinations of image descriptors. Prevailing features of a dominant image are selected through evolution. Genetic operators introduce the essential diversity preventing sample impoverishment. At the second stage, transitional probability between the estimated segments is computed and stored. The proposed scheme is applied to extract dependencies in an episode of a TV drama. We demonstrate performance by comparing to human estimations. Jun Hee Yoo, Ho-Sik Seok, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Hierarchical Slow-Feature Models of Gesture Conversation
Jiseob Kim, Sooyong Jang, Eun-Sol Kim, Byoung-Tak Zhang |
CogSci | 4 |
| 2012 | 'Is this right?' or 'Is that wrong?': Evidence from Dynamic Eye-Hand Movement in Decision Making
Eun-Sol Kim, Jiseob Kim, Thies Pfeiffer, Ipke Wachsmuth, Byoung-Tak Zhang |
CogSci | 5 |
| 2012 | Neural Correlates of Episodic Memory Formation in Audio-Visual Pairing Tasks
Chung-Yeon Lee, Beom-Jin Lee, Joon Shik Kim, Byoung-Tak Zhang |
CogSci | 4 |
| 2012 | Effect of Saliency-Based Masking in Scene Classification
Tae-Suh Park, Byoung-Tak Zhang |
CogSci | 2 |
| 2012 | Sparse Population Code Models of Word Learning in Concept Drift
Byoung-Tak Zhang, Jung-Woo Ha 0001, Myunggu Kang |
CogSci | 1 |
| 2012 | Active stabilization of a humanoid robot for impact motions with unknown reaction forcesabstractDuring heavy work, humans utilize whole body motions in order to generate large forces. In extreme cases, exaggerated weight shifts are used to impart large impact forces. There have been approaches to design stable whole body impact motions based on precise dynamic models of the robot and the target object, but they have practical limitations as the uncertainty in the ensuing reaction forces can lead to instability. In the current work, we describe a motion controller for a humanoid robot that generates impacts at an end effector while keeping the robot body balanced before and after the impact. Instead of relying on the accuracy of the impact dynamics model, we use a simplified model of the robot and biomechanically motivated push recovery controllers to reactively stabilize the robot against unknown perturbations from the impact. We demonstrate our approach in physically realistic simulations, as well as experimentally on a small humanoid robot platform. Seung-Joon Yi, Byoung-Tak Zhang, Dennis W. Hong, Daniel D. Lee |
IROS | 2 |
| 2012 | Text-to-image retrieval based on incremental association via multimodal hypernetworksabstractText-to-image retrieval is to retrieve the images associated with the textual queries. A text-to-image retrieval model requires an incremental learning method for its practical use since the multimodal data grow up dramatically. Here we propose an incremental text-to-image retrieval method using a multimodal association model. The association model is based on a hypernetwork (HN) where a vertex corresponds to a textual word or a visual patch and a hyperedge represents a higher-order multimodal association. Using the HN incrementally learned by a sequential Bayesian sampling, in the multimodal hypernetwork-based text-to-image retrieval, a given text query is crossmodally expanded to the visual query and then similar images are retrieved to the expanded visual query. We evaluated the proposed method using 3,000 images with textual description from Flickr.com. The experimental results present that the proposed method achieves very competitive retrieval performances compared to a baseline method. Moreover, we demonstrate that our method provides robust text-to-image retrieval results for the increasing data. Jung-Woo Ha 0001, Beom-Jin Lee, Byoung-Tak Zhang |
SMC | 3 |
| 2012 | A probabilistic coevolutionary biclustering algorithm for discovering coherent patterns in gene expression datasetabstractBACKGROUND: Biclustering has been utilized to find functionally important patterns in biological problem. Here a bicluster is a submatrix that consists of a subset of rows and a subset of columns in a matrix, and contains homogeneous patterns. The problem of finding biclusters is still challengeable due to computational complex trying to capture patterns from two-dimensional features. RESULTS: We propose a Probabilistic COevolutionary Biclustering Algorithm (PCOBA) that can cluster the rows and columns in a matrix simultaneously by utilizing a dynamic adaptation of multiple species and adopting probabilistic learning. In biclustering problems, a coevolutionary search is suitable since it can optimize interdependent subcomponents formed of rows and columns. Furthermore, acquiring statistical information on two populations using probabilistic learning can improve the ability of search towards the optimum value. We evaluated the performance of PCOBA on synthetic dataset and yeast expression profiles. The results demonstrated that PCOBA outperformed previous evolutionary computation methods as well as other biclustering methods. CONCLUSIONS: Our approach for searching particular biological patterns could be valuable for systematically understanding functional relationships between genes and other biological components at a genome-wide level. Je-Gun Joung, Soo-Jin Kim, Soo-Yong Shin, Byoung-Tak Zhang |
BMC Bioinform. | 4 |
| 2011 | Mutual information-based evolution of hypernetworks for brain data analysisabstractCortical analysis becomes increasingly important for brain research and clinical diagnosis. This problem involves a combinatorial search to find the essential modules among a large number of brain regions. Despite several statistical approaches, cortical analysis remains a formidable challenge due to high dimensionality and sparsity of data. Here we describe an evolutionary method for finding significant modules from cortical data. The method uses a hypernetwork which is encoded as a population of hyperedges, where hyperedges represent building blocks or potential modules. We develop an efficient method for evolving the hypernetwork using mutual information to generate essential hyperedges. We evaluate the method on predicting intelligence quotient (IQ) levels and finding potential significant modules on IQ from brain MRI data consisting of 62 healthy adults with over 80,000 measured points (variables). The experimental results show that our information-theoretic evolutionary hypernetworks improve the classification accuracy by 5-15%. Moreover, it extracts significant cortical modules that distinguish high IQ from low IQ groups. Eun-Sol Kim, Jung-Woo Ha 0001, Wi Hoon Jung, Joon Hwan Jang, Jun Soo Kwon, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 6 |
| 2011 | A molecular evolutionary algorithm for learning hypernetworks on simulated DNA computersabstractAbstract—We describe a “molecular ” evolutionary algorithm that can be implemented in DNA computing in vitro to learn the recently-proposed hypernetwork model of cognitive memory. The molecular learning process is designed to make it possible to perform wet-lab experiments using DNA molecules and bio-lab tools. We present the bio-experimental protocols for selection, amplification and mutation operators for evolving hypernetworks. We analyze the convergence properties of the molecular evolutionary algorithms on simulated DNA computers. The performance of the algorithms is demonstrated on the task of simulating the cognitive process of learning a language model from a drama corpus to identify the style of an unknown drama. We also discuss other applications of the molecular evolutionary algorithms. In addition to their feasibility in DNA computing, which opens a new horizon of in vitro evolutionary computing, the molecular evolutionary algorithm provides unique properties that are distinguished from conventional evolutionary algorithms and makes a new addition to the arsenal of tools in evolutionary computation. Keywords-molecular evolutionary algorithms, molecular evolutionary learning; DNA computing; hypernetwork; cognitive memory simualtion I. Bado Lee, Joon Shik Kim, Russell J. Deaton, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 5 |
| 2011 | Evolving a population code for multimodal concept learningabstractWe describe an evolutionary method for learning concepts of objects from multimodal data. The proposed method uses a population code (hypernetwork representation), i.e. a col lection of codewords (hyperedges) and associated weights, which is adapted by evolutionary computation based on observations of positive and negative examples. The goal of evolution is to find the best compositions and weights of hyperedges to estimate the underlying distribution of the target concepts. We discuss the relationship of this method with estimation of distribution algorithms (EDAs), classifier systems, and ensemble learning methods. We evaluate the method on a suite of image/text benchmarks. The experimental results demonstrate that the evolutionary process successfully discovers salient codewords representing multi-modal feature combinations for describing and distinguishing different concepts. We also analyze how the complexity of the population code evolves as learning proceeds. Bado Lee, Ho-Sik Seok, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2011 | Learning full body push recovery control for small humanoid robotsabstractDynamic bipedal walking is susceptible to external disturbances and surface irregularities, requiring robust feedback control to remain stable. In this work, we present a practical hierarchical push recovery strategy that can be readily implemented on a wide range of humanoid robots. Our method consists of low level controllers that perform simple, biomechanically motivated push recovery actions and a high level controller that combines the low level controllers according to proprioceptive and inertial sensory signals and the current robot state. Reinforcement learning is used to optimize the parameters of the controllers in order to maximize the stability of the robot over a broad range of external disturbances. The controllers are learned on a physical simulation and implemented on the Darwin-HP humanoid robot platform, and the resulting experiments demonstrate effective full body push recovery behaviors during dynamic walking. Seung-Joon Yi, Byoung-Tak Zhang, Dennis W. Hong, Daniel D. Lee |
ICRA | 2 |
| 2011 | Practical bipedal walking control on uneven terrain using surface learning and push recoveryabstractBipedal walking in human environments is made difficult by the unevenness of the terrain and by external disturbances. Most approaches to bipedal walking in such environments either rely upon a precise model of the surface or special hardware designed for uneven terrain. In this paper, we present an alternative approach to stabilize the walking of an inexpensive, commercially-available, position-controlled humanoid robot in difficult environments. We use electrically compliant swing foot dynamics and onboard sensors to estimate the inclination of the local surface, and use a online learning algorithm to learn an adaptive surface model. Perturbations due to external disturbances or model errors are rejected by a hierarchical push recovery controller, which modulates three biomechanically motivated push recovery controllers according to the current estimated state. We use a physically realistic simulation with an articulated robot model and reinforcement learning algorithm to train the push recovery controller, and implement the learned controller on a commercial DARwIn-OP small humanoid robot. Experimental results show that this combined approach enables the robot to walk over unknown, uneven surfaces without falling down. Seung-Joon Yi, Byoung-Tak Zhang, Dennis W. Hong, Daniel D. Lee |
IROS | 2 |
| 2011 | Ensemble Learning with Active Example Selection for Imbalanced Biomedical Data ClassificationabstractIn biomedical data, the imbalanced data problem occurs frequently and causes poor prediction performance for minority classes. It is because the trained classifiers are mostly derived from the majority class. In this paper, we describe an ensemble learning method combined with active example selection to resolve the imbalanced data problem. Our method consists of three key components: 1) an active example selection algorithm to choose informative examples for training the classifier, 2) an ensemble learning method to combine variations of classifiers derived by active example selection, and 3) an incremental learning scheme to speed up the iterative training procedure for active example selection. We evaluate the method on six real-world imbalanced data sets in biomedical domains, showing that the proposed method outperforms both the random under sampling and the ensemble with under sampling methods. Compared to other approaches to solving the imbalanced data problem, our method excels by 0.03-0.15 points in AUC measure. Sangyoon Oh 0001, Minsu Lee 0001, Byoung-Tak Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2011 | Feature Relevance Network-Based Transfer Learning for Indoor Location EstimationabstractWe present a new machine learning framework for indoor location estimation. In many cases, locations could be easily estimated using various traditional positioning methods and conventional machine learning approaches based on signalling devices, e.g., access points (APs). When there exist environmental changes, however, such traditional methods cannot be employed due to data distribution change. In order to circumvent this difficulty, we introduce feature relevance network-based method, which focuses on interrelatedness among features. Feature relevance networks are connected graphs representing concurrency of the signalling devices such as APs. In the newly created relevance network, a test instance and the prototype of a location are expanded until convergence. The expansion cost corresponds to distance between the test instance and the prototype. Unlike other methods, our model is nonparametric making no assumptions about signal distributions. The proposed method is applied to the 2007 IEEE International Conference on Data Mining Data Mining Contest Task #2 (transfer learning), which is a typical example situation where the training and test datasets have been gathered during different periods. Using the proposed method, we accomplish the estimation accuracy of 0.3238, which is better than the best result of the contest. Ho-Sik Seok, Kyu-Baek Hwang, Byoung-Tak Zhang |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2010 | Online Learning of Uneven Terrain for Humanoid Bipedal WalkingabstractWe present a novel method to control a biped humanoid robot to walk on unknown inclined terrains, using an online learning algorithm to estimate in real-time the local terrain from proprioceptive and inertial sensors. Compliant controllers for the ankle joints are used to actively probe the surrounding surface, and the measured sensor data are combined to explicitly learn the global inclination and local disturbances of the terrain. These estimates are then used to adaptively modify the robot locomotion and control parameters. Results from both a physically-realistic computer simulation and experiments on a commercially available small humanoid robot show that our method can rapidly adapt to changing surface conditions to ensure stable walking on uneven surfaces. Seung-Joon Yi, Byoung-Tak Zhang, Daniel D. Lee |
AAAI | 2 |
| 2010 | Subjective Document Classification Using Network AnalysisabstractNetwork analysis methods have been applied in many areas such as computer science, social science, biology and physics. In this paper, we apply network analysis methods to the linguistic domain for classifying subjective documents. Particularly, we view that subjective documents are related to one another according to some common subjective words and build a subjective document network of which nodes are documents and of which links represent the similarity between two documents. In addition, we consider that adjectives and adverbs are the two representatives carrying sentimental polarities among parts-of-speeches, and perform experiments for three cases, using adjectives only, adverbs only, and both adjectives and adverbs together. In conclusion, this paper proposes a new method to the subjective document classification problem by applying network analysis methods without requiring linguistic domain knowledge and suggests the possibility of detecting themes among documents rather than binary classification. Minkyoung Kim, Byoung-Tak Zhang, June-Sup Lee |
ASONAM | 2 |
| 2010 | Evolutionary layered hypernetworks for identifying microRNA-mRNA regulatory modulesabstractExploring micro RNA (miRNA) and mRNA regulatory interactions may give new insights into diverse biological phenomena. While elucidating complex miRNA-mRNA interactions has been studied with experimental and computational approaches, it is still difficult to infer miRNA-mRNA regulatory modules. Here we present a novel method for identifying functional miRNA-mRNA modules from heterogeneous expression data. The proposed approach is layered hypernetworks consisting of two layers which are the layer of modality-dependent hypernetworks and of an integrating hypernetwork. The layered hypernetwork model is suitable for detecting relationships between heterogeneous modalities. Applied to the analysis of miRNA and mRNA expression profiles on multiple human cancers, the proposed model identifies oncogenic miRNA-mRNA regulatory modules. The experimental results show that our method provides a competitive performance to support vector machines, and outperforms other standard machine learning algorithms. The biological significance of the discovered miRNA-mRNA modules were validated by literature reviews. Soo-Jin Kim, Jung-Woo Ha 0001, Bado Lee, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 4 |
| 2010 | Generative Local Metric Learning for Nearest Neighbor ClassificationabstractWe consider the problem of learning a local metric to enhance the performance of nearest neighbor classification. Conventional metric learning methods attempt to separate data distributions in a purely discriminative manner; here we show how to take advantage of information from parametric generative models. We focus on the bias in the information-theoretic error arising from finite sampling effects, and find an appropriate local metric that maximally reduces the bias based upon knowledge from generative models. As a byproduct, the asymptotic theoretical analysis in this work relates metric learning with dimensionality reduction, which was not understood from previous discriminative approaches. Empirical experiments show that this learned local metric enhances the discriminative nearest neighbor performance on various datasets using simple class conditional generative models. Yung-Kyun Noh, Byoung-Tak Zhang, Daniel D. Lee |
NIPS | 2 |
| 2010 | MMG: A Learning Game Platform for Understanding and Predicting Human Recall Memory
Umer Fareed, Byoung-Tak Zhang |
PKAW | 2 |
| 2010 | Layered Hypernetwork Models for Cross-Modal Associative Text and Image Keyword Generation in Multimodal Information Retrieval
Jung-Woo Ha 0001, Byoung-Hee Kim, Bado Lee, Byoung-Tak Zhang |
PRICAI | 4 |
| 2010 | Visual Query Expansion via Incremental Hypernetwork Models of Image and Text
Min-Oh Heo, Myunggu Kang, Byoung-Tak Zhang |
PRICAI | 3 |
| 2009 | Dynamic and Static Influence Models on Starbucks NetworksabstractIn this paper, we consider Starbucks stores in Korea, which have been spread most rapidly in the world, as an exemplary social network. We define two kinds of influence models and analyze the diffusive power of each node in a network. One isDynamic Influence Modelfor time-varying diffusion analysis and the other isStatic Influence Modelfor cumulative spreading power analysis. The physical features of the most influential stores of the simulation results show that our social influence models are convincing. Minkyoung Kim, Byoung-Tak Zhang, June-Sup Lee |
ASONAM | 2 |
| 2009 | AESNB: Active Example Selection with Naive Bayes Classifier for Learning from Imbalanced Biomedical DataabstractVarious real-world biomedical classification tasks suffer from the imbalanced data problem which tends to make the prediction performance of some classes significantly decrease. In this paper, we present an active example selection method with naiumlve Bayes classifier (AESNB) as a solution for the imbalanced data problem. The proposed method starts with a small balanced subset of training examples. A naive Bayes classifier is trained incrementally by actively selecting and adding informative examples regardless of the original class distribution. Informative examples are defined as examples that produce high error scores by the current classifier. We examined the performance of AESNB algorithm by using five imbalanced biomedical datasets. Our experimental results show that the naiumlve Bayes classifier with our active example selection method achieves a competitive classification performance compared to the classifier with sampling or cost-sensitive methods. Minsu Lee 0001, Je-Keun Rhee, Byoung-Hee Kim, Byoung-Tak Zhang |
BIBE | 4 |
| 2009 | Ensemble Learning Based on Active Example Selection for Solving Imbalanced Data Problem in Biomedical DataabstractThe imbalanced data problem is popular in biomedical classification tasks. Since trained classifiers using imbalanced data are mostly derived from the majority class, their prediction performance is poor for the minority class. In this paper, we propose a novel ensemble learning method based on an active example selection algorithm to resolve the imbalanced data problem. To compensate a possible sub-optimal classifier, our proposed ensemble learning methods aggregates classifiers built by the active example selection algorithm. We implement this ensemble learning method based on the active example selection algorithm using incremental naive Bayes classifiers. Our empirical results show that we greatly improve the performance of classification models trained by five real world imbalanced biomedical data. The proposed ensemble learning methods outperforms other approaches by 0.03~0.15 in terms of AUC which solve imbalanced data problem. Minsu Lee 0001, Sangyoon Oh 0001, Byoung-Tak Zhang |
BIBM | 3 |
| 2009 | Evolving hypernetwork models of binary time series for forecasting price movements on stock marketsabstractThe paper proposes a hypernetwork-based method for stock market prediction through a binary time series problem. Hypernetworks are a random hypergraph structure of higher-order probabilistic relations of data. The problem we tackle concerns the prediction of price movements (up/down) on stock markets. Compared to previous approaches, the proposed method discovers a large population of variable subpatterns, i.e. local and global patterns, using a novel evolutionary hypernetwork. An output is obtained from combining these patterns. In the paper, we describe two methods for assessing the prediction quality of the hypernetwork approach. Applied to the Dow Jones Industrial Average Index and the Korea Composite Stock Price Index data, the experimental results show that the proposed method effectively learns and predicts the time series information. In particular, the hypernetwork approach outperforms other machine learning methods such as support vector machines, naive Bayes, multilayer perceptrons, and k-nearest neighbors. Elena Bautu, Sun Kim, Andrei Bautu, Henri Luchian, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 5 |
| 2009 | Gender classification with cortical thickness measurement from magnetic resonance imaging by using a feature selection method based on evolutionary hypernetworksabstractHypernetworks are a weighted hypergraph where evolutionary methods are learning the model structure and parameters. The evolutionary methods enable the hypernetwork model to conserve significant features implicitly during the learning process. In this study, we propose a novel feature selection method based on occurrence frequencies of attributes in hyperedges by analyzing the structure of a hypernetwork. We also apply the evolutionary hypernetwork with the proposed feature selection method to the gender classification based on cortical thickness measurement on healthy young adults from Magnetic Resonance Imaging (MRI). The experimental results show that the proposed selection method improves the classification accuracy by approximately 20%. Also, a comparative study on four classification algorithms and three feature selection methods shows that the hypernetwork model with the proposed feature selection method achieves a competitive classification performance. Jung-Woo Ha 0001, Joon Hwan Jang, Do-Hyung Kang, Wi Hoon Jung, Jun Soo Kwon, Byoung-Tak Zhang |
FUZZ-IEEE | 6 |
| 2009 | Evolutionary hypernetworks for learning to generate music from examplesabstractEvolutionary hypernetworks (EHNs) are recently introduced models for learning higher-order probabilistic relations of data by an evolutionary self-organizing process. We present a method that enables EHNs to learn and generate music from examples. Short-term and long-term sequential patterns can be extracted and combined to generate music with various styles by our method. Based on a music corpus consisting of several genres and artists, an EHN generates genre-specific or artist-dependent music fragments when a fraction of score is given as a cue. Our method shows about 88% of success rate in partial music completion task. By inspecting hyperedges in the trained hypernetworks, we can extract a set of arguments that constitutes melodic structures in music. Byoung-Hee Kim, Byoung-Tak Zhang |
FUZZ-IEEE | 3 |
| 2009 | Evolutionary hypernetwork classifiers for protein-proteininteraction sentence filteringabstractProtein-Protein Interaction (PPI) extraction, among ongoing biomedical text mining challenges, is becoming a topic in focus because of its crucial role in providing a starting point to understand biological processes. Machine learning (ML) techniques have been applied to extract the PPI information from biomedical literature. Although they have provided reasonable performance so far, more features are required for real use. In particular, many ML-approaches lack human understandability for learned models. Here, we propose a novel method for classifying PPI sentences. Our approach utilizes the modified hypernetwork model, a hypergraph with weighted hyperedges that are calibrated via an evolutionary learning method. The evolutionary hypernetwork memorizes fragments of training patterns while self-adjusting its own structure for detecting PPI sentences. For experiments, we show that our approach provides competitive performance compared to other ML methods. Apart from its superior classification performance, the evolving hypernetwork model comes with a highly interpretable structure. We show how significant PPI patterns can be naturally extracted from the learned model. We also analyze the discovered patterns. Jakramate Bootkrajang, Sun Kim, Byoung-Tak Zhang |
GECCO | 3 |
| 2009 | EvoOligo: Oligonucleotide Probe Design With Multiobjective Evolutionary AlgorithmsabstractProbe design is one of the most important tasks in successful deoxyribonucleic acid microarray experiments. We propose a multiobjective evolutionary optimization method for oligonucleotide probe design based on the multiobjective nature of the probe design problem. The proposed multiobjective evolutionary approach has several distinguished features, compared with previous methods. First, the evolutionary approach can find better probe sets than existing simple filtering methods with fixed threshold values. Second, the multiobjective approach can easily incorporate the user's custom criteria or change the existing criteria. Third, our approach tries to optimize the combination of probes for the given set of genes, in contrast to other tools that independently search each gene for qualifying probes. Lastly, the multiobjective optimization method provides various sets of probe combinations, among which the user can choose, depending on the target application. The proposed method is implemented as a platform called EvoOligo and is available for service on the web. We test the performance of EvoOligo by designing probe sets for 19 types of Human Papillomavirus and 52 genes in the Arabidopsis Calmodulin multigene family. The design results from EvoOligo are proven to be superior to those from well-known existing probe design tools, such as OligoArray and OligoWiz. Soo-Yong Shin, In-Hee Lee 0001, Kyung-Ae Yang, Byoung-Tak Zhang |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2008 | Cognitive learning and the multimodal memory game: Toward human-level machine learningabstractMachine learning has made great progress during the last decades and is being deployed in a wide range of applications. However, current machine learning techniques are far from sufficient for achieving human-level intelligence. Here we identify the properties of learners required for human-level intelligence and suggest a new direction of machine learning research, i.e. the cognitive learning approach, that takes into account the recent findings in brain and cognitive sciences. In particular, we suggest two fundamental principles to achieve human-level machine learning: continuity (forming a lifelong memory continuously) and glocality (organizing a plastic structure of localized micromodules connected globally). We then propose a multimodal memory game as a research platform to study cognitive learning architectures and algorithms, where the machine learner and two human players question and answer about the scenes and dialogues after watching the movies. Concrete experimental results are presented to illustrate the usefulness of the game and the cognitive learning framework for studying human-level learning and intelligence. Byoung-Tak Zhang |
IJCNN | 1 |
| 2008 | AptaCDSS-E: A classifier ensemble-based clinical decision support system for cardiovascular disease level prediction
Jae-Hong Eom, Sung-Chun Kim, Byoung-Tak Zhang |
Expert Syst. Appl. | 3 |
| 2007 | Finding Cancer-Related Gene Combinations Using a Molecular Evolutionary AlgorithmabstractHigh-throughput data such as microarrays make it possible to investigate the molecular-level mechanism of cancer more efficiently. Computational methods boost the microarray analysis by managing large and complex data systematically. However, combinatorial interactions among genes have not been considered as a unit of the analysis since previous methods mainly focus on a whole gene or a single isolated gene. Here, we introduce a molecular evolutionary algorithm called probabilistic library model (PLM). In the PLM, library elements are generated from gene combinations. An evolutionary procedure is adopted to learn the probabilistic distribution of training samples. We apply the PLM to prostate cancer microarray data. The experimental results show that the PLM classifiers perform better than conventional methods such as neural networks and decision trees in accuracy. We also examine the evolved library to find cancer-related gene combinations. Chan-Hoon Park, Soo-Jin Kim, Sun Kim, Dong-Yeon Cho, Byoung-Tak Zhang |
BIBE | 5 |
| 2007 | Evolving hypernetwork classifiers for microRNA expression profile analysisabstractHigh-throughput microarrays inform us on different outlooks of the molecular mechanisms underlying the function of cells and organisms. While computational analysis for the microarrays show good performance, it is still difficult to infer modules of multiple co-regulated genes. Here, we present a novel classification method to identify the gene modules associated with cancers from microarray data. The proposed approach is based on 'hypernetworks', a hypergraph model consisting of vertices and weighted hyperedges. The hypernetwork model is inspired by biological networks and its learning process is suitable for identifying interacting gene modules. Applied to the analysis of microRNA (miRNA) expression profiles on multiple human cancers, the hypernetwork classifiers identified cancer-related miRNA modules. The results show that our method performs better than decision trees and naive Bayes. The biological meaning of the discovered miRNA modules has been examined by literature search. Sun Kim, Soo-Jin Kim, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2007 | Evolving hypernetworks for pattern classificationabstractHypernetworks consist of a large number of hyperedges that represent higher-order features sampled from training patterns. Evolutionary algorithms have been used as a method for evolving hypernetworks. The order of a hyperedge is defined as the number of feature variables in the hyperedge and it is an important parameter of the hypernetwork model. Previous studies used fixed-order hyperedges which limit model spaces and, thus, the best performance achievable by hypernetworks. Here, we present a method for evolving variable-order hypernetwork models. To find the proper orders automatically, the fitness values are calculated for each hyperedge and the hyperedges with low fitness values are substituted by new hyperedges. The method was tested on three data sets from UCI machine learning repository. The results show that the evolutionary hypernetworks show classification accuracies comparable to those of other conventional algorithms, find appropriate orders of hyperedges automatically, and extract important rules in the hyperedges for the given pattern classification problems. Joo-Kyung Kim, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | Multiplex PCR Assay Design by Hybrid Multiobjective Evolutionary Algorithm
In-Hee Lee 0001, Soo-Yong Shin, Byoung-Tak Zhang |
EMO | 3 |
| 2007 | Discovery of microRNA-mRNA modules via population-based probabilistic learningabstractMOTIVATION: MicroRNAs (miRNAs) and mRNAs constitute an important part of gene regulatory networks, influencing diverse biological phenomena. Elucidating closely related miRNAs and mRNAs can be an essential first step towards the discovery of their combinatorial effects on different cellular states. Here, we propose a probabilistic learning method to identify synergistic miRNAs involving regulation of their condition-specific target genes (mRNAs) from multiple information sources, i.e. computationally predicted target genes of miRNAs and their respective expression profiles. RESULTS: We used data sets consisting of miRNA-target gene binding information and expression profiles of miRNAs and mRNAs on human cancer samples. Our method allowed us to detect functionally correlated miRNA-mRNA modules involved in specific biological processes from multiple data sources by using a balanced fitness function and efficient searching over multiple populations. The proposed algorithm found two miRNA-mRNA modules, highly correlated with respect to their expression and biological function. Moreover, the mRNAs included in the same module showed much higher correlations when the related miRNAs were highly expressed, demonstrating our method's ability for finding coherent miRNA-mRNA modules. Most members of these modules have been reported to be closely related with cancer. Consequently, our method can provide a primary source of miRNA and target sets presumed to constitute closely related parts of gene regulatory pathways. Je-Gun Joung, Kyu-Baek Hwang, Jin-Wu Nam, Soo-Jin Kim, Byoung-Tak Zhang |
Bioinform. | 5 |
| 2007 | A Global Minimization Algorithm Based on a Geodesic of a Lagrangian Formulation of Newtonian Dynamics
Joon Shik Kim, Jangmin O, Byoung-Tak Zhang |
Neural Process. Lett. | 4 |
| 2006 | Text Classifiers Evolved on a Simulated DNA ComputerabstractThe use of synthetic DNA molecules for computing provides various insights to evolutionary computation. A molecular computing algorithm to evolve DNA-encoded genetic patterns has been previously reported in [1], [2]. Here we improve on the previous work by studying the convergence behavior of the molecular evolutionary algorithm in the context of text classification problems. In particular, we study the error reduction behavior of the evolutionary learning algorithm, both theoretically and experimentally. The individuals represent decision lists of variable length and the whole population takes part in making probabilistic decisions. The evolutionary process is to change each individual towards correct classification of training data, which is based on an error minimization strategy. The evolved molecular classifiers show a performance competitive to the standard algorithms such as naïve Bayes and neural network classifiers on the data set we studied. The possibility of molecular implementation by use of DNA-encoded individuals combined with simple molecular operations on a very big population distinguishes this approach from other existing evolutionary algorithms. Sun Kim, Min-Oh Heo, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2006 | DNA Hypernetworks for Information Storage and Retrieval
Byoung-Tak Zhang, Joo-Kyung Kim |
DNA | 1 |
| 2006 | Prediction of Protein Interaction with Neural Network-Based Feature Association Rule Mining
Jae-Hong Eom, Byoung-Tak Zhang |
ICONIP (3) | 2 |
| 2006 | Learning Hierarchical Bayesian Networks for Large-Scale Data Analysis
Kyu-Baek Hwang, Byoung-Hee Kim, Byoung-Tak Zhang |
ICONIP (1) | 3 |
| 2006 | Mining Protein Interaction from Biomedical Literature with Relation Kernel Method
Jae-Hong Eom, Byoung-Tak Zhang |
ISNN (2) | 2 |
| 2006 | Identification of biochemical networks by S-tree based genetic programmingabstractMOTIVATION: Most previous approaches to model biochemical networks have focused either on the characterization of a network structure with a number of components or on the estimation of kinetic parameters of a network with a relatively small number of components. For system-level understanding, however, we should examine both the interactions among the components and the dynamic behaviors of the components. A key obstacle to this simultaneous identification of the structure and parameters is the lack of data compared with the relatively large number of parameters to be estimated. Hence, there are many plausible networks for the given data, but most of them are not likely to exist in the real system. RESULTS: We propose a new representation named S-trees for both the structural and dynamical modeling of a biochemical network within a unified scheme. We further present S-tree based genetic programming to identify the structure of a biochemical network and to estimate the corresponding parameter values at the same time. While other evolutionary algorithms require additional techniques for sparse structure identification, our approach can automatically assemble the sparse primitives of a biochemical network in an efficient way. We evaluate our algorithm on the dynamic profiles of an artificial genetic network. In 20 trials for four settings, we obtain the true structure and their relative squared errors are <5% regardless of releasing constraints about structural sparseness. In addition, we confirm that the proposed algorithm is robust within +/-10% noise ratio. Furthermore, the proposed approach ensures a reasonable estimate of a real yeast fermentation pathway. The comparatively less important connections with non-zero parameters can be detected even though their orders are below 10(-2). To demonstrate the usefulness of the proposed algorithm for real experimental biological data, we provide an additional example on the transcriptional network of SOS response to DNA damage in Escherichia coli. We confirm that the proposed algorithm can successfully identify the true structure except only one relation. Dong-Yeon Cho, Kwang-Hyun Cho, Byoung-Tak Zhang |
Bioinform. | 3 |
| 2006 | Identification of regulatory modules by co-clustering latent variable models: stem cell differentiationabstractMOTIVATION: An important issue in stem cell biology is to understand how to direct differentiation towards a specific cell type. To elucidate the mechanism, previous studies have focused on identifying the responsible gene regulators, which have, however, failed to provide a systemic view of regulatory modules. To obtain a unified description of the regulatory modules, we characterized major stem cell species by employing a co-clustering latent variable model (LVM). The LVM-based method allowed us to elucidate the cell type-specific transcription factors, using genomic sequences as well as expression profiles. RESULTS: We used a list of genes enriched in each of 21 stem cell subpopulations, and their upstream genomic sequences. The LVM-based study allowed us to uncover the regulatory modules for each stem cell cluster, e.g. GABP and E2F for the proliferation phase, and Ap2alpha and Ap2gamma for the quiescence phase. Furthermore, the identities of the stem cell clusters were well revealed by the constituent genes that were directly targeted by the modules. Consequently, our analytical framework was demonstrated to be useful through a detailed case study of stem cell differentiation and can be applied to problems with similar characteristics. Je-Gun Joung, Rho Hyun Seong, Byoung-Tak Zhang |
Bioinform. | 4 |
| 2006 | miTarget: microRNA target gene prediction using a support vector machineabstractBACKGROUND: MicroRNAs (miRNAs) are small noncoding RNAs, which play significant roles as posttranscriptional regulators. The functions of animal miRNAs are generally based on complementarity for their 5' components. Although several computational miRNA target-gene prediction methods have been proposed, they still have limitations in revealing actual target genes. RESULTS: We implemented miTarget, a support vector machine (SVM) classifier for miRNA target gene prediction. It uses a radial basis function kernel as a similarity measure for SVM features, categorized by structural, thermodynamic, and position-based features. The latter features are introduced in this study for the first time and reflect the mechanism of miRNA binding. The SVM classifier produces high performance with a biologically relevant data set obtained from the literature, compared with previous tools. We predicted significant functions for human miR-1, miR-124a, and miR-373 using Gene Ontology (GO) analysis and revealed the importance of pairing at positions 4, 5, and 6 in the 5' region of a miRNA from a feature selection experiment. We also provide a web interface for the program. CONCLUSION: miTarget is a reliable miRNA target gene prediction tool and is a successful application of an SVM classifier. Compared with previous tools, its predictions are meaningful by GO analysis and its performance can be improved given more training examples. Jin-Wu Nam, Je-Keun Rhee, Wha-Jin Lee, Byoung-Tak Zhang |
BMC Bioinform. | 5 |
| 2006 | Construction of phylogenetic trees by kernel-based comparative analysis of metabolic networksabstractBACKGROUND: To infer the tree of life requires knowledge of the common characteristics of each species descended from a common ancestor as the measuring criteria and a method to calculate the distance between the resulting values of each measure. Conventional phylogenetic analysis based on genomic sequences provides information about the genetic relationships between different organisms. In contrast, comparative analysis of metabolic pathways in different organisms can yield insights into their functional relationships under different physiological conditions. However, evaluating the similarities or differences between metabolic networks is a computationally challenging problem, and systematic methods of doing this are desirable. Here we introduce a graph-kernel method for computing the similarity between metabolic networks in polynomial time, and use it to profile metabolic pathways and to construct phylogenetic trees. RESULTS: To compare the structures of metabolic networks in organisms, we adopted the exponential graph kernel, which is a kernel-based approach with a labeled graph that includes a label matrix and an adjacency matrix. To construct the phylogenetic trees, we used an unweighted pair-group method with arithmetic mean, i.e., a hierarchical clustering algorithm. We applied the kernel-based network profiling method in a comparative analysis of nine carbohydrate metabolic networks from 81 biological species encompassing Archaea, Eukaryota, and Eubacteria. The resulting phylogenetic hierarchies generally support the tripartite scheme of three domains rather than the two domains of prokaryotes and eukaryotes. CONCLUSION: By combining the kernel machines with metabolic information, the method infers the context of biosphere development that covers physiological events required for adaptation by genetic reconstruction. The results show that one may obtain a global view of the tree of life by comparing the metabolic pathway structures using meta-level information rather than sequence information. This method may yield further information about biological evolution, such as the history of horizontal transfer of each gene, by studying the detailed structure of the phylogenetic tree constructed by the kernel-based method. Sok June Oh, Je-Gun Joung, Jeong Ho Chang, Byoung-Tak Zhang |
BMC Bioinform. | 4 |
| 2006 | Adaptive stock trading with dynamic asset allocation using reinforcement learning
Jangmin O, Jongwoo Lee, Jae Won Lee, Byoung-Tak Zhang |
Inf. Sci. | 4 |
| 2005 | A Kernel Method for MicroRNA Target Prediction Using Sensible Data and Position-Based Features
Jin-Wu Nam, Wha-Jin Lee, Byoung-Tak Zhang |
CIBCB | 4 |
| 2005 | Molecular programming: evolving genetic programs in a test tubeabstractWe present a molecular computing algorithm for evolving DNA-encoded genetic programs in a test tube. The use of synthetic DNA molecules combined with biochemical techniques for variation and selection allows for various possibilities for building novel evolvable hardware. Also, the possibility of maintaining a huge number of individuals and their massively parallel manipulation allows us to make robust decisions by the "molecular" genetic programs evolved within a single population. We evaluate the potentials of this "molecular programming" approach by solving a medical diagnosis problem on a simulated DNA computer. Here the individual genetic program represents a decision list of variable length and the whole population takes part in making probabilistic decisions. Tested on a real-life leukemia diagnosis data, the evolved molecular genetic programs showed a comparable performance to decision trees. The molecular evolutionary algorithm can be adapted to solve problems in bio-technology and nano-technology where the physico-chemical evolution of target molecules is of pressing importance. Byoung-Tak Zhang, Ha-Young Jang |
GECCO | 1 |
| 2005 | Prediction of Yeast Protein-Protein Interactions by Neural Feature Association Rule
Jae-Hong Eom, Byoung-Tak Zhang |
ICANN (2) | 2 |
| 2005 | Extraction of Gene/Protein Interaction from Text Documents with Relation Kernel
Jae-Hong Eom, Byoung-Tak Zhang |
KES (2) | 2 |
| 2005 | CrossChip: a system supporting comparative analysis of different generations of Affymetrix arraysabstractSUMMARY: To increase compatibility between different generations of Affymetrix GeneChip arrays, we propose a method of filtering probes based on their sequences. Our method is implemented as a web-based service for downloading necessary materials for converting the raw data files (*.CEL) for comparative analysis. The user can specify the appropriate level of filtering by setting the criteria for the minimum overlap length between probe sequences and the minimum number of usable probe pairs per probe set. Our website supports a within-species comparison for human and mouse GeneChip arrays. AVAILABILITY: http://www.crosschip.org Sek Won Kong, Kyu-Baek Hwang, Richard D. Kim, Byoung-Tak Zhang, Steven A. Greenberg, Isaac S. Kohane, Peter J. Park |
Bioinform. | 4 |
| 2005 | Multiobjective evolutionary optimization of DNA sequences for reliable DNA computingabstractDNA computing relies on biochemical reactions of DNA molecules and may result in incorrect or undesirable computations. Therefore, much work has focused on designing the DNA sequences to make the molecular computation more reliable. Sequence design involves with a number of heterogeneous and conflicting design criteria and traditional optimization methods may face difficulties. In this paper, we formulate the DNA sequence design as a multiobjective optimization problem and solve it using a constrained multiobjective evolutionary algorithm (EA). The method is implemented into the DNA sequence design system, NACST/Seq, with a suite of sequence-analysis tools to help choose the best solutions among many alternatives. The performance of NACST/Seq is compared with other sequence design methods, and analyzed on a traveling salesman problem solved by bio-lab experiments. Our experimental results show that the evolutionary sequence design by NACST/Seq outperforms in its reliability the existing sequence design techniques such as conventional EAs, simulated annealing, and specialized heuristic methods. Soo-Yong Shin, In-Hee Lee 0001, Byoung-Tak Zhang |
IEEE Trans. Evol. Comput. | 4 |
| 2005 | Bayesian model averaging of Bayesian network classifiers over multiple node-orders: application to sparse datasetsabstractBayesian model averaging (BMA) can resolve the overfitting problem by explicitly incorporating the model uncertainty into the analysis procedure. Hence, it can be used to improve the generalization performance of Bayesian network classifiers. Until now, BMA of Bayesian network classifiers has only been performed in some restricted forms, e.g., the model is averaged given a single node-order, because of its heavy computational burden. However, it can be hard to obtain a good node-order when the available training dataset is sparse. To alleviate this problem, we propose BMA of Bayesian network classifiers over several distinct node-orders obtained using the Markov chain Monte Carlo sampling technique. The proposed method was examined using two synthetic problems and four real-life datasets. First, we show that the proposed method is especially effective when the given dataset is very sparse. The classification accuracy of averaging over multiple node-orders was higher in most cases than that achieved using a single node-order in our experiments. We also present experimental results for test datasets with unobserved variables, where the quality of the averaged node-order is more important. Through these experiments, we show that the difference in classification performance between the cases of multiple node-orders and single node-order is related to the level of noise, confirming the relative benefit of averaging over multiple node-orders for incomplete data. We conclude that BMA of Bayesian network classifiers over multiple node-orders has an apparent advantage when the given dataset is sparse and noisy, despite the method's heavy computational cost. Kyu-Baek Hwang, Byoung-Tak Zhang |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Korean Compound Noun Decomposition Using Syllabic Information Only
Seong-Bae Park, Jeong Ho Chang, Byoung-Tak Zhang |
CICLing | 3 |
| 2004 | BioPubMiner: Machine Learning Component-Based Biomedical Information Analysis Platform
Jae-Hong Eom, Byoung-Tak Zhang |
CIT | 2 |
| 2004 | Adaptive Neural Network-Based Clustering of Yeast Protein-Protein Interactions
Jae-Hong Eom, Byoung-Tak Zhang |
CIT | 2 |
| 2004 | Dynamic Asset Allocation Exploiting Predictors in Reinforcement Learning Framework
Jangmin O, Jae Won Lee, Jongwoo Lee, Byoung-Tak Zhang |
ECML | 4 |
| 2004 | Genetic Mining of DNA Sequence Structures for Effective Classification of the Risk Types of Human Papillomavirus (HPV)
Jae-Hong Eom, Seong-Bae Park, Byoung-Tak Zhang |
ICONIP | 3 |
| 2004 | Prediction of Implicit Protein-Protein Interaction by Optimal Associative Feature Mining
Jae-Hong Eom, Jeong Ho Chang, Byoung-Tak Zhang |
IDEAL | 3 |
| 2004 | PromSearch: A Hybrid Approach to Human Core-Promoter Prediction
Byoung-Hee Kim, Seong-Bae Park, Byoung-Tak Zhang |
IDEAL | 3 |
| 2004 | Stock Trading by Modelling Price Trend with Dynamic Bayesian Networks
Jangmin O, Jae Won Lee, Sung-Bae Park, Byoung-Tak Zhang |
IDEAL | 4 |
| 2004 | Evolutionary Continuous Optimization by Distribution Estimation with Variational Bayesian Independent Component Analyzers Mixture Model
Dong-Yeon Cho, Byoung-Tak Zhang |
PPSN | 2 |
| 2004 | Searching Transcriptional Modules Using Evolutionary Algorithms
Je-Gun Joung, Sok June Oh, Byoung-Tak Zhang |
PPSN | 3 |
| 2004 | Prediction of the Risk Types of Human Papillomaviruses by Support Vector Machines
Je-Gun Joung, Sok June Oh, Byoung-Tak Zhang |
PRICAI | 3 |
| 2004 | Multi-objective Evolutionary Probe Design Based on Thermodynamic Criteria for HPV Detection
In-Hee Lee 0001, Sun Kim, Byoung-Tak Zhang |
PRICAI | 3 |
| 2004 | Computational Methods for Identification of Human microRNA Precursors
Jin-Wu Nam, Wha-Jin Lee, Byoung-Tak Zhang |
PRICAI | 3 |
| 2004 | Co-trained support vector machines for large scale unstructured document classification using unlabeled data and syntactic information
Seong-Bae Park, Byoung-Tak Zhang |
Inf. Process. Manag. | 2 |
| 2004 | Development, evaluation and benchmarking of simulation software for biomolecule-based computing
Derrel Blain, Max H. Garzon, Soo-Yong Shin, Byoung-Tak Zhang, Satoshi Kashiwamura, Masahito Yamamoto, Atsushi Kameda, Azuma Ohuchi |
Nat. Comput. | 4 |
| 2003 | Text Chunking by Combining Hand-Crafted Rules and Memory-Based LearningabstractThis paper proposes a hybrid of hand-crafted rules and a machine learning method for chunking Korean. In the partially free word-order languages such as Korean and Japanese, a small number of rules dominate the performance due to their well-developed postpositions and endings. Thus, the proposed method is primarily based on the rules, and then the residual errors are corrected by adopting a memory-based machine learning method. Since the memory-based learning is an efficient method to handle exceptions in natural language processing, it is good at checking whether the estimates are exceptional cases of the rules and revising them. An evaluation of the method yields the improvement in F-score over the rules or various machine learning methods alone. Seong-Bae Park, Byoung-Tak Zhang |
ACL | 2 |
| 2003 | Molecular immunocomputing with application to alphabetical pattern recognition mimics the characterization of ABO blood typeabstractWe propose the concept of molecular immunocomputing which is a kind of peptide computing. Molecular immunocomputing is basically implemented by direct antigen-antibody biomolecular recognition on the basis of Fab fragments diversity. In this paper, we consider how molecular immunocomputing can tell the alphabet "O" from the characters "A" and "B" similar to the characterization of ABO blood type. To implement molecular immunocomputing, the two-dimensional figures are coded on the one-dimensional DNA strings. The information coded on the virtual DNA strings are transcribed to virtual RNA sequences, and then translated to the polypeptide sequences. The resulting peptide sequences are artificially synthesized and coupled to the carrier proteins. The resulting conjugate proteins are injected as input antigens to immunize the experimental animals. After immunization, we could purify the corresponding antibodies. These antibodies can be arrayed onto the protein microarray chips. The recent developments in the field of protein microarrays show the feasibility of molecular immunocomputing. Su Dong Kim, Ki-Roo Shin, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2003 | DNA sequence optimization using constrained multi-objective evolutionary algorithmabstractGenerating a set of the good DNA sequences needs to optimize multiple objectives and to satisfy several constraints. Therefore, it can be regarded as an instance of constrained multiobjective optimization problem. We apply the controlled elitist nondominating sorting genetic algorithm with constrained tournament selection to this problem. First, multiobjective approach and constrained multiobjective approach are compared in terms of the effectiveness in finding feasible the solutions. Then the performance is evaluated by comparing with the good sequences published in literature. In-Hee Lee 0001, Soo-Yong Shin, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 3 |
| 2003 | Mining the Risk Types of Human Papillomavirus (HPV) by AdaCost
Seong-Bae Park, Sohyun Hwang, Byoung-Tak Zhang |
DEXA | 3 |
| 2003 | Classification of the Risk Types of Human Papillomavirus by Decision Trees
Seong-Bae Park, Sohyun Hwang, Byoung-Tak Zhang |
IDEAL | 3 |
| 2003 | Automatic Webpage Classification Enhanced by Unlabeled Data
Seong-Bae Park, Byoung-Tak Zhang |
IDEAL | 2 |
| 2003 | An Empirical Study on Dimensionality Optimization in Text Mining for Linguistic Knowledge Acquisition
Yuseop Kim, Jeong Ho Chang, Byoung-Tak Zhang |
PAKDD | 3 |
| 2003 | Large Scale Unstructured Document Classification Using Unlabeled Data and Syntactic Information
Seong-Bae Park, Byoung-Tak Zhang |
PAKDD | 2 |
| 2003 | Genetic Mining of HTML Structures for Effective Web-Document Retrieval
Sun Kim, Byoung-Tak Zhang |
Appl. Intell. | 2 |
| 2003 | Word Sense Disambiguation by Learning Decision Trees from Unlabeled Data
Seong-Bae Park, Byoung-Tak Zhang, Yung Taek Kim |
Appl. Intell. | 2 |
| 2003 | Self-Organizing Latent Lattice Models for Temporal Gene Expression Profiling
Byoung-Tak Zhang, Jinsan Yang, Sung Wook Chi |
Mach. Learn. | 1 |
| 2002 | Evolutionary optimization by distribution estimation with mixtures of factor analyzersabstractEvolutionary optimization algorithms based on the probability models have been studied to capture the relationship between variables in the given problems and finally to find the optimal solutions more efficiently. However, premature convergence to local optima still happens in these algorithms. Many researchers have used the multiple populations to prevent this ill behavior since the key point is to ensure the diversity of the population. In this paper, we propose a new estimation of distribution algorithm by using the mixture of factor analyzers (MFA) which can cluster similar individuals in a group and explain the high order interactions with the latent variables for each group concurrently. We also adopt a stochastic selection method based on the evolutionary Markov chain Monte Carlo (eMCMC). Our experimental results support that the presented estimation of distribution algorithms with MFA and eMCMC-like selection scheme can achieve better performance for continuous optimization problems. Dong-Yeon Cho, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 2 |
| 2002 | Evolutionary sequence generation for reliable DNA computingabstractSince DNA computing technologies use the bio-molecules as basic computing materials, DNA computing involves the possibilities for errors caused by the chemical characteristics of bio-molecules. To overcome these drawbacks, many researchers have studied the design of DNA sequences to reduce the possibilities for illegal reactions. We developed an evolutionary sequence generation system to minimize the potential errors in DNA sequences for reliable DNA computing. We verified our system by investigating the sequences designed by another sequence generator, and generated the sequences for solving travelling salesman problems. Soo-Yong Shin, In-Hee Lee 0001, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 4 |
| 2002 | A Comparative Evaluation of Data-driven Models in Translation Selection of Machine Translation
Yuseop Kim, Jeong Ho Chang, Byoung-Tak Zhang |
COLING | 3 |
| 2002 | Stock Trading System Using Reinforcement Learning with Cooperative Agents
Jangmin O, Jae Won Lee, Byoung-Tak Zhang |
ICML | 3 |
| 2002 | A Boosted Maximum Entropy Model for Learning Text Chunking
Seong-Bae Park, Byoung-Tak Zhang |
ICML | 2 |
| 2002 | Topic Extraction from Text Documents Using Multiple-Cause Networks
Jeong Ho Chang, Jae Won Lee, Yuseop Kim, Byoung-Tak Zhang |
PRICAI | 4 |
| 2002 | Construction of Large-Scale Bayesian Networks by Local to Global Search
Kyu-Baek Hwang, Jae Won Lee, Seung-Woo Chung, Byoung-Tak Zhang |
PRICAI | 4 |
| 2002 | Target Word Selection Using WordNet and Data-Driven Models in Machine Translation
Yuseop Kim, Jeong Ho Chang, Byoung-Tak Zhang |
PRICAI | 3 |
| 2001 | Continuous estimation of distribution algorithms with probabilistic principal component analysisabstractMany evolutionary algorithms have been studied to build and use a probability distribution model of the population for optimization problems. Most of these methods tried to represent explicitly the relationship between variables in the problem with factorization techniques or a graphical model such as Bayesian or Gaussian networks. Thus enormous computational cost is required for constructing those models when the problem size is large. We propose a new estimation of distribution algorithm by using probabilistic principal component analysis (PPCA) which can explain the high order interactions with the latent variables. Since there are no explicit search procedures for the probability density structure, it is possible to rapidly estimate the distribution and readily sample the new individuals from it. Our experimental results support that the presented estimation of distribution algorithms with PPCA can find good solutions more efficiently than other EDAs for the continuous spaces. Dong-Yeon Cho, Byoung-Tak Zhang |
CEC | 2 |
| 2001 | Actively searching for committees of RBF networks using Bayesian evolutionary computationabstractCommittee machines are known to improve generalization performance by combining the predictions of many different individual learners. Evolutionary algorithms generate multiple models that can be combined to build a committee machine. This paper uses Bayesian evolutionary algorithms (BEAs) as a solution to evolve individual learners and build a committee machine. BEAs are based on the Bayesian evolutionary framework in which evolutionary computation is the process of repeatedly updating the posterior distribution of a population to find an individual with the maximum posteriori probability. BEAs evolve the number of centroids and the centroids' positions and widths for radial basis function (RBF) networks which are individual learners, and then the algorithms find an optimal committee from many different individuals. Empirical results show the machine's convergence characteristics and accuracy. Je-Gun Joung, Byoung-Tak Zhang |
CEC | 2 |
| 2001 | Evolutionary learning of Web-document structure for information retrievalabstractWeb documents have a number of tags indicating the structure of documents. The tag information can be utilized to improve the performance of document retrieval systems. The authors propose an approach to retrieve Web documents using HTML tags and then use a genetic algorithm to adapt the tag weights. This method uses a modified similarity measure based on the tag weights. A genetic learning method is used to select the tags for retrieval and get the optimal tag weights. We evaluate our method via experiments on conference pages and TREC document sets. The experimental results show that the tag weights are well trained by the proposed algorithm in accordance with the importance factors for retrieval. The proposed method has achieved about 10% improvement in retrieval accuracy. Sun Kim, Byoung-Tak Zhang |
CEC | 2 |
| 2001 | Convergence properties of Bayesian evolutionary algorithms with population size greater than 1abstractA Bayesian evolutionary algorithm is a probabilistic model of evolutionary computation for learning and optimization. It explicitly estimates the posterior distribution of the individuals and then samples offspring from the distribution. In the previous paper, using the asymptotic results from Markov chain Monte Carlo and annealing techniques, the asymptotic convergence of Bayesian evolutionary algorithms was shown for the case of population size 1. This paper presents convergence properties of Bayesian evolutionary algorithms with population size greater than 1. The basic idea is that BEAs can be reduced to Bayesian particle filters. The Bayesian particle filter approximates the posterior distribution of individuals at each generation. As the individuals evolve, the approximated posterior distribution also evolves. Then using the convergence properties of particle filters under some mild conditions, it is shown that as the number of individuals increases, a BEA converges to the posterior distribution. Si-Eun Lee, Byoung-Tak Zhang, Arnaud Doucet |
CEC | 2 |
| 2001 | Evolutionary calibration of sensors using genetic programming on evolvable hardwareabstractIn order to retain some degree of decision-making ability in a complex and dynamic environment, there have been many attempts to build autonomous mobile robots. However, conventional methods pay little attention to the unreliability of sensors. Because of corruption by noise and differences in sensitivity, even the same kinds of sensors show different observations under the same conditions. This causes a problem in that a minor change to the environment of the sensor system has a great influence on the perceptual ability of the robot. To improve the reliability of the sensors, we present a method for the evolutionary calibration of sensors using genetic programming as the calibration mechanism. In our approach, the sensor calibration logic is implemented on evolvable hardware. Therefore, as the learning goes on, the sensor interpretation circuit reconfigures itself to a more suitable form at run-time. Through two experiments on different tasks, we confirmed that our method significantly improved the correctness of interpretation. Ho-Sik Seok, Byoung-Tak Zhang |
CEC | 2 |
| 2001 | Bayesian evolutionary algorithms for continuous function optimizationabstractRecently many researchers have studied the estimation of distribution algorithms (EDAs) as an optimization method. While most EDAs focus on solving combinatorial optimization problems, only a few algorithms have been proposed for continuous function optimization. In previous work, we developed a Bayesian evolutionary algorithm (BEA) for combinatorial optimization problems using a probabilistic graphical model known as a Helmholtz machine. Since BEA is a general framework for evolutionary computation based on the Bayesian inductive principle, we improved BEA for continuous function optimization problems. By using the nature of the neural network and availability of the wake-sleep learning algorithm, the Helmholtz machine can capture the continuous distribution with a small modification. The proposed method has been applied to a suite of benchmark functions and compared with a real-coded genetic algorithm and previous experimental results. Soo-Yong Shin, Byoung-Tak Zhang |
CEC | 2 |
| 2001 | System identification using evolutionary Markov chain Monte Carlo
Byoung-Tak Zhang, Dong-Yeon Cho |
J. Syst. Archit. | 1 |
| 2001 | Collocation Dictionary Optimization Using WordNet and k-Nearest Neighbor Learning
Yuseop Kim, Byoung-Tak Zhang, Yung Taek Kim |
Mach. Transl. | 2 |
| 2001 | Learning-based Intrasentence Segmentation for Efficient Translation of Long Sentences
Sung Dong Kim, Byoung-Tak Zhang, Yung Taek Kim |
Mach. Transl. | 2 |
| 2000 | Word Sense Disambiguation by Learning from Unlabeled DataabstractMost corpus-based approaches to natural language processing suffer from lack of training data. This is because acquiring a large number of labeled data is expensive. This paper describes a learning method that exploits unlabeled data to tackle data sparseness problem. The method uses committee learning to predict the labels of unlabeled data that augment the existing training data. Our experiments on word sense disambiguation show that predictive accuracy is significantly improved by using additional unlabeled data. Seong-Bae Park, Byoung-Tak Zhang, Yung Taek Kim |
ACL | 2 |
| 2000 | Bayesian evolutionary algorithms for evolving neural tree models of time series dataabstractModel induction plays an important role in many fields of science and engineering to analyze data. Specifically, the performance of time series prediction whose objectives are to find out the dynamics of the underlying process in given data is greatly affected by the model. Bayesian evolutionary algorithms have been proposed as a method for automatic model induction from data. We apply Bayesian evolutionary algorithms (BEAs) to evolving neural tree models of time series data. The performances of various BEAs are compared on two time series prediction problems by varying the population size and the type of variation operations. Our experimental results support that population based BEAs with unlimited crossover find good models more efficiently than single individual BEAs, parallelized individual based BEAs, and population based BEAs with limited crossover. Dong-Yeon Cho, Byoung-Tak Zhang |
CEC | 2 |
| 2000 | Convergence properties of incremental Bayesian evolutionary algorithms with single Markov chainsabstractBayesian evolutionary algorithms (BEAs) are a probabilistic model of evolutionary computation for learning and optimization. Starting from a population of individuals drawn from a prior distribution, a Bayesian evolutionary algorithm iteratively generates a new population by estimating the posterior fitness distribution of parent individuals and then sampling from the distribution offspring individuals by variation and selection operators. Due to the non-homogeneity of their Markov chains, the convergence properties of the full BEAs are difficult to analyze. However, recent developments in Markov chain analysis for dynamic Monte Carlo methods provide a useful tool for studying asymptotic behaviors of adaptive Markov chain Monte Carlo methods including evolutionary algorithms. We apply these results to Investigate the convergence properties of Bayesian evolutionary algorithms with incremental data growth. We study the case of BEAs that generate single chains or have populations of size one. It is shown that under regularity conditions the incremental BEA asymptotically converges to a maximum a posteriori (MAP) estimate which is concentrated around the maximum likelihood estimate. This result relies on the observation that increasing the number of data items has an equivalent effect of reducing the temperature in simulated annealing. Byoung-Tak Zhang, Gerhard Paass, Heinz Mühlenbein |
CEC | 1 |
| 2000 | Reducing Parsing Complexity by Intra-Sentence Segmentation based on Maximum Entropy ModelabstractLong sentence analysis has been a critical problem because of high complexity. This paper addresses the reduction of parsing complexity by intra-sentence segmentation, and presents maximum entropy model for determining segmentation positions. The model features lexical contexts of segmentation positions, giving a probability to each potential position. Segmentation coverage and accuracy of the proposed method are 96% and 88% respectively. The parsing efficiency is improved by 77% in time and 71% in space. Sung Dong Kim, Byoung-Tak Zhang, Yung Tack Kim |
EMNLP | 2 |
| 2000 | A reinforcement learning agent for personalized information filteringabstractThis paper describes a method for learning user's interests in the Web-based personalized information filtering system called WAIR. The proposed method analyzes user's reactions to the presented documents and learns from them the profiles for the individual users. Reinforcement learning is used to adapt the term weights in the user profile so that user's preferences are best represented. In contrast to conventional relevance feedback methods which require explicit user feedbacks, our approach learns user preferences implicitly from direct observations of user behaviors during interaction. Field tests have been made which involved 7 users reading a total of 7,700 HTML documents during 4 weeks. The proposed method showed superior performance in personalized information filtering compared to the existing relevance feedback methods. Young-Woo Seo, Byoung-Tak Zhang |
IUI | 2 |
| 2000 | Building Optimal Committees of Genetic Programs
Byoung-Tak Zhang, Je-Gun Joung |
PPSN | 1 |
| 2000 | Bayesian Evolutionary Optimization Using Helmholtz Machines
Byoung-Tak Zhang, Soo-Yong Shin |
PPSN | 1 |
| 2000 | Text filtering by boosting naive bayes classifiersabstractSeveral machine learning algorithms have recently been used for text categorization and filtering. In particular, boosting methods such as AdaBoost have shown good performance applied to real text data. However, most of existing boosting algorithms are based on classifiers that use binary-valued features. Thus, they do not fully make use of the weight information provided by standard term weighting methods. In this paper, we present a boosting-based learning method for text filtering that uses naive Bayes classifiers as a weak learner. The use of naive Bayes allows the boosting algorithm to utilize term frequency information while maintaining probabilistically accurate confidence ratio. Applied to TREC-7 and TREC-8 filtering track documents, the proposed method obtained a significant improvement in LF1, LF2, F1 and F3 measures compared to the best results submitted by other TREC entries. Yu-Hwan Kim, Shang-Yoon Hahn, Byoung-Tak Zhang |
SIGIR | 3 |
| 2000 | Behavior evolution of autonomous mobile robot using genetic programming based on evolvable hardwareabstractThis paper presents a genetic programming based evolutionary strategy for on-line adaptive learnable evolvable hardware. Genetic programming can be a useful control method for evolvable hardware for its unique tree structured chromosome. However it is difficult to represent the tree structured chromosome in hardware, and it is difficult to use the crossover operator in hardware. Therefore, genetic programming is not as popular as genetic algorithms in the evolvable hardware community in spite of its possible strengths. We propose a chromosome representation method and a hardware implementation method that can be helpful for this situation. Our method uses a context switchable identical block structure to implement a genetic tree in evolvable hardware. We compose an evolutionary strategy for evolvable hardware by combining the proposed method with other research results. The proposed method is applied to the autonomous mobile robots cooperation problem to verify its usefulness. Chang-Bong Ban, Kwee-Bo Sim, Ho-Sik Seok, Kwang-Ju Lee, Byoung-Tak Zhang |
SMC | 6 |
| 1999 | Effects of selection schemes in genetic programming for time series predictionabstractThe problem of time series prediction provides a practical benchmark for testing the performance of evolutionary algorithms. In this paper, we compare various selection methods for genetic programming, an evolutionary computation with variable-size tree representations, with application to time series data. Selection is an important operator that controls the dynamics of evolutionary computation. A number of selection operators have been so far proposed and tested in evolutionary algorithms with fixed-size chromosomes. However, the effect of selection schemes remains relatively unexplored in evolutionary algorithms with variable-size representations. We analyze the evolutionary dynamics of genetic programming by means of the selection to response and the selection differential proposed in the breeder genetic algorithm (BGA). The empirical analysis using the laser time-series data suggests that hard selection is more preferable than soft selection. This seems due to the lack of heritability in genetic programming. Jung-Jib Kim, Byoung-Tak Zhang |
CEC | 2 |
| 1999 | Solving traveling salesman problems using molecular programmingabstractMolecular programming (MP) has been proposed as an evolutionary computation algorithm at the molecular level [Zhang and Shin, 1998]. MP are different from other evolutionary algorithms in its representation of solutions using DNA molecular structures and its use of bio-lab techniques for recombination of partial solutions. In this paper, molecular programming is applied to traveling salesman problems (TSPs) whose solution requires encoding of real-values in DNA strands. We propose a new encoding scheme for real values that is biologically plausible and has a fixed code length. The effectiveness of the proposed method is verified by simulations and by comparison with Narayanan and Zorbalas'. 1 Introduction The field of DNA computing was pioneered by Adleman [1] who showed the potential of using biomolecules for solving computational problems. He solved the hamiltonian path problem using DNA molecules, and Lipton came up with a method using DNA computing to solve the satisfiability (SA... Soo-Yong Shin, Byoung-Tak Zhang, Sung-Soo Jun |
CEC | 2 |
| 1999 | A Bayesian framework for evolutionary computationabstractA Bayesian framework for evolutionary computation is presented. Given a data set for fitness evaluation the best (fittest) individual is defined as the most probable model of the data with respect to the prior knowledge on the problem domain. In each generation, Bayes theorem is used to estimate the posterior fitness of individuals from their prior fitness values. Offspring individuals are then generated by sampling from the posterior distribution combined with the transition probabilities formed by variation operators. The evolutionary inference steps from the prior via posterior distribution of parent fitness to the expected fitness distribution of offspring are essential elements in Bayesian evolutionary computation. One of the most interesting aspects of Bayesian evolution is that it provides principled techniques for controlling evolutionary dynamics. Specifically, we describe two examples of the application of the Bayesian framework. One is a Bayesian evolutionary algorithm (BEA) designed to evolve parsimonious individuals in evolutionary computation with variable-size representation. We show that the adaptive Occam method for program growth control is a special form of Bayesian evolution. The other example is an evolutionary algorithm with incremental data inheritance (IDI). In this BEA, the fitness of individuals is estimated on incrementally chosen data subsets, rather than on the whole data set, and thus the convergence is accelerated by reducing the effective number of fitness evaluations. Experimental results are provided to show the effectiveness of the BEAs. Byoung-Tak Zhang |
CEC | 1 |
| 1999 | Time series prediction using committee machines of evolutionary neural treesabstractEvolutionary neural trees (ENTs) are tree-structured neural networks constructed by evolutionary algorithms. We use ENTs to build predictive models of time series data. Time series data are typically characterized by dynamics of the underlying process and thus the robustness of predictions is crucial. We describe a method for making more robust predictions by building committees of ENTs, i.e. CENTs. The method extends the concept of mixing genetic programming (MGP) which makes use of the fact that evolutionary computation produces multiple models as output instead of just one best. Experiments have been performed on the laser time series in which the CENTs outperformed the single best ENTs. We also discuss a theoretical foundation of CENTs using the Bayesian framework for evolutionary computation. Byoung-Tak Zhang, Je-Gun Joung |
CEC | 1 |
| 1999 | Combining locally trained neural networks by introducing a reject classabstractThis paper presents a new strategy for building and combining a local committee when a dataset is given. Training local committees is performed in two stages: active data partitioning and recombination by introducing an additional reject class. Active data partitioning is a preprocessing step that partitions the given dataset into several similar subsets using active learning. Additional reject class in this strategy plays an important role in assigning a focused area to each individual network of the committee. For combining the outputs of each individual network, we use a kind of sum rule criteria, assuming that the outputs of the individuals are equivalent to a posteriori Bayesian probabilities. All the learning procedures are based on the active learning paradigm. Experiments are performed on the two real-world datasets from the UCI machine learning database. The results show that the active data partitioning and recombining strategy is very successful for building a local committee and the combined result outperforms other algorithms, but the combined result can be affected by the training error level /spl epsiv/. Suk-Joon Kim, Byoung-Tak Zhang |
IJCNN | 2 |
| 1999 | Temporal pattern recognition using a spiking neural network with delaysabstractSpiking neural networks have been shown to have powerful computation capability, but most results have been restricted to theoretical work. In this paper, we apply a spiking neural network to a time-series prediction problem, i.e., laser amplitude fluctuation data. We formulate the time-series problem as a spatio-temporal pattern recognition problem and present a learning method in which spatio-temporal patterns are recorded as synaptic delays. Experimental results show that the presented model is useful for temporal pattern recognition. Jeong-woo Sohn, Byoung-Tak Zhang, Bong-Kiun Kaang |
IJCNN | 2 |
| 1999 | Compound noun decomposition using a Markov modelabstractA statistical method for compound noun decomposition is presented. Previous studies on this problem showed some statistical information are helpful. But applying statistical information was not so systemic that performance depends heavily on the algorithm and some algorithms usually have many separated steps. In our work statistical information is collected from manually decomposed compound noun corpus to build a Markov model for composition. Two Markov chains representing statistical information are assumed independent: one for the sequence of participants’ lengths and another for the sequence of participants ’ features. Besides Markov assumptions, least participants preference assumption also is used. These two assumptions enable the decomposition algorithm to be a kind of conditional dynamic programming so that efficient and systemic computation can be performed. When applied to test data of size 5027, we obtained a precision of 98.4%. Jongwoo Lee, Byoung-Tak Zhang, Yung Taek Kim |
MTSummit | 2 |
| 1998 | Active Data Partitioning for Building Mixture Models
Suk-Joon Kim, Byoung-Tak Zhang |
ICONIP | 2 |
| 1997 | Evolutionary Induction of Sparse Neural TreesabstractThis paper is concerned with the automatic induction of parsimonious neural networks. In contrast to other program induction situations, network induction entails parametric learning as well as structural adaptation. We present a novel representation scheme called neural trees that allows efficient learning of both network architectures and parameters by genetic search. A hybrid evolutionary method is developed for neural tree induction that combines genetic programming and the breeder genetic algorithm under the unified framework of the minimum description length principle. The method is successfully applied to the induction of higher order neural trees while still keeping the resulting structures sparse to ensure good generalization performance. Empirical results are provided on two chaotic time series prediction problems of practical interest. Byoung-Tak Zhang, Peter Ohm, Heinz Mühlenbein |
Evol. Comput. | 1 |
| 1995 | Balancing Accuracy and Parsimony in Genetic ProgrammingabstractGenetic programming is distinguished from other evolutionary algorithms in that it uses tree representations of variable size instead of linear strings of fixed length. The flexible representation scheme is very important because it allows the underlying structure of the data to be discovered automatically. One primary difficulty, however, is that the solutions may grow too big without any improvement of their generalization ability. In this article we investigate the fundamental relationship between the performance and complexity of the evolved structures. The essence of the parsimony problem is demonstrated empirically by analyzing error landscapes of programs evolved for neural network synthesis. We consider genetic programming as a statistical inference problem and apply the Bayesian model-comparison framework to introduce a class of fitness functions with error and complexity terms. An adaptive learning method is then presented that automatically balances the model-complexity factor to evolve parsimonious programs without losing the diversity of the population needed for achieving the desired training accuracy. The effectiveness of this approach is empirically shown on the induction of sigma-pi neural networks for solving a real-world medical diagnosis problem as well as benchmark tasks. Byoung-Tak Zhang, Heinz Mühlenbein |
Evol. Comput. | 1 |
| 1994 | Effects of Occam's Razor in Evolving Sigma-Pi Neural Nets
Byoung-Tak Zhang |
PPSN | 1 |
| 1994 | Accelerated Learning by Active Example SelectionabstractMuch previous work on training multilayer neural networks has attempted to speed up the backpropagation algorithm using more sophisticated weight modification rules, whereby all the given training examples are used in a random or predetermined sequence. In this paper we investigate an alternative approach in which the learning proceeds on an increasing number of selected training examples, starting with a small training set. We derive a measure of criticality of examples and present an incremental learning algorithm that uses this measure to select a critical subset of given examples for solving the particular task. Our experimental results suggest that the method can significantly improve training speed and generalization performance in many real applications of neural networks. This method can be used in conjunction with other variations of gradient descent algorithms. Byoung-Tak Zhang |
Int. J. Neural Syst. | 1 |
| 1990 | Morphological Analysis and Synthesis by Automated Discovery and Acquisition of Linguistic Rules
Byoung-Tak Zhang, Yung Taek Kim |
COLING | 1 |