Yi-Ren Yeh

dblp:41/4830 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-4264-523XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unseen documents.Nonetheless, they struggle with high-level conceptual understanding and holistic comprehension due to limited context windows, which constrain their ability to perform deep reasoning over long-form, domainspecific content such as full-length books.To solve this problem, knowledge graphs (KGs) have been leveraged to provide entity-centric structure and hierarchical summaries, offering more structured support for reasoning.However, existing KG-based RAG solutions remain restricted to text-only inputs and fail to leverage the complementary insights provided by other modalities such as vision.On the other hand, reasoning from visual documents requires textual, visual, and spatial cues into structured, hierarchical concepts.To address this issue, we introduce a multimodal knowledge graphbased RAG that enables cross-modal reasoning for better content understanding.Our method incorporates visual cues into the construction of knowledge graphs, the retrieval phase, and the answer generation process.Experimental results across both global and fine-grained question answering tasks show that our approach consistently outperforms existing approaches on both textual and multimodal benchmarks.
Chi-Hsiang Hsiao, Yi-Cheng Wang, Tzung-Sheng Lin, Yi-Ren Yeh, Chu-Song Chen
ACL (1)4
2026 HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation
abstract
Graph-based Retrieval-Augmented Generation (RAG) typically operates on binary Knowledge Graphs (KGs). However, decomposing complex facts into binary triples often leads to semantic fragmentation and longer reasoning paths, increasing the risk of retrieval drift and computational overhead. In contrast, n-ary hypergraphs preserve high-order relational integrity, enabling shallower and more semantically cohesive inference. To exploit this topology, we propose HyperRAG, a framework tailored for n-ary hypergraphs featuring two complementary retrieval paradigms: (i) HyperRetriever learns structural-semantic reasoning over n-ary facts to construct query-conditioned relational chains. It enables accurate factual tracking, adaptive high-order traversal, and interpretable multi-hop reasoning under context constraints. (ii) HyperMemory leverages the LLM's parametric memory to guide beam search, dynamically scoring n-ary facts and entities for query-aware path expansion. Extensive evaluations on WikiTopics (11 closed-domain datasets) and three open-domain QA benchmarks (HotpotQA, MuSiQue, and 2WikiMultiHopQA) validate HyperRAG's effectiveness. HyperRetriever achieves the highest answer accuracy overall, with average gains of 2.95% in MRR and 1.23% in Hits@10 over the strongest baseline. Qualitative analysis further shows that HyperRetriever bridges reasoning gaps through adaptive and interpretable n-ary chain construction, benefiting both open and closed-domain QA. Our codes are publicly available at https://github.com/Vincent-Lien/HyperRAG.git.
Wen-Sheng Lien, Yu-Kai Chan, Hao-Lung Hsiao, Bo-Kai Ruan, Meng-Fen Chiang, Chien-An Chen, Yi-Ren Yeh, Hong-Han Shuai
WWW7
2025 Relation-Rich Visual Document Generator for Visual Information Extraction
abstract
Despite advances in Large Language Models (LLMs) and Multimodal LLMs (MLLMs) for visual document understanding (VDU), visual information extraction (VIE) from relation-rich documents remains challenging due to the layout diversity and limited training data. While existing synthetic document generators attempt to address data scarcity, they either rely on manually designed layouts and templates, or adopt rule-based approaches that limit layout diversity. Besides, current layout generation methods focus solely on topological patterns without considering textual content, making them impractical for generating documents with complex associations between the contents and layouts. In this paper, we propose a Relation-rIch visual Document GEnerator (RIDGE) that addresses these limitations through a two-stage approach: (1) Content Generation, which leverages LLMs to generate document content using a carefully designed Hierarchical Structure Text format which captures entity categories and relationships, and (2) Content-driven Layout Generation, which learns to create diverse, plausible document layouts solely from easily available Optical Character Recognition (OCR) results, requiring no human labeling or annotations efforts. Experimental results have demonstrated that our method significantly enhances the performance of document understanding models on various VIE benchmarks.
Zi-Han Jiang, Chien-Wei Lin, Hsuan-Tung Liu, Yi-Ren Yeh, Chu-Song Chen
CVPR5
2025 Time Series Classification with Language Models
abstract
Time series classification is a crucial yet challenging task due to the diverse and complex nature of time series data. While deep learning methods, particularly Convolutional Neural Networks (CNNs), have been widely explored for this problem, transformer-based approaches remain relatively underutilized. With the rise of large language models, recent research has begun to investigate how various data types, including time series, can be integrated into language models. In this paper, we explore the transformation of time series data into symbolic representations for use with language models such as BERT. Symbolizing time series data is a foundational and non-trivial step, as the choice of symbolization method can significantly affect downstream performance. We examine three different symbolization techniques and analyze their impact on classification accuracy. Additionally, we investigate the effect of task-adaptive pretraining using an expanded symbolic time series corpus. By combining these elements, our study provides a analysis of applying language models to time series classification, contributing valuable insights into the effective utilization of symbolic representations and pretrained transformers in this domain.
Chun-Chieh Lin, Yi-Ren Yeh
KES2
2024 Audio Classification with Semi-Supervised Contrastive Loss and Consistency Regularization
abstract
Deep learning models have shown remarkable success across various domains, but their effectiveness often depends on access to extensive labeled datasets. Acquiring large amounts of labeled data, however, can be challenging. Contrastive learning, using contrastive loss, has emerged as a promising approach for learning useful representations. In this work, we address a scenario where we have access to both labeled and unlabeled data from the same domain. While labeled data may be limited, unlabeled data is often more abundant. To tackle this challenge, we propose an end-to-end audio classification model in a semi-supervised learning setting, integrating contrastive loss, supervised contrastive loss, and consistency regularization. Our approach effectively combines information from both labeled and unlabeled data, enhancing classification performance even with limited labeled data. Experimental results demonstrate the efficacy of our model, which outperforms baseline models and the FixMatch method, highlighting the potential of leveraging both labeled and unlabeled data for audio classification tasks.
Juan-Wei Xu, Yi-Ren Yeh
COMPSAC2
2024 DetailSemNet: Elevating Signature Verification Through Detail-Semantic Integration
Meng-Cheng Shih, Tsai-Ling Huang, Yu-Heng Shih, Hong-Han Shuai, Hsuan-Tung Liu, Yi-Ren Yeh
ECCV (25)6
2023 Domain-Generalized Face Anti-Spoofing with Unknown Attacks
abstract
Although face anti-spoofing (FAS) methods have achieved remarkable performance on specific domains or attack types, few studies have focused on the simultaneous presence of domain changes and unknown attacks, which is closer to real application scenarios. To handle domain-generalized unknown attacks, we introduce a new method, DGUA-FAS, which consists of a Transformer-based feature extractor and a synthetic unknown attack sample generator (SUASG). The SUASG network simulates unknown attack samples to assist the training of the feature extractor. Experimental results show that our method achieves superior performance on domain generalization FAS with known or unknown attacks.
Zong-Wei Hong, Hsuan-Tung Liu, Yi-Ren Yeh, Chu-Song Chen
ICIP4
2023 Domain Invariant Vision Transformer Learning for Face Anti-spoofing
abstract
Existing face anti-spoofing (FAS) models have achieved high performance on specific datasets. However, for the application of real-world systems, the FAS model should generalize to the data from unknown domains rather than only achieve good results on a single baseline. As vision transformer models have demonstrated astonishing performance and strong capability in learning discriminative information, we investigate applying transformers to distinguish the face presentation attacks over unknown domains. In this work, we propose the Domain-invariant Vision Transformer (DiVT) for FAS, which adopts two losses to improve the generalizability of the vision transformer. First, a concentration loss is employed to learn a domain-invariant representation that aggregates the features of real face data. Second, a separation loss is utilized to union each type of attack from different domains. The experimental results show that our proposed method achieves state-of-the-art performance on the protocols of domain-generalized FAS tasks. Compared to previous domain generalization FAS models, our proposed method is simpler but more effective.
Chen-Hao Liao, Wen-Cheng Chen, Hsuan-Tung Liu, Yi-Ren Yeh, Min-Chun Hu 0001, Chu-Song Chen
WACV4
2022 Smile: Sequence-to-Sequence Domain Adaptation with Minimizing Latent Entropy for Text Image Recognition
abstract
Excellent text recognition results have been obtained by training recognition models with synthetic images. However, recognizing text from real-world images still faces challenges due to the domain shift between synthetic and real-world text images. One strategy to eliminate this domain difference without manual annotation is unsupervised domain adaptation (UDA). Due to the characteristics of sequential labeling tasks, most popular UDA methods cannot be directly applied to text recognition. To tackle this problem, we proposed a UDA method that minimizes latent entropy on sequence-to-sequence attention-based models with class-balanced self-paced learning. Experimental results show that our proposed framework achieves better recognition results than the existing methods on most UDA text recognition benchmarks. All codes are publicly available1.
Yen-Cheng Chang, Yi-Chang Chen, Yu-Chuan Chang, Yi-Ren Yeh
ICIP4
2022 g2pW: A Conditional Weighted Softmax BERT for Polyphone Disambiguation in Mandarin
abstract
Polyphone disambiguation is the most crucial task in Mandarin grapheme-to-phoneme (g2p) conversion.Previous studies have approached this problem using pre-trained language models, restricted output, and extra information from Part-Of-Speech (POS) tagging.Inspired by these strategies, we propose a novel approach, called g2pW, which adapts learnable softmax-weights to condition the outputs of BERT with the polyphonic character of interest and its POS tagging.Rather than using the hard mask as in previous works, our experiments show that learning a soft-weighting function for the candidate phonemes benefits performance.In addition, our proposed g2pW does not require extra pre-trained POS tagging models while using POS tags as auxiliary features since we train the POS tagging model simultaneously with the unified encoder.Experimental results show that our g2pW outperforms existing methods on the public CPP dataset.All codes, model weights, and a user-friendly package are publicly available.
Yi-Chang Chen, Yu-Chuan Steven, Yen-Cheng Chang, Yi-Ren Yeh
INTERSPEECH4
2022 Towards Understanding Cross Resolution Feature Matching for Surveillance Face Recognition
abstract
Cross-resolution face recognition (CRFR) in an open-set setting is a practical application for surveillance scenarios where low-resolution (LR) probe faces captured via surveillance cameras require being matched to a watchlist of high-resolution (HR) galleries. Although CRFR is to be of practical use, it sees a performance drop of more than 10% compared to that of high-resolution face recognition protocols. The challenges of CRFR are multifold, including the domain gap induced by the HR and LR images, the pose/texture variations, etc. To this end, this work systematically discusses possible issues and their solutions that affect the accuracy of CRFR. First, we explore the effect of resolution changes and conclude that resolution matching is the key for CRFR. Even simply downscaling the HR faces to match the LR ones brings a performance gain. Next, to further boost the accuracy of matching cross-resolution faces, we found that a well-designed super-resolution network, which can (a) represent the images continuously, is (b) suitable for real-world degradation kernel, (c) adaptive to different input resolutions, and (d) guided by an identity-preserved loss, is necessary to upsample the LR faces with discriminative enhancement. Here, the proposed identity-preserved loss plays the role of reconciling the objective discrepancy of super-resolution between human perception and machine recognition. Finally, we emphasize that removing the pose variations is an essential step before matching faces for recognition in the super-resolved feature space. Our method is evaluated on benchmark datasets, including SCface, cross-resolution LFW, and QMUL-Tinyface. The results show that the proposed method outperforms the SOTA methods by a clear margin and narrows the performance gap compared to the high-resolution face recognition protocol.
Chiawei Kuo, Yi-Ting Tsai, Hong-Han Shuai, Yi-Ren Yeh
ACM Multimedia4
2018 A Malware Beacon of Botnet by Local Periodic Communication Behavior
abstract
Botnets are one of most serious threats in cyber security. Many previous studies have been proposed for botnet detection. Among those approaches, one of main tracks focuses on extracting informative features from network traffic flows. Nevertheless, most features of interest are extracted from the information of a single connection, such as flow duration, flow packet size etc. In this paper, we proposed an novel feature, which is able to detect a long-term behavior of botnets. More specifically, we aim to extract a malware beacon from the periodic communication between bots and bot master. Besides the regular communication pattern, we also explore several types of botnet behavior to leverage the effectiveness of the proposed feature. Our experimental results show that our proposed periodic communication signature could be one of effective features for detecting compromised devices.
Yi-Ren Yeh, Ming-Kung Sun, C.-Y. Huang
COMPSAC (2)1
2017 Semantics-Preserving Locality Embedding for Zero-Shot Learning
Shi-Yen Tao, Yi-Ren Yeh, Yu-Chiang Frank Wang
BMVC2
2016 Domain-Constraint Transfer Coding for Imbalanced Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) deals with the task that labeled training and unlabeled test data collected from source and target domains, respectively. In this paper, we particularly address the practical and challenging scenario of imbalanced cross-domain data. That is, we do not assume the label numbers across domains to be the same, and we also allow the data in each domain to be collected from multiple datasets/sub-domains. To solve the above task of imbalanced domain adaptation, we propose a novel algorithm of Domain-constraint Transfer Coding (DcTC). Our DcTC is able to exploit latent subdomains within and across data domains, and learns a common feature space for joint adaptation and classification purposes. Without assuming balanced cross-domain data as most existing UDA approaches do, we show that our method performs favorably against state-of-the-art methods on multiple cross-domain visual classification tasks.
Yao-Hung Tsai, Cheng-An Hou, Wei-Yu Chen, Yi-Ren Yeh, Yu-Chiang Frank Wang
AAAI4
2016 Learning Cross-Domain Landmarks for Heterogeneous Domain Adaptation
abstract
While domain adaptation (DA) aims to associate the learning tasks across data domains, heterogeneous domain adaptation (HDA) particularly deals with learning from cross-domain data which are of different types of features. In other words, for HDA, data from source and target domains are observed in separate feature spaces and thus exhibit distinct distributions. In this paper, we propose a novel learning algorithm of Cross-Domain Landmark Selection (CDLS) for solving the above task. With the goal of deriving a domain-invariant feature subspace for HDA, our CDLS is able to identify representative cross-domain data, including the unlabeled ones in the target domain, for performing adaptation. In addition, the adaptation capabilities of such cross-domain landmarks can be determined accordingly. This is the reason why our CDLS is able to achieve promising HDA performance when comparing to state-of-the-art HDA methods. We conduct classification experiments using data across different features, domains, and modalities. The effectiveness of our proposed method can be successfully verified.
Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
CVPR2
2016 Heterogeneous domain adaptation with label and structure consistency
abstract
Domain adaptation is a challenging task, since it associates data collected from different domains or exhibiting distinct distributions. In this paper, we particularly focus on adapting cross-domain data with distinct feature dimensions or representations. Thus, this is referred to as the task of heterogeneous domain adaptation (HDA). To solve HDA, we propose Label and Structure-consistent Unilateral Projection (LS-UP) that transforms source-domain data to the target domain, with the goal of matching cross-domain data distribution and preserving data structure after projection. The main contribution of our work is its ability in relating cross-domain data with different feature representations. We evaluate our LS-UP for HDA on two different cross-domain classification problems, and we show that our method would perform favorably against state-of-the-art approaches.
Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICASSP2
2016 Recognizing heterogeneous cross-domain data via generalized joint distribution adaptation
abstract
In this paper, we propose a novel algorithm of Generalized Joint Distribution Adaptation (G-JDA) for heterogeneous domain adaptation (HDA), which associates and recognizes cross-domain data observed in different feature spaces (and thus with different dimensionality). With the objective to derive a domain-invariant feature subspace for relating source and target-domain data, our G-JDA learns a pair of feature projection matrices (one for each domain), which allows us to eliminate the difference between projected cross-domain heterogeneous data by matching their marginal and class-conditional distributions. We conduct experiments on cross-domain classification tasks using data across different features, datasets, and modalities. We confirm that our G-JDA would perform favorably against state-of-the-art HDA approaches.
Yuan-Ting Hsieh, Shi-Yen Tao, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICME4
2016 Unsupervised Domain Adaptation With Label and Structural Consistency
abstract
Unsupervised domain adaptation deals with scenarios in which labeled data are available in the source domain, but only unlabeled data can be observed in the target domain. Since the classifiers trained by source-domain data would not be expected to generalize well in the target domain, how to transfer the label information from source to target-domain data is a challenging task. A common technique for unsupervised domain adaptation is to match cross-domain data distributions, so that the domain and distribution differences can be suppressed. In this paper, we propose to utilize the label information inferred from the source domain, while the structural information of the unlabeled target-domain data will be jointly exploited for adaptation purposes. Our proposed model not only reduces the distribution mismatch between domains, improved recognition of target-domain data can be achieved simultaneously. In the experiments, we will show that our approach performs favorably against the state-of-the-art unsupervised domain adaptation methods on benchmark data sets. We will also provide convergence, sensitivity, and robustness analysis, which support the use of our model for cross-domain classification.
Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
IEEE Trans. Image Process.3
2015 An unsupervised domain adaptation approach for cross-domain visual classification
abstract
For cross-view action recognition and many real-world visual classification problems, one needs to recognize test data at a particular target domain of interest, while training data are collected at a different source domain. Without eliminating such domain differences, recognition of test data using classifiers trained in the source domain will not be expected to produce satisfactory performance. In this paper, we propose a novel domain adaptation approach, which is able to learn a common feature space relating cross-domain data. In particular, we not only aim at matching cross-domain data marginal distributions during adaptation, we also exploit the structure of target domain data and update class-conditional distributions accordingly. Experiments on various cross-domain visual classification tasks would verify the effectiveness and robustness of our proposed method.
Cheng-An Hou, Yi-Ren Yeh, Yu-Chiang Frank Wang
AVSS2
2015 Unsupervised Domain Adaptation with Imbalanced Cross-Domain Data
abstract
We address a challenging unsupervised domain adaptation problem with imbalanced cross-domain data. For standard unsupervised domain adaptation, one typically obtains labeled data in the source domain and only observes unlabeled data in the target domain. However, most existing works do not consider the scenarios in which either the label numbers across domains are different, or the data in the source and/or target domains might be collected from multiple datasets. To address the aforementioned settings of imbalanced cross-domain data, we propose Closest Common Space Learning (CCSL) for associating such data with the capability of preserving label and structural information within and across domains. Experiments on multiple cross-domain visual classification tasks confirm that our method performs favorably against state-of-the-art approaches, especially when imbalanced cross-domain data are presented.
Tzu-Ming Harry Hsu, Wei-Yu Chen, Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICCV5
2015 Connecting the dots without clues: Unsupervised domain adaptation for cross-domain visual classification
abstract
Many real-world visual classification tasks require one to recognize test data in a particular domain of interest, while the training data can only be collected from a different domain. This can be viewed as the problem of unsupervised domain adaptation, in which the domain difference and the lack of cross-domain label/correspondence information make the recognition task very difficult. In this paper, we propose to exploit the cross-domain data correspondence using both observed data similarity and labels transferred from the source domain. This allows us to perform distribution matching for cross-domain data with recognition guarantees. Our experiments on three different cross-domain visual classification tasks would confirm the effectiveness of our method, which is shown to perform favorably against state-of-the-art unsupervised domain adaptation approaches.
Wei-Yu Chen, Tzu-Ming Harry Hsu, Cheng-An Hou, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICIP4
2014 Heterogeneous Domain Adaptation and Classification by Exploiting the Correlation Subspace
abstract
We present a novel domain adaptation approach for solving cross-domain pattern recognition problems, i.e., the data or features to be processed and recognized are collected from different domains of interest. Inspired by canonical correlation analysis (CCA), we utilize the derived correlation subspace as a joint representation for associating data across different domains, and we advance reduced kernel techniques for kernel CCA (KCCA) if nonlinear correlation subspace are desirable. Such techniques not only makes KCCA computationally more efficient, potential over-fitting problems can be alleviated as well. Instead of directly performing recognition in the derived CCA subspace (as prior CCA-based domain adaptation methods did), we advocate the exploitation of domain transfer ability in this subspace, in which each dimension has a unique capability in associating cross-domain data. In particular, we propose a novel support vector machine (SVM) with a correlation regularizer, named correlation-transfer SVM, which incorporates the domain adaptation ability into classifier design for cross-domain recognition. We show that our proposed domain adaptation and classification approach can be successfully applied to a variety of cross-domain recognition tasks such as cross-view action recognition, handwritten digit recognition with different features, and image-to-text or text-to-image classification. From our empirical results, we verify that our proposed method outperforms state-of-the-art domain adaptation approaches in terms of recognition performance.
Yi-Ren Yeh, Chun-Hao Huang, Yu-Chiang Frank Wang
IEEE Trans. Image Process.1
2013 Locality-sensitive dictionary learning for sparse representation based classification
Chia-Po Wei, Yu-Wei Chao, Yi-Ren Yeh, Yu-Chiang Frank Wang
Pattern Recognit.3
2013 A rank-one update method for least squares linear discriminant analysis with concept drift
Yi-Ren Yeh, Yu-Chiang Frank Wang
Pattern Recognit.1
2013 Anomaly Detection via Online Oversampling Principal Component Analysis
abstract
Anomaly detection has been an important research topic in data mining and machine learning. Many real-world applications such as intrusion or credit card fraud detection require an effective and efficient framework to identify deviated data instances. However, most anomaly detection methods are typically implemented in batch mode, and thus cannot be easily extended to large-scale problems without sacrificing computation and memory requirements. In this paper, we propose an online oversampling principal component analysis (osPCA) algorithm to address this problem, and we aim at detecting the presence of outliers from a large amount of data via an online updating technique. Unlike prior principal component analysis (PCA)-based approaches, we do not store the entire data matrix or covariance matrix, and thus our approach is especially of interest in online or large-scale problems. By oversampling the target instance and extracting the principal direction of the data, the proposed osPCA allows us to determine the anomaly of the target instance according to the variation of the resulting dominant eigenvector. Since our osPCA need not perform eigen analysis explicitly, the proposed framework is favored for online applications which have computation or memory limitations. Compared with the well-known power method for PCA and other popular anomaly detection algorithms, our experimental results verify the feasibility of our proposed method in terms of both accuracy and efficiency.
Yuh-Jye Lee, Yi-Ren Yeh, Yu-Chiang Frank Wang
IEEE Trans. Knowl. Data Eng.2
2012 A Novel Multiple Kernel Learning Framework for Heterogeneous Feature Fusion and Variable Selection
abstract
We propose a novel multiple kernel learning (MKL) algorithm with a group lasso regularizer, called group lasso regularized MKL (GL-MKL), for heterogeneous feature fusion and variable selection. For problems of feature fusion, assigning a group of base kernels for each feature type in an MKL framework provides a robust way in fitting data extracted from different feature domains. Adding a mixed norm constraint (i.e., group lasso) as the regularizer, we can enforce the sparsity at the group/feature level and automatically learn a compact feature set for recognition purposes. More precisely, our GL-MKL determines the optimal base kernels, including the associated weights and kernel parameters, and results in improved recognition performance. Besides, our GL-MKL can also be extended to address heterogeneous variable selection problems. For such problems, we aim to select a compact set of variables (i.e., feature attributes) for comparable or improved performance. Our proposed method does not need to exhaustively search for the entire variable space like prior sequential-based variable selection methods did, and we do not require any prior knowledge on the optimal size of the variable subset either. To verify the effectiveness and robustness of our GL-MKL, we conduct experiments on video and image datasets for heterogeneous feature fusion, and perform variable selection on various UCI datasets.
Yi-Ren Yeh, Ting-Chu Lin, Yung-Yu Chung, Yu-Chiang Frank Wang
IEEE Trans. Multim.1
2011 Locality-constrained group sparse representation for robust face recognition
abstract
This paper presents a novel sparse representation for robust face recognition. We advance both group sparsity and data locality and formulate a unified optimization framework, which produces a locality and group sensitive sparse representation (LGSR) for improved recognition. Empirical results confirm that our LGSR not only outperforms state-of-the-art sparse coding based image classification methods, our approach is robust to variations such as lighting, pose, and facial details (glasses or not), which are typically seen in real-world face recognition problems.
Yu-Wei Chao, Yi-Ren Yeh, Yuh-Jye Lee, Yu-Chiang Frank Wang
ICIP2
2011 Group lasso regularized multiple kernel learning for heterogeneous feature selection
abstract
We propose a novel multiple kernel learning (MKL) algorithm with a group lasso regularizer, called group lasso regularized MKL (GL-MKL), for heterogeneous feature selection. We extend the existing MKL algorithm and impose a mixed ℓ1and ℓ2norm constraint (known as group lasso) as the regularizer. Our GL-MKL determines the optimal base kernels, including the associated weights and kernel parameters, and results in a compact set of features for comparable or improved recognition performance. The use of our GL-MKL avoids the problem of choosing the proper technique to normalize the feature attributes collected from heterogeneous domains (and thus with different properties and distribution ranges). Our approach does not need to exhaustively search for the entire feature space when performing feature selection like prior sequential-based feature selection methods did, and we do not require any prior knowledge on the optimal size of the feature subset either. Comparisons with existing MKL or sequential-based feature selection methods on a variety of datasets confirm the effectiveness of our method in selecting a compact feature subset for comparable or improved classification performance.
Yi-Ren Yeh, Yung-Yu Chung, Ting-Chu Lin, Yu-Chiang Frank Wang
IJCNN1
2011 An iterative algorithm for robust kernel principal component analysis
Hsin-Hsiung Huang, Yi-Ren Yeh
Neurocomputing2
2009 Robust Kernel Principal Component Analysis
abstract
This letter discusses the robustness issue of kernel principal component analysis. A class of new robust procedures is proposed based on eigenvalue decomposition of weighted covariance. The proposed procedures will place less weight on deviant patterns and thus be more resistant to data contamination and model deviation. Theoretical influence functions are derived, and numerical examples are presented as well. Both theoretical and numerical results indicate that the proposed robust method outperforms the conventional approach in the sense of being less sensitive to outliers. Our robust method and results also apply to functional principal component analysis.
Su-Yun Huang, Yi-Ren Yeh, Shinto Eguchi
Neural Comput.2
2009 Nonlinear Dimension Reduction with Kernel Sliced Inverse Regression
abstract
Sliced inverse regression (SIR) is a renowned dimension reduction method for finding an effective low-dimensional linear subspace. Like many other linear methods, SIR can be extended to nonlinear setting via the ldquokernel trick.rdquo The main purpose of this paper is two-fold. We build kernel SIR in a reproducing kernel Hilbert space rigorously for a more intuitive model explanation and theoretical development. The second focus is on the implementation algorithm of kernel SIR for fast computation and numerical stability. We adopt a low-rank approximation to approximate the huge and dense full kernel covariance matrix and a reduced singular value decomposition technique for extracting kernel SIR directions. We also explore kernel SIR's ability to combine with other linear learning algorithms for classification and regression including multiresponse regression. Numerical experiments show that kernel SIR is an effective kernel tool for nonlinear dimension reduction and it can easily combine with other linear algorithms to form a powerful toolkit for nonlinear data analysis.
Yi-Ren Yeh, Su-Yun Huang, Yuh-Jye Lee
IEEE Trans. Knowl. Data Eng.1