Yao-Hung Tsai

dblp:154/3702 · also Yao-Hung Hubert Tsai · DBLP profile ↗
← Back
36ranked-venue papers
17as first author
14since 2021 · last 2024
0000-0001-5312-1875ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 14 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2024 KPConvX: Modernizing Kernel Point Convolution with Kernel Attention
abstract
In the field of deep point cloud understanding, KP-Conv is a unique architecture that uses kernel points to locate convolutional weights in space, instead of relying on Multi-Layer Perceptron (MLP) encodings. While it initially achieved success, it has since been surpassed by recent MLP networks that employ updated designs and training strategies. Building upon the kernel point principle, we present two novel designs: KPConvD (depthwise KP-Conv), a lighter design that enables the use of deeper architectures, and KPConvX, an innovative design that scales the depthwise convolutional weights of KPConvD with kernel attention values. Using KPConvX with a modern architecture and training strategy, we are able to outperform current state-of-the-art approaches on the ScanObjectNN, Scannetv2, and S3DIS datasets. We validate our design choices through ablation studies and release our code and models.
Hugues Thomas, Yao-Hung Tsai, Tim D. Barfoot, Jian Zhang 0050
CVPR2
2024 Efficient Modality Selection in Multimodal Learning
abstract
Multimodal learning aims to learn from data of different modalities by fusing information from heterogeneous sources. Although it is beneficial to learn from more modalities, it is often infeasible to use all available modalities under limited computational resources. Modeling with all available modalities can also be inefficient and unnecessary when information across input modalities overlaps. In this paper, we study the modality selection problem, which aims to select the most useful subset of modalities for learning under a cardinality constraint. To that end, we propose a unified theoretical framework to quantify the learning utility of modalities, and we identify dependence assumptions to flexibly model the heterogeneous nature of multimodal data, which also allows efficient algorithm design. Accordingly, we derive a greedy modality selection algorithm via submodular maximization, which selects the most useful modalities with an optimality guarantee on learning performance. We also connect marginal-contribution-based feature importance scores, such as Shapley value, from the feature selection domain to the context of modality selection, to efficiently compute the importance of individual modality. We demonstrate the efficacy of our theoretical results and modality selection algorithms on 2 synthetic and 4 real-world data sets on a diverse range of multimodal data.
Runxiang Cheng, Gargi Balasubramaniam, Yao-Hung Tsai, Han Zhao 0002
J. Mach. Learn. Res.4
2023 Self-Supervised Object Goal Navigation with In-Situ Finetuning
abstract
A household robot should be able to navigate to target objects without requiring users to first annotate everything in their home. Most current approaches to object navigation do not test on real robots and rely solely on reconstructed scans of houses and their expensively labeled semantic 3D meshes. In this work, our goal is to build an agent that builds self-supervised models of the world via exploration, the same as a child might - thus we (1) eschew the expense of labeled 3D mesh and (2) enable self-supervised in-situ finetuning in the real world. We identify a strong source of self-supervision (Location Consistency - LocCon) that can train all components of an ObjectNav agent, using unannotated simulated houses. Our key insight is that embodied agents can leverage location consistency as a self-supervision signal - collecting images from different views/angles and applying contrastive learning. We show that our agent can perform competitively in the real world and simulation. Our results also indicate that supervised training with 3D mesh annotations causes models to learn simulation artifacts, which are not transferrable to the real world. In contrast, our LocCon shows the most robust transfer in the real world among the set of models we compare to, and that the real-world performance of all models can be further improved with self-supervised LocCon in-situ training.
So Yeon Min, Yao-Hung Tsai, Ali Farhadi, Ruslan Salakhutdinov, Yonatan Bisk, Jian Zhang 0050
IROS2
2022 Learning Weakly-supervised Contrastive Representations
Yao-Hung Tsai, Tianqin Li, Peiyuan Liao, Ruslan Salakhutdinov, Louis-Philippe Morency
ICLR1
2022 Conditional Contrastive Learning with Kernel
Yao-Hung Tsai, Tianqin Li, Martin Q. Ma, Han Zhao 0002, Kun Zhang 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR1
2022 Paraphrasing Is All You Need for Novel Object Captioning
abstract
Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequence-to-sequence training or CIDEr optimization. As a result, we present Paraphrasing-to-Captioning (P2C), a two-stage learning framework for NOC, which would heuristically optimize the output captions via paraphrasing. With P2C, the captioning model first learns paraphrasing from a language model pre-trained on text-only corpus, allowing expansion of the word bank for improving linguistic fluency. To further enforce the output caption sufficiently describing the visual content of the input image, we perform self-paraphrasing for the captioning model with fidelity and adequacy objectives introduced. Since no ground truth captions are available for novel object images during training, our P2C leverages cross-modality (image-text) association modules to ensure the above caption characteristics can be properly preserved. In the experiments, we not only show that our P2C achieves state-of-the-art performances on nocaps and COCO Caption datasets, we also verify the effectiveness and flexibility of our learning framework by replacing language and cross-modality association models for NOC. Implementation details and code are available in the supplementary materials.
Cheng-Fu Yang, Yao-Hung Tsai, Wan-Cyuan Fan, Ruslan Salakhutdinov, Louis-Philippe Morency, Frank Wang
NeurIPS2
2022 Feature-Robust Optimal Transport for High-Dimensional Data
Mathis Petrovich, Chao Liang 0002, Ryoma Sato, Yanbin Liu 0003, Yao-Hung Tsai, Linchao Zhu, Yi Yang 0001, Ruslan Salakhutdinov, Makoto Yamada
ECML/PKDD (5)5
2022 Greedy modality selection via approximate submodular maximization
abstract
Multimodal learning considers learning from multi-modality data, aiming to fuse heterogeneous sources of information. However, it is not always feasible to leverage all available modalities due to memory constraints. Further, training on all the modalities may be inefficient when redundant information exists within data, such as different subsets of modalities providing similar performance. In light of these challenges, we study modality selection, intending to efficiently select the most informative and complementary modalities under certain computational constraints. We formulate a theoretical framework for optimizing modality selection in multimodal learning and introduce a utility measure to quantify the benefit of selecting a modality. For this optimization problem, we present efficient algorithms when the utility measure exhibits monotonicity and approximate submodularity. We also connect the utility measure with existing Shapley-value-based feature importance scores. Last, we demonstrate the efficacy of our algorithm on synthetic (Patch-MNIST) and real-world (PEMS-SF, CMU-MOSI) datasets.
Runxiang Cheng, Gargi Balasubramaniam, Yao-Hung Tsai, Han Zhao 0002
UAI4
2022 A 0.0067-mm2 12-bit 20-MS/s SAR ADC Using Digital Place-and-Route Tools in 40-nm CMOS
abstract
A 12-bit 20-MS/s asynchronous successive approximation register (SAR) analog-to-digital converter (ADC) is presented by using the digital place-and-route (DPR) tools. The macrocells for the capacitive digital-to-analog converter, the bootstrapped switch, and the dynamic comparator are presented. The custom standard cells for the dynamic SAR logic are also presented. By using the macrocells and the custom standard ones, the layout of this SAR ADC is completed by using the DPR tools. Several techniques are presented to improve the parasitic capacitances, the current density of the metal interconnections, and the nonideal effects caused by the DPR tools. This SAR ADC is fabricated in 40-nm CMOS technology and its active area is 0.0067 mm2. To compare with the full-custom method, the proposed DPR flow has speeded up by a factor of 288 to complete the interconnection wires. Its power dissipation is 363$\mu \text{W}$at 20 MS/s and the calculated Walden FoM is 23 fJ/c. step at Nyquist frequency.
Yao-Hung Tsai, Shen-Iuan Liu
IEEE Trans. Very Large Scale Integr. Syst.1
2021 Hubert: How Much Can a Bad Teacher Benefit ASR Pre-Training?
abstract
Compared to vision and language applications, self-supervised pre-training approaches for ASR are challenged by three unique problems: (1) There are multiple sound units in each input utterance, (2) With audio-only pre-training, there is no lexicon of sound units, and (3) Sound units have variable lengths with no explicit segmentation. In this paper, we propose the Hidden-Unit BERT (HUBERT) model which utilizes a cheap k-means clustering step to provide aligned target labels for pre-training of a BERT model. A key ingredient of our approach is applying the predictive loss over the masked regions only. This allows the pre-training stage to benefit from the consistency of the unsupervised teacher rather that its intrinsic quality. Starting with a simple k-means teacher of 100 cluster, and using two iterations of clustering, the HUBERT model matches the state-of-the-art wav2vec 2.0 performance on the ultra low-resource Libri-light 10h, 1h, 10min supervised subsets.
Wei-Ning Hsu, Yao-Hung Tsai, Benjamin Bolte, Ruslan Salakhutdinov, Abdel-rahman Mohamed
ICASSP2
2021 Self-supervised Learning from a Multi-view Perspective
Yao-Hung Tsai, Yue Wu 0001, Ruslan Salakhutdinov, Louis-Philippe Morency
ICLR1
2021 Self-supervised Representation Learning with Relative Predictive Coding
Yao-Hung Tsai, Martin Q. Ma, Muqiao Yang, Han Zhao 0002, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR1
2021 LSMI-Sinkhorn: Semi-supervised Mutual Information Estimation with Optimal Transport
Yanbin Liu 0003, Makoto Yamada, Yao-Hung Tsai, Tam Le, Ruslan Salakhutdinov, Yi Yang 0001
ECML/PKDD (1)3
2021 HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
abstract
Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound units have variable lengths with no explicit segmentation. To deal with these three problems, we propose the Hidden-Unit BERT (HuBERT) approach for self-supervised speech representation learning, which utilizes an offline clustering step to provide aligned target labels for a BERT-like prediction loss. A key ingredient of our approach is applying the prediction loss over the masked regions only, which forces the model to learn a combined acoustic and language model over the continuous inputs. HuBERT relies primarily on the consistency of the unsupervised clustering step rather than the intrinsic quality of the assigned cluster labels. Starting with a simple k-means teacher of 100 clusters, and using two iterations of clustering, the HuBERT model either matches or improves upon the state-of-the-art wav2vec 2.0 performance on the Librispeech (960h) and Libri-light (60,000h) benchmarks with 10min, 1h, 10h, 100h, and 960h fine-tuning subsets. Using a 1B parameter model, HuBERT shows up to 19% and 13% relative WER reduction on the more challenging dev-other and test-other evaluation subsets.
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, Abdel-rahman Mohamed
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis
abstract
, which dynamically adjusts weights between input modalities and output representations differently for each input sample. Multimodal routing can identify relative importance of both individual modalities and cross-modality features. Moreover, the weight assignment by routing allows us to interpret modality-prediction relationships not only globally (i.e. general trends over the whole dataset), but also locally for each single input sample, mean-while keeping competitive performance compared to state-of-the-art methods.
Yao-Hung Tsai, Martin Q. Ma, Muqiao Yang, Ruslan Salakhutdinov, Louis-Philippe Morency
EMNLP (1)1
2020 Complex Transformer: A Framework for Modeling Complex-Valued Sequence
abstract
While deep learning has received a surge of interest in a variety of fields in recent years, major deep learning models barely use complex numbers. However, speech, signal and audio data are naturally complex-valued after Fourier Transform, and studies have shown a potentially richer representation of complex nets. In this paper, we propose a Complex Transformer, which incorporates the transformer model as a backbone for sequence modeling; we also develop attention and encoder-decoder network operating for complex input. The model achieves state-of-the-art performance on the MusicNet dataset and an In-phase Quadrature (IQ) signal dataset. The GitHub implementation to reproduce the experimental results is available at https://github.com/muqiaoy/dl_signal.
Muqiao Yang, Martin Q. Ma, Dongyu Li, Yao-Hung Tsai, Ruslan Salakhutdinov
ICASSP4
2020 Capsules with Inverted Dot-Product Attention Routing
Yao-Hung Tsai, Nitish Srivastava, Hanlin Goh, Ruslan Salakhutdinov
ICLR1
2020 Neural Methods for Point-wise Dependency Estimation
abstract
Since its inception, the neural estimation of mutual information (MI) has demonstrated the empirical success of modeling expected dependency between high-dimensional random variables. However, MI is an aggregate statistic and cannot be used to measure point-wise dependency between different events. In this work, instead of estimating the expected dependency, we focus on estimating point-wise dependency (PD), which quantitatively measures how likely two outcomes co-occur. We show that we can naturally obtain PD when we are optimizing MI neural variational bounds. However, optimizing these bounds is challenging due to its large variance in practice. To address this issue, we develop two methods (free of optimizing MI variational bounds): Probabilistic Classifier and Density-Ratio Fitting. We demonstrate the effectiveness of our approaches in 1) MI estimation, 2) self-supervised representation learning, and 3) cross-modal retrieval task.
Yao-Hung Tsai, Han Zhao 0002, Makoto Yamada, Louis-Philippe Morency, Ruslan Salakhutdinov
NeurIPS1
2019 Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization
abstract
There has been an increased interest in multimodal language processing including multimodal dialog, question answering, sentiment analysis, and speech recognition.However, naturally occurring multimodal data is often imperfect as a result of imperfect modalities, missing entries or noise corruption.To address these concerns, we present a regularization method based on tensor rank minimization.Our method is based on the observation that high-dimensional multimodal time series data often exhibit correlations across time and modalities which leads to low-rank tensor representations.However, the presence of noise or incomplete values breaks these correlations and results in tensor representations of higher rank.We design a model to learn such tensor representations and effectively regularize their rank.Experiments on multimodal language data show that our model achieves good results across various levels of imperfection.
Paul Pu Liang, Zhun Liu, Yao-Hung Tsai, Qibin Zhao, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL (1)3
2019 Multimodal Transformer for Unaligned Multimodal Language Sequences
abstract
Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data non-alignment due to variable sampling rates for the sequences from each modality; and 2) long-range dependencies between elements across modalities. In this paper, we introduce the Multimodal Transformer (MulT) to generically address the above issues in an end-to-end manner without explicitly aligning the data. At the heart of our model is the directional pairwise cross-modal attention, which attends to interactions between multimodal sequences across distinct time steps and latently adapt streams from one modality to another. Comprehensive experiments on both aligned and non-aligned multimodal time-series show that our model outperforms state-of-the-art methods by a large margin. In addition, empirical analysis suggests that correlated crossmodal signals are able to be captured by the proposed crossmodal attention mechanism in MulT.
Yao-Hung Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov
ACL (1)1
2019 Video Relationship Reasoning Using Gated Spatio-Temporal Energy Graph
abstract
Visual relationship reasoning is a crucial yet challenging task for understanding rich interactions across visual concepts. For example, a relationship {man, open, door} involves a complex relation {open} between concrete entities {man, door}. While much of the existing work has studied this problem in the context of still images, understanding visual relationships in videos has received limited attention. Due to their temporal nature, videos enable us to model and reason about a more comprehensive set of visual relationships, such as those requiring multiple (temporal) observations (e.g., {man, lift up, box} vs. {man, put down, box}), as well as relationships that are often correlated through time (e.g., {woman, pay, money} followed by {woman, buy, coffee}). In this paper, we construct a Conditional Random Field on a fully-connected spatiotemporal graph that exploits the statistical dependency between relational entities spatially and temporally. We introduce a novel gated energy function parametrization that learns adaptive relations conditioned on visual observations. Our model optimization is computationally efficient, and its space computation complexity is significantly amortized through our proposed parameterization. Experimental results on benchmark video datasets (ImageNet Video and Charades) demonstrate state-of-the-art performance across three standard relationship reasoning tasks: Detection, Tagging, and Recognition.
Yao-Hung Tsai, Santosh Kumar Divvala, Louis-Philippe Morency, Ruslan Salakhutdinov, Ali Farhadi
CVPR1
2019 Transformer Dissection: An Unified Understanding for Transformer's Attention via the Lens of Kernel
abstract
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, Ruslan Salakhutdinov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yao-Hung Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, Ruslan Salakhutdinov
EMNLP/IJCNLP (1)1
2019 Learning Factorized Multimodal Representations
Yao-Hung Tsai, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR (Poster)1
2019 Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator
Makoto Yamada, Denny Wu, Yao-Hung Tsai, Hirofumi Ohta, Ruslan Salakhutdinov, Ichiro Takeuchi, Kenji Fukumizu
ICLR (Poster)3
2019 Learning Neural Networks with Adaptive Regularization
abstract
Feed-forward neural networks can be understood as a combination of an intermediate representation and a linear hypothesis. While most previous works aim to diversify the representations, we explore the complementary direction by performing an adaptive and data-dependent regularization motivated by the empirical Bayes method. Specifically, we propose to construct a matrix-variate normal prior (on weights) whose covariance matrix has a Kronecker product structure. This structure is designed to capture the correlations in neurons through backpropagation. Under the assumption of this Kronecker factorization, the prior encourages neurons to borrow statistical strength from one another. Hence, it leads to an adaptive and data-dependent regularization when training networks on small datasets. To optimize the model, we present an efficient block coordinate descent algorithm with analytical solutions. Empirically, we demonstrate that the proposed method helps networks converge to local optima with smaller stable ranks and spectral norms. These properties suggest better generalizations and we present empirical results to support this expectation. We also verify the effectiveness of the approach on multiclass classification and multitask regression problems with various network structures. Our code is publicly available at:~\url{https://github.com/yaohungt/Adaptive-Regularization-Neural-Network}.
Han Zhao 0002, Yao-Hung Tsai, Ruslan Salakhutdinov, Geoffrey J. Gordon
NeurIPS2
2019 Transfer Neural Trees: Semi-Supervised Heterogeneous Domain Adaptation and Beyond
abstract
Heterogeneous domain adaptation (HDA) addresses the task of associating data not only across dissimilar domains but also described by different types of features. Inspired by the recent advances of neural networks and deep learning, we propose a deep leaning model of Transfer Neural Trees (TNT), which jointly solves cross-domain feature mapping, adaptation, and classification in a unified architecture. As the prediction layer in TNT, we introduce Transfer Neural Decision Forest (Transfer- NDF), which is able to learn the neurons in TNT for adaptation by stochastic pruning. In order to handle semi-supervised HDA, a unique embedding loss term is introduced to TNT for preserving prediction and structural consistency between labeled and unlabeled target-domain data. We further show that our TNT can be extended to zero shot learning for associating image and attribute data with promising performance. Finally, experiments on different classification tasks across features, datasets, and modalities would verify the effectiveness of our TNT.
Wei-Yu Chen, Tzu-Ming Harry Hsu, Yao-Hung Tsai, Ming-Syan Chen, Yu-Chiang Frank Wang
IEEE Trans. Image Process.3
2017 Learning Robust Visual-Semantic Embeddings
abstract
Many of the existing methods for learning joint embedding of images and text use only supervised information from paired images and its textual attributes. Taking advantage of the recent success of unsupervised learning in deep neural networks, we propose an end-to-end learning framework that is able to extract more robust multi-modal representations across domains. The proposed method combines representation learning models (i.e., auto-encoders) together with cross-domain learning criteria (i.e., Maximum Mean Discrepancy loss) to learn joint embeddings for semantic and visual features. A novel technique of unsupervised-data adaptation inference is introduced to construct more comprehensive embeddings for both labeled and unlabeled data. We evaluate our method on Animals with Attributes and Caltech-UCSD Birds 200-2011 dataset with a wide range of applications, including zero and few-shot image recognition and retrieval, from inductive to transductive settings. Empirically, we show that our frame-work improves over the current state of the art on many of the considered tasks.
Yao-Hung Tsai, Liang-Kang Huang, Ruslan Salakhutdinov
ICCV1
2016 Domain-Constraint Transfer Coding for Imbalanced Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) deals with the task that labeled training and unlabeled test data collected from source and target domains, respectively. In this paper, we particularly address the practical and challenging scenario of imbalanced cross-domain data. That is, we do not assume the label numbers across domains to be the same, and we also allow the data in each domain to be collected from multiple datasets/sub-domains. To solve the above task of imbalanced domain adaptation, we propose a novel algorithm of Domain-constraint Transfer Coding (DcTC). Our DcTC is able to exploit latent subdomains within and across data domains, and learns a common feature space for joint adaptation and classification purposes. Without assuming balanced cross-domain data as most existing UDA approaches do, we show that our method performs favorably against state-of-the-art methods on multiple cross-domain visual classification tasks.
Yao-Hung Tsai, Cheng-An Hou, Wei-Yu Chen, Yi-Ren Yeh, Yu-Chiang Frank Wang
AAAI1
2016 Learning Cross-Domain Landmarks for Heterogeneous Domain Adaptation
abstract
While domain adaptation (DA) aims to associate the learning tasks across data domains, heterogeneous domain adaptation (HDA) particularly deals with learning from cross-domain data which are of different types of features. In other words, for HDA, data from source and target domains are observed in separate feature spaces and thus exhibit distinct distributions. In this paper, we propose a novel learning algorithm of Cross-Domain Landmark Selection (CDLS) for solving the above task. With the goal of deriving a domain-invariant feature subspace for HDA, our CDLS is able to identify representative cross-domain data, including the unlabeled ones in the target domain, for performing adaptation. In addition, the adaptation capabilities of such cross-domain landmarks can be determined accordingly. This is the reason why our CDLS is able to achieve promising HDA performance when comparing to state-of-the-art HDA methods. We conduct classification experiments using data across different features, domains, and modalities. The effectiveness of our proposed method can be successfully verified.
Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
CVPR1
2016 Transfer Neural Trees for Heterogeneous Domain Adaptation
Wei-Yu Chen, Tzu-Ming Harry Hsu, Yao-Hung Tsai, Yu-Chiang Frank Wang, Ming-Syan Chen
ECCV (5)3
2016 Heterogeneous domain adaptation with label and structure consistency
abstract
Domain adaptation is a challenging task, since it associates data collected from different domains or exhibiting distinct distributions. In this paper, we particularly focus on adapting cross-domain data with distinct feature dimensions or representations. Thus, this is referred to as the task of heterogeneous domain adaptation (HDA). To solve HDA, we propose Label and Structure-consistent Unilateral Projection (LS-UP) that transforms source-domain data to the target domain, with the goal of matching cross-domain data distribution and preserving data structure after projection. The main contribution of our work is its ability in relating cross-domain data with different feature representations. We evaluate our LS-UP for HDA on two different cross-domain classification problems, and we show that our method would perform favorably against state-of-the-art approaches.
Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICASSP1
2016 Recognizing heterogeneous cross-domain data via generalized joint distribution adaptation
abstract
In this paper, we propose a novel algorithm of Generalized Joint Distribution Adaptation (G-JDA) for heterogeneous domain adaptation (HDA), which associates and recognizes cross-domain data observed in different feature spaces (and thus with different dimensionality). With the objective to derive a domain-invariant feature subspace for relating source and target-domain data, our G-JDA learns a pair of feature projection matrices (one for each domain), which allows us to eliminate the difference between projected cross-domain heterogeneous data by matching their marginal and class-conditional distributions. We conduct experiments on cross-domain classification tasks using data across different features, datasets, and modalities. We confirm that our G-JDA would perform favorably against state-of-the-art HDA approaches.
Yuan-Ting Hsieh, Shi-Yen Tao, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICME3
2016 Unsupervised Domain Adaptation With Label and Structural Consistency
abstract
Unsupervised domain adaptation deals with scenarios in which labeled data are available in the source domain, but only unlabeled data can be observed in the target domain. Since the classifiers trained by source-domain data would not be expected to generalize well in the target domain, how to transfer the label information from source to target-domain data is a challenging task. A common technique for unsupervised domain adaptation is to match cross-domain data distributions, so that the domain and distribution differences can be suppressed. In this paper, we propose to utilize the label information inferred from the source domain, while the structural information of the unlabeled target-domain data will be jointly exploited for adaptation purposes. Our proposed model not only reduces the distribution mismatch between domains, improved recognition of target-domain data can be achieved simultaneously. In the experiments, we will show that our approach performs favorably against the state-of-the-art unsupervised domain adaptation methods on benchmark data sets. We will also provide convergence, sensitivity, and robustness analysis, which support the use of our model for cross-domain classification.
Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
IEEE Trans. Image Process.2
2015 Unsupervised Domain Adaptation with Imbalanced Cross-Domain Data
abstract
We address a challenging unsupervised domain adaptation problem with imbalanced cross-domain data. For standard unsupervised domain adaptation, one typically obtains labeled data in the source domain and only observes unlabeled data in the target domain. However, most existing works do not consider the scenarios in which either the label numbers across domains are different, or the data in the source and/or target domains might be collected from multiple datasets. To address the aforementioned settings of imbalanced cross-domain data, we propose Closest Common Space Learning (CCSL) for associating such data with the capability of preserving label and structural information within and across domains. Experiments on multiple cross-domain visual classification tasks confirm that our method performs favorably against state-of-the-art approaches, especially when imbalanced cross-domain data are presented.
Tzu-Ming Harry Hsu, Wei-Yu Chen, Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang
ICCV4
2014 Person-specific domain adaptation with applications to heterogeneous face recognition
abstract
Heterogeneous face recognition (HFR) is a practical yet challenging task in which gallery and probe face images are collected in terms of different modalities or features (e.g., sketch vs. photo). In this paper, we present a person-specific domain adaptation framework for HFR. By utilizing the subjects not of interest (i.e., those not to be recognized), we first derive a common feature space using their cross-domain face images, with the goal of eliminating differences between image modalities. To generalize our feature space for representing and recognizing the subjects of interest, we advocate the construction of person-specific domain adaptation model in this space, so that the classifiers (trained by the gallery images) are able to achieve satisfactory recognition performance. In our experiments, we consider sketch-to-photo and near-infrared (NIR) to visible spectrum (VIS) face recognition problems for evaluating the effectiveness of our method.
Yao-Hung Tsai, Hung-Ming Hsu, Cheng-An Hou, Yu-Chiang Frank Wang
ICIP1
2014 Stable pose tracking from a planar target with an analytical motion model in real-time applications
abstract
Object pose tracking from a camera is a well-developed method in computer vision. In theory, the pose can be determined uniquely from a calibrated camera. However, in practice, most real-time pose estimation algorithms experience pose ambiguity. We consider that pose ambiguity, i.e., the detection of two distinct local minima according to an error function, is caused by a geometric illusion. In this case, both ambiguous poses are plausible, but we cannot select the pose with the minimum error as the final pose. Thus, we developed a real-time algorithm for correct pose estimation for a planar target object using an analytical motion model. Our experimental results showed that the proposed algorithm effectively reduced the effects of pose jumping and pose jittering. To the best of our knowledge, this is the first approach to address the pose ambiguity problem using an analytical motion model in real-time applications.
Po-Chen Wu, Yao-Hung Tsai, Shao-Yi Chien
MMSP2