EDBT 2026 Demo / reviewers in the wild / expert
Yi-Hsuan Tsai
dblp:142/2924
· DBLP profile ↗
83ranked-venue papers
11as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 8 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 10 first-author · 28 since 2021Systems, architecture and hardware · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
Yi-Lun Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
ICPR (9) | 2 |
| 2026 | Editing 3D Scenes via Text Prompts Without RetrainingabstractNumerous diffusion models have been developed for 2D image synthesis and editing, and recently they are extended to 3D scene editing tasks. However, editing 3D scenes is still in its early stages, and the challenges of scene representations and multi-view consistency need to be addressed. A notable limitation of existing approaches is the need for specific modules for different edits and model retraining for each scene. To tackle these issues, we propose a novel and versatile text-driven 3D scene editing method, termed DN2N, which allows for the direct acquisition of the editing results without the requirement for retraining. Our method employs off-the-shelf text-based editing models of 2D images to modify the multi-view images of a 3D scene. A content filtering process is then applied to discard poorly edited images that disrupt 3D consistency. We consider the remaining inconsistency as a problem of removing noise perturbations and solve it by generating data with similar perturbation characteristics for training. We develop a versatile NeRF model structure and propose two novel cross-view regularization terms to help the DN2N mitigate these perturbations. Empirical results show that our method achieves multiple editing types based solely on text prompts, including but not limited to appearance editing, weather transition, object changing, and style transfer. Most importantly, DN2N exhibits a versatility of editing capabilities, eliminating the need to customize or retrain editing models for specific scenes or editing types. Namely, DN2N achieves comparable total editing time to the 3DGS-based editing method, enhancing its practical value. Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang 0033, Shuchang Zhou 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian SplattingabstractBird’s-eye view (BEV) perception has gained significant attention because it provides a unified representation to fuse multiple view images and enables a wide range of downstream autonomous driving tasks, such as forecasting and planning. Recent state-of-the-art models utilize projection-based methods which formulate BEV perception as query learning to bypass explicit depth estimation. While we observe promising advancements in this paradigm, they still fall short of real-world applications because of the lack of uncertainty modeling and expensive computational requirement. In this work, we introduce GaussianLSS, an uncertainty-aware BEV perception framework that revisits the unprojection-based method, specifically the Lift-Splat-Shoot (LSS) paradigm, and enhances it with depth uncertainty modeling. Our GaussianLSS represents spatial dispersion by learning a soft depth mean and computing the variance of the depth distribution, which implicitly captures object extents. We then transform the depth distribution into 3D Gaussians and rasterize them to construct uncertainty-aware BEV features. We evaluate GaussianLSS on the nuScenes dataset, achieving state-of-the-art performance compared to unprojection-based methods. In particular, it provides significant advantages in speed, running 2x faster, and in memory efficiency, using 0.3x less memory compared to projection-based methods, while achieving competitive performance with only a 0.7% IoU difference. See our project page for more details: https://hcis-lab.github.io/GaussianLSS/. Shu-Wei Lu, Yi-Hsuan Tsai, Yi-Ting Chen 0001 |
CVPR | 2 |
| 2025 | MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Shuchang Zhou 0001, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang 0001 |
ICCV | 4 |
| 2025 | What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningabstractUnderstanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling the temporal order of actions, but has not explicitly learned the state changes (scene transformations). In this work, we study procedure-aware video representation learning by incorporating state-change descriptions generated by Large Language Models (LLMs) as supervision signals for video encoders. Moreover, we generate state-change counterfactuals that simulate hypothesized failure outcomes, allowing models to learn by imagining unseen "What if" scenarios. This counterfactual reasoning facilitates the model's ability to understand the cause and effect of each step in an activity. We conduct extensive experiments on procedure-aware tasks, including temporal action segmentation, error detection, action phase classification, frame retrieval, multi-instance retrieval, and action recognition. Our results demonstrate the effectiveness of the proposed state-change descriptions and their counterfactuals, and achieve significant improvements on multiple tasks. Chi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen 0001, David Crandall, Yi-Hsuan Tsai |
ICCV | 6 |
| 2025 | Ranking-aware adapter for text-driven image ordering with CLIPabstractRecent advances in vision-language models (VLMs) have made significant progress in downstream tasks that require quantitative concepts such as facial age estimation and image quality assessment, enabling VLMs to explore applications like image ranking and retrieval. However, existing studies typically focus on the reasoning based on a single image and heavily depend on text prompting, limiting their ability to learn comprehensive understanding from multiple images. To address this, we propose an effective yet efficient approach that reframes the CLIP model into a learning-to-rank task and introduces a lightweight adapter to augment CLIP for text-guided image ranking. Specifically, our approach incorporates learnable prompts to adapt to new instructions for ranking purposes and an auxiliary branch with ranking-aware attention, leveraging text-conditioned visual differences for additional supervision in image ranking. Our ranking-aware adapter consistently outperforms fine-tuned CLIPs on various tasks and achieves competitive results compared to state-of-the-art models designed for specific tasks like facial age estimation and image quality assessment. Overall, our approach primarily focuses on ranking images with a single instruction, which provides a natural and generalized way of learning from visual differences across images, bypassing the need for extensive text prompts tailored to individual tasks. Wei-Hsiang Yu, Yen-Yu Lin, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
ICLR | 4 |
| 2025 | Low-Rank Head Avatar Personalization with RegistersabstractWe introduce a novel method for low-rank personalization of a generic model for head avatar generation. Prior work proposes generic models that achieve high-quality face animation by leveraging large-scale datasets of multiple identities. However, such generic models usually fail to synthesize unique identity-specific details, since they learn a general domain prior. To adapt to specific subjects, we find that it is still challenging to capture high-frequency facial details via popular solutions like low-rank adaptation (LoRA). This motivates us to propose a specific architecture, a Register Module, that enhances the performance of LoRA, while requiring only a small number of parameters to adapt to an unseen identity. Our module is applied to intermediate features of a pre-trained model, storing and re-purposing information in a learnable 3D feature space. To demonstrate the efficacy of our personalization method, we collect a dataset of talking videos of individuals with distinctive facial details, such as wrinkles and tattoos. Our approach faithfully captures unseen faces, outperforming existing methods quantitatively and qualitatively. Sai Tanmay Reddy Chakkera, Aggelina Chatziagapi, Chen-Ping Yu, Yi-Hsuan Tsai, Dimitris Samaras |
NeurIPS | 5 |
| 2025 | IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video GenerationabstractAlthough diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appearance across frames. To address this limitation, we propose IllumiCraft, an end-to-end diffusion framework accepting three complementary inputs: (1) high-dynamic-range (HDR) video maps for detailed lighting control; (2) synthetically relit frames with randomized illumination changes (optionally paired with a static background reference image) to provide appearance cues; and (3) 3D point tracks that capture precise 3D geometry information. By integrating the lighting, appearance, and geometry cues within a unified diffusion architecture, IllumiCraft generates temporally coherent videos aligned with user-defined prompts. It supports the background-conditioned and text-conditioned video relighting and provides better fidelity than existing controllable video generation methods. Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark, Ming-Hsuan Yang 0001 |
NeurIPS | 3 |
| 2025 | uLayout: Unified Room Layout Estimation for Perspective and Panoramic ImagesabstractWe present uLayout, a unified model for estimating room layout geometries from both perspective and panoramic images, whereas traditional solutions require different model designs for each image type. The key idea of our solution is to unify both domains into the equirectangular projection, particularly, allocating perspective images into the most suitable latitude coordinate to effectively exploit both domains seamlessly. To address the Field-of-View (FoV) difference between the input domains, we design uLayout with a shared feature extractor with an extra 1D-Convolution layer to condition each domain input differently. This conditioning allows us to efficiently formulate a column-wise feature regression problem regardless of the FoV input. This simple yet effective approach achieves competitive performance with current state-of-the-art solutions and shows for the first time a single end-to-end model for both domains. Extensive experiments in the real-world datasets, LSUN, Matterport3D, PanoContext, and Stanford 2D-3D evidence the contribution of our approach. Code is available at https://github.com/JonathanLee112/uLayout Bolivar Solarte, Chin-Hsuan Wu, Jin-Cheng Jhang, Fu-En Wang, Yi-Hsuan Tsai, Min Sun 0001 |
WACV | 6 |
| 2024 | PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionabstractRecent temporal LiDAR-based 3D object detectors achieve promising performance based on the two-stage proposal-based approach. They generate 3D box candidates from the first-stage dense detector, followed by different temporal aggregation methods. However, these approaches require per-frame objects or whole point clouds, posing challenges related to memory bank utilization. Moreover, point clouds and trajectory features are combined solely based on concatenation, which may neglect effective interactions between them. In this paper, we propose a point-trajectory transformer with long short-term memory for efficient temporal 3D object detection. To this end, we only utilize point clouds of current-frame objects and their historical trajectories as input to minimize the memory bank storage requirement. Furthermore, we introduce modules to encode trajectory features, focusing on long short-term and future-aware perspectives, and then effectively aggregate them with point cloud features. We conduct extensive experiments on the large-scale Waymo dataset to demon-strate that our approach performs well against state-of-the-art methods. Code and models will be made publicly available at https://github.com/kuanchihhuang/Ptt. Kuan-Chih Huang, Weijie Lyu, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
CVPR | 4 |
| 2024 | Action-Slot: Visual Action-Centric Representations for Multi-Label Atomic Activity Recognition in Traffic ScenesabstractIn this paper, we study multi-label atomic activity recognition. Despite the notable progress in action recognition, it is still challenging to recognize atomic activities due to a deficiency in holistic understanding of both multiple road users' motions and their contextual information. In this paper, we introduce Action-slot, a slot attention-based approach that learns visual action-centric representations, capturing both motion and contextual information. Our key idea is to design action slots that are capable of paying attention to regions where atomic activities occur, without the need for explicit perception guidance. To further enhance slot attention, we introduce a background slot that competes with action slots, aiding the training process in avoiding un-necessary focus on background regions devoid of activities. Yet, the imbalanced class distribution in the existing dataset hampers the assessment of rare activities. To address the limitation, we collect a synthetic dataset called TACO, which is four times larger than OATS and features a balanced distribution of atomic activities. To validate the effectiveness of our method, we conduct comprehensive experiments and ablation studies against various action recognition baselines. We also show that the performance of multi-label atomic activity recognition on real-world datasets can be improved by pretraining representations on TACO. Our source code, dataset, and visualization videos are available at https://hcis-lab.github.io/Action-slot/. Chi-Hsi Kung, Shu-Wei Lu, Yi-Hsuan Tsai, Yi-Ting Chen 0001 |
CVPR | 3 |
| 2024 | Text-Driven Image Editing via Learnable RegionsabstractLanguage has emerged as a natural interface for image editing. In this paper, we introduce a method for region- based image editing driven by textual prompts, without the need for user-provided masks or sketches. Specifically, our approach leverages an existing pre-trained text-to-image model and introduces a bounding box generator to iden-tify the editing regions that are aligned with the textual prompts. We show that this simple approach enables flex-ible editing that is compatible with current image generation models, and is able to handle complex prompts featuring multiple objects, complex sentences, or lengthy para- graphs. We conduct an extensive user study to compare our method against state-of-the-art methods. The experiments demonstrate the competitive performance of our method in manipulating images with high fidelity and realism that correspond to the provided language descriptions. Our project webpage can be found at: https://yuanzelin.me/LearnableRegions_page. Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Lu Jiang 0004, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2024 | Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001, Ming-Hsuan Yang 0001 |
ECCV (42) | 3 |
| 2024 | Weakly Supervised 3D Object Detection via Multi-level Visual Guidance
Kuan-Chih Huang, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
ECCV (1) | 2 |
| 2024 | Self-training Room Layout Estimation via Geometry-Aware Ray-Casting
Bolivar Solarte, Chin-Hsuan Wu, Jin-Cheng Jhang, Yi-Hsuan Tsai, Min Sun 0001 |
ECCV (45) | 5 |
| 2023 | Multimodal Prompting with Missing Modalities for Visual RecognitionabstractIn this paper, we tackle two challenges in multimodal learning for visual recognition: 1) when missing-modality occurs either during training or testing in real-world situations; and 2) when the computation resources are not available to finetune on heavy transformer models. To this end, we propose to utilize prompt learning and mitigate the above two challenges together. Specifically, our modality-missing-aware prompts can be plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 1% learnable parameters compared to training the entire model. We further explore the effect of different prompt configurations and analyze the robustness to missing modality. Extensive experiments are conducted to show the effectiveness of our prompt learning framework that improves the performance under various missing-modality cases, while alleviating the requirement of heavy model retraining. Code is available.11https://github.com/YiLunLee/missing_aware_prompts Yi-Lun Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Chen-Yu Lee |
CVPR | 2 |
| 2023 | Delving into Motion-Aware Matching for Monocular 3D Object TrackingabstractRecent advances of monocular 3D object detection facilitate the 3D multi-object tracking task based on lowcost camera sensors. In this paper, we find that the motion cue of objects along different time frames is critical in 3D multi-object tracking, which is less explored in existing monocular-based approaches. To this end, we propose MoMA-M3T, a framework that mainly consists of three motion-aware components. First, we represent the possible movement of an object related to all object tracklets in the feature space as its motion features. Then, we further model the historical object tracklet along the time frame in a spatial-temporal perspective via a motion transformer. Finally, we propose a motion-aware matching module to associate historical object tracklets and current observations as final tracking results. We conduct extensive experiments on the nuScenes and KITTI datasets to demonstrate that our MoMA-M3T achieves competitive performance against state-of-the-art methods. Moreover, the proposed tracker is flexible and can be easily plugged into existing image-based 3D object detectors without re-training. Code and models are available at https://github.com/kuanchihhuang/MoMA-M3T. Kuan-Chih Huang, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
ICCV | 3 |
| 2023 | Diffusion-SS3D: Diffusion Model for Semi-supervised 3D Object DetectionabstractSemi-supervised object detection is crucial for 3D scene understanding, efficiently addressing the limitation of acquiring large-scale 3D bounding box annotations. Existing methods typically employ a teacher-student framework with pseudo-labeling to leverage unlabeled point clouds. However, producing reliable pseudo-labels in a diverse 3D space still remains challenging. In this work, we propose Diffusion-SS3D, a new perspective of enhancing the quality of pseudo-labels via the diffusion model for semi-supervised 3D object detection. Specifically, we include noises to produce corrupted 3D object size and class label distributions, and then utilize the diffusion model as a denoising process to obtain bounding box outputs. Moreover, we integrate the diffusion model into the teacher-student framework, so that the denoised bounding boxes can be used to improve pseudo-label generation, as well as the entire semi-supervised learning process. We conduct experiments on the ScanNet and SUN RGB-D benchmark datasets to demonstrate that our approach achieves state-of-the-art performance against existing methods. We also present extensive analysis to understand how our diffusion model design affects performance in semi-supervised learning. The source code will be available at https://github.com/luluho1208/Diffusion-SS3D. Cheng-Ju Ho, Chen-Hsuan Tai, Yen-Yu Lin, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
NeurIPS | 5 |
| 2023 | Data Efficient Incremental Learning via Attentive Knowledge ReplayabstractClass-incremental learning (CIL) tackles the problem of continuously optimizing a classification model to support growing number of classes, where the data of novel classes arrive in streams. Recent works propose to use representative exemplars of learnt classes, and replay the knowledge of them afterward under certain memory constraints. However, training on a fixed set of exemplars with an imbalanced proportion to the new data leads to strong biases in the trained models. In this paper, we propose an attentive knowledge replay framework to refresh the knowledge of previously learnt classes during incremental learning, which generates virtual training samples by blending between pairs of data. Particularly, we design an attention module that learns to predict the adaptive blending weights in accordance with their relative importance to the overall objective, where the importance is derived from the change of the image features over incremental phases. Our strategy of attentive knowledge replay encourages the model to learn smoother decision boundaries and thus improves its generalization beyond memorizing the exemplars. We validate our design in a standard class-incremental learning setup and demonstrate its flexibility in various settings. Yi-Lun Lee, Dian-Shan Chen, Chen-Yu Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
SMC | 4 |
| 2023 | BiFuse++: Self-Supervised and Efficient Bi-Projection Fusion for 360° Depth EstimationabstractDue to the rise of spherical cameras, monocular 360$^\circ$depth estimation becomes an important technique for many applications (e.g., autonomous systems). Thus, state-of-the-art frameworks for monocular 360$^\circ$depth estimation such as bi-projection fusion in BiFuse are proposed. To train such a framework, a large number of panoramas along with the corresponding depth ground truths captured by laser sensors are required, which highly increases the cost of data collection. Moreover, since such a data collection procedure is time-consuming, the scalability of extending these methods to different scenes becomes a challenge. To this end, self-training a network for monocular depth estimation from 360$^\circ$videos is one way to alleviate this issue. However, there are no existing frameworks that incorporate bi-projection fusion into the self-training scheme, which highly limits the self-supervised performance since bi-projection fusion can leverage information from different projection types. In this paper, we propose BiFuse++ to explore the combination of bi-projection fusion and the self-training scenario. To be specific, we propose a new fusion module and Contrast-Aware Photometric Loss to improve the performance of BiFuse and increase the stability of self-training on real-world videos. We conduct both supervised and self-supervised experiments on benchmark datasets and achieve state-of-the-art performance. Fu-En Wang, Yu-Hsuan Yeh, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning Object-level Point Augmentor for Semi-supervised 3D Object Detection
Cheng-Ju Ho, Chen-Hsuan Tai, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
BMVC | 3 |
| 2022 | Learning to Learn across Diverse Data Biases in Deep Face RecognitionabstractConvolutional Neural Networks have achieved remarkable success in face recognition, in part due to the abundant availability of data. However, the data used for training CNNs is often imbalanced. Prior works largely focus on the long-tailed nature of face datasets in data volume per identity, or focus on single bias variation. In this paper, we show that many bias variations such as ethnicity, head pose, occlusion and blur can jointly affect the accuracy significantly. We propose a sample level weighting approach termed Multi-variation Cosine Margin (MvCoM), to simultaneously consider the multiple variation factors, which orthogonally enhances the face recognition losses to incorporate the importance of training samples. Further, we leverage a learning to learn approach, guided by a held-out meta learning set and use an additive modeling to predict the MvCoM. Extensive experiments on challenging face recognition benchmarks demonstrate the advantages of our method in jointly handling imbalances due to multiple variations. Chang Liu 0022, Xiang Yu 0002, Yi-Hsuan Tsai, Masoud Faraki, Ramin Moslemi, Manmohan Krishna Chandraker, Yun Fu 0001 |
CVPR | 3 |
| 2022 | MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic SegmentationabstractTest-time adaptation approaches have recently emerged as a practical solution for handling domain shift without access to the source domain data. In this paper, we propose and explore a new multi-modal extension of test-time adaptation for 3D semantic segmentation. We find that, directly applying existing methods usually results in performance instability at test time, because multi-modal input is not considered jointly. To design a framework that can take full advantage of multi-modality, where each modality provides regularized self-supervisory signals to other modalities, we propose two complementary modules within and across the modalities. First, Intra-modal Pseudo-label Generation (Intra-PG) is introduced to obtain reliable pseudo labels within each modality by aggregating information from two models that are both pre-trained on source data but updated with target data at different paces. Second, Inter-modal Pseudo-label Refinement (Inter-PR) adaptively selects more reliable pseudo labels from different modalities based on a proposed consistency scheme. Experiments demonstrate that our regularized pseudo labels produce stable self-learning signals in numerous multi-modal test-time adaptation scenarios for 3D semantic segmentation. Visit our project website at https://www.nec-labs.com/~mas/MM-TTA Inkyu Shin, Yi-Hsuan Tsai, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Sparsh Garg, In-So Kweon, Kuk-Jin Yoon |
CVPR | 2 |
| 2022 | On Generalizing Beyond Domains in Cross-Domain Continual LearningabstractHumans have the ability to accumulate knowledge of new tasks in varying conditions, but deep neural networks of-ten suffer from catastrophic forgetting of previously learned knowledge after learning a new task. Many recent methods focus on preventing catastrophic forgetting under the assumption of train and test data following similar distributions. In this work, we consider a more realistic scenario of continual learning under domain shifts where the model must generalize its inference to an unseen domain. To this end, we encourage learning semantically meaningful features by equipping the classifier with class similarity metrics as learning parameters which are obtained through Mahalanobis similarity computations. Learning of the backbone representation along with these extra parameters is done seamlessly in an end-to-end manner. In addition, we propose an approach based on the exponential moving average of the parameters for better knowledge distillation. We demonstrate that, to a great extent, existing continual learning algorithms fail to handle the forgetting issue under multiple distributions, while our proposed approach learns new tasks under domain shift with accuracy boosts up to 10% on challenging datasets such as DomainNet and OfficeHome. Christian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu 0002, Samuel Schulter, Yumin Suh, Mehrtash Harandi, Manmohan Krishna Chandraker |
CVPR | 3 |
| 2022 | Learning Semantic Segmentation from Multiple Datasets with Label Shifts
Dongwan Kim, Yi-Hsuan Tsai, Yumin Suh, Masoud Faraki, Sparsh Garg, Manmohan Krishna Chandraker, Bohyung Han |
ECCV (28) | 2 |
| 2022 | Learning Phase Mask for Privacy-Preserving Passive Depth Estimation
Zaid Tasneem, Giovanni Milione, Yi-Hsuan Tsai, Xiang Yu 0002, Ashok Veeraraghavan, Manmohan Krishna Chandraker, Francesco Pittaluga |
ECCV (7) | 3 |
| 2022 | 3D-PL: Domain Adaptive Depth Estimation with 3D-Aware Pseudo-Labeling
Yu-Ting Yen, Chia-Ni Lu, Walon Wei-Chen Chiu, Yi-Hsuan Tsai |
ECCV (27) | 4 |
| 2022 | Self-Supervised Feature Learning from Partial Point Clouds via Pose DisentanglementabstractSelf-supervised learning on point clouds has gained a lot of attention recently, since it addresses the label-efficiency and domain-gap problems on point cloud tasks. In this paper, we propose a novel self-supervised framework to learn informative features from partial point clouds. We leverage partial point clouds scanned by LiDAR that contain both content and pose attributes, and we show that disentangling such two factors from partial point clouds enhances feature learning. To this end, our framework consists of three main parts: 1) a completion network to capture holistic semantics of point clouds; 2) a pose regression network to understand the viewing angle where partial data is scanned from; 3) a partial reconstruction network to encourage the model to learn content and pose features. To demonstrate the robustness of the learnt feature representations, we conduct several downstream tasks including classification, part segmentation, and registration, with comparisons against state-of-the-art methods. Our method not only outperforms existing self-supervised methods, but also shows a better generalizability across synthetic and real-world datasets. Meng-Shiun Tsai, Pei-Ze Chiang, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
IROS | 3 |
| 2022 | 360-MLC: Multi-view Layout Consistency for Self-training and Hyper-parameter TuningabstractWe present 360-MLC, a self-training method based on multi-view layout consistency for finetuning monocular room-layout models using unlabeled 360-images only. This can be valuable in practical scenarios where a pre-trained model needs to be adapted to a new data domain without using any ground truth annotations. Our simple yet effective assumption is that multiple layout estimations in the same scene must define a consistent geometry regardless of their camera positions. Based on this idea, we leverage a pre-trained model to project estimated layout boundaries from several camera views into the 3D world coordinate. Then, we re-project them back to the spherical coordinate and build a probability function, from which we sample the pseudo-labels for self-training. To handle unconfident pseudo-labels, we evaluate the variance in the re-projected boundaries as an uncertainty value to weight each pseudo-label in our loss function during training. In addition, since ground truth annotations are not available during training nor in testing, we leverage the entropy information in multiple layout estimations as a quantitative metric to measure the geometry consistency of the scene, allowing us to evaluate any layout estimator for hyper-parameter tuning, including model selection without ground truth annotations. Experimental results show that our solution achieves favorable performance against state-of-the-art methods when self-training from three publicly available source datasets to a unique, newly labeled dataset consisting of multi-view images of the same scenes. Bolivar Solarte, Chin-Hsuan Wu, Yueh-Cheng Liu, Yi-Hsuan Tsai, Min Sun 0001 |
NeurIPS | 4 |
| 2022 | Intelligent DNA methylation biomarker selection for colorectal cancerabstractDNA methylation plays an important role in the regulation of gene expression, and aberrant changes in epigenetic regulation can be detected in cancer development or progress. In this study, genome-wide DNA methylation profiles and electronic medical records were combined to identify effective DNA methylation biomarkers for colorectal cancer (CRC). Statistical analytics and deep learning approaches were integrated to explore associated novel biomarkers. These identified biomarkers could facilitate accurate diagnosis of CRC and monitor disease progression for a testing subject, and provide suggestions for appropriate medical treatments. Firstly, DNA methylation profiles were analyzed through a standard pipeline to discover significantly differential methylation biomarkers as primary biomarkers. Incorporating Taiwan’s National Health Insurance Research Database (NHIRD), associated comorbidities of CRC were identified and associated disease genes were considered as secondary biomarkers. The intersection of primary and secondary biomarkers was performed to obtain CRC-relevant biomarker candidates, and gene ontology annotations were used to calculate candidate biomarker-to-biomarker distances. Based on the formulated gene-pair distance matrix for all candidate biomarkers, we applied clustering algorithms to categorize candidate biomarkers into different functional groups. Corresponding coefficients for each biomarker within a functional cluster were ranked through an attention-based recurrent neural network. The weighting coefficient for each biomarker was further taken as a reference for designing a practical testing toolkit. Finally, we obtained 11 important biomarkers from 5 functional clustered groups which were established by the 141 biomarker candidates screened from 450k DNA methylation probes. Yi-Hsuan Tsai, Lucy Yi-Hsuan Lai, Yun-Te Liao, Shu-Jen Chen, Tun-Wen Pai, Margaret Dah-Tsyr Chang |
SMC | 1 |
| 2022 | Semi-supervised Multi-task Learning for Semantics and DepthabstractMulti-Task Learning (MTL) aims to enhance the model generalization by sharing representations between related tasks for better performance. Typical MTL methods are jointly trained with the complete multitude of ground-truths for all tasks simultaneously. However, one single dataset may not contain the annotations for each task of interest. To address this issue, we propose the Semi-supervised Multi-Task Learning (SemiMTL) method to leverage the available supervisory signals from different datasets, particularly for semantic segmentation and depth estimation tasks. To this end, we design an adversarial learning scheme in our semi-supervised training by leveraging unlabeled data to optimize all the task branches simultaneously and accomplish all tasks across datasets with partial annotations. We further present a domain-aware discriminator structure with various alignment formulations to mitigate the domain discrepancy issue among datasets. Finally, we demonstrate the effectiveness of the proposed method to learn across different datasets on challenging street view and remote sensing benchmarks. Yufeng Wang 0004, Yi-Hsuan Tsai, Wei-Chih Hung, Wenrui Ding, Ming-Hsuan Yang 0001 |
WACV | 2 |
| 2022 | Understanding Synonymous Referring Expressions via Contrastive FeaturesabstractAbstract Referring expression comprehension aims to localize objects identified by natural language descriptions. This is a challenging task as it requires understanding of both visual and language domains. One nature is that each object can be described by synonymous sentences with paraphrases, and such varieties in languages have critical impact on learning a comprehension model. While prior work usually treats each sentence and attends it to an object separately, we focus on learning a referring expression comprehension model that considers the property in synonymous sentences. To this end, we develop an end-to-end trainable framework to learn contrastive features on the image and object instance levels, where features extracted from synonymous sentences to describe the same object should be closer to each other after mapping to the visual domain. We conduct extensive experiments to evaluate the proposed algorithm on several benchmark datasets, and demonstrate that our method performs favorably against the state-of-the-art approaches. Furthermore, since the varieties in expressions become larger across datasets when they describe objects in different ways, we present the cross-dataset and transfer learning settings to validate the ability of our learned transferable features. Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Learning to Hide Residual for Boosting Image Compression
Yi-Lun Lee, Yen-Chung Chen, Min-Yuan Tseng, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
BMVC | 4 |
| 2021 | Cross-Domain Similarity Learning for Face Recognition in Unseen DomainsabstractFace recognition models trained under the assumption of identical training and test distributions often suffer from poor generalization when faced with unknown variations, such as a novel ethnicity or unpredictable individual make-ups during test time. In this paper, we introduce a novel cross-domain metric learning loss, which we dub Cross-Domain Triplet (CDT) loss, to improve face recognition in unseen domains. The CDT loss encourages learning semantically meaningful features by enforcing compact feature clusters of identities from one domain, where the compactness is measured by underlying similarity metrics that belong to another training domain with different statistics. Intuitively, it discriminatively correlates explicit metrics derived from one domain, with triplet samples from another domain in a unified loss function to be minimized within a network, which leads to better alignment of the training domains. The network parameters are further enforced to learn generalized features under domain shift, in a model-agnostic learning pipeline. Unlike the recent work of Meta Face Recognition [18], our method does not require careful hard-pair sample mining and filtering strategy during training. Extensive experiments on various face recognition benchmarks show the superiority of our method in handling variations, compared to baseline and the state-of-the-art methods. Masoud Faraki, Xiang Yu 0002, Yi-Hsuan Tsai, Yumin Suh, Manmohan Krishna Chandraker |
CVPR | 3 |
| 2021 | LED2-Net: Monocular 360deg Layout Estimation via Differentiable Depth RenderingabstractAlthough significant progress has been made in room layout estimation, most methods aim to reduce the loss in the 2D pixel coordinate rather than exploiting the room structure in the 3D space. Towards reconstructing the room layout in 3D, we formulate the task of 360° layout estimation as a problem of predicting depth on the horizon line of a panorama. Specifically, we propose the Differentiable Depth Rendering procedure to make the conversion from layout to depth prediction differentiable, thus making our proposed model end-to-end trainable while leveraging the 3D geometric information, without the need of providing the ground truth depth. Our method achieves state-of-the-art performance on numerous 360° layout benchmark datasets. Moreover, our formulation enables a pre-training step on the depth dataset, which further improves the generalizability of our layout estimation model. Fu-En Wang, Yu-Hsuan Yeh, Min Sun 0001, Walon Wei-Chen Chiu, Yi-Hsuan Tsai |
CVPR | 5 |
| 2021 | Learning Cross-Modal Contrastive Features for Video Domain AdaptationabstractLearning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data is usually associated with multi-modal information, e.g., RGB and optical flow, and thus it remains a challenge to design a better method that considers the cross-modal inputs under the cross-domain adaptation setting. To this end, we propose a unified framework for video domain adaptation, which simultaneously regularizes cross-modal and cross-domain feature representations. Specifically, we treat each modality in a domain as a view and leverage the contrastive learning technique with properly designed sampling strategies. As a result, our objectives regularize feature spaces, which originally lack the connection across modalities or have less alignment across domains. We conduct experiments on domain adaptive action recognition benchmark datasets, i.e., UCF, HMDB, and EPIC-Kitchens, and demonstrate the effectiveness of our components against state-of-the-art algorithms. Donghyun Kim 0006, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu 0002, Stan Sclaroff, Kate Saenko, Manmohan Krishna Chandraker |
ICCV | 2 |
| 2021 | Towards Interpretable Deep Networks for Monocular Depth EstimationabstractDeep networks for Monocular Depth Estimation (MDE) have achieved promising performance recently and it is of great importance to further understand the interpretability of these networks. Existing methods attempt to provide post-hoc explanations by investigating visual cues, which may not explore the internal representations learned by deep networks. In this paper, we find that some hidden units of the network are selective to certain ranges of depth, and thus such behavior can be served as a way to interpret the internal representations. Based on our observations, we quantify the interpretability of a deep MDE network by the depth selectivity of its hidden units. Moreover, we then propose a method to train interpretable MDE deep networks without changing their original architectures, by assigning a depth range for each unit to select. Experimental results demonstrate that our method is able to enhance the interpretability of deep MDE networks by largely improving the depth selectivity of their units, while not harming or even improving the depth estimation accuracy. We further provide comprehensive analysis to show the reliability of selective units, the applicability of our method on different layers, models, and datasets, and a demonstration on analysis of model error. Source code and models are available at https://github.com/youzunzhi/InterpretableMDE. Zunzhi You, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Guanbin Li |
ICCV | 2 |
| 2021 | Robust 360-8PA: Redesigning The Normalized 8-point Algorithm for 360-FoV ImagesabstractIn this paper, we present a novel preconditioning strategy for the classic 8-point algorithm (8-PA) for estimating an essential matrix from 360-FoV images (i.e., equirectangular images) in spherical projection. To alleviate the effect of uneven key-feature distributions and outlier correspondences, which can potentially decrease the accuracy of an essential matrix, our method optimizes a non-rigid transformation to deform a spherical camera into a new spatial domain, defining a new constraint and a more robust and accurate solution for an essential matrix. Through several experiments using random synthetic points, 360-FoV, and fish-eye images, we demonstrate that our normalization can increase the camera pose accuracy about 20% without significantly overhead the computation time. In addition, we present further benefits of our method through both a constant weighted least-square optimization that improves further the well known Gold Standard Method (GSM) (i.e., the non-linear optimization by using epipolar errors); and a relaxation of the number of RANSAC iterations, both showing that our normalization outcomes a more reliable, robust, and accurate solution. Bolivar Solarte, Chin-Hsuan Wu, Kuan-Wei Lu 0001, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001 |
ICRA | 4 |
| 2021 | End-to-end Multi-modal Video Temporal GroundingabstractWe address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as visual features, we propose a multi-modal framework to extract complementary information from videos. Specifically, we adopt RGB images for appearance, optical flow for motion, and depth maps for image structure. While RGB images provide abundant visual cues of certain events, the performance may be affected by background clutters. Therefore, we use optical flow to focus on large motion and depth maps to infer the scene configuration when the action is related to objects recognizable with their shapes. To integrate the three modalities more effectively and enable inter-modal learning, we design a dynamic fusion scheme with transformers to model the interactions between modalities. Furthermore, we apply intra-modal self-supervised learning to enhance feature representations across videos for each modality, which also facilitates multi-modal learning. We conduct extensive experiments on the Charades-STA and ActivityNet Captions datasets, and show that the proposed method performs favorably against state-of-the-art approaches. Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
NeurIPS | 2 |
| 2021 | Dual-Stream Fusion Network for Spatiotemporal Video Super-ResolutionabstractVisual data upsampling has been an important research topic for improving the perceptual quality and benefiting various computer vision applications. In recent years, we have witnessed remarkable progresses brought by the re-naissance of deep learning techniques for video or image super-resolution. However, most existing methods focus on advancing super-resolution at either spatial or temporal direction, i.e, to increase the spatial resolution or the video frame rate. In this paper, we instead turn to discuss both directions jointly and tackle the spatiotemporal upsampling problem. Our method is based on an important observation that: even the direct cascade of prior research in spatial and temporal super-resolution can achieve the spatiotemporal upsampling, changing orders for combining them would lead to results with a complementary property. Thus, we propose a dual-stream fusion network to adaptively fuse the intermediate results produced by two spatiotemporal up-sampling streams, where the first stream applies the spatial super-resolution followed by the temporal super-resolution, while the second one is with the reverse order of cascade. Extensive experiments verify the efficacy of the proposed method against several baselines. Moreover, we investigate various spatial and temporal upsampling methods as the basis in our two-stream model and demonstrate the flexibility with wide applicability of the proposed framework. Min-Yuan Tseng, Yen-Chung Chen, Yi-Lun Lee, Wei-Sheng Lai, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
WACV | 5 |
| 2021 | Learning to Caricature via Semantic Shape TransformabstractAbstract Caricature is an artistic drawing created to abstract or exaggerate facial features of a person. Rendering visually pleasing caricatures is a difficult task that requires professional skills, and thus it is of great interest to design a method to automatically generate such drawings. To deal with large shape changes, we propose an algorithm based on a semantic shape transform to produce diverse and plausible shape exaggerations. Specifically, we predict pixel-wise semantic correspondences and perform image warping on the input photo to achieve dense shape transformation. We show that the proposed framework is able to render visually pleasing shape exaggerations while maintaining their facial structures. In addition, our model allows users to manipulate the shape via the semantic map. We demonstrate the effectiveness of our approach on a large photograph-caricature benchmark dataset with comparisons to the state-of-the-art methods. Wenqing Chu, Wei-Chih Hung, Yi-Hsuan Tsai, Yu-Ting Chang, Yijun Li 0001, Deng Cai 0001, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Adversarial Learning of Privacy-Preserving and Task-Oriented RepresentationsabstractData privacy has emerged as an important issue as data-driven deep learning has been an essential component of modern machine learning systems. For instance, there could be a potential privacy risk of machine learning systems via the model inversion attack, whose goal is to reconstruct the input data from the latent representation of deep networks. Our work aims at learning a privacy-preserving and task-oriented representation to defend against such model inversion attacks. Specifically, we propose an adversarial reconstruction learning framework that prevents the latent representations decoded into original input data. By simulating the expected behavior of adversary, our framework is realized by minimizing the negative pixel reconstruction loss or the negative feature reconstruction (i.e., perceptual distance) loss. We validate the proposed method on face attribute prediction, showing that our method allows protecting visual privacy with a small decrease in utility performance. In addition, we show the utility-privacy trade-off with different choices of hyperparameter for negative perceptual distance loss at training, allowing service providers to determine the right level of privacy-protection with a certain utility performance. Moreover, we provide an extensive study with different selections of features, tasks, and the data to further analyze their influence on privacy protection. Taihong Xiao, Yi-Hsuan Tsai, Kihyuk Sohn, Manmohan Krishna Chandraker, Ming-Hsuan Yang 0001 |
AAAI | 2 |
| 2020 | Regularizing Meta-learning via Gradient Dropout
Hung-Yu Tseng, Yi-Wen Chen, Yi-Hsuan Tsai, Sifei Liu, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
ACCV (4) | 3 |
| 2020 | Mixup-CAM: Weakly-supervised Semantic Segmentation via Uncertainty Regularization
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
BMVC | 5 |
| 2020 | Boosting Image and Video Compression via Learning Latent Residual Patterns
Yen-Chung Chen, Keng-Jui Chang, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
BMVC | 3 |
| 2020 | Adaptation Across Extreme Variations using Unlabeled Bridges
Shuyang Dai, Kihyuk Sohn, Yi-Hsuan Tsai, Lawrence Carin, Manmohan Krishna Chandraker |
BMVC | 3 |
| 2020 | Weakly-Supervised Semantic Segmentation via Sub-Category ExplorationabstractExisting weakly-supervised semantic segmentation methods using image-level annotations typically rely on initial responses to locate object regions. However, such response maps generated by the classification network usually focus on discriminative object parts, due to the fact that the network does not need the entire object for optimizing the objective function. To enforce the network to pay attention to other parts of an object, we propose a simple yet effective approach that introduces a self-supervised task by exploiting the sub-category information. Specifically, we perform clustering on image features to generate pseudo sub-categories labels within each annotated parent class, and construct a sub-category objective to assign the network to a more challenging task. By iteratively clustering image features, the training process does not limit itself to the most discriminative object parts, hence improving the quality of the response maps. We conduct extensive analysis to validate the proposed method and show that our approach performs favorably against the state-of-the-art approaches. Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
CVPR | 5 |
| 2020 | BiFuse: Monocular 360 Depth Estimation via Bi-Projection FusionabstractDepth estimation from a monocular 360 image is an emerging problem that gains popularity due to the availability of consumer-level 360 cameras and the complete surrounding sensing capability. While the standard of 360 imaging is under rapid development, we propose to predict the depth map of a monocular 360 image by mimicking both peripheral and foveal vision of the human eye. To this end, we adopt a two-branch neural network leveraging two common projections: equirectangular and cubemap projections. In particular, equirectangular projection incorporates a complete field-of-view but introduces distortion, whereas cubemap projection avoids distortion but introduces discontinuity at the boundary of the cube. Thus we propose a bi-projection fusion scheme along with learnable masks to balance the feature map from the two projections. Moreover, for the cubemap projection, we propose a spherical padding procedure which mitigates discontinuity at the boundary of each face. We apply our method to four panorama datasets and show favorable results against the existing state-of-the-art methods. Fu-En Wang, Yu-Hsuan Yeh, Min Sun 0001, Walon Wei-Chen Chiu, Yi-Hsuan Tsai |
CVPR | 5 |
| 2020 | Every Pixel Matters: Center-Aware Feature Alignment for Domain Adaptive Object Detector
Cheng-Chun Hsu, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
ECCV (9) | 2 |
| 2020 | Colorization of Depth Map via Disentanglement
Chung-Sheng Lai, Zunzhi You, Chingchun Huang, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
ECCV (7) | 4 |
| 2020 | Domain Adaptive Semantic Segmentation Using Weak Labels
Sujoy Paul, Yi-Hsuan Tsai, Samuel Schulter, Amit K. Roy-Chowdhury, Manmohan Krishna Chandraker |
ECCV (9) | 2 |
| 2020 | Object Detection with a Unified Label Space from Multiple Datasets
Xiangyun Zhao, Samuel Schulter, Yi-Hsuan Tsai, Manmohan Krishna Chandraker, Ying Wu 0001 |
ECCV (14) | 4 |
| 2020 | 360SD-Net: 360° Stereo Depth Estimation with Learnable Cost VolumeabstractRecently, end-to-end trainable deep neural networks have significantly improved stereo depth estimation for perspective images. However, 360° images captured under equirectangular projection cannot benefit from directly adopting existing methods due to distortion introduced (i.e., lines in 3D are not projected onto lines in 2D). To tackle this issue, we present a novel architecture specifically designed for spherical disparity using the setting of top-bottom 360° camera pairs. Moreover, we propose to mitigate the distortion issue by (1) an additional input branch capturing the position and relation of each pixel in the spherical coordinate, and (2) a cost volume built upon a learnable shifting filter. Due to the lack of 360° stereo data, we collect two 360° stereo datasets from Matterport3D and Stanford3D for training and evaluation. Extensive experiments and ablation study are provided to validate our method against existing algorithms. Finally, we show promising results on real-world environments capturing images with two consumer-level cameras. Our project page is at https://albert100121.github.io/360SD-Net-Project-Page. Ning-Hsu Wang, Bolivar Solarte, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001 |
ICRA | 3 |
| 2020 | Image Hashing via Linear Discriminant LearningabstractHashing has attracted attention in recent years due to the rapid growth of image and video data on the web. Benefiting from recent advances in deep learning, deep supervised hashing has achieved promising results for image retrieval. However, existing methods are either less efficient in data usage or incapable of learning linearly discriminative binary codes. In this paper, we revisit linear discriminative analysis and propose a linear discriminative hashing (LDH) objective that is efficient in training and achieves better accuracy in retrieval. With the joint supervision of a classification loss, we design a robust deep network to obtain binary codes that are inter-class separable and intra-class compact, which provides better representations for image retrieval. We conduct extensive experiments on three benchmark datasets, and our LDH algorithm performs favorably against existing state-of-the-art deep supervised hashing methods. Weixiang Hong 0001, Yu-Ting Chang, Haifang Qin, Wei-Chih Hung, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
WACV | 5 |
| 2020 | Progressive Domain Adaptation for Object DetectionabstractRecent deep learning methods for object detection rely on a large amount of bounding box annotations. Collecting these annotations is laborious and costly, yet supervised models do not generalize well when testing on images from a different distribution. Domain adaptation provides a solution by adapting existing labels to the target testing data. However, a large gap between domains could make adaptation a challenging task, which leads to unstable training processes and sub-optimal results. In this paper, we propose to bridge the domain gap with an intermediate domain and progressively solve easier adaptation subtasks. This intermediate domain is constructed by translating the source images to mimic the ones in the target domain. To tackle the domain-shift problem, we adopt adversarial learning to align distributions at the feature level. In addition, a weighted task loss is applied to deal with unbalanced image quality in the intermediate domain. Experimental results show that our method performs favorably against the state-of-the-art method in terms of the performance on the target domain. Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai, Wei-Chih Hung, Hung-Yu Tseng, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001 |
WACV | 3 |
| 2020 | Active Adversarial Domain AdaptationabstractWe propose an active learning approach for transferring representations across domains. Our approach, active adversarial domain adaptation (AADA), explores a duality between two related problems: adversarial domain alignment and importance sampling for adapting models across domains. The former uses a domain discriminative model to align domains, while the latter utilizes the model to weigh samples to account for distribution shifts. Specifically, our importance weight promotes unlabeled samples with large uncertainty in classification and diversity compared to la-beled examples, thus serving as a sample selection scheme for active learning. We show that these two views can be unified in one framework for domain adaptation and transfer learning when the source domain has many labeled examples while the target domain does not. AADA provides significant improvements over fine-tuning based approaches and other sampling methods when the two domains are closely related. Results on challenging domain adaptation tasks such as object detection demonstrate that the advantage over baseline approaches is retained even after hundreds of examples being actively annotated. Jong-Chyi Su, Yi-Hsuan Tsai, Kihyuk Sohn, Buyu Liu, Subhransu Maji, Manmohan Krishna Chandraker |
WACV | 2 |
| 2020 | VOSTR: Video Object Segmentation via Transferable Representations
Yi-Wen Chen, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Transfer learning in computer vision tasks: Remember where you come from
Xuhong Li 0002, Yves Grandvalet, Franck Davoine, Jingchun Cheng, Yin Cui, Han Zhang 0005, Serge J. Belongie, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
Image Vis. Comput. | 8 |
| 2019 | Guide Your Eyes: Learning Image Manipulation under Saliency Guidance
Yen-Chung Chen, Keng-Jui Chang, Yi-Hsuan Tsai, Yu-Chiang Frank Wang, Walon Wei-Chen Chiu |
BMVC | 3 |
| 2019 | Referring Expression Object Segmentation with Caption-Aware Consistency
Yi-Wen Chen, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
BMVC | 2 |
| 2019 | A Top-Down Unified Framework for Instance-level Human Parsing
Haifang Qin, Weixiang Hong 0001, Wei-Chih Hung, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
BMVC | 4 |
| 2019 | Bridging Stereo Matching and Optical Flow via Spatiotemporal CorrespondenceabstractStereo matching and flow estimation are two essential tasks for scene understanding, spatially in 3D and temporally in motion. Existing approaches have been focused on the unsupervised setting due to the limited resource to obtain the large-scale ground truth data. To construct a self-learnable objective, co-related tasks are often linked together to form a joint framework. However, the prior work usually utilizes independent networks for each task, thus not allowing to learn shared feature representations across models. In this paper, we propose a single and principled network to jointly learn spatiotemporal correspondence for stereo matching and flow estimation, with a newly designed geometric connection as the unsupervised signal for temporally adjacent stereo pairs. We show that our method performs favorably against several state-of-the-art baselines for both unsupervised depth and flow estimation on the KITTI benchmark dataset. Hsueh-Ying Lai, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
CVPR | 2 |
| 2019 | Domain Adaptation for Structured Output via Discriminative Patch RepresentationsabstractPredicting structured outputs such as semantic segmentation relies on expensive per-pixel annotations to learn supervised models like convolutional neural networks. However, models trained on one data domain may not generalize well to other domains without annotations for model finetuning. To avoid the labor-intensive process of annotation, we develop a domain adaptation method to adapt the source data to the unlabeled target domain. We propose to learn discriminative feature representations of patches in the source domain by discovering multiple modes of patch-wise output distribution through the construction of a clustered space. With such representations as guidance, we use an adversarial learning scheme to push the feature representations of target patches in the clustered space closer to the distributions of source patches. In addition, we show that our framework is complementary to existing domain adaptation techniques and achieves consistent improvements on semantic segmentation. Extensive ablations and results are demonstrated on numerous benchmark datasets with various settings, such as synthetic-to-real and cross-city scenarios. Yi-Hsuan Tsai, Kihyuk Sohn, Samuel Schulter, Manmohan Krishna Chandraker |
ICCV | 1 |
| 2019 | Weakly-Supervised Caricature Face Parsing Through Domain AdaptationabstractA caricature is an artistic form of a person's picture in which certain striking characteristics are abstracted or exaggerated in order to create a humor or sarcasm effect. For numerous caricature related applications such as attribute recognition and caricature editing, face parsing is an essential pre-processing step that provides a complete facial structure understanding. However, current state-of-the-art face parsing methods require large amounts of labeled data on the pixel-level and such process for caricature is tedious and labor-intensive. For real photos, there are numerous labeled datasets for face parsing. Thus, we formulate caricature face parsing as a domain adaptation problem, where real photos play the role of the source domain, adapting to the target caricatures. Specifically, we first leverage a spatial transformer based network to enable shape domain shifts. A feed-forward style transfer network is then utilized to capture texture-level domain gaps. With these two steps, we synthesize face caricatures from real photos, and thus we can use parsing ground truths of the original photos to learn the parsing model. Experimental results on the synthetic and real caricatures demonstrate the effectiveness of the proposed domain adaptation algorithm. Code is available at: https://github.com/ZJULearning/CariFaceParsing. Wenqing Chu, Wei-Chih Hung, Yi-Hsuan Tsai, Deng Cai 0001, Ming-Hsuan Yang 0001 |
ICIP | 3 |
| 2019 | Plug-and-Play: Improve Depth Prediction via Sparse Data PropagationabstractWe propose a novel plug-and-play (PnP) module for improving depth prediction with taking arbitrary patterns of sparse depths as input. Given any pre-trained depth prediction model, our PnP module updates the intermediate feature map such that the model outputs new depths consistent with the given sparse depths. Our method requires no additional training and can be applied to practical applications such as leveraging both RGB and sparse LiDAR points to robustly estimate dense depth map. Our approach achieves consistent improvements on various state-of-the-art methods on indoor (i.e., NYU-v2) and outdoor (i.e., KITTI) datasets. Various types of LiDARs are also synthesized in our experiments to verify the general applicability of our PnP module in practice. Tsun-Hsuan Wang, Fu-En Wang, Juan-Ting Lin, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001 |
ICRA | 4 |
| 2019 | 3D LiDAR and Stereo Fusion using Stereo Matching Network with Conditional Cost Volume NormalizationabstractThe complementary characteristics of active and passive depth sensing techniques motivate the fusion of the LiDAR sensor and stereo camera for improved depth perception. Instead of directly fusing estimated depths across LiDAR and stereo modalities, we take advantages of the stereo matching network with two enhanced techniques: Input Fusion and Conditional Cost Volume Normalization (CCVNorm) on the LiDAR information. The proposed framework is generic and closely integrated with the cost volume component that is commonly utilized in stereo matching neural networks. We experimentally verify the efficacy and robustness of our method on the KITTI Stereo and Depth Completion datasets, obtaining favorable performance against various fusion strategies. Moreover, we demonstrate that, with a hierarchical extension of CCVNorm, the proposed method brings only slight overhead to the stereo matching network in terms of computation time and model size. Tsun-Hsuan Wang, Hou-Ning Hu, Chieh Hubert Lin, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001 |
IROS | 4 |
| 2018 | Learning Binary Residual Representations for Domain-Specific Video StreamingabstractWe study domain-specific video streaming. Specifically, we target a streaming setting where the videos to be streamed from a server to a client are all in the same domain and they have to be compressed to a small size for low-latency transmission. Several popular video streaming services, such as the video game streaming services of GeForce Now and Twitch, fall in this category. While conventional video compression standards such as H.264 are commonly used for this task, we hypothesize that one can leverage the property that the videos are all in the same domain to achieve better video quality. Based on this hypothesis, we propose a novel video compression pipeline. Specifically, we first apply H.264 to compress domain-specific videos. We then train a novel binary autoencoder to encode the leftover domain-specific residual information frame-by-frame into binary representations. These binary representations are then compressed and sent to the client together with the H.264 stream. In our experiments, we show that our pipeline yields consistent gains over standard H.264 compression across several benchmark datasets while using the same channel bandwidth. Yi-Hsuan Tsai, Ming-Yu Liu 0001, Deqing Sun, Ming-Hsuan Yang 0001, Jan Kautz |
AAAI | 1 |
| 2018 | Unseen Object Segmentation in Videos via Transferable Representations
Yi-Wen Chen, Yi-Hsuan Tsai, Chu-Ya Yang, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
ACCV (4) | 2 |
| 2018 | Adversarial Learning for Semi-supervised Semantic Segmentation
Wei-Chih Hung, Yi-Hsuan Tsai, Yan-Ting Liou, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
BMVC | 2 |
| 2018 | Fast and Accurate Online Video Object Segmentation via Tracking PartsabstractOnline video object segmentation is a challenging task as it entails to process the image sequence timely and accurately. To segment a target object through the video, numerous CNN-based methods have been developed by heavily finetuning on the object mask in the first frame, which is time-consuming for online applications. In this paper, we propose a fast and accurate video object segmentation algorithm that can immediately start the segmentation process once receiving the images. We first utilize a part-based tracking method to deal with challenging factors such as large deformation, occlusion, and cluttered background. Based on the tracked bounding boxes of parts, we construct a region-of-interest segmentation network to generate part masks. Finally, a similarity-based scoring function is adopted to refine these object parts by comparing them to the visual information in the first frame. Our method performs favorably against state-of-the-art algorithms in accuracy on the DAVIS benchmark dataset, while achieving much faster runtime performance. Jingchun Cheng, Yi-Hsuan Tsai, Wei-Chih Hung, Shengjin Wang, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2018 | Learning to Adapt Structured Output Space for Semantic SegmentationabstractConvolutional neural network-based approaches for semantic segmentation rely on supervision with pixel-level ground truth, but may not generalize well to unseen image domains. As the labeling process is tedious and labor intensive, developing algorithms that can adapt source ground truth labels to the target domain is of great interest. In this paper, we propose an adversarial learning method for domain adaptation in the context of semantic segmentation. Considering semantic segmentations as structured outputs that contain spatial similarities between the source and target domains, we adopt adversarial learning in the output space. To further enhance the adapted model, we construct a multi-level adversarial network to effectively perform output space domain adaptation at different feature levels. Extensive experiments and ablation study are conducted under various domain adaptation settings, including synthetic-to-real and cross-city scenarios. We show that the proposed method performs favorably against the state-of-the-art methods in terms of accuracy and visual quality. Yi-Hsuan Tsai, Wei-Chih Hung, Samuel Schulter, Kihyuk Sohn, Ming-Hsuan Yang 0001, Manmohan Krishna Chandraker |
CVPR | 1 |
| 2018 | Learning Video-Story Composition via Recurrent Neural NetworkabstractIn this paper, we propose a learning-based method to compose a video-story from a group of video clips that describe an activity or experience. We learn the coherence between video clips from real videos via the Recurrent Neural Network (RNN) that jointly incorporates the spatial-temporal semantics and motion dynamics to generate smooth and relevant compositions. We further rearrange the results generated by the RNN to make the overall video-story compatible with the storyline structure via a submodular ranking optimization process. Experimental results on the video-story dataset show that the proposed algorithm outperforms the state-of-the-art approach. Guangyu Zhong, Yi-Hsuan Tsai, Sifei Liu, Zhixun Su, Ming-Hsuan Yang 0001 |
WACV | 2 |
| 2017 | Deep Image HarmonizationabstractCompositing is one of the most common operations in photo editing. To generate realistic composites, the appearances of foreground and background need to be adjusted to make them compatible. Previous approaches to harmonize composites have focused on learning statistical relationships between hand-crafted appearance features of the foreground and background, which is unreliable especially when the contents in the two layers are vastly different. In this work, we propose an end-to-end deep convolutional neural network for image harmonization, which can capture both the context and semantic information of the composite images during harmonization. We also introduce an efficient way to collect large-scale and high-quality training data that can facilitate the training process. Experiments on the synthesized dataset and real composite images show that the proposed network outperforms previous state-of-the-art methods. Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Xin Lu 0006, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2017 | SegFlow: Joint Learning for Video Object Segmentation and Optical FlowabstractThis paper proposes an end-to-end trainable network, SegFlow, for simultaneously predicting pixel-wise object segmentation and optical flow in videos. The proposed SegFlow has two branches where useful information of object segmentation and optical flow is propagated bidirectionally in a unified framework. The segmentation branch is based on a fully convolutional network, which has been proved effective in image segmentation task, and the optical flow branch takes advantage of the FlowNet model. The unified framework is trained iteratively offline to learn a generic notion, and fine-tuned online for specific objects. Extensive experiments on both the video object segmentation and optical flow datasets demonstrate that introducing optical flow improves the performance of segmentation and vice versa, against the state-of-the-art algorithms. Jingchun Cheng, Yi-Hsuan Tsai, Shengjin Wang, Ming-Hsuan Yang 0001 |
ICCV | 2 |
| 2017 | Scene Parsing with Global Context EmbeddingabstractWe present a scene parsing method that utilizes global context information based on both the parametric and nonparametric models. Compared to previous methods that only exploit the local relationship between objects, we train a context network based on scene similarities to generate feature representations for global contexts. In addition, these learned features are utilized to generate global and spatial priors for explicit classes inference. We then design modules to embed the feature representations and the priors into the segmentation network as additional global context cues. We show that the proposed method can eliminate false positives that are not compatible with the global context representations. Experiments on both the MIT ADE20K and PASCAL Context datasets show that the proposed method performs favorably against existing methods. Wei-Chih Hung, Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Xin Lu 0006, Ming-Hsuan Yang 0001 |
ICCV | 2 |
| 2016 | Weakly-Supervised Video Scene Co-parsing
Guangyu Zhong, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
ACCV (1) | 2 |
| 2016 | Video Segmentation via Object FlowabstractVideo object segmentation is challenging due to fast moving objects, deforming shapes, and cluttered backgrounds. Optical flow can be used to propagate an object segmentation over time but, unfortunately, flow is often inaccurate, particularly around object boundaries. Such boundaries are precisely where we want our segmentation to be accurate. To obtain accurate segmentation across time, we propose an efficient algorithm that considers video segmentation and optical flow estimation simultaneously. For video segmentation, we formulate a principled, multiscale, spatio-temporal objective function that uses optical flow to propagate information between frames. For optical flow estimation, particularly at object boundaries, we compute the flow independently in the segmented regions and recompose the results. We call the process object flow and demonstrate the effectiveness of jointly optimizing optical flow and video segmentation using an iterative scheme. Experiments on the SegTrack v2 and Youtube-Objects datasets show that the proposed algorithm performs favorably against the other state-of-the-art methods. Yi-Hsuan Tsai, Ming-Hsuan Yang 0001, Michael J. Black |
CVPR | 1 |
| 2016 | Semantic Co-segmentation in Videos
Yi-Hsuan Tsai, Guangyu Zhong, Ming-Hsuan Yang 0001 |
ECCV (4) | 1 |
| 2016 | Sky is not the limit: semantic-aware sky replacementabstractSkies are common backgrounds in photos but are often less interesting due to the time of photographing. Professional photographers correct this by using sophisticated tools with painstaking efforts that are beyond the command of ordinary users. In this work, we propose an automatic background replacement algorithm that can generate realistic, artifact-free images with a diverse styles of skies. The key idea of our algorithm is to utilize visual semantics to guide the entire process including sky segmentation, search and replacement. First we train a deep convolutional neural network for semantic scene parsing, which is used as visual prior to segment sky regions in a coarse-to-fine manner. Second, in order to find proper skies for replacement, we propose a data-driven sky search scheme based on semantic layout of the input image. Finally, to re-compose the stylized sky with the original foreground naturally, an appearance transfer method is developed to match statistics locally and semantically. We show that the proposed algorithm can automatically generate a set of visually pleasing results. In addition, we demonstrate the effectiveness of the proposed algorithm with extensive user studies. Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Ming-Hsuan Yang 0001 |
ACM Trans. Graph. | 1 |
| 2015 | Adaptive region pooling for object detectionabstractLearning models for object detection is a challenging problem due to the large intra-class variability of objects in appearance, viewpoints, and rigidity. We address this variability by a novel feature pooling method that is adaptive to segmented regions. The proposed detection algorithm automatically discovers a diverse set of exemplars and their distinctive parts which are used to encode the region structure by the proposed feature pooling method. Based on each exemplar and its parts, a regression model is learned with samples selected by a coarse region matching scheme. The proposed algorithm performs favorably on the PASCAL VOC 2007 dataset against existing algorithms. We demonstrate the benefits of our feature pooling method when compared to conventional spatial pyramid pooling features. We also show that object information can be transferred through exemplars for detected objects. Yi-Hsuan Tsai, Onur C. Hamsici, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2014 | Locality preserving hashingabstractThe spectral hashing algorithm relaxes and solves an objective function for generating hash codes such that data similarity is preserved in the Hamming space. However, the assumption of uniform global data distribution limits its applicability. In the paper, we introduce locality preserving projection to determine the data distribution adaptively, and a spectral method is adopted to estimate the eigenfunctions of the underlying graph Laplacian. Furthermore, pairwise label similarity can be further incorporated in the weight matrix to bridge the semantic gap between data and hash codes. Experiments on three benchmark datasets show the proposed algorithm performs favorably against state-of-the-art hashing methods. Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
ICIP | 1 |
| 2013 | Decomposed Learning for Joint Segmentation and Categorization
Yi-Hsuan Tsai, Jimei Yang, Ming-Hsuan Yang 0001 |
BMVC | 1 |
| 2013 | Exemplar CutabstractWe present a hybrid parametric and nonparametric algorithm, exemplar cut, for generating class-specific object segmentation hypotheses. For the parametric part, we train a pylon model on a hierarchical region tree as the energy function for segmentation. For the nonparametric part, we match the input image with each exemplar by using regions to obtain a score which augments the energy function from the pylon model. Our method thus generates a set of highly plausible segmentation hypotheses by solving a series of exemplar augmented graph cuts. Experimental results on the Graz and PASCAL datasets show that the proposed algorithm achieves favorable segmentation performance against the state-of-the-art methods in terms of visual quality and accuracy. Jimei Yang, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
ICCV | 2 |