Walon Wei-Chen Chiu

dblp:148/9413 · also Wei-Chen Chiu · DBLP profile ↗
← Back
81ranked-venue papers
5as first author
53since 2021 · last 2026
0000-0001-7715-8306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 5 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 5 first-author · 32 since 2021Systems, architecture and hardware · 12 · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
Yi-Lun Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
ICPR (9)3
2026 MCPNet++: Interpretable Classification Models via Multi-Level Concept Prototypes
abstract
Post-hoc and inherently interpretable methods have shown great success in uncovering the inner workings of black-box models, whether by examining them after training or by explicitly designing for interpretability. While these approaches effectively narrow the semantic gap between a model's latent space and human understanding, they typically extract only high-level semantics from the model's final feature map. As a result, they provide a limited perspective on the decision-making process. We argue that explanations lacking insight into both lower- and mid-level semantics cannot be considered fully faithful or genuinely useful. To address this issue, we introduce the Multi-Level Concept Prototypes Classifier (MCPNet), which offers a more holistic interpretation by drawing on information from multiple levels within the model. Rather than relying on predefined concept labels, MCPNet autonomously discovers meaningful concepts from feature maps. To increase versatility, we further propose MCPNet++, which can be seamlessly applied to both CNN and transformer backbones, allowing it to learn meaningful concepts from their respective features. Building on these learned concepts, we also introduce a large language model (LLM)-based method to bridge the gap between these concepts and human perception. Experimental results show that MCPNet++ provides more comprehensive explanations without sacrificing model performance, with the discovered concepts aligning closely with human understanding.
Bor-Shiun Wang, Chien-Yi Wang, Walon Wei-Chen Chiu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 RC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration Network
abstract
This paper presents a groundbreaking approach - the first online automatic geometric calibration method for radar and camera systems. Given the significant data sparsity and measurement uncertainty in radar height data, achieving automatic calibration during system operation has long been a challenge. To address the sparsity issue, we propose a Dual-Perspective representation that gathers features from both frontal and bird’s-eye views. The frontal view contains rich but sensitive height information, whereas the bird’s-eye view provides robust features against height uncertainty. We thereby propose a novel Selective Fusion Mechanism to identify and fuse reliable features from both perspectives, reducing the effect of height uncertainty. Moreover, for each view, we incorporate a Multi-Modal Cross-Attention Mechanism to explicitly find location correspondences through cross-modal matching. During the training phase, we also design a Noise-Resistant Matcher to provide better supervision and enhance the robustness of the matching mechanism against sparsity and height uncertainty. Our experimental results, tested on the nuScenes dataset, demonstrate that our method significantly outperforms previous radar-camera auto-calibration methods, as well as existing state-of-the-art LiDAR-camera calibration techniques, establishing a new benchmark for future research. The code is available at https://github.com/nycu-acm/RC-AutoCalib
Van-Tin Luu, Yon-Lin Cai, Vu-Hoang Tran, Walon Wei-Chen Chiu
CVPR4
2025 DynFaceRestore: Balancing Fidelity and Quality in Diffusion-Guided Blind Face Restoration with Dynamic Blur-Level Mapping and Guidance
abstract
Blind Face Restoration aims to recover high-fidelity, detail-rich facial images from unknown degraded inputs, presenting significant challenges in preserving both identity and detail. Pre-trained diffusion models have been increasingly used as image priors to generate fine details. Still, existing methods often use fixed diffusion sampling timesteps and a global guidance scale, assuming uniform degradation. This limitation and potentially imperfect degradation kernel estimation frequently lead to under- or over-diffusion, resulting in an imbalance between fidelity and quality. We propose DynFaceRestore, a novel blind face restoration approach that learns to map any blindly degraded input to Gaussian blurry images. By leveraging these blurry images and their respective Gaussian kernels, we dynamically select the starting timesteps for each blurry image and apply closed-form guidance during the diffusion sampling process to maintain fidelity. Additionally, we introduce a dynamic guidance scaling adjuster that modulates the guidance strength across local regions, enhancing detail generation in complex areas while preserving structural fidelity in contours. This strategy effectively balances the trade-off between fidelity and quality. DynFaceRestore achieves state-of-the-art performance in both quantitative and qualitative evaluations, demonstrating robustness and effectiveness in blind face restoration. Project page at https://nycu-acm.github.io/DynFaceRestore/
Huu-Phu Do, Yi-Cheng Liao, Chi-Wei Hsiao, Walon Wei-Chen Chiu
ICCV6
2025 StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusions
abstract
3D scene representation methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have significantly advanced novel view synthesis. As these methods become prevalent, addressing their vulnerabilities becomes critical. We analyze 3DGS robustness against image-level poisoning attacks and propose a novel density-guided poisoning method. Our method strategically injects Gaussian points into low-density regions identified via Kernel Density Estimation (KDE), embedding viewpoint-dependent illusory objects clearly visible from poisoned views while minimally affecting innocent views. Additionally, we introduce an adaptive noise strategy to disrupt multi-view consistency, further enhancing attack effectiveness. We propose a KDE-based evaluation protocol to assess attack difficulty systematically, enabling objective benchmarking for future research. Extensive experiments demonstrate our method's superior performance compared to state-of-the-art techniques. Project page: https://hentci.github.io/stealthattack/
Bo-Hsu Ke, You-Zhe Xie, Yu-Lun Liu 0001, Walon Wei-Chen Chiu
ICCV4
2025 Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and Approximation
abstract
Although larger datasets are crucial for training large deep models, the rapid growth of dataset size has brought a significant challenge in terms of considerable training costs, which even results in prohibitive computational expenses. Dataset Distillation becomes a popular technique recently to reduce the dataset size via learning a highly compact set of representative exemplars, where the model trained with these exemplars ideally should have comparable performance with respect to the one trained with the full dataset. While most of existing works upon dataset distillation focus on supervised datasets, \todo{we instead aim to distill images and their self-supervisedly trained representations into a distilled set. This procedure, named as Self-Supervised Dataset Distillation, effectively extracts rich information from real datasets, yielding the distilled sets with enhanced cross-architecture generalizability.} Particularly, in order to preserve the key characteristics of original dataset more faithfully and compactly, several novel techniques are proposed: 1) we introduce an innovative parameterization upon images and representations via distinct low-dimensional bases, where the base selection for parameterization is experimentally shown to play a crucial role; 2) we tackle the instability induced by the randomness of data augmentation -- a key component in self-supervised learning but being underestimated in the prior work of self-supervised dataset distillation -- by utilizing predetermined augmentations; 3) we further leverage a lightweight network to model the connections among the representations of augmented views from the same image, leading to more compact pairs of distillation. Extensive experiments conducted on various datasets validate the superiority of our approach in terms of distillation efficiency, cross-architecture generalization, and transfer learning performance.
Sheng-Feng Yu, Jia-Jiun Yao, Walon Wei-Chen Chiu
ICLR3
2025 RMSeg-UDA: Unsupervised Domain Adaptation for Road Marking Segmentation Under Adverse Conditions
abstract
The segmentation of road markings plays a crucial role in visual perception for the autonomous driving system. It enables vehicles to recognize road markings at the pixel-level, and facilitates subsequent path planning, localization, and map construction tasks. Current techniques mainly focus on normal driving scenes (i.e., clear daytime), and the performance would decrease significantly for adverse weather conditions. This work proposes RMSeg-UDA: an unsupervised domain adaptive road marking segmentation framework. By combining schedule self-training and class-conditioned adversarial training, the network utilizes both labeled normal data and unlabeled data from other domains to train a road marking segmentation model. For the evaluation on adverse conditions, a new image dataset, RLMDAC, is established with rainy and nighttime driving scenes. The experiments conducted using both public and our datasets have demonstrated the effectiveness of the proposed technique. Code and dataset are available at https://github.com/stu9113611/RMSeg-UDA.
Yi-Chang Cai, Heng-Chih Hsiao, Walon Wei-Chen Chiu, Huei-Yung Lin, Chiao-Tung Chan
ICRA3
2025 Stands on Shoulders of Giants: Learning to Lift 2D Detection to 3D with Geometry-Driven Objectives
abstract
3D detection of vehicles is an essential component for autonomous driving applications. Nevertheless, collecting the supervised training data for learning 3D vehicle detectors would be costly (e.g. utilization of expensive LiDAR sensors) and labor-intensive (for human annotation). In comparison to 3D detection, 2D object detection has achieved a welldeveloped status, boosting stable and robust performance with widespread application in numerous fields, thanks to the large scale (i.e. amount of samples) of existing training datasets of 2D object detection. Hence, in our work, we propose to realize 3D detection via leveraging the robustness of 2D detectors and developing a network that lifts 2D detections to 3D. With the flexibility of building upon various backbone models (e.g. the models which take image regions detected by 2D detector as inputs to predict their corresponding 3D bounding boxes, or the existing monocular 3D detection models which have the intermediate output of 2 D bounding boxes), we propose several geometry-driven objectives, including projection consistency loss, geometry depth loss, and opposite bin loss, to improve the training upon 2D-to-3D lifting. Our extensive experimental results demonstrate that our proposed geometrydriven objectives not only contribute to the superior results of 3D detection but also provide better generalizability across datasets.
Jhih-Rong Chen, Che-Yuan Chang, Szu-Han Tseng, Chih-Sheng Huang, Yong-Sheng Chen, Walon Wei-Chen Chiu
ICRA6
2025 FuseRoad: Enhancing Lane Shape Prediction Through Semantic Knowledge Integration and Cross-Dataset Training
abstract
The rapid evolution of advanced driver assistance systems (ADAS) has been driven by the advances of deep neural networks, and multi-tasking is essential for autonomous driving systems. This paper presents FuseRoad, a new multi-task model that leverages cross-dataset learning to address the dependency on specific multi-task datasets and reduce the annotation costs. It integrates semantic segmentation and lane detection into an end-to-end framework while providing an effective approach to utilize multiple single-task datasets. By incorporating Semantic Road Knowledge Extractor (SRKE) to direct more attentions on the roadway, FuseRoad enhances the accuracy and reliability of lane detection. The model also employs the logit normalization loss to address the issue of overconfidence commonly faced by conventional lane detection methods. In experiments, FuseRoad outperforms state-of-the-art approaches in both accuracy and F-1 score. The evaluation on semantic segmentation metrics also demonstrates that the proposed technique is highly effective for multi-task road scene analysis. Code and datasets are available at https://github.com/HengChihHsiao/FuseRoad.
Heng-Chih Hsiao, Yi-Chang Cai, Huei-Yung Lin, Walon Wei-Chen Chiu, Chiao-Tung Chan, Chieh-Chih Wang
IV4
2025 Boosting Diffusion Guidance via Learning Degradation-Aware Models for Blind Super Resolution
abstract
Recently, diffusion-based blind super-resolution (SR) methods have shown great ability to generate high-resolution images with abundant high-frequency detail, but the detail is often achieved at the expense of fidelity. Meanwhile, another line of research focusing on rectifying the reverse process of diffusion models (i.e., diffusion guidance), has demonstrated the power to generate high-fidelity results for non-blind SR. However, these methods rely on known degradation kernels, making them difficult to apply to blind SR. To address these issues, we introduce degradation-aware models that can be integrated into the diffusion guidance framework, eliminating the need to know degradation kernels. Additionally, we propose two novel techniques-input perturbation and guidance scalar-to further improve our performance. Extensive experimental results show that our proposed method has superior performance over state-of-the-art methods on blind SR benchmarks.
Shao-Hao Lu, Ren Wang 0014, Walon Wei-Chen Chiu
WACV4
2024 Improving Robustness for Joint Optimization of Camera Pose and Decomposed Low-Rank Tensorial Radiance Fields
abstract
In this paper, we propose an algorithm that allows joint refinement of camera pose and scene geometry represented by decomposed low-rank tensor, using only 2D images as supervision. First, we conduct a pilot study based on a 1D signal and relate our findings to 3D scenarios, where the naive joint pose optimization on voxel-based NeRFs can easily lead to sub-optimal solutions. Moreover, based on the analysis of the frequency spectrum, we propose to apply convolutional Gaussian filters on 2D and 3D radiance fields for a coarse-to-fine training schedule that enables joint camera pose optimization. Leveraging the decomposition property in decomposed low-rank tensor, our method achieves an equivalent effect to brute-force 3D convolution with only incurring little computational overhead. To further improve the robustness and stability of joint optimization, we also propose techniques of smoothed 2D supervision, randomly scaled kernel parameters, and edge-guided loss mask. Extensive quantitative and qualitative evaluations demonstrate that our proposed framework achieves superior performance in novel view synthesis as well as rapid convergence for optimization. The source code is available at https://github.com/Nemo1999/Joint-TensoRF.
Bo-Yu Chen, Walon Wei-Chen Chiu, Yu-Lun Liu 0001
AAAI2
2024 A Recipe for CAC: Mosaic-Based Generalized Loss for Improved Class-Agnostic Counting
Tsung-Han Chou, Walon Wei-Chen Chiu, Jun-Cheng Chen
ACCV (6)3
2024 MCPNet: An Interpretable Classifier via Multi-Level Concept Prototypes
abstract
Recent advancements in post-hoc and inherently inter-pretable methods have markedly enhanced the explanations of black box classifier models. These methods operate either through post-analysis or by integrating concept learning during model training. Although being effective in bridging the semantic gap between a model's latent space and human interpretation, these explanation methods only partially reveal the model's decision-making process. The outcome is typically limited to high-level semantics derived from the last feature map. We argue that the expla-nations lacking insights into the decision processes at low and mid-level features are neither fully faithful nor useful. Addressing this gap, we introduce the Multi-Level Concept Prototypes Classifier (MCPNet), an inherently interpretable model. MCPNet autonomously learns meaningful concept prototypes across multiple feature map levels using Cen-tered Kernel Alignment (CKA) loss and an energy-based weighted PCA mechanism, and it does so without reliance on predefined concept labels. Further, we propose a novel classifier paradigm that learns and aligns multilevel concept prototype distributions for classification purposes via Class-aware Concept Distribution (CCD) loss. Our experiments reveal that our proposed MCPNet while being adapt-able to various model architectures, offers comprehensive multilevel explanations while maintaining classification accuracy. Additionally, its concept distribution-based classification approach shows improved generalization capabilities in few-shot classification scenarios. Project page is available11https://eddie221.github.io/MCPNet/.
Bor-Shiun Wang, Chien-Yi Wang, Walon Wei-Chen Chiu
CVPR3
2024 Two Heads Better Than One: Dual Degradation Representation for Blind Super-Resolution
abstract
Previous methods have demonstrated remarkable performance in single image super-resolution (SISR) tasks with known and fixed degradation (e.g., bicubic downsampling). However, when the actual degradation deviates from these assumptions, these methods may experience significant declines in performance. In this paper, we propose a Dual Branch Degradation Extractor Network to address the blind SR problem. While some blind SR methods assume noisefree degradation and others do not explicitly consider the presence of noise in the degradation model, our approach predicts two unsupervised degradation embeddings that represent blurry and noisy information. The SR network can then be adapted to blur embedding and noise embedding in distinct ways. Furthermore, we treat the degradation extractor as a regularizer to capitalize on differences between SR and HR images. Extensive experiments on several benchmarks demonstrate our method achieves SOTA performance in the blind SR problem.
Hsuan Yuan, Shao-Yu Weng, I-Hsuan Lo, Walon Wei-Chen Chiu, Yu-Syuan Xu, Hao-Chien Hsueh, Jen-Hui Chuang
ICIP4
2024 Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
abstract
Text-to-image diffusion models, e.g. Stable Diffusion (SD), lately have shown remarkable ability in high-quality content generation, and become one of the representatives for the recent wave of transformative AI. Nevertheless, such advance comes with an intensifying concern about the misuse of this generative technology, especially for producing copyrighted or NSFW (i.e. not safe for work) images. Although efforts have been made to filter inappropriate images/prompts or remove undesirable concepts/styles via model fine-tuning, the reliability of these safety mechanisms against diversified problematic prompts remains largely unexplored. In this work, we propose Prompting4Debugging (P4D) as a debugging and red-teaming tool that automatically finds problematic prompts for diffusion models to test the reliability of a deployed safety mechanism. We demonstrate the efficacy of our P4D tool in uncovering new vulnerabilities of SD models with safety mechanisms. Particularly, our result shows that around half of prompts in existing safe prompting benchmarks which were originally considered "safe" can actually be manipulated to bypass many deployed safety mechanisms, including concept removal, negative prompt, and safety guidance. Our findings suggest that, without comprehensive testing, the evaluations on limited safe prompting benchmarks can lead to a false sense of safety for text-to-image models.
Zhi-Yi Chin, Chieh-Ming Jiang, Walon Wei-Chen Chiu
ICML5
2024 Best of Both Sides: Integration of Absolute and Relative Depth Sensing Modalities Based on iToF and RGB Cameras
I-Sheng Fang, Walon Wei-Chen Chiu, Yong-Sheng Chen
ICPR (16)2
2024 Skin the sheep not only once: Reusing Various Depth Datasets to Drive the Learning of Optical Flow
abstract
Optical flow estimation is crucial for various applications in vision and robotics. As the difficulty of collecting ground truth optical flow in real-world scenarios, most of the existing methods of learning optical flow still adopt synthetic dataset for supervised training or utilize photometric consistency across temporally adjacent video frames to drive the unsupervised learning, where the former typically has issues of generalizability while the latter usually performs worse than the supervised ones. To tackle such challenges, we propose to leverage the geometric connection between optical flow estimation and stereo matching (based on the similarity upon finding pixel correspondences across images) to unify various real-world depth estimation datasets for generating supervised training data upon optical flow. Specifically, we turn the monocular depth datasets into stereo ones via synthesizing virtual disparity, thus leading to the flows along the horizontal direction; moreover, we introduce virtual camera motion into stereo data to produce additional flows along the vertical direction. Furthermore, we propose applying geometric augmentations on one image of an optical flow pair, encouraging the optical flow estimator to learn from more challenging cases. Lastly, as the optical flow maps under different geometric augmentations actually exhibit distinct characteristics, an auxiliary classifier which trains to identify the type of augmentation from the appearance of the flow map is utilized to further enhance the learning of the optical flow estimator. Our proposed method is general and is not tied to any particular flow estimator, where extensive experiments based on various datasets and optical flow estimation models verify its efficacy and superiority.
Sheng-Chi Huang, Walon Wei-Chen Chiu
IROS2
2024 T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition
abstract
While widespread access to the Internet and the rapid advancement of generative models boost people's creativity and productivity, the risk of encountering inappropriate or harmful content also increases. To address the aforementioned issue, researchers managed to incorporate several harmful contents datasets with machine learning methods to detect harmful concepts. However, existing harmful datasets are curated by the presence of a narrow range of harmful objects, and only cover real harmful content sources. This restricts the generalizability of methods based on such datasets and leads to the potential misjudgment in certain cases. Therefore, we propose a comprehensive and extensive harmful dataset, VHD11K, consisting of 10,000 images and 1,000 videos, crawled from the Internet and generated by 4 generative models, across a total of 10 harmful categories covering a full spectrum of harmful concepts with non-trival definition. We also propose a novel annotation framework by formulating the annotation process as a multi-agent Visual Question Answering (VQA) task, having 3 different VLMs "debate" about whether the given image/video is harmful, and incorporating the in-context learning strategy in the debating process. Therefore, we can ensure that the VLMs consider the context of the given image/video and both sides of the arguments thoroughly before making decisions, further reducing the likelihood of misjudgments in edge cases. Evaluation and experimental results demonstrate that (1) the great alignment between the annotation from our novel annotation framework and those from human, ensuring the reliability of VHD11K;(2) our full-spectrum harmful dataset successfully identifies the inability of existing harmful content detection methods to detect extensive harmful contents and improves the performance of existing harmfulness recognition methods;(3) our dataset outperforms the baseline dataset, SMID, as evidenced by the superior improvement in harmfulness recognition methods.The entire dataset is publicly available: https://huggingface.co/datasets/denny3388/VHD11K
Chen Yeh, You-Ming Chang, Walon Wei-Chen Chiu
NeurIPS3
2024 Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where
abstract
While image data starts to enjoy the simple-but-effective self-supervised learning scheme built upon masking and self-reconstruction objective thanks to the introduction of tokenization procedure and vision transformer backbone, convolutional neural networks as another important and widely-adopted architecture for image data, though having contrastive-learning techniques to drive the self-supervised learning, still face the difficulty of leveraging such straightforward and general masking operation to benefit their learning process significantly. In this work, we aim to alleviate the burden of including masking operation into the contrastive-learning framework for convolutional neural networks as an extra augmentation method. In addition to the additive but unwanted edges (between masked and unmasked regions) as well as other adverse effects caused by the masking operations for ConvNets, which have been discussed by prior works, we particularly identify the potential problem where for one view in a contrastive sample-pair the randomly-sampled masking regions could be overly concentrated on important/salient objects thus resulting in misleading contrastiveness to the other view. To this end, we propose to explicitly take the saliency constraint into consideration in which the masked regions are more evenly distributed among the foreground and background for realizing the masking-based augmentation. Moreover, we introduce hard negative samples by masking larger regions of salient patches in an input image. Extensive experiments conducted on various datasets, contrastive learning mechanisms, and downstream tasks well verify the efficacy as well as the superior performance of our proposed method with respect to several state-of-the-art baselines. Our code is publicly available at: https://github.com/joycenerd/Saliency-Guided-Masking-for-ConvNets
Zhi-Yi Chin, Chieh-Ming Jiang, Walon Wei-Chen Chiu
WACV5
2024 Best of Both Worlds: Learning Arbitrary-scale Blind Super-Resolution via Dual Degradation Representations and Cycle-Consistency
abstract
Single image super-resolution (SISR) for reconstructing from a low-resolution (LR) input image its corresponding high-resolution (HR) output is a widely-studied research problem in the field of multimedia applications and computer vision. Despite the magic leap brought by recent development of deep neural networks for SISR, such problem is still considered to be quite challenging and non-scalable for the real-world data due to its ill-posed nature, where the degradations happened to the input LR images are usually complex and even unknown (in which the degradations in the test data could be unseen or different from the ones shown in the training dataset). To this end, two branches of SISR methods have emerged: blind super-resolution (blindSR) and arbitrary-scale super-resolution (ASSR), where the former aims to reconstruct SR images under the unknown degradations, while the latter improves the scalability via learning to handle arbitrary up-sampling ratios. In this paper, we propose a holistic framework to take both blind-SR and ASSR tasks (accordingly named as arbitrary-scale blind-SR) into consideration with two main designs: 1) learning dual degradation representations where the implicit and explicit representations of degradation are sequentially extracted from the input LR image, and 2) modeling both upsampling (i.e. LR→HR) and downsampling (i.e. HR→LR) processes at the same time, where they utilize the implicit and explicit degradation representations respectively, in order to enable the cycle-consistency objective and further improve the training. We conduct extensive experiments on various datasets where the results well verify the effectiveness of our proposed framework in handling complex degradations as well as its superiority with respect to several state-of-the-art baselines.
Shao-Yu Weng, Hsuan Yuan, Yu-Syuan Xu, Walon Wei-Chen Chiu
WACV5
2023 Scalable Spatial Memory for Scene Rendering and Navigation
abstract
Neural scene representation and rendering methods have shown promise in learning the implicit form of scene structure without supervision. However, the implicit representation learned in most existing methods is non-expandable and cannot be inferred online for novel scenes, which makes the learned representation difficult to be applied across different reinforcement learning (RL) tasks. In this work, we introduce Scene Memory Network (SMN) to achieve online spatial memory construction and expansion for view rendering in novel scenes. SMN models the camera projection and back-projection as spatially aware memory control processes, where the memory values store the information of the partial 3D area, and the memory keys indicate the position of that area. The memory controller can learn the geometry property from observations without the camera's intrinsic parameters and depth supervision. We further apply the memory constructed by SMN to exploration and navigation tasks. The experimental results reveal the generalization ability of our proposed SMN in large-scale scene synthesis and its potential to improve the performance of spatial RL tasks.
Wen-Cheng Chen, Chu-Song Chen, Walon Wei-Chen Chiu, Min-Chun Hu 0001
AAAI3
2023 Are You Killing Time? Predicting Smartphone Users' Time-killing Moments via Fusion of Smartphone Sensor Data and Screenshots
abstract
Time-killing on smartphones has become a pervasive activity, and could be opportune for delivering content to their users. This research is believed to be the first attempt at time-killing detection, which leverages the fusion of phone-sensor and screenshot data. We collected nearly one million user-annotated screenshots from 36 Android users. Using this dataset, we built a deep-learning fusion model, which achieved a precision of 0.83 and an AUROC of 0.72. We further employed a two-stage clustering approach to separate users into four groups according to the patterns of their phone-usage behaviors, and then built a fusion model for each group. The performance of the four models, though diverse, yielded better average precision of 0.87 and AUROC of 0.76, and was superior to that of the general/unified model shared among all users. We investigated and discussed the features of the four time-killing behavior clusters that explain why the models’ performance differ.
Yu-Chun Chen, Yu-Jen Lee, Kuei-Chun Kao, Jie Tsai, En-Chi Liang, Walon Wei-Chen Chiu, Faye Shih, Yung-Ju Chang
CHI6
2023 Multimodal Prompting with Missing Modalities for Visual Recognition
abstract
In this paper, we tackle two challenges in multimodal learning for visual recognition: 1) when missing-modality occurs either during training or testing in real-world situations; and 2) when the computation resources are not available to finetune on heavy transformer models. To this end, we propose to utilize prompt learning and mitigate the above two challenges together. Specifically, our modality-missing-aware prompts can be plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 1% learnable parameters compared to training the entire model. We further explore the effect of different prompt configurations and analyze the robustness to missing modality. Extensive experiments are conducted to show the effectiveness of our prompt learning framework that improves the performance under various missing-modality cases, while alleviating the requirement of heavy model retraining. Code is available.11https://github.com/YiLunLee/missing_aware_prompts
Yi-Lun Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Chen-Yu Lee
CVPR3
2023 Decontamination Transformer For Blind Image Inpainting
abstract
Blind image inpainting aims at recovering the content from a corrupted image in which the mask indicating the corrupted regions is not available in inference time. Inspired that most existing methods for inpainting suffer from complex contamination, we propose a model that explicitly predicts the realvalued alpha mask and contaminant to eliminate the contamination from the corrupted image, thus improving the inpainting performance. To enhance the overall semantic consistency, the attention mechanism of transformers is exploited and integrated into our inpainting network. We conduct extensive experiments to verify our method against blind and non-blind inpainting models and demonstrate its effectiveness and generalizability to different sources of contaminant.
Chun-Yi Li, Yen-Yu Lin, Walon Wei-Chen Chiu
ICASSP3
2023 TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine Perception
abstract
This work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Transformer-based image compression framework, termed TransTIC. Inspired by visual prompt tuning, TransTIC adopts an instance-specific prompt generator to inject instance-specific prompts to the encoder and task-specific prompts to the decoder. Extensive experiments show that our proposed method is capable of transferring the base codec to various machine tasks and outperforms the competing methods significantly. To our best knowledge, this work is the first attempt to utilize prompting on the low-level image compression task.
Yi-Hsin Chen, Ying-Chieh Weng, Chia-Hao Kao, Cheng Chien, Walon Wei-Chen Chiu, Wen-Hsiao Peng
ICCV5
2023 Transformer-Based Variable-Rate Image Compression with Region-of-Interest Control
abstract
This paper proposes a transformer-based learned image compression system. It is capable of achieving variable-rate compression with a single model while supporting the region-of-interest (ROI) functionality. Inspired by prompt tuning, we introduce prompt generation networks to condition the transformer-based autoencoder of compression. Our prompt generation networks generate content-adaptive tokens according to the input image, an ROI mask, and a rate parameter. The separation of the ROI mask and the rate parameter allows an intuitive way to achieve variable-rate and ROI coding simultaneously. Extensive experiments validate the effectiveness of our proposed method and confirm its superiority over the other competing methods.
Chia-Hao Kao, Ying-Chieh Weng, Yi-Hsin Chen, Walon Wei-Chen Chiu, Wen-Hsiao Peng
ICIP4
2023 MENTOR: Multilingual Text Detection Toward Learning by Analogy
abstract
Text detection is frequently used in vision-based mobile robots when they need to interpret texts in their surroundings to perform a given task. For instance, delivery robots in multilingual cities need to be capable of doing multilingual text detection so that the robots can read traffic signs and road markings. Moreover, the target languages change from region to region, implying the need of efficiently re-training the models to recognize the novel/new languages. However, collecting and labeling training data for novel languages are cumbersome, and the efforts to re-train an existing/trained text detector are considerable. Even worse, such a routine would repeat whenever a novel language appears. This motivates us to propose a new problem setting for tackling the aforementioned challenges in a more efficient way: “We ask for a generalizable multilingual text detection framework to detect and identify both seen and unseen language regions inside scene images without the requirement of collecting supervised training data for unseen languages as well as model re-training”. To this end, we propose “MENTOR”, the first work to realize a learning strategy between zero-shot learning and few-shot learning for multilingual scene text detection. During the training phase, we leverage the “zero-cost” synthesized printed texts and the available training/seen languages to learn the meta-mapping from printed texts to language-specific kernel weights. Meanwhile, dynamic convolution networks guided by the language-specific kernel are trained to realize a detection-by-feature-matching scheme. In the inference phase, “zero-cost” printed texts are synthesized given a new target language. By utilizing the learned meta-mapping and the matching network, our “MENTOR” can freely identify the text regions of the new language. Experiments show our model can achieve comparable results with supervised methods for seen languages and outperform other methods in detecting unseen languages.
Hsin-Ju Lin, Tsu-Chun Chung, Ching-Chun Hsiao, Walon Wei-Chen Chiu
IROS5
2023 Continually-Adapted Margin and Multi-Anchor Distillation for Class-Incremental Learning
abstract
This paper addresses the problem of class-incremental learning. The model is trained to recognize the classes added incrementally. It thus suffers from the challenging issue of catastrophic forgetting. Stemming from the knowledge distillation idea of attempting to retain the model's knowledge on seen classes while learning the newly-added ones, we advance to further alleviate the catastrophic forgetting via our proposed multi-anchor distillation objective, which is realized by constraining the spatial relationship between the input data and the multiple class embeddings of each seen class in the feature space while training the model. Moreover, since the knowledge distillation for incremental learning generally relies on keeping a replay buffer to store the samples of seen classes, the buffer of limited size brings another issue of class imbalance: the number of samples from each seen class decreases gradually, thus being much smaller than the number of samples from each new class. We therefore propose to introduce the continually-adapted margin into the classification objective for tackling the prediction bias towards new classes caused by the class imbalance. Experiments are conducted on various datasets and settings to demonstrate the effectiveness and superior performance of our proposed techniques in comparison to several state-of-the-art baselines.
Yi-Hsin Chen, Dian-Shan Chen, Ying-Chieh Weng, Wen-Hsiao Peng, Walon Wei-Chen Chiu
SMC5
2023 Data Efficient Incremental Learning via Attentive Knowledge Replay
abstract
Class-incremental learning (CIL) tackles the problem of continuously optimizing a classification model to support growing number of classes, where the data of novel classes arrive in streams. Recent works propose to use representative exemplars of learnt classes, and replay the knowledge of them afterward under certain memory constraints. However, training on a fixed set of exemplars with an imbalanced proportion to the new data leads to strong biases in the trained models. In this paper, we propose an attentive knowledge replay framework to refresh the knowledge of previously learnt classes during incremental learning, which generates virtual training samples by blending between pairs of data. Particularly, we design an attention module that learns to predict the adaptive blending weights in accordance with their relative importance to the overall objective, where the importance is derived from the change of the image features over incremental phases. Our strategy of attentive knowledge replay encourages the model to learn smoother decision boundaries and thus improves its generalization beyond memorizing the exemplars. We validate our design in a standard class-incremental learning setup and demonstrate its flexibility in various settings.
Yi-Lun Lee, Dian-Shan Chen, Chen-Yu Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
SMC5
2023 Mitigating Forgetting in Continual Learning via Contrasting Semantically Distinct Augmentations
abstract
Online continual learning (OCL) aims to enable model learning from a non-stationary data stream to continuously acquire new knowledge as well as retain the learnt one. Under the constraints of having limited system size and computational cost, in which the main challenge comes from the “catastrophic forgetting” issue - the inability to well remember the learnt knowledge while learning the new ones. With the specific focus on the class-incremental OCL scenario, i.e. OCL for classification, the recent advance incorporates the contrastive learning technologies for learning more generalised feature representation to achieve the state-of-the-art performance but is still unable to fully resolve the catastrophic forgetting. In this paper, we follow the strategy of adopting contrastive learning but further introduce the semantically distinct augmentation technique, in which it leverages strong augmentation to generate more data samples, and we show that considering these samples semantically different from their original classes (thus being related to the out-of-distribution samples) in the contrastive learning mechanisms contributes to alleviate forgetting and facilitate model stability. Moreover, in addition to contrastive learning, the typical classification mechanism and objective (i.e. softmax classifier and cross-entropy loss) are included in our model design for utilising the label information, but particularly equipped with a sampling strategy to tackle the tendency of favouring the new classes (i.e. model bias towards the recently learnt classes). Upon conducting extensive experiments on CIFAR-10, CIFAR-100, and Mini-Imagenet datasets, our proposed method is shown to achieve superior performance against various baselines.
Sheng-Feng Yu, Walon Wei-Chen Chiu
SMC2
2023 Adaptively-Realistic Image Generation from Stroke and Sketch with Diffusion Model
abstract
Generating images from hand-drawings is a crucial and fundamental task in content creation. The translation is difficult as there exist infinite possibilities and the different users usually expect different outcomes. Therefore, we propose a unified framework supporting a three-dimensional control over the image synthesis from sketches and strokes based on diffusion models. Users can not only decide the level of faithfulness to the input strokes and sketches, but also the degree of realism, as the user inputs are usually not consistent with the real images. Qualitative and quantitative experiments demonstrate that our framework achieves state-of-the-art performance while providing flexibility in generating customized images with control over shape, color, and realism. Moreover, our method unleashes applications such as editing on real images, generation with partial sketches and strokes, and multi-domain multi-modal synthesis.
Shin-I Cheng, Yu-Jie Chen, Walon Wei-Chen Chiu, Hung-Yu Tseng, Hsin-Ying Lee 0001
WACV3
2023 BiFuse++: Self-Supervised and Efficient Bi-Projection Fusion for 360° Depth Estimation
abstract
Due to the rise of spherical cameras, monocular 360$^\circ$depth estimation becomes an important technique for many applications (e.g., autonomous systems). Thus, state-of-the-art frameworks for monocular 360$^\circ$depth estimation such as bi-projection fusion in BiFuse are proposed. To train such a framework, a large number of panoramas along with the corresponding depth ground truths captured by laser sensors are required, which highly increases the cost of data collection. Moreover, since such a data collection procedure is time-consuming, the scalability of extending these methods to different scenes becomes a challenge. To this end, self-training a network for monocular depth estimation from 360$^\circ$videos is one way to alleviate this issue. However, there are no existing frameworks that incorporate bi-projection fusion into the self-training scheme, which highly limits the self-supervised performance since bi-projection fusion can leverage information from different projection types. In this paper, we propose BiFuse++ to explore the combination of bi-projection fusion and the self-training scenario. To be specific, we propose a new fusion module and Contrast-Aware Photometric Loss to improve the performance of BiFuse and increase the stability of self-training on real-world videos. We conduct both supervised and self-supervised experiments on benchmark datasets and achieve state-of-the-art performance.
Fu-En Wang, Yu-Hsuan Yeh, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Vector Quantized Image-to-Image Translation
Yu-Jie Chen, Shin-I Cheng, Walon Wei-Chen Chiu, Hung-Yu Tseng, Hsin-Ying Lee 0001
ECCV (16)3
2022 3D-PL: Domain Adaptive Depth Estimation with 3D-Aware Pseudo-Labeling
Yu-Ting Yen, Chia-Ni Lu, Walon Wei-Chen Chiu, Yi-Hsuan Tsai
ECCV (27)3
2022 Find The Way Back: Invertible Kernel Estimator For Blind Image Super-Resolution
abstract
We address the task of zero-shot blind image super-resolution, where it aims to recover the high-resolution details from the low-resolution input image under a challenging problem setting of having no external training data, no prior assumption on the downsampling kernel, and no pre-training components used for estimating the downsampling kernel. While existing zero-shot blind super-resolution works follow the strategy of firstly estimating the downsampling kernel via cross-scale recurrence and then learning the non-blind upsampling model, we in turn propose a carefully-designed invertible network for modeling both the downsampling and upsampling operations at once. Specifically, the invertible property enables the use of cross-scale recurrence across more scales and thus further benefits the overall model training. We conduct extensive experiments to demonstrate our proposed method’s superior performance over several baselines and its effectiveness in handling the images downsampled by nonlinear kernels.
Ting-Wei Chang, Walon Wei-Chen Chiu
ICASSP2
2022 MAML is a Noisy Contrastive Learner in Classification
Chia-Hsiang Kao, Walon Wei-Chen Chiu
ICLR2
2022 RPG: Learning Recursive Point Cloud Generation
abstract
In this paper we propose a novel point cloud generator that is able to reconstruct and generate 3D point clouds composed of semantic parts. Given a latent representation of the target 3D model, the generation starts from a single point and gets expanded recursively to produce the high-resolution point cloud via a sequence of point expansion stages. During the recursive procedure of generation, we not only obtain the coarse-to-fine point clouds for the target 3D model from every expansion stage, but also unsupervisedly discover the semantic segmentation of the target model according to the hierarchical/parent-child relation between the points across expansion stages. Moreover, the expansion modules and other elements used in our recursive generator are mostly sharing weights thus making the overall framework light and efficient. Extensive experiments are conducted to show that our point cloud generator has comparable or even superior performance on both generation and reconstruction tasks in comparison to various baselines, and provides the consistent co-segmentation among instances of the same object class.
Wei-Jan Ko, Chen-Yi Chiu, Yu-Liang Kuo, Walon Wei-Chen Chiu
IROS4
2022 Improving Single-View Mesh Reconstruction for Unseen Categories via Primitive-Based Representation and Mesh Augmentation
abstract
As most existing works of single-view 3D reconstruction aim at learning the better mapping functions to directly transform the 2D observation into the corresponding 3D shape for achieving state-of-the-art performance, there often comes a potential concern on having the implicit bias towards the seen classes learnt in their models (i.e. reconstruction intertwined with the classification) thus leading to poor generalizability for the unseen object categories. Moreover, such implicit bias typically stemmed from adopting the object-centered coordinate in their model designs, in which the reconstructed 3D shapes of the same class are all aligned to the same canonical pose regardless of different view-angles in the 2D observations. To this end, we propose an end-to-end framework to reconstruct the 3D mesh from a single image, where the reconstructed mesh is not only view-centered (i.e. its 3D pose respects the viewpoint of the 2D observation) but also preliminarily represented as a composition of volumetric 3D primitives before being further deformed into the fine-grained mesh to capture the shape details. In particular, the usage of volumetric primitives is motivated from the assumption that there generally exists some similar shape parts shared across various object categories, learning to estimate the primitive-based 3D model thus becomes more generalizable to the unseen categories. Furthermore, we advance to propose a novel mesh augmentation strategy, CvxRearrangement, to enrich the distribution of training shapes, which contributes to increasing the robustness of our proposed model and achieves better generalization. Extensive experiments demonstrate that our proposed method provides superior performance on both unseen and seen classes in comparison to several representative baselines of single-view 3D reconstruction.
Yu-Liang Kuo, Wei-Jan Ko, Chen-Yi Chiu, Walon Wei-Chen Chiu
IROS4
2022 Self-Supervised Feature Learning from Partial Point Clouds via Pose Disentanglement
abstract
Self-supervised learning on point clouds has gained a lot of attention recently, since it addresses the label-efficiency and domain-gap problems on point cloud tasks. In this paper, we propose a novel self-supervised framework to learn informative features from partial point clouds. We leverage partial point clouds scanned by LiDAR that contain both content and pose attributes, and we show that disentangling such two factors from partial point clouds enhances feature learning. To this end, our framework consists of three main parts: 1) a completion network to capture holistic semantics of point clouds; 2) a pose regression network to understand the viewing angle where partial data is scanned from; 3) a partial reconstruction network to encourage the model to learn content and pose features. To demonstrate the robustness of the learnt feature representations, we conduct several downstream tasks including classification, part segmentation, and registration, with comparisons against state-of-the-art methods. Our method not only outperforms existing self-supervised methods, but also shows a better generalizability across synthetic and real-world datasets.
Meng-Shiun Tsai, Pei-Ze Chiang, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
IROS4
2022 Make an Omelette with Breaking Eggs: Zero-Shot Learning for Novel Attribute Synthesis
abstract
Most of the existing algorithms for zero-shot classification problems typically rely on the attribute-based semantic relations among categories to realize the classification of novel categories without observing any of their instances. However, training the zero-shot classification models still requires attribute labeling for each class (or even instance) in the training dataset, which is also expensive. To this end, in this paper, we bring up a new problem scenario: ''Can we derive zero-shot learning for novel attribute detectors/classifiers and use them to automatically annotate the dataset for labeling efficiency?'' Basically, given only a small set of detectors that are learned to recognize some manually annotated attributes (i.e., the seen attributes), we aim to synthesize the detectors of novel attributes in a zero-shot learning manner. Our proposed method, Zero-Shot Learning for Attributes (ZSLA), which is the first of its kind to the best of our knowledge, tackles this new research problem by applying the set operations to first decompose the seen attributes into their basic attributes and then recombine these basic attributes into the novel ones. Extensive experiments are conducted to verify the capacity of our synthesized detectors for accurately capturing the semantics of the novel attributes and show their superior performance in terms of detection and localization compared to other baseline approaches. Moreover, we demonstrate the application of automatic annotation using our synthesized detectors on Caltech-UCSD Birds-200-2011 dataset. Various generalized zero-shot classification algorithms trained upon the dataset re-annotated by ZSLA shows comparable performance with those trained with the manual ground-truth annotations.
Yu Hsuan Li, Tzu-Yin Chao, Walon Wei-Chen Chiu
NeurIPS5
2022 Stylizing 3D Scene via Implicit Representation and HyperNetwork
abstract
In this work, we aim to address the 3D scene stylization problem - generating stylized images of the scene at arbitrary novel view angles. A straightforward solution is to combine existing novel view synthesis and image/video style transfer approaches, which often leads to blurry results or inconsistent appearance. Inspired by the high-quality results of the neural radiance fields (NeRF) method, we propose a joint framework to directly render novel views with the desired style. Our framework consists of two components: an implicit representation of the 3D scene with the neural radiance fields model, and a hypernetwork to transfer the style information into the scene representation. To alleviate the training difficulties and memory burden, we propose a two-stage training procedure and a patch sub-sampling approach to optimize the style and content losses with the neural radiance fields model. After optimization, our model is able to render consistent novel views at arbitrary view angles with arbitrary style. Both quantitative evaluation and human subject study have demonstrated that the proposed method generates faithful stylization results with consistent appearance across different views.
Pei-Ze Chiang, Meng-Shiun Tsai, Hung-Yu Tseng, Wei-Sheng Lai, Walon Wei-Chen Chiu
WACV5
2021 Learning to Hide Residual for Boosting Image Compression
Yi-Lun Lee, Yen-Chung Chen, Min-Yuan Tseng, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
BMVC5
2021 Bridging the Visual Gap: Wide-Range Image Blending
abstract
In this paper we propose a new problem scenario in image processing, wide-range image blending, which aims to smoothly merge two different input photos into a panorama by generating novel image content for the intermediate region between them. Although such problem is closely related to the topics of image inpainting, image outpainting, and image blending, none of the approaches from these topics is able to easily address it. We introduce an effective deep-learning model to realize wide-range image blending, where a novel Bidirectional Content Transfer module is proposed to perform the conditional prediction for the feature representation of the intermediate region via recurrent neural networks. In addition to ensuring the spatial and semantic consistency during the blending, we also adopt the contextual attention mechanism as well as the adversarial learning scheme in our proposed method for improving the visual quality of the resultant panorama. We experimentally demonstrate that our proposed method is not only able to produce visually appealing results for wide-range image blending, but also able to provide superior performance with respect to several baselines built upon the state-of-the-art image inpainting and outpainting approaches.
Chia-Ni Lu, Ya-Chu Chang, Walon Wei-Chen Chiu
CVPR3
2021 LED2-Net: Monocular 360deg Layout Estimation via Differentiable Depth Rendering
abstract
Although significant progress has been made in room layout estimation, most methods aim to reduce the loss in the 2D pixel coordinate rather than exploiting the room structure in the 3D space. Towards reconstructing the room layout in 3D, we formulate the task of 360° layout estimation as a problem of predicting depth on the horizon line of a panorama. Specifically, we propose the Differentiable Depth Rendering procedure to make the conversion from layout to depth prediction differentiable, thus making our proposed model end-to-end trainable while leveraging the 3D geometric information, without the need of providing the ground truth depth. Our method achieves state-of-the-art performance on numerous 360° layout benchmark datasets. Moreover, our formulation enables a pre-training step on the depth dataset, which further improves the generalizability of our layout estimation model.
Fu-En Wang, Yu-Hsuan Yeh, Min Sun 0001, Walon Wei-Chen Chiu, Yi-Hsuan Tsai
CVPR4
2021 Domain Adaptation for Learning Generator From Paired Few-Shot Data
abstract
We propose a Paired Few-shot GAN (PFS-GAN) model for learning generators with sufficient source data and a few target data. While generative model learning typically needs large-scale training data, our PFS-GAN not only uses the concept of few-shot learning but also domain shift to transfer the knowledge across domains, which alleviates the issue of obtaining low-quality generator when only trained with target domain data. The cross-domain datasets are assumed to have two properties: (1) each target-domain sample has its source-domain correspondence and (2) two domains share similar content information but different appearance. Our PFS-GAN aims to learn the disentangled representation from images, which composed of domain-invariant content features and domain-specific appearance features. Furthermore, a relation loss is introduced on the content features while shifting the appearance features to increase the structural diversity. Extensive experiments show that our method has better quantitative and qualitative results on the generated target-domain data with higher diversity in comparison to several baselines.
Chun-Chih Teng, Walon Wei-Chen Chiu
ICASSP3
2021 Learning Facial Representations from the Cycle-consistency of Face
abstract
Faces manifest large variations in many aspects, such as identity, expression, pose, and face styling. Therefore, it is a great challenge to disentangle and extract these characteristics from facial images, especially in an unsupervised manner. In this work, we introduce cycle-consistency in facial characteristics as free supervisory signal to learn facial representations from unlabeled facial images. The learning is realized by superimposing the facial motion cycle-consistency and identity cycle-consistency constraints. The main idea of the facial motion cycle-consistency is that, given a face with expression, we can perform de-expression to a neutral face via the removal of facial motion and further perform re-expression to reconstruct back to the original face. The main idea of the identity cycle-consistency is to exploit both de-identity into mean face by depriving the given neutral face of its identity via feature re-normalization and re-identity into neutral face by adding the personal attributes to the mean face. At training time, our model learns to disentangle two distinct facial representations to be useful for performing cycle-consistent face reconstruction. At test time, we use the linear protocol scheme for evaluating facial representations on various tasks, including facial expression recognition and head pose regression. We also can directly apply the learnt facial representations to person recognition, frontalization and image-to-image translation. Our experiments show that the results of our approach is competitive with those of existing methods, demonstrating the rich and unique information embedded in the dis-entangled representations. Code is available at https://github.com/JiaRenChang/FaceCycle.
Jia-Ren Chang, Yong-Sheng Chen, Walon Wei-Chen Chiu
ICCV3
2021 Towards Interpretable Deep Networks for Monocular Depth Estimation
abstract
Deep networks for Monocular Depth Estimation (MDE) have achieved promising performance recently and it is of great importance to further understand the interpretability of these networks. Existing methods attempt to provide post-hoc explanations by investigating visual cues, which may not explore the internal representations learned by deep networks. In this paper, we find that some hidden units of the network are selective to certain ranges of depth, and thus such behavior can be served as a way to interpret the internal representations. Based on our observations, we quantify the interpretability of a deep MDE network by the depth selectivity of its hidden units. Moreover, we then propose a method to train interpretable MDE deep networks without changing their original architectures, by assigning a depth range for each unit to select. Experimental results demonstrate that our method is able to enhance the interpretability of deep MDE networks by largely improving the depth selectivity of their units, while not harming or even improving the depth estimation accuracy. We further provide comprehensive analysis to show the reliability of selective units, the applicability of our method on different layers, models, and datasets, and a demonstration on analysis of model error. Source code and models are available at https://github.com/youzunzhi/InterpretableMDE.
Zunzhi You, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Guanbin Li
ICCV3
2021 Inverse Halftone Colorization: Making Halftone Prints Color Photos
abstract
We propose to address the Inverse Halftone Colorization task, which tries to recover colorful images from black and white halftone prints, and can be treated as the joint problem of inverse halftone and colorization. Although the inverse halftone colorization seems to be achievable via applying inverse halftone followed by colorization, our proposed method advances from two perspectives: (1) we empirically discover that the orders of cascading inverse halftone and colorization (i.e first inverse halftone then colorization versus the reverse order) would lead to results with different properties, hence a fusion scheme is proposed to integrate their results; (2) we introduce several novel losses to encourage the realness, diversity, and the structural coherence of the colorization. Moreover, our model is flexible to support both exemplar-based and random colorization. We conduct extensive experiments to demonstrate the efficacy of our method as well as verify the contributions of our design choices.
Yu-Ting Yen, Chia-Chi Cheng, Walon Wei-Chen Chiu
ICIP3
2021 Robust 360-8PA: Redesigning The Normalized 8-point Algorithm for 360-FoV Images
abstract
In this paper, we present a novel preconditioning strategy for the classic 8-point algorithm (8-PA) for estimating an essential matrix from 360-FoV images (i.e., equirectangular images) in spherical projection. To alleviate the effect of uneven key-feature distributions and outlier correspondences, which can potentially decrease the accuracy of an essential matrix, our method optimizes a non-rigid transformation to deform a spherical camera into a new spatial domain, defining a new constraint and a more robust and accurate solution for an essential matrix. Through several experiments using random synthetic points, 360-FoV, and fish-eye images, we demonstrate that our normalization can increase the camera pose accuracy about 20% without significantly overhead the computation time. In addition, we present further benefits of our method through both a constant weighted least-square optimization that improves further the well known Gold Standard Method (GSM) (i.e., the non-linear optimization by using epipolar errors); and a relaxation of the number of RANSAC iterations, both showing that our normalization outcomes a more reliable, robust, and accurate solution.
Bolivar Solarte, Chin-Hsuan Wu, Kuan-Wei Lu 0001, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
ICRA5
2021 Demystifying T1-MRI to FDG18-PET Image Translation via Representational Similarity
Chia-Hsiang Kao, Yong-Sheng Chen, Li-Fen Chen, Walon Wei-Chen Chiu
MICCAI (3)4
2021 An unsupervised video game playstyle metric via state discretization
abstract
On playing video games, different players usually have their own playstyles. Recently, there have been great improvements for the video game AIs on the playing strength. However, past researches for analyzing the behaviors of players still used heuristic rules or the behavior features with the game-environment support, thus being exhausted for the developers to define the features of discriminating various playstyles. In this paper, we propose the first metric for video game playstyles directly from the game observations and actions, without any prior specification on the playstyle in the target game. Our proposed method is built upon a novel scheme of learning discrete representations that can map game observations into latent discrete states, such that playstyles can be exhibited from these discrete states. Namely, we measure the playstyle distance based on game observations aligned to the same states. We demonstrate high playstyle accuracy of our metric in experiments on some video game platforms, including TORCS, RGSK, and seven Atari games, and for different agents including rule-based AI bots, learning-based AI bots, and human players.
Chiu-Chou Lin, Walon Wei-Chen Chiu, I-Chen Wu
UAI2
2021 Single Image Reflection Removal with Edge Guidance, Reflection Classifier, and Recurrent Decomposition
abstract
Removing undesired reflection from an image captured through a glass window is a notable task in computer vision. In this paper, we propose a novel model with auxiliary techniques to tackle the problem of single image reflection removal. Our model takes a reflection contaminated image as input, and decomposes it into the reflection layer and the transmission layer. In order to ensure quality of the transmission layer, we introduce three auxiliary techniques into our architecture, including the edge guidance, a reflection classifier, and the recurrent decomposition. The contributions and the efficacy of these techniques are investigated and verified in the ablation study. Furthermore, in comparison to the state-of-the-art baselines of reflection removal, both quantitative and qualitative results demonstrate that our proposed method is able to deal with different kinds of images, achieving the best results in average.
Ya-Chu Chang, Chia-Ni Lu, Chia-Chi Cheng, Walon Wei-Chen Chiu
WACV4
2021 Dual-Stream Fusion Network for Spatiotemporal Video Super-Resolution
abstract
Visual data upsampling has been an important research topic for improving the perceptual quality and benefiting various computer vision applications. In recent years, we have witnessed remarkable progresses brought by the re-naissance of deep learning techniques for video or image super-resolution. However, most existing methods focus on advancing super-resolution at either spatial or temporal direction, i.e, to increase the spatial resolution or the video frame rate. In this paper, we instead turn to discuss both directions jointly and tackle the spatiotemporal upsampling problem. Our method is based on an important observation that: even the direct cascade of prior research in spatial and temporal super-resolution can achieve the spatiotemporal upsampling, changing orders for combining them would lead to results with a complementary property. Thus, we propose a dual-stream fusion network to adaptively fuse the intermediate results produced by two spatiotemporal up-sampling streams, where the first stream applies the spatial super-resolution followed by the temporal super-resolution, while the second one is with the reverse order of cascade. Extensive experiments verify the efficacy of the proposed method against several baselines. Moreover, we investigate various spatial and temporal upsampling methods as the basis in our two-stream model and demonstrate the flexibility with wide applicability of the proposed framework.
Min-Yuan Tseng, Yen-Chung Chen, Yi-Lun Lee, Wei-Sheng Lai, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
WACV6
2020 Class-Incremental Learning with Rectified Feature-Graph Preservation
Cheng-Hsun Lei, Yi-Hsin Chen, Wen-Hsiao Peng, Walon Wei-Chen Chiu
ACCV (6)4
2020 Boosting Image and Video Compression via Learning Latent Residual Patterns
Yen-Chung Chen, Keng-Jui Chang, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
BMVC4
2020 Time Flies: Animating a Still Image With Time-Lapse Video As Reference
abstract
Time-lapse videos usually perform eye-catching appearances but are often hard to create. In this paper, we propose a self-supervised end-to-end model to generate the time-lapse video from a single image and a reference video. Our key idea is to extract both the style and the features of temporal variation from the reference video, and transfer them onto the input image. To ensure both the temporal consistency and realness of our resultant videos, we introduce several novel designs in our architecture, including classwise NoiseAdaIN, flow loss, and the video discriminator. In comparison to the baselines of state-of-the-art style transfer approaches, our proposed method is not only efficient in computation but also able to create more realistic and temporally smooth time-lapse video of a still image, with its temporal variation consistent to the reference.
Chia-Chi Cheng, Hung-Yu Chen, Walon Wei-Chen Chiu
CVPR3
2020 BiFuse: Monocular 360 Depth Estimation via Bi-Projection Fusion
abstract
Depth estimation from a monocular 360 image is an emerging problem that gains popularity due to the availability of consumer-level 360 cameras and the complete surrounding sensing capability. While the standard of 360 imaging is under rapid development, we propose to predict the depth map of a monocular 360 image by mimicking both peripheral and foveal vision of the human eye. To this end, we adopt a two-branch neural network leveraging two common projections: equirectangular and cubemap projections. In particular, equirectangular projection incorporates a complete field-of-view but introduces distortion, whereas cubemap projection avoids distortion but introduces discontinuity at the boundary of the cube. Thus we propose a bi-projection fusion scheme along with learnable masks to balance the feature map from the two projections. Moreover, for the cubemap projection, we propose a spherical padding procedure which mitigates discontinuity at the boundary of each face. We apply our method to four panorama datasets and show favorable results against the existing state-of-the-art methods.
Fu-En Wang, Yu-Hsuan Yeh, Min Sun 0001, Walon Wei-Chen Chiu, Yi-Hsuan Tsai
CVPR4
2020 Colorization of Depth Map via Disentanglement
Chung-Sheng Lai, Zunzhi You, Chingchun Huang, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
ECCV (7)5
2020 Real-time Monocular Depth Estimation with Extremely Light-Weight Neural Network
abstract
Obstacle avoidance and environment sensing are crucial applications in autonomous driving and robotics. Among all types of sensors, RGB camera is widely used in these applications as it can offer rich visual contents with relatively low-cost, and using a single image to perform depth estimation has become one of the main focuses in resent research works. However, prior works usually rely on highly complicated computation and power-consuming GPU to achieve such task; therefore, we focus on developing a real-time light-weight system for depth prediction in this paper. Based on the well-known encoder-decoder architecture, we propose a supervised learning-based CNN with detachable decoders that produce depth predictions with different scales. We also formulate a novel log-depth loss function that computes the difference of predicted depth map and ground truth depth map in log space, so as to increase the prediction accuracy for nearby locations. To train our model efficiently, we generate depth map and semantic segmentation with complex teacher models. Via a series of ablation studies and experiments, it is validated that our model can efficiently performs real-time depth prediction with only 0.32M parameters, with the best trained model outperforms previous works on KITTI dataset for various evaluation matrices.
Mian-Jhong Chiu, Walon Wei-Chen Chiu, Hua-Tsung Chen, Jen-Hui Chuang
ICPR2
2020 Learning Low-Shot Generative Networks for Cross-Domain Data
abstract
We tackle a novel problem of learning generators for cross-domain data under a specific scenario of low-shot learning. Basically, given a source domain with sufficient amount of training data, we aim to transfer the knowledge of its generative process to another target domain, which not only has few data samples but also contains the domain shift with respect to the source domain. This problem has great potential in practical use and is different from the well-known image translation task, as the target-domain data can be generated without requiring any source-domain ones and the large data consumption for learning target-domain generator can be alleviated. Built upon a cross-domain dataset where (1) each of the low shots in the target domain has its correspondence in the source and (2) these two domains share the similar content information but different appearance, two approaches are proposed: a Latent-Disentanglement-Orientated model (LaDo) and a Generative-Hierarchy-Oriented (GenHo) model. Our LaDo and GenHo approaches address the problem from different perspectives, where the former relies on learning the disentangled representation composed of domain-invariant content features and domain-specific appearance ones; while the later decomposes the generative process of a generator into two parts for synthesizing the content and appearance sequentially. We perform extensive experiments under various settings of cross-domain data and show the efficacy of our models for generating target-domain data with the abundant content variance as in the source domain, which lead to the favourable performance in comparison to several baselines.
Hsuan-Kai Kao, Cheng-Che Lee, Walon Wei-Chen Chiu
ICPR3
2020 DEN: Disentangling and Exchanging Network for Depth Completion
abstract
In this paper, we tackle the depth completion problem. Conventional depth sensors usually produce incomplete depth maps due to the property of surface reflection, especially for the window areas, metal surfaces, and object boundaries. However, we observe that the corresponding RGB images are still dense and preserve all of the useful structural information. The observation brings us to the question of whether we can borrow this structural information from RGB images to inpaint the corresponding incomplete depth maps. In this paper, we answer that question by proposing a Disentangling and Exchanging Network (DEN) for depth completion. The network is designed based on the assumption that after suitable feature disentanglement, RGB images and depth maps share a common domain for representing structural information. So we firstly disentangle both RGB and depth images into domain-invariant content parts, which contain structural information, and domain-specific style parts. Then, by exchanging the complete structural information extracted from the RGB image with incomplete information extracted from the depth map, we can generate the complete version of the depth map. Furthermore, to address the mixed-depth problem, a newly proposed depth representation is applied. By modeling depth estimation as a classification problem coupled with coefficient estimation, blurry edges are enhanced in the depth map. At last, we have implemented ablation experiments to verify the effectiveness of the proposed DEN model. The results also demonstrate the superiority of DEN over some state-of-the-art approaches.
You-Feng Wu, Vu-Hoang Tran, Ting-Wei Chang, Walon Wei-Chen Chiu
ICPR4
2020 Learning Face Recognition Unsupervisedly by Disentanglement and Self-Augmentation
abstract
As the growth of smart home, healthcare, and home robot applications, learning a face recognition system which is specific for a particular environment and capable of self-adapting to the temporal changes in appearance (e.g., caused by illumination or camera position) is nowadays an important topic. In this paper, given a video of a group of people, which simulates the surveillance video in a smart home environment, we propose a novel approach which unsuper- visedly learns a face recognition model based on two main components: (1) a triplet network that extracts identity-aware feature from face images for performing face recognition by clustering, and (2) an augmentation network that is conditioned on the identity-aware features and aims at synthesizing more face samples. Particularly, the training data for the triplet network is obtained by using the spatiotemporal characteristic of face samples within a video, while the augmentation network learns to disentangle a face image into identity-aware and identity-irrelevant features thus is able to generate new faces of the same identity but with variance in appearance. With taking the richer training data produced by augmentation network, the triplet network is further fine-tuned and achieves better performance in face recognition. Extensive experiments not only show the efficacy of our model in learning an environment- specific face recognition model unsupervisedly, but also verify its adaptability to various appearance changes.
Yi-Lun Lee, Min-Yuan Tseng, Yu-Cheng Luo, Dung-Ru Yu, Walon Wei-Chen Chiu
ICRA5
2020 360SD-Net: 360° Stereo Depth Estimation with Learnable Cost Volume
abstract
Recently, end-to-end trainable deep neural networks have significantly improved stereo depth estimation for perspective images. However, 360° images captured under equirectangular projection cannot benefit from directly adopting existing methods due to distortion introduced (i.e., lines in 3D are not projected onto lines in 2D). To tackle this issue, we present a novel architecture specifically designed for spherical disparity using the setting of top-bottom 360° camera pairs. Moreover, we propose to mitigate the distortion issue by (1) an additional input branch capturing the position and relation of each pixel in the spherical coordinate, and (2) a cost volume built upon a learnable shifting filter. Due to the lack of 360° stereo data, we collect two 360° stereo datasets from Matterport3D and Stanford3D for training and evaluation. Extensive experiments and ablation study are provided to validate our method against existing algorithms. Finally, we show promising results on real-world environments capturing images with two consumer-level cameras. Our project page is at https://albert100121.github.io/360SD-Net-Project-Page.
Ning-Hsu Wang, Bolivar Solarte, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
ICRA4
2020 Self-Contained Stylization via Steganography for Reverse and Serial Style Transfer
abstract
Style transfer has been widely applied to give real-world images a new artistic look. However, given a stylized image, the attempts to use typical style transfer methods for de-stylization or transferring it again into another style usually lead to artifacts or undesired results. We realize that these issues are originated from the content inconsistency between the original image and its stylized output. Therefore, in this paper we advance to keep the content information of the input image during the process of style transfer by the power of steganography, with two approaches proposed: a two-stage model and an end-to-end model. We conduct extensive experiments to successfully verify the capacity of our models, in which both of them are able to not only generate stylized images of quality comparable with the ones produced by typical style transfer methods, but also effectively eliminate the artifacts introduced in reconstructing original input from a stylized image as well as performing multiple times of style transfer in series.
Hung-Yu Chen, I-Sheng Fang, Chia-Ming Cheng, Walon Wei-Chen Chiu
WACV4
2019 Guide Your Eyes: Learning Image Manipulation under Saliency Guidance
Yen-Chung Chen, Keng-Jui Chang, Yi-Hsuan Tsai, Yu-Chiang Frank Wang, Walon Wei-Chen Chiu
BMVC5
2019 All About Structure: Adapting Structural Information Across Domains for Boosting Semantic Segmentation
abstract
In this paper we tackle the problem of unsupervised domain adaptation for the task of semantic segmentation, where we attempt to transfer the knowledge learned upon synthetic datasets with ground-truth labels to real-world images without any annotation. With the hypothesis that the structural content of images is the most informative and decisive factor to semantic segmentation and can be readily shared across domains, we propose a Domain Invariant Structure Extraction (DISE) framework to disentangle images into domain-invariant structure and domain-specific texture representations, which can further realize image-translation across domains and enable label transfer to improve segmentation performance. Extensive experiments verify the effectiveness of our proposed DISE model and demonstrate its superiority over several state-of-the-art approaches.
Wei-Lun Chang, Hui-Po Wang, Wen-Hsiao Peng, Walon Wei-Chen Chiu
CVPR4
2019 Bridging Stereo Matching and Optical Flow via Spatiotemporal Correspondence
abstract
Stereo matching and flow estimation are two essential tasks for scene understanding, spatially in 3D and temporally in motion. Existing approaches have been focused on the unsupervised setting due to the limited resource to obtain the large-scale ground truth data. To construct a self-learnable objective, co-related tasks are often linked together to form a joint framework. However, the prior work usually utilizes independent networks for each task, thus not allowing to learn shared feature representations across models. In this paper, we propose a single and principled network to jointly learn spatiotemporal correspondence for stereo matching and flow estimation, with a newly designed geometric connection as the unsupervised signal for temporally adjacent stereo pairs. We show that our method performs favorably against several state-of-the-art baselines for both unsupervised depth and flow estimation on the KITTI benchmark dataset.
Hsueh-Ying Lai, Yi-Hsuan Tsai, Walon Wei-Chen Chiu
CVPR3
2019 Learning Pose-aware 3D Reconstruction via 2D-3D Self-consistency
abstract
3D reconstruction, inferring 3D shape information from a single 2D image, has drawn attention from learning and vision communities. In this paper, we propose a framework for learning pose-aware 3D shape reconstruction. Our proposed model learns deep representation for recovering the 3D object, with the ability to extract camera pose information but without any direct supervision of ground truth camera pose. This is realized by exploitation of 2D-3D self-consistency between 2D masks and 3D voxels. Experiments qualitatively and quantitatively demonstrate the effectiveness and robustness of our model, which performs favorably against state-of-the-art methods.
Yi-Lun Liao, Yao-Cheng Yang, Yuan-Fang Lin, Pin-Jung Chen, Chia-Wen Kuo, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang
ICASSP6
2019 Plug-and-Play: Improve Depth Prediction via Sparse Data Propagation
abstract
We propose a novel plug-and-play (PnP) module for improving depth prediction with taking arbitrary patterns of sparse depths as input. Given any pre-trained depth prediction model, our PnP module updates the intermediate feature map such that the model outputs new depths consistent with the given sparse depths. Our method requires no additional training and can be applied to practical applications such as leveraging both RGB and sparse LiDAR points to robustly estimate dense depth map. Our approach achieves consistent improvements on various state-of-the-art methods on indoor (i.e., NYU-v2) and outdoor (i.e., KITTI) datasets. Various types of LiDARs are also synthesized in our experiments to verify the general applicability of our PnP module in practice.
Tsun-Hsuan Wang, Fu-En Wang, Juan-Ting Lin, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
ICRA5
2019 3D LiDAR and Stereo Fusion using Stereo Matching Network with Conditional Cost Volume Normalization
abstract
The complementary characteristics of active and passive depth sensing techniques motivate the fusion of the LiDAR sensor and stereo camera for improved depth perception. Instead of directly fusing estimated depths across LiDAR and stereo modalities, we take advantages of the stereo matching network with two enhanced techniques: Input Fusion and Conditional Cost Volume Normalization (CCVNorm) on the LiDAR information. The proposed framework is generic and closely integrated with the cost volume component that is commonly utilized in stereo matching neural networks. We experimentally verify the efficacy and robustness of our method on the KITTI Stereo and Depth Completion datasets, obtaining favorable performance against various fusion strategies. Moreover, we demonstrate that, with a hierarchical extension of CCVNorm, the proposed method brings only slight overhead to the stereo matching network in terms of computation time and model size.
Tsun-Hsuan Wang, Hou-Ning Hu, Chieh Hubert Lin, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
IROS5
2018 Detach and Adapt: Learning Cross-Domain Disentangled Deep Representation
abstract
While representation learning aims to derive interpretable features for describing visual data, representation disentanglement further results in such features so that particular image attributes can be identified and manipulated. However, one cannot easily address this task without observing ground truth annotation for the training data. To address this problem, we propose a novel deep learning model of Cross-Domain Representation Disentangler (CDRD). By observing fully annotated source-domain data and unlabeled target-domain data of interest, our model bridges the information across data domains and transfers the attribute information accordingly. Thus, cross-domain feature disentanglement and adaptation can be jointly performed. In the experiments, we provide qualitative results to verify our disentanglement capability. Moreover, we further confirm that our model can be applied for solving classification tasks of unsupervised domain adaptation, and performs favorably against state-of-the-art image disentanglement and translation methods.
Yen-Cheng Liu, Yu-Ying Yeh, Tzu-Chien Fu, Sheng-De Wang, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang
CVPR5
2018 Summarizing First-Person Videos from Third Persons' Points of Views
Hsuan-I Ho, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang
ECCV (15)2
2017 STD2P: RGBD Semantic Segmentation Using Spatio-Temporal Data-Driven Pooling
abstract
We propose a novel superpixel-based multi-view convolutional neural network for semantic image segmentation. The proposed network produces a high quality segmentation of a single image by leveraging information from additional views of the same scene. Particularly in indoor videos such as captured by robotic platforms or handheld and bodyworn RGBD cameras, nearby video frames provide diverse viewpoints and additional context of objects and scenes. To leverage such information, we first compute region correspondences by optical flow and image boundary-based superpixels. Given these region correspondences, we propose a novel spatio-temporal pooling layer to aggregate information over space and time. We evaluate our approach on the NYU-Depth-V2 and the SUN3D datasets and compare it to various state-of-the-art single-view and multi-view approaches. Besides a general improvement over the state-of-the-art, we also show the benefits of making use of unlabeled frames during training for multi-view as well as single-view prediction.
Yang He 0005, Walon Wei-Chen Chiu, Margret Keuper, Mario Fritz
CVPR2
2016 Towards Segmenting Consumer Stereo Videos: Benchmark, Baselines and Ensembles
Walon Wei-Chen Chiu, Fabio Galasso, Mario Fritz
ACCV (5)1
2015 See the Difference: Direct Pre-Image Reconstruction and Pose Estimation by Differentiating HOG
abstract
The Histogram of Oriented Gradient (HOG) descriptor has led to many advances in computer vision over the last decade and is still part of many state of the art approaches. We realize that the associated feature computation is piecewise differentiable and therefore many pipelines which build on HOG can be made differentiable. This lends to advanced introspection as well as opportunities for end-to-end optimization. We present our implementation of ΔHOG based on the auto-differentiation toolbox Chumpy [18] and show applications to pre-image visualization and pose estimation which extends the existing differentiable renderer OpenDR [19] pipeline. Both applications improve on the respective state-of-the-art HOG approaches.
Walon Wei-Chen Chiu, Mario Fritz
ICCV1
2015 Joint segmentation and activity discovery using semantic and temporal priors
abstract
We introduce a hierarchical nonparametric topic modeling approach to infer activity routines from context sensor data streams based on a distance dependent Chinese restaurant process (ddCRP). Our approach does not require labeled data at any stage. Neither does our approach depend on time-invariant sliding windows to sample context word statistics. Our activity discovery approach builds on the idea that context words occurring within one activity are semantically similar, whereas context words of different activities are less similar. Context word streams are segmented into supersamples and then semantic and temporal features are obtained to construct a segmentation prior that relates supersamples via its context words. Our hierarchical model uses the segmentation prior and ddCRP to group supersamples and the Chinese restaurant process (CRP) to discover activities. We evaluate our approach using the Opportunity dataset that contains activities of daily living. Besides being nonparametric, our ddCRP based model outperforms both, classic parametric latent Dirichlet allocation (LDA) and the nonparametric Chinese restaurant franchise (CRF). We conclude that ddCRP+CRP is an adequate approach for fully unsupervised activity discovery from context sensor data.
Julia Seiter 0001, Walon Wei-Chen Chiu, Mario Fritz, Oliver Amft, Gerhard Tröster
PerCom2
2014 Object Disambiguation for Augmented Reality Applications
Walon Wei-Chen Chiu, Gregory Johnson, Daniel McCulley, Oliver Grau, Mario Fritz
BMVC1
2013 Multi-class Video Co-segmentation with a Generative Multi-video Model
abstract
Video data provides a rich source of information that is available to us today in large quantities e.g. from on-line resources. Tasks like segmentation benefit greatly from the analysis of spatio-temporal motion patterns in videos and recent advances in video segmentation has shown great progress in exploiting these addition cues. However, observing a single video is often not enough to predict meaningful segmentations and inference across videos becomes necessary in order to predict segmentations that are consistent with objects classes. Therefore the task of video co-segmentation is being proposed, that aims at inferring segmentation from multiple videos. But current approaches are limited to only considering binary foreground/background segmentation and multiple videos of the same object. This is a clear mismatch to the challenges that we are facing with videos from online resources or consumer videos. We propose to study multi-class video co-segmentation where the number of object classes is unknown as well as the number of instances in each frame and video. We achieve this by formulating a non-parametric Bayesian model across videos sequences that is based on a new videos segmentation prior as well as a global appearance model that links segments of the same class. We present the first multi-class video co-segmentation evaluation. We show that our method is applicable to real video data from online resources and outperforms state-of-the-art video segmentation and image co-segmentation baselines.
Walon Wei-Chen Chiu, Mario Fritz
CVPR1
2011 Improving the Kinect by Cross-Modal Stereo
Walon Wei-Chen Chiu, Ulf Blanke, Mario Fritz
BMVC1
2010 Probabilistic Modeling of Dynamic Traffic Flow across Non-overlapping Camera Views
abstract
In this paper, we propose a probabilistic method to model the dynamic traffic flow across non-overlapping camera views. By assuming the transition time of object movement follows a certain global model, we may infer the time-varying traffic status in the unseen region without performing explicit object correspondence between camera views. In this paper, we model object correspondence and parameter estimation as a unified problem under the proposed Expectation-Maximization (EM) based framework. By treating object correspondence as a latent random variable, the proposed framework can iteratively search for the optimal model parameters with the implicit consideration of object correspondence.
Chingchun Huang, Walon Wei-Chen Chiu, Sheng-Jyh Wang, Jen-Hui Chuang
ICPR2
2007 Robust Parking Space Detection Considering Inter-Space Correlation
abstract
A major problem in metropolitan areas is searching for parking spaces. In this paper, we propose a novel method for parking space detection. Given input video captured by a camera, we can distinguish the empty spaces from the occupied spaces by using an 8-class support vector machine (SVM) classifier with probabilistic outputs. Considering the inter-space correlation, the outputs of the SVM classifier are fused together using a Markov random field (MRF) framework. The result is much improved detection performance, even when there are significant occlusion and shadowing effects in the scene. Experimental results are given to show the robustness of the proposed approach.
Qi Wu 0014, Chingchun Huang, Shih-yu Wang, Walon Wei-Chen Chiu, Tsuhan Chen
ICME4