EDBT 2026 Demo / reviewers in the wild / expert
Yu-Chiang Frank Wang
dblp:30/1690 · also Yu-Chiang Wang 0001
· DBLP profile ↗
202ranked-venue papers
6as first author
72since 2021 · last 2026
0000-0002-2333-157XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 167 · 55 since 2021Artificial intelligence and machine learning · 99 · 6 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni PerceptionabstractZhen Wan, Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye, Ankita Pasad, Szu-Wei Fu, Arushi Goel, Ryo Hachiuma, Shizhe Diao, Kunal Dhawan, Sreyan Ghosh, Yusuke Hirota, Zhehuai Chen, Rafael Valle, Chenhui Chu, Shinji Watanabe, Boris Ginsburg, Yu-Chiang Frank Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye, Ankita Pasad, Szu-Wei Fu, Arushi Goel, Ryo Hachiuma, Shizhe Diao, Kunal Dhawan, Sreyan Ghosh, Yusuke Hirota, Zhehuai Chen, Rafael Valle, Chenhui Chu, Shinji Watanabe 0001, Boris Ginsburg, Yu-Chiang Frank Wang |
ACL (1) | 18 |
| 2026 | Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive AlignmentabstractRecent advancement in multimodal LLMs (MLLMs) has demonstrated their remarkable capability to generate descriptive captions for input videos. However, these models suffer from factual inaccuracies in the generated descriptions, causing severe hallucination issues. While prior works have explored alleviating hallucinations for static images, jointly mitigating visual object and temporal action hallucinations for dynamic videos remains a challenging and unsolved task. To tackle this challenge, we propose a Self-Augmented Contrastive Alignment (SANTA) framework for enabling object and action faithfulness by exempting the spurious correlations and enforcing the emphasis on visual facts. SANTA employs a hallucinative self-augmentation scheme to identify the potential hallucinations that lie in the MLLM and transform the original captions to the contrasted negatives. Furthermore, we develop a tracklet-phrase contrastive alignment to match the regional objects and relation-guided actions with their corresponding visual and temporal phrases. Extensive experiments demonstrate that SANTA outperforms existing methods in alleviating object and action hallucinations, yielding superior performance on the hallucination examination benchmarks. Kai-Po Chang, Wei-Yuan Cheng, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang |
WACV | 5 |
| 2026 | TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal AnchorsabstractDense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data. However, existing VideoLLMs remain challenging in identifying precise event boundaries in untrimmed videos, causing the generated captions to be not properly grounded. In this paper, we propose TA-Prompting, which enhances VideoLLMs via Temporal Anchors that learn to precisely localize events and prompt the VideoLLMs to perform temporal-aware video event understanding. During inference, in order to properly determine the output caption sequence from an arbitrary number of events presented within a video, we introduce an event coherent sampling strategy to select event captions with sufficient coherence across temporal events and cross-modal similarity with the given video. Through extensive experiments on benchmark datasets, we show that our TA-Prompting is favorable against state-of-the-art VideoLLMs, yielding superior performance on dense video captioning and temporal understanding tasks including moment retrieval and temporalQA. Wei-Yuan Cheng, Kai-Po Chang, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang |
WACV | 5 |
| 2025 | Serial Lifelong Editing via Mixture of Knowledge ExpertsabstractIt is challenging to update Large language models (LLMs) since real-world knowledge evolves.While existing Lifelong Knowledge Editing (LKE) methods efficiently update sequentially incoming edits, they often struggle to precisely overwrite the outdated knowledge with the latest one, resulting in conflicts that hinder LLMs from determining the correct answer.To address this Serial Lifelong Knowledge Editing (sLKE) problem, we propose a novel Mixture-of-Knowledge-Experts scheme with an Activation-guided Routing Mechanism (ARM), which assigns specialized experts to store domain-specific knowledge and ensures that each update completely overwrites old information with the latest data.Furthermore, we introduce a novel sLKE benchmark where answers to the same concept are updated repeatedly, to assess the ability of editing methods to refresh knowledge accurately.Experimental results on both LKE and sLKE benchmarks show that our ARM performs favorably against SOTA knowledge editing methods. YuJu Cheng, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank Wang |
ACL (1) | 4 |
| 2025 | Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field RenderingabstractWe propose an efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. There are two key contributions coupled with the proposed system. The first is to adaptively and explicitly allocate sparse voxels to different levels of detail within scenes, faithfully reproducing scene details with 655363grid resolution while achieving high rendering frame rates. Second, we customize a rasterizer for efficient adaptive sparse voxels rendering. We render voxels in the correct depth order by using ray direction-dependent Morton ordering, which avoids the well-known popping artifact found in Gaussian splat- ting. Our method improves the previous neural-free voxel model by over 4db PSNR and more than 10x FPS speedup, achieving state-of-the-art comparable novel-view synthesis results. Additionally, our voxel representation is seamlessly compatible with grid-based 3D processing techniques such as Volume Fusion, Voxel Pooling, and Marching Cubes, enabling a wide range of future extensions and applications. Code: github.com/NVlabs/svraster Cheng Sun 0004, Jaesung Choe, Charles Loop, Wei-Chiu Ma, Yu-Chiang Frank Wang |
CVPR | 5 |
| 2025 | Omni-RGPT: Unifying Image and Video Region-level Understanding via Token MarksabstractWe present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we introduce Token Mark, a set of tokens highlighting the target regions within the visual feature space. These tokens are directly embedded into spatial regions using region prompts (e.g., boxes or masks) and simultaneously incorporated into the text prompt to specify the target, establishing a direct connection between visual and text tokens. To further support robust video understanding without requiring tracklets, we introduce an auxiliary task that guides Token Mark by leveraging the consistency of the tokens, enabling stable region interpretation across the video. Additionally, we introduce a large-scale region-level video instruction dataset (RegVID300k). Omni-RGPT achieves state-of-the-art results on image and video-based commonsense reasoning benchmarks while showing strong performance in captioning and referring expression comprehension tasks. Miran Heo, Min-Hung Chen, De-An Huang, Sifei Liu, Subhashree Radhakrishnan, Seon Joo Kim, Yu-Chiang Frank Wang, Ryo Hachiuma |
CVPR | 7 |
| 2025 | 3D Gaussian Inpainting with Depth-Guided Cross-View ConsistencyabstractWhen performing 3D inpainting using novel-view rendering methods like Neural Radiance Field (NeRF) or 3D Gaussian Splatting (3DGS), how to achieve texture and geometry consistency across camera views has been a challenge. In this paper, we propose a framework of 3D Gaussian Inpainting with Depth-Guided Cross-View Consistency (3DGIC) for cross-view consistent 3D inpainting. Guided by the rendered depth information from each training view, our 3DGIC exploits background pixels visible across different views for updating the inpainting mask, allowing us to refine the 3DGS for inpainting purposes. Through extensive experiments on benchmark datasets, we confirm that our 3DGIC outperforms current state-of-the-art 3D inpainting methods quantitatively and qualitatively. Sheng-Yu Huang, Zi-Ting Chou, Yu-Chiang Frank Wang |
CVPR | 3 |
| 2025 | VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsabstractCustomized text-to-video generation aims to produce high-quality videos that incorporate user-specified subject identities or motion patterns. However, existing methods mainly focus on personalizing a single concept, either subject identity or motion pattern, limiting their effectiveness for multiple subjects with the desired motion patterns. To tackle this challenge, we propose a unified framework Video-Mage for video customization over both multiple subjects and their interactive motions. VideoMage employs subject and motion LoRAs to capture personalized content from user-provided images and videos, along with an appearance-agnostic motion learning approach to disentangle motion patterns from visual appearance. Furthermore, we develop a spatial-temporal composition scheme to guide interactions among subjects within the desired motion patterns. Extensive experiments demonstrate that VideoMage outperforms existing methods, generating coherent, user-controlled videos with consistent subject identities and interactions. Project Page: https://jasper0314-huang.github.io/videomage-customization/ Chi-Pin Huang, Yen-Siang Wu, Hung-Kai Chung, Kai-Po Chang, Fu-En Yang, Yu-Chiang Frank Wang |
CVPR | 6 |
| 2025 | Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding RegistrationabstractWe introduce Dr. Splat, a novel approach for open-vocabulary 3D scene understanding leveraging 3D Gaussian Splatting. Unlike existing language-embedded 3DGS methods, which rely on a rendering process, our method directly associates language-aligned CLIP embeddings with 3D Gaussians for holistic 3D scene understanding. The key of our method is a language feature registration technique where CLIP embeddings are assigned to the dominant Gaussians intersected by each pixel-ray. Moreover, we integrate Product Quantization (PQ) trained on general large-scale image data to compactly represent embeddings without per-scene optimization. Experiments demonstrate that our approach significantly outperforms existing approaches in 3D perception benchmarks, such as openvocabulary 3D semantic segmentation, 3D object localization, and 3D object selection tasks. For video results, please visit : https://drsplat.github.io/ Kim Jun-Seong, GeonU Kim, Kim Yu-Ji, Yu-Chiang Frank Wang, Jaesung Choe, Tae-Hyun Oh |
CVPR | 4 |
| 2025 | UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingabstractAudio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost of annotating training data with the class labels of all these events, along with their start and end times, imposes constraints on the scalability of AVVP techniques unless they can be trained in a weakly-supervised setting, where only modality-agnostic, video-level labels are available in the training data. To this end, recently proposed approaches seek to generate segment-level pseudo-labels to better guide model training. However, the absence of inter-segment dependencies when generating these pseudo-labels and the general bias towards predicting labels that are absent in a segment limit their performance. This work proposes a novel approach towards overcoming these weaknesses called Uncertainty-Weighted Weakly-Supervised Audio-Visual Video Parsing (UWAV). Additionally, our innovative approach factors in the uncertainty associated with these estimated pseudo-labels and incorporates a feature mixup based training regularization for improved training. Empirical results show that UWAV outperforms state-of-the-art methods for the AVVP task on multiple metrics, across two different datasets, attesting to its effectiveness and generalizability.1 Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain, Michael J. Jones 0001, Moitreya Chatterjee |
CVPR | 3 |
| 2025 | VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language ModelsabstractThe recent surge in high-quality visual instruction tuning samples from closed-source vision-language models (VLMs) such as GPT-4V has accelerated the release of open-source VLMs across various model sizes. However, scaling VLMs to improve performance using larger models brings significant computational challenges, especially for deployment on resource-constrained devices like mobile platforms and robots. To address this, we propose VLsI: Verbalized Layers-to-Interactions, a new VLM family in 2B and 7B model sizes, which prioritizes efficiency without compromising accuracy. VLsI leverages a unique, layer-wise distillation process, introducing intermediate "verbalizers" that map features from each layer to natural language space, allowing smaller VLMs to flexibly align with the reasoning processes of larger VLMs. This approach mitigates the training instability often encountered in output imitation and goes beyond typical final-layer tuning by aligning the small VLMs’ layer-wise progression with that of the large ones. We validate VLsI across ten challenging vision-language benchmarks, achieving notable performance gains (11.0% for 2B and 17.4% for 7B) over GPT-4V without the need for model scaling, merging, or architectural changes. Project Page. Ryo Hachiuma, Yu-Chiang Frank Wang, Yong Man Ro, Yueh-Hua Wu |
CVPR | 3 |
| 2025 | Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationabstractWe tackle open-vocabulary 3D scene segmentation tasks by introducing a novel data generation pipeline and training framework. Our work targets three essential aspects required for an effective dataset: precise 3D region segmentation, comprehensive textual descriptions, and sufficient dataset scale. By leveraging state-of-the-art open-vocabulary image segmentation models and region-aware vision-language models (VLM), we develop an automatic pipeline capable of producing high-quality 3D mask-text pairs. Applying this pipeline to multiple 3D scene datasets, we create Mosaic3D-5.6M, a dataset of more than 30K annotated scenes with 5.6M mask-text pairs - significantly larger than existing datasets. Building on these data, we propose Mosaic3D, a 3D visiual foundation model (3D-VFM) combining a 3D encoder trained with contrastive learning and a lightweight mask decoder for open-vocabulary 3D semantic and instance segmentation. Our approach achieves state-of-the-art results on open-vocabulary 3D semantic and instance segmentation benchmarks including ScanNet200, Matterport3D, and ScanNet++, with ablation studies validating the effectiveness of our large-scale training data. https://nvlabs.github.io/Mosaic3D/ Junha Lee, Chunghyun Park, Jaesung Choe, Yu-Chiang Frank Wang, Jan Kautz, Minsu Cho, Christopher B. Choy |
CVPR | 4 |
| 2025 | Segment Anything, Even OccludedabstractAmodal instance segmentation, which aims to detect and segment both visible and invisible parts of objects in images, plays a crucial role in various applications, including autonomous driving, robotic manipulation, and scene understanding. While existing methods require training both front-end detectors and mask decoders jointly, this approach lacks flexibility and fails to leverage the strengths of pre-existing modal detectors. To address this limitation, we propose SAMEO, a novel framework that adapts the Segment Anything Model (SAM) as a versatile mask decoder capable of interfacing with various front-end detectors to enable mask prediction even for partially occluded objects. Acknowledging the constraints of limited amodal segmentation datasets, we introduce Amodal-LVIS, a large-scale synthetic dataset comprising 300K images derived from the modal LVIS and LVVIS datasets. This dataset significantly expands the training data available for amodal segmentation research. Our experimental results demonstrate that our approach, when trained on the newly extended dataset, including Amodal-LVIS, achieves remarkable zero-shot performance on both COCOA-cls and D2SA benchmarks, highlighting its potential for generalization to unseen scenarios. Wei-En Tai, Yu-Lin Shih, Cheng Sun 0004, Yu-Chiang Frank Wang, Hwann-Tzong Chen |
CVPR | 4 |
| 2025 | Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning DataabstractRecent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires significant annotation efforts and risks catastrophic forgetting of the original language capabilities. In this work, we present a simple yet effective automatic process for creating speech-text pair data that carefully injects speech paralinguistic understanding abilities into SLMs while preserving the inherent language capabilities of the text-based LLM. Our model demonstrates general capabilities for speech-related tasks without the need for speech instruction-tuning data, achieving impressive performance on Dynamic-SUPERB and AIR-Bench-Chat benchmarks. Furthermore, our model exhibits the ability to follow complex instructions derived from LLMs, such as specific output formatting and chain-of-thought reasoning. Our approach not only enhances the versatility and effectiveness of SLMs but also reduces reliance on extensive annotated datasets, paving the way for more efficient and capable speech understanding systems.1 Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, Chao-Han Huck Yang, Jagadeesh Balam, Boris Ginsburg, Yu-Chiang Frank Wang, Hung-yi Lee |
ICASSP | 7 |
| 2025 | Bias in Gender Bias Benchmarks: How Spurious Features Distort EvaluationabstractGender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such as objects and backgrounds, we identify a critical oversight in gender bias evaluation: Do spurious features distort gender bias evaluation? To address this question, we systematically perturb non-gender features across four widely used benchmarks (COCO-gender, FACET, MIAP, and PHASE) and various VLMs to quantify their impact on bias evaluation. Our findings reveal that even minimal perturbations, such as masking just 10% of objects or weakly blurring backgrounds, can dramatically alter bias scores, shifting metrics by up to 175% in generative VLMs and 43% in CLIP variants. This suggests that current bias evaluations often reflect model responses to spurious features rather than gender bias, undermining their reliability. Since creating spurious feature-free benchmarks is fundamentally challenging, we recommend reporting bias metrics alongside feature-sensitivity measurements to enable a more reliable bias assessment. Yusuke Hirota, Ryo Hachiuma, Boyi Li 0001, Ximing Lu, Michael Ross Boone, Boris Ivanovic, Yejin Choi 0001, Marco Pavone 0001, Yu-Chiang Frank Wang, Noa Garcia, Yuta Nakashima, Chao-Han Huck Yang |
ICCV | 9 |
| 2025 | Continual Personalization for Diffusion ModelsabstractUpdating diffusion models in an incremental setting would be practical in real-world applications yet computationally challenging. We present a novel learning strategy of Concept Neuron Selection (CNS), a simple yet effective approach to perform personalization in a continual learning scheme. CNS uniquely identifies neurons in diffusion models that are closely related to the target concepts. In order to mitigate catastrophic forgetting problems while preserving zero-shot text-to-image generation ability, CNS finetunes concept neurons in an incremental manner and jointly preserves knowledge learned of previous concepts. Evaluation of real-world datasets demonstrates that CNS achieves state-of-the-art performance with minimal parameter adjustments, outperforming previous methods in both single and multi-concept personalization works. CNS also achieves fusion-free operation, reducing memory storage and processing time for continual personalization. Yu-Chien Liao, Jr-Jen Chen, Chi-Pin Huang, Ci-Siang Lin, Meng-Lin Wu, Yu-Chiang Frank Wang |
ICCV | 6 |
| 2025 | SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIPabstractLarge-scale vision-language models, such as CLIP, are known to contain societal bias regarding protected attributes (e.g., gender, age). This paper aims to address the problems of societal bias in CLIP. Although previous studies have proposed to debias societal bias through adversarial learning or test-time projecting, our comprehensive study of these works identifies two critical limitations: 1) loss of attribute information when it is explicitly disclosed in the input and 2) use of the attribute annotations during debiasing process. To mitigate societal bias in CLIP and overcome these limitations simultaneously, we introduce a simple-yet-effective debiasing method called SANER (societal attribute neutralizer) that eliminates attribute information from CLIP text features only of attribute-neutral descriptions. Experimental results show that SANER, which does not require attribute annotations and preserves original information for attribute-specific descriptions, demonstrates superior debiasing ability than the existing methods. Yusuke Hirota, Min-Hung Chen, Chien-Yi Wang, Yuta Nakashima, Yu-Chiang Frank Wang, Ryo Hachiuma |
ICLR | 5 |
| 2025 | UniWav: Towards Unified Pre-training for Speech Representation Learning and GenerationabstractPre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or generative tasks. In this work, we make the first attempt at building a unified pre-training framework for both types of tasks in speech. We show that with the appropriate design choices for pre-training, one can jointly learn a representation encoder and generative audio decoder that can be applied to both types of tasks. We propose UniWav, an encoder-decoder framework designed to unify pre-training representation learning and generative tasks. On speech recognition, text-to-speech, and speech tokenization, UniWav achieves comparable performance to different existing foundation models, each trained on a specific task. Our findings suggest that a single general-purpose foundation model for speech can be built to replace different foundation models, reducing the overhead and cost of pre-training. Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong 0001, Yu-Chiang Frank Wang, James R. Glass, Rafael Valle, Bryan Catanzaro |
ICLR | 5 |
| 2025 | Universal Speech Enhancement with Regression and Generative Mamba
Rong Chao, Rauf Nasretdinov, Yu-Chiang Frank Wang, Ante Jukic, Szu-Wei Fu, Yu Tsao 0001 |
INTERSPEECH | 3 |
| 2025 | VoiceNoNG: Robust High-Quality Speech Editing Model without Hallucinations
Sung-Feng Huang, Heng-Cheng Kuo, Zhehuai Chen, Xuesong Yang, Pin-Jui Ku, Ante Jukic, Chao-Han Huck Yang, Yu Tsao 0001, Yu-Chiang Frank Wang, Hung-yi Lee, Szu-Wei Fu |
INTERSPEECH | 9 |
| 2025 | ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningabstractVision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks. Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang, Fu-En Yang |
NeurIPS | 4 |
| 2025 | Unified Reinforcement and Imitation Learning for Vision-Language ModelsabstractVision-Language Models (VLMs) have achieved remarkable progress, yet their large scale often renders them impractical for resource-constrained environments. This paper introduces Unified Reinforcement and Imitation Learning (RIL), a novel and efficient training algorithm designed to create powerful, lightweight VLMs. RIL distinctively combines the strengths of reinforcement learning with adversarial imitation learning. This enables smaller student VLMs not only to mimic the sophisticated text generation of large teacher models but also to systematically improve their generative capabilities through reinforcement signals. Key to our imitation framework is a LLM-based discriminator that adeptly distinguishes between student and teacher outputs, complemented by guidance from multiple large teacher VLMs to ensure diverse learning. This unified learning strategy, leveraging both reinforcement and imitation, empowers student models to achieve significant performance gains, making them competitive with leading closed-source VLMs. Extensive experiments on diverse vision-language benchmarks demonstrate that RIL significantly narrows the performance gap with state-of-the-art open- and closed-source VLMs and, in several instances, surpasses them. Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu |
NeurIPS | 4 |
| 2025 | EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA CorrectionabstractOpen-source foundation models have seen rapid adoption and development, enabling powerful general-purpose capabilities across diverse domains. However, fine-tuning large foundation models for domain-specific or personalized tasks remains prohibitively expensive for most users due to the significant memory overhead beyond that of inference. We introduce EMLoC, an Emulator-based Memory-efficient fine-tuning framework with LoRA Correction, which enables model fine-tuning within the same memory budget required for inference. EMLoC constructs a task-specific light-weight emulator using activation-aware singular value decomposition (SVD) on a small downstream calibration set. Fine-tuning then is performed on this lightweight emulator via LoRA. To tackle the misalignment between the original model and the compressed emulator, we propose a novel compensation algorithm to correct the fine-tuned LoRA module, which thus can be merged into the original model for inference. EMLoC supports flexible compression ratios and standard training pipelines, making it adaptable to a wide range of applications. Extensive experiments demonstrate that EMLoC outperforms other baselines across multiple datasets and modalities. Moreover, without quantization, EMLoC enables fine-tuning of a 38B model, which originally required 95GB of memory, on a single 24GB consumer GPU—bringing efficient and practical model adaptation to individual users. Hsi-Che Lin, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank Wang |
NeurIPS | 4 |
| 2025 | Semantic Prompt Learning for Weakly-Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) aims to train segmentation models using image data with only image-level supervision. Since precise pixel-level annotations are not accessible, existing methods typically focus on producing pseudo masks for training segmentation models by refining CAM-like heatmaps. However, the produced heatmaps may capture only the discriminative image regions of object categories or the associated co-occurring backgrounds. To address the issues, we propose a Semantic Prompt Learning for WSSS (SemPLeS) framework, which learns to effectively prompt the CLIP latent space to enhance the semantic alignment between the segmented regions and the target object categories. More specifically, we propose Contrastive Prompt Learning and Prompt-guided Semantic Refinement to learn the prompts that adequately describe and suppress the co-occurring backgrounds associated with each object category. In this way, SemPLeS can perform better semantic alignment between object regions and class labels, resulting in desired pseudo masks for training segmentation models. The proposed SemPLeS framework achieves competitive performance on standard WSSS benchmarks, PASCAL VOC 2012 and MS COCO 2014, and shows compatibility with other WSSS methods. Project page: https://projectdisr.github.io/semples/ Ci-Siang Lin, Chien-Yi Wang, Yu-Chiang Frank Wang, Min-Hung Chen |
WACV | 3 |
| 2025 | Data-Efficient 3D Visual Grounding via Order-Aware Referringabstract3D visual grounding aims to identify the target object within a 3D point cloud scene referred to by a natural language description. Previous works usually require significant data relating to point color and their descriptions to exploit the corresponding complicated verbo-visual relations. In our work, we introduce Vigor, a novel Data-Efficient 3D Visual Grounding framework via Order-aware Referring. Vigor leverages LLM to produce a desirable referential order from the input description for 3D visual grounding. With the proposed stacked object-referring blocks, the predicted anchor objects in the above order allow one to locate the target object progressively with-out supervision on the identities of anchor objects or exact relations between anchor/target objects. We also present an order-aware warm-up training strategy, which augments referential orders for pre-training the visual grounding framework, allowing us to better capture the complex verbo-visual relations and benefit the desirable data-efficient learning scheme. Experimental results on the NR3D and ScanRefer datasets demonstrate our superiority in low-resource scenarios. In particular, Vigor surpasses current state-of-the-art frameworks by 9.3% and 7.6% grounding accuracy under 1% data and 10% data settings on the NR3D dataset, respectively. Tung-Yu Wu, Sheng-Yu Huang, Yu-Chiang Frank Wang |
WACV | 3 |
| 2025 | Learning Shape-Color Diffusion Priors for Text-Guided 3D Object GenerationabstractGenerating 3D shapes according to specific textual input is a crucial topic in the multimedia application, with its potential enhancement to the VR/AR/XR usage that enables more diverse virtual scenes. Due to the recent success of diffusion models, text-guided 3D object generation has drawn a lot of attention recently. However, current latent diffusion-based methods are restricted to shape-only generation, requiring time-consuming and computationally expensive post-processing to obtain colored objects. In this paper, we propose an end-to-endShape-Color Diffusion Prior framework (SCDiff)to achieve colored text-to-3D object generation. Given a general text description as input, our SCDiff is able to distinguish shape and color-related priors in the text and generate a shape latent and a color latent for a pre-trained 3D object auto-encoder to derive colored 3D objects. Our SCDiff contains two 3D latent diffusion models (LDM), where one generates the shape latent from the input text and the other generates the color latent. To help the two LDMs focus on shape/color-related information, we further adopt a Large Language Model (LLM) to separate the input text into a shape phrase and a color phrase via an in-context learning technique so that our shape/color LDM would not be influenced by irrelevant information. Due to the separation of shape and color latent, we are able to manipulate the color of an object by giving different color phrases while maintaining the original shape. Experiments on a benchmark dataset would quantitatively and qualitatively verify the effectiveness and practicality of our proposed model. As an extension, we show the capability of our SCDiff on 3D object generation and manipulation based on various modality conditions, which further confirms the scalability and applications in multimedia of our proposed framework. Sheng-Yu Huang, Chi-Pin Huang, Kai-Po Chang, Zi-Ting Chou, I-Jieh Liu, Yu-Chiang Frank Wang |
IEEE Trans. Multim. | 6 |
| 2024 | Language-Guided Transformer for Federated Multi-Label ClassificationabstractFederated Learning (FL) is an emerging paradigm that enables multiple users to collaboratively train a robust model in a privacy-preserving manner without sharing their private data. Most existing approaches of FL only consider traditional single-label image classification, ignoring the impact when transferring the task to multi-label image classification. Nevertheless, it is still challenging for FL to deal with user heterogeneity in their local data distribution in the real-world FL scenario, and this issue becomes even more severe in multi-label image classification. Inspired by the recent success of Transformers in centralized settings, we propose a novel FL framework for multi-label classification. Since partial label correlation may be observed by local clients during training, direct aggregation of locally updated models would not produce satisfactory performances. Thus, we propose a novel FL framework of Language-Guided Transformer (FedLGT) to tackle this challenging task, which aims to exploit and transfer knowledge across different clients for learning a robust global model. Through extensive experiments on various multi-label datasets (e.g., FLAIR, MS-COCO, etc.), we show that our FedLGT is able to achieve satisfactory performance and outperforms standard FL techniques under multi-label FL scenarios. Code is available at https://github.com/Jack24658735/FedLGT. I-Jieh Liu, Ci-Siang Lin, Fu-En Yang, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2024 | Seg2Reg: Differentiable 2D Segmentation to 1D Regression Rendering for 360 Room Layout ReconstructionabstractState-of-the-art single-view 360° room layout reconstruction methods formulate the problem as a high-level 1D (per-column) regression task. On the other hand, traditional low-level 2D layout segmentation is simpler to learn and can represent occluded regions, but it requires complex post-processing for the targeting layout polygon and sacrifices accuracy. We present Seg2Reg to render 1D layout depth regression from the 2D segmentation map in a differentiable and occlusion-aware way, marrying the merits of both sides. Specifically, our model predicts floor-plan density for the input equirectangular 360° image. Formulating the 2D layout representation as a density field enables us to employ ‘flattened’ volume rendering to form 1D layout depth regression. In addition, we propose a novel 3D warping augmentation on layout to improve generalization. Finally, we re-implement recent room layout reconstruction methods into our codebase for benchmarking and explore modern backbones and training techniques to serve as the strong baseline. The code is at https://PanoLayoutStudio.github.io. Cheng Sun 0004, Wei-En Tai, Yu-Lin Shih, Kuan-Wei Chen, Yong-Jing Syu, Kent Selwyn The, Yu-Chiang Frank Wang, Hwann-Tzong Chen |
CVPR | 7 |
| 2024 | GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene UnderstandingabstractUtilizing multi-view inputs to synthesize novel-view images, Neural Radiance Fields (NeRF) have emerged as a popular research topic in 3D vision. In this work, we introduce a Generalizable Semantic Neural Radiance Fields (GSNeRF), which uniquely takes image semantics into the synthesis process so that both novel view image and the associated semantic maps can be produced for unseen scenes. Our GSNeRF is composed of two stages: Semantic Geo-Reasoning and Depth-Guided Visual rendering. The former is able to observe multi-view image inputs to extract semantic and geometry features from a scene. Guided by the resulting image geometry information, the latter performs both image and semantic rendering with improved performances. Our experiments not only confirm that GSNeRF performs favorably against prior works on both novel-view image and semantic segmentation synthesis but the effectiveness of our sampling strategy for visual rendering is further verified. Zi-Ting Chou, Sheng-Yu Huang, I-Jieh Liu, Yu-Chiang Frank Wang |
CVPR | 4 |
| 2024 | SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
Yi-Chia Chen, Cheng Sun 0004, Yu-Chiang Frank Wang, Chu-Song Chen |
ECCV (81) | 4 |
| 2024 | Receler: Reliable Concept Erasing of Text-to-Image Diffusion Models via Lightweight Erasers
Chi-Pin Huang, Kai-Po Chang, Chung-Ting Tsai 0002, Yung-Hsuan Lai, Fu-En Yang, Yu-Chiang Frank Wang |
ECCV (40) | 6 |
| 2024 | TPA3D: Triplane Attention for Fast Text-to-3D Generation
Bin-Shih Wu, Hong-En Chen, Sheng-Yu Huang, Yu-Chiang Frank Wang |
ECCV (18) | 4 |
| 2024 | Select and Distill: Selective Dual-Teacher Knowledge Transfer for Continual Learning on Vision-Language Models
Yu-Chu Yu, Chi-Pin Huang, Jr-Jen Chen, Kai-Po Chang, Yung-Hsuan Lai, Fu-En Yang, Yu-Chiang Frank Wang |
ECCV (26) | 7 |
| 2024 | Enhancing Violin Fingering Generation through Audio-Symbolic FusionabstractThe selection of violin fingerings is influenced by factors such as musical context, skill level, and personal taste. Current deep-learningbased models, relying solely on symbolic data, are able to generate playable fingerings but struggle to capture the personal nuances of musical performance, which only lie in the audio data. To address this limitation, we introduce a novel model that incorporates both audio and symbolic data, allowing users to upload music scores and their corresponding violinist recordings to obtain personalized fingerings related to the audio data. To simulate such a real-world application scenario, we also collect a new dataset from online audios. The experiment results demonstrate the superiority of our proposed method over previous symbolic-based methods, even in the situations involving multiple instruments in audio. Wei-Yang Lin, Yu-Chiang Frank Wang |
ICASSP | 2 |
| 2024 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question AnsweringabstractNatural Language Explanation (NLE) in vision and language tasks aims to provide human-understandable explanations for the associated decision-making process. In practice, one might encounter explanations which lack informativeness or contradict visual-grounded facts, known as implausibility and hallucination problems, respectively. To tackle these challenging issues, we consider the task of visual question answering (VQA) and introduce Rapper, a two-stage Reinforced Rationale-Prompted Paradigm. By knowledge distillation, the former stage of Rapper infuses rationale-prompting via large language models (LLMs), encouraging the rationales supported by language-based facts. As for the latter stage, a unique Reinforcement Learning from NLE Feedback (RLNF) is introduced for injecting visual facts into NLE generation. Finally, quantitative and qualitative experiments on two VL-NLE benchmarks show that Rapper surpasses state-of-the-art VQA-NLE methods while providing plausible and faithful NLE. Kai-Po Chang, Chi-Pin Huang, Wei-Yuan Cheng, Fu-En Yang, Chien-Yi Wang, Yung-Hsuan Lai, Yu-Chiang Frank Wang |
ICLR | 7 |
| 2024 | Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean SpeechabstractSpeech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code and pre-trained models will be released Szu-Wei Fu, Kuo-Hsuan Hung, Yu Tsao 0001, Yu-Chiang Frank Wang |
ICLR | 4 |
| 2024 | DoRA: Weight-Decomposed Low-Rank AdaptationabstractAmong the widely used parameter-efficient fine-tuning (PEFT) methods, LoRA and its variants have gained considerable popularity because of avoiding additional inference costs. However, there still often exists an accuracy gap between these methods and full fine-tuning (FT). In this work, we first introduce a novel weight decomposition analysis to investigate the inherent differences between FT and LoRA. Aiming to resemble the learning capacity of FT from the findings, we propose Weight-Decomposed Low-Rank Adaptation (DoRA). DoRA decomposes the pre-trained weight into two components, magnitude and direction, for fine-tuning, specifically employing LoRA for directional updates to efficiently minimize the number of trainable parameters. By employing DoRA, we enhance both the learning capacity and training stability of LoRA while avoiding any additional inference overhead. DoRA consistently outperforms LoRA on fine-tuning LLaMA, LLaVA, and VL-BART on various downstream tasks, such as commonsense reasoning, visual instruction tuning, and image/video-text understanding. The code is available at https://github.com/NVlabs/DoRA. Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 0001, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen |
ICML | 5 |
| 2024 | DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, He Huang 0012, Boris Ginsburg, Yu-Chiang Frank Wang, Hung-yi Lee |
INTERSPEECH | 6 |
| 2024 | ReXTime: A Benchmark Suite for Reasoning-Across-Time in VideosabstractWe introduce ReXTime, a benchmark designed to rigorously test AI models' ability to perform temporal reasoning within video events.Specifically, ReXTime focuses on reasoning across time, i.e. human-like understanding when the question and its corresponding answer occur in different video segments. This form of reasoning, requiring advanced understanding of cause-and-effect relationships across video segments, poses significant challenges to even the frontier multimodal large language models. To facilitate this evaluation, we develop an automated pipeline for generating temporal reasoning question-answer pairs, significantly reducing the need for labor-intensive manual annotations. Our benchmark includes 921 carefully vetted validation samples and 2,143 test samples, each manually curated for accuracy and relevance. Evaluation results show that while frontier large language models outperform academic models, they still lag behind human performance by a significant 14.3\% accuracy gap. Additionally, our pipeline creates a training dataset of 9,695 machine generated samples without manual effort, which empirical studies suggest can enhance the across-time reasoning via fine-tuning. Jr-Jen Chen, Yu-Chien Liao, Hsi-Che Lin, Yu-Chu Yu, Yen-Chun Chen 0001, Yu-Chiang Frank Wang |
NeurIPS | 6 |
| 2024 | Diffusion-Reward Adversarial Imitation LearningabstractImitation learning aims to learn a policy from observing expert demonstrations without access to reward signals from environments. Generative adversarial imitation learning (GAIL) formulates imitation learning as adversarial learning, employing a generator policy learning to imitate expert behaviors and discriminator learning to distinguish the expert demonstrations from agent trajectories. Despite its encouraging results, GAIL training is often brittle and unstable. Inspired by the recent dominance of diffusion models in generative modeling, we propose Diffusion-Reward Adversarial Imitation Learning (DRAIL), which integrates a diffusion model into GAIL, aiming to yield more robust and smoother rewards for policy learning. Specifically, we propose a diffusion discriminative classifier to construct an enhanced discriminator, and design diffusion rewards based on the classifier’s output for policy learning. Extensive experiments are conducted in navigation, manipulation, and locomotion, verifying DRAIL’s effectiveness compared to prior imitation learning methods. Moreover, additional experimental results demonstrate the generalizability and data efficiency of DRAIL. Visualized learned reward functions of GAIL and DRAIL suggest that DRAIL can produce more robust and smoother rewards. Project page: https://nturobotlearninglab.github.io/DRAIL/ Chun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Yu-Chiang Frank Wang, Min-Hung Chen, Shao-Hua Sun |
NeurIPS | 4 |
| 2024 | Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech EditsabstractNeural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like $\mathrm{A}^{3} \mathrm{~T}$ and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made available at: https://jasonswfu.github.io/SINE_dataset/index.html Sung-Feng Huang, Heng-Cheng Kuo, Zhehuai Chen, Xuesong Yang, Chao-Han Huck Yang, Yu Tsao 0001, Yu-Chiang Frank Wang, Hung-yi Lee, Szu-Wei Fu |
SLT | 7 |
| 2023 | Frido: Feature Pyramid Diffusion for Complex Scene Image SynthesisabstractDiffusion models (DMs) have shown great potential for high-quality image synthesis. However, when it comes to producing images with complex scenes, how to properly describe both image global structures and object details remains a challenging task. In this paper, we present Frido, a Feature Pyramid Diffusion model performing a multi-scale coarse-to-fine denoising process for image synthesis. Our model decomposes an input image into scale-dependent vector quantized features, followed by a coarse-to-fine gating for producing image output. During the above multi-scale representation learning stage, additional input conditions like text, scene graph, or image layout can be further exploited. Thus, Frido can be also applied for conditional or cross-modality image synthesis. We conduct extensive experiments over various unconditioned and conditional image generation tasks, ranging from text-to-image synthesis, layout-to-image, scene-graph-to-image, to label-to-image. More specifically, we achieved state-of-the-art FID scores on five benchmarks, namely layout-to-image on COCO and OpenImages, scene-graph-to-image on COCO and Visual Genome, and label-to-image on COCO. Wan-Cyuan Fan, Yen-Chun Chen 0001, Dongdong Chen 0001, Yu Cheng 0001, Lu Yuan 0001, Yu-Chiang Frank Wang |
AAAI | 6 |
| 2023 | Target-Free Text-Guided Image ManipulationabstractWe tackle the problem of target-free text-guided image manipulation, which requires one to modify the input reference image based on the given text instruction, while no ground truth target image is observed during training. To address this challenging task, we propose a Cyclic-Manipulation GAN (cManiGAN) in this paper, which is able to realize where and how to edit the image regions of interest. Specifically, the image editor in cManiGAN learns to identify and complete the input image, while cross-modal interpreter and reasoner are deployed to verify the semantic correctness of the output image based on the input instruction. While the former utilizes factual/counterfactual description learning for authenticating the image semantics, the latter predicts the "undo" instruction and provides pixel-level supervision for the training of cManiGAN. With the above operational cycle-consistency, our cManiGAN can be trained in the above weakly supervised setting. We conduct extensive experiments on the datasets of CLEVR and COCO datasets, and the effectiveness and generalizability of our proposed method can be successfully verified. Project page: sites.google.com/view/wancyuanfan/projects/cmanigan. Wan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2023 | Bias-Eliminating Augmentation Learning for Debiased Federated LearningabstractLearning models trained on biased datasets tend to observe correlations between categorical and undesirable features, which result in degraded performances. Most existing debiased learning models are designed for centralized machine learning, which cannot be directly applied to distributed settings like federated learning (FL), which collects data at distinct clients with privacy preserved. To tackle the challenging task of debiased federated learning, we present a novel FL framework of Bias-Eliminating Augmentation Learning (FedBEAL), which learns to deploy Bias-Eliminating Augmenters (BEA) for producing client-specific bias-conflicting samples at each client. Since the bias types or attributes are not known in advance, a unique learning strategy is presented to jointly train BEA with the proposed FL framework. Extensive image classification experiments on datasets with various bias types confirm the effectiveness and applicability of our FedBEAL, which performs favorably against state-of-the-art debiasing and FL methods for debiased FL. Yuan-Yi Xu, Ci-Siang Lin, Yu-Chiang Frank Wang |
CVPR | 3 |
| 2023 | LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction FollowingabstractEnd-to-end Transformers have demonstrated an impressive success rate for Embodied Instruction Following when the environment has been seen in training.However, they tend to struggle when deployed in an unseen environment.This lack of generalizability is due to the agent's insensitivity to subtle changes in natural language instructions.To mitigate this issue, we propose explicitly aligning the agent's hidden states with the instructions via contrastive learning.Nevertheless, the semantic gap between high-level language instructions and the agent's low-level action space remains an obstacle.Therefore, we further introduce a novel concept of meta-actions to bridge the gap.Meta-actions are ubiquitous action patterns that can be parsed from the original action sequence.These patterns represent higher-level semantics that are intuitively aligned closer to the instructions.When meta-actions are applied as additional training signals, the agent generalizes better to unseen environments.Compared to a strong multi-modal Transformer baseline, we achieve a significant 4.5% absolute gain in success rate in unseen environments of ALFRED Embodied Instruction Following.Additional analysis shows that the contrastive objective and meta-actions are complementary in achieving the best results, and the resulting agent better aligns its states with corresponding instructions, making it more suitable for real-world embodied agents. 1 Cheng-Fu Yang, Yen-Chun Chen 0001, Xiyang Dai, Lu Yuan 0001, Yu-Chiang Frank Wang, Kai-Wei Chang 0001 |
EMNLP | 6 |
| 2023 | Semantics-Aware Gamma Correction for Unsupervised Low-Light Image EnhancementabstractLow-light image enhancement aims to improve the visual quality of images captured under poor lighting conditions. While recent works have successfully developed deep learning-based solutions, a large number of existing works require ground-truth normal-light images during training, and most methods are not designed to exploit and preserve semantic information in the low-light inputs. In this paper, we propose a semantics-aware yet unsupervised low-light enhancement model based on gamma correction. Without observing ground-truth images or semantic annotations of the low-light inputs, our model learns via the introduced semantics-aware adversarial learning scheme with the associated objectives given a set of unpaired reference images of interest. Guided by such high-quality reference images and the inherent semantic practicality, our proposed method performs favorably against recent unsupervised low-light enhancement approaches. Fu-Cheng Pan, Yu-Chien Liao, Jao-Hong Kao, Yu-Chiang Frank Wang |
ICASSP | 5 |
| 2023 | Efficient Model Personalization in Federated Learning via Client-Specific Prompt GenerationabstractFederated learning (FL) emerges as a decentralized learning framework which trains models from multiple distributed clients without sharing their data to preserve privacy. Recently, large-scale pre-trained models (e.g., Vision Transformer) have shown a strong capability of deriving robust representations. However, the data heterogeneity among clients, the limited computation resources, and the communication bandwidth restrict the deployment of large-scale models in FL frameworks. To leverage robust representations from large-scale models while enabling efficient model personalization for heterogeneous clients, we propose a novel personalized FL framework of client-specific Prompt Generation (pFedPG), which learns to deploy a personalized prompt generator at the server for producing client-specific visual prompts that efficiently adapts frozen backbones to local data distributions. Our proposed framework jointly optimizes the stages of personalized prompt adaptation locally and personalized prompt generation globally. The former aims to train visual prompts that adapt foundation models to each client, while the latter observes local optimization directions to generate personalized prompts for all clients. Through extensive experiments on benchmark datasets, we show that our pFedPG is favorable against state-of-the-art personalized FL methods under various types of data heterogeneity, allowing computation and communication efficient model personalization. Fu-En Yang, Chien-Yi Wang, Yu-Chiang Frank Wang |
ICCV | 3 |
| 2023 | Interpreting Latent Representation in Neural Radiance Fields for Manipulating Object SemanticsabstractManipulating 3D objects has been among the active research topic for 3D vision. With the development and success of neural radiance field (NeRF) [1] on scene modeling, synthesizing and manipulating 3D objects using such a representation becomes desirable. In this paper, we introduce a semantic-aware generative NeRF, which is able to interpret the latent representation learned by category-specific generative NeRFs and to achieve editing of particular part attributes. With pretrained generative NeRF, we propose to deploy a semantic segmentor for performing part segmentation on the object category. This allows the rendering of the 2D image and prediction of the corresponding segmentation mask. Our proposed scheme learns to manipulate the resulting latent representation, optimized to edit the object part of interest with varying degrees. We conduct experiments on various object categories on benchmark datasets, and the results successfully verify the effectiveness and practicality of our proposed model. Yu-Shan Huang, Sheng-Yu Huang, Hao-Yu Hsu, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2023 | Consistent and Multi-Scale Scene Graph Transformer for Semantic-Guided Image OutpaintingabstractThe task of image outpainting extends an image beyond its boundaries with semantically plausible content. Recently, Scene Graph Transformer (SGT) introduced a transformer architecture to leverage scene graph guidance for image outpainting. Despite its success, we identified two shortcomings: (a) SGT uses a positional encoding that was originally proposed for 1D signal; (b) SGT uses a scene graph attention layer that propagates information between neighboring nodes which limited the model to learning local graph features. To address these issues, we propose incorporating Laplacian positional encoding and introducing a multiscale scene graph attention into SGT. Extensive results on MS-COCO and Visual Genome show that our proposed approach generates more plausible outpainted images with higher quality. Chiao-An Yang, Meng-Lin Wu, Raymond A. Yeh, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2023 | Self-Supervised Pyramid Representation Learning for Multi-Label Visual Analysis and BeyondabstractWhile self-supervised learning has been shown to benefit a number of vision tasks, existing techniques mainly focus on image-level manipulation, which may not generalize well to downstream tasks at patch or pixel levels. Moreover, existing SSL methods might not sufficiently describe and associate the above representations within and across image scales. In this paper, we propose a Self-Supervised Pyramid Representation Learning (SS-PRL) framework. The proposed SS-PRL is designed to derive pyramid representations at patch levels via learning proper prototypes, with additional learners to observe and relate inherent semantic information within an image. In particular, we present a cross-scale patch-level correlation learning in SS-PRL, which allows the model to aggregate and associate information learned across patch scales. We show that, with our proposed SS-PRL for model pre-training, one can easily adapt and fine-tune the models for a variety of applications including multi-label classification, object detection, and instance segmentation. Cheng-Yen Hsieh, Chih-Jung Chang, Fu-En Yang, Yu-Chiang Frank Wang |
WACV | 4 |
| 2023 | Semantics-Guided Intra-Category Knowledge Transfer for Generalized Zero-Shot Learning
Fu-En Yang, Yuan-Hao Lee, Chia-Ching Lin, Yu-Chiang Frank Wang |
Int. J. Comput. Vis. | 4 |
| 2023 | Describe, Spot and Explain: Interpretable Representation Learning for Discriminative Visual ReasoningabstractDespite the recent success achieved by deep neural networks (DNNs), it remains challenging to disclose/explain the decision-making process from the numerous parameters and complex non-linear functions. To address the problem, explainable AI (XAI) aims to provide explanations corresponding to the learning and prediction processes for deep learning models. In this paper, we propose a novel representation learning framework of Describe, Spot and eXplain (DSX). Based on the architecture of Transformer, our proposed DSX framework is composed of two learning stages, descriptive prototype learning and discriminative prototype discovery. Given an input image, the former stage is designed to derive a set of descriptive representations, while the latter stage further identifies a discriminative subset, offering semantic interpretability for the corresponding classification tasks. While our DSX does not require any ground truth attribute supervision during training, the derived visual representations can be practically associated with physical attributes provided by domain experts. Extensive experiments on fine-grained classification and person re-identification tasks qualitatively and quantitatively verify the use our DSX model for offering semantically practical interpretability with satisfactory recognition performances. Ci-Siang Lin, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 2 |
| 2022 | Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and ManipulationabstractAs a key characteristic in audio-visual speech recognition (AVSR), relating linguistic information observed across visual and audio data has been a challenge, benefiting not only audio/visual speech recognition (ASR/VSR) but also for manipulating data within/across modalities. In this paper, we present a feature disentanglement-based framework for jointly addressing the above tasks. By advancing cross-modal mutual learning strategies, our model is able to convert visual or audio-based linguistic features into modality-agnostic representations. Such derived linguistic representations not only allow one to perform ASR, VSR, and AVSR, but also to manipulate audio and visual data output based on the desirable subject identity and linguistic content information. We perform extensive experiments on different recognition and synthesis tasks to show that our model performs favorably against state-of-the-art approaches on each individual task, while ours is a unified solution that is able to jointly tackle the aforementioned audio-visual learning tasks. Chih-Chun Yang 0004, Wan-Cyuan Fan, Cheng-Fu Yang, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2022 | NeurMiPs: Neural Mixture of Planar Experts for View SynthesisabstractWe present Neural Mixtures of Planar Experts (Neur-MiPs), a novel planar-based scene representation for modeling geometry and appearance. NeurMiPs leverages a collection of local planar experts in 3D space as the scene representation. Each planar expert consists of the parameters of the local rectangular shape representing geometry and a neural radiance field modeling the color and opacity. We render novel views by calculating ray-plane intersections and composite output colors and densities at intersected points to the image. NeurMiPs blends the efficiency of explicit mesh rendering and flexibility of the neural radiance field. Experiments demonstrate superior performance and speed of our proposed method, compared to other 3D representations in novel view synthesis. Zhi-Hao Lin, Wei-Chiu Ma, Hao-Yu Hsu, Yu-Chiang Frank Wang, Shenlong Wang |
CVPR | 4 |
| 2022 | Scene Graph Expansion for Semantics-Guided Image OutpaintingabstractIn this paper, we address the task of semantics-guided image outpainting, which is to complete an image by generating semantically practical content. Different from most existing image outpainting works, we approach the above task by understanding and completing image semantics at the scene graph level. In particular, we propose a novel network of Scene Graph Transformer (SGT), which is designed to take node and edge features as inputs for modeling the associated structural information. To better understand and process graph-based inputs, our SGT uniquely performs feature attention at both node and edge levels. While the former views edges as relationship regularization, the latter observes the co-occurrence of nodes for guiding the attention process. We demonstrate that, given a partial input image with its layout and scene graph, our SGT can be applied for scene graph expansion and its conversion to a complete layout. Following state-of-the-art layout-to-image conversions works, the task of image outpainting can be completed with sufficient and practical semantics introduced. Extensive experiments are conducted on the datasets of MS-COCO and Visual Genome, which quantitatively and qualitatively confirm the effectiveness of our proposed SGT and outpainting frameworks. Chiao-An Yang, Cheng-Yo Tan, Wan-Cyuan Fan, Cheng-Fu Yang, Meng-Lin Wu, Yu-Chiang Frank Wang |
CVPR | 6 |
| 2022 | Domain-Agnostic Meta-Learning for Cross-Domain Few-Shot ClassificationabstractFew-shot classification requires one to classify instances of novel classes, given only a few examples of each class. Although promising meta-learning methods have been proposed recently, there is no guarantee that existing solutions would generalize to novel classes from an unseen domain. In this paper, we tackle the challenging task of cross-domain few-shot classification and propose Domain-Agnostic Meta-Learning (DAML) algorithm. Our DAML, serving as an optimization strategy, learns to adapt the model to novel classes in both seen and unseen domains by data sampled from multiple domains with desirable task settings. In our experiments, we apply DAML on three popular metric-based models under cross-domain settings. Experiments on several benchmarks (mini-ImageNet, CUB, Cars, Places, Plantae and META-DATASET) show that DAML significantly improves the generalization ability of learning models, and addresses cross-domain few-shot classification with promising results. Wei-Yu Lee, Jheng-Yu Wang, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2022 | 3D-Selfcutmix: Self-Supervised Learning for 3D Point Cloud AnalysisabstractPoint clouds have been widely applied to represent 3D data, with a variety of applications such as autonomous driving, augmented reality, and robotics. Since collecting a large amount of labeled 3D point cloud data for training deep learning models might not always be applicable, we propose the novel learning strategy of 3D-SelfCutMix, which advances mixed-sample data augmentation techniques while exploiting the spatial and semantic consistencies between point cloud data. Depending on the availability of label supervision, the proposed network can be realized in either self-supervised or fully-supervised manners, while both versions are shown to benefit downstream tasks. In our experiments, we consider a variety of tasks including classification and part-segmentation tasks, which sufficiently support the use of the proposed method for 3D point cloud analysis. Yuan-Yi Xu, Yan-Yang Ji, Sheng-Yu Huang, Zhi-Hao Lin, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2022 | Domain-Generalized Textured Surface Anomaly DetectionabstractAnomaly detection aims to identify abnormal data that deviates from the normal ones, while typically requiring a sufficient amount of normal data to train the model for performing this task. Despite the success of recent anomaly detection methods, performing anomaly detection in an unseen domain remain a challenging task. In this paper, we address the task of domain-generalized textured surface anomaly detection. By observing normal and abnormal surface data across multiple source domains, our model is expected to be generalized to an unseen textured surface of interest, in which only a small number of normal data can be observed during testing. Although with only image-level labels observed in the training data, our patch-based meta-learning model exhibits promising generalization ability: not only can it generalize to unseen image domains, but it can also localize abnormal regions in the query image. Our experiments verify that our model performs favorably against state-of-the-art anomaly detection and domain generalization approaches in various settings. Shang-Fu Chen, Yu-Min Liu, Chia-Ching Lin, Trista Pei-Chun Chen, Yu-Chiang Frank Wang |
ICME | 5 |
| 2022 | Learning Facial Liveness Representation for Domain Generalized Face Anti-SpoofingabstractFace anti-spoofing (FAS) aims at distinguishing face spoof attacks from the authentic ones, which is typically approached by learning proper models for performing the associated classification task. In practice, one would expect such models to be generalized to FAS in different image domains. Moreover, it is not practical to assume that the type of spoof attacks would be known in advance. In this paper, we propose a deep learning model for addressing the aforementioned domain-generalized face anti-spoofing task. In particular, our proposed network is able to disentangle facial liveness representation from the irrelevant ones (i.e., facial content and image domain features). The resulting liveness representation exhibits sufficient domain invariant properties, and thus it can be applied for performing domain-generalized FAS. In our experiments, we conduct experiments on five benchmark datasets with various settings, and we verify that our model performs favorably against state-of-the-art approaches in identifying novel types of spoof attacks in unseen image domains. Zih-Ching Chen, Lin-Hsi Tsao, Chin-Lun Fu, Shang-Fu Chen, Yu-Chiang Frank Wang |
ICME | 5 |
| 2022 | SPoVT: Semantic-Prototype Variational Transformer for Dense Point Cloud Semantic CompletionabstractPoint cloud completion is an active research topic for 3D vision and has been widelystudied in recent years. Instead of directly predicting missing point cloud fromthe partial input, we introduce a Semantic-Prototype Variational Transformer(SPoVT) in this work, which takes both partial point cloud and their semanticlabels as the inputs for semantic point cloud object completion. By observingand attending at geometry and semantic information as input features, our SPoVTwould derive point cloud features and their semantic prototypes for completionpurposes. As a result, our SPoVT not only performs point cloud completion withvarying resolution, it also allows manipulation of different semantic parts of anobject. Experiments on benchmark datasets would quantitatively and qualitativelyverify the effectiveness and practicality of our proposed model. Sheng-Yu Huang, Hao-Yu Hsu, Yu-Chiang Frank Wang |
NeurIPS | 3 |
| 2022 | A Pixel-Level Meta-Learner for Weakly Supervised Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation addresses the learning task in which only few images with ground truth pixel-level labels are available for the novel classes of interest. One is typically required to collect a large mount of data (i.e., base classes) with such ground truth information, followed by meta-learning strategies to address the above learning task. When only image-level semantic labels can be observed during both training and testing, it is considered as an even more challenging task of weakly supervised few-shot semantic segmentation. To address this problem, we propose a novel meta-learning framework, which predicts pseudo pixel-level segmentation masks from a limited amount of data and their semantic labels. More importantly, our learning scheme further exploits the produced pixel-level information for query image inputs with segmentation guarantees. Thus, our proposed learning model can be viewed as a pixel-level meta-learner. Through extensive experiments on benchmark datasets, we show that our model achieves satisfactory performances under fully supervised settings, yet performs favorably against state-of-the-art methods under weakly supervised settings. Yuan-Hao Lee, Fu-En Yang, Yu-Chiang Frank Wang |
WACV | 3 |
| 2022 | Learning of 3D Graph Convolution Networks for Point Cloud AnalysisabstractPoint clouds are among the popular geometry representations in 3D vision. However, unlike 2D images with pixel-wise layouts, such representations containing unordered data points which make the processing and understanding the associated semantic information quite challenging. Although a number of previous works attempt to analyze point clouds and achieve promising performances, their performances would degrade significantly when data variations like shift and scale changes are presented. In this paper, we propose 3D graph convolution networks (3D-GCN), which uniquely learns 3D kernels with graph max-pooling mechanisms for extracting geometric features from point cloud data across different scales. We show that, with the proposed 3D-GCN, satisfactory shift and scale invariance can be jointly achieved. We show that 3D-GCN can be applied to point cloud classification and segmentation tasks, with ablation studies and visualizations verifying the design of 3D-GCN. Zhi-Hao Lin, Sheng-Yu Huang, Yu-Chiang Frank Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio GenerationabstractHuman perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which would degrade the user experience due to the lack of ambient information. To address this issue, we propose an audio spatialization framework to convert a monaural video into a binaural one exploiting the relationship across audio and visual components. By preserving the left-right consistency in both audio and visual modalities, our learning strategy can be viewed as a self-supervised learning technique, and alleviates the dependency on a large amount of video data with ground truth binaural audio data during training. Experiments on benchmark datasets confirm the effectiveness of our proposed framework in both semi-supervised and fully supervised scenarios, with ablation studies and visualization further support the use of our model for audio spatialization. Yan-Bo Lin, Yu-Chiang Frank Wang |
AAAI | 2 |
| 2021 | LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityabstractWhen translating text inputs into layouts or images, existing works typically require explicit descriptions of each object in a scene, including their spatial information or the associated relationships. To better exploit the text input, so that implicit objects or relationships can be properly inferred during layout generation, we propose a LayoutTransformer Network (LT-Net) in this paper. Given a scene-graph input, our LT-Net uniquely encodes the semantic features for exploiting their co-occurrences and implicit relationships. This allows one to manipulate conceptually diverse yet plausible layout outputs. Moreover, the decoder of our LT-Net translates the encoded contextual features into bounding boxes with self-supervised relation consistency preserved. By fitting their distributions to Gaussian mixture models, spatially-diverse layouts can be additionally produced by LT-Net. We conduct extensive experiments on the datasets of MS-COCO and Visual Genome, and confirm the effectiveness and plausibility of our LT-Net over recent layout generation models. Codes will be released at LayoutTransformer. Cheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank Wang |
CVPR | 4 |
| 2021 | Representation Decomposition For Image Manipulation And BeyondabstractRepresentation disentanglement aims at learning interpretable features, so that the output can be recovered or manipulated accordingly. While existing works like infoGAN [1] and ACGAN [2] exist, they choose to derive disjoint attribute code for feature disentanglement, which is not applicable for existing/trained generative models. In this paper, we propose a decomposition-GAN (dec-GAN), which is able to achieve the decomposition of an existing latent representation into content and attribute features. Guided by the classifier pre-trained on the attributes of interest, our dec-GAN decomposes the attributes of interest from the latent representation, while data recovery and feature consistency objectives enforce the learning of our proposed method. Our experiments on multiple image datasets confirm the effectiveness and robustness of our dec-GAN over recent representation disentanglement models. Shang-Fu Chen, Jia-Wei Yan, Ya-Fan Su, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2021 | Few-Shot Classification in Unseen Domains by Episodic Meta-Learning Across Visual DomainsabstractFew-shot classification aims to carry out classification given only few labeled examples for the categories of interest. Though several approaches have been proposed, most existing few-shot learning (FSL) models assume that base and novel classes are drawn from the same data domain. When it comes to recognizing novel-class data in an unseen domain, this becomes an even more challenging task of domain generalized few-shot classification. In this paper, we present a unique learning framework for domain-generalized few-shot classification, where base classes are from homogeneous multiple source domains, while novel classes to be recognized are from target domains which are not seen during training. By advancing meta-learning strategies, our learning framework exploits data across multiple source domains to capture domain-invariant features, with FSL ability introduced by metric-learning based mechanisms across support and query data. We conduct extensive experiments to verify the effectiveness of our proposed learning framework and show learning from small yet homogeneous source data is able to perform preferably against learning from large-scale one. Moreover, we provide insights into choices of backbone models for domain-generalized few-shot classification. Yuan-Chia Cheng, Ci-Siang Lin, Fu-En Yang, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2021 | Self-Supervised Bodymap-to-Appearance Co-Attention for Partial Person Re-IdentificationabstractPerson re-identification (re-ID) aims at recognizing the same person across distinct camera views. Although a series of works have been proposed to tackle re-ID, most existing methods are based on the assumption that full body detection is available. In practice, detection of pedestrians may not be perfect due to partial occlusion or background clutter, which would result in the challenging task of partial re-ID. To tackle this problem, we propose a novel deep learning framework which jointly performs image rescaling and bodymap-to-appearance co-attention, followed by image matching for re-ID purposes. The image rescaler presented in our framework learns to produce distortion-free images in a self-supervised manner, which allows the subsequent image matching based on the associated image regions via body part attention. Our quantitative and qualitative results on two benchmark partial re-ID datasets confirm the effectiveness of our approach and its superiority over state-of-the-art partial re-ID approaches. Ci-Siang Lin, Yu-Chiang Frank Wang |
ICIP | 2 |
| 2021 | Meta-Learned Feature Critics for Domain Generalized Semantic SegmentationabstractHow to handle domain shifts when recognizing or segmenting visual data across domains has been studied by learning and vision communities. In this paper, we address domain generalized semantic segmentation, in which the segmentation model is trained on multiple source domains and is expected to generalize to unseen data domains. We propose a novel meta-learning scheme with feature disentanglement ability, which derives domain-invariant features for semantic segmentation with domain generalization guarantees. In particular, we introduce a class-specific feature critic module in our framework, enforcing the disentangled visual features with domain generalization guarantees. Finally, our quantitative results on benchmark datasets confirm the effectiveness and robustness of our proposed model, performing favorably against state-of-the-art domain adaptation and generalization methods in segmentation. Zu-Yun Shiau, Ci-Siang Lin, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2021 | Robust Image Outpainting With Learnable Image MarginsabstractGiven a partial image input, image outpainting is to produce the desirable output by recovering or extending the surrounding image regions. While existing image outpainting methods achieve impressive results based on the recent advances of deep learning, they either lack the ability to extend image regions in arbitrary directions or require the filling image margins to be given in advance. To address this challenging task, we propose a unique deep learning framework for robust image outpainting, which consists of a margin prediction network and a teacher-student-based network for producing outpainted images. Our proposed model does not require image filling margins to be known beforehand, while both image appearance and perceptual feature consistencies can be jointly enforced. Our experiments quantitatively and qualitatively verify the effectiveness of our method, which is shown to perform favorably against baseline and state-of-the-art image outpainting works. Cheng-Yo Tan, Chiao-An Yang, Shang-Fu Chen, Meng-Lin Wu, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2021 | Leveraging Auxiliary Information from EMR for Weakly Supervised Pulmonary Nodule Detection
Hao-Hsiang Yang, Fu-En Wang, Cheng Sun 0004, Kuan-Chih Huang, Hung-Wei Chen, Hung-Chih Chen, Chun-Yu Liao, Shih-Hsuan Kao, Yu-Chiang Frank Wang, Chou-Chin Lan |
MICCAI (7) | 10 |
| 2021 | Adversarial Teacher-Student Representation Learning for Domain GeneralizationabstractDomain generalization (DG) aims to transfer the learning task from a single or multiple source domains to unseen target domains. To extract and leverage the information which exhibits sufficient generalization ability, we propose a simple yet effective approach of Adversarial Teacher-Student Representation Learning, with the goal of deriving the domain generalizable representations via generating and exploring out-of-source data distributions. Our proposed framework advances Teacher-Student learning in an adversarial learning manner, which alternates between knowledge-distillation based representation learning and novel-domain data augmentation. The former progressively updates the teacher network for deriving domain-generalizable representations, while the latter synthesizes data out-of-source yet plausible distributions. Extensive image classification experiments on benchmark datasets in multiple and single source DG settings confirm that, our model exhibits sufficient generalization ability and performs favorably against state-of-the-art DG methods. Fu-En Yang, Yuan-Chia Cheng, Zu-Yun Shiau, Yu-Chiang Frank Wang |
NeurIPS | 4 |
| 2021 | Joint Feature Disentanglement and Hallucination for Few-Shot Image ClassificationabstractFew-shot learning (FSL) refers to the learning task that generalizes from base to novel concepts with only few examples observed during training. One intuitive FSL approach is to hallucinate additional training samples for novel categories. While this is typically done by learning from a disjoint set of base categories with sufficient amount of training data, most existing works did not fully exploit the intra-class information from base categories, and thus there is no guarantee that the hallucinated data would represent the class of interest accordingly. In this paper, we propose Feature Disentanglement and Hallucination Network (FDH-Net), which jointly performs feature disentanglement and hallucination for FSL purposes. More specifically, our FDH-Net is able to disentangle input visual data into class-specific and appearance-specific features. With both data recovery and classification constraints, hallucination of image features for novel categories using appearance information extracted from base categories can be achieved. We perform extensive experiments on two fine-grained datasets (CUB and FLO) and two coarse-grained ones (mini-ImageNet and CIFAR-100). The results confirm that our framework performs favorably against state-of-the-art metric-learning and hallucination-based FSL models. Chia-Ching Lin, Hsin-Li Chu, Yu-Chiang Frank Wang, Chin-Laung Lei |
IEEE Trans. Image Process. | 3 |
| 2020 | Audiovisual Transformer with Instance Attention for Audio-Visual Event Localization
Yan-Bo Lin, Yu-Chiang Frank Wang |
ACCV (6) | 2 |
| 2020 | Transforming Multi-concept Attention into Video Summarization
Yen-Ting Liu, Yu-Jhe Li, Yu-Chiang Frank Wang |
ACCV (5) | 3 |
| 2020 | Learning Identity-Invariant Motion Representations for Cross-ID Face ReenactmentabstractHuman face reenactment aims at transferring motion patterns from one face (from a source-domain video) to an-other (in the target domain with the identity of interest).While recent works report impressive results, they are notable to handle multiple identities in a unified model. In this paper, we propose a unique network of CrossID-GAN to perform multi-ID face reenactment. Given a source-domain video with extracted facial landmarks and a target-domain image, our CrossID-GAN learns the identity-invariant motion patterns via the extracted landmarks and such information to produce the videos whose ID matches that of the target domain. Both supervised and unsupervised settings are proposed to train and guide our model during training.Our qualitative/quantitative results confirm the robustness and effectiveness of our model, with ablation studies confirming our network design. Po-Hsiang Huang, Fu-En Yang, Yu-Chiang Frank Wang |
CVPR | 3 |
| 2020 | Convolution in the Cloud: Learning Deformable Kernels in 3D Graph Convolution Networks for Point Cloud AnalysisabstractPoint clouds are among the popular geometry representations for 3D vision applications. However, without regular structures like 2D images, processing and summarizing information over these unordered data points are very challenging. Although a number of previous works attempt to analyze point clouds and achieve promising performances, their performances would degrade significantly when data variations like shift and scale changes are presented. In this paper, we propose 3D Graph Convolution Networks (3D-GCN), which is designed to extract local 3D features from point clouds across scales, while shift and scale-invariance properties are introduced. The novelty of our 3D-GCN lies in the definition of learnable kernels with a graph max-pooling mechanism. We show that 3D-GCN can be applied to 3D classification and segmentation tasks, with ablation studies and visualizations verifying the design of 3D-GCN. Zhi-Hao Lin, Sheng-Yu Huang, Yu-Chiang Frank Wang |
CVPR | 3 |
| 2020 | Learning to Learn in a Semi-supervised Fashion
Yun-Chun Chen, Chao-Te Chou, Yu-Chiang Frank Wang |
ECCV (18) | 3 |
| 2020 | Self-Supervised Deep Learning for Fisheye Image RectificationabstractTo rectify fisheye distortion from a single image, we advance self-supervised learning strategies and propose a unique deep learning model of Fisheye GAN (FE-GAN). Our FE-GAN learns pixel-level distortion flow from sets of fisheye distorted images and distortion-free ones (but not requiring such correspondences), with unique cross-rotation and intra-warping consistency introduced. With such novel self-supervised learning techniques, our FEGAN is able to recover the distortion-free image directly from the single fisheye image input. Our experiments quantitatively and qualitative confirm the effectiveness and robustness of our proposed model, which performs favorably against recent GAN-based image translation models. Chun-Hao Chao, Pin-Lun Hsu, Hung-yi Lee, Yu-Chiang Frank Wang |
ICASSP | 4 |
| 2020 | Face Feature Recovery via Temporal Fusion for Person SearchabstractSearching actors from videos by a single portrait image is a challenging task, due to large variations of video scenes and intra-person appearance. To tackle this problem, most recent works apply deep neural networks for detecting and extracting robust facial features for matching. However, when the face of an actor is not detected due to occlusion, such image-matching based strategies would not be applicable. To address the issue, we propose a unique framework of "Face Feature Recovery via Temporal Fusion" to synthesize virtual facial features by observing both temporal and contextual information. Once such face features are extracted, a simple extension to the k-nearest neighbors for re-ranking, "Iterative k-nearest Multi-fusion", is presented to utilize both face and body features for improved person search. We conduct extensive experiments to evaluate the performance of our framework on the challenging extended version of the Cast Search in Movies (ECSM) dataset [1]. Without utilizing tracklet information during training, the proposed approach still performs favorably against recent works in searching actors of interest from movie videos. Besides, we also show the proposed approach can be fused with them to further improve the performance. Cheng-Yu Fan, Chao-Peng Liu, Kuan-Chun Wang, Jiun-Hao Jhan, Yu-Chiang Frank Wang, Jun-Cheng Chen |
ICASSP | 5 |
| 2020 | Learning Disentangled Feature Representations For Anomaly DetectionabstractAnomaly detection is a challenging task that requires identifying anomalous data by observing only normal data during training. Previous works typically approached this problem by assessing data recovery with properly selected thresholds. However, the performance would be affected by content variants or background clutter. Hence, in this paper, we propose a novel deep learning based method, which learns disentangled feature representations for separating semantic and visual appearance information, so that the anomaly of the input data can be determined based on its semantic features. Our qualitative results demonstrate the feasibility of the feature disentanglement, and the quantitative experiments confirm that our method outperforms other methods. Wei-Yu Lee, Yu-Chiang Frank Wang |
ICIP | 2 |
| 2020 | Space-Time Guided Association Learning For Unsupervised Person Re-IdentificationabstractPerson re-identification (Re-ID) aims to match images of the same person across distinct camera views. In this paper, we propose the Space-Time Guided Association Learning (STGAL) for unsupervised Re-ID without ground truth identity nor image correspondence observed during training. By exploiting the spatial-temporal information presented in pedestrian data, our STGAL is able to identify positive and negative image pairs for learning Re-ID feature representations. Experiments on a variety of datasets confirm the effectiveness of our approach, which achieves promising performance when comparing to the state-of-the-art methods. Chih-Wei Wu, Chih-Ting Liu, Wei-Chih Tu, Yu Tsao 0001, Yu-Chiang Frank Wang, Shao-Yi Chien |
ICIP | 5 |
| 2020 | Wavelet Channel Attention Module With A Fusion Network For Single Image DerainingabstractSingle image deraining is a crucial problem because rain severely degenerates the visibility of images and affects the performance of computer vision tasks like outdoor surveillance systems and intelligent vehicles. In this paper, we propose the new convolutional neural network (CNN) called the wavelet channel attention module with a fusion network. Wavelet transform and the inverse wavelet transform are substituted for down-sampling and up-sampling so feature maps from the wavelet transform and convolutions contain different frequencies and scales. Furthermore, feature maps are integrated by channel attention. Our proposed network learns confidence maps of four sub-band images derived from the wavelet transform of the original images. Finally, the clear image can be well restored via the wavelet reconstruction and fusion of the low-frequency part and high-frequency parts. Several experimental results on synthetic and real images present that the proposed algorithm outperforms state-of-the-art methods. Hao-Hsiang Yang, Chao-Han Huck Yang, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2020 | Domain Generalized Person Re-Identification via Cross-Domain Episodic LearningabstractAiming at recognizing images of the same person across distinct camera views, person re-identification (re-ID) has been among active research topics in computer vision. Most existing re-ID works require collection of a large amount of labeled image data from the scenes of interest. When the data to be recognized are different from the source-domain training ones, a number of domain adaptation approaches have been proposed. Nevertheless, one still needs to collect labeled or unlabelled target-domain data during training. In this paper, we tackle an even more challenging and practical setting, domain generalized (DG) person re-ID. That is, while a number of labeled source-domain datasets are available, we do not have access to any target-domain training data. In order to learn domain-invariant features without knowing the target domain of interest, we present an episodic learning scheme which advances meta learning strategies to exploit the observed source-domain labeled data. The learned features would exhibit sufficient domain-invariant properties while not overfitting the source-domain data or ID labels. Our experiments on four benchmark datasets confirm the superiority of our method over the state-of-the-arts. Ci-Siang Lin, Yuan-Chia Cheng, Yu-Chiang Frank Wang |
ICPR | 3 |
| 2020 | Learning Interpretable Representation for 3D Point CloudsabstractPoint clouds have emerged as a popular representation of 3D visual data. With a set of unordered 3D points, one typically needs to transform them into latent representation before further classification and segmentation tasks. However, one cannot easily interpret such encoded latent representation. To address this issue, we propose a unique deep learning framework for disentangling body-type and pose information from 3D point clouds. Extending from autoencoder, we advance adversarial learning a selected feature type, while classification and data recovery can be additionally observed. Our experiments confirm that our model can be successfully applied to perform a wide range of 3D applications like shape synthesis, action translation, shape/action interpolation, and synchronization. Feng-Guang Su, Ci-Siang Lin, Yu-Chiang Frank Wang |
ICPR | 3 |
| 2020 | Semantics-Guided Representation Learning with Applications to Visual Synthesis
Jia-Wei Yan, Ci-Siang Lin, Fu-En Yang, Yu-Jhe Li, Yu-Chiang Frank Wang |
ICPR | 5 |
| 2020 | Dual-MTGAN: Stochastic and Deterministic Motion Transfer for Image-to-Video SynthesisabstractGenerating videos with content and motion variations is a challenging task in computer vision. While the recent development of GAN allows video generation from latent representations, it is not easy to produce videos with particular content of motion patterns of interest. In this paper, we propose Dual Motion Transfer GAN (Dual-MTGAN), which takes image and video data as inputs while learning disentangled content and motion representations. Our Dual-MTGAN is able to perform deterministic motion transfer and stochastic motion generation. Based on a given image, the former preserves the input content and transfers motion patterns observed from another video sequence, and the latter directly produces videos with plausible yet diverse motion patterns based on the input image. The proposed model is trained in an end-to-end manner, without the need to utilize pre-defined motion features like pose or facial landmarks. Our quantitative and qualitative results would confirm the effectiveness and robustness of our model in addressing such conditioned image-to-video tasks. Fu-En Yang, Jing-Cheng Chang, Yuan-Hao Lee, Yu-Chiang Frank Wang |
ICPR | 4 |
| 2020 | A Multi-Domain and Multi-Modal Representation Disentangler for Cross-Domain Image Manipulation and ClassificationabstractLearning interpretable data representation has been an active research topic in deep learning and computer vision. While representation disentanglement is an effective technique for addressing this task, existing works cannot easily handle the problems in which manipulating and recognizing data across multiple domains are desirable. In this paper, we present a unified network architecture of Multi-domain and Multi-modal Representation Disentangler (M2RD), with the goal of learning domain-invariant content representation with the associated domain-specific representation observed. By advancing adversarial learning and disentanglement techniques, the proposed model is able to perform continuous image manipulation across data domains with multiple modalities. More importantly, the resulting domain-invariant feature representation can be applied for unsupervised domain adaptation. Finally, our quantitative and qualitative results would confirm the effectiveness and robustness of the proposed model over state-of-the-art methods on the above tasks. Fu-En Yang, Jing-Cheng Chang, Chung-Chi Tsai, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 4 |
| 2019 | Learning Resolution-Invariant Deep Representations for Person Re-IdentificationabstractPerson re-identification (re-ID) solves the task of matching images across cameras and is among the research topics in vision community. Since query images in real-world scenarios might suffer from resolution loss, how to solve the resolution mismatch problem during person re-ID becomes a practical problem. Instead of applying separate image super-resolution models, we propose a novel network architecture of Resolution Adaptation and re-Identification Network (RAIN) to solve cross-resolution person re-ID. Advancing the strategy of adversarial learning, we aim at extracting resolution-invariant representations for re-ID, while the proposed model is learned in an end-to-end training fashion. Our experiments confirm that the use of our model can recognize low-resolution query images, even if the resolution is not seen during training. Moreover, the extension of our model for semi-supervised re-ID further confirms the scalability of our proposed method for real-world scenarios and applications. Yun-Chun Chen, Yu-Jhe Li, Xiaofei Du 0001, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2019 | Guide Your Eyes: Learning Image Manipulation under Saliency Guidance
Yen-Chung Chen, Keng-Jui Chang, Yi-Hsuan Tsai, Yu-Chiang Frank Wang, Walon Wei-Chen Chiu |
BMVC | 4 |
| 2019 | Spatially and Temporally Efficient Non-local Attention Network for Video-based Person Re-Identification
Chih-Ting Liu, Chih-Wei Wu, Yu-Chiang Frank Wang, Shao-Yi Chien |
BMVC | 3 |
| 2019 | Element-Embedded Style Transfer Networks for Style Harmonization
Hwai-Jin Peng, Chia-Ming Wang, Yu-Chiang Frank Wang |
BMVC | 3 |
| 2019 | Towards Scene Understanding: Unsupervised Monocular Depth Estimation With Semantic-Aware RepresentationabstractMonocular depth estimation is a challenging task in scene understanding, with the goal to acquire the geometric properties of 3D space from 2D images. Due to the lack of RGB-depth image pairs, unsupervised learning methods aim at deriving depth information with alternative supervision such as stereo pairs. However, most existing works fail to model the geometric structure of objects, which generally results from considering pixel-level objective functions during training. In this paper, we propose SceneNet to overcome this limitation with the aid of semantic understanding from segmentation. Moreover, our proposed model is able to perform region-aware depth estimation by enforcing semantics consistency between stereo pairs. In our experiments, we qualitatively and quantitatively verify the effectiveness and robustness of our model, which produces favorable results against the state-of-the-art approaches do. Po-Yi Chen, Alexander H. Liu, Yen-Cheng Liu, Yu-Chiang Frank Wang |
CVPR | 4 |
| 2019 | Spot and Learn: A Maximum-Entropy Patch Sampler for Few-Shot Image ClassificationabstractFew-shot learning (FSL) requires one to learn from object categories with a small amount of training data (as novel classes), while the remaining categories (as base classes) contain a sufficient amount of data for training. It is often desirable to transfer knowledge from the base classes and derive dominant features efficiently for the novel samples. In this work, we propose a sampling method that de-correlates an image based on maximum entropy reinforcement learning, and extracts varying sequences of patches on every forward-pass with discriminative information observed. This can be viewed as a form of "learned" data augmentation in the sense that we search for different sequences of patches within an image and performs classification with aggregation of the extracted features, resulting in improved FSL performances. In addition, our positive and negative sampling policies along with a newly defined reward function would favorably improve the effectiveness of our model. Our experiments on two benchmark datasets confirm the effectiveness of our framework and its superiority over recent FSL approaches. Wen-Hsuan Chu, Yu-Jhe Li, Jing-Cheng Chang, Yu-Chiang Frank Wang |
CVPR | 4 |
| 2019 | Perceptual Quality Preserving Image Super-resolution via Channel AttentionabstractGenerative Adversarial Network (GAN) has been widely applied on Single Image Super-Resolution (SISR) problems. However, there can be quite a variability in the results from the GAN-based methods. In some cases, the GAN-based methods might cause structure distortion, which can be easily distinguished by human beings, especially for artificial structures, because the methods only focus on the perceptual quality of the whole image. On the other hand, PSNR-oriented methods can prevent structure distortion but with overly smoothed context. To overcome these problems, we propose a deep neural net refiner for SISR methods, not only improving perceptual quality but also preserving context structures. In the experiments, our model qualitatively and quantitatively performs favorably against the state-of-the-art SISR methods. Wei-Yu Lee, Po-Yu Chuang, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2019 | Learning Pose-aware 3D Reconstruction via 2D-3D Self-consistencyabstract3D reconstruction, inferring 3D shape information from a single 2D image, has drawn attention from learning and vision communities. In this paper, we propose a framework for learning pose-aware 3D shape reconstruction. Our proposed model learns deep representation for recovering the 3D object, with the ability to extract camera pose information but without any direct supervision of ground truth camera pose. This is realized by exploitation of 2D-3D self-consistency between 2D masks and 3D voxels. Experiments qualitatively and quantitatively demonstrate the effectiveness and robustness of our model, which performs favorably against state-of-the-art methods. Yi-Lun Liao, Yao-Cheng Yang, Yuan-Fang Lin, Pin-Jung Chen, Chia-Wen Kuo, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang |
ICASSP | 7 |
| 2019 | Dual-modality Seq2Seq Network for Audio-visual Event LocalizationabstractAudio-visual event localization requires one to identify the event which is both visible and audible in a video (either at a frame or video level). To address this task, we propose a deep neural network named Audio-Visual sequence-to-sequence dual network (AVSDN). By jointly taking both audio and visual features at each time segment as inputs, our proposed model learns global and local event information in a sequence to sequence manner, which can be realized in either fully supervised or weakly supervised settings. Empirical results confirm that our proposed method performs favorably against recent deep learning approaches in both settings. Yan-Bo Lin, Yu-Jhe Li, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2019 | Recover and Identify: A Generative Dual Model for Cross-Resolution Person Re-IdentificationabstractPerson re-identification (re-ID) aims at matching images of the same identity across camera views. Due to varying distances between cameras and persons of interest, resolution mismatch can be expected, which would degrade person re-ID performance in real-world scenarios. To overcome this problem, we propose a novel generative adversarial network to address cross-resolution person re-ID, allowing query images with varying resolutions. By advancing adversarial learning techniques, our proposed model learns resolution-invariant image representations while being able to recover the missing details in low-resolution input images. The resulting features can be jointly applied for improving person re-ID performance due to preserving resolution invariance and recovering re-ID oriented discriminative details. Our experiments on five benchmark datasets confirm the effectiveness of our approach and its superiority over the state-of-the-art methods, especially when the input resolutions are unseen during training. Yu-Jhe Li, Yun-Chun Chen, Yen-Yu Lin, Xiaofei Du 0001, Yu-Chiang Frank Wang |
ICCV | 5 |
| 2019 | Cross-Dataset Person Re-Identification via Unsupervised Pose Disentanglement and AdaptationabstractPerson re-identification (re-ID) aims at recognizing the same person from images taken across different cameras. To address this challenging task, existing re-ID models typically rely on a large amount of labeled training data, which is not practical for real-world applications. To alleviate this limitation, researchers now targets at cross-dataset re-ID which focuses on generalizing the discriminative ability to the unlabeled target domain when given a labeled source domain dataset. To achieve this goal, our proposed Pose Disentanglement and Adaptation Network (PDA-Net) aims at learning deep image representation with pose and domain information properly disentangled. With the learned cross-domain pose invariant feature space, our proposed PDA-Net is able to perform pose disentanglement across domains without supervision in identities, and the resulting features can be applied to cross-dataset re-ID. Both of our qualitative and quantitative results on two benchmark datasets confirm the effectiveness of our approach and its superiority over the state-of-the-art cross-dataset Re-ID approaches. Yu-Jhe Li, Ci-Siang Lin, Yan-Bo Lin, Yu-Chiang Frank Wang |
ICCV | 4 |
| 2019 | Semantics-Guided Data Hallucination for Few-Shot Visual ClassificationabstractFew-shot learning (FSL) addresses learning tasks in which only few samples are available for selected object categories. In this paper, we propose a deep learning framework for data hallucination, which overcomes the above limitation and alleviate possible overfitting problems. In particular, our method exploits semantic information into the hallucination process, and thus the augmented data would be able to exhibit semantics-oriented modes of variation for improved FSL performances. Very promising performances on CIFAR-100 and AwA datasets confirm the effectiveness of our proposed method for FSL. Chia-Ching Lin, Yu-Chiang Frank Wang, Chin-Laung Lei, Kuan-Ta Chen |
ICIP | 2 |
| 2019 | Learning Hierarchical Self-Attention for Video SummarizationabstractVideo summarization still remains a challenging task. Due to sufficient video data on the Internet, such task draws significant attention in the vision community and benefits a wide range of applications, e.g., video retrieval, search, etc. To effectively perform video summarization by deriving the keyframes which represent the given input video, we propose a novel framework named Hierarchical Multi-Attention Network (H-MAN) which comprises the shot-level reconstruction model and multi-head attention model. While our designed attention model is two-stage hierarchical structure for producing various attention maps, we are among the first to utilize the multi-attention mechanism in the video summarization task, which brings improved performance. The quantitative and qualitative results demonstrate the effectiveness of our model, which performs favorably against state-of-the-art approaches. Yen-Ting Liu, Yu-Jhe Li, Fu-En Yang, Shang-Fu Chen, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2019 | Weakly-Supervised Learning for Attention-Guided Skull Fracture Classification In Computed Tomography ImagingabstractWe propose a novel attention-guided deep learning framework for image classification, with the goal of predicting both image and pixel-level labels for input images. While the training images are with either positive or negative labels, we do not assume that each training image is with annotated pixel-level ground truth information, and thus our method can be realized in such a weakly supervised setting. Our proposed module can be easily combined with standard CNN architectures with no extra parameter needed. With the above advantages, we evaluate our model on a skull CT dataset, and the experimental results confirm the effectiveness and robustness of our approach over recent popular CNN architectures. Cheng-Yen Yang, Chi-Hsin Lo, Huan-Chih Wang, Jen-Hai Chou, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2019 | A Closer Look at Few-shot Classification
Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, Jia-Bin Huang 0001 |
ICLR (Poster) | 4 |
| 2019 | Transfer Neural Trees: Semi-Supervised Heterogeneous Domain Adaptation and BeyondabstractHeterogeneous domain adaptation (HDA) addresses the task of associating data not only across dissimilar domains but also described by different types of features. Inspired by the recent advances of neural networks and deep learning, we propose a deep leaning model of Transfer Neural Trees (TNT), which jointly solves cross-domain feature mapping, adaptation, and classification in a unified architecture. As the prediction layer in TNT, we introduce Transfer Neural Decision Forest (Transfer- NDF), which is able to learn the neurons in TNT for adaptation by stochastic pruning. In order to handle semi-supervised HDA, a unique embedding loss term is introduced to TNT for preserving prediction and structural consistency between labeled and unlabeled target-domain data. We further show that our TNT can be extended to zero shot learning for associating image and attribute data with promising performance. Finally, experiments on different classification tasks across features, datasets, and modalities would verify the effectiveness of our TNT. Wei-Yu Chen, Tzu-Ming Harry Hsu, Yao-Hung Tsai, Ming-Syan Chen, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 5 |
| 2018 | Order-Free RNN With Visual Attention for Multi-Label ClassificationabstractWe propose a recurrent neural network (RNN) based model for image multi-label classification. Our model uniquely integrates and learning of visual attention and Long Short Term Memory (LSTM) layers, which jointly learns the labels of interest and their co-occurrences, while the associated image regions are visually attended. Different from existing approaches utilize either model in their network architectures, training of our model does not require pre-defined label orders. Moreover, a robust inference process is introduced so that prediction errors would not propagate and thus affect the performance. Our experiments on NUS-WISE and MS-COCO datasets confirm the design of our network and its effectiveness in solving multi-label classification problems. Shang-Fu Chen, Chih-Kuan Yeh, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2018 | Learning a Code-Space Predictor by Exploiting Intra-Image-Dependencies
Jan Klopp, Yu-Chiang Frank Wang, Shao-Yi Chien, Liang-Gee Chen |
BMVC | 2 |
| 2018 | Multi-Label Zero-Shot Learning With Structured Knowledge GraphsabstractIn this paper, we propose a novel deep learning architecture for multi-label zero-shot learning (ML-ZSL), which is able to predict multiple unseen class labels for each input instance. Inspired by the way humans utilize semantic knowledge between objects of interests, we propose a framework that incorporates knowledge graphs for describing the relationships between multiple labels. Our model learns an information propagation mechanism from the semantic label space, which can be applied to model the interdependencies between seen and unseen class labels. With such investigation of structured knowledge graphs for visual reasoning, we show that our model can be applied for solving multi-label classification and ML-ZSL tasks. Compared to state-of-the-art approaches, comparable or improved performances can be achieved by our method. Chung-wei Lee, Chih-Kuan Yeh, Yu-Chiang Frank Wang |
CVPR | 4 |
| 2018 | Detach and Adapt: Learning Cross-Domain Disentangled Deep RepresentationabstractWhile representation learning aims to derive interpretable features for describing visual data, representation disentanglement further results in such features so that particular image attributes can be identified and manipulated. However, one cannot easily address this task without observing ground truth annotation for the training data. To address this problem, we propose a novel deep learning model of Cross-Domain Representation Disentangler (CDRD). By observing fully annotated source-domain data and unlabeled target-domain data of interest, our model bridges the information across data domains and transfers the attribute information accordingly. Thus, cross-domain feature disentanglement and adaptation can be jointly performed. In the experiments, we provide qualitative results to verify our disentanglement capability. Moreover, we further confirm that our model can be applied for solving classification tasks of unsupervised domain adaptation, and performs favorably against state-of-the-art image disentanglement and translation methods. Yen-Cheng Liu, Yu-Ying Yeh, Tzu-Chien Fu, Sheng-De Wang, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang |
CVPR | 6 |
| 2018 | Deep Generative Models for Weakly-Supervised Multi-Label Classification
Hong-Min Chu, Chih-Kuan Yeh, Yu-Chiang Frank Wang |
ECCV (2) | 3 |
| 2018 | Summarizing First-Person Videos from Third Persons' Points of Views
Hsuan-I Ho, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang |
ECCV (15) | 3 |
| 2018 | Learning Semantics-Guided Visual Attention for Few-Shot Image ClassificationabstractWe propose a deep learning framework for few-shot image classification, which exploits information across label semantics and image domains, so that regions of interest can be properly attended for improved classification. The proposed semantics-guided attention module is able to focus on most relevant regions in an image, while the attended image samples allow data augmentation and alleviate possible overfitting during FSL training. Promising performances are presented in our experiments, in which we consider both closed and open-world settings. The former considers the test input belong to the categories of few shots only, while the latter requires recognition of all categories of interest. Wen-Hsuan Chu, Yu-Chiang Frank Wang |
ICIP | 2 |
| 2018 | Deep Reinforcement Learning for Playing 2.5D Fighting GamesabstractDeep reinforcement learning has shown its success in game playing. However, 2.5D fighting games would be a challenging task to handle due to ambiguity in visual appearances like height or depth of the characters. Moreover, actions in such games typically involve particular sequential action orders, which also makes the network design very difficult. Based on the network of Asynchronous Advantage Actor-Critic (A3C), we create an OpenAI-gym-like gaming environment with the game of Little Fighter 2 (LF2), and present a novel A3C+ network for learning RL agents. The introduced model includes a Recurrent Info network, which utilizes game-related info features with recurrent layers to observe combo skills for fighting. In the experiments, we consider LF2 in different settings, which successfully demonstrates the use of our proposed model for learning 2.5D fighting games. Yu-Jhe Li, Hsin-Yu Chang, Yu-Jing Lin, Po-Wei Wu, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2018 | A Unified Feature Disentangler for Multi-Domain Image Translation and ManipulationabstractWe present a novel and unified deep learning framework which is capable of learning domain-invariant representation from data across multiple domains. Realized by adversarial training with additional ability to exploit domain-specific information, the proposed network is able to perform continuous cross-domain image translation and manipulation, and produces desirable output images accordingly. In addition, the resulting feature representation exhibits superior performance of unsupervised domain adaptation, which also verifies the effectiveness of the proposed model in learning disentangled features for describing cross-domain data. Alexander H. Liu, Yen-Cheng Liu, Yu-Ying Yeh, Yu-Chiang Frank Wang |
NeurIPS | 4 |
| 2018 | Edge-Preserving Depth Map Upsampling by Joint Trilateral FilterabstractCompared to the color images, their associated depth images captured by the RGB-D sensors are typically with lower resolution. The task of depth map super-resolution (SR) aims at increasing the resolution of the range data by utilizing the high-resolution (HR) color image, while the details of the depth information are to be properly preserved. In this paper, we present a joint trilateral filtering (JTF) algorithm for depth image SR. The proposed JTF first observes context information from the HR color image. In addition to the extracted spatial and range information of local pixels, our JTF further integrates local gradient information of the depth image, which allows the prediction and refinement of HR depth image outputs without artifacts like textural copies or edge discontinuities. Quantitative and qualitative experimental results demonstrate the effectiveness and robustness of our approach over prior depth map upsampling works. Kai-Han Lo, Yu-Chiang Frank Wang, Kai-Lung Hua |
IEEE Trans. Cybern. | 2 |
| 2017 | Learning Deep Latent Space for Multi-Label ClassificationabstractMulti-label classification is a practical yet challenging task in machine learning related fields, since it requires the prediction of more than one label category for each input instance. We propose a novel deep neural networks (DNN) based model, Canonical Correlated AutoEncoder (C2AE), for solving this task. Aiming at better relating feature and label domain data for improved classification, we uniquely perform joint feature and label embedding by deriving a deep latent space, followed by the introduction of label-correlation sensitive loss function for recovering the predicted label outputs. Our C2AE is achieved by integrating the DNN architectures of canonical correlation analysis and autoencoder, which allows end-to-end learning and prediction with the ability to exploit label dependency. Moreover, our C2AE can be easily extended to address the learning problem with missing labels. Our experiments on multiple datasets with different scales confirm the effectiveness and robustness of our proposed method, which is shown to perform favorably against state-of-the-art methods for multi-label classification. Chih-Kuan Yeh, Wei-Chieh Wu, Wei-Jen Ko, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2017 | Semantics-Preserving Locality Embedding for Zero-Shot Learning
Shi-Yen Tao, Yi-Ren Yeh, Yu-Chiang Frank Wang |
BMVC | 3 |
| 2017 | Enhanced canonical correlation analysis with local density for cross-domain visual classificationabstractReal-world visual classification tasks typically need to deal with data observed from different domains. Inspired by canonical correlation analysis (CCA), we propose an enhanced CCA with local density for associating and recognizing cross-domain data. In addition to maximizing the correlation of the projected cross-domain data, our CCA model further exploits the local density information observed from each domain. As a result, our CCA not only exhibits excellent abilities in identifying representative data, noisy data like outliers can be further suppressed during the derivation of our CCA subspace. In our experiments, we successfully apply the proposed methods for solving two cross-domain classification tasks: person re-identification and cross-view action recognition. Wei-Jen Ko, Jheng-Ying Yu, Wei-Yu Chen, Yu-Chiang Frank Wang |
ICASSP | 4 |
| 2017 | Learning Grassmann manifolds for object state discoveryabstractIn this paper, we advocate the use of Grassmann manifolds for discovering object images in different states (e.g., unripe, peeled, etc.). We propose a novel dictionary learning algorithm, which derives the subspaces on a Grassmann manifold for describing each object state. By our introduced geodesic-flow constraint, our Grassmann manifold exhibits excellent capabilities in relating objects in distinct states (i.e., subspaces on the derived manifold), while the geodesics connecting different states can be viewed as transformations between the associated states. This is the reason why the use of our proposed Grassmann manifold can be applied to perform object state classification with improved performance. In our experiments, we provide quantitative and qualitative results to verify the effectiveness of our proposed method. Hao-Wei Lee, Chia-Po Wei, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2017 | Partial image blur detection and segmentation from a single snapshotabstractIn this paper, we address the problem of detecting and segmenting partial image blur from a single input image. Instead of assuming particular image priors or requiring additional user annotation, we propose a novel learning framework which jointly solves the tasks of blur kernel estimation and image blur segmentation, so that partial image blur can be automatically separated from the remaining parts of the input image. By alternating between the two learning tasks, we show that our proposed method would achieve promising detection and segmentation performance, which would benefit further processing or analysis tasks of interest. We also verify that, via both qualitative and quantitative evaluation, our approach would perform favorably against state-of-the-art blur detection or segmentation works. Te-Li Wang, Kuan-Yun Lee, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2017 | No More Discrimination: Cross City Adaptation of Road Scene SegmentersabstractDespite the recent success of deep-learning based semantic segmentation, deploying a pre-trained road scene segmenter to a city whose images are not presented in the training set would not achieve satisfactory performance due to dataset biases. Instead of collecting a large number of annotated images of each city of interest to train or refine the segmenter, we propose an unsupervised learning approach to adapt road scene segmenters across different cities. By utilizing Google Street View and its time-machine feature, we can collect unannotated images for each road scene at different times, so that the associated static-object priors can be extracted accordingly. By advancing a joint global and class-specific domain adversarial learning framework, adaptation of pre-trained segmenters to that city can be achieved without the need of any user annotation or interaction. We show that our method improves the performance of semantic segmentation in multiple cities across continents, while it performs favorably against state-of-the-art approaches requiring annotated training data. Yi-Hsin Chen, Wei-Yu Chen, Bo-Cheng Tsai, Yu-Chiang Frank Wang |
ICCV | 5 |
| 2017 | Occlusion-aware face inpainting via generative adversarial networksabstractFace inpainting aims to restore the corrupted regions of face images due to extreme lighting variations, occlusion, or even disguise. This task becomes especially challenging, when the face images are taken in an unconstrained environment (i.e., with pose, illumination, and expression variations) and the type of corruption is not known in advance. In this paper, we propose a deep-learning based approach of occlusion-aware generative adversarial networks (GAN) for solving this problem. By utilizing GAN pre-trained on occlusion-free images, we are able to detect corrupted image regions automatically with the associated image pixels properly recovered. We produce promising performances on images from the benchmark dataset of LFW, and show that recognition of such face images would be benefited from our proposed approach. Yu-An Chen, Wei-Che Chen, Chia-Po Wei, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2016 | Domain-Constraint Transfer Coding for Imbalanced Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) deals with the task that labeled training and unlabeled test data collected from source and target domains, respectively. In this paper, we particularly address the practical and challenging scenario of imbalanced cross-domain data. That is, we do not assume the label numbers across domains to be the same, and we also allow the data in each domain to be collected from multiple datasets/sub-domains. To solve the above task of imbalanced domain adaptation, we propose a novel algorithm of Domain-constraint Transfer Coding (DcTC). Our DcTC is able to exploit latent subdomains within and across data domains, and learns a common feature space for joint adaptation and classification purposes. Without assuming balanced cross-domain data as most existing UDA approaches do, we show that our method performs favorably against state-of-the-art methods on multiple cross-domain visual classification tasks. Yao-Hung Tsai, Cheng-An Hou, Wei-Yu Chen, Yi-Ren Yeh, Yu-Chiang Frank Wang |
AAAI | 5 |
| 2016 | Learning Cross-Domain Landmarks for Heterogeneous Domain AdaptationabstractWhile domain adaptation (DA) aims to associate the learning tasks across data domains, heterogeneous domain adaptation (HDA) particularly deals with learning from cross-domain data which are of different types of features. In other words, for HDA, data from source and target domains are observed in separate feature spaces and thus exhibit distinct distributions. In this paper, we propose a novel learning algorithm of Cross-Domain Landmark Selection (CDLS) for solving the above task. With the goal of deriving a domain-invariant feature subspace for HDA, our CDLS is able to identify representative cross-domain data, including the unlabeled ones in the target domain, for performing adaptation. In addition, the adaptation capabilities of such cross-domain landmarks can be determined accordingly. This is the reason why our CDLS is able to achieve promising HDA performance when comparing to state-of-the-art HDA methods. We conduct classification experiments using data across different features, domains, and modalities. The effectiveness of our proposed method can be successfully verified. Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang |
CVPR | 3 |
| 2016 | Transfer Neural Trees for Heterogeneous Domain Adaptation
Wei-Yu Chen, Tzu-Ming Harry Hsu, Yao-Hung Tsai, Yu-Chiang Frank Wang, Ming-Syan Chen |
ECCV (5) | 4 |
| 2016 | Style-centric image summarization from photographic views of a cityabstractVisual summarization addresses the task of selecting images from an image collection, so that the sampled images would contain representative information which sufficiently highlights the collected visual data. In this paper, we solve the problem of style-centric visual summarization using photographic landmark images of a city. Different from existing works which typically retrieve landmark images based on salient visual appearances, our proposed method is able to produce different sets of summarized images, while each set corresponds to a particular image style. This is achieved by performing unsupervised clustering on images within and across landmark categories, which discovers the common photographic styles from the input image collection. Our experiments will confirm that, compared to standard clustering algorithms, our approach is able to achieve satisfactory summarization outputs with style consistency. Wei-Yi Chang, Yu-Chiang Frank Wang |
ICASSP | 2 |
| 2016 | Heterogeneous domain adaptation with label and structure consistencyabstractDomain adaptation is a challenging task, since it associates data collected from different domains or exhibiting distinct distributions. In this paper, we particularly focus on adapting cross-domain data with distinct feature dimensions or representations. Thus, this is referred to as the task of heterogeneous domain adaptation (HDA). To solve HDA, we propose Label and Structure-consistent Unilateral Projection (LS-UP) that transforms source-domain data to the target domain, with the goal of matching cross-domain data distribution and preserving data structure after projection. The main contribution of our work is its ability in relating cross-domain data with different feature representations. We evaluate our LS-UP for HDA on two different cross-domain classification problems, and we show that our method would perform favorably against state-of-the-art approaches. Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2016 | Style retrieval from natural imagesabstractIt has been a challenging task to identify and distinguish between images of different styles. The challenges mainly come from the extraction of high-level image semantic information, and the presence of the associated ambiguity. In this work, we propose a ranking model for style identification. Given training images of different styles, we learn a pointwise ranking model for each style based on random forests. To handle the high dimensionality of visual features and to prevent against possible ambiguity, we further introduce dimension reduction and pruning techniques for our random forests. In our experiments, we provide quantitative evaluation for style categorization in terms of mean square error (MSE) and relative ranking accuracy. Moreover, our visualization and qualitative results support the use of the proposed method for style retrieval of natural images. Ting-En Tseng, Wei-Yi Chang, Chu-Song Chen, Yu-Chiang Frank Wang |
ICASSP | 4 |
| 2016 | Recognizing heterogeneous cross-domain data via generalized joint distribution adaptationabstractIn this paper, we propose a novel algorithm of Generalized Joint Distribution Adaptation (G-JDA) for heterogeneous domain adaptation (HDA), which associates and recognizes cross-domain data observed in different feature spaces (and thus with different dimensionality). With the objective to derive a domain-invariant feature subspace for relating source and target-domain data, our G-JDA learns a pair of feature projection matrices (one for each domain), which allows us to eliminate the difference between projected cross-domain heterogeneous data by matching their marginal and class-conditional distributions. We conduct experiments on cross-domain classification tasks using data across different features, datasets, and modalities. We confirm that our G-JDA would perform favorably against state-of-the-art HDA approaches. Yuan-Ting Hsieh, Shi-Yen Tao, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang |
ICME | 5 |
| 2016 | With one look: 3D face shape estimation from a single snapshotabstractEstimating the 3D shape information of a face from a single image is a challenging task, especially when the input image is captured under unconstrained scenarios (e.g., variations of pose, illumination, expression, or even disguise). Previous approaches to this problem typically require careful initialization, registration, or segmentation of the face image regions. With the objective to match the detected landmarks of the input image with those of a set of reference 3D models, we propose a non-negative least squares (NNLS) based algorithm for joint pose and shape estimation. With the additional imposed pose regularization, our method is able to perform person-specific shape estimation, while the camera pose can be simultaneously recovered. We show that our method is robust, effective, and computationally feasible. Moreover, it would perform favorably against existing approaches to 3D shape estimation from a single unconstrained image. Chia-Po Wei, Yu-Chiang Frank Wang |
ICME | 2 |
| 2016 | Learning patch-based anchors for face hallucinationabstractWith the goal of increasing the resolution of face images, recent face hallucination methods advance learning techniques which observe training low and high-resolution patches for recovering the output image of interest. Since most existing patch-based face hallucination approaches do not consider the location information of the patches to be hallucinated, the resulting performance might be limited. In this paper, we propose an anchored patch-based hallucination method, which is able to exploit and identify image patches exhibiting structurally and spatially similar information. With these representative anchors observed, improved performance and computation efficiency can be achieved. Experimental results demonstrate that our proposed method achieves satisfactory performance and performs favorably against recent face hallucination approaches. Wei-Jen Ko, Yu-Chiang Frank Wang, Shao-Yi Chien |
MMSP | 2 |
| 2016 | Robust propagated filtering with applications to image texture filtering and beyondabstractExtracting meaningful structures from an image is an important task and benefits a wide range of image application tasks. However, it is typically very challenging to distinguish between noisy or textural patterns from image structures, especially when such patterns do not exhibit regularity (e.g., irregular textural patterns or those with varying scales). While existing edge-preserving image filters like bilateral, guided, or propagation filters aim at observing strong image edges, they cannot be easily applied to solve the above texture filtering tasks. In this paper, we propose robust propagated filter, which is an extension to propagation filters while exhibiting excellent ability in eliminating the aforementioned textural patterns when performing filtering. We show in our experimental results that our filter provides promising results on image filtering. Additional experiments on inverse image half toning and detail enhancement further verify the effectiveness of our proposed method. Hsin-Yuan Dennis Wen, Yu-Chiang Frank Wang |
MMSP | 2 |
| 2016 | Learning graph fusion for query and database specific image retrievalabstractIn this paper, we propose a graph-based image retrieval algorithm via query and database specific feature fusion. While existing feature fusion approaches exist for image retrieval, they typically do not consider the image database of interest (i.e., to be retrieved) for observing the associated feature contributions. In the offline learning stage, our proposed method first identifies representative features for describing images to be retrieved. Given a query input, we further exploit and integrate its visual information and utilize graph-based fusion for performing query-database specific retrieval. In our experiments, we show that our proposed method achieves promising performance on the benchmark database of UKbench, and performs favorably against recent fusion-based image retrieval approaches. Chih-Kuan Yeh, Wei-Chieh Wu, Yu-Chiang Frank Wang |
MMSP | 3 |
| 2016 | Context-aware joint dictionary learning for color image demosaicking
Kai-Lung Hua, Shintami Chusnul Hidayati, Fang-Lin He, Chia-Po Wei, Yu-Chiang Frank Wang |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Unsupervised Domain Adaptation With Label and Structural ConsistencyabstractUnsupervised domain adaptation deals with scenarios in which labeled data are available in the source domain, but only unlabeled data can be observed in the target domain. Since the classifiers trained by source-domain data would not be expected to generalize well in the target domain, how to transfer the label information from source to target-domain data is a challenging task. A common technique for unsupervised domain adaptation is to match cross-domain data distributions, so that the domain and distribution differences can be suppressed. In this paper, we propose to utilize the label information inferred from the source domain, while the structural information of the unlabeled target-domain data will be jointly exploited for adaptation purposes. Our proposed model not only reduces the distribution mismatch between domains, improved recognition of target-domain data can be achieved simultaneously. In the experiments, we will show that our approach performs favorably against the state-of-the-art unsupervised domain adaptation methods on benchmark data sets. We will also provide convergence, sensitivity, and robustness analysis, which support the use of our model for cross-domain classification. Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 4 |
| 2016 | RedEye: Preventing Collisions Caused by Red-Light Running Scooters With SmartphonesabstractIn this paper, we present a scooter collision avoidance system that can identify red-light runners (RLRs) at intersections. When the RLR behavior is detected, the system would advise the RLR to slow down immediately and warn nearby vehicles on the intersecting road in real time. In particular, we do not consider infrastructure-based solutions such as those utilizing a radar or a camera. This is because, in addition to high implementation costs, collisions can be only avoided at intersections where such infrastructure configurations are deployed. Instead, we advance an on-scooter solution using smartphones carried by scooter riders. Smartphones provide a useful platform that has a high penetration rate, more than sufficient computational power, inertial sensors to reflect the driving behavior, and the communication capability to transmit or receive information from other vehicles. In our system, we utilize a support vector machine and design an RLR classifier for learning and predicting RLR behaviors. The evaluation results show that our system is able to achieve over 70% recognition rates when distinguishing between RLR and non-RLR cases, as compared with approximately 80% recognition rates of the infrastructure-based (and higher cost) solution using a laser range finder (LADAR). Kuang-Shih Huang, Po-Jui Chiu, Hsin-Mu Tsai, Chih-Chung Kuo, Hui-Yu Lee, Yu-Chiang Frank Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2015 | An unsupervised domain adaptation approach for cross-domain visual classificationabstractFor cross-view action recognition and many real-world visual classification problems, one needs to recognize test data at a particular target domain of interest, while training data are collected at a different source domain. Without eliminating such domain differences, recognition of test data using classifiers trained in the source domain will not be expected to produce satisfactory performance. In this paper, we propose a novel domain adaptation approach, which is able to learn a common feature space relating cross-domain data. In particular, we not only aim at matching cross-domain data marginal distributions during adaptation, we also exploit the structure of target domain data and update class-conditional distributions accordingly. Experiments on various cross-domain visual classification tasks would verify the effectiveness and robustness of our proposed method. Cheng-An Hou, Yi-Ren Yeh, Yu-Chiang Frank Wang |
AVSS | 3 |
| 2015 | Propagated image filteringabstractWe propose the propagation filter as a novel image filtering operator, with the goal of smoothing over neighboring image pixels while preserving image context like edges or textural regions. In particular, our filter does not to utilize explicit spatial kernel functions as bilateral and guided filters do. We will show that our propagation filter can be viewed as a robust estimator, which minimizes the expected difference between the filtered and desirable image outputs. We will also relate propagation filtering to belief propagation, and suggest techniques if further speedup of the filtering process is necessary. In our experiments, we apply our propagation filter to a variety of applications such as image denoising, smoothing, fusion, and high-dynamic-range (HDR) compression. We will show that improved performance over existing image filters can be achieved. Jen-Hao Rick Chang, Yu-Chiang Frank Wang |
CVPR | 2 |
| 2015 | Unsupervised Domain Adaptation with Imbalanced Cross-Domain DataabstractWe address a challenging unsupervised domain adaptation problem with imbalanced cross-domain data. For standard unsupervised domain adaptation, one typically obtains labeled data in the source domain and only observes unlabeled data in the target domain. However, most existing works do not consider the scenarios in which either the label numbers across domains are different, or the data in the source and/or target domains might be collected from multiple datasets. To address the aforementioned settings of imbalanced cross-domain data, we propose Closest Common Space Learning (CCSL) for associating such data with the capability of preserving label and structural information within and across domains. Experiments on multiple cross-domain visual classification tasks confirm that our method performs favorably against state-of-the-art approaches, especially when imbalanced cross-domain data are presented. Tzu-Ming Harry Hsu, Wei-Yu Chen, Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang |
ICCV | 6 |
| 2015 | Connecting the dots without clues: Unsupervised domain adaptation for cross-domain visual classificationabstractMany real-world visual classification tasks require one to recognize test data in a particular domain of interest, while the training data can only be collected from a different domain. This can be viewed as the problem of unsupervised domain adaptation, in which the domain difference and the lack of cross-domain label/correspondence information make the recognition task very difficult. In this paper, we propose to exploit the cross-domain data correspondence using both observed data similarity and labels transferred from the source domain. This allows us to perform distribution matching for cross-domain data with recognition guarantees. Our experiments on three different cross-domain visual classification tasks would confirm the effectiveness of our method, which is shown to perform favorably against state-of-the-art unsupervised domain adaptation approaches. Wei-Yu Chen, Tzu-Ming Harry Hsu, Cheng-An Hou, Yi-Ren Yeh, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2015 | Graph regularized low-rank matrix recovery for robust person re-identificationabstractRobust person re-identification (PRID) refers to the problem of matching individuals across non-overlapping camera views, while the images captured by either camera might be occluded or even missing. To address this challenging task, we propose a low-rank matrix recovery (LR) based approach in this paper. In addition to observing the global structure of cross-camera images via LR, we further exploit their local geometrical information via graph regularization, which preserves the recovered images with recognition guarantees. Our experiments verify the effectiveness and robustness of our approach, which is shown to perform favorably against state-of-the-art PRID methods. Ming-Chia Tsai, Chia-Po Wei, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2015 | Seeing through the appearance: Body shape estimation using multi-view clothing imagesabstractWe propose a learning-based algorithm for body shape estimation, which only requires 2D clothing images taken in multiple views as the input data. Compared with the use of 3D scanners or depth cameras, although our setting is more user friendly, it also makes the learning and estimation problems more challenging. In addition to utilizing ground truth body images for constructing human body models at each view of interest, our work uniquely associates the anthropometric measurements (e.g., body height or leg length) across different views. For performing body shape estimation using multi-view clothing images, the proposed algorithm solves an optimization task which recovers the body shape with image and measurement reconstruction guarantees. In the experiments, we will show that the use of our proposed method would achieve satisfactory estimation results, and performs favorably against single-view or other baseline approaches for both body shape and measurement estimation. Wei-Yi Chang, Yu-Chiang Frank Wang |
ICME | 2 |
| 2015 | Undersampled face recognition with one-pass dictionary learningabstractUndersampled face recognition deals with the problem in which, for each subject to be recognized, only one or few images are available in the gallery (training) set. Thus, it is very difficult to handle large intra-class variations for face images. In this paper, we propose a one-pass dictionary learning algorithm to derive an auxiliary dictionary from external data, which consists of image variants of the subjects not of interest (not to be recognized). The proposed algorithm not only allows us to efficiently model intra-class variations such as illumination and expression changes, it also exhibits excellent abilities in recognizing corrupted images due to occlusion. In our experiments, we will show that our method would perform favorably against existing sparse representation or dictionary learning based approaches. Moreover, our computation time is remarkably less than that of recent dictionary learning based face recognition methods. Therefore, the effectiveness and efficiency of our proposed algorithm can be successfully verified. Chia-Po Wei, Yu-Chiang Frank Wang |
ICME | 2 |
| 2015 | R2P: Recomposition and Retargeting of Photographic ImagesabstractIn this paper, we propose a novel approach for performing joint recomposition and retargeting of photographic images (R2P). Given a reference image of interest, our method is able to automatically alter the composition of the input source image accordingly, while the recomposed output will be jointly retargeted to fit the reference. This is achieved by recomposing the visual components of the source image via graph matching, followed by solving a constrained mesh-warping based optimization problem for retargeting. As a result, the recomposed output image would fit the reference while suppressing possible distortion. Our experiments confirm that our proposed R2P method is able to achieve visually satisfactory results, without the need to use pre-collected labeled data or predetermined aesthetics rules. Hui-Tang Chang, Po-Cheng Pan, Yu-Chiang Frank Wang, Ming-Syan Chen |
ACM Multimedia | 3 |
| 2015 | ObjectMinutiae: Fingerprinting for Object AuthenticationabstractIn this work, we present \emph{ObjectMinutiae}, which is a framework for authenticating different objects or materials via extracting and matching their fingerprints. Unlike biometrics fingerprinting processes, which use patterns such as ridge ending and bifurcation points as the interest points, our work applies stereo photometric techniques for reconstructing objects' local image regions that contain the surface texture information. The interest points of the recovered image regions can be detected and described by state-of-the-art computer vision algorithms. Together with dimension reduction and hashing techniques, our proposed system is able to perform object verification using compact image features. With neutral and different torturing conditions, preliminary results on multiple types of papers support the use of our framework for practical object authentication tasks. Tzu-Yun Lin, Yu-Chiang Frank Wang, Sean Moss-Pultz |
ACM Multimedia | 2 |
| 2015 | Optimizing the decomposition for multiple foreground cosegmentation
Haw-Shiuan Chang, Yu-Chiang Frank Wang |
Comput. Vis. Image Underst. | 2 |
| 2015 | Query-Adaptive Multiple Instance Learning for Video Instance RetrievalabstractGiven a query image containing the object of interest (OOI), we propose a novel learning framework for retrieving relevant frames from the input video sequence. While techniques based on object matching have been applied to solve this task, their performance would be typically limited due to the lack of capabilities in handling variations in visual appearances of the OOI across video frames. Our proposed framework can be viewed as a weakly supervised approach, which only requires a small number of (randomly selected) relevant and irrelevant frames from the input video for performing satisfactory retrieval performance. By utilizing frame-level label information of such video frames together with the query image, we propose a novel query-adaptive multiple instance learning algorithm, which exploits the visual appearance information of the OOI from the query and that of the aforementioned video frames. As a result, the derived learning model would exhibit additional discriminating abilities while retrieving relevant instances. Experiments on two real-world video data sets would confirm the effectiveness and robustness of our proposed approach. Ting-Chu Lin, Min-Chun Yang, Chia-Yin Tsai, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 4 |
| 2015 | Undersampled Face Recognition via Robust Auxiliary Dictionary LearningabstractIn this paper, we address the problem of robust face recognition with undersampled training data. Given only one or few training images available per subject, we present a novel recognition approach, which not only handles test images with large intraclass variations such as illumination and expression. The proposed method is also to handle the corrupted ones due to occlusion or disguise, which is not present during training. This is achieved by the learning of a robust auxiliary dictionary from the subjects not of interest. Together with the undersampled training data, both intra and interclass variations can thus be successfully handled, while the unseen occlusions can be automatically disregarded for improved recognition. Our experiments on four face image datasets confirm the effectiveness and robustness of our approach, which is shown to outperform state-of-the-art sparse representation-based methods. Chia-Po Wei, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 2 |
| 2014 | Simple-to-Complex Discriminative Clustering for Hierarchical Image Segmentation
Haw-Shiuan Chang, Yu-Chiang Frank Wang |
ACCV (3) | 2 |
| 2014 | Extended-bag-of-features for translation, rotation, and scale-invariant image retrievalabstractWhile bag-of-features (BOF) models have been widely applied for addressing image retrieval problems, the resulting performance is typically limited due to its disregard of spatial information of local image descriptors (and the associated visual words). In this paper, we present a novel spatial pooling scheme, called extended bag-of-features (EBOF), for solving the above task. Besides improving image representation capability, the incorporation of the our EBOF model with a proposed circular-correlation based similarity measure allows us to perform translation, rotation, and scale-invariant image retrieval. We conduct experiments on two benchmark image datasets, and the performance confirms the effectiveness and robustness of our proposed approach. Chia-Yin Tsai, Ting-Chu Lin, Chia-Po Wei, Yu-Chiang Frank Wang |
ICASSP | 4 |
| 2014 | Exploiting low-rank structures from cross-camera images for robust person re-identificationabstractMatching individuals across non-overlapping camera views is known as the problem of person re-identification. In addition to significant visual appearance variations due to lighting, view angle, etc. changes, one might encounter corrupted data due to background clutter and occlusion, or even missing data at some camera views in practical scenarios. To address the above challenges, we present a novel approach to robust person re-identification, particularly aiming at handling missing and corrupted image data across camera views. Based on the technique of low-rank matrix decomposition, our proposed algorithm observes the low-rank structure of cross-view data, which is able to disregard extreme/sparse errors while the missing instances can be recovered automatically. Our experiments will confirm the effectiveness and robustness of our method, which is shown to outperform several baseline and state-of-the-art person re-identification approaches. Ming-Hang Fu, Yu-Chiang Frank Wang, Chu-Song Chen |
ICIP | 2 |
| 2014 | Exploiting image structural similarity for single image rain removalabstractWithout any prior knowledge or user interaction, single image rain removal has been a challenging task. Typically, one needs to disregard image components associated with the rain patterns, so that rain removal can be achieved via image reconstruction. By observing the limitations of standard batch-mode learning-based methods, we propose to exploit the structural similarity of the image bases for solving this task. By formulating the basis selection as an optimization problem, we are able to disregard those associated with rain patterns while the detailed image information can be preserved. Experiments on both synthetic and real-world images will verify the effectiveness of our proposed method. Shao-Hua Sun, Shang-Pu Fan, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2014 | Person-specific domain adaptation with applications to heterogeneous face recognitionabstractHeterogeneous face recognition (HFR) is a practical yet challenging task in which gallery and probe face images are collected in terms of different modalities or features (e.g., sketch vs. photo). In this paper, we present a person-specific domain adaptation framework for HFR. By utilizing the subjects not of interest (i.e., those not to be recognized), we first derive a common feature space using their cross-domain face images, with the goal of eliminating differences between image modalities. To generalize our feature space for representing and recognizing the subjects of interest, we advocate the construction of person-specific domain adaptation model in this space, so that the classifiers (trained by the gallery images) are able to achieve satisfactory recognition performance. In our experiments, we consider sketch-to-photo and near-infrared (NIR) to visible spectrum (VIS) face recognition problems for evaluating the effectiveness of our method. Yao-Hung Tsai, Hung-Ming Hsu, Cheng-An Hou, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2014 | Multi-view Nonnegative Matrix Factorization for Clothing Image CharacterizationabstractDue to the ambiguity in describing and discriminating between clothing images of different styles, it has been a challenging task to solve clothing image characterization problems. Based on the use of multiple types of visual features, we propose a novel multi-view nonnegative matrix factorization (NMF) algorithm for solving the above task. Our multi-view NMF not only observes image representations for describing clothing images in terms of visual appearances, an optimal combination of such features for each clothing image style would also be learned, while the separation between different image styles can be preserved. To verify the effectiveness of our method, we conduct experiments on two image datasets, and we confirm that our method produces satisfactory performance in terms of both clustering and categorization. Wei-Yi Chang, Chia-Po Wei, Yu-Chiang Frank Wang |
ICPR | 3 |
| 2014 | Domain Adaptive Self-Taught Learning for Heterogeneous Face RecognitionabstractRecognizing image data across different domains has been a challenging task. For biometrics, heterogeneous face recognition (HFR) deals with recognition problems in which training/gallery images are collected in terms of one modality (e.g., photos), while test/probe images are observed in the other (e.g., sketches). In this paper, we present a domain adaptation approach for solving HFR problems. By utilizing external face images (i.e., those collected from the subjects not of interest) from both source and target domains, we propose a novel Domain-independent Component Analysis (DiCA) algorithm for deriving a common subspace for relating and representing cross-domain image data. In order to introduce improved representation ability, we further advance the self-taught learning strategy for learning a domain-independent dictionary in our DiCA subspace, which can be applied to both gallery and probe images of interest to improve representation and recognition. Different from some prior domain-adaptation approaches, we do not require the data correspondences (i.e., data pairs) when collecting external cross-domain image data, nor the label information is needed for learning the common feature space when associating different domains. Thus, our method is practical for real-world cross-domain classification problems. In our experiments, we consider sketch-to-photo and near-infrared (NIR) to visible spectrum (VIS) face recognition problems for evaluating the performance of our proposed approach. Cheng-An Hou, Min-Chun Yang, Yu-Chiang Frank Wang |
ICPR | 3 |
| 2014 | Transfer in Photography CompositionabstractIn this paper, we propose novel photography recomposition method, which aims at transferring the photography composition of a reference image to an input image automatically. Without any user interaction, our approach first identifies the salient foreground objects or image regions of interest, and the recomposition is performed by solving a graph-matching based optimization task. With additional post-processing step to preserve the locality and boundary information of the recomposed visual components, we can solve the task of photography recomposition without the uses of any prior knowledge on photography or predetermined image aesthetics rules. Experiments on a variety of images, including transferring the photography composition from real photos, sketches or even paintings, would confirm the effectiveness of our proposed method. Hui-Tang Chang, Yu-Chiang Frank Wang, Ming-Syan Chen |
ACM Multimedia | 2 |
| 2014 | Robust Face Recognition With Structurally Incoherent Low-Rank Matrix DecompositionabstractFor the task of robust face recognition, we particularly focus on the scenario in which training and test image data are corrupted due to occlusion or disguise. Prior standard face recognition methods like Eigenfaces or state-of-the-art approaches such as sparse representation-based classification did not consider possible contamination of data during training, and thus their recognition performance on corrupted test data would be degraded. In this paper, we propose a novel face recognition algorithm based on low-rank matrix decomposition to address the aforementioned problem. Besides the capability of decomposing raw training data into a set of representative bases for better modeling the face images, we introduce a constraint of structural incoherence into the proposed algorithm, which enforces the bases learned for different classes to be as independent as possible. As a result, additional discriminating ability is added to the derived base matrices for improved recognition performance. Experimental results on different face databases with a variety of variations verify the effectiveness and robustness of our proposed method. Chia-Po Wei, Chih-Fan Chen, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 3 |
| 2014 | Heterogeneous Domain Adaptation and Classification by Exploiting the Correlation SubspaceabstractWe present a novel domain adaptation approach for solving cross-domain pattern recognition problems, i.e., the data or features to be processed and recognized are collected from different domains of interest. Inspired by canonical correlation analysis (CCA), we utilize the derived correlation subspace as a joint representation for associating data across different domains, and we advance reduced kernel techniques for kernel CCA (KCCA) if nonlinear correlation subspace are desirable. Such techniques not only makes KCCA computationally more efficient, potential over-fitting problems can be alleviated as well. Instead of directly performing recognition in the derived CCA subspace (as prior CCA-based domain adaptation methods did), we advocate the exploitation of domain transfer ability in this subspace, in which each dimension has a unique capability in associating cross-domain data. In particular, we propose a novel support vector machine (SVM) with a correlation regularizer, named correlation-transfer SVM, which incorporates the domain adaptation ability into classifier design for cross-domain recognition. We show that our proposed domain adaptation and classification approach can be successfully applied to a variety of cross-domain recognition tasks such as cross-view action recognition, handwritten digit recognition with different features, and image-to-text or text-to-image classification. From our empirical results, we verify that our proposed method outperforms state-of-the-art domain adaptation approaches in terms of recognition performance. Yi-Ren Yeh, Chun-Hao Huang, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 3 |
| 2014 | Self-Learning Based Image Decomposition With Applications to Single Image DenoisingabstractDecomposition of an image into multiple semantic components has been an effective research topic for various image processing applications such as image denoising, enhancement, and inpainting. In this paper, we present a novel self-learning based image decomposition framework. Based on the recent success of sparse representation, the proposed framework first learns an over-complete dictionary from the high spatial frequency parts of the input image for reconstruction purposes. We perform unsupervised clustering on the observed dictionary atoms (and their corresponding reconstructed image versions) via affinity propagation, which allows us to identify image-dependent components with similar context information. While applying the proposed method for the applications of image denoising, we are able to automatically determine the undesirable patterns (e.g., rain streaks or Gaussian noise) from the derived image components directly from the input image, so that the task of single-image denoising can be addressed. Different from prior image processing works with sparse representation, our method does not need to collect training image data in advance, nor do we assume image priors such as the relationship between input and output image dictionaries. We conduct experiments on two denoising problems: single-image denoising with Gaussian noise and rain removal. Our empirical results confirm the effectiveness and robustness of our approach, which is shown to outperform state-of-the-art image denoising algorithms. De-An Huang, Li-Wei Kang, Yu-Chiang Frank Wang, Chia-Wen Lin |
IEEE Trans. Multim. | 3 |
| 2013 | Depth map super-resolution via Markov Random Fields without texture-copying artifactsabstractThe use of time-of-flight sensors enables the record of full-frame depth maps at video frame rate, which benefits a variety of 3D image or video processing applications. However, such depth maps are typically corrupted by noise and with limited resolution. In this paper, we present a learning-based depth map super-resolution framework by solving a MRF labeling optimization problem. With the captured depth map and the associated high-resolution color image, our proposed method exhibits the capability of preserving the edges of range data while suppressing the artifacts of texture copying due to color discontinuities. Quantitative and qualitative experimental results demonstrate the effectiveness and robustness of our approach over prior depth map upsampling works. Kai-Han Lo, Kai-Lung Hua, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2013 | Coupled Dictionary and Feature Space Learning with Applications to Cross-Domain Image Synthesis and RecognitionabstractCross-domain image synthesis and recognition are typically considered as two distinct tasks in the areas of computer vision and pattern recognition. Therefore, it is not clear whether approaches addressing one task can be easily generalized or extended for solving the other. In this paper, we propose a unified model for coupled dictionary and feature space learning. The proposed learning model not only observes a common feature space for associating cross-domain image data for recognition purposes, the derived feature space is able to jointly update the dictionaries in each image domain for improved representation. This is why our method can be applied to both cross-domain image synthesis and recognition problems. Experiments on a variety of synthesis and recognition tasks such as single image super-resolution, cross-view action recognition, and sketch-to-photo face recognition would verify the effectiveness of our proposed learning model. De-An Huang, Yu-Chiang Frank Wang |
ICCV | 2 |
| 2013 | Superpixel-based large displacement optical flowabstractIt has been a challenging task to estimate optical flow for videos in which either foreground or background exhibits remarkable motion information (i.e., large displacement), or those with insufficient resolution due to artifacts like motion blur or noise. We present a novel optical flow algorithm, which approaches the above problem as solving the task of energy minimization, which exploits image data and smoothness terms at the superpixel level. Our proposed method can be considered as an extended mean-shift algorithm, which advances color and gradient information of superpixels across consecutive frames with smoothness guarantees. Since we do not require assumptions of linearlization during optimization (as standard optical flow approaches do), we are able to alleviate local minimum problems and thus produce improved estimation results. Empirical results on the MPI-Sintel video dataset verify the effectiveness of our proposed method. Haw-Shiuan Chang, Yu-Chiang Frank Wang |
ICIP | 2 |
| 2013 | A discriminative domain adaptation model for cross-domain image classificationabstractTechniques of domain adaptation have been applied to address cross-domain recognition problems. In particular, such techniques favor the scenarios in which labeled data can be obtained at the source domain, but only few labeled target domain data are available during the training stage. In this paper, we propose a domain adaptation approach which is able to transfer source domain labeled data to the target domain, so that one can collect a sufficient amount of training data at that domain for recognition purposes. By advancing low-rank matrix decomposition for obtaining representative cross-domain data, our proposed model aims at transferring source domain labeled data to the target domain while preserving class label information. This introduces additional discriminating ability into our model, and thus improved recognition can be expected. Empirical results on cross-domain image datasets confirm the use of our proposed model for solving cross-domain recognition problems. Yen-Cheng Chou, Chia-Po Wei, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2013 | Cross-view action recognition via low-rank based domain adaptationabstractCross-view action recognition is a challenging problem, since one typically does not have sufficient training data at the target view of interest. With recent developments of domain adaptation, we propose a novel low-rank based domain adaptation model for mapping labeled data from the original source view to the target view, so that training and testing can be performed at that domain. Our model not only provides an effective way for associating image data across different domains, we further advocate the structural incoherence between transformed data of different categories. As a result, additional data discriminating ability is introduced to our domain adaptation model, and thus improved recognition can be expected. Experimental results on the IXMAS dataset verify the effectiveness of our proposed method, which is shown to outperform state-of-the-art domain adaptation approaches. Wen-Sheng Tseng, Lun-Kai Hsu, Li-Wei Kang, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2013 | Seeing through the expression: Bridging the gap between expression and emotion recognitionabstractIn this paper, we propose a novel approach for visualizing and recognizing different emotion categories using facial expression images. Extended by the unsupervised nonlinear dimension reduction technique of locally linear embedding (LLE), we propose a supervised LLE (sLLE) algorithm utilizing emotion labels of face expression images. While existing works typically aim at training on such labeled data for emotion recognition, our approach allows one to derive subspaces for visualizing facial expression images within and across different emotion categories, and thus emotion recognition can be properly performed. In our work, we relate the resulting two-dimensional subspace to the valence-arousal emotion space, in which our method is observed to automatically identify and discriminate emotions in different degrees. Experimental results on two facial emotion datasets verify the effectiveness of our algorithm. With reduced numbers of feature dimensions (2D or beyond), our approach is shown to achieve promising emotion recognition performance. Lun-Kai Hsu, Wen-Sheng Tseng, Li-Wei Kang, Yu-Chiang Frank Wang |
ICME | 4 |
| 2013 | Learning auxiliary dictionaries for undersampled face recognitionabstractIn this paper, we address the problem of robust face recognition using undersampled data. Given only one or few face images per class, our proposed method not only handles test images with large intra-class variations such as illumination and expression, it is also able to recognize the corrupted ones due to occlusion or disguise. In our work, we advocate the learning of auxiliary dictionaries from the subjects not of interest. With the proposed optimization algorithm which jointly solves the tasks of auxiliary dictionary learning and sparsere-presentation based face recognition, our approach is able to model the above intra-class variations and corruptions for improved recognition. Our experiments on two face image datasets confirm the effectiveness and robustness of our approach, which is shown to outperform state-of-the-art sparse representation based methods. Chia-Po Wei, Yu-Chiang Frank Wang |
ICME | 2 |
| 2013 | With one look: robust face recognition using single sample per personabstractIn this paper, we address the problem of robust face recognition using single sample per person. Given only one training image per subject of interest, our proposed method is able to recognize query images with illumination or expression changes, or even the corrupted ones due to occlusion. In order to model the above intra-class variations, we advocate the use of external data (i.e., images of subjects not of interest) for learning an exemplar-based dictionary. This dictionary provides auxiliary yet representative information for handling intra-class variation, while the gallery set containing one training image per class preserves separation between different subjects for recognition purposes. Our experiments on two face datasets confirm the effectiveness and robustness of our approach, which is shown to outperform state-of-the-art sparse representation based methods. De-An Huang, Yu-Chiang Frank Wang |
ACM Multimedia | 2 |
| 2013 | Joint trilateral filtering for depth map super-resolutionabstractDepth map super-resolution is an emerging topic due to the increasing needs and applications using RGB-D sensors. Together with the color image, the corresponding range data provides additional information and makes visual analysis tasks more tractable. However, since the depth maps captured by such sensors are typically with limited resolution, it is preferable to enhance its resolution for improved recognition. In this paper, we present a novel joint trilateral filtering (JTF) algorithm for solving depth map super-resolution (SR) problems. Inspired by bilateral filtering, our JTF utilizes and preserves edge information from the associated high-resolution (HR) image by taking spatial and range information of local pixels. Our proposed further integrates local gradient information of the depth map when synthesizing its HR output, which alleviates textural artifacts like edge discontinuities. Quantitative and qualitative experimental results demonstrate the effectiveness and robustness of our approach over prior depth map upsampling works. Kai-Han Lo, Yu-Chiang Frank Wang, Kai-Lung Hua |
VCIP | 2 |
| 2013 | Moving foreground object detection via robust SIFT trajectories
Shih-Wei Sun, Yu-Chiang Frank Wang, Fay Huang, Hong-Yuan Mark Liao |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Locality-sensitive dictionary learning for sparse representation based classification
Chia-Po Wei, Yu-Wei Chao, Yi-Ren Yeh, Yu-Chiang Frank Wang |
Pattern Recognit. | 4 |
| 2013 | A rank-one update method for least squares linear discriminant analysis with concept drift
Yi-Ren Yeh, Yu-Chiang Frank Wang |
Pattern Recognit. | 2 |
| 2013 | Exploring Visual and Motion Saliency for Automatic Video Object ExtractionabstractThis paper presents a saliency-based video object extraction (VOE) framework. The proposed framework aims to automatically extract foreground objects of interest without any user interaction or the use of any training data (i.e., not limited to any particular type of object). To separate foreground and background regions within and across video frames, the proposed method utilizes visual and motion saliency information extracted from the input video. A conditional random field is applied to effectively combine the saliency induced features, which allows us to deal with unknown pose and scale variations of the foreground object (and its articulated parts). Based on the ability to preserve both spatial continuity and temporal consistency in the proposed VOE framework, experiments on a variety of videos verify that our method is able to produce quantitatively and qualitatively satisfactory VOE results. Wei-Te Li, Haw-Shiuan Chang, Kuo-Chin Lien, Hui-Tang Chang, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 5 |
| 2013 | Anomaly Detection via Online Oversampling Principal Component AnalysisabstractAnomaly detection has been an important research topic in data mining and machine learning. Many real-world applications such as intrusion or credit card fraud detection require an effective and efficient framework to identify deviated data instances. However, most anomaly detection methods are typically implemented in batch mode, and thus cannot be easily extended to large-scale problems without sacrificing computation and memory requirements. In this paper, we propose an online oversampling principal component analysis (osPCA) algorithm to address this problem, and we aim at detecting the presence of outliers from a large amount of data via an online updating technique. Unlike prior principal component analysis (PCA)-based approaches, we do not store the entire data matrix or covariance matrix, and thus our approach is especially of interest in online or large-scale problems. By oversampling the target instance and extracting the principal direction of the data, the proposed osPCA allows us to determine the anomaly of the target instance according to the variation of the resulting dominant eigenvector. Since our osPCA need not perform eigen analysis explicitly, the proposed framework is favored for online applications which have computation or memory limitations. Compared with the well-known power method for PCA and other popular anomaly detection algorithms, our experimental results verify the feasibility of our proposed method in terms of both accuracy and efficiency. Yuh-Jye Lee, Yi-Ren Yeh, Yu-Chiang Frank Wang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Robust Texture Analysis Using Multi-Resolution Gray-Scale Invariant Features for Breast Sonographic Tumor DiagnosisabstractComputer-aided diagnosis (CAD) systems in gray-scale breast ultrasound images have the potential to reduce unnecessary biopsy of breast masses. The purpose of our study is to develop a robust CAD system based on the texture analysis. First, gray-scale invariant features are extracted from ultrasound images via multi-resolution ranklet transform. Thus, one can apply linear support vector machines (SVMs) on the resulting gray-level co-occurrence matrix (GLCM)-based texture features for discriminating the benign and malignant masses. To verify the effectiveness and robustness of the proposed texture analysis, breast ultrasound images obtained from three different platforms are evaluated based on cross-platform training/testing and leave-one-out cross-validation (LOO-CV) schemes. We compare our proposed features with those extracted by wavelet transform in terms of receiver operating characteristic (ROC) analysis. The AUC values derived from the area under the curve for the three databases via ranklet transform are 0.918 (95% confidence interval [CI], 0.848 to 0.961), 0.943 (95% CI, 0.906 to 0.968), and 0.934 (95% CI, 0.883 to 0.961), respectively, while those via wavelet transform are 0.847 (95% CI, 0.762 to 0.910), 0.922 (95% CI, 0.878 to 0.958), and 0.867 (95% CI, 0.798 to 0.914), respectively. Experiments with cross-platform training/testing scheme between each database reveal that the diagnostic performance of our texture analysis using ranklet transform is less sensitive to the sonographic ultrasound platforms. Also, we adopt several co-occurrence statistics in terms of quantization levels and orientations (i.e., descriptor settings) for computing the co-occurrence matrices with 0.632+ bootstrap estimators to verify the use of the proposed texture analysis. These experiments suggest that the texture analysis using multi-resolution gray-scale invariant features via ranklet transform is useful for designing a robust CAD system. Min-Chun Yang, Woo Kyung Moon, Yu-Chiang Frank Wang, Min Sun Bae, Chiun-Sheng Huang, Jeon-Hor Chen, Ruey-Feng Chang |
IEEE Trans. Medical Imaging | 3 |
| 2013 | A Self-Learning Approach to Single Image Super-ResolutionabstractLearning-based approaches for image super-resolution (SR) have attracted the attention from researchers in the past few years. In this paper, we present a novel self-learning approach for SR. In our proposed framework, we advance support vector regression (SVR) with image sparse representation, which offers excellent generalization in modeling the relationship between images and their associated SR versions. Unlike most prior SR methods, our proposed framework does not require the collection of training low and high-resolution image data in advance, and we do not assume the reoccurrence (or self-similarity) of image patches within an image or across image scales. With theoretical supports of Bayes decision theory, we verify that our SR framework learns and selects the optimal SVR model when producing an SR image, which results in the minimum SR reconstruction error. We evaluate our method on a variety of images, and obtain very promising SR results. In most cases, our method quantitatively and qualitatively outperforms bicubic interpolation and state-of-the-art learning-based SR approaches. Min-Chun Yang, Yu-Chiang Frank Wang |
IEEE Trans. Multim. | 2 |
| 2012 | Low-rank matrix recovery with structural incoherence for robust face recognitionabstractWe address the problem of robust face recognition, in which both training and test image data might be corrupted due to occlusion and disguise. From standard face recognition algorithms such as Eigenfaces to recently proposed sparse representation-based classification (SRC) methods, most prior works did not consider possible contamination of data during training, and thus the associated performance might be degraded. Based on the recent success of low-rank matrix recovery, we propose a novel low-rank matrix approximation algorithm with structural incoherence for robust face recognition. Our method not only decomposes raw training data into a set of representative basis with corresponding sparse errors for better modeling the face images, we further advocate the structural incoherence between the basis learned from different classes. These basis are encouraged to be as independent as possible due to the regularization on structural incoherence. We show that this provides additional discriminating ability to the original low-rank models for improved performance. Experimental results on public face databases verify the effectiveness and robustness of our method, which is also shown to outperform state-of-the-art SRC based approaches. Chih-Fan Chen, Chia-Po Wei, Yu-Chiang Frank Wang |
CVPR | 3 |
| 2012 | Self-learning approach to color demosaicking via support vector regressionabstractMost digital cameras capture one primary color at each pixel by a single sensor overlaid with a color filter array. To recover a full color image from incomplete color samples, one needs to restore the two missing color values for each pixel. This restoration process is known as color demosaicking. In this paper, we present a novel self-learning approach to this problem via support vector regression. Unlike prior learning-based demosaicking methods, our approach aims at extracting image-dependent information in constructing the learning model, and we do not require any additional training data. Experimental results show that our proposed method outperforms many state-of-the-art techniques in both subjective and objective image quality measures. Fang-Lin He, Yu-Chiang Frank Wang, Kai-Lung Hua |
ICIP | 2 |
| 2012 | Automatic saliency inspired foreground object extraction from videosabstractIn this paper, we propose a saliency inspired video object extraction (VOE) method to extract and segment foreground objects of interest from videos captured by freely moving cameras. Our method aims at detecting visual and motion salient regions from an input video, and thus we integrate such cosaliency information with the associated foreground and background color models to achieve VOE. A conditional random field (CRF) is applied in our framework to automatically identify the foreground object regions based on the above features, while our method does not need any prior knowledge on the foreground objects of interest or any interaction from the users. Experiments on a variety of videos confirm that our method is able to provide quantitatively and qualitatively more satisfactory results when comparing to state-of-the-art VOE approaches. Wei-Te Li, Hui-Tang Chang, Hermes Shing Lyu, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2012 | Context-Aware Single Image Rain RemovalabstractRain removal from a single image is one of the challenging image denoising problems. In this paper, we present a learning-based framework for single image rain removal, which focuses on the learning of context information from an input image, and thus the rain patterns present in it can be automatically identified and removed. We approach the single image rain removal problem as the integration of image decomposition and self-learning processes. More precisely, our method first performs context-constrained image segmentation on the input image, and we learn dictionaries for the high-frequency components in different context categories via sparse coding for reconstruction purposes. For image regions with rain streaks, dictionaries of distinct context categories will share common atoms which correspond to the rain patterns. By utilizing PCA and SVM classifiers on the learned dictionaries, our framework aims at automatically identifying the common rain patterns present in them, and thus we can remove rain streaks as particular high-frequency components from the input image. Different from prior works on rain removal from images/videos which require image priors or training image data from multiple frames, our proposed self-learning approach only requires the input image itself, which would save much pre-training effort. Experimental results demonstrate the subjective and objective visual quality improvement with our proposed method. De-An Huang, Li-Wei Kang, Min-Chun Yang, Chia-Wen Lin, Yu-Chiang Frank Wang |
ICME | 5 |
| 2012 | Self-Learning of Edge-Preserving Single Image Super-Resolution via Contourlet TransformabstractWe present a self-learning approach for single image super-resolution (SR), with the ability to preserve high frequency components such as edges in resulting high resolution (HR) images. Given a low-resolution (LR) input image, we construct its image pyramid and produce a super pixel dataset. By extracting context information from the super-pixels, we propose to deploy context-specific contour let transform on them in order to model the relationship (via support vector regression) between the input patches and their associated directional high-frequency responses. These learned models are applied to predict the SR output with satisfactory quality. Unlike prior learning-based SR methods, our approach advances a self-learning technique and does not require the self similarity of image patches within or across image scales. More importantly, we do not need to collect training LR/HR image data in advance and only require a single LR input image. Empirical results verify the effectiveness of our approach, which quantitatively and qualitatively outperforms existing interpolation or learning-based SR methods. Min-Chun Yang, De-An Huang, Chih-Yun Tsai, Yu-Chiang Frank Wang |
ICME | 4 |
| 2012 | Context-aware single image super-resolution using locality-constrained group sparse representationabstractWe present a novel learning-based method for single image super-resolution (SR). Given a single input low-resolution (LR) image (and its image pyramid), we propose to learn context-specific image sparse representation, which aims at modeling the relationship between low and high-resolution image patch pairs of different context categories in terms of the learned dictionaries. To predict the SR image, we derive the context-specific sparse representation of each image patch in the LR input with additional locality and group sparsity constraints. While the locality constraint searches for the most similar image patches and uses the corresponding highresolution outputs for SR, the group sparsity constraint allows us to utilize the information from most relevant context categories for predicting the final SR output. Experimental results show the proposed method is able to quantitatively and qualitatively achieve state-of-the-art performance. Chih-Yun Tsai, De-An Huang, Min-Chun Yang, Li-Wei Kang, Yu-Chiang Frank Wang |
VCIP | 5 |
| 2012 | A Novel Multiple Kernel Learning Framework for Heterogeneous Feature Fusion and Variable SelectionabstractWe propose a novel multiple kernel learning (MKL) algorithm with a group lasso regularizer, called group lasso regularized MKL (GL-MKL), for heterogeneous feature fusion and variable selection. For problems of feature fusion, assigning a group of base kernels for each feature type in an MKL framework provides a robust way in fitting data extracted from different feature domains. Adding a mixed norm constraint (i.e., group lasso) as the regularizer, we can enforce the sparsity at the group/feature level and automatically learn a compact feature set for recognition purposes. More precisely, our GL-MKL determines the optimal base kernels, including the associated weights and kernel parameters, and results in improved recognition performance. Besides, our GL-MKL can also be extended to address heterogeneous variable selection problems. For such problems, we aim to select a compact set of variables (i.e., feature attributes) for comparable or improved performance. Our proposed method does not need to exhaustively search for the entire variable space like prior sequential-based variable selection methods did, and we do not require any prior knowledge on the optimal size of the variable subset either. To verify the effectiveness and robustness of our GL-MKL, we conduct experiments on video and image datasets for heterogeneous feature fusion, and perform variable selection on various UCI datasets. Yi-Ren Yeh, Ting-Chu Lin, Yung-Yu Chung, Yu-Chiang Frank Wang |
IEEE Trans. Multim. | 4 |
| 2011 | Locality-constrained group sparse representation for robust face recognitionabstractThis paper presents a novel sparse representation for robust face recognition. We advance both group sparsity and data locality and formulate a unified optimization framework, which produces a locality and group sensitive sparse representation (LGSR) for improved recognition. Empirical results confirm that our LGSR not only outperforms state-of-the-art sparse coding based image classification methods, our approach is robust to variations such as lighting, pose, and facial details (glasses or not), which are typically seen in real-world face recognition problems. Yu-Wei Chao, Yi-Ren Yeh, Yuh-Jye Lee, Yu-Chiang Frank Wang |
ICIP | 5 |
| 2011 | Joint source-channel coding optimization with packet loss resilience for video transmissionabstractWe propose a novel joint source-channel coding (JSCC) framework that jointly optimizes encoding modes of macroblocks and unequal error protection (UEP) of packets for error-resilient video transmission, and address the problem of source packet loss due to bit errors incurred in error-prone channels. In the proposed framework, we consider the source packet loss probability as a function of the encoding configuration of a slice. The task of optimization is complicated by the inherent interdependency between macroblocks in the same slice, since the encoding mode optimization for each macroblock requires the encoding configuration of entire slice to be known in advance. We resolve such interdependency via an iterative method that alternates between the optimization of the slice size and the associated encoding modes, and we also utilize adaptive quantization for coding residual for source rate control in different channel conditions. Simulation results show that our framework outperforms existing approaches. Ching-Hui Chen, Wei-Ho Chung, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2011 | Cross-layer design for video streaming with dynamic antenna selectionabstractSpatial multiplexing is an efficient transmission technique for video streaming in multiple-input multiple-output (MIMO) wireless communication links due to its capability to support the high transmit rate. Nevertheless, the video quality is sensitive to packet losses resulting from channel fading and co-interference between transmit antennas in spatial multiplexing systems. Antenna selection techniques have been investigated to balance the spatial multiplexing gain and the link reliability. Herein, a novel cross-layer framework using dynamic antenna selection is proposed to control the multiplexing-diversity tradeoff with joint error-resilient source and channel coding. The cross-layer optimization is performed to optimize the physical, link, and application layers, and to minimize the end-to-end video distortion. Ching-Hui Chen, Wei-Ho Chung, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2011 | Learning context-aware sparse representation for single image super-resolutionabstractThis paper presents a novel learning-based method for single image super-resolution (SR). Given an input low-resolution image and its image pyramid, we propose to perform context-constrained image segmentation and construct an image segment dataset with different context categories. By learning context-specific image sparse representation, our method aims to model the relationship between the interpolated image patches and their ground truth pixel values from different context categories via support vector regression (SVR). To synthesize the final SR output, we upsample the input image by bicubic interpolation, followed by the refinement of each image patch using the SVR model learned from the associated context category. Unlike prior learning-based SR methods, our approach does not require the reoccurrence of similar image patches (within or across image scales), and we do not need to collect training low and high-resolution image data in advance either. Empirical results show that our proposed method is quantitatively and qualitatively more effective than existing interpolation or learning-based SR approaches. Ming-Chun Yang, Chang-Heng Wang, Ting-Yao Hu, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2011 | Automatic object extraction in single-concept videosabstractWe propose a motion-driven video object extraction (VOE) method, which is able to model and segment foreground objects in single-concept videos, i.e. videos which have only one object category of interest but may have multiple object instances with pose, scale, etc. variations. Given such a video, we construct a compact shape model induced by motion cues, and extract the foreground and background color information accordingly. We integrate these feature models into a unified framework via a conditional random field (CRF), and this CRF can be applied to video object segmentation and further video editing and retrieval applications. One of the advantages of our method is that we do not require the prior knowledge of the object of interest, and thus no training data or predetermined object detectors are needed; this makes our approach robust and practical to real-world problems. Very attractive empirical results on a variety of videos with highly articulated objects support the feasibility of our proposed method. Kuo-Chin Lien, Yu-Chiang Frank Wang |
ICME | 2 |
| 2011 | Automatic annotation of Web videosabstractMost Web videos are captured in uncontrolled environments (e.g. videos captured by freely-moving cameras with low resolution); this makes automatic video annotation very difficult. To address this problem, we present a robust moving foreground object detection method followed by the integration of features collected from heterogeneous domains. We advance SIFT feature matching and present a probabilistic framework to construct consensus foreground object templates (CFOT). The CFOT can detect moving foreground objects of interest across video frames, and this allows us to extract visual features from foreground regions of interest. Together with the use of audio features, we are able to improve resulting annotation accuracy. We conduct experiments and achieve promising results on a Web video dataset collected from YouTube. Shih-Wei Sun, Yu-Chiang Frank Wang, Yao-Ling Hung, Chia-Ling Chang, Kuan-Chieh Chen, Shih-Sian Cheng, Hsin-Min Wang, Hong-Yuan Mark Liao |
ICME | 2 |
| 2011 | Group lasso regularized multiple kernel learning for heterogeneous feature selectionabstractWe propose a novel multiple kernel learning (MKL) algorithm with a group lasso regularizer, called group lasso regularized MKL (GL-MKL), for heterogeneous feature selection. We extend the existing MKL algorithm and impose a mixed ℓ1and ℓ2norm constraint (known as group lasso) as the regularizer. Our GL-MKL determines the optimal base kernels, including the associated weights and kernel parameters, and results in a compact set of features for comparable or improved recognition performance. The use of our GL-MKL avoids the problem of choosing the proper technique to normalize the feature attributes collected from heterogeneous domains (and thus with different properties and distribution ranges). Our approach does not need to exhaustively search for the entire feature space when performing feature selection like prior sequential-based feature selection methods did, and we do not require any prior knowledge on the optimal size of the feature subset either. Comparisons with existing MKL or sequential-based feature selection methods on a variety of datasets confirm the effectiveness of our method in selecting a compact feature subset for comparable or improved classification performance. Yi-Ren Yeh, Yung-Yu Chung, Ting-Chu Lin, Yu-Chiang Frank Wang |
IJCNN | 4 |
| 2011 | Exploring self-similarities of bag-of-features for image classificationabstractThe use of bag-of-features (BOF) models has been a popular technique for image classification and retrieval. In order to better represent and discriminate images from different classes, we advance BOF and explore the self-similarities of visual words for improved performance. The proposed self-similarity hypercubes (SSH) model, which observes the concurrent occurrences of visual words in an image, is able to describe the structural information of the BOF in an image. Our experiments confirm that our SSH provides additional and complementary information to BOF and thus results in improved classification performance. Unlike most prior methods requiring extraction or integration of multiple types of features for similar improvements, our SSH works in the same domain as the BOF does. Moreover, we do not limit the use of our SSH to any particular type of image descriptors, and its generalization is also verified. Chih-Fan Chen, Yu-Chiang Frank Wang |
ACM Multimedia | 2 |
| 2011 | Learning of context-aware single image super-resolutionabstractWe propose a novel learning-based method for single image super-resolution (SR). Given a low-resolution input image and its image pyramid, we advance a context-constrained image segmentation to construct a super-pixel database with different context categories for learning purposes. By utilizing context-specific image sparse representation, our method aims at modeling the relationship between the interpolated image patches and their ground truth pixels from different context categories via support vector regression (SVR). To produce the final SR output, we upsample the low-resolution input, followed by the refinement of each image patch using the SVR models observed from the associated context categories. Unlike prior learning-based SR methods, our approach advances a self-learning technique and does not assume the reoccurrence of image patches (within or across image scales). We do not need to collect training low/high-resolution image data in advance either. Empirical results verify the effectiveness of our SR approach, which quantitatively and qualitatively outperforms existing interpolation or learning-based SR methods in most cases. Ming-Chun Yang, Ting-Yao Hu, Chang-Heng Wang, Yu-Chiang Frank Wang |
VCIP | 4 |
| 2010 | A Multi-Scale Learning Framework for Visual Categorization
Shao-Chuan Wang 0001, Yu-Chiang Frank Wang |
ACCV (1) | 2 |
| 2010 | Learning Dense Optical-Flow Trajectory Patterns for Video Object ExtractionabstractWe proposes an unsupervised method to address video object extraction (VOE) in uncontrolled videos, i.e. videos captured by low-resolution and freely moving cameras. We advocate the use of dense optical-flow trajectories (DOTs), which are obtained by propagating the optical flow information at the pixel level. Therefore, no interest point extraction is required in our framework. To integrate color and and shape information of moving objects, we group the DOTs at the super-pixel level to extract co-motion regions, and use the associated pyramid histogram of oriented gradients (PHOG) descriptors to extract objects of interest across video frames. Our approach for VOE is easy to implement, and the use of DOTs for both motion segmentation and object tracking is more robust than existing trajectory-based methods. Experiments on several video sequences exhibit the feasibility of our proposed VOE framework. Wang-Chou Lu, Yu-Chiang Frank Wang, Chu-Song Chen |
AVSS | 2 |
| 2010 | Simultaneous Object Recognition and Localization in Image CollectionsabstractThis papers presents a weakly supervised method to simultaneously address object localization and recognition problems. Unlike prior work using exhaustive search methods such as sliding windows, we propose to learn category and image-specific visual words in image collections by extracting discriminating feature information via two different types of support vector machines: the standard L2-regularized L1-loss SVM, and the one with L1 regularization and L2 loss. The selected visual words are used to construct visual attention maps, which provide descriptive information for each object category. To preserve local spatial information, we further refine these maps by Gaussian smoothing and cross bilateral filtering, and thus both appearance and spatial information can be utilized for visual categorization applications. Our method is not limited to any specific type of image descriptors, or any particular codebook learning and feature encoding techniques. In this paper, we conduct preliminary experiments on a subset of the Caltech-256 dataset using bag-of-feature (BOF) models with SIFT descriptors. We show that the use of our visual attention maps improves the recognition performance, while the one selected by L1-regularized L2-loss SVMs exhibits the best recognition and localization results. Shao-Chuan Wang 0001, Yu-Chiang Frank Wang |
AVSS | 2 |
| 2010 | Learning sparse image representation with support vector regression for single-image super-resolutionabstractLearning-based approaches for super-resolution (SR) have been studied in the past few years. In this paper, a novel single-image SR framework based on the learning of sparse image representation with support vector regression (SVR) is presented. SVR is known to offer excellent generalization ability in predicting output class labels for input data. Given a low resolution image, we approach the SR problem as the estimation of pixel labels in its high resolution version. The feature considered in this work is the sparse representation of different types of image patches. Prior studies have shown that this feature is robust to noise and occlusions present in image data. Experimental results show that our method is quantitatively more effective than prior work using bicubic interpolation or SVR methods, and our computation time is significantly less than that of existing SVR-based methods due to the use of sparse image representations. Ming-Chun Yang, Chao-Tsung Chu, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2010 | Segmentation-free object localization in image collectionsabstractWe propose a novel method to address object localization in a weakly supervised framework. Unlike prior work using exhaustive search methods such as sliding windows, we advocate the use of visual attention maps which are constructed by class-specific visual words. Based on dense SIFT descriptors, these visual words are selected by support vector machines and feature ranking techniques. Therefore, discriminative information is learned and embedded in these visual words. We further refine the constructed map by Gaussian smoothing and cross bilateral filtering to preserve local spatial information of the objects. Very promising localization results are reported on a subset of the Caltech-256 dataset, and our method is shown to improve the state-of-the-art recognition performance using the bag-of-feature (BOF) model. Shao-Chuan Wang 0001, Yu-Chiang Frank Wang |
ICME | 2 |
| 2009 | A support vector hierarchical method for multi-class classification and rejectionabstractWe address both recognition of true classes and rejection of unseen false classes inputs, as occurs in many realistic pattern recognition problems. we advance a hierarchical binary-decision classifier and produce analog outputs at each node, with yields a new soft-decision hierarchical is designed by our new support vector clustering method, which selects the classes to be separated at each node in the hierarchy. Use of our SVRDM (support vector representation and discrimination machine) classifiers at each node provides generalization and rejection ability. The soft-decision SVRDM output allows use of the confidence score for each class at each node; this is shown to improve classification (for true classes) and rejection (for false classes) performance. New aspects of this paper are that we provide remarks on our hierarchical design method, including our hierarchical clustering rule, and discuss the meaning and the use of probabilities in our soft-decision hierarchical SVRDM classifiers. We also provide initial tests results on a new database (COIL) that allows large class problem to be addressed. No prior work considered rejection of false classes on this database. Yu-Chiang Frank Wang, David P. Casasent |
IJCNN | 1 |
| 2008 | Soft-decision hierarchical classification using SVM-type classifiersabstractIn this paper, we address both recognition of true object classes and rejection of false (non-object) classes as occurs in many realistic pattern recognition problems. We modified our hierarchical binary-decision classifier to produce analog outputs at each node, with values proportional to the class conditional probabilities at that node. This yields a new soft-decision hierarchical system. The hierarchical classification structure is designed by our weighted support vector k-means clustering method, which selects the classes to be separated at each node in the hierarchy. Use of our SVRDM (support vector representation and discrimination machine) classifiers at each node provides generalization and rejection ability. Compared to the standard SVM, use of the Gaussian kernel function and a looser constraint in the classifier design give our SVRDM an improved rejection ability. The soft-decision SVRDM output allows us to use the confidence level of each class to improve the classification (for true class inputs) and rejection (for false class inputs) performance of the hierarchical classifier. False class rejection is a major new aspect of this work. It is not present in most prior work. Excellent test results on a real infra-red (IR) database are presented. Yu-Chiang Frank Wang, David P. Casasent |
IJCNN | 1 |
| 2008 | New support vector-based design method for binary hierarchical classifiers for multi-class classification problems
Yu-Chiang Frank Wang, David P. Casasent |
Neural Networks | 1 |
| 2007 | New Weighted Support Vector K-means Clustering for Hierarchical Multi-class ClassificationabstractWe propose a binary hierarchical classification structure to address the multi-class classification problem with a new hierarchical design method, weighted support vector k-means clustering, which automatically separates a set of classes into two smaller groups at each node in the hierarchy. This method is able to visualize and cluster high-dimensional support vector data; therefore, it greatly improves upon prior hierarchical classifier design. At each node in the hierarchy, we apply an SVRDM (support vector representation and discrimination machine) classifier, which offers generalization and good rejection of unseen false objects, which is not achieved by the standard SVM classifier. We provide a new theoretical basis for the good SVRDM rejection obtained, due to its looser constrained optimization problem, compared to that of an SVM. New classification and rejection test results are presented on a real IR (infra-red) database. Yu-Chiang Frank Wang, David P. Casasent |
IJCNN | 1 |
| 2006 | Hierarchical K-means Clustering Using New Support Vector Machines for Multi-class ClassificationabstractWe propose a binary hierarchical classification structure to address the multi-class classification problem with a new hierarchical design method, k-means SVRM (support vector representation machine) clustering. This greatly improves upon our prior IJCNN hierarchical design. At each node in the hierarchy, we apply the SVRDM (support vector representation and discrimination machine) classifier, which offers generalization and good rejection ability. We also provide new theoretical bases and methods for our choice of the kernel function and new SVRDM parameter selection rules. Classification and rejection test results are presented on new databases of both simulated and real infra-red (IR) data. Yu-Chiang Frank Wang, David P. Casasent |
IJCNN | 1 |
| 2005 | A Hierarchical Classifier Using New Support Vector MachineabstractA binary hierarchical classifier is proposed to solve the multi-class classification problem. We also require rejection of non-target inputs, which produces a very difficult problem. The SVRDM (support vector representation and discrimination machine) classifier is considered at each node in the hierarchy, since it offers good generalization and rejection ability. Using this hierarchical SVRDM classifier with magnitude Fourier transform features, initial recognition and rejection test results on simulated infrared data are excellent. Yu-Chiang Frank Wang, David P. Casasent |
ICDAR | 1 |
| 2005 | Automatic target recognition using new support vector machineabstractA hierarchical classifier using a new SVRDM (support vector representation and discrimination machine) is proposed for automatic target recognition. An accuracy and distance-based method is used to design a hierarchical classifier. Our SVRDM hierarchical classifier has the ability to reject unseen non-object classes and clutter inputs. Uses of both iconic and spatial frequency domain features are considered. Initial recognition and rejection test results on infrared (IR) data are excellent. David P. Casasent, Yu-Chiang Frank Wang |
IJCNN | 2 |
| 2005 | A hierarchical classifier using new support vector machines for automatic target recognition
David P. Casasent, Yu-Chiang Frank Wang |
Neural Networks | 2 |