EDBT 2026 Demo / reviewers in the wild / expert
Yen-Yu Lin
dblp:44/4894
· DBLP profile ↗
118ranked-venue papers
10as first author
42since 2021 · last 2026
0000-0002-7183-6070ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 104 · 8 first-author · 35 since 2021Artificial intelligence and machine learning · 68 · 8 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDR Reconstruction Boosting with Training-Free and Exposure-Consistent Diffusion
Yo-Tin Lin, Su-Kai Chen, Hou-Ning Hu, Yen-Yu Lin |
WACV | 4 |
| 2026 | CSGaussian: Progressive Rate-Distortion Compression and Segmentation for 3D Gaussian SplattingabstractWe present the first unified framework for rate-distortion-optimized compression and segmentation of 3D Gaussian Splatting (3DGS). While 3DGS has proven effective for both real-time rendering and semantic scene understanding, prior works have largely treated these tasks independently, leaving their joint consideration unexplored. Inspired by recent advances in rate-distortion-optimized 3DGS compression, this work integrates semantic learning into the compression pipeline to support decoder-side applications–such as scene editing and manipulation–that extend beyond traditional scene reconstruction and view synthesis. Our scheme features a lightweight implicit neural representation-based hyperprior, enabling efficient entropy coding of both color and semantic attributes while avoiding costly grid-based hyperprior as seen in many prior works. To facilitate compression and segmentation, we further develop compression-guided segmentation learning, consisting of quantization-aware training to enhance feature separability and a quality-aware weighting mechanism to suppress unreliable Gaussian primitives. Extensive experiments on the LERF and 3D-OVS datasets demonstrate that our approach significantly reduces transmission cost while preserving high rendering quality and strong segmentation performance. Yu-Jen Tseng, Chia-Hao Kao, Jing-Zhong Chen, Alessandro Gnutti, Shao-Yuan Lo, Yen-Yu Lin, Wen-Hsiao Peng |
WACV | 6 |
| 2026 | GOT-JEPA: Generic Object Tracking With Model Adaptation and Occlusion Handling Using Joint-Embedding Predictive ArchitectureabstractThe human visual system tracks objects by integrating current observations with previously observed information, adapting to target and scene changes, and reasoning about occlusion at fine granularity. In contrast, recent generic object trackers are often optimized for training targets, which limits robustness and generalization in unseen scenarios, and their occlusion reasoning remains coarse, lacking detailed modeling of occlusion patterns. To address these limitations in generalization and occlusion perception, we propose GOT-JEPA, a model-predictive pretraining framework that extends JEPA from predicting image features to predicting tracking models. Given identical historical information, a teacher predictor generates pseudo-tracking models from a clean current frame, and a student predictor learns to predict the same pseudo-tracking models from a corrupted version of the current frame. This design provides stable pseudo supervision and explicitly trains the predictor to produce reliable tracking models under occlusions, distractors, and other adverse observations, improving generalization to dynamic environments. Building on GOT-JEPA, we further propose OccuSolver to enhance occlusion perception for object tracking. OccuSolver adapts a point-centric point tracker for object-aware visibility estimation and detailed occlusion-pattern capture. Conditioned on object priors iteratively generated by the tracker, OccuSolver incrementally refines visibility states, strengthens occlusion handling, and produces higher-quality reference labels that progressively improve subsequent model predictions. Extensive evaluations on seven benchmarks show that our method effectively enhances tracker generalization and robustness. The code will be available at https://github.com/chenshihfang/GOT. Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene InpaintingabstractThree-dimensional scene inpainting is crucial for applications from virtual reality to architectural visualization, yet existing methods struggle with view consistency and geometric accuracy in 360° unbounded scenes. We present AuraFusion360, a novel reference-based method that enables high-quality object removal and hole filling in 3D scenes represented by Gaussian Splatting. Our approach introduces (1) depth-aware unseen mask generation for accurate occlusion identification, (2) Adaptive Guided Depth Diffusion, a zero-shot method for accurate initial point placement without requiring additional training, and (3) SDEdit-based detail enhancement for multi-view coherence. We also introduce 360-USID, the first comprehensive dataset for 360° unbounded scene inpainting with ground truth. Extensive experiments demonstrate that AuraFusion360 significantly outperforms existing methods, achieving superior perceptual quality while maintaining geometric accuracy across dramatic viewpoint changes. Chung-Ho Wu, Yang-Jung Chen, Ying-Huan Chen, Jie-Ying Lee, Bo-Hsu Ke, Chun-Wei Tuan Mu, Yi-Chuan Huang, Chin-Yang Lin, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu 0001 |
CVPR | 10 |
| 2025 | LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long VideosabstractLongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitations. To address these issues, we introduce LongSplat, a robust unposed 3D Gaussian Splatting framework featuring: (1) Incremental Joint Optimization that concurrently optimizes camera poses and 3D Gaussians to avoid local minima and ensure global consistency; (2) a robust Pose Estimation Module leveraging learned 3D priors; and (3) an efficient Octree Anchor Formation mechanism that converts dense point clouds into anchors based on spatial density. Extensive experiments on challenging benchmarks demonstrate that LongSplat achieves state-of-the-art results, substantially improving rendering quality, pose accuracy, and computational efficiency compared to prior approaches. Project page: https://linjohnss.github.io/longsplat/ Chin-Yang Lin, Cheng Sun 0004, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu 0001 |
ICCV | 5 |
| 2025 | PHATNet: A Physics-Guided Haze Transfer Network for Domain-Adaptive Real-World Image DehazingabstractImage dehazing aims to remove unwanted hazy artifacts in images. Although previous research has collected paired real-world hazy and haze-free images to improve dehazing models' performance in real-world scenarios, these models often experience significant performance drops when handling unseen real-world hazy images due to limited training data. This issue motivates us to develop a flexible domain adaptation method to enhance dehazing performance during testing. Observing that predicting haze patterns is generally easier than recovering clean content, we propose the Physics-guided Haze Transfer Network (PHATNet) which transfers haze patterns from unseen target domains to source-domain haze-free images, creating domain-specific fine-tuning sets to update dehazing models for effective domain adaptation. Additionally, we introduce a Haze-Transfer-Consistency loss and a Content-Leakage Loss to enhance PHATNet's disentanglement ability. Experimental results demonstrate that PHATNet significantly boosts state-of-the-art dehazing models on benchmark real-world image dehazing datasets. Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin, Chia-Wen Lin |
ICCV | 3 |
| 2025 | Consistent View Synthesis with Bidirectional Epipolar Attention and ReconstructionabstractNovel view synthesis from a single image aims to generate novel scene views given a reference image and a sequence of camera poses. Its primary difficulty lies in effectively leveraging a generative model to achieve high-quality image generation while simultaneously ensuring consistency and faithfulness across synthesized views. In this paper, we propose a novel approach to address the consistency and faithfulness issues in view synthesis. Specifically, we develop a new attention layer, termed bidirectional epipolar attention, which utilizes a pair of complementary epipolar lines to guide the associations between features from different viewpoints. Each bidirectional epipolar layer calculates forward and backward epipolar lines, enabling geometrically constrained attention that improves cross-view consistency. To ensure faithful synthesis, we introduce an epipolar-aware reconstruction module that prevents creating novel content in regions where the newly generated image overlaps with existing ones. Extensive experimental results demonstrate that our method outperforms previous approaches to novel view synthesis, achieving superior performance in both image quality and consistency. The source code is available at https://github.com/fallantbell/Bidirectional-Epipolar-Synthesis. I-Chung Chiu, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
ICIP | 4 |
| 2025 | Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and ComprehensionabstractReferring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. Moreover, REG methods frequently struggle to bridge the visual and textual domains due to the limited capacity, leading to low-quality and restricted diversity in expression generation. To address these issues, we propose a novel VIsion-guided Expression Diffusion Model (VIE-DM) for the REG task, where diverse synonymous expressions adhering to both image and text contexts of the target object are generated to augment REC datasets. VIE-DM consists of a vision-text condition (VTC) module and a transformer decoder. Our VTC and token selection design effectively addresses the feature discrepancy problem prevalent in existing REG methods. This enables us to generate high-quality, diverse synonymous expressions that can serve as augmented data for REC model learning. Extensive experiments on five datasets demonstrate the high quality and large diversity of our generated expressions. Furthermore, the augmented image-expression pairs consistently enhance the performance of existing REC models, achieving state-of-the-art results. Jingcheng Ke, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
ICLR | 5 |
| 2025 | Ranking-aware adapter for text-driven image ordering with CLIPabstractRecent advances in vision-language models (VLMs) have made significant progress in downstream tasks that require quantitative concepts such as facial age estimation and image quality assessment, enabling VLMs to explore applications like image ranking and retrieval. However, existing studies typically focus on the reasoning based on a single image and heavily depend on text prompting, limiting their ability to learn comprehensive understanding from multiple images. To address this, we propose an effective yet efficient approach that reframes the CLIP model into a learning-to-rank task and introduces a lightweight adapter to augment CLIP for text-guided image ranking. Specifically, our approach incorporates learnable prompts to adapt to new instructions for ranking purposes and an auxiliary branch with ranking-aware attention, leveraging text-conditioned visual differences for additional supervision in image ranking. Our ranking-aware adapter consistently outperforms fine-tuned CLIPs on various tasks and achieves competitive results compared to state-of-the-art models designed for specific tasks like facial age estimation and image quality assessment. Overall, our approach primarily focuses on ranking images with a single instruction, which provides a natural and generalized way of learning from visual differences across images, bypassing the need for extensive text prompts tailored to individual tasks. Wei-Hsiang Yu, Yen-Yu Lin, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
ICLR | 2 |
| 2025 | BlurDM: A Blur Diffusion Model for Image DeblurringabstractDiffusion models show promise for dynamic scene deblurring; however, existing studies often fail to leverage the intrinsic nature of the blurring process within diffusion models, limiting their full potential. To address it, we present a Blur Diffusion Model (BlurDM), which seamlessly integrates the blur formation process into diffusion for image deblurring. Observing that motion blur stems from continuous exposure, BlurDM implicitly models the blur formation process through a dual-diffusion forward scheme, diffusing both noise and blur onto a sharp image. During the reverse generation process, we derive a dual denoising and deblurring formulation, enabling BlurDM to recover the sharp image by simultaneously denoising and deblurring, given pure Gaussian noise conditioned on the blurred image as input. Additionally, to efficiently integrate BlurDM into deblurring networks, we perform BlurDM in the latent space, forming a flexible prior generation network for deblurring. Extensive experiments demonstrate that BlurDM significantly and consistently enhances existing deblurring methods on four benchmark datasets. The project page is available at https://jin-ting-he.github.io/BlurDM/. Jin-Ting He, Fu-Jen Tsai, Yan-Tsung Peng, Min-Hung Chen, Chia-Wen Lin, Yen-Yu Lin |
NeurIPS | 6 |
| 2025 | ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark DetectionabstractAlthough facial landmark detection (FLD) has gained significant progress, existing FLD methods still suffer from performance drops on partially non-visible faces, such as faces with occlusions or under extreme lighting conditions or poses. To address this issue, we introduce ORFormer, a novel transformer-based method that can detect non-visible regions and recover their missing features from vis-ible parts. Specifically, ORFormer associates each image patch token with one additional learnable token called the messenger token. The messenger token aggregates features from all but its patch. This way, the consensus between a patch and other patches can be assessed by referring to the similarity between its regular and messenger embeddings, enabling non-visible region identification. Our method then recovers occluded patches with features aggregated by the messenger tokens. Leveraging the recovered features, OR-Former compiles high-quality heatmaps for the downstream FLD task. Extensive experiments show that our method generates heatmaps resilient to partial occlusions. By inte-grating the resultant heatmaps into existing FLD methods, our method performs favorably against the state of the arts on challenging datasets such as WFLWand COFW. Jui-Che Chiang, Hou-Ning Hu, Bo-Syuan Hou, Chia-Yu Tseng, Yu-Lun Liu 0001, Min-Hung Chen, Yen-Yu Lin |
WACV | 7 |
| 2025 | CorrFill: Enhancing Faithfulness in Reference-Based Inpainting with Correspondence Guidance in Diffusion ModelsabstractIn the task of reference-based image inpainting, an additional reference image is provided to restore a damaged target image to its original state. The advancement of diffusion models, particularly Stable Diffusion, allows for simple formulations in this task. However, existing diffusion-based methods often lack explicit constraints on the correlation between the reference and damaged images, resulting in lower faithfulness to the reference images in the inpainting results. In this work, we propose CorrFill, a training-free module designed to enhance the awareness of geometric correlations between the reference and target images. This enhancement is achieved by guiding the inpainting process with correspondence constraints estimated during inpainting, utilizing attention masking in self-attention layers and an objective function to update the input tensor according to the constraints. Experimental results demonstrate that CorrFill significantly enhances the performance of multiple baseline diffusion-based methods, including state-of-the-art approaches, by emphasizing faithfulness to the reference images. Kuan-Hung Liu, Cheng-Kun Yang, Min-Hung Chen, Yu-Lun Liu 0001, Yen-Yu Lin |
WACV | 5 |
| 2025 | Improving Visual Object Tracking Through Visual PromptingabstractLearning a discriminative model to distinguish a target from its surrounding distractors is essential to generic visual object tracking. Dynamic target representation adaptation against distractors is challenging due to the limited discriminative capabilities of prevailing trackers. We present a new visual Prompting mechanism for generic Visual Object Tracking (PiVOT) to address this issue. PiVOT proposes a prompt generation network with the pre-trained foundation model CLIP to automatically generate and refine visual prompts, enabling the transfer of foundation model knowledge for tracking. While CLIP offers broad category-level knowledge, the tracker, trained on instance-specific data, excels at recognizing unique object instances. Thus, PiVOT first compiles a visual prompt highlighting potential target locations. To transfer the knowledge of CLIP to the tracker, PiVOT leverages CLIP to refine the visual prompt based on the similarities between candidate objects and the reference templates across potential targets. Once the visual prompt is refined, it can better highlight potential target locations, thereby reducing irrelevant prompt information. With the proposed prompting mechanism, the tracker can generate improved instance-aware feature maps through the guidance of the visual prompt, thus effectively reducing distractors. The proposed method does not involve CLIP during training, thereby keeping the same training complexity and preserving the generalization capability of the pretrained foundation model. Extensive experiments across multiple benchmarks indicate that PiVOT, using the proposed prompting method can suppress distracting objects and enhance the tracker. Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
IEEE Trans. Multim. | 4 |
| 2025 | Make Graph-Based Referring Expression Comprehension Great Again Through Expression-Guided Dynamic Gating and RegressionabstractOne common belief is that with complex models and pre-training on large-scale datasets, transformer-based methods for referring expression comprehension (REC) perform much better than existing graph-based methods. We observe that since most graph-based methods adopt an off-the-shelf detector to locate candidate objects (i.e., regions detected by the object detector), they face two challenges that result in subpar performance: (1) the presence of significant noise caused by numerous irrelevant objects during reasoning, and (2) inaccurate localization outcomes attributed to the provided detector. To address these issues, we introduce a plug-and-adapt module guided by sub-expressions, called dynamic gate constraint (DGC), which can adaptively disable irrelevant proposals and their connections in graphs during reasoning. We further introduce an expression-guided regression strategy (EGR) to refine location prediction. Extensive experimental results on the RefCOCO, RefCOCO+, RefCOCOg, Flickr30 K, RefClef, and Ref-reasoning datasets demonstrate the effectiveness of the DGC module and the EGR strategy in consistently boosting the performances of various graph-based REC methods. Without any pretaining, the proposed graph-based method achieves better performance than the state-of-the-art (SOTA) transformer-based methods. Jingcheng Ke, Dele Wang, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
IEEE Trans. Multim. | 6 |
| 2024 | PartDistill: 3D Shape Part Segmentation by Vision-Language Model DistillationabstractThis paper proposes a cross-modal distillation frame-work, PartDistill, which transfers 2D knowledge from vision-language models (VLMs) to facilitate 3D shape part segmentation. PartDistill addresses three major challenges in this task: the lack of 3D segmentation in invisible or undetected regions in the 2D projections, inconsistent 2D predictions by VLMs, and the lack of knowledge accumu-lation across different 3D shapes. PartDistill consists of a teacher network that uses a VLM to make 2D predictions and a student network that learns from the 2D pre-dictions while extracting geometrical features from multi-ple 3D shapes to carry out 3D part segmentation. A bi-directional distillation, including forward and backward distillations, is carried out within the framework, where the former forward distills the 2D predictions to the student net-work, and the latter improves the quality of the 2D predictions, which subsequently enhances the final 3D segmen-tation. Moreover, PartDistill can exploit generative mod-els that facilitate effortless 3D shape creation for generating knowledge sources to be distilled. Through extensive experiments, PartDistill boosts the existing methods with substantial margins on widely used ShapeNetPart and Part-NetE datasets, by more than 15% and 12% higher mIoU scores, respectively. The code for this work is available at https://github.com/ardianumam/PartDistill. Ardian Umam, Cheng-Kun Yang, Min-Hung Chen, Jen-Hui Chuang, Yen-Yu Lin |
CVPR | 5 |
| 2024 | Image-Text Co-Decomposition for Text-Supervised Semantic SegmentationabstractThis paper addresses text-supervised semantic segmentation, aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively aligns visual segments with the meanings of texts. We notice that there is a discrepancy between text alignment and semantic segmentation: A text often consists of multiple semantic concepts, whereas semantic segmentation strives to create semantically homogeneous segments. To address this issue, we propose a novel framework, Image-Text Co-Decomposition (CoDe), where the paired image and text are jointly decomposed into a set of image regions and a set of word segments, respectively, and contrastive learning is developed to enforce region-word alignment. To work with a vision-language model, we present a prompt learning mechanism that derives an extra representation to highlight an image segment or a word segment of interest, with which more effective features can be extracted from that segment. Comprehensive experimental results demonstrate that our method performs favorably against existing text-supervised semantic segmentation methods on six benchmark datasets. The code is available at https://github.com/072jiajia/image-text-co-decomposition. Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin |
CVPR | 9 |
| 2024 | ID-Blau: Image Deblurring by Implicit Diffusion-Based reBLurring AUgmentationabstractImage deblurring aims to remove undesired blurs from an image captured in a dynamic scene. Much research has been dedicated to improving deblurring performance through model architectural designs. However, there is little work on data augmentation for image deblurring. Since continuous motion causes blurred artifacts during image exposure, we aspire to develop a groundbreaking blur augmentation method to generate diverse blurred images by simulating motion trajectories in a continuous space. This paper proposes Implicit Diffusion-based reBLurring AUgmentation (ID-Blau), utilizing a sharp image paired with a controllable blur condition map to produce a corresponding blurred image. We parameterize the blur patterns of a blurred image with their orientations and magnitudes as a pixel-wise blur condition map to simulate motion trajectories and implicitly represent them in a continuous space. By sampling diverse blur conditions, ID-Blau can generate various blurred images unseen in the training set. Experimental results demonstrate that ID-Blau can produce realistic blurred images for training and thus significantly improve performance for state-of-the-art deblurring models. The source code is available at https://github.com/plusgood-steven/ID-Blau. Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Chia-Wen Lin, Yen-Yu Lin |
CVPR | 6 |
| 2024 | Domain-Adaptive Video Deblurring via Test-Time Blurring
Jin-Ting He, Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Chia-Wen Lin, Yen-Yu Lin |
ECCV (30) | 7 |
| 2024 | CLIPREC: Graph-Based Domain Adaptive Network for Zero-Shot Referring Expression ComprehensionabstractReferring expression comprehension (REC) is a cross-modal matching task that aims to localize the target object in an image specified by a text description. Most existing approaches for this task focus on identifying only objects whose categories are covered by training data. This restricts their generalization to unseen categories and practical usage. To address this issue, we propose a domain adaptive network called CLIPREC for zero-shot REC, which integrates the Contrastive Language-Image Pretraining (CLIP) model for graph-based REC. The proposed CLIPREC is composed of a graph collaborative attention module with two directed graphs: one for objects in an image and the other for their corresponding categorical labels. To carry out zero-shot REC, we leverage the strong common image-text feature space from the CLIP model to correlate the two graphs. Furthermore, a multilayer perceptron is introduced to enable feature alignment so that the CLIP model is adapted to the expression representation from the language parser, resulting in effective reasoning from expressions involving both seen and unseen object categories. Extensive experimental and ablation results on several widely-adopted benchmarks show that the proposed approach performs favorably against state-of-the-art approaches for zero-shot REC. Jingcheng Ke, Jia Wang 0020, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
IEEE Trans. Multim. | 6 |
| 2024 | Unsupervised Point Cloud Co-Part Segmentation via Co-Attended Superpoint Generation and AggregationabstractWe propose a co-part segmentation method that takes a set of point clouds of the same category as input where neither a ground truth label nor a prior network is required. With difficulties caused by the label absence, we formulate the co-part segmentation task into two subtasks, including superpoint generation and part aggregation. In the first subtask, our superpoint generation network divides each point cloud into homogeneous partitions, each called superpoint, while in the second subtask, these superpoints are further aggregated into a few semantic parts via our part aggregation network. We introduce the coupled attention blocks in the part aggregation network to explicitly enforce semantic consistency in the segmentation by exploiting intra-, inter-, and paired-cloud geometrical information by minimizing the devised intra-, inter-, and paired-cloud losses, respectively. The intra-cloud loss triggers a semantic segmentation in each point cloud, while the inter-cloud loss considers all clouds to enforce their semantic consistency. The paired-cloud loss is designed to ensure that each part of one point cloud can be discriminatively reconstructed from the superpoints of another point cloud. We perform experiments on two benchmark datasets, ShapeNet part and COSEG, and provide quantitative and qualitative results to demonstrate the superiority of our method over existing methods. We also show that the proposed method can help several downstream tasks, including semi-supervised part segmentation and data augmentation for shape classification. The code for this work will be publicly available upon the paper's publication. Ardian Umam, Cheng-Kun Yang, Jen-Hui Chuang, Yen-Yu Lin |
IEEE Trans. Multim. | 4 |
| 2023 | Decontamination Transformer For Blind Image InpaintingabstractBlind image inpainting aims at recovering the content from a corrupted image in which the mask indicating the corrupted regions is not available in inference time. Inspired that most existing methods for inpainting suffer from complex contamination, we propose a model that explicitly predicts the realvalued alpha mask and contaminant to eliminate the contamination from the corrupted image, thus improving the inpainting performance. To enhance the overall semantic consistency, the attention mechanism of transformers is exploited and integrated into our inpainting network. We conduct extensive experiments to verify our method against blind and non-blind inpainting models and demonstrate its effectiveness and generalizability to different sources of contaminant. Chun-Yi Li, Yen-Yu Lin, Walon Wei-Chen Chiu |
ICASSP | 2 |
| 2023 | MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-ResolutionabstractThis work addresses continuous space-time video super-resolution (C-STVSR) that aims to up-scale an input video both spatially and temporally by any scaling factors. One key challenge of C-STVSR is to propagate information temporally among the input video frames. To this end, we introduce a space-time local implicit neural function. It has the striking feature of learning forward motion for a continuum of pixels. We motivate the use of forward motion from the perspective of learning individual motion trajectories, as opposed to learning a mixture of motion trajectories with backward motion. To ease motion interpolation, we encode sparsely sampled forward motion extracted from the input video as the contextual input. Along with a reliability-aware splatting and decoding scheme, our framework, termed MoTIF, achieves the state-of-the-art performance on C-STVSR. The source code of MoTIF is available at https://github.com/sichun233746/MoTIF. Si-Cun Chen, Yi-Hsin Chen, Yen-Yu Lin, Wen-Hsiao Peng |
ICCV | 4 |
| 2023 | Learning Continuous Exposure Value Representations for Single-Image HDR ReconstructionabstractDeep learning is commonly used to reconstruct HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning-generated LDR stack. However, current methods generate the stack with predetermined exposure values (EVs), which may limit the quality of HDR reconstruction. To address this, we propose the continuous exposure value representation (CEVR), which uses an implicit function to generate LDR images with arbitrary EVs, including those unseen during training. Our approach generates a continuous stack with more images containing diverse EVs, significantly improving HDR reconstruction. We use a cycle training strategy to supervise the model in generating continuous EV LDR images without corresponding ground truths. Our CEVR model outperforms existing methods, as demonstrated by experimental results. Su-Kai Chen, Hung-Lin Yen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Wen-Hsiao Peng, Yen-Yu Lin |
ICCV | 7 |
| 2023 | 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level SupervisionabstractWe present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to achieve 2D-3D information fusion. Considering the high annotation cost of point clouds, effective 2D and 3D feature fusion based on weakly supervised learning is in great demand. To this end, we propose a transformer model with two encoders and one decoder for weakly supervised point cloud segmentation using only scene-level class tags. Specifically, the two encoders compute the self-attended features for 3D point clouds and 2D multi-view images, respectively. The decoder implements interlaced 2D-3D cross-attention and carries out implicit 2D and 3D feature fusion. We alternately switch the roles of queries and key-value pairs in the decoder layers. It turns out that the 2D and 3D features are iteratively enriched by each other. Experiments show that it performs favorably against existing weakly supervised point cloud segmentation methods by a large margin on the S3DIS and ScanNet benchmarks. The project page will be available at https://jimmy15923.github.io/mit_web/. Cheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu Lin |
ICCV | 4 |
| 2023 | Task-Adaptive Feature Matching Loss for Image DeblurringabstractImage deblurring is a highly challenging and ill-posed image restoration problem. Contemporary deep learning-based approaches usually tackle this problem by exploiting the encoder-decoder-based models trained by the commonly used mean squared error loss with the feature matching loss as a regularization to obtain perceptual consistent restored results as the ground truths. We argue that since the general backbone models for computing feature matching loss are usually not trained on the image deblurring task, the loss lacks specific knowledge of blur and usually leads to suboptimal performance. To address this issue, we propose a task-adaptive feature matching loss for image deblurring where we synthesize blurred images in different blur extents and employ triplet loss to finetune the backbone model for learning specific blur priors. Then, we leverage the finetuned backbone to compute feature matching loss which can greatly enhance the existing image deblurring models for better perceptual results. With extensive experiments on the GoPro and RealBlur datasets, both qualitative and quantitative results show that the SOTA deblurring models trained with the proposed loss can effectively obtain better and sharper restored images in terms of various perceptual image quality metrics than the original models while maintaining comparable PSNR and SSIM performances. Chiao-Chang Chang, Bo-Cheng Yang, Yi-Ting Liu, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
ICIP | 6 |
| 2023 | Diffusion-SS3D: Diffusion Model for Semi-supervised 3D Object DetectionabstractSemi-supervised object detection is crucial for 3D scene understanding, efficiently addressing the limitation of acquiring large-scale 3D bounding box annotations. Existing methods typically employ a teacher-student framework with pseudo-labeling to leverage unlabeled point clouds. However, producing reliable pseudo-labels in a diverse 3D space still remains challenging. In this work, we propose Diffusion-SS3D, a new perspective of enhancing the quality of pseudo-labels via the diffusion model for semi-supervised 3D object detection. Specifically, we include noises to produce corrupted 3D object size and class label distributions, and then utilize the diffusion model as a denoising process to obtain bounding box outputs. Moreover, we integrate the diffusion model into the teacher-student framework, so that the denoised bounding boxes can be used to improve pseudo-label generation, as well as the entire semi-supervised learning process. We conduct experiments on the ScanNet and SUN RGB-D benchmark datasets to demonstrate that our approach achieves state-of-the-art performance against existing methods. We also present extensive analysis to understand how our diffusion model design affects performance in semi-supervised learning. The source code will be available at https://github.com/luluho1208/Diffusion-SS3D. Cheng-Ju Ho, Chen-Hsuan Tai, Yen-Yu Lin, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
NeurIPS | 3 |
| 2023 | Unsupervised sound localization via iterative contrastive learning
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee 0001, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2022 | TANet: Triplet Attention Network for All-In-One Adverse Weather Image Restoration
Hsing-Hua Wang, Fu-Jen Tsai, Yen-Yu Lin, Chia-Wen Lin |
ACCV (4) | 3 |
| 2022 | Learning Object-level Point Augmentor for Semi-supervised 3D Object Detection
Cheng-Ju Ho, Chen-Hsuan Tai, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
BMVC | 4 |
| 2022 | Meta Transferring for Deblurring
Po-Sheng Liu, Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Chia-Wen Lin, Yen-Yu Lin |
BMVC | 6 |
| 2022 | An MIL-Derived Transformer for Weakly Supervised Point Cloud SegmentationabstractWe address weakly supervised point cloud segmentation by proposing a new model, MIL-derived transformer, to mine additional supervisory signals. First, the transformer model is derived based on multiple instance learning (MIL) to explore pair-wise cloud-level supervision, where two clouds of the same category yield a positive bag while two of different classes produce a negative bag. It leverages not only individual cloud annotations but also pair-wise cloud semantics for model optimization. Second, Adaptive global weighted pooling (AdaGWP) is integrated into our transformer model to replace max pooling and average pooling. It introduces learnable weights to re-scale logits in the class activation maps. It is more robust to noise while discovering more complete foreground points under weak supervision. Third, we perform point subsampling and enforce feature equivariance between the original and subsampled point clouds for regularization. The proposed method is end-to-end trainable and is general because it can work with different backbones with diverse types of weak supervision signals, including sparsely annotated points and cloud-level labels. The experiments show that it achieves state-of-the-art performance on the S3DIS and ScanNet benchmarks. The source code will be available at https://github.com/jimmy15923/wspss_mil_transformer. Cheng-Kun Yang, Ji-Jia Wu, Kai-Syun Chen, Yung-Yu Chuang, Yen-Yu Lin |
CVPR | 5 |
| 2022 | Stripformer: Strip Transformer for Fast Image Deblurring
Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin, Chung-Chi Tsai, Chia-Wen Lin |
ECCV (19) | 3 |
| 2022 | Point MixSwap: Attentional Point Cloud Mixing via Swapping Matched Structural Divisions
Ardian Umam, Cheng-Kun Yang, Yung-Yu Chuang, Jen-Hui Chuang, Yen-Yu Lin |
ECCV (29) | 5 |
| 2022 | AQT: Adversarial Query Transformers for Domain Adaptive Object DetectionabstractAdversarial feature alignment is widely used in domain adaptive object detection. Despite the effectiveness on CNN-based detectors, its applicability to transformer-based detectors is less studied. In this paper, we present AQT (adversarial query transformers) to integrate adversarial feature alignment into detection transformers. The generator is a detection transformer which yields a sequence of feature tokens, and the discriminator consists of a novel adversarial token and a stack of cross-attention layers. The cross-attention layers take the adversarial token as the query and the feature tokens from the generator as the key-value pairs. Through adversarial learning, the adversarial token in the discriminator attends to the domain-specific feature tokens, while the generator produces domain-invariant features, especially on the attended tokens, hence realizing adversarial feature alignment on transformers. Thorough experiments over several domain adaptive object detection benchmarks demonstrate that our approach performs favorably against the state-of-the-art methods. Source code is available at https://github.com/weii41392/AQT. Wei-Jie Huang, Yu-Lin Lu, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin |
IJCAI | 5 |
| 2022 | BANet: A Blur-Aware Attention Network for Dynamic Scene DeblurringabstractImage motion blur results from a combination of object motions and camera shakes, and such blurring effect is generally directional and non-uniform. Previous research attempted to solve non-uniform blurs using self-recurrent multi-scale, multi-patch, or multi-temporal architectures with self-attention to obtain decent results. However, using self-recurrent frameworks typically leads to a longer inference time, while inter-pixel or inter-channel self-attention may cause excessive memory usage. This paper proposes a Blur-aware Attention Network (BANet), that accomplishes accurate and efficient deblurring via a single forward pass. Our BANet utilizes region-based self-attention with multi-kernel strip pooling to disentangle blur patterns of different magnitudes and orientations and cascaded parallel dilated convolution to aggregate multi-scale content features. Extensive experimental results on the GoPro and RealBlur benchmarks demonstrate that the proposed BANet performs favorably against the state-of-the-arts in blurred image restoration and can provide deblurred results in real-time. Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Yen-Yu Lin, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2021 | Unsupervised Point Cloud Object Co-segmentation by Co-contrastive Learning and Mutual Attention SamplingabstractThis paper presents a new task, point cloud object co-segmentation, aiming to segment the common 3D objects in a set of point clouds. We formulate this task as an object point sampling problem, and develop two techniques, the mutual attention module and co-contrastive learning, to enable it. The proposed method employs two point samplers based on deep neural networks, the object sampler and the background sampler. The former targets at sampling points of common objects while the latter focuses on the rest. The mutual attention module explores point-wise correlation across point clouds. It is embedded in both samplers and can identify points with strong cross-cloud correlation from the rest. After extracting features for points selected by the two samplers, we optimize the networks by developing the co-contrastive loss, which minimizes feature discrepancy of the estimated object points while maximizing feature separation between the estimated object and back-ground points. Our method works on point clouds of an arbitrary object class. It is end-to-end trainable and does not need point-level annotations. It is evaluated on the ScanObjectNN and S3DIS datasets and achieves promising results. The source code will be available at https://github.com/jimmy15923/unsup_point_coseg. Cheng-Kun Yang, Yung-Yu Chuang, Yen-Yu Lin |
ICCV | 3 |
| 2021 | Soft Ranking Threshold Losses For Image RetrievalabstractThis paper proposes a novel loss, soft ranking threshold loss, for driving deep networks to learn better representations for image retrieval. Instead of working in the metric space, our loss works in the rank space which has a more uniform distribution and explicit scale and bounds. Our loss reduces the ranks of the distances between anchor-positive pairs below the threshold while increasing the ones between anchor-negative pairs above the threshold. In addition to the basic form, two extensions are proposed for improving the effectiveness: hard thresholds and ranking margin. Experiments show that the proposed loss outperforms the state-of-the-art losses on image retrieval applications. Chiao-An Yang, Zhixiang Wang 0001, Yen-Yu Lin, Yung-Yu Chuang |
ICIP | 3 |
| 2021 | Stain Mix-Up: Unsupervised Domain Generalization for Histopathology Images
Jia-Ren Chang, Min-Sheng Wu, Wei-Hsiang Yu, Chi-Chung Chen, Cheng-Kung Yang, Yen-Yu Lin, Chao-Yuan Yeh |
MICCAI (3) | 6 |
| 2021 | Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingabstractThe audio-visual video parsing task aims to temporally parse a video into audio or visual event categories. However, it is labor intensive to temporally annotate audio and visual events and thus hampers the learning of a parsing model. To this end, we propose to explore additional cross-video and cross-modality supervisory signals to facilitate weakly-supervised audio-visual video parsing. The proposed method exploits both the common and diverse event semantics across videos to identify audio or visual events. In addition, our method explores event co-occurrence across audio, visual, and audio-visual streams. We leverage the explored cross-modality co-occurrence to localize segments of target events while excluding irrelevant ones. The discovered supervisory signals across different videos and modalities can greatly facilitate the training with only video-level annotations. Quantitative and qualitative results demonstrate that the proposed method performs favorably against existing methods on weakly-supervised audio-visual video parsing. Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee 0001, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
NeurIPS | 4 |
| 2021 | MVHM: A Large-Scale Multi-View Hand Mesh Benchmark for Accurate 3D Hand Pose EstimationabstractEstimating 3D hand poses from a single RGB image is challenging because depth ambiguity leads the problem ill-posed. Training hand pose estimators with 3D hand mesh annotations and multi-view images often results in significant performance gains. However, existing multi-view datasets are relatively small with hand joints annotated by off-the-shelf trackers or automated through model predictions, both of which may be inaccurate and can introduce biases. Collecting a large-scale multi-view 3D hand pose images with accurate mesh and joint annotations is valuable but strenuous. In this paper, we design a spin match algorithm that enables a rigid mesh model matching with any target mesh ground truth. Based on the match algorithm, we propose an efficient pipeline to generate a large-scale multi-view hand mesh (MVHM) dataset with accurate 3D hand mesh and joint labels. We further present a multi-view hand pose estimation approach to verify that training a hand pose estimator with our generated dataset greatly enhances the performance. Experimental results show that our approach achieves the performance of 0.990 in AUC20-50 on the MHP dataset compared to the previous state-of-the-art of 0.939 on this dataset. Our datasset is available at https://github.com/Kuzphi/MVHM. Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin, Xiaohui Xie |
WACV | 4 |
| 2021 | Temporal-Aware Self-Supervised Learning for 3D Hand Pose and Mesh Estimation in VideosabstractEstimating 3D hand pose directly from RGB images is challenging but has gained steady progress recently by training deep models with annotated 3D poses. However annotating 3D poses is difficult and as such only a few 3D hand pose datasets are available, all with limited sample sizes. In this study, we propose a new framework of training 3D pose estimation models from RGB images without using explicit 3D annotations, i.e., trained with only 2D information. Our framework is motivated by two observations: 1) Videos provide richer information for estimating 3D poses as opposed to static images; 2) Estimated 3D poses ought to be consistent whether the videos are viewed in the forward order or reverse order. We leverage these two observations to develop a self-supervised learning model called temporal-aware self-supervised network (TASSN). By enforcing temporal consistency constraints, TASSN learns 3D hand poses and meshes from videos with only 2D keypoint position annotations. Experiments show that our model achieves surprisingly good results, with 3D estimation accuracy on par with the state-of-the-art models trained with 3D annotations, highlighting the benefit of the temporal consistency in constraining 3D prediction models. Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin, Xiaohui Xie |
WACV | 4 |
| 2021 | Show, Match and Segment: Joint Weakly Supervised Learning of Semantic Matching and Object Co-SegmentationabstractWe present an approach for jointly matching and segmenting object instances of the same category within a collection of images. In contrast to existing algorithms that tackle the tasks of semantic matching and object co-segmentation in isolation, our method exploits the complementary nature of the two tasks. The key insights of our method are two-fold. First, the estimated dense correspondence fields from semantic matching provide supervision for object co-segmentation by enforcing consistency between the predicted masks from a pair of images. Second, the predicted object masks from object co-segmentation in turn allow us to reduce the adverse effects due to background clutters for improving semantic matching. Our model is end-to-end trainable and does not require supervision from manually annotated correspondences and object masks. We validate the efficacy of our approach on five benchmark datasets: TSS, Internet, PF-PASCAL, PF-WILLOW, and SPair-71k, and show that our algorithm performs favorably against the state-of-the-art methods on both semantic matching and object co-segmentation tasks. Yun-Chun Chen, Yen-Yu Lin, Ming-Hsuan Yang 0001, Jia-Bin Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Weakly-Supervised Video Re-Localization with Multiscale Attention ModelabstractVideo re-localization aims to localize a sub-sequence, called target segment, in an untrimmed reference video that is similar to a given query video. In this work, we propose an attention-based model to accomplish this task in a weakly supervised setting. Namely, we derive our CNN-based model without using the annotated locations of the target segments in reference videos. Our model contains three modules. First, it employs a pre-trained C3D network for feature extraction. Second, we design an attention mechanism to extract multiscale temporal features, which are then used to estimate the similarity between the query video and a reference video. Third, a localization layer detects where the target segment is in the reference video by determining whether each frame in the reference video is consistent with the query video. The resultant CNN model is derived based on the proposed co-attention loss which discriminatively separates the target segment from the reference video. This loss maximizes the similarity between the query video and the target segment while minimizing the similarity between the target segment and the rest of the reference video. Our model can be modified to fully supervised re-localization. Our method is evaluated on a public dataset and achieves the state-of-the-art performance under both weakly supervised and fully supervised settings. Yung-Han Huang, Kuang-Jui Hsu, Shyh-Kang Jeng, Yen-Yu Lin |
AAAI | 4 |
| 2020 | Regularizing Meta-learning via Gradient Dropout
Hung-Yu Tseng, Yi-Wen Chen, Yi-Hsuan Tsai, Sifei Liu, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
ACCV (4) | 5 |
| 2020 | Every Pixel Matters: Center-Aware Feature Alignment for Domain Adaptive Object Detector
Cheng-Chun Hsu, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
ECCV (9) | 3 |
| 2020 | Spatiotemporal Super-Resolution with Cross-Task Consistency and Its Semi-supervised ExtensionabstractSpatiotemporal super-resolution (SR) aims to upscale both the spatial and temporal dimensions of input videos, and produces videos with higher frame resolutions and rates. It involves two essential sub-tasks: spatial SR and temporal SR. We design a two-stream network for spatiotemporal SR in this work. One stream contains a temporal SR module followed by a spatial SR module, while the other stream has the same two modules in the reverse order. Based on the interchangeability of performing the two sub-tasks, the two network streams are supposed to produce consistent spatiotemporal SR results. Thus, we present a cross-stream consistency to enforce the similarity between the outputs of the two streams. In this way, the training of the two streams is correlated, which allows the two SR modules to share their supervisory signals and improve each other. In addition, the proposed cross-stream consistency does not consume labeled training data and can guide network training in an unsupervised manner. We leverage this property to carry out semi-supervised spatiotemporal SR. It turns out that our method makes the most of training data, and can derive an effective model with few high-resolution and high-frame-rate videos, achieving the state-of-the-art performance. The source code of this work is available at https://hankweb.github.io/STSRwithCrossTask/. Han-Yi Lin, Pi-Cheng Hsiu, Tei-Wei Kuo, Yen-Yu Lin |
IJCAI | 4 |
| 2020 | Learning From Music to Visual Storytelling of Shots: A Deep Interactive Learning MechanismabstractLearning from music to visual storytelling of shots is an interesting and emerging task. It produces a coherent visual story in the form of a shot type sequence, which not only expands the storytelling potential for a song but also facilitates automatic concert video mashup process and storyboard generation. In this study, we present a deep interactive learning (DIL) mechanism for building a compact yet accurate sequence-to-sequence model to accomplish the task. Different from the one-way transfer between a pre-trained teacher network (or ensemble network) and a student network in knowledge distillation (KD), the proposed method enables collaborative learning between an ensemble teacher network and a student network. Namely, the student network also teaches. Specifically, our method first learns a teacher network that is composed of several assistant networks to generate a shot type sequence and produce the soft target (shot types) distribution accordingly through KD. It then constructs the student network that learns from both the ground truth label (hard target) and the soft target distribution to alleviate the difficulty of optimization and improve generalization capability. As the student network gradually advances, it turns to feed back knowledge to the assistant networks, thereby improving the teacher network in each iteration. Owing to such interactive designs, the DIL mechanism bridges the gap between the teacher and student networks and produces more superior capability for both networks. Objective and subjective experimental results demonstrate that both the teacher and student networks can generate more attractive shot sequences from music, thereby enhancing the viewing and listening experience. Jen-Chun Lin, Wen-Li Wei, Yen-Yu Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao |
ACM Multimedia | 3 |
| 2020 | MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose SynthesisabstractEstimating the 3D hand pose from a monocular RGB image is important but challenging. A solution is training on large-scale RGB hand images with accurate 3D hand keypoint annotations. However, it is too expensive in practice. Instead, we develop a learning-based approach to synthesize realistic, diverse, and 3D pose-preserving hand images under the guidance of 3D pose information. We propose a 3D-aware multi-modal guided hand generative network (MM-Hand), together with a novel geometry-based curriculum learning strategy. Our extensive experimental results demonstrate that the 3D-annotated images generated by MM-Hand qualitatively and quantitatively outperform existing options. Moreover, the augmented data can consistently improve the quantitative performance of the state-of-the-art 3D hand pose estimators on two benchmark datasets. The code will be available at https://github.com/ScottHoang/mm-hand. Zhenyu Wu 0002, Duc Hoang, Shih-Yao Lin 0001, Yusheng Xie, Liangjian Chen, Yen-Yu Lin, Zhangyang Wang, Wei Fan 0001 |
ACM Multimedia | 6 |
| 2020 | DGGAN: Depth-image Guided Generative Adversarial Networks for Disentangling RGB and Depth Images in 3D Hand Pose EstimationabstractEstimating 3D hand poses from RGB images is essential to a wide range of potential applications, but is challenging owing to substantial ambiguity in the inference of depth information from RGB images. State-of-the-art estimators address this problem by regularizing 3D hand pose estimation models during training to enforce the consistency between the predicted 3D poses and the ground-truth depth maps. However, these estimators rely on both RGB images and the paired depth maps during training. In this study, we propose a conditional generative adversarial network (GAN) model, called Depth-image Guided GAN (DGGAN), to generate realistic depth maps conditioned on the input RGB image, and use the synthesized depth maps to regularize the 3D hand pose estimation model, therefore eliminating the need for ground-truth depth maps. Experimental results on multiple benchmark datasets show that the synthesized depth maps produced by DGGAN are quite effective in regularizing the pose estimation model, yielding new state-of-the-art results in estimation accuracy, notably reducing the mean 3D endpoint errors (EPE) by 4.7%, 16.5%, and 6.8% on the RHD, STB and MHP datasets, respectively. Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin, Wei Fan 0001, Xiaohui Xie |
WACV | 4 |
| 2020 | VOSTR: Video Object Segmentation via Transferable Representations
Yi-Wen Chen, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Deep Co-Saliency Detection via Stacked Autoencoder-Enabled Fusion and Self-Trained CNNsabstractImage co-saliency detection via fusion-based or learning-based methods faces cross-cutting issues. Fusion-based methods often combine saliency proposals using a majority voting rule. Their performance hence highly depends on the quality and coherence of individual proposals. Learning-based methods typically require ground-truth annotations for training, which are not available for co-saliency detection. In this work, we present a two-stage approach to address these issues jointly. At the first stage, an unsupervised deep learning model with stacked autoencoder (SAE) is proposed to evaluate the quality of saliency proposals. It employs latent representations for image foregrounds, and auto-encodes foreground consistency and foreground-background distinctiveness in a discriminative way. The resultant model, SAE-enabled fusion (SAEF), can combine multiple saliency proposals to yield a more reliable saliency map. At the second stage, motivated by the fact that fusion often leads to over-smoothed saliency maps, we develop self-trained convolutional neural networks (STCNN) to alleviate this negative effect.STCNNtakes the saliency maps produced bySAEFas inputs. It propagates information from regions of high confidence to those of low confidence. During propagation, feature representations are distilled, resulting in sharper and better co-saliency maps. Our approach is comprehensively evaluated on three benchmarks, including MSRC, iCoseg, and Cosal2015, and performs favorably against the state-of-the-arts. In addition, we demonstrate that our method can be applied to object co-segmentation and object co-localization, achieving the state-of-the-art performance in both applications. Chung-Chi Tsai, Kuang-Jui Hsu, Yen-Yu Lin, Xiaoning Qian, Yung-Yu Chuang |
IEEE Trans. Multim. | 3 |
| 2019 | Deep Video Frame Interpolation Using Cyclic Frame GenerationabstractVideo frame interpolation algorithms predict intermediate frames to produce videos with higher frame rates and smooth view transitions given two consecutive frames as inputs. We propose that: synthesized frames are more reliable if they can be used to reconstruct the input frames with high quality. Based on this idea, we introduce a new loss term, the cycle consistency loss. The cycle consistency loss can better utilize the training data to not only enhance the interpolation results, but also maintain the performance better with less training data. It can be integrated into any frame interpolation network and trained in an end-to-end manner. In addition to the cycle consistency loss, we propose two extensions: motion linearity loss and edge-guided training. The motion linearity loss approximates the motion between two input frames to be linear and regularizes the training. By applying edge-guided training, we further improve results by integrating edge information into training. Both qualitative and quantitative experiments demonstrate that our model outperforms the state-of-the-art methods. The source codes of the proposed method and more experimental results will be available at https://github.com/alex04072000/CyclicGen. Yu-Lun Liu 0001, Yi-Tung Liao, Yen-Yu Lin, Yung-Yu Chuang |
AAAI | 3 |
| 2019 | What Makes You Look Like You: Learning an Inherent Feature Representation for Person Re-IdentificationabstractIn this work, we address person re-identification (ReID) by learning an inherent feature representation (inherent code) that is unique to each individual. This task is difficult because the appearance of a person may vary dramatically due to diverse factors, such as illuminations, viewpoints, and human pose changes. To tackle this issue, we propose new learning objectives to learn the inherent code for each person based on deep learning. Specifically, the proposed deep-net model is trained by jointly optimizing the multiple objectives that pulls the instances of the same person closer while pushing the instances belonging to different persons far from each other. Owing to such complementary designs, the deep-net model yields a robust code for each individual and hence better solve person ReID. Promising experimental results demonstrate the robustness and effectiveness of our proposed method. Wen-Li Wei, Jen-Chun Lin, Yen-Yu Lin, Hong-Yuan Mark Liao |
AVSS | 3 |
| 2019 | TAGAN: Tonality Aligned Generative Adversarial Networks for Realistic Hand Pose Synthesis
Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yufan Xue, Yen-Yu Lin, Xiaohui Xie, Wei Fan 0001 |
BMVC | 6 |
| 2019 | Referring Expression Object Segmentation with Caption-Aware Consistency
Yi-Wen Chen, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
BMVC | 4 |
| 2019 | CrDoCo: Pixel-Level Domain Transfer With Cross-Domain ConsistencyabstractUnsupervised domain adaptation algorithms aim to transfer the knowledge learned from one domain to another (e.g., synthetic to real images). The adapted representations often do not capture pixel-level domain shifts that are crucial for dense prediction tasks (e.g., semantic segmentation). In this paper, we present a novel pixel-wise adversarial domain adaptation algorithm. By leveraging image-to-image translation methods for data augmentation, our key insight is that while the translated images between domains may differ in styles, their predictions for the task should be consistent. We exploit this property and introduce a cross-domain consistency loss that enforces our adapted model to produce consistent predictions. Through extensive experimental results, we show that our method compares favorably against the state-of-the-art on a wide variety of unsupervised domain adaptation tasks. Yun-Chun Chen, Yen-Yu Lin, Ming-Hsuan Yang 0001, Jia-Bin Huang 0001 |
CVPR | 2 |
| 2019 | DeepCO3: Deep Instance Co-Segmentation by Co-Peak Search and Co-Saliency DetectionabstractIn this paper, we address a new task called instance co-segmentation. Given a set of images jointly covering object instances of a specific category, instance co-segmentation aims to identify all of these instances and segment each of them, i.e. generating one mask for each instance. This task is important since instance-level segmentation is preferable for humans and many vision applications. It is also challenging because no pixel-wise annotated training data are available and the number of instances in each image is unknown. We solve this task by dividing it into two sub-tasks, co-peak search and instance mask segmentation. In the former sub-task, we develop a CNN-based network to detect the co-peaks as well as co-saliency maps for a pair of images. A co-peak has two endpoints, one in each image, that are local maxima in the response maps and similar to each other. Thereby, the two endpoints are potentially covered by a pair of instances of the same category. In the latter subtask, we design a ranking function that takes the detected co-peaks and co-saliency maps as inputs and can select the object proposals to produce the final results. Our method for instance co-segmentation and its variant for object colocalization are evaluated on four datasets, and achieve favorable performance against the state-of-the-art methods. The source codes and the collected datasets are available at https://github.com/KuangJuiHsu/DeepCO3/ Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 2 |
| 2019 | FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation From a Single ImageabstractThis paper proposes a method for head pose estimation from a single image. Previous methods often predict head poses through landmark or depth estimation and would require more computation than necessary. Our method is based on regression and feature aggregation. For having a compact model, we employ the soft stagewise regression scheme. Existing feature aggregation methods treat inputs as a bag of features and thus ignore their spatial relationship in a feature map. We propose to learn a fine-grained structure mapping for spatially grouping features before aggregation. The fine-grained structure provides part-based information and pooled values. By utilizing learnable and non-learnable importance over the spatial location, different model variants can be generated and form a complementary ensemble. Experiments show that our method outperforms the state-of-the-art methods including both the landmark-free ones and the ones based on landmark or depth estimation. With only a single RGB frame as input, our method even outperforms methods utilizing multi-modality information (RGB-D, RGB-Time) on estimating the yaw angle. Furthermore, the memory overhead of our model is 100 times smaller than those of previous methods. Tsun-Yi Yang, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 3 |
| 2019 | LSIM: Ultra Lightweight Similarity Measurement for Mobile Graphics ApplicationsabstractPerceptual similarity measurement allows mobile applications to eliminate unnecessary computations without compromising visual experience. Existing pixel-wise measures incur significant overhead with increasing display resolutions and frame rates. This paper presents an ultra lightweight similarity measure called LSIM, which assesses the similarity between frames based on the transformation matrices of graphics objects. To evaluate its efficacy, we integrate LSIM into the Open Graphics Library and conduct experiments on an Android smartphone with various mobile 3D games. The results show that LSIM is highly correlated with the most widely used pixel-wise measure SSIM, yet three to five orders of magnitude faster. We also apply LSIM to a CPU-GPU governor to suppress the rendering of similar frames, thereby further reducing computation energy consumption by up to 27.3% while maintaining satisfactory visual quality. Yu-Chuan Chang, Wei-Ming Chen, Pi-Cheng Hsiu, Yen-Yu Lin, Tei-Wei Kuo |
DAC | 4 |
| 2019 | Recover and Identify: A Generative Dual Model for Cross-Resolution Person Re-IdentificationabstractPerson re-identification (re-ID) aims at matching images of the same identity across camera views. Due to varying distances between cameras and persons of interest, resolution mismatch can be expected, which would degrade person re-ID performance in real-world scenarios. To overcome this problem, we propose a novel generative adversarial network to address cross-resolution person re-ID, allowing query images with varying resolutions. By advancing adversarial learning techniques, our proposed model learns resolution-invariant image representations while being able to recover the missing details in low-resolution input images. The resulting features can be jointly applied for improving person re-ID performance due to preserving resolution invariance and recovering re-ID oriented discriminative details. Our experiments on five benchmark datasets confirm the effectiveness of our approach and its superiority over the state-of-the-art methods, especially when the input resolutions are unseen during training. Yu-Jhe Li, Yun-Chun Chen, Yen-Yu Lin, Xiaofei Du 0001, Yu-Chiang Frank Wang |
ICCV | 3 |
| 2019 | Weakly Supervised Instance Segmentation using the Bounding Box Tightness PriorabstractThis paper presents a weakly supervised instance segmentation method that consumes training data with tight bounding box annotations. The major difficulty lies in the uncertain figure-ground separation within each bounding box since there is no supervisory signal about it. We address the difficulty by formulating the problem as a multiple instance learning (MIL) task, and generate positive and negative bags based on the sweeping lines of each bounding box. The proposed deep model integrates MIL into a fully supervised instance segmentation network, and can be derived by the objective consisting of two terms, i.e., the unary term and the pairwise term. The former estimates the foreground and background areas of each bounding box while the latter maintains the unity of the estimated object masks. The experimental results show that our method performs favorably against existing weakly supervised methods and even surpasses some fully supervised methods for instance segmentation on the PASCAL VOC dataset. Cheng-Chun Hsu, Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Yung-Yu Chuang |
NeurIPS | 4 |
| 2019 | Weakly Supervised Salient Object Detection by Learning A Classifier-Driven Map GeneratorabstractTop-down saliency detection aims to highlight the regions of a specific object category, and typically relies on pixel-wise annotated training data. In this paper, we address the high cost of collecting such training data by a weakly supervised approach to object saliency detection, where only image-level labels, indicating the presence or absence of a target object in an image, are available. The proposed framework is composed of two collaborative CNN modules, an image-level classifier and a pixel-level map generator. While the former distinguishes images with objects of interest from the rest, the latter is learned to generate saliency maps by which the images masked by the maps can be better predicted by the former. In addition to the top-down guidance from class labels, the map generator is derived by also exploring other cues, including the background prior, superpixel- and object proposal-based evidence. The background prior is introduced to reduce false positives. Evidence from superpixels helps preserve sharp object boundaries. The clue from object proposals improves the integrity of highlighted objects. These different types of cues greatly regularize the training process and reduces the risk of overfitting, which happens frequently when learning CNN models with few training data. Experiments show that our method achieves superior results, even outperforming fully supervised methods. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
IEEE Trans. Image Process. | 2 |
| 2019 | Image Co-Saliency Detection and Co-Segmentation via Progressive Joint OptimizationabstractWe present a novel computational model for simultaneous image co-saliency detection and co-segmentation that concurrently explores the concepts of saliency and objectness in multiple images. It has been shown that the co-saliency detection via aggregating multiple saliency proposals by diverse visual cues can better highlight the salient objects; however, the optimal proposals are typically region-dependent and the fusion process often leads to blurred results. Co-segmentation can help preserve object boundaries, but it may suffer from complex scenes. To address these issues, we develop a unified method that addresses co-saliency detection and co-segmentation jointly via solving an energy minimization problem over a graph. Our method iteratively carries out the region-wise adaptive saliency map fusion and object segmentation to transfer useful information between the two complementary tasks. Through the optimization iterations, sharp saliency maps are gradually obtained to recover entire salient objects by referring to object segmentation, while these segmentations are progressively improved owing to the better saliency prior. We evaluate our method on four public benchmark data sets while comparing it to the state-of-the-art methods. Extensive experiments demonstrate that our method can provide consistently higher-quality results on both co-saliency detection and co-segmentation. Chung-Chi Tsai, Weizhi Li, Kuang-Jui Hsu, Xiaoning Qian, Yen-Yu Lin |
IEEE Trans. Image Process. | 5 |
| 2018 | Learning Adaptive Hidden Layers for Mobile Gesture RecognitionabstractThis paper addresses two obstacles hindering advances in accurate gesture recognition on mobile devices. First, gesture recognition performance is highly dependent on feature selection, but optimal features typically vary from gesture to gesture. Second, diverse user behaviors and mobile environments result in extremely large intra-class variations. We tackle these issues by introducing a new network layer, called an adaptive hidden layer (AHL), to generalize a hidden layer in deep neural networks and dynamically generate an activation map conditioned on the input. To this end, an AHL is composed of multiple neuron groups and an extra selector. The former compiles multi-modal features captured by mobile sensors, while the latter adaptively picks a plausible group for each input sample. The AHL is end-to-end trainable and can generalize an arbitrary subset of hidden layers. Through a series of AHLs, the great expressive power from exponentially many forward paths allows us to choose proper multi-modal features in a sample-specific fashion and resolve the problems caused by the unfavorable variations in mobile gesture recognition. The proposed approach is evaluated on a benchmark for gesture recognition and a newly collected dataset. Superior performance demonstrates its effectiveness. Ting-Kuei Hu, Yen-Yu Lin, Pi-Cheng Hsiu |
AAAI | 2 |
| 2018 | Deep Semantic Matching with Foreground Detection and Cycle-Consistency
Yun-Chun Chen, Po-Hsiang Huang, Li-Yu Yu, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001, Yen-Yu Lin |
ACCV (3) | 6 |
| 2018 | Unseen Object Segmentation in Videos via Transferable Representations
Yi-Wen Chen, Yi-Hsuan Tsai, Chu-Ya Yang, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
ACCV (4) | 4 |
| 2018 | Action Recognition with the Augmented MoCap Data using Neural Data Translation
Shih-Yao Lin 0001, Yen-Yu Lin |
BMVC | 2 |
| 2018 | Adversarial Learning for Semi-supervised Semantic Segmentation
Wei-Chih Hung, Yi-Hsuan Tsai, Yan-Ting Liou, Yen-Yu Lin, Ming-Hsuan Yang 0001 |
BMVC | 4 |
| 2018 | Unsupervised CNN-Based Co-saliency Detection with Graphical Optimization
Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Xiaoning Qian, Yung-Yu Chuang |
ECCV (5) | 3 |
| 2018 | Clustering Trajectories in Heterogeneous Representations for Video Event DetectionabstractTrajectories have been shown to be robust and widely used in surveillance video event analysis. They encode spatial and temporal evidence simultaneously. Hence, clustering trajectories in a video can detect representative events. How to effectively represent trajectories is thus essential to video event detection. However, no a single representation of trajectories suffices in increasingly complex video analysis tasks. To address this issue, this paper presents a hierarchical clustering algorithm for grouping trajectories in multiple heterogeneous representations. It turns out that our method can not only group trajectories of highly similar events but also identify rare events from the dominant events. Experimental results show that our method can retrieve both dominant events and rare events compared with the state-of-the-art methods, leading to a better performance. Wei-Cheng Wang, Yen-Yu Lin, Hsin-Wei Cheng, Chun-Rong Huang |
ICIP | 2 |
| 2018 | Co-attention CNNs for Unsupervised Object Co-segmentationabstractObject co-segmentation aims to segment the common objects in images. This paper presents a CNN-based method that is unsupervised and end-to-end trainable to better solve this task. Our method is unsupervised in the sense that it does not require any training data in the form of object masks but merely a set of images jointly covering objects of a specific class. Our method comprises two collaborative CNN modules, a feature extractor and a co-attention map generator. The former module extracts the features of the estimated objects and backgrounds, and is derived based on the proposed co-attention loss which minimizes inter-image object discrepancy while maximizing intra-image figure-ground separation. The latter module is learned to generated co-attention maps by which the estimated figure-ground segmentation can better fit the former module. Besides, the co-attention loss, the mask loss is developed to retain the whole objects and remove noises. Experiments show that our method achieves superior results, even outperforming the state-of-the-art, supervised methods. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
IJCAI | 2 |
| 2018 | SSR-Net: A Compact Soft Stagewise Regression Network for Age EstimationabstractThis paper presents a novel CNN model called Soft Stagewise Regression Network (SSR-Net) for age estimation from a single image with a compact model size. Inspired by DEX, we address age estimation by performing multi-class classification and then turning classification results into regression by calculating the expected values. SSR-Net takes a coarse-to-fine strategy and performs multi-class classification with multiple stages. Each stage is only responsible for refining the decision of its previous stage for more accurate age estimation. Thus, each stage performs a task with few classes and requires few neurons, greatly reducing the model size. For addressing the quantization issue introduced by grouping ages into classes, SSR-Net assigns a dynamic range to each age class by allowing it to be shifted and scaled according to the input face image. Both the multi-stage strategy and the dynamic range are incorporated into the formulation of soft stagewise regression. A novel network architecture is proposed for carrying out soft stagewise regression. The resultant SSR-Net model is very compact and takes only 0.32 MB. Despite its compact size, SSR-Net’s performance approaches those of the state-of-the-art methods whose model sizes are often more than 1500× larger. Tsun-Yi Yang, Yi-Hsuan Huang, Yen-Yu Lin, Pi-Cheng Hsiu, Yung-Yu Chuang |
IJCAI | 3 |
| 2018 | USEAQ: Ultra-Fast Superpixel Extraction via Adaptive Sampling From Quantized RegionsabstractWe present a novel and highly efficient superpixel extraction method called USEAQ to generate regular and compact superpixels in an image. To reduce the computational cost of iterative optimization procedures adopted in most recent approaches, the proposed USEAQ for superpixel generation works in a one-pass fashion. It firstly performs joint spatial and color quantizations and groups pixels into regions. It then takes into account the variations between regions, and adaptively samples one or a few superpixel candidates for each region. It finally employs maximum a posteriori (MAP) estimation to assign pixels to the most spatially consistent and perceptually similar superpixels. It turns out that the proposed USEAQ is quite efficient, and the extracted superpixels can precisely adhere to boundaries of objects. Experimental results show that USEAQ achieves better or equivalent performance compared to the stateof- the-art superpixel extraction approaches in terms of boundary recall, undersegmentation error, achievable segmentation accuracy, the average miss rate, average undersegmentation error, and average unexplained variation, and it is significantly faster than these approaches. Chun-Rong Huang, Wei-Cheng Wang, Wei-An Wang, Szu-Yu Lin, Yen-Yu Lin |
IEEE Trans. Image Process. | 5 |
| 2017 | Weakly Supervised Saliency Detection with A Category-Driven Map Generator
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
BMVC | 2 |
| 2017 | Deep Co-occurrence Feature Learning for Visual Object RecognitionabstractThis paper addresses three issues in integrating part-based representations into convolutional neural networks (CNNs) for object recognition. First, most part-based models rely on a few pre-specified object parts. However, the optimal object parts for recognition often vary from category to category. Second, acquiring training data with part-level annotation is labor-intensive. Third, modeling spatial relationships between parts in CNNs often involves an exhaustive search of part templates over multiple network streams. We tackle the three issues by introducing a new network layer, called co-occurrence layer. It can extend a convolutional layer to encode the co-occurrence between the visual parts detected by the numerous neurons, instead of a few pre-specified parts. To this end, the feature maps serve as both filters and images, and mutual correlation filtering is conducted between them. The co-occurrence layer is end-to-end trainable. The resultant co-occurrence features are rotation-and translation-invariant, and are robust to object deformation. By applying this new layer to the VGG-16 and ResNet-152, we achieve the recognition rates of 83.6% and 85.8% on the Caltech-UCSD bird benchmark, respectively. The source code is available at https://github.com/yafangshih/Deep-COOC. Ya-Fang Shih, Yang-Ming Yeh, Yen-Yu Lin, Ming-Fang Weng, Yi-Chang Lu, Yung-Yu Chuang |
CVPR | 3 |
| 2017 | Learning and inferring human actions with temporal pyramid features based on conditional random fieldsabstractFinding an effective way to represent human actions is yet an open problem because it usually requires taking evidences extracted from various temporal resolutions into account. A conventional way of representing an action employs temporally ordered fine-grained movements, e.g., key poses or subtle motions. Many existing approaches model actions by directly learning the transitional relationships between those fine-grained features. Yet, an action data may have many similar observations with occasional and irregular changes, which make commonly used fine-grained features less reliable. This paper presents a set of temporal pyramid features that enriches action representation with various levels of semantic granularities. For learning and inferring the proposed pyramid features, we adopt a discriminative model with latent variables to capture the hidden dynamics in each layer of the pyramid. Our method is evaluated on a Tai-Chi Chun dataset and a daily activities dataset. Both of them are collected by us. Experimental results demonstrate that our approach achieves more favorable performance than existing methods. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ICASSP | 2 |
| 2017 | Image co-saliency detection via locally adaptive saliency map fusionabstractCo-saliency detection aims at discovering the common and salient objects in multiple images. It explores not only intra-image but extra inter-image visual cues, and hence compensates the shortages in single-image saliency detection. The performance of co-saliency detection substantially relies on the explored visual cues. However, the optimal cues typically vary from region to region. To address this issue, we develop an approach that detects co-salient objects by region-wise saliency map fusion. Specifically, our approach takes intra-image appearance, inter-image correspondence, and spatial consistence into account, and accomplishes saliency detection with locally adaptive saliency map fusion via solving an energy optimization problem over a graph. It is evaluated on a benchmark dataset and compared to the state-of-the-art methods. Promising results demonstrate its effectiveness and superiority. Chung-Chi Tsai, Xiaoning Qian, Yen-Yu Lin |
ICASSP | 3 |
| 2017 | DeepCD: Learning Deep Complementary Descriptors for Patch RepresentationsabstractThis paper presents the DeepCD framework which learns a pair of complementary descriptors jointly for image patch representation by employing deep learning techniques. It can be achieved by taking any descriptor learning architecture for learning a leading descriptor and augmenting the architecture with an additional network stream for learning a complementary descriptor. To enforce the complementary property, a new network layer, called data-dependent modulation (DDM) layer, is introduced for adaptively learning the augmented network stream with the emphasis on the training data that are not well handled by the leading stream. By optimizing the proposed joint loss function with late fusion, the obtained descriptors are complementary to each other and their fusion improves performance. Experiments on several problems and datasets show that the proposed method1 is simple yet effective, outperforming state-of-the-art methods. Tsun-Yi Yang, Jo-Han Hsu, Yen-Yu Lin, Yung-Yu Chuang |
ICCV | 3 |
| 2017 | Deep dictionary learning for fine-grained image classificationabstractFine-grained image classification is quite challenging due to high inter-class similarity and large intra-class variations. Another issue is the small amount of training images with a large number of classes to be identified. To address the challenges, we propose a model for fine-grained image classification with its application to bird species recognition. Based on the features extracted by bilinear convolutional neural network (BCNN), we propose an on-line dictionary learning algorithm where the principle of sparsity is integrated into classification. The features extracted by BCNN encode pairwise neuron interaction in a translation-invariant manner. This property is valuable to fine-grained classification. The proposed algorithm for dictionary learning further carries out sparsity based classification, where training data can be represented with a less number of dictionary atoms. It alleviates the problems caused by insufficient training data, and makes classification much more efficient. Our approach is evaluated and compared with the state-of-the-art approaches on the CUB-200-2011 dataset. The promising experimental results demonstrate its efficacy and superiority. Yen-Yu Lin, Hong-Yuan Mark Liao |
ICIP | 2 |
| 2017 | Recognizing offensive tactics in broadcast basketball videos via key player detectionabstractWe address offensive tactic recognition in broadcast basketball videos. As a crucial component towards basketball video content understanding, tactic recognition is quite challenging because it involves multiple independent players, each of which has respective spatial and temporal variations. Motivated by the observation that most intra-class variations are caused by non-key players, we present an approach that integrates key player detection into tactic recognition. To save the annotation cost, our approach can work on training data with only video-level tactic annotation, instead of key players labeling. Specifically, this task is formulated as an MIL (multiple instance learning) problem where a video is treated as a bag with its instances corresponding to subsets of the five players. We also propose a representation to encode the spatio-temporal interaction among multiple players. It turns out that our approach not only effectively recognizes the tactics but also precisely detects the key players. Tsung-Yu Tsai, Yen-Yu Lin, Hong-Yuan Mark Liao, Shyh-Kang Jeng |
ICIP | 2 |
| 2017 | Learning deep and sparse feature representation for fine-grained object recognitionabstractIn this paper, we address fine-grained classification which is quite challenging due to high intra-class variations and subtle inter-class variations. Most modern approaches to fine-grained recognition are established based on convolutional neural networks (CNN). Despite the effectiveness, these approaches still suffer from two major problems. First, they highly rely on large sets of training data, but manually annotating numerous training data is expensive. Second, the learned feature presentations by these approaches are often of high dimensions, leading to less efficiency. To tackle the two problems, we present an approach where on-line dictionary learning is integrated into CNN. The dictionaries can be incrementally learned by leveraging a vast amount of weakly labeled data on the Internet. With these dictionaries, all the training and testing data can be sparsely represented. Our approach is evaluated and compared with the state-of-the-art approaches on the benchmark dataset, CUB-200-2011. The promising results demonstrate its superiority in both efficiency and accuracy. Yen-Yu Lin, Hong-Yuan Mark Liao |
ICME | 2 |
| 2017 | Segmentation guided local proposal fusion for co-saliency detectionabstractWe address two issues hindering existing image co-saliency detection methods. First, it has been shown that object boundaries can help improve saliency detection; But segmentation may suffer from significant intra-object variations. Second, aggregating the strength of different saliency proposals via fusion helps saliency detection covering entire object areas; However, the optimal saliency proposal fusion often varies from region to region, and the fusion process may lead to blurred results. Object segmentation and region-wise proposal fusion are complementary to help address the two issues if we can develop a unified approach. Our proposed segmentation-guided locally adaptive proposal fusion is the first of such efforts for image co-saliency detection to the best of our knowledge. Specifically, it leverages both object-aware segmentation evidence and region-wise consensus among saliency proposals via solving a joint co-saliency and co-segmentation energy optimization problem over a graph. Our approach is evaluated on a benchmark dataset and compared to the state-of-the-art methods. Promising results demonstrate its effectiveness and superiority. Chung-Chi Tsai, Xiaoning Qian, Yen-Yu Lin |
ICME | 3 |
| 2017 | Recognizing Human Actions with Outlier Frames by Observation Filtering and CompletionabstractThis article addresses the problem of recognizing partially observed human actions. Videos of actions acquired in the real world often contain corrupt frames caused by various factors. These frames may appear irregularly, and make the actions only partially observed. They change the appearance of actions and degrade the performance of pretrained recognition systems. In this article, we propose an approach to address the corrupt-frame problem without knowing their locations and durations in advance. The proposed approach includes two key components: outlier filtering and observation completion . The former identifies and filters out unobserved frames, and the latter fills up the filtered parts by retrieving coherent alternatives from training data. Hidden Conditional Random Fields (HCRFs) are then used to recognize the filtered and completed actions. Our approach has been evaluated on three datasets, which contain both fully observed actions and partially observed actions with either real or synthetic corrupt frames. The experimental results show that our approach performs favorably against the other state-of-the-art methods, especially when corrupt frames are present. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2016 | Progressive Feature Matching with Alternate Descriptor Selection and Correspondence EnrichmentabstractWe address two difficulties in establishing an accurate system for image matching. First, image matching relies on the descriptor for feature extraction, but the optimal descriptor often varies from image to image, or even patch to patch. Second, conventional matching approaches carry out geometric checking on a small set of correspondence candidates due to the concern of efficiency. It may result in restricted performance in recall. We aim at tackling the two issues by integrating adaptive descriptor selection and progressive candidate enrichment into image matching. We consider that the two integrated components are complementary: The high-quality matching yielded by adaptively selected descriptors helps in exploring more plausible candidates, while the enriched candidate set serves as a better reference for descriptor selection. It motivates us to formulate image matching as a joint optimization problem, in which adaptive descriptor selection and progressive correspondence enrichment are alternately conducted. Our approach is comprehensively evaluated and compared with the state-of-the-art approaches on two benchmarks. The promising results manifest its effectiveness. Yuan-Ting Hu, Yen-Yu Lin |
CVPR | 2 |
| 2016 | Accumulated Stability Voting: A Robust Descriptor from Descriptors of Multiple ScalesabstractThis paper proposes a novel local descriptor through accumulated stability voting (ASV). The stability of feature dimensions is measured by their differences across scales. To be more robust to noise, the stability is further quantized by thresholding. The principle of maximum entropy is utilized for determining the best thresholds for maximizing discriminant power of the resultant descriptor. Accumulating stability renders a real-valued descriptor and it can be converted into a binary descriptor by an additional thresholding process. The real-valued descriptor attains high matching accuracy while the binary descriptor makes a good compromise between storage and accuracy. Our descriptors are simple yet effective, and easy to implement. In addition, our descriptors require no training. Experiments on popular benchmarks demonstrate the effectiveness of our descriptors and their superiority to the state-of-the-art descriptors. Tsun-Yi Yang, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 2 |
| 2016 | Precise player segmentation in team sports videos using contrast-aware co-segmentationabstractPlayer segmentation in team sports videos is challenging but crucial to video semantic understanding, such as player interaction identification and tactic analysis. We leverage the appearance similarity among players of the same team, and cast this task as a co-segmentation problem. In this way, the extra knowledge shared across players significantly reduces unfavorable uncertainty in segmenting individual players. We are also aware that the performance of co-segmentation highly depends on the used features, and further propose a contrast-based approach to estimate the discriminant power of each feature in an unsupervised manner. It turns out that our approach can properly fuse features by assigning higher weights to discriminant ones, and result in remarkable performance gains. The promising results on segmenting basketball players manifest the effectiveness of our approach. Tsung-Yu Tsai, Yen-Yu Lin, Hong-Yuan Mark Liao, Shyh-Kang Jeng |
ICASSP | 2 |
| 2016 | USEQ: Ultra-fast superpixel extraction via quantizationabstractWe propose a novel superpixel extraction method named USEQ to generate regular and compact superpixels. To reduce the computational burden of iterative optimization procedures used in most recent approaches, the spatial and color quantizations are performed in advance to represent pixels and superpixels. Maximum a posteriori estimation in both pixel and region levels is then adopted to aggregate pixels into spatially and visually coherent superpixels. The resultant superpixels are extremely efficient to generate and can more precisely adhere to object boundaries. Compared to the state-of-the-art approaches to superpixel extraction, USEQ can achieve better or competitive performance in terms of boundary recall, undersegmentation error and achievable segmentation accuracy, and is significantly faster than these approaches. Chun-Rong Huang, Wei-An Wang, Szu-Yu Lin, Yen-Yu Lin |
ICPR | 4 |
| 2016 | Learning Discriminatively Reconstructed Source Data for Object Recognition With Few ExamplesabstractWe aim at improving the object recognition with few training data in the target domain by leveraging abundant auxiliary data in the source domain. The major issue obstructing knowledge transfer from source to target is the limited correlation between the two domains. Transferring irrelevant information from the source domain usually leads to performance degradation in the target domain. To address this issue, we propose a transfer learning framework with the two key components, such as discriminative source data reconstruction and dual-domain boosting. The former correlates the two domains via reconstructing source data by target data in a discriminative manner. The latter discovers and delivers only knowledge shared by the target data and the reconstructed source data. Hence, it facilitates recognition in the target. The promising experimental results on three benchmarks of object recognition demonstrate the effectiveness of our approach. Pai-Heng Hsiao, Feng-Ju Chang, Yen-Yu Lin |
IEEE Trans. Image Process. | 3 |
| 2015 | Robust image alignment with multiple feature descriptors and matching-guided neighborhoodsabstractThis paper addresses two issues hindering the advances in accurate image alignment. First, he performance of descriptor-based approaches to image alignment relies on the chosen descriptor, but the optimal descriptor typically varies from image to image, or even pixel to pixel. Second, the neighborhood structure for smoothness enforcement is usually predefined before alignment. However, object boundaries are often better discovered during alignment. The proposed approach tackles the two issues by adaptive descriptor selection and dynamic neighborhood construction. Specifically we associate each pixel to be aligned with an affine transformation, and integrate the learning of the pixel-specific transformations into image alignment. The transformations serve as the common domain for descriptor fusion, since the local consensus of each descriptor can be estimated by accessing the corresponding affine transformation t allows us to pick the most plausible descriptor for aligning each pixel. On the other hand more object-aware neighborhoods can be produced by referencing the consistency between the learned affine transformations of neighboring pixels. The promising results on popular image alignment benchmarks manifests the effectiveness of our approach. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 2 |
| 2015 | Blur kernel estimation using normalized color-line priorsabstractThis paper proposes a single-image blur kernel estimation algorithm that utilizes the normalized color-line prior to restore sharp edges without altering edge structures or enhancing noise. The proposed prior is derived from the color-line model, which has been successfully applied to non-blind deconvolution and many computer vision problems. In this paper, we show that the original color-line prior is not effective for blur kernel estimation and propose a normalized color-line prior which can better enhance edge contrasts. By optimizing the proposed prior, our method gradually enhances the sharpness of the intermediate patches without using heuristic filters or external patch priors. The intermediate patches can then guide the estimation of the blur kernel. A comprehensive evaluation on a large image deblurring dataset shows that our algorithm achieves the state-of-the-art results. Wei-Sheng Lai, Jian-Jiun Ding, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 3 |
| 2015 | Co-Segmentation Guided Hough Transform for Robust Feature MatchingabstractWe present an algorithm that integrates image co-segmentation into feature matching, and can robustly yield accurate and dense feature correspondences. Inspired by the fact that correct feature correspondences on the same object typically have coherent transformations, we cast the task of feature matching as a density estimation problem in the homography space. Specifically, we project the homographies of correspondence candidates into the parametric Hough space, in which geometric verification of correspondences can be activated by voting. The precision of matching is then boosted. On the other hand, we leverage image co-segmentation, which discovers object boundaries, to determine relevant voters and speed up Hough voting. In addition, correspondence enrichment can be achieved by inferring the concerted homographies that are propagated between the features within the same segments. The recall is hence increased. In our approach, feature matching and image co-segmentation are tightly coupled. Through an iterative optimization process, more and more correct correspondences are detected owing to object boundaries revealed by co-segmentation. The proposed approach is comprehensively evaluated. Promising experimental results on four datasets manifest its effectiveness. Hsin-Yi Chen, Yen-Yu Lin, Bing-Yu Chen 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Linear Spectral Mixture Analysis via Multiple-Kernel Learning for Hyperspectral Image ClassificationabstractLinear spectral mixture analysis (LSMA) has received wide interests for spectral unmixing in the remote sensing community. This paper introduces a framework called multiplekernel learning-based spectral mixture analysis (MKL-SMA) that integrates a newly proposed MKL method into the training process of LSMA. MKL-SMA allows us to adopt a set of nonlinear basis kernels to better characterize the data so that it can enrich the discriminant capability in classification. Because a single kernel is often insufficient to well present all the data characteristics, MKL-SMA has the advantage of providing a broader range of representation flexibilities; it also eases the kernel selection process because the kernel combination parameters can be learned automatically. Unlike most MKL approaches where complex nonlinear optimization problems are involved in their training process, we derived a closed-form solution of the kernel combination parameters in MKL-SMA. Our method is thus efficient for training and easy to implement. The usefulness of MKL-SMA is demonstrated by conducting real hyperspectral image experiments for performance evaluation. Promising results manifest the effectiveness of the proposed MKL-SMA. Keng-Hao Liu, Yen-Yu Lin, Chu-Song Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Matching Images With Multiple Descriptors: An Unsupervised Approach for Locally Adaptive Descriptor SelectionabstractWith the aim to improve the performance of feature matching, we present an unsupervised approach for adaptive description selection in the space of homographies. Inspired by the observation that the homographies of correct feature correspondences vary smoothly along the spatial domain, our approach stands on the unsupervised nature of feature matching, and can choose a good descriptor locally for matching each feature point, instead of using one global descriptor. To this end, the homography space serves as the domain for selecting various heterogeneous descriptors. Correspondences obtained by any descriptors are considered as points in the space, and their geometric coherence and spatial continuity are measured via computing the geodesic distances. In this way, mutual verification across different descriptors is allowed, and correct correspondences will be highlighted with a high degree of consistency short geodesic distances here. It follows that one-class SVM can be applied to identifying these correct correspondences, and achieves adaptive descriptor selection. The proposed approach is comprehensively compared with the state-of-the-art approaches, and evaluated on five benchmarks of image matching. The promising results manifest its effectiveness. Yuan-Ting Hu, Yen-Yu Lin, Hsin-Yi Chen, Kuang-Jui Hsu, Bing-Yu Chen 0004 |
IEEE Trans. Image Process. | 2 |
| 2015 | Robust Action Recognition via Borrowing Information Across Video ModalitiesabstractThe recent advances in imaging devices have opened the opportunity of better solving the tasks of video content analysis and understanding. Next-generation cameras, such as the depth or binocular cameras, capture diverse information, and complement the conventional 2D RGB cameras. Thus, investigating the yielded multimodal videos generally facilitates the accomplishment of related applications. However, the limitations of the emerging cameras, such as short effective distances, expensive costs, or long response time, degrade their applicability, and currently make these devices not online accessible in practical use. In this paper, we provide an alternative scenario to address this problem, and illustrate it with the task of recognizing human actions. In particular, we aim at improving the accuracy of action recognition in RGB videos with the aid of one additional RGB-D camera. Since RGB-D cameras, such as Kinect, are typically not applicable in a surveillance system due to its short effective distance, we instead offline collect a database, in which not only the RGB videos but also the depth maps and the skeleton data of actions are available jointly. The proposed approach can adapt the interdatabase variations, and activate the borrowing of visual knowledge across different video modalities. Each action to be recognized in RGB representation is then augmented with the borrowed depth and skeleton features. Our approach is comprehensively evaluated on five benchmark data sets of action recognition. The promising results manifest that the borrowed information leads to remarkable boost in recognition accuracy. Nick C. Tang, Yen-Yu Lin, Ju-Hsuan Hua, Shih-En Wei, Ming-Fang Weng, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 2 |
| 2015 | Cross-Camera Knowledge Transfer for Multiview People CountingabstractWe present a novel two-pass framework for counting the number of people in an environment, where multiple cameras provide different views of the subjects. By exploiting the complementary information captured by the cameras, we can transfer knowledge between the cameras to address the difficulties of people counting and improve the performance. The contribution of this paper is threefold. First, normalizing the perspective of visual features and estimating the size of a crowd are highly correlated tasks. Hence, we treat them as a joint learning problem. The derived counting model is scalable and it provides more accurate results than existing approaches. Second, we introduce an algorithm that matches groups of pedestrians in images captured by different cameras. The results provide a common domain for knowledge transfer, so we can work with multiple cameras without worrying about their differences. Third, the proposed counting system is comprised of a pair of collaborative regressors. The first one determines the people count based on features extracted from intracamera visual information, whereas the second calculates the residual by considering the conflicts between intercamera predictions. The two regressors are elegantly coupled and provide an accurate people counting system. The results of experiments in various settings show that, overall, our approach outperforms comparable baseline methods. The significant performance improvement demonstrates the effectiveness of our two-pass regression framework. Nick C. Tang, Yen-Yu Lin, Ming-Fang Weng, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 2 |
| 2014 | Multiple Structured-Instance Learning for Semantic Segmentation with Uncertain Training DataabstractWe present an approach MSIL-CRF that incorporates multiple instance learning (MIL) into conditional random fields (CRFs). It can generalize CRFs to work on training data with uncertain labels by the principle of MIL. In this work, it is applied to saving manual efforts on annotating training data for semantic segmentation. Specifically, we consider the setting in which the training dataset for semantic segmentation is a mixture of a few object segments and an abundant set of objects' bounding boxes. Our goal is to infer the unknown object segments enclosed by the bounding boxes so that they can serve as training data for semantic segmentation. To this end, we generate multiple segment hypotheses for each bounding box with the assumption that at least one hypothesis is close to the ground truth. By treating a bounding box as a bag with its segment hypotheses as structured instances, MSIL-CRF selects the most likely segment hypotheses by leveraging the knowledge derived from both the labeled and uncertain training data. The experimental results on the Pascal VOC segmentation task demonstrate that MSIL-CRF can provide effective alternatives to manually labeled segments for semantic segmentation. Feng-Ju Chang, Yen-Yu Lin, Kuang-Jui Hsu |
CVPR | 2 |
| 2014 | Depth and Skeleton Associated Action Recognition without Online Accessible RGB-D CamerasabstractThe recent advances in RGB-D cameras have allowed us to better solve increasingly complex computer vision tasks. However, modern RGB-D cameras are still restricted by the short effective distances. The limitation may make RGB-D cameras not online accessible in practice, and degrade their applicability. We propose an alternative scenario to address this problem, and illustrate it with the application to action recognition. We use Kinect to offline collect an auxiliary, multi-modal database, in which not only the RGB videos but also the depth maps and skeleton structures of actions of interest are available. Our approach aims to enhance action recognition in RGB videos by leveraging the extra database. Specifically, it optimizes a feature transformation, by which the actions to be recognized can be concisely reconstructed by entries in the auxiliary database. In this way, the inter-database variations are adapted. More importantly, each action can be augmented with additional depth and skeleton images retrieved from the auxiliary database. The proposed approach has been evaluated on three benchmarks of action recognition. The promising results manifest that the augmented depth and skeleton features can lead to remarkable boost in recognition accuracy. Yen-Yu Lin, Ju-Hsuan Hua, Nick C. Tang, Min-Hung Chen, Hong-Yuan Mark Liao |
CVPR | 1 |
| 2014 | Human action recognition using associated depth and skeleton informationabstractThe recent advances in imaging devices have opened the opportunity of better solving computer vision tasks. The next-generation cameras, such as the depth or binocular cameras, capture diverse information, and complement the conventional 2D RGB cameras. Thus, investigating the yielded multi-modal images generally facilitates the accomplishment of related applications. However, the limitations of these devices, such as short effective distances, expensive costs, or long response time, degrade their applicability in practical use. Addressing this problem in this work, we aim at action recognition in RGB videos with the aid of Kinect. We improve recognition accuracy by leveraging information derived from an offline collected database, in which not only the RGB but also the depth and skeleton images of actions are available. Our approach adapts the inter-database variations, and enables the sharing of visual knowledge across different image modalities. Each action instance for recognition in RGB representation is then augmented with the borrowed depth and skeleton features. Nick C. Tang, Yen-Yu Lin, Ju-Hsuan Hua, Ming-Fang Weng, Hong-Yuan Mark Liao |
ICASSP | 2 |
| 2014 | Video Saliency Map Detection by Dominant Camera Motion RemovalabstractWe present a trajectory-based approach to detect salient regions in videos by dominant camera motion removal. Our approach is designed in a general way so that it can be applied to videos taken by either stationary or moving cameras without any prior information. Moreover, multiple salient regions of different temporal lengths can also be detected. To this end, we extract a set of spatially and temporally coherent trajectories of keypoints in a video. Then, velocity and acceleration entropies are proposed to represent the trajectories. In this way, long-term object motions are exploited to filter out short-term noises, and object motions of various temporal lengths can be represented in the same way. On the other hand, we are inspired by the observation that the trajectories in backgrounds, i.e., the nonsalient trajectories, are usually consistent with the dominant camera motion no matter whether the camera is stationary or not. We make use of this property to develop a unified approach to saliency generation for both stationary and moving cameras. Specifically, one-class SVM is employed to remove the consistent trajectories in motion. It follows that the salient regions could be highlighted by applying a diffusion process to the remaining trajectories. In addition, we create a set of manually annotated ground truth on the collected videos. The annotated videos are then used for performance evaluation and comparison. The promising results on various types of videos demonstrate the effectiveness and great applicability of our approach. Chun-Rong Huang, Yun-Jung Chang, Zhi-Xiang Yang, Yen-Yu Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Augmented Multiple Instance Regression for Inferring Object Contours in Bounding BoxesabstractIn this paper, we address the problem of the high annotation cost of acquiring training data for semantic segmentation. Most modern approaches to semantic segmentation are based upon graphical models, such as the conditional random fields, and rely on sufficient training data in form of object contours. To reduce the manual effort on pixel-wise annotating contours, we consider the setting in which the training data set for semantic segmentation is a mixture of a few object contours and an abundant set of bounding boxes of objects. Our idea is to borrow the knowledge derived from the object contours to infer the unknown object contours enclosed by the bounding boxes. The inferred contours can then serve as training data for semantic segmentation. To this end, we generate multiple contour hypotheses for each bounding box with the assumption that at least one hypothesis is close to the ground truth. This paper proposes an approach, called augmented multiple instance regression (AMIR), that formulates the task of hypothesis selection as the problem of multiple instance regression (MIR), and augments information derived from the object contours to guide and regularize the training process of MIR. In this way, a bounding box is treated as a bag with its contour hypotheses as instances, and the positive instances refer to the hypotheses close to the ground truth. The proposed approach has been evaluated on the Pascal VOC segmentation task. The promising results demonstrate that AMIR can precisely infer the object contours in the bounding boxes, and hence provide effective alternatives to manually labeled contours for semantic segmentation. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
IEEE Trans. Image Process. | 2 |
| 2014 | Per-Cluster Ensemble Kernel Learning for Multi-Modal Image Clustering With Group-Dependent Feature SelectionabstractIn this paper, we present a clustering approach, MK-SOM, that carries out cluster-dependent feature selection, and partitions images with multiple feature representations into clusters. This work is motivated by the observations that human visual systems (HVS) can receive various kinds of visual cues for interpreting the world. Images identified by HVS as the same category are typically coherent to each other in certain crucial visual cues, but the crucial cues vary from category to category. To account for this observation and bridge the semantic gap, the proposed MK-SOM integrates multiple kernel learning (MKL) into the training process of self-organizing map (SOM), and associates each cluster with a learnable, ensemble kernel. Hence, it can leverage information captured by various image descriptors, and discoveries the cluster-specific characteristics via learning the per-cluster ensemble kernels. Through the optimization iterations, cluster structures are gradually revealed via the features specified by the learned ensemble kernels, while the quality of these ensemble kernels is progressively improved owing to the coherent clusters by enforcing SOM. Besides, MK-SOM allows the introduction of side information to improve performance, and it hence provides a new perspective of applying MKL to address both unsupervised and semi-supervised clustering tasks. Our approach is comprehensively evaluated in the two applications. The superior and promising results manifest its effectiveness. Jeng-Tsung Tsai, Yen-Yu Lin, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 2 |
| 2013 | Robust Feature Matching with Alternate Hough and Inverted Hough TransformsabstractWe present an algorithm that carries out alternate Hough transform and inverted Hough transform to establish feature correspondences, and enhances the quality of matching in both precision and recall. Inspired by the fact that nearby features on the same object share coherent homographies in matching, we cast the task of feature matching as a density estimation problem in the Hough space spanned by the hypotheses of homographies. Specifically, we project all the correspondences into the Hough space, and determine the correctness of the correspondences by their respective densities. In this way, mutual verification of relevant correspondences is activated, and the precision of matching is boosted. On the other hand, we infer the concerted homographies propagated from the locally grouped features, and enrich the correspondence candidates for each feature. The recall is hence increased. The two processes are tightly coupled. Through iterative optimization, plausible enrichments are gradually revealed while more correct correspondences are detected. Promising experimental results on three benchmark datasets manifest the effectiveness of the proposed approach. Hsin-Yi Chen, Yen-Yu Lin, Bing-Yu Chen 0004 |
CVPR | 2 |
| 2013 | Multi-view face detection in videos with online adaptationabstractMost learning-based approaches to face detection suffer from the problem of performance degradation on faces that are not covered by training data. However, including all variations of faces in training is practically infeasible due to the scalability restriction of machine learning algorithms and expensive manual labeling. In this work, we focus on face detection in videos, and alleviate this problem by exploiting strong correlation among video frames. We augment a pre-trained multiview face detection with an incrementally derived Gaussian process regressor. The regressor can extract and propagate visual knowledge across frames, and adapts the detector to handle unseen faces. Testing on two datasets, the promising results manifest the effectiveness of the proposed approach. Yao-Chuan Chang, Yen-Yu Lin, Hong-Yuan Mark Liao |
ICIP | 2 |
| 2012 | Cross-Database Transfer Learning via Learnable and Discriminant Error-Correcting Output Codes
Feng-Ju Chang, Yen-Yu Lin, Ming-Fang Weng |
ACCV (1) | 2 |
| 2012 | Knowledge Leverage from Contours to Bounding Boxes: A Concise Approach to Annotation
Jie-Zhi Cheng, Feng-Ju Chang, Kuang-Jui Hsu, Yen-Yu Lin |
ACCV (1) | 4 |
| 2012 | Action recognition using instance-specific and class-consistent cuesabstractWe aim to resolve the difficulties of action recognition arising from the large intra-class variations. These unfavorable variations make it infeasible to represent one action instance by other ones of the same action. We hence propose to extract both instance-specific and class-consistent features to facilitate action recognition. Specifically, the instance-specific features explore the self-similarities among frames of each video instance, while class-consistent features summarize within-class similarities. We introduce a generative formulation to combine the two diverse types of features. The experimental results demonstrate the effectiveness of our approach. Chin-An Lin, Yen-Yu Lin, Hong-Yuan Mark Liao, Shyh-Kang Jeng |
ICIP | 2 |
| 2012 | Cluster-dependent feature selection by multiple kernel self-organizing map
Kuan-Chieh Huang, Yen-Yu Lin, Jie-Zhi Cheng |
ICPR | 2 |
| 2012 | The acousticvisual emotion guassians model for automatic generation of music videoabstractThis paper presents a novel content-based system that utilizes the perceived emotion of multimedia content as a bridge to connect music and video. Specifically, we propose a novel machine learning framework, called Acousticvisual Emotion Guassians (AVEG), to jointly learn the tripartite relationship among music, video, and emotion from an emotion-annotated corpus of music videos. For a music piece (or a video sequence), the AVEG model is applied to predict its emotion distribution in a stochastic emotion space from the corresponding low-level acoustic (resp. visual) features. Finally, music and video are matched by measuring the similarity between the two corresponding emotion distributions, based on a distance measure such as KL divergence. Ju-Chiang Wang, Yi-Hsuan Yang, I-Hong Jhuo, Yen-Yu Lin, Hsin-Min Wang |
ACM Multimedia | 4 |
| 2012 | Visual knowledge transfer among multiple cameras for people counting with occlusion handlingabstractWe present a framework to count the number of people in an environment where multiple cameras with different angles of view are available. We consider the visual cues captured by each camera as a knowledge source, and carry out cross-camera knowledge transfer to alleviate the difficulties of people counting, such as partial occlusions, low-quality images, clutter backgrounds, and so on. Specifically, this work distinguishes itself with the following contributions. First, we overcome the variations of multiple heterogeneous cameras with different perspective settings by matching the same groups of pedestrians taken by these cameras, and present an algorithm for accomplishing cross-camera correspondence. Second, the proposed counting model is composed of a pair of collaborative regressors. While one regressor measures people counts by the features extracted from intra-camera visual evidences, the other recovers the yielded residual by taking the conflicts among inter-camera predictions into account. The two regressors are elegantly coupled, and jointly lead to an accurate counting system. Additionally, we provide a set of manually annotated pedestrian labels on the PETS 2010 videos for performance evaluation. Our approach is comprehensively tested in various settings and compared with competitive baselines. The significant improvement in performance manifests the effectiveness of the proposed approach. Ming-Fang Weng, Yen-Yu Lin, Nick C. Tang, Hong-Yuan Mark Liao |
ACM Multimedia | 2 |
| 2011 | Multiple Kernel Learning for Dimensionality ReductionabstractIn solving complex visual learning tasks, adopting multiple descriptors to more precisely characterize the data has been a feasible way for improving performance. The resulting data representations are typically high-dimensional and assume diverse forms. Hence, finding a way of transforming them into a unified space of lower dimension generally facilitates the underlying tasks such as object recognition or clustering. To this end, the proposed approach (termed MKL-DR) generalizes the framework of multiple kernel learning for dimensionality reduction, and distinguishes itself with the following three main contributions: first, our method provides the convenience of using diverse image descriptors to describe useful characteristics of various aspects about the underlying data. Second, it extends a broad set of existing dimensionality reduction techniques to consider multiple kernel learning, and consequently improves their effectiveness. Third, by focusing on the techniques pertaining to dimensionality reduction, the formulation introduces a new class of applications with the multiple kernel learning framework to address not only the supervised learning problems but also the unsupervised and semi-supervised ones. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Clustering Complex Data with Group-Dependent Feature Selection
Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (6) | 1 |
| 2009 | Efficient discriminative local learning for object recognitionabstractAlthough object recognition methods based on local learning can reasonably resolve the difficulties caused by the large variations in images from the same category, the high risk of overfitting and the heavy computational cost in training numerous local models (classifiers or distance functions) often limit their applicability. To address these two unpleasant issues, we cast the multiple, independent training processes of local models as a correlative multi-task learning problem, and design a new boosting algorithm to accomplish it. Specifically, we establish a parametric space where these local models lie and spread as a manifold-like structure, and use boosting to perform local model training by completing the manifold embedding. Via sharing the common embedding space, the learning of each local model can be properly regularized by the extra knowledge from other models, while the training time is also significantly reduced. Experimental results on two benchmark datasets, Caltech-101 and VOC 2007, support that our approach not only achieves promising recognition rates but also gives a two order speed-up in realizing local learning. Yen-Yu Lin, Jyun-Fan Tsai, Tyng-Luh Liu |
ICCV | 1 |
| 2008 | Dimensionality Reduction for Data in Multiple Feature RepresentationsabstractIn solving complex visual learning tasks, adopting multiple descriptors to more precisely characterize the data has been a feasible way for improving performance. These representations are typically high dimensional and assume diverse forms. Thus finding a way to transform them into a unified space of lower dimension generally facilitates the underlying tasks, such as object recognition or clustering. We describe an approach that incorporates multiple kernel learning with dimensionality reduction (MKL-DR). While the proposed framework is flexible in simultaneously tackling data in various feature representations, the formulation itself is general in that it is established upon graph embedding. It follows that any dimensionality reduction techniques explainable by graph embedding can be generalized by our method to consider data in multiple feature representations. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
NIPS | 1 |
| 2007 | Local Ensemble Kernel Learning for Object Category RecognitionabstractThis paper describes a local ensemble kernel learning technique to recognize/classify objects from a large number of diverse categories. Due to the possibly large intraclass feature variations, using only a single unified kernel-based classifier may not satisfactorily solve the problem. Our approach is to carry out the recognition task with adaptive ensemble kernel machines, each of which is derived from proper localization and regularization. Specifically, for each training sample, we learn a distinct ensemble kernel constructed in a way to give good classification performance for data falling within the corresponding neighborhood. We achieve this effect by aligning each ensemble kernel with a locally adapted target kernel, followed by smoothing out the discrepancies among kernels of nearby data. Our experimental results on various image databases manifest that the technique to optimize local ensemble kernels is effective and consistent for object recognition. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
CVPR | 1 |
| 2005 | Robust Face Detection with Multi-Class BoostingabstractWith the aim to design a general learning framework for detecting faces of various poses or under different lighting conditions, we are motivated to formulate the task as a classification problem over data of multiple classes. Specifically, our approach focuses on a new multi-class boosting algorithm, called MBHboost, and its integration with a cascade structure for effectively performing face detection. There are three main advantages of using MBHboost: 1) each MBH weak learner is derived by sharing a good projection direction such that each class of data has its own decision boundary; 2) the proposed boosting algorithm is established based on an optimal criterion for multi-class classification; and 3) since MBHboost is flexible with respect to the number of classes, it turns out that it is possible to use only one single boosted cascade for the multi-class detection. All these properties give rise to a robust system to detect faces efficiently and accurately. Yen-Yu Lin, Tyng-Luh Liu |
CVPR (1) | 1 |
| 2005 | Shape recognition using fast boosted filteringabstractWe address the problem of recognizing 2-D shapes in images via multi-class classifications. Our approach has three key elements. First, a signed distance transform is introduced to represent a shape more informatively. Second, a filter bank is generated such that its filters can capture multiple-scale local and global features between two shapes of different classes. We then apply boosting to combine useful filters to construct discriminant classifiers. Third, in implementing our system, a new classification architecture is developed to accomplish multi-class recognition. To examine the claimed efficiencies, we consider an example of document recognition by pinpointing the strengths of our method through experimental results and comparisons. Yen-Yu Lin, Tyng-Luh Liu |
ICIP (2) | 1 |
| 2005 | Semantic manifold learning for image retrievalabstractLearning the user's semantics for CBIR involves two different sources of information: the similarity relations entailed by the content-based features, and the relevance relations specified in the feedback. Given that, we propose an augmented relation embedding (ARE) to map the image space into a semantic manifold that faithfully grasps the user's preferences. Besides ARE, we also look into the issues of selecting a good feature set for improving the retrieval performance. With these two aspects of efforts we have established a system that yields far better results than those previously reported. Overall, our approach can be characterized by three key properties: 1) The framework uses one relational graph to describe the similarity relations, and the other two to encode the relevant/irrelevant relations indicated in the feedback. 2) With the relational graphs so defined, learning a semantic manifold can be transformed into solving a constrained optimization problem, and is reduced to the ARE algorithm accounting for both the representation and the classification points of views. 3) An image representation based on augmented features is introduced to couple with the ARE learning. The use of these features is significant in capturing the semantics concerning different scales of image regions. We conclude with experimental results and comparisons to demonstrate the effectiveness of our method. Yen-Yu Lin, Tyng-Luh Liu, Hwann-Tzong Chen |
ACM Multimedia | 1 |
| 2004 | Fast Object Detection with Occlusions
Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (1) | 1 |