EDBT 2026 Demo / reviewers in the wild / expert
Jun-Cheng Chen
dblp:77/251
· DBLP profile ↗
79ranked-venue papers
9as first author
50since 2021 · last 2026
0000-0002-0209-8932ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 8 first-author · 43 since 2021Artificial intelligence and machine learning · 30 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StyleDiT: A Unified Framework for Diverse Child and Partner Faces Synthesis with Style Latent Diffusion TransformerabstractKinship face synthesis is a challenging problem due to the scarcity and low quality of the available kinship data. Existing methods often struggle to generate descendants with both high diversity and fidelity while precisely controlling facial attributes such as age and gender. To address these issues, we propose the Style Latent Diffusion Transformer (StyleDiT), a novel framework that integrates the strengths of StyleGAN with the diffusion model to generate high-quality and diverse kinship faces. In this framework, the rich facial priors of StyleGAN enable fine-grained attribute control, while our conditional diffusion model is used to sample a StyleGAN latent aligned with the kinship relationship of conditioning images by utilizing the advantage of modeling complex kinship relationship distribution. StyleGAN then handles latent decoding for final face generation. Additionally, we introduce the Relational Trait Guidance (RTG) mechanism, enabling independent control of influencing conditions, such as each parent's facial image. RTG also enables a fine-grained adjustment between the diversity and fidelity in synthesized faces. Furthermore, we extend the application to an unexplored domain: predicting a partner's facial images using a child's image and one parent's image within the same framework. Extensive experiments demonstrate that our StyleDiT outperforms existing methods by striking an excellent balance between generating diverse and high-fidelity kinship faces. Pin-Yen Chiu, Dai-Jie Wu, Po-Hsun Chu, Chia-Hsuan Hsu, Hsiang-Chen Chiu, Chih-Yu Wang 0001, Jun-Cheng Chen |
FG | 7 |
| 2026 | Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA AdaptersabstractRecent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over free-form text prompts. However, these methods not only require intensive training time and GPU memory usage to learn the sliders or embeddings but also need to be retrained for different diffusion backbones, limiting their scalability and adaptability. To address these limitations, we introduce Text Slider, a lightweight, efficient and plug-and-play framework that identifies low-rank directions within a pre-trained text encoder, enabling continuous control of visual concepts while significantly reducing training time, GPU memory consumption, and the number of trainable parameters. Furthermore, Text Slider supports multi-concept composition and continuous control, enabling fine-grained and flexible manipulation in both image and video synthesis. We show that Text Slider enables smooth and continuous modulation of specific attributes while preserving the original spatial layout and structure of the input. Text Slider achieves significantly better efficiency: 5× faster training than Concept Slider and 47× faster than Attribute Control, while reducing GPU memory usage by nearly 2× and 4×, respectively. Project page: https://textslider.github.io Pin-Yen Chiu, I-Sheng Fang, Jun-Cheng Chen |
WACV | 3 |
| 2026 | CropAT: Leveraging Diffusion-Generated Target-Like Cropped Objects for Pseudo-Label Refinement in Domain-Adaptive Object Detection
Chen-Che Huang, Tzuhsuan Huang, Jun-Cheng Chen |
WACV | 3 |
| 2026 | M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion ModelsabstractText-to-image diffusion models may generate harmful or copyrighted content, motivating research on concept erasure. However, existing approaches primarily focus on erasing concepts from text prompts, overlooking other input modalities that are increasingly critical in real-world applications such as image editing and personalized generation. These modalities can become attack surfaces, where erased concepts re-emerge despite defenses. To bridge this gap, we introduce M-ErasureBench, a novel multimodal evaluation framework that systematically benchmarks concept erasure methods across three input modalities: text prompts, learned embeddings, and inverted latents. For the latter two, we evaluate both white-box and black-box access, yielding five evaluation scenarios. Our analysis shows that existing methods achieve strong erasure performance against text prompts but largely fail under learned embeddings and inverted latents, with Concept Reproduction Rate (CRR) exceeding 90% in the white-box setting. To address these vulnerabilities, we propose IRECE (Inference-time Robustness Enhancement for Concept Erasure), a plug-and-play module that localizes target concepts via cross-attention and perturbs the associated latents during denoising. Experiments demonstrate that IRECE consistently restores robustness, reducing CRR by up to 40% under the most challenging white-box latent inversion scenario, while preserving visual quality. To the best of our knowledge, M-ErasureBench provides the first comprehensive benchmark of concept erasure beyond text prompts. Together with IRECE, our benchmark offers practical safeguards for building more reliable protective generative models. Ju-Hsuan Weng, Jia-Wei Liao, Cheng-Fu Chou, Jun-Cheng Chen |
WACV | 4 |
| 2026 | KMOPS: Keypoint-Driven Method for Multi-Object Pose and Metric Size Estimation from Stereo ImagesabstractThe six-degree-of-freedom (6-DoF) pose and metric size estimation of multiple objects from RGB images alone remains a challenging task, particularly due to significant variations in object shape, appearance, and frequent occlusions in complex scenes. To address these challenges, we introduce KMOPS, a Keypoint-driven method tailored specifically for Multi-Object Pose and metric Size estimation from a single calibrated stereo image pair. Leveraging the stereo input, our approach first extracts the 2D keypoints of the enclosing bounding boxes of the objects across both views, and subsequently triangulates them to acquire metric 3D positions. Then, we obtain each object's rotation, translation, and dimensions by aligning the triangulated 3D keypoints to the canonical ones using a closed-form solution. Our formulation eliminates the need for predefined 3D search spaces or volumetric anchors, which are often required by other methods to constrain the vast 3D solution space. With extensive experiments on the challenging dataset Transparent Object Dataset (TOD) and StereOBJ-1M, we show that our method outperforms all competing methods with a simple and effective architecture. Ying-Kun Wu, Tzuhsuan Huang, I-Sheng Fang, Jun-Cheng Chen |
WACV | 5 |
| 2026 | GOT-JEPA: Generic Object Tracking With Model Adaptation and Occlusion Handling Using Joint-Embedding Predictive ArchitectureabstractThe human visual system tracks objects by integrating current observations with previously observed information, adapting to target and scene changes, and reasoning about occlusion at fine granularity. In contrast, recent generic object trackers are often optimized for training targets, which limits robustness and generalization in unseen scenarios, and their occlusion reasoning remains coarse, lacking detailed modeling of occlusion patterns. To address these limitations in generalization and occlusion perception, we propose GOT-JEPA, a model-predictive pretraining framework that extends JEPA from predicting image features to predicting tracking models. Given identical historical information, a teacher predictor generates pseudo-tracking models from a clean current frame, and a student predictor learns to predict the same pseudo-tracking models from a corrupted version of the current frame. This design provides stable pseudo supervision and explicitly trains the predictor to produce reliable tracking models under occlusions, distractors, and other adverse observations, improving generalization to dynamic environments. Building on GOT-JEPA, we further propose OccuSolver to enhance occlusion perception for object tracking. OccuSolver adapts a point-centric point tracker for object-aware visibility estimation and detailed occlusion-pattern capture. Conditioned on object priors iteratively generated by the tracker, OccuSolver incrementally refines visibility states, strengthens occlusion handling, and produces higher-quality reference labels that progressively improve subsequent model predictions. Extensive evaluations on seven benchmarks show that our method effectively enhances tracker generalization and robustness. The code will be available at https://github.com/chenshihfang/GOT. Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion ModelsabstractDiffusion Models have emerged as powerful generative models for high-quality image synthesis, with many subsequent image editing techniques based on them. However, the ease of text-based image editing introduces significant risks, such as malicious editing for scams or intellectual property infringement. Previous works have attempted to safeguard images from diffusion-based editing by adding imperceptible perturbations. These methods are costly and specifically target prevalent Latent Diffusion Models (LDMs), while Pixel-domain Diffusion Models (PDMs) remain largely unexplored and robust against such attacks. Our work addresses this gap by proposing a novel attack framework, AtkPDM. AtkPDM is mainly composed of a feature representation attacking loss that exploits vulnerabilities in denoising UNets and a latent optimization strategy to enhance the naturalness of adversarial images. Extensive experiments demonstrate the effectiveness of our approach in attacking dominant PDM-based editing methods (e.g., SDEdit) while maintaining reasonable fidelity and robustness against common defense methods. Additionally, our framework is extensible to LDMs, achieving comparable performance to existing approaches. Chun-Yen Shih, Li-Xuan Peng, Jia-Wei Liao, Ernie Chu, Cheng-Fu Chou, Jun-Cheng Chen |
AAAI | 6 |
| 2025 | CSS: Clustering-guided SliderSpace for Under-represented Visual Concept Exploration of Diffusion ModelsabstractRecent advances in diffusion models have enabled users to generate high-quality images from simple text prompts. However, efficiently steering these models to produce images with diverse styles and appearances remains a significant challenge. SliderSpace was recently introduced as a promising solution, leveraging principal component analysis (PCA) on generated image features to find meaningful semantic directions and training low-rank adaptation-based (LoRA) sliders of these directions for controllable image generation. However, we observe that existing pre-trained diffusion models, such as Stable Diffusion, exhibit substantial biases in their generation results which limit the visual exploration capability of SliderSpace for rare attributes or concepts. To address this issue, we propose clustering-guided SliderSpace, which introduces an initial clustering step to explicitly identify subconcepts from generated images followed by applying PCA and the remaining SliderSpace procedure to each cluster. Experimental results demonstrate that this clustering step is useful and essential for the discovery of rare and under-represented concepts. In addition, the proposed method complements the original SliderSpace, when used together, it enables more comprehensive visual concept exploration. Jun-Cheng Chen |
AVSS | 1 |
| 2025 | Invisible Backdoor Triggers in Image Editing Model via Deep Watermarking
Tzuhsuan Huang, Pin-Yen Chiu, Jun-Cheng Chen |
AVSS | 4 |
| 2025 | Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation ModelabstractThe current deep generative models have enabled the creation of synthetic facial images with remarkable photorealism, raising significant societal concerns over their potential misuse. Despite rapid advancements in the field of deepfake detection, developing an efficient and effective approach for the generalized deepfake detection of unseen forgery samples remains challenging. To address this challenge, we leverage the rich semantic priors of foundation models and propose a novel side-network-based decoder that extracts spatial and temporal cues using the CLIP image encoder for generalized video-based Deepfake detection. Additionally, we introduce Facial Component Guidance (FCG) to enhance spatial learning generalizability by encouraging the model to focus on key facial regions. By leveraging the generic features of a vision-language foundation model, our approach demonstrates promising generalizability on challenging Deepfake datasets while also exhibiting superiority in training data efficiency, parameter efficiency, and model robustness. The source code is available at: https://github.com/aiiu-lab/DFD-FCG Yue-Hua Han, Tai-Ming Huang, Kai-Lung Hua, Jun-Cheng Chen |
CVPR | 4 |
| 2025 | Consistent View Synthesis with Bidirectional Epipolar Attention and ReconstructionabstractNovel view synthesis from a single image aims to generate novel scene views given a reference image and a sequence of camera poses. Its primary difficulty lies in effectively leveraging a generative model to achieve high-quality image generation while simultaneously ensuring consistency and faithfulness across synthesized views. In this paper, we propose a novel approach to address the consistency and faithfulness issues in view synthesis. Specifically, we develop a new attention layer, termed bidirectional epipolar attention, which utilizes a pair of complementary epipolar lines to guide the associations between features from different viewpoints. Each bidirectional epipolar layer calculates forward and backward epipolar lines, enabling geometrically constrained attention that improves cross-view consistency. To ensure faithful synthesis, we introduce an epipolar-aware reconstruction module that prevents creating novel content in regions where the newly generated image overlaps with existing ones. Extensive experimental results demonstrate that our method outperforms previous approaches to novel view synthesis, achieving superior performance in both image quality and consistency. The source code is available at https://github.com/fallantbell/Bidirectional-Epipolar-Synthesis. I-Chung Chiu, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
ICIP | 2 |
| 2025 | Enhancing Image Deraining Through VLM-Based Data Refinement and ClassificationabstractImage deraining has gained significant attention in recent years due to its essential role in applications like autonomous driving and surveillance systems. Despite their good performance in rain removal, most image-deraining models are trained on synthetic datasets, leading to a performance gap when applied to real-world scenarios due to differences between synthetic and actual rain patterns. Although some real-world datasets exist, they may contain unsuitable images for training image-deraining models, such as those without clear rain or containing haze, which could lead to suboptimal model performance. Therefore, we propose Automated Understanding and Refinement Agents (AURA) to refine deraining datasets by leveraging large vision-language models to filter out inappropriate training data and categorize images based on the rain density (e.g., light, moderate, and heavy rain). This automated refinement and annotation process enhances dataset quality, resulting in improved deraining model performance in real-world scenarios. Experimental results demonstrate that training deraining models on datasets refined and annotated by AURA can significantly enhance their deraining performance. Shih-Jui Liang, Zihao Chen 0003, Jun-Cheng Chen, Yan-Tsung Peng |
ICIP | 3 |
| 2025 | Diffusion to Confusion: Naturalistic Adversarial Patch Generation Based on Diffusion Model for Object DetectorabstractMany physical adversarial patch generation methods are widely proposed to protect personal privacy from malicious monitoring using object detectors. However, they usually fail to generate satisfactory patch images in terms of both stealthiness and attack performance without making huge efforts on careful hyperparameter tuning. To address this issue, we propose a novel naturalistic diffusion model-based (DM) adversarial patch generation method. Through sampling the optimal image from the pretrained DM model upon natural images, it allows us to stably craft high-quality physical adversarial patches without suffering serious mode collapse problems as other deep generative models. Moreover, the generated patches are not only visually pleasing but also robust against the state-of-the-art adversarial patch removal algorithm due to different image statistics from the traditional ones. In addition, to resolve the huge memory requirement of the diffusion model during backpropagation, we also utilize adjoint method for patch generation. With extensive experiments, the results demonstrate the effectiveness of the proposed approach to generate better-quality and more stealthy adversarial patches than other approaches against the adversarial patch removal algorithm while achieving comparable attack performance than other state-of-the-art patch generation methods. Shuo-Yen Lin, Ernie Chu, Po-Hung Yeh, Jun-Cheng Chen, Jia-Ching Wang |
ICIP | 4 |
| 2025 | EXDF: Explainable Deepfake Detection with Vision-Language ModelabstractAlthough many deepfake detection methods have been proposed to fight against severe misuse of generative AI, none provide detailed human-interpretable explanations beyond simple real/fake responses. This limitation makes it challenging for humans to assess the accuracy of detection results, especially when the models encounter unseen deepfakes. To address this issue, we propose a novel deepfake detector based on a large Vision-Language Model (VLM), capable of explaining manipulated facial regions. We frame the deepfake detection task as Visual Question Answering (VQA) and perform visual instruction tuning to train the model on our collected Explainable Deepfake Face (ExDF) dataset. The dataset consists of fake images from diverse generative adversarial networks (GANs) and diffusion models (DMs), with explanations produced by GPT-4o guided by the corresponding ground-truth masks of the manipulated regions. Moreover, a facial mask encoder is introduced to guide the model to focus on key facial features, thereby improving the detection and explanation performances. Extensive experiments demonstrate that training the proposed model on the full ExDF dataset not only enhances detection accuracy compared to baseline methods but also provides detailed, human-interpretable explanations. To our knowledge, ExDF is the first explainable deepfake face dataset covering both GANs and DMs with comprehensive descriptions of altered facial regions. Our code and dataset are available at https://github.com/aiiu-lab/ExDF. Shu-Tzu Lo, Tai-Ming Huang, Yue-Hua Han, Kai-Lung Hua, Jun-Cheng Chen |
ICIP | 5 |
| 2025 | Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and ComprehensionabstractReferring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. Moreover, REG methods frequently struggle to bridge the visual and textual domains due to the limited capacity, leading to low-quality and restricted diversity in expression generation. To address these issues, we propose a novel VIsion-guided Expression Diffusion Model (VIE-DM) for the REG task, where diverse synonymous expressions adhering to both image and text contexts of the target object are generated to augment REC datasets. VIE-DM consists of a vision-text condition (VTC) module and a transformer decoder. Our VTC and token selection design effectively addresses the feature discrepancy problem prevalent in existing REG methods. This enables us to generate high-quality, diverse synonymous expressions that can serve as augmented data for REC model learning. Extensive experiments on five datasets demonstrate the high quality and large diversity of our generated expressions. Furthermore, the augmented image-expression pairs consistently enhance the performance of existing REC models, achieving state-of-the-art results. Jingcheng Ke, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
ICLR | 2 |
| 2025 | Training-Free Diffusion Model Alignment with Sampling DemonsabstractAligning diffusion models with user preferences has been a key challenge.
Existing methods for aligning diffusion models either require retraining or are limited to differentiable reward functions.
To address these limitations, we propose a stochastic optimization approach, dubbed *Demon*, to guide the denoising process at inference time without backpropagation through reward functions or model retraining.
Our approach works by controlling noise distribution in denoising steps to concentrate density on regions corresponding to high rewards through stochastic optimization.
We provide comprehensive theoretical and empirical evidence to support and validate our approach, including experiments that use non-differentiable sources of rewards such as Visual-Language Model (VLM) APIs and human judgements.
To the best of our knowledge, the proposed approach is the first inference-time, backpropagation-free preference alignment method for diffusion models.
Our method can be easily integrated with existing diffusion models without further training.
Our experiments show that the proposed approach significantly improves the average aesthetics scores for text-to-image generation. Implementation is available at [this URL](https://github.com/aiiu-lab/DemonSampling). Po-Hung Yeh, Kuang-Huei Lee, Jun-Cheng Chen |
ICLR | 3 |
| 2025 | DRAG: Data Reconstruction Attack using Guided DiffusionabstractWith the rise of large foundation models, split inference (SI) has emerged as a popular computational paradigm for deploying models across lightweight edge devices and cloud servers, addressing data privacy and computational cost concerns. However, most existing data reconstruction attacks have focused on smaller CNN classification models, leaving the privacy risks of foundation models in SI settings largely unexplored. To address this gap, we propose a novel data reconstruction attack based on guided diffusion, which leverages the rich prior knowledge embedded in a latent diffusion model (LDM) pre-trained on a large-scale dataset. Our method performs iterative reconstruction on the LDM’s learned image prior, effectively generating high-fidelity images resembling the original data from their intermediate representations (IR). Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, both qualitatively and quantitatively, in reconstructing data from deep-layer IRs of the vision foundation model. The results highlight the urgent need for more robust privacy protection mechanisms for large models in SI scenarios. Wa-Kin Lei, Jun-Cheng Chen, Shang-Tse Chen |
ICML | 2 |
| 2025 | Restore Anything Anywhere: Targeted Image Restoration with Object Segmentation and Text GuidanceabstractImage restoration techniques are widely used in various fields, such as autonomous driving, medical imaging, and satellite imagery. These techniques typically aim to restore an entire degraded image to its original state. However, in many cases, users may wish to focus on restoring specific areas to achieve desired effects in the image, tailored to their preferences. In this work, we propose Restore Anything anyWhere (RAW), a framework that enables users to restore specific types of degradation on any selected object in an image by specifying point or text prompts. Our framework first employs the Multimodal Segmentation module to generate mask priors for the target objects. Then, the text-guided restoration model performs targeted restoration on the areas defined by the object mask priors. To improve segmentation performance, we propose a Text-To-Segmentation Refinement method that combines existing techniques with the capabilities of Segment Anything, enhancing text-to-segmentation accuracy. Additionally, we introduce a streamlined text-guided restoration method to provide better control over degradation restoration, delivering highly effective results. RAW provides excellent object-level control in restoration and significantly improves the Laion Aesthetics score. Yen-Ku Yeh, Chun-Hao Yang, Kun-Tai Wu, Yan-Tsung Peng, Chun-Rong Huang, Jun-Cheng Chen |
MMSP | 6 |
| 2025 | FIPER: Factorized Features for Robust Image Super-Resolution and CompressionabstractIn this work, we propose using a unified representation, termed **Factorized Features**, for low-level vision tasks, where we test on **Single Image Super-Resolution (SISR)** and **Image Compression**. Motivated by the shared principles between these tasks, they require recovering and preserving fine image details, whether by enhancing resolution for SISR or reconstructing compressed data for Image Compression. Unlike previous methods that mainly focus on network architecture, our proposed approach utilizes a basis-coefficient decomposition as well as an explicit formulation of frequencies to capture structural components and multi-scale visual features in images, which addresses the core challenges of both tasks.
We replace the representation of prior models from simple feature maps with Factorized Features to validate the potential for broad generalizability.
In addition, we further optimize the compression pipeline by leveraging the mergeable-basis property of our Factorized Features, which consolidates shared structures on multi-frame compression.
Extensive experiments show that our unified representation delivers state-of-the-art performance, achieving an average relative improvement of 204.4\% in PSNR over the baseline in Super-Resolution (SR) and 9.35\% BD-rate reduction in Image Compression compared to the previous SOTA. Yang-Che Sun, Cheng Yu Yeo, Ernie Chu, Jun-Cheng Chen, Yu-Lun Liu 0001 |
NeurIPS | 4 |
| 2025 | DiffQRCoder: Diffusion-Based Aesthetic QR Code Generation with Scanning Robustness Guided Iterative RefinementabstractWith the success of Diffusion Models for image generation, the technologies also have revolutionized the aesthetic Quick Response (QR) code generation. Despite significant improvements in visual attractiveness for the beautified codes, their scannabilities are usually sacrificed and thus hinder their practical uses in real-world scenarios. To address this issue, we propose a novel training-free Diffusion-based QR Code generator (DiffQRCoder) to effectively craft both scannable and visually pleasing QR codes. The proposed approach introduces Scanning-Robust Perceptual Guidance (SRPG), a new diffusion guidance for Diffusion Models to guarantee the generated aesthetic codes to obey the ground-truth QR codes while maintaining their attractiveness during the denoising process. Additionally, we present another post-processing technique, Scanning Robust Manifold Projected Gradient Descent (SR-MPGD), to further enhance their scanning robustness through iterative latent space optimization. With extensive experiments, the results demonstrate that our approach not only outperforms other compared methods in Scanning Success Rate (SSR) with better or comparable CLIP aesthetic score (CLIP-aes.) but also significantly improves the SSR of the ControlNet-only approach from 60% to 99%. The subjective evaluation indicates that our approach achieves promising visual attractiveness to users as well. Finally, even with different scanning angles and the most rigorous error tolerance settings, our approach robustly achieves over 95% SSR, demonstrating its capability for real-world applications. Our project page is available at https://jwliao1209.github.io/DiffQRCoder. Jia-Wei Liao, Winston Wang, Tzu-Sian Wang, Li-Xuan Peng, Ju-Hsuan Weng, Cheng-Fu Chou, Jun-Cheng Chen |
WACV | 7 |
| 2025 | Improving Visual Object Tracking Through Visual PromptingabstractLearning a discriminative model to distinguish a target from its surrounding distractors is essential to generic visual object tracking. Dynamic target representation adaptation against distractors is challenging due to the limited discriminative capabilities of prevailing trackers. We present a new visual Prompting mechanism for generic Visual Object Tracking (PiVOT) to address this issue. PiVOT proposes a prompt generation network with the pre-trained foundation model CLIP to automatically generate and refine visual prompts, enabling the transfer of foundation model knowledge for tracking. While CLIP offers broad category-level knowledge, the tracker, trained on instance-specific data, excels at recognizing unique object instances. Thus, PiVOT first compiles a visual prompt highlighting potential target locations. To transfer the knowledge of CLIP to the tracker, PiVOT leverages CLIP to refine the visual prompt based on the similarities between candidate objects and the reference templates across potential targets. Once the visual prompt is refined, it can better highlight potential target locations, thereby reducing irrelevant prompt information. With the proposed prompting mechanism, the tracker can generate improved instance-aware feature maps through the guidance of the visual prompt, thus effectively reducing distractors. The proposed method does not involve CLIP during training, thereby keeping the same training complexity and preserving the generalization capability of the pretrained foundation model. Extensive experiments across multiple benchmarks indicate that PiVOT, using the proposed prompting method can suppress distracting objects and enhance the tracker. Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
IEEE Trans. Multim. | 2 |
| 2025 | Make Graph-Based Referring Expression Comprehension Great Again Through Expression-Guided Dynamic Gating and RegressionabstractOne common belief is that with complex models and pre-training on large-scale datasets, transformer-based methods for referring expression comprehension (REC) perform much better than existing graph-based methods. We observe that since most graph-based methods adopt an off-the-shelf detector to locate candidate objects (i.e., regions detected by the object detector), they face two challenges that result in subpar performance: (1) the presence of significant noise caused by numerous irrelevant objects during reasoning, and (2) inaccurate localization outcomes attributed to the provided detector. To address these issues, we introduce a plug-and-adapt module guided by sub-expressions, called dynamic gate constraint (DGC), which can adaptively disable irrelevant proposals and their connections in graphs during reasoning. We further introduce an expression-guided regression strategy (EGR) to refine location prediction. Extensive experimental results on the RefCOCO, RefCOCO+, RefCOCOg, Flickr30 K, RefClef, and Ref-reasoning datasets demonstrate the effectiveness of the DGC module and the EGR strategy in consistently boosting the performances of various graph-based REC methods. Without any pretaining, the proposed graph-based method achieves better performance than the state-of-the-art (SOTA) transformer-based methods. Jingcheng Ke, Dele Wang, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
IEEE Trans. Multim. | 3 |
| 2024 | MeDM: Mediating Image Diffusion Models for Video-to-Video Translation with Temporal Correspondence GuidanceabstractThis study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position information, such as a normal G-buffer, or perform text-guided editing on videos captured in real-world scenarios. We employ explicit optical flows to construct a practical coding that enforces physical constraints on generated frames and mediates independent frame-wise scores. By leveraging this coding, maintaining temporal consistency in the generated videos can be framed as an optimization problem with a closed-form solution. To ensure compatibility with Stable Diffusion, we also suggest a workaround for modifying observation-space scores in latent Diffusion Models. Notably, MeDM does not require fine-tuning or test-time optimization of the Diffusion Models. Through extensive qualitative, quantitative, and subjective experiments on various benchmarks, the study demonstrates the effectiveness and superiority of the proposed approach. Our project page can be found at https://medm2023.github.io Ernie Chu, Tzuhsuan Huang, Shuo-Yen Lin, Jun-Cheng Chen |
AAAI | 4 |
| 2024 | A Recipe for CAC: Mosaic-Based Generalized Loss for Improved Class-Agnostic Counting
Tsung-Han Chou, Walon Wei-Chen Chiu, Jun-Cheng Chen |
ACCV (6) | 4 |
| 2024 | Blenda: Domain Adaptive Object Detection Through Diffusion-Based BlendingabstractUnsupervised domain adaptation (UDA) aims to transfer a model learned using labeled data from the source domain to unlabeled data in the target domain. To address the large domain gap issue between the source and target domains, we propose a novel regularization method for domain adaptive object detection, BlenDA, by generating the pseudo samples of the intermediate domains and their corresponding soft domain labels for adaptation training. The intermediate samples are generated by dynamically blending the source images with their corresponding translated images using an off-the-shelf pre-trained text-to-image diffusion model which takes the text label of the target domain as input and has demonstrated superior image-to-image translation quality. Based on experimental results from two adaptation benchmarks, our proposed approach can significantly enhance the performance of the state-of-the-art domain adaptive object detector, Adversarial Query Transformer (AQT). Particularly, in the Cityscapes to Foggy Cityscapes adaptation, we achieve an impressive 53.4% mAP on the Foggy Cityscapes dataset, surpassing the previous state-of-the-art by 1.5%. It is worth noting that our proposed method is also applicable to various paradigms of domain adaptive object detection. The code is available at https://github.com/aiiu-lab/BlenDA Tzuhsuan Huang, Chen-Che Huang, Chung-Hao Ku, Jun-Cheng Chen |
ICASSP | 4 |
| 2024 | Generalized Image-Based Deepfake Detection Through Foundation Model Adaptation
Tai-Ming Huang, Yue-Hua Han, Ernie Chu, Shu-Tzu Lo, Kai-Lung Hua, Jun-Cheng Chen |
ICPR (21) | 6 |
| 2024 | Adversarially Robust Deepfake Detection via Adversarial Feature Similarity Learning
Sarwar Khan, Jun-Cheng Chen, Wen-Hung Liao, Chu-Song Chen |
MMM (3) | 2 |
| 2024 | Camera Settings as Tokens: Modeling Photography on Latent Diffusion Models
I-Sheng Fang, Yue-Hua Han, Jun-Cheng Chen |
SIGGRAPH Asia | 3 |
| 2024 | Towards Validating Face Editing Ability in Generative ModelsabstractFace editing has recently blossomed into a highly active and significant domain, impacting numerous applications from entertainment to security. Despite its rapid growth, the field still faces challenges in establishing a universally accepted and robust evaluation mechanism that can comprehensively assess the performances of face editing techniques and their underlying generative models. Our paper introduces a well-defined evaluation protocol that seamlessly combines systematic experimental methodologies with thorough subjective evaluations. This collaborative approach ensures a more in-depth and unbiased examination of face editing techniques. Based on our extensive studies, we observe that traditional metrics, notably the Fréchet Inception Distance (FID), serve well in measuring perceptual attributes of edited images. However, they might fall short in covering all aspects of a face editing method’s capabilities. To bridge this gap, we have incorporated additional metrics that assess disentanglement and editing effectiveness, leading to the creation of a holistic assessment framework that promises a more comprehensive evaluation upon face editing capability of different deep generative models. Dai-Jie Wu, Pin-Yen Chiu, Chih-Yu Wang 0001, Jun-Cheng Chen |
VCIP | 4 |
| 2024 | CLIPREC: Graph-Based Domain Adaptive Network for Zero-Shot Referring Expression ComprehensionabstractReferring expression comprehension (REC) is a cross-modal matching task that aims to localize the target object in an image specified by a text description. Most existing approaches for this task focus on identifying only objects whose categories are covered by training data. This restricts their generalization to unseen categories and practical usage. To address this issue, we propose a domain adaptive network called CLIPREC for zero-shot REC, which integrates the Contrastive Language-Image Pretraining (CLIP) model for graph-based REC. The proposed CLIPREC is composed of a graph collaborative attention module with two directed graphs: one for objects in an image and the other for their corresponding categorical labels. To carry out zero-shot REC, we leverage the strong common image-text feature space from the CLIP model to correlate the two graphs. Furthermore, a multilayer perceptron is introduced to enable feature alignment so that the CLIP model is adapted to the expression representation from the language parser, resulting in effective reasoning from expressions involving both seen and unseen object categories. Extensive experimental and ablation results on several widely-adopted benchmarks show that the proposed approach performs favorably against state-of-the-art approaches for zero-shot REC. Jingcheng Ke, Jia Wang 0020, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
IEEE Trans. Multim. | 3 |
| 2023 | Improved Photometric Stereo through Efficient and Differentiable Shadow Estimation
Po-Hung Yeh, Pei Yuan Wu, Jun-Cheng Chen |
BMVC | 3 |
| 2023 | Hearing and Seeing Abnormality: Self-Supervised Audio-Visual Mutual Learning for Deepfake DetectionabstractThe recent development of deepfakes has resulted in serious threats to society, such as spreading misinformation, defamation, etc. Although recent deepfake detection methods are capable of achieving satisfactory results for seen forgeries, the performance drops significantly for unseen ones. With proper supervised pretraining on auxiliary tasks as prior, the situation can be improved, but the requirement to collect a large number of additional annotations for these tasks may restrict the further development of a generalized deep-fake detector. To address this issue, we propose an Audio-Visual Temporal Synchronization for Deepfake Detection framework for detecting deepfakes that maintains reasonable detection capabilities for unseen ones. The primary objective of our framework is to determine whether there has been a forgery by evaluating the consistency between the sound and the faces in a video clip, together with the relationship between the two features. First, the spatiotemporal feature extraction network is pretrained in a self-supervised manner by exploiting the audio-visual temporal synchronization task to build up a rich representation based on the temporal synchronization relationship between the audio and its corresponding video. For pretraining, we use only real data and carefully selected negative samples with contrastive loss to train the model. A temporal classifier network is used to determine whether or not the video has been manipulated using the representations obtained from the pretrained feature extraction networks. To prevent the model from overfitting to certain manipulation-specific artifacts, we froze the feature extraction networks and only trained the final classifier network on forged data. Extensive experiments on unseen forgery categories and unseen datasets have shown the effectiveness of our method to achieve state-of-the-art results. Chang-Sung Sung, Jun-Cheng Chen, Chu-Song Chen |
ICASSP | 2 |
| 2023 | Task-Adaptive Feature Matching Loss for Image DeblurringabstractImage deblurring is a highly challenging and ill-posed image restoration problem. Contemporary deep learning-based approaches usually tackle this problem by exploiting the encoder-decoder-based models trained by the commonly used mean squared error loss with the feature matching loss as a regularization to obtain perceptual consistent restored results as the ground truths. We argue that since the general backbone models for computing feature matching loss are usually not trained on the image deblurring task, the loss lacks specific knowledge of blur and usually leads to suboptimal performance. To address this issue, we propose a task-adaptive feature matching loss for image deblurring where we synthesize blurred images in different blur extents and employ triplet loss to finetune the backbone model for learning specific blur priors. Then, we leverage the finetuned backbone to compute feature matching loss which can greatly enhance the existing image deblurring models for better perceptual results. With extensive experiments on the GoPro and RealBlur datasets, both qualitative and quantitative results show that the SOTA deblurring models trained with the proposed loss can effectively obtain better and sharper restored images in terms of various perceptual image quality metrics than the original models while maintaining comparable PSNR and SSIM performances. Chiao-Chang Chang, Bo-Cheng Yang, Yi-Ting Liu, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
ICIP | 4 |
| 2023 | Multi-Task Self-Blended Images for Face Forgery DetectionabstractDeepfake detection has attracted extensive attention due to widespread forged images on social media. Recently, self-supervised learning (SSL) based Deepfake detection approaches have outperformed supervised methods in terms of model generalization. However, we notice that most SSL-based methods do not take the manipulation strength levels of synthesized forgery samples into consideration according to different synthesis parameters and result in suboptimal detection performances. To address this issue, we introduce several auxiliary losses to the state-of-the-art SSL-based method based on different synthesis sub-tasks during data generation by inferring their synthesis parameters where the ground-truth labels are obtained from the synthesis pipeline for free. With comprehensive evaluations on various benchmarks, our approach has achieved noticeable performance improvement. Specifically, for the cross-dataset evaluation, the proposed approach outperforms the state-of-the-art method in terms of AUC on various datasets with improvements of 3.4%, 1.47%, 1.56%, and 1.3% on the CDF, DFDC, DFDCP, and FFIW datasets and achieves competitive performance on the DFD dataset. This further demonstrates the effectiveness of the proposed approach in its generalization ability. Yue-Hua Han, Ernie Chu, Jun-Cheng Chen, Kai-Lung Hua |
MMAsia | 4 |
| 2023 | DEFAEK: Domain Effective Fast Adaptive Network for Face Anti-Spoofing
Jiun-Da Lin, Yue-Hua Han, Julianne Tan, Jun-Cheng Chen, Muhammad Tanveer 0001, Kai-Lung Hua |
Neural Networks | 5 |
| 2022 | KinStyle: A Strong Baseline Photorealistic Kinship Face Synthesis with an Optimized StyleGAN Encoder
Li-Chen Cheng, Shu-Chuan Hsu, Pin-Hua Lee, Hsiu-Chieh Lee, Che-Hsien Lin, Jun-Cheng Chen, Chih-Yu Wang 0001 |
ACCV (4) | 6 |
| 2022 | Continual Learning for Visual Search with Backward Consistent Feature EmbeddingabstractIn visual search, the gallery set could be incrementally growing and added to the database in practice. However, existing methods rely on the model trained on the entire dataset, ignoring the continual updating of the model. Besides, as the model updates, the new model must re-extract features for the entire gallery set to maintain compatible feature space, imposing a high computational cost for a large gallery set. To address the issues of long-term visual search, we introduce a continual learning (CL) approach that can handle the incrementally growing gallery set with backward embedding consistency. We enforce the losses of inter-session data coherence, neighbor-session model coherence, and intra-session discrimination to conduct a continual learner. In addition to the disjoint setup, our CL solution also tackles the situation of increasingly adding new classes for the blurry boundary without assuming all categories known in the beginning and during model update. To our knowledge, this is the first CL method both tackling the issue of backward-consistent feature embedding and allowing novel classes to occur in the new sessions. Extensive experiments on various benchmarks show the efficacy of our approach under a wide range of setups11Code: https://github.com/ivclab/CVS. Timmy S. T. Wan, Jun-Cheng Chen, Tzer-Yi Wu, Chu-Song Chen |
CVPR | 2 |
| 2022 | CLIPCAM: A Simple Baseline For Zero-Shot Text-Guided Object And Action LocalizationabstractThe key for the contemporary deep learning-based object and action localization algorithms to work is the large-scale annotated data. However, in real-world scenarios, since there are infinite amounts of unlabeled data beyond the categories of publicly available datasets, it is not only time- and manpower-consuming to annotate all the data but also requires a lot of computational resources to train the detectors. To address these issues, we show a simple and reliable baseline that can be easily obtained and work directly for the zero-shot text-guided object and action localization tasks without introducing additional training costs by using Grad-CAM, the widely used class visual saliency map generator, with the help of the recently released Contrastive Language-Image Pre-Training (CLIP) model by OpenAI, which is trained contrastively using the dataset of 400 million image-sentence pairs with rich cross-modal information between text semantics and image appearances. With extensive experiments on the Open Images and HICO-DET datasets, the results demonstrate the effectiveness of the proposed approach for the text-guided unseen object and action localization tasks for images. Hsuan-An Hsia, Che-Hsien Lin, Bo-Han Kung, Jhao-Ting Chen, Daniel Stanley Tan, Jun-Cheng Chen, Kai-Lung Hua |
ICASSP | 6 |
| 2022 | Residual Graph Attention Network and Expression-Respect Data Augmentation Aided Visual GroundingabstractVisual grounding aims to localize a target object in an image based on a given text description. Due to the innate complexity of language, it is still a challenging problem to perform reasoning of complex expressions and to infer the underlying relationship between the expression and the object in an image. To address these issues, we propose a residual graph attention network for visual grounding. The proposed approach first builds an expression-guided relation graph and then performs multi-step reasoning followed by matching the target object. It allows performing better visual grounding with complex expressions by using deeper layers than other graph network approaches. Moreover, to increase the diversity of training data, we perform an expression-respect data augmentation based on copy-paste operations to pairs of source and target images. The proposed approach achieves better performance with extensive experiments than other state-of-the-art graph network-based approaches and demonstrates its effectiveness. Jia Wang 0020, Hung-Yi Wu, Jun-Cheng Chen, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 3 |
| 2022 | Enhancing the Robustness of Deep Learning Based Fingerprinting to Improve Deepfake AttributionabstractArtificial Fingerprinting (AF or the so-called digital watermarking) is a technique that can be used to conduct Deepfake attribution by ensuring media authenticity. However, AF does not prioritize its robustness to certain kinds of distortions, making the embedded watermarks vulnerable to some standard image processing operations. Insufficient robustness reduces the practicality of digital watermarking techniques. To address this issue, we propose an enhanced distortion agnostic artificial fingerprinting (EDA-AF) framework which introduces a novel noise layer consisting of an attack booster followed by a convolutional network-based attacker. The attacker simulates various distortions by exploiting adversarial learning with AF for distortion agnostic robustness. Meanwhile, due to the modeling limitation of the convolutional network, we also employ the attack booster to apply a set of differentiable image distortions which cannot be well simulated by the attacker. Extensive experimental results show that the proposed approach improves the quality of the extracted fingerprints. EDA-AF can improve the bitwise accuracy by up to 36%, which takes another step forward on the road of Deepfake attribution. Chieh-Yin Liao, Chen-Hsiu Huang, Jun-Cheng Chen, Ja-Ling Wu |
MMAsia | 3 |
| 2022 | A Comparative Study of Cross-Model Universal Adversarial Perturbation for Face ForgeryabstractAlthough the rapid development of deep generative models (DGM) enables diverse applications of content creation, increasing illegal uses of the technologies also severely threaten the privacy and security of personal information, especially for faces. Several previous works have been proposed to leverage adversarial attacks to fight against these malicious manipulations by adding an imperceptible perturbation to each input image to disrupt the output. In addition, to improve its scalability, a sequential cross-model universal perturbation attack has been proposed to learn a common adversarial perturbation to defend the images from the manipulation of multiple DGMs. However, we find that the order of DGMs for the adversarial perturbation generation does matter and influence the final defense performance. To address this issue, we propose to generate the universal perturbation through joint optimization of multiple DGMs. From the extensive experimental results, we find that the universal perturbation generated by the proposed method can successfully disrupt the output faces of multiple DGMs at the same time and achieves higher attack success rates than the previous state-of-the-art method based on the sequential generation, even under the situations where the model robustness of DGMs are enhanced by random perturbations. Shuo-Yen Lin, Jun-Cheng Chen, Jia-Ching Wang |
VCIP | 2 |
| 2022 | Pose and Joint-Aware Action RecognitionabstractRecent progress on action recognition has mainly focused on RGB and optical flow features. In this paper, we approach the problem of joint-based action recognition. Unlike other modalities, constellation of joints and their motion generate models with succinct human motion information for activity recognition. We present a new model for joint-based action recognition, which first extracts motion features from each joint separately through a shared motion encoder before performing collective reasoning. Our joint selector module re-weights the joint information to select the most discriminative joints for the task. We also propose a novel joint-contrastive loss that pulls together groups of joint features which convey the same action. We strengthen the joint-based representations by using a geometry-aware data augmentation technique which jitters pose heatmaps while retaining the dynamics of the action. We show large improvements over the current state-of-the-art joint-based approaches on JHMDB, HMDB, Charades, AVA action recognition datasets. A late fusion with RGB and Flow-based approaches yields additional improvements. Our model also outperforms the existing baseline on Mimetics, a dataset with out-of-context actions. Anshul Shah 0001, Shlok Kumar Mishra, Ankan Bansal, Jun-Cheng Chen, Rama Chellappa, Abhinav Shrivastava |
WACV | 4 |
| 2022 | A cooperative image object recognition framework and task offloading optimization in edge computing
Chu-Fu Wang, Yih-Kai Lin, Jun-Cheng Chen |
J. Netw. Comput. Appl. | 3 |
| 2022 | Lightweight Face Anti-Spoofing Network for Telehealth ApplicationsabstractOnline healthcare applications have grown more popular over the years. For instance, telehealth is an online healthcare application that allows patients and doctors to schedule consultations, prescribe medication, share medical documents, and monitor health conditions conveniently. Apart from this, telehealth can also be used to store a patient's personal and medical information. With its rise in usage due to COVID-19, given the amount of sensitive data it stores, security measures are necessary. A simple way of making these applications more secure is through user authentication. One of the most common and often used authentications is face recognition. It is convenient and easy to use. However, face recognition systems are not foolproof. They are prone to malicious attacks like printed photos, paper cutouts, replayed videos, and 3D masks. The goal of face anti-spoofing is to differentiate real users (live) from attackers (spoof). Although effective in terms of performance, existing methods use a significant amount of parameters, making them resource-heavy and unsuitable for handheld devices. Apart from this, they fail to generalize well to new environments like changes in lighting or background. This paper proposes a lightweight face anti-spoofing framework that does not compromise on performance. Our proposed method achieves good performance with the help of an ArcFace Classifier (AC). The AC encourages differentiation between spoof and live samples by making clear boundaries between them. With clear boundaries, classification becomes more accurate. We further demonstrate our model's capabilities by comparing the number of parameters, FLOPS, and performance with other state-of-the-art methods. Jiun-Da Lin, Hung-Hsiang Lin, Jilyan Bianca Dy, Jun-Cheng Chen, Muhammad Tanveer 0001, Muhammad Imran Razzak, Kai-Lung Hua |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Class-Aware Robust Adversarial Training for Object DetectionabstractObject detection is an important computer vision task with plenty of real-world applications; therefore, how to enhance its robustness against adversarial attacks has emerged as a crucial issue. However, most of the previous defense methods focused on the classification task and had few analysis in the context of the object detection task. In this work, to address the issue, we present a novel class-aware robust adversarial training paradigm for the object detection task. For a given image, the proposed approach generates an universal adversarial perturbation to simultaneously attack all the occurred objects in the image through jointly maximizing the respective loss for each object. Meanwhile, instead of normalizing the total loss with the number of objects, the proposed approach decomposes the total loss into class-wise losses and normalizes each class loss using the number of objects for the class. The adversarial training based on the class weighted loss can not only balances the influence of each class but also effectively and evenly improves the adversarial robustness of trained models for all the object classes as compared with the previous defense methods. Furthermore, with the recent development of fast adversarial training, we provide a fast version of the proposed algorithm which can be trained faster than the traditional adversarial training while keeping comparable performance. With extensive experiments on the challenging PASCAL-VOC and MS-COCO datasets, the evaluation results demonstrate that the proposed defense methods can effectively enhance the robustness of the object detection models. Pin-Chun Chen, Bo-Han Kung, Jun-Cheng Chen |
CVPR | 3 |
| 2021 | StyleDNA: A High-Fidelity Age and Gender Aware Kinship Face SynthesizerabstractHigh-fidelity kinship face synthesis receives increasing interest in this technology for visual kinship applications, including law enforcement, social media analysis, finding lost children, etc. However, it is a challenging task because of the unresolved ambiguities by the limited amount of available kinship data and severe data noise. To address these issues, we leverage the pretrained state-of-the-art face synthesis model, StyleGAN2, to assist the synthesis. With StyleGAN2, we develop three different kinship face synthesis strategies: (1) synthesis based on the kinship statistics, (2) synthesis using the latent code interpolation of the parents, and (3) synthesis based on the latent code interpolation from the disentangled age and gender independent latent space. The first two methods synthesize kinship faces through the direct manipulation of the original StyleGAN2 latent codes. The third one, on the other hand, is a two-stage synthesis method which first learns an age and gender invariant latent representation upon the one of StyleGAN2 to represent the genes. Combining with the maximal selection process to fuse the corresponding representation of parents, we form the genes of the child followed by them feeding back to StyleGAN2 for the final synthesis. With extensive ablation studies and experiments, we observe that all three methods can generate more photo-realistic and clearer faces than the previous state-of-the-art method. In addition, the third method achieves the best kinship verification results on the FIW dataset. Surprisingly, the subjective evaluation results of the three proposed methods are very close because typical humans are not good at recognizing unfamiliar kinship faces. Che-Hsien Lin, Hung-Chun Chen, Li-Chen Cheng, Shu-Chuan Hsu, Jun-Cheng Chen, Chih-Yu Wang 0001 |
FG | 5 |
| 2021 | Naturalistic Physical Adversarial Patch for Object DetectorsabstractMost prior works on physical adversarial attacks mainly focus on the attack performance but seldom enforce any restrictions over the appearance of the generated adversarial patches. This leads to conspicuous and attention-grabbing patterns for the generated patches which can be easily identified by humans. To address this issue, we pro-pose a method to craft physical adversarial patches for object detectors by leveraging the learned image manifold of a pretrained generative adversarial network (GAN) (e.g., BigGAN and StyleGAN) upon real-world images. Through sampling the optimal image from the GAN, our method can generate natural looking adversarial patches while maintaining high attack performance. With extensive experiments on both digital and physical domains and several independent subjective surveys, the results show that our proposed method produces significantly more realistic and natural looking patches than several state-of-the-art base-lines while achieving competitive attack performance.1 Yu-Chih-Tuan Hu, Jun-Cheng Chen, Bo-Han Kung, Kai-Lung Hua, Daniel Stanley Tan |
ICCV | 2 |
| 2021 | Squeeze And Reconstruct: Improved Practical Adversarial Defense Using Paired Image Compression And ReconstructionabstractAs shown in the previous literature, non-robust features of an image such as texture are both the secrets why deep neural networks achieve outstanding classification performance and the sources of adversarial examples. Image compression methods such as JPEG can be used to effectively defend against diverse adversarial attacks by eliminating these non-robust features in the pre-processing stage while significantly sacrificing clean accuracy. To address this issue, we present a squeeze-and-reconstruct framework which first performs image compression followed by image reconstruction to recover necessary details for the improved clean and robust accuracies. With extensive experiments on the challenging ImageNet dataset, the evaluation results demonstrate the effectiveness of the proposed method to defend against the Fast Gradient Sign Method and the powerful Projected Gradient Descent attacks in the white-box scenarios. In addition, the proposed approach also outperforms other common and off-the-shelf defense models in terms of both clean and robust accuracies. Bo-Han Kung, Pin-Chun Chen, Jun-Cheng Chen |
ICIP | 4 |
| 2021 | A Multi-Factor Combinations Enhanced Reversible Privacy Protection System for Facial ImagesabstractWith the abuse of deepfake and other deep learning technologies, anonymization and deanonymization for face images have become one of the essential tasks for privacy protection. Thus, we propose a novel reversible privacy protection framework for facial images based on conditional encoder and de-coder framework. For the purpose to increase the diversity and controllability over the anonymized faces, we also introduce facial attributes and a style vector from a reference back-ground face dataset and pretrained face recognition model and thus name the proposed framework as the Multi-factor Modifier (MfM) to achieve multi-factor facial de/re-identification. Specifically, with the correct password, our method produces near-original reconstructed images. Otherwise, it can generate photo-realistic and diverse anonymized images. With extensive experiments, it shows that the proposed approach can successfully anonymize face images in high fidelity according to the given conditions as compared with other methods and deanonymize without altering the facial data distributions. Yi-Lun Pan, Jun-Cheng Chen, Ja-Ling Wu |
ICME | 2 |
| 2021 | Heterogeneous Federated Learning Through Multi-Branch NetworkabstractRecently, federated learning has gained increasing attention for privacy-preserving computation since the learning paradigm allows to train models without the need for exchanging the data across different institutions distributively. However, heterogeneity of computational capabilities of edge devices is seldom discussed and analyzed in the current literature for heterogeneous federated learning. To address this issue, we propose a novel heterogeneous federated learning framework based on multi-branch deep neural network models which enable the selection of a proper sub-branch model for the client devices according to their computational capabilities. Meanwhile, we also present an aggregation method for model training, MFedAvg, that performs branch-wise averaging-based aggregation. With extensive experiments on MNIST, FashionMNIST, MedMNIST, and CIFAR-10, it demonstrates that our proposed approaches can achieve satisfactory performance with guaranteed convergence and effectively utilize all the available resources for training across different devices with lower communication cost than its homogeneous counterpart. Ching-Hao Wang, Kang-Yang Huang, Jun-Cheng Chen, Hong-Han Shuai, Wen-Huang Cheng |
ICME | 3 |
| 2020 | The Devil Is in the Details: Self-supervised Attention for Vehicle Re-identification
Pirazh Khorramshahi, Neehar Peri, Jun-Cheng Chen, Rama Chellappa |
ECCV (14) | 3 |
| 2020 | Face Feature Recovery via Temporal Fusion for Person SearchabstractSearching actors from videos by a single portrait image is a challenging task, due to large variations of video scenes and intra-person appearance. To tackle this problem, most recent works apply deep neural networks for detecting and extracting robust facial features for matching. However, when the face of an actor is not detected due to occlusion, such image-matching based strategies would not be applicable. To address the issue, we propose a unique framework of "Face Feature Recovery via Temporal Fusion" to synthesize virtual facial features by observing both temporal and contextual information. Once such face features are extracted, a simple extension to the k-nearest neighbors for re-ranking, "Iterative k-nearest Multi-fusion", is presented to utilize both face and body features for improved person search. We conduct extensive experiments to evaluate the performance of our framework on the challenging extended version of the Cast Search in Movies (ECSM) dataset [1]. Without utilizing tracklet information during training, the proposed approach still performs favorably against recent works in searching actors of interest from movie videos. Besides, we also show the proposed approach can be fused with them to further improve the performance. Cheng-Yu Fan, Chao-Peng Liu, Kuan-Chun Wang, Jiun-Hao Jhan, Yu-Chiang Frank Wang, Jun-Cheng Chen |
ICASSP | 6 |
| 2020 | Label Reuse for Efficient Semi-Supervised LearningabstractIn this paper, we propose a new learning strategy for semi-supervised deep learning algorithms, called label reuse, aiming to significantly reduce the expensive computational cost of pseudo label generation and the like for each unlabeled training instance since pseudo labels require to be repeatedly evaluated through the whole training process. For label reuse, we first divide the unlabeled training data into several partitions, replicate each partition in several copies, and place them consecutively in the training queue so as to reuse the pseudo labels computed at first time before invalidation. To evaluate the effectiveness of the proposed approach, we conduct extensive experiments on CIFAR-10 [1] and SVHN [2] by applying it upon the recent state-of-the-art semi-supervised deep learning approach, MixMatch [3]. The results demonstrate the proposed approach can not only significantly reduce the cost of pseudo label computation of MixMatch by a large amount but also keep comparable classification performance. Tsung-Hung Hsieh, Jun-Cheng Chen, Chu-Song Chen |
ICASSP | 2 |
| 2020 | An Experimental Evaluation of Recent Face Recognition Losses for Deepfake DetectionabstractDue to the recent breakthroughs of deep generative models, the fake faces, also known as deepfake which has been abused to deceive the general public, can be easily produced at scale and in very high fidelity. Many works focus on exploring various network architectures or various artifacts produced by deep generative models. Instead, in this work, we focus on the loss functions which have been shown to play a significant role in the context of face recognition. We perform a thorough study of several recent state-of-the-art losses commonly used in face recognition task for deepfake classification and detection since the current deepfake is highly related to face generation. With extensive experiments on the challenging FaceForensic++ and Celeb-DF datasets, the evaluation results provide a clear overview of the performance comparisons of different loss functions and generalization capability across different deepfake data. I-Hsuan Chen, Yu-Ru Ku, Jun-Cheng Chen |
ICPR | 5 |
| 2020 | S2SiamFC: Self-supervised Fully Convolutional Siamese Network for Visual TrackingabstractTo exploit rich information from unlabeled data, in this work, we propose a novel self-supervised framework for visual tracking which can easily adapt the state-of-the-art supervised Siamese-based trackers into unsupervised ones by utilizing the fact that an image and any cropped region of it can form a natural pair for self-training. Besides common geometric transformation-based data augmentation and hard negative mining, we also propose adversarial masking which helps the tracker to learn other context information by adaptively blacking out salient regions of the target. The proposed approach can be trained offline using images only without any requirement of manual annotations and temporal information from multiple consecutive frames. Thus, it can be used with any kind of unlabeled data, including images and video frames. For evaluation, we take SiamFC as the base tracker and name the proposed self-supervised method as S2SiamFC. Extensive experiments and ablation studies on the challenging VOT2016 and VOT2018 datasets are provided to demonstrate the effectiveness of the proposed method which not only achieves comparable performance to its supervised counterpart and other unsupervised methods requiring multiple frames. Chon-Hou Sio, Yu-Jen Ma, Hong-Han Shuai, Jun-Cheng Chen, Wen-Huang Cheng |
ACM Multimedia | 4 |
| 2019 | Unsupervised Domain-Specific Deblurring via Disentangled RepresentationsabstractImage deblurring aims to restore the latent sharp images from the corresponding blurred ones. In this paper, we present an unsupervised method for domain-specific, single-image deblurring based on disentangled representations. The disentanglement is achieved by splitting the content and blur features in a blurred image using content encoders and blur encoders. We enforce a KL divergence loss to regularize the distribution range of extracted blur attributes such that little content information is contained. Meanwhile, to handle the unpaired training data, a blurring branch and the cycle-consistency loss are added to guarantee that the content structures of the deblurred results match the original images. We also add an adversarial loss on deblurred results to generate visually realistic images and a perceptual loss to further mitigate the artifacts. We perform extensive experiments on the tasks of face and text deblurring using both synthetic datasets and real images, and achieve improved results compared to recent state-of-the-art deblurring methods. Boyu Lu, Jun-Cheng Chen, Rama Chellappa |
CVPR | 2 |
| 2019 | A Dual-Path Model With Adaptive Attention for Vehicle Re-IdentificationabstractIn recent years, attention models have been extensively used for person and vehicle re-identification. Most re-identification methods are designed to focus attention on key-point locations. However, depending on the orientation, the contribution of each key-point varies. In this paper, we present a novel dual-path adaptive attention model for vehicle re-identification (AAVER). The global appearance path captures macroscopic vehicle features while the orientation conditioned part appearance path learns to capture localized discriminative features by focusing attention on the most informative key-points. Through extensive experimentation, we show that the proposed AAVER method is able to accurately re-identify vehicles in unconstrained scenarios, yielding state of the art results on the challenging dataset VeRi-776. As a byproduct, the proposed system is also able to accurately predict vehicle key-points and shows an improvement of more than 7% over state of the art. The code for key-point estimation model is available at https://github.com/Pirazh/Vehicle_Key_ Point_Orientation_Estimation. Pirazh Khorramshahi, Amit Kumar 0013, Neehar Peri, Sai Saketh Rambhatla, Jun-Cheng Chen, Rama Chellappa |
ICCV | 5 |
| 2019 | Uncertainty Modeling of Contextual-Connections Between Tracklets for Unconstrained Video-Based Face RecognitionabstractUnconstrained video-based face recognition is a challenging problem due to significant within-video variations caused by pose, occlusion and blur. To tackle this problem, an effective idea is to propagate the identity from high-quality faces to low-quality ones through contextual connections, which are constructed based on context such as body appearance. However, previous methods have often propagated erroneous information due to lack of uncertainty modeling of the noisy contextual connections. In this paper, we propose the Uncertainty-Gated Graph (UGG), which conducts graph-based identity propagation between tracklets, which are represented by nodes in a graph. UGG explicitly models the uncertainty of the contextual connections by adaptively updating the weights of the edge gates according to the identity distributions of the nodes during inference. UGG is a generic graphical model that can be applied at only inference time or with end-to-end training. We demonstrate the effectiveness of UGG with state-of-the-art results in the recently released challenging Cast Search in Movies and IARPA Janus Surveillance Video Benchmark dataset. Jingxiao Zheng, Ruichi Yu, Jun-Cheng Chen, Boyu Lu, Carlos Domingo Castillo, Rama Chellappa |
ICCV | 3 |
| 2019 | A Proposal-Based Solution to Spatio-Temporal Action Detection in Untrimmed VideosabstractExisting approaches for spatio-temporal action detection in videos are limited by the spatial extent and temporal duration of the actions. In this paper, we present a modular system for spatio-temporal action detection in untrimmed surveillance videos. We propose a two stage approach. The first stage generates dense spatio-temporal proposals using hierarchical clustering and temporal jittering techniques on frame-wise object detections. The second stage is a Temporal Refinement I3D (TRI-3D) network that performs action classification and temporal refinement on the generated proposals. The object detection-based proposal generation step helps in detecting actions occurring in a small spatial region of a video frame, while temporal jittering and refinement helps in detecting actions of variable lengths. Experimental results on an unconstrained surveillance action detection dataset - DIVA - show the effectiveness of our system. For comparison, the performance of our system is also evaluated on the THUMOS'14 temporal action detection dataset. Joshua Gleason, Rajeev Ranjan 0003, Steven Schwarcz, Carlos Domingo Castillo, Jun-Cheng Chen, Rama Chellappa |
WACV | 5 |
| 2018 | Deep Density Clustering of Unconstrained FacesabstractIn this paper, we consider the problem of grouping a collection of unconstrained face images in which the number of subjects is not known. We propose an unsupervised clustering algorithm called Deep Density Clustering (DDC) which is based on measuring density affinities between local neighborhoods in the feature space. By learning the minimal covering sphere for each neighborhood, information about the underlying structure is encapsulated. The encapsulation is also capable of locating high-density region of the neighborhood, which aids in measuring the neighborhood similarity. We theoretically show that the encapsulation asymptotically converges to a Parzen window density estimator. Our experiments show that DDC is a superior candidate for clustering unconstrained faces when the number of subjects is unknown. Unlike conventional linkage and density-based methods that are sensitive to the selection operating points, DDC attains more consistent and improved performance. Furthermore, the density-aware property reduces the difficulty in finding appropriate operating points. Wei-An Lin, Jun-Cheng Chen, Carlos Domingo Castillo, Rama Chellappa |
CVPR | 2 |
| 2018 | A Real-Time Multi-Task Single Shot Face DetectorabstractFace, fiducial detection, and 3D head pose estimation are important face preprocessing modules for face recognition which are usually performed separately and loosely coupled. In this paper, we propose a unifying framework to simultaneously detect face, fiducial points, and head pose in real-time. In addition, since no single dataset contains all the required and best annotations, we develop a progressive training strategy to overcome the annotation discrepancy across different datasets. Extensive experiments on face detection, fiducial detection, and pose estimation benchmarks demonstrate the proposed approach can achieve comparable performance to a state-of-the-art system [1] but runs 60 times faster. (i.e., 20 frames per second). Jun-Cheng Chen, Wei-An Lin, Jingxiao Zheng, Rama Chellappa |
ICIP | 1 |
| 2018 | Unconstrained Still/Video-Based Face Verification with Deep Convolutional Neural Networks
Jun-Cheng Chen, Rajeev Ranjan 0003, Swami Sankaranarayanan, Amit Kumar 0013, Ching-Hui Chen, Vishal M. Patel, Carlos Domingo Castillo, Rama Chellappa |
Int. J. Comput. Vis. | 1 |
| 2018 | Proximity-Aware Hierarchical Clustering of unconstrained faces
Wei-An Lin, Jun-Cheng Chen, Rajeev Ranjan 0003, Ankan Bansal, Swami Sankaranarayanan, Carlos Domingo Castillo, Rama Chellappa |
Image Vis. Comput. | 2 |
| 2017 | Video-Based Face Association and IdentificationabstractIn this paper, we present a new video-based face identification algorithm, where the target (i.e., person of interest) in the probe video is only annotated once with a face bounding box in a frame and the video may consist of multiple shots. Most video face identification techniques assume that the video is of single shot, and thus the bounding boxes of the target face can be extracted by tracking a face across the video frames. Nevertheless, such automatic annotation is vulnerable to the drifting of the face tracker, and the face tracking algorithm is inadequate to associate the face images of the target across multiple shots. In this paper, we propose a target face association (TFA) technique that retrieves a set of representative face images in a given video that are likely to have the same identity as the target face. These face images are then utilized to construct a robust face representation of the target face for searching the corresponding subject in the gallery. Since two faces that appear in the same video frame cannot belong to the same person, such cannot-link constraints are utilized for learning a target-specific linear classifier for establishing the intra/inter-shot face association of the target. Experimental results on the newly released JANUS challenge set 3 (JANUS CS3) dataset show that our method generates robust representations from target-annotated videos and demonstrates good performance for the task of video-based face identification problem. Ching-Hui Chen, Jun-Cheng Chen, Carlos Domingo Castillo, Rama Chellappa |
FG | 2 |
| 2017 | A Proximity-Aware Hierarchical Clustering of FacesabstractIn this paper, we propose an unsupervised face clustering algorithm called “Proximity-Aware Hierarchical Clustering” (PAHC) that exploits the local structure of deep representations. In the proposed method, a similarity measure between deep features is computed by evaluating linear SVM margins. SVMs are trained using nearest neighbors of sample data, and thus do not require any external training data. Clus- ters are then formed by thresholding the similarity scores. We evaluate the clustering performance using three challenging un- constrained face datasets, including Celebrity in Frontal-Profile (CFP), IARPA JANUS Benchmark A (IJB-A), and JANUS Challenge Set 3 (JANUS CS3) datasets. Experimental results demonstrate that the proposed approach can achieve significant improvements over state-of-the-art methods. Moreover, we also show that the proposed clustering algorithm can be applied to curate a set of large-scale and noisy training dataset while maintaining sufficient amount of images and their variations due to nuisance factors. The face verification performance on JANUS CS3 improves significantly by finetuning a DCNN model with the curated MS-Celeb-1M dataset which contains over three million face images. Wei-An Lin, Jun-Cheng Chen, Rama Chellappa |
FG | 2 |
| 2017 | Face and Image Representation in Deep CNN FeaturesabstractFace recognition algorithms based on deep convolutional neural networks (DCNNs) have made progress on the task of recognizing faces in unconstrained viewing conditions. These networks operate with compact feature-based face representations derived from learning a very large number of face images. Although the learned feature sets produced by DCNNs can be highly robust to changes in viewpoint, illumination, and appearance, little is known about the nature of the face code that emerges at the top level of these networks. We analyzed the DCNN features produced by two recent face recognition algorithms. In the first set of experiments, we used the top-level features from the DCNNs as input into linear classifiers aimed at predicting metadata about the images. The results showed that the DCNN features contained surprisingly accurate information about the yaw and pitch of a face, and about whether the input face came from a still image or a video frame. In the second set of experiments, we measured the extent to which individual DCNN features operated in a view-dependent or view-invariant manner for different identities. We found that view-dependent coding was a characteristic of the identities rather than the DCNN features– with some identities coded consistently in a view-dependent way and others in a view-independent way. In our third analysis, we visualized the DCNN feature space for 24,000+ images of 500 identities. Images in the center of the space were uniformly of low quality (e.g., extreme views, face occlusion, poor contrast, low resolution). Image quality increased monotonically as a function of distance from the origin. This result suggests that image quality information is available in the DCNN features, such that consistently average feature values reflect coding failures that reliably indicate poor or unusable images. Combined, the results offer insight into the coding mechanisms that support robust representation of faces in DCNNs. Connor J. Parde, Carlos Domingo Castillo, Matthew Q. Hill, Y. Ivette Colon, Swami Sankaranarayanan, Jun-Cheng Chen, Alice J. O'Toole |
FG | 6 |
| 2017 | Deep Heterogeneous Feature Fusion for Template-Based Face RecognitionabstractAlthough deep learning has yielded impressive performance for face recognition, many studies have shown that different networks learn different feature maps: while some networks are more receptive to pose and illumination others appear to capture more local information. Thus, in this work, we propose a deep heterogeneous feature fusion network to exploit the complementary information present in features generated by different deep convolutional neural networks (DCNNs) for template-based face recognition, where a template refers to a set of still face images or video frames from different sources which introduces more blur, pose, illumination and other variations than traditional face datasets. The proposed approach efficiently fuses the discriminative information of different deep features by 1) jointly learning the non-linear high-dimensional projection of the deep features and 2) generating a more discriminative template representation which preserves the inherent geometry of the deep features in the feature space. Experimental results on the IARPA Janus Challenge Set 3 (Janus CS3) dataset demonstrate that the proposed method can effectively improve the recognition performance. In addition, we also present a series of covariate experiments on the face verification task for in-depth qualitative evaluations for the proposed approach. Navaneeth Bodla, Jingxiao Zheng, Hongyu Xu, Jun-Cheng Chen, Carlos Domingo Castillo, Rama Chellappa |
WACV | 4 |
| 2017 | Pose-Robust Face Verification by Exploiting Competing TasksabstractIn this paper, we propose a pose-robust metric learning framework for unconstrained face verification by jointly optimizing face and pose verification tasks. We learn a joint model for these two tasks and explicitly discourage the information sharing between pose and identity verification metrics so as to mitigate the information contained in the pose verification task leading to making the identity metrics for face verification more pose-robust. Specifically, we use the joint Bayesian metric learning framework to learn the metrics for both tasks and enforce an orthogonal regularization constraint on the learned projection matrices for the two tasks. The pose labels used for training the joint model are automatically estimated and do not require extra annotations. An efficient stochastic gradient descent (SGD) algorithm is used to solve the optimization problem. We conduct extensive experiments on three challenging unconstrained face datasets and show promising results compared to state-of-the-art methods. Boyu Lu, Jingxiao Zheng, Jun-Cheng Chen, Rama Chellappa |
WACV | 3 |
| 2016 | Fisher vector encoded deep convolutional features for unconstrained face verificationabstractWe present a method to combine the Fisher vector representation and the Deep Convolutional Neural Network (DCNN) features to generate a rerpesentation, called the Fisher vector encoded DCNN (FV-DCNN) features, for unconstrained face verification. One of the key features of our method is that spatial and appearance information are simultaneously processed when learning the Gaussian mixture model to encode the DCNN features. Evaluations on two challenging verification datasets show that the proposed FV-DCNN method is able to capture the salient local features and also performs well when compared to many state-of-the-art face verification methods. Jun-Cheng Chen, Jingxiao Zheng, Vishal M. Patel, Rama Chellappa |
ICIP | 1 |
| 2016 | Regularized metric adaptation for unconstrained face verificationabstractIn this work, we propose a metric adaptation method for set-based face verification and evaluate it on the newly released IARPA Janus Benchmark A (IJB-A) dataset and its extended version, the Janus Challenging Set 2 (CS2). A template-specific metric is trained to adaptively learn the discriminative information in test templates and the negative training set, which contains subjects that are mutually exclusive to subjects in test templates. The proposed regularized joint Bayesian metric learning framework not only alleviates the over-fitting problem but also provides a way to efficiently reduce the model size. We also analyze the selection of the compact and representative negative set to speed up the training time and to reduce storage space. Experiments on the IJB-A and CS2 datasets yield promising results. Boyu Lu, Jun-Cheng Chen, Rama Chellappa |
ICPR | 2 |
| 2016 | VLAD encoded Deep Convolutional features for unconstrained face verificationabstractWe present a method for combining the Vector of Locally Aggregated Descriptor (VLAD) feature encoding with Deep Convolutional Neural Network (DCNN) features for unconstrained face verification. One of the key features of our method, called the VLAD-encoded DCNN (VLAD-DCNN) features, is that spatial and appearance information are simultaneously processed to learn an improved discriminative representation. Evaluations on the challenging IARPA Janus Benchmark A (IJB-A) face dataset show that the proposed VLAD-DCNN method is able to capture the salient local features and yield promising results for face verification. Furthermore, we show that additional performance gains can be achieved by simply fusing the VLAD-DCNN features that capture the local variations with the traditional DCNN features which characterize more global features. Jingxiao Zheng, Jun-Cheng Chen, Navaneeth Bodla, Vishal M. Patel, Rama Chellappa |
ICPR | 2 |
| 2016 | Unconstrained face verification using deep CNN featuresabstractIn this paper, we present an algorithm for unconstrained face verification based on deep convolutional features and evaluate it on the newly released IARPA Janus Benchmark A (IJB-A) dataset as well as on the traditional Labeled Face in the Wild (LFW) dataset. The IJB-A dataset includes real-world unconstrained faces from 500 subjects with full pose and illumination variations which are much harder than the LFW and Youtube Face (YTF) datasets. The deep convolutional neural network (DCNN) is trained using the CASIA-WebFace dataset. Results of experimental evaluations on the IJB-A and the LFW datasets are provided. Jun-Cheng Chen, Vishal M. Patel, Rama Chellappa |
WACV | 1 |
| 2016 | Frontal to profile face verification in the wildabstractWe have collected a new face data set that will facilitate research in the problem of frontal to profile face verification `in the wild'. The aim of this data set is to isolate the factor of pose variation in terms of extreme poses like profile, where many features are occluded, along with other `in the wild' variations. We call this data set the Celebrities in Frontal-Profile (CFP) data set. We find that human performance on Frontal-Profile verification in this data set is only slightly worse (94.57% accuracy) than that on Frontal-Frontal verification (96.24% accuracy). However we evaluated many state-of-the-art algorithms, including Fisher Vector, Sub-SML and a Deep learning algorithm. We observe that all of them degrade more than 10% from Frontal-Frontal to Frontal-Profile verification. The Deep learning implementation, which performs comparable to humans on Frontal-Frontal, performs significantly worse (84.91% accuracy) on Frontal-Profile. This suggests that there is a gap between human performance and automatic face recognition methods for large pose variation in unconstrained images. Roni Sengupta, Jun-Cheng Chen, Carlos Domingo Castillo, Vishal M. Patel, Rama Chellappa, David Jacobs 0001 |
WACV | 2 |
| 2015 | Landmark-based fisher vector representation for video-based face verificationabstractUnconstrained video-based face verification is a challenging problem because of dramatic variations in pose, illumination, and image quality of each face in a video. In this paper, we propose a landmark-based Fisher vector representation for video-to-video face verification. The proposed representation encodes dense multi-scale SIFT features extracted from patches centered at detected facial landmarks, and face similarity is computed with the distance measure learned from joint Bayesian metric learning. Experimental results demonstrate that our approach achieves significantly better performance than other competitive video-based face verification algorithms on two challenging unconstrained video face dataseis, Multiple Biometric Grand Challenge (MBGC) and Face and Ocular Challenge Series (FOCS). Jun-Cheng Chen, Vishal M. Patel, Rama Chellappa |
ICIP | 1 |
| 2014 | Dictionary-based video face recognition using dense multi-scale facial landmark featuresabstractIn video-based face recognition, different video sequences of the same subject contain variations in pose, illumination, and expression which contribute to the challenges in designing an effective video-based face-recognition system. In this paper, we propose a dictionary-based approach using dense and high-dimensional features extracted from multi-scale patches centered at detected facial landmarks for video-to-video face identification and verification. Experiments using unconstrained video sequences from Multiple Biometric Grand Challenge (MBGC) and Face and Ocular Challenge Series (FOCS) datasets show that our method performs significantly better than many state-of-the-art video-based face recognition algorithms. Jun-Cheng Chen, Vishal M. Patel, Huy Tho Ho, Rama Chellappa |
ICIP | 1 |
| 2007 | Film Narrative Exploration Through the Analysis of Aesthetic Elements
Chia-Wei Wang, Wen-Huang Cheng, Jun-Cheng Chen, Shu-Sian Yang, Ja-Ling Wu |
MMM (1) | 3 |
| 2006 | Tiling slideshowabstractThis paper presents a new medium, called tiling slideshow, to display photos in a tile-like manner, coordinating with the pace of background music. In contrast to the conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Motivated by the concepts of technical writing, each displaying layout is composed of a larger topic photo and several small-size supportive photos. Based on this idea, the proposed tiling slideshow system consists of three major components: image clustering, music analyzer, and layout organizer. Given the limited displaying space, we consider the context and relationship between photos and model the layout organization as a constrainted optimization problem. Experiments on real consumer photograph collections show that the novel displaying method gives users more pleasant browsing experience than the methods that focus only on single photograph display. Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu |
ACM Multimedia | 1 |
| 2006 | Audiovisual slideshow: present your journey by photosabstractThis demonstration presents a novel way to systematically display photos and enhance the viewing experience of photo browsing. In contrast to conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Moreover, the displaying pace is coordinated with the beat of the user-selected incidental music. To automatically generate the audiovisual slideshow, we develop a system that consists of three main components: photo analysis, music analysis, and audiovisual composition. Audiovisual content analysis and cross-media synchronization issues are addressed in this work. This novel demonstration is especially suitable to present photos taken in a journey. It vigorously presents the delights of traveling and helps us recall or experience the trip. Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu |
ACM Multimedia | 1 |
| 2005 | An adaptive edge detection based colorization algorithm and its applicationsabstractColorization is a computer-assisted process for adding colors to grayscale images or movies. It can be viewed as a process for assigning a three-dimensional color vector (YUV or RGB) to each pixel of a grayscale image. In previous works, with some color hints the resultant chrominance value varies linearly with that of the luminance. However, it is easy to find that existing methods may introduce obvious color bleeding, especially, around region boundaries. It then needs extra human-assistance to fix these artifacts, which limits its practicability. Facing such a challenging issue, we introduce a general and fast colorization methodology with the aid of an adaptive edge detection scheme. By extracting reliable edge information, the proposed approach may prevent the colorization process from bleeding over object boundaries. Next, integration of the proposed fast colorization scheme to a scribble-based colorization system, a modified color transferring system and a novel chrominance coding approach are investigated. In our experiments, each system exhibits obvious improvement as compared to those corresponding previous works. Yi-Chin Huang, Yi-Shin Tung, Jun-Cheng Chen, Sung-Wen Wang, Ja-Ling Wu |
ACM Multimedia | 3 |