Hung-Jen Chen 0001

dblp:95/7000-1 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-3129-1595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Foreground Focus: Enhancing Coherence and Fidelity in Camouflaged Image Generation
abstract
Camouflaged image generation is emerging as a solution to data scarcity in camouflaged vision perception, offering a cost-effective alternative to data collection and labeling. Recently, the state-of-the-art approach successfully generates camouflaged images using only foreground objects. However, it faces two critical weaknesses: 1) the background knowledge does not integrate effectively with foreground features, resulting in a lack of foreground-background coherence (e.g., color discrepancy); 2) the generation process does not prioritize the fidelity of foreground objects, which leads to distortion, particularly for small objects. To address these issues, we propose a Foreground-Aware Camouflaged Image Generation (FACIG) model. Specifically, we introduce a Foreground-Aware Feature Integration Module (FAFIM) to strengthen the integration between foreground features and background knowledge. In addition, a Foreground-Aware Denoising Loss is designed to enhance foreground reconstruction supervision. Experiments on various datasets show our method outperforms previous methods in overall camouflaged image quality and foreground fidelity.
Pei-Chi Chen, Chan-Feng Hsu, Hung-Jen Chen 0001, Hong-Han Shuai, Wen-Huang Cheng
ICME5
2024 EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
abstract
Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing tasks but is still unexplored in vision emotion understanding. In this work, we focus on enhancing the model's proficiency in understanding and adhering to instructions related to emotional contexts. Initially, we identify key visual clues critical to visual emotion recognition. Subsequently, we introduce a novel GPT-assisted pipeline for generating emotion visual instruction data, effectively addressing the scarcity of annotated instruction data in this domain. Expanding on the groundwork established by InstructBLIP, our proposed EmoVIT architecture incorporates emotion-specific instruction data, leveraging the powerful capabilities of Large Language Models to enhance performance. Through extensive experiments, our model showcases its proficiency in emotion classification, adeptness in affective reasoning, and competence in comprehending humor. The comparative analysis provides a robust benchmark for Emotion Visual Instruction Tuning in the era of LLMs, providing valuable insights and opening avenues for future exploration in this domain. Our code is available at https://github.com/aimmemotion/EmoVIT.
Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen 0001, Chan-Feng Hsu, Hong-Han Shuai, Wen-Huang Cheng
CVPR4
2024 Hierarchically Aggregated Identification Transformer Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) targets the segmentation of objects hidden in intricate environments, a task complicated by the pronounced similarities between objects and their surroundings. The diverse appearances of camouflaged objects, such as different view angles, partial visibilities, and ambiguous forms, further exacerbate this challenge. To address these issues, we introduce the Hierarchically Aggregated Identification Transformer Network (HAIT-Net). HAITNet harnesses local and global features to refine object localization by employing multi-scale transformer features unified through the Feature Cascaded Fusion Module (FCFM). To tackle ambiguity from indistinct textures, we present the Graph-based Low-level Feature Enhancement Module (GLFEM) and Graph-based Feature Aggregation Module (GFAM). GLFEM enhances texture representation in ambiguous areas, while GFAM reduces false positives and refines prediction maps by discerning contextual relationships. Experimental results on three widely used datasets demonstrate that the proposed HAITNet outperforms the state-of-the-art approaches. Our code is available at https://github.com/underlmao/HAITNet.
Thanh Hai Phung, Hung-Jen Chen 0001, Hong-Han Shuai
ICME2
2023 Most Important Person-guided Dual-branch Cross-Patch Attention for Group Affect Recognition
abstract
Group affect refers to the subjective emotion that is evoked by an external stimulus in a group, which is an important factor that shapes group behavior and outcomes. Recognizing group affect involves identifying important individuals and salient objects among a crowd that can evoke emotions. However, most existing methods lack attention to affective meaning in group dynamics and fail to account for the contextual relevance of faces and objects in group-level images. In this work, we propose a solution by incorporating the psychological concept of the Most Important Person (MIP), which represents the most noteworthy face in a crowd and has affective semantic meaning. We present the Dual-branch Cross-Patch Attention Transformer (DCAT) which uses global image and MIP together as inputs. Specifically, we first learn the informative facial regions produced by the MIP and the global context separately. Then, the Cross-Patch Attention module is proposed to fuse the features of MIP and global context together to complement each other. Our proposed method outperforms state-of-the-art methods on GAF 3.0, GroupEmoW, and HECO datasets. Moreover, we demonstrate the potential for broader applications by showing that our proposed model can be transferred to another group affect task, group cohesion, and achieve comparable results.
Ming-Xian Lee, Tzu-Jui Chen, Hung-Jen Chen 0001, Hou-I Liu, Hong-Han Shuai, Wen-Huang Cheng
ICCV4
2022 Finding the Achilles Heel: Progressive Identification Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to segment objects assimilating into their surroundings. The key challenge for COD is that there are existing high intrinsic similarities between the target object and the background. To solve this challenging problem, we propose the Cascaded Decamouflage Module to progressively improve the prediction map, where each decamouflage module is composed of the region enhancement block and the reverse attention mining block to accurately detect the camouflaged object and obtain complete target objects. In addition, we introduce the classification-based label reweighting to produce the gated label maps as the supervision for assisting the network to capture the most conspicuous region of a camouflaged object and obtain the target object entirely. Extensive experiments on three challenging datasets demonstrate that the proposed model outperforms state-of-the-art methods under different evaluation metrics.
Mu-Chun Chou, Hung-Jen Chen 0001, Hong-Han Shuai
ICME2
2021 Re-Attention Is All You Need: Memory-Efficient Scene Text Detection via Re-Attention on Uncertain Regions
abstract
Scene text detection plays an important role on vision-based robot navigation to many potential landmarks such as nameplates, information signs, floor button in the elevators. Recently, scene text detection with segmentation-based methods has been receiving more and more attention. The segmentation results can be used to efficiently predict scene text of various shapes, such as irregular text in most scene text images. However, two kinds of texts remain unsolved: 1) tiny and 2) blurry instances. Moreover, the annotations for tiny/blurry texts are usually ignored during training, while tiny/blurry texts can still offer visual auxiliaries for robots to understand the world. Therefore, in this paper, we propose a new approach to effectively detect both clear and blurry texts. Specifically, we propose a re-attention module without increasing the learnable parameters, which first predicts the region of texts as the candidate region and leverages the same network to detect the candidate region again for reducing the required memory. Moreover, to avoid the errors from the first detection propagating to the re-attended area, we propose a new fusion module that learns to integrate the results of the re-attended regions and the first prediction. Experimental results manifest that the proposed method outperforms state-of-the-art methods on four challenging datasets.
Hsiang-Chun Chang, Hung-Jen Chen 0001, Yu-Chia Shen, Hong-Han Shuai, Wen-Huang Cheng
IROS2
2020 Character-Preserving Coherent Story Visualization
Yun-Zhu Song, Zhi Rui Tam, Hung-Jen Chen 0001, Huiao-Han Lu, Hong-Han Shuai
ECCV (17)3
2019 BeautyGlow: On-Demand Makeup Transfer Framework With Reversible Generative Network
abstract
As makeup has been widely-adopted for beautification, finding suitable makeup by virtual makeup applications becomes popular. Therefore, a recent line of studies proposes to transfer the makeup from a given reference makeup image to the source non-makeup one. However, it is still challenging due to the massive number of makeup combinations. To facilitate on-demand makeup transfer, in this work, we propose BeautyGlow that decompose the latent vectors of face images derived from the Glow model into makeup and non-makeup latent vectors. Since there is no paired dataset, we formulate a new loss function to guide the decomposition. Afterward, the non-makeup latent vector of a source image and makeup latent vector of a reference image and are effectively combined and revert back to the image domain to derive the results. Experimental results show that the transfer quality of BeautyGlow is comparable to the state-of-the-art methods, while the unique ability to manipulate latent vectors allows BeautyGlow to realize on-demand makeup transfer.
Hung-Jen Chen 0001, Ka-Ming Hui, Szu-Yu Wang, Li-Wu Tsao, Hong-Han Shuai, Wen-Huang Cheng
CVPR1