Chun-Hsiao Yeh

dblp:203/3213 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
abstract
Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown impressive advances in high-level reasoning and planning, they frequently fall short when confronted with multi-view geometric consistency and cross-view correspondence. To comprehensively evaluate the challenges of MLLMs in multi-view scene reasoning, we introduce All-Angles Bench, a human carefully benchmark with over 2,100 question-answer pairs from 90 diverse, real-world scenes. Our broad evaluation across 38 general-purpose and 3D spatial reasoning MLLMs reveals a substantial performance gap compared to humans. More critically, our analysis identifies two root failure modes: (1) cross-view object mismatch—the inability to establish consistent object correspondence across views; and (2) cross-view spatial misalignment—the failure to infer accurate camera poses and spatial layouts. These findings underscore a lack of multi-view awareness in current MLLMs, calling for architectural innovations beyond prompt tuning alone. We believe that our benchmark offers valuable insights toward building spatially-intelligent MLLMs.
Chun-Hsiao Yeh, Shengbang Tong, Ta Ying Cheng, Ruoyu Wang 0014, Tianzhe Chu, Yuexiang Zhai, Yubei Chen, Shenghua Gao, Yi Ma 0001
AAAI1
2026 Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
abstract
Recent diffusion-based image editing methods have made great strides in text-guided tasks but often struggle with complex, indirect instructions. Additionally, current models frequently exhibit poor identity preservation, unintended edits, or rely on manual masks. To overcome these limitations, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that bridges user intent with editing model capabilities. X-Planner uses chain-of-thought reasoning to systematically break down complex instructions into simpler sub-instructions. For each one, X-Planner automatically generates precise edit types and segmentation masks, enabling localized, identity-preserving edits without applying external tools or models during inference. To enable the training of such a planner, we also introduce a fully automated, reproducible pipeline to generate large-scale, high-quality training data. Our complete system achieves state-of-the-art results on both existing and newly proposed complex instruction-based editing benchmarks.
Chun-Hsiao Yeh, Yilin Wang 0002, Nanxuan Zhao, Hao (Richard) Zhang, Krishna Kumar Singh
AAAI1
2024 Insight: A Multi-modal Diagnostic Pipeline Using LLMs for Ocular Surface Disease Diagnosis
Chun-Hsiao Yeh, Andrew D. Graham, Andrea J. Liu, Yubei Chen, Yi Ma 0001, Meng C. Lin
MICCAI (1)1
2023 Meta-Personalizing Vision-Language Models to Find Named Instances in Video
abstract
Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently struggle with personalized searches for moments in a video where a specific object instance such as “My dog Biscuit” appears. We present the following three contributions to address this problem. First, we describe a method to meta-personalize a pre-trained VLM, i.e., learning how to learn to personalize a VLM at test time to search in video. Our method extends the VLM's token vocabulary by learning novel word embeddings specific to each instance. To capture only instance-specific features, we represent each instance embedding as a combination of shared and learned global category features. Second, we propose to learn such personalization without explicit human supervision. Our approach automatically identifies moments of named visual instances in video using transcripts and vision-language similarity in the VLM's embedding space. Finally, we introduce This-Is-My, a personal video instance retrieval benchmark. We evaluate our approach on This-Is-My and Deep-Fashion2 and show that we obtain a 15% relative improvement over the state of the art on the latter dataset.
Chun-Hsiao Yeh, Bryan C. Russell, Josef Sivic, Fabian Caba Heilbron, Simon Jenni
CVPR1
2022 Decoupled Contrastive Learning
Chun-Hsiao Yeh, Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu, Yubei Chen, Yann LeCun
ECCV (26)1
2022 SAGA: Self-Augmentation with Guided Attention for Representation Learning
abstract
Self-supervised training that elegantly couples contrastive learning with a wide spectrum of data augmentation techniques has been shown to be a successful paradigm for representation learning. However, current methods implicitly maximize the agreement between differently augmented views of the same sample, which may perform poorly in certain situations. For example, considering an image comprising a boat on the sea, one augmented view is cropped solely from the boat and the other from the sea, whereas linking these two to form a positive pair could be misleading. To resolve this issue, we introduce a Self-Augmentation with Guided Attention (SAGA) strategy, which augments input data based on predictive attention to learn representations rather than simply applying off-the-shelf augmentation schemes. As a result, the proposed self-augmentation framework enables feature learning to enhance the robustness of representation.
Chun-Hsiao Yeh, Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu
ICASSP1
2022 Face anti-spoofing detection based on multi-scale image quality assessment
Herng-Hua Chang, Chun-Hsiao Yeh
Image Vis. Comput.2
2018 Face Liveness Detection Based on Perceptual Image Quality Assessment Features with Multi-scale Analysis
abstract
Vulnerability of recognition systems to spoofing attacks (presentation attacks) is still an open security issue in the biometrics domain. Among all biometric traits, face is exposed to the most serious threat since it is particularly easy to access and reproduce. In this paper, an effective approach against face spoofing attacks based on perceptual image quality assessment features with multiscale analysis is presented. First, we demonstrate that the recently proposed blind image quality evaluator (BIQE) is effective in detecting spoofing attacks. Next, we combine the BIQE with an image quality assessment model called effective pixel similarity deviation (EPSD), which we propose to obtain the standard deviation of the gradient magnitude similarity map by selecting effective pixels in the image. A total number of 21 features acquired from the BIQE and EPSD constitute the multi-scale descriptor for classification. Extensive experiments based on both intradataset and cross-dataset protocols were performed using three existing benchmarks, namely, Replay-Attack, CASIA, and UVAD. The proposed algorithm demonstrated its superiority in detecting face spoofing attacks over many state of the art methods. We believe that the incorporation of the image quality assessment knowledge into face liveness detection is promising to improve the overall accuracy.
Chun-Hsiao Yeh, Herng-Hua Chang
WACV1