EDBT 2026 Demo / reviewers in the wild / expert
Chiao-An Yang
dblp:312/7959
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-1947-1331ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Local Scale Equivariance with Latent Deep Equilibrium CanonicalizerabstractScale variation is a fundamental challenge in computer vision. Objects of the same class can have different sizes, and their perceived size is further affected by the distance from the camera. These variations are local to the objects, i.e., different object sizes may change differently within the same image. To effectively handle scale variations, we present a deep equilibrium canonicalizer (DEC) to improve the local scale equivariance of a model. DEC can be easily incorporated into existing network architectures and can be adapted to a pre-trained model. Notably, we show that on the competitive ImageNet benchmark, DEC improves both model performance and local scale consistency across four popular pre-trained deep-nets, e.g., ViT, DeiT, Swin, and BEiT. Our code is available at https://github.com/ashiq24/local-scale-equivariance. Chiao-An Yang, Michael N. Cheng, Lim Jun Hao, Jeremiah Jiang, Teck-Yian Lim, Raymond A. Yeh |
ICCV | 2 |
| 2025 | Toward Long-Tailed Online Anomaly Detection Through Class-Agnostic ConceptsabstractAnomaly detection (AD) identifies the defect regions of a given image. Recent works have studied AD, focusing on learning AD without abnormal images, with long-tailed distributed training data, and using a unified model for all classes. In addition, online AD learning has also been explored. In this work, we expand in both directions to a realistic setting by considering the novel task of long-tailed online AD (LTOAD). We first identified that the offline state-of-the-art LTAD methods cannot be directly applied to the online setting. Specifically, LTAD is class-aware, requiring class labels that are not available in the online setting. To address this challenge, we propose a class-agnostic framework for LTAD and then adapt it to our online learning setting. Our method outperforms the SOTA baselines in most offline LTAD settings, including both the industrial manufacturing and the medical domain. In particular, we observe +4.63% image-AUROC on MVTec even compared to methods that have access to class labels and the number of classes. In the most challenging long-tailed online setting, we achieve +0.53% image-AUROC compared to baselines. Our LTOAD benchmark is released here: https://doi.org/10.5281/zenodo.16283852 . Chiao-An Yang, Kuan-Chuan Peng, Raymond A. Yeh |
ICCV | 1 |
| 2025 | Heatmap Regression without Soft-Argmax for Facial Landmark DetectionabstractFacial landmark detection is an important task in computer vision with numerous applications, such as head pose estimation, expression analysis, face swapping, etc. Heatmap regression-based methods have been widely used to achieve state-of-the-art results in this task. These methods involve computing the argmax over the heatmaps to predict a landmark. Since argmax is not differentiable, these methods use a differentiable approximation, Soft-argmax, to enable end-to-end training on deep-nets. In this work, we revisit this long-standing choice of using Soft-argmax and demonstrate that it is not the only way to achieve strong performance. Instead, we propose an alternative training objective based on the classic structured prediction framework. Empirically, our method achieves state-of-the-art performance on three facial landmark benchmarks (WFLW, COFW, and 300W), converging 2.2x faster during training while maintaining better/competitive accuracy. Our code is available here: https://github.com/ca-joe-yang/regression-without-softarg. Chiao-An Yang, Raymond A. Yeh |
ICCV | 1 |
| 2024 | Deep Nets with Subsampling Layers Unwittingly Discard Useful Activations at Test-Time
Chiao-An Yang, Ziwei Liu 0002, Raymond A. Yeh |
ECCV (21) | 1 |
| 2024 | Learning to Obstruct Few-Shot Image Classification over Restricted Classes
Amber Yijia Zheng, Chiao-An Yang, Raymond A. Yeh |
ECCV (20) | 2 |
| 2024 | Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial AttentionabstractMulti-object 3D Grounding involves locating 3D boxes based on a given query phrase from a point cloud. It is a challenging and significant task that has numerous applications in visual understanding, human-computer interaction, and robotics. To tackle this challenge, we introduce D-LISA, a two-stage approach that incorporates three innovations. First, a dynamic vision module that enables a variable and learnable number of box proposals. Second, a dynamic camera positioning that extracts features for each proposal. Third, a language-informed spatial attention module that better reasons over the proposals to output the final prediction. Empirically, experiments show that our method outperforms the state-of-the-art methods on multi-object 3D grounding by 12.8% (absolute) and is competitive in single-object 3D grounding. Haomeng Zhang, Chiao-An Yang, Raymond A. Yeh |
NeurIPS | 2 |
| 2023 | Target-Free Text-Guided Image ManipulationabstractWe tackle the problem of target-free text-guided image manipulation, which requires one to modify the input reference image based on the given text instruction, while no ground truth target image is observed during training. To address this challenging task, we propose a Cyclic-Manipulation GAN (cManiGAN) in this paper, which is able to realize where and how to edit the image regions of interest. Specifically, the image editor in cManiGAN learns to identify and complete the input image, while cross-modal interpreter and reasoner are deployed to verify the semantic correctness of the output image based on the input instruction. While the former utilizes factual/counterfactual description learning for authenticating the image semantics, the latter predicts the "undo" instruction and provides pixel-level supervision for the training of cManiGAN. With the above operational cycle-consistency, our cManiGAN can be trained in the above weakly supervised setting. We conduct extensive experiments on the datasets of CLEVR and COCO datasets, and the effectiveness and generalizability of our proposed method can be successfully verified. Project page: sites.google.com/view/wancyuanfan/projects/cmanigan. Wan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank Wang |
AAAI | 3 |
| 2023 | Consistent and Multi-Scale Scene Graph Transformer for Semantic-Guided Image OutpaintingabstractThe task of image outpainting extends an image beyond its boundaries with semantically plausible content. Recently, Scene Graph Transformer (SGT) introduced a transformer architecture to leverage scene graph guidance for image outpainting. Despite its success, we identified two shortcomings: (a) SGT uses a positional encoding that was originally proposed for 1D signal; (b) SGT uses a scene graph attention layer that propagates information between neighboring nodes which limited the model to learning local graph features. To address these issues, we propose incorporating Laplacian positional encoding and introducing a multiscale scene graph attention into SGT. Extensive results on MS-COCO and Visual Genome show that our proposed approach generates more plausible outpainted images with higher quality. Chiao-An Yang, Meng-Lin Wu, Raymond A. Yeh, Yu-Chiang Frank Wang |
ICIP | 1 |
| 2022 | Scene Graph Expansion for Semantics-Guided Image OutpaintingabstractIn this paper, we address the task of semantics-guided image outpainting, which is to complete an image by generating semantically practical content. Different from most existing image outpainting works, we approach the above task by understanding and completing image semantics at the scene graph level. In particular, we propose a novel network of Scene Graph Transformer (SGT), which is designed to take node and edge features as inputs for modeling the associated structural information. To better understand and process graph-based inputs, our SGT uniquely performs feature attention at both node and edge levels. While the former views edges as relationship regularization, the latter observes the co-occurrence of nodes for guiding the attention process. We demonstrate that, given a partial input image with its layout and scene graph, our SGT can be applied for scene graph expansion and its conversion to a complete layout. Following state-of-the-art layout-to-image conversions works, the task of image outpainting can be completed with sufficient and practical semantics introduced. Extensive experiments are conducted on the datasets of MS-COCO and Visual Genome, which quantitatively and qualitatively confirm the effectiveness of our proposed SGT and outpainting frameworks. Chiao-An Yang, Cheng-Yo Tan, Wan-Cyuan Fan, Cheng-Fu Yang, Meng-Lin Wu, Yu-Chiang Frank Wang |
CVPR | 1 |
| 2021 | Robust Image Outpainting With Learnable Image MarginsabstractGiven a partial image input, image outpainting is to produce the desirable output by recovering or extending the surrounding image regions. While existing image outpainting methods achieve impressive results based on the recent advances of deep learning, they either lack the ability to extend image regions in arbitrary directions or require the filling image margins to be given in advance. To address this challenging task, we propose a unique deep learning framework for robust image outpainting, which consists of a margin prediction network and a teacher-student-based network for producing outpainted images. Our proposed model does not require image filling margins to be known beforehand, while both image appearance and perceptual feature consistencies can be jointly enforced. Our experiments quantitatively and qualitatively verify the effectiveness of our method, which is shown to perform favorably against baseline and state-of-the-art image outpainting works. Cheng-Yo Tan, Chiao-An Yang, Shang-Fu Chen, Meng-Lin Wu, Yu-Chiang Frank Wang |
ICIP | 2 |
| 2021 | Soft Ranking Threshold Losses For Image RetrievalabstractThis paper proposes a novel loss, soft ranking threshold loss, for driving deep networks to learn better representations for image retrieval. Instead of working in the metric space, our loss works in the rank space which has a more uniform distribution and explicit scale and bounds. Our loss reduces the ranks of the distances between anchor-positive pairs below the threshold while increasing the ones between anchor-negative pairs above the threshold. In addition to the basic form, two extensions are proposed for improving the effectiveness: hard thresholds and ranking margin. Experiments show that the proposed loss outperforms the state-of-the-art losses on image retrieval applications. Chiao-An Yang, Zhixiang Wang 0001, Yen-Yu Lin, Yung-Yu Chuang |
ICIP | 1 |