Cheng-Fu Yang

dblp:51/8564 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6916-2142ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Verbalized Representation Learning for Interpretable Few-Shot Generalization
abstract
Humans recognize objects after observing only a few examples, a remarkable capability enabled by their inherent language understanding of the real-world environment. Developing verbalized and interpretable representation can significantly improve model generalization in low-data settings. In this work, we propose Verbalized Representation Learning (VRL), a novel approach for automatically extracting human-interpretable features for object recognition using few-shot data. Our method uniquely captures inter-class differences and intra-class commonalities in the form of natural language by employing a Vision-Language Model (VLM) to identify key discriminative features between different classes and shared characteristics within the same class. These verbalized features are then mapped to numeric vectors through the VLM. The resulting feature vectors can be further utilized to train and infer with downstream classifiers. Experimental results show that, at the same model scale, VRL achieves a 24% absolute improvement over prior state-of-the-art methods while using 95% less data and a smaller mode. Furthermore, compared to human-labeled attributes, the features learned by VRL exhibit a 20% absolute gain when used for downstream classification tasks. Code is available at: https://github.com/joeyy5588/VRL/tree/main.
Cheng-Fu Yang, Da Yin, Wenbo Hu 0006, Heng Ji 0001, Nanyun Peng 0001, Bolei Zhou, Kai-Wei Chang 0001
ICCV1
2024 Re-ReST: Reflection-Reinforced Self-Training for Language Agents
abstract
Finetuning language agents with reasoningaction trajectories is effective, but obtaining these trajectories from human annotations or stronger models is costly and sometimes impractical.In this paper, we investigate the use of self-training in language agents, which can generate supervision from the agent itself, offering a promising alternative without relying on human or stronger model demonstrations.Self-training, however, requires high-quality model-generated samples, which are hard to obtain for challenging language agent tasks.To address this, we present Reflection-Reinforced Self-Training (Re-ReST), which uses a reflector to refine low-quality generated samples during self-training.The reflector takes the agent's output and feedback from an external environment (e.g., unit test results in code generation) to produce improved samples.This technique enhances the quality of inferior samples and efficiently enriches the self-training dataset with higher-quality samples.We conduct extensive experiments on open-source language agents across tasks, including multi-hop question answering, sequential decision-making, code generation, visual question answering, and text-toimage generation.The results demonstrate the effectiveness of self-training and Re-ReST in language agent tasks, with self-training improving baselines by 7.6% on HotpotQA and 28.4% on AlfWorld, and Re-ReST further boosting performance by 2.0% and 14.1%, respectively.Our studies also confirm the efficiency of using a reflector to generate high-quality samples for self-training.Moreover, we demonstrate a method to employ reflection during inference without ground-truth feedback, addressing the limitation of previous reflection work.
Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP2
2023 Target-Free Text-Guided Image Manipulation
abstract
We tackle the problem of target-free text-guided image manipulation, which requires one to modify the input reference image based on the given text instruction, while no ground truth target image is observed during training. To address this challenging task, we propose a Cyclic-Manipulation GAN (cManiGAN) in this paper, which is able to realize where and how to edit the image regions of interest. Specifically, the image editor in cManiGAN learns to identify and complete the input image, while cross-modal interpreter and reasoner are deployed to verify the semantic correctness of the output image based on the input instruction. While the former utilizes factual/counterfactual description learning for authenticating the image semantics, the latter predicts the "undo" instruction and provides pixel-level supervision for the training of cManiGAN. With the above operational cycle-consistency, our cManiGAN can be trained in the above weakly supervised setting. We conduct extensive experiments on the datasets of CLEVR and COCO datasets, and the effectiveness and generalizability of our proposed method can be successfully verified. Project page: sites.google.com/view/wancyuanfan/projects/cmanigan.
Wan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank Wang
AAAI2
2023 LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following
abstract
End-to-end Transformers have demonstrated an impressive success rate for Embodied Instruction Following when the environment has been seen in training.However, they tend to struggle when deployed in an unseen environment.This lack of generalizability is due to the agent's insensitivity to subtle changes in natural language instructions.To mitigate this issue, we propose explicitly aligning the agent's hidden states with the instructions via contrastive learning.Nevertheless, the semantic gap between high-level language instructions and the agent's low-level action space remains an obstacle.Therefore, we further introduce a novel concept of meta-actions to bridge the gap.Meta-actions are ubiquitous action patterns that can be parsed from the original action sequence.These patterns represent higher-level semantics that are intuitively aligned closer to the instructions.When meta-actions are applied as additional training signals, the agent generalizes better to unseen environments.Compared to a strong multi-modal Transformer baseline, we achieve a significant 4.5% absolute gain in success rate in unseen environments of ALFRED Embodied Instruction Following.Additional analysis shows that the contrastive objective and meta-actions are complementary in achieving the best results, and the resulting agent better aligns its states with corresponding instructions, making it more suitable for real-world embodied agents. 1
Cheng-Fu Yang, Yen-Chun Chen 0001, Xiyang Dai, Lu Yuan 0001, Yu-Chiang Frank Wang, Kai-Wei Chang 0001
EMNLP1
2022 Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and Manipulation
abstract
As a key characteristic in audio-visual speech recognition (AVSR), relating linguistic information observed across visual and audio data has been a challenge, benefiting not only audio/visual speech recognition (ASR/VSR) but also for manipulating data within/across modalities. In this paper, we present a feature disentanglement-based framework for jointly addressing the above tasks. By advancing cross-modal mutual learning strategies, our model is able to convert visual or audio-based linguistic features into modality-agnostic representations. Such derived linguistic representations not only allow one to perform ASR, VSR, and AVSR, but also to manipulate audio and visual data output based on the desirable subject identity and linguistic content information. We perform extensive experiments on different recognition and synthesis tasks to show that our model performs favorably against state-of-the-art approaches on each individual task, while ours is a unified solution that is able to jointly tackle the aforementioned audio-visual learning tasks.
Chih-Chun Yang 0004, Wan-Cyuan Fan, Cheng-Fu Yang, Yu-Chiang Frank Wang
AAAI3
2022 Scene Graph Expansion for Semantics-Guided Image Outpainting
abstract
In this paper, we address the task of semantics-guided image outpainting, which is to complete an image by generating semantically practical content. Different from most existing image outpainting works, we approach the above task by understanding and completing image semantics at the scene graph level. In particular, we propose a novel network of Scene Graph Transformer (SGT), which is designed to take node and edge features as inputs for modeling the associated structural information. To better understand and process graph-based inputs, our SGT uniquely performs feature attention at both node and edge levels. While the former views edges as relationship regularization, the latter observes the co-occurrence of nodes for guiding the attention process. We demonstrate that, given a partial input image with its layout and scene graph, our SGT can be applied for scene graph expansion and its conversion to a complete layout. Following state-of-the-art layout-to-image conversions works, the task of image outpainting can be completed with sufficient and practical semantics introduced. Extensive experiments are conducted on the datasets of MS-COCO and Visual Genome, which quantitatively and qualitatively confirm the effectiveness of our proposed SGT and outpainting frameworks.
Chiao-An Yang, Cheng-Yo Tan, Wan-Cyuan Fan, Cheng-Fu Yang, Meng-Lin Wu, Yu-Chiang Frank Wang
CVPR4
2022 Paraphrasing Is All You Need for Novel Object Captioning
abstract
Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequence-to-sequence training or CIDEr optimization. As a result, we present Paraphrasing-to-Captioning (P2C), a two-stage learning framework for NOC, which would heuristically optimize the output captions via paraphrasing. With P2C, the captioning model first learns paraphrasing from a language model pre-trained on text-only corpus, allowing expansion of the word bank for improving linguistic fluency. To further enforce the output caption sufficiently describing the visual content of the input image, we perform self-paraphrasing for the captioning model with fidelity and adequacy objectives introduced. Since no ground truth captions are available for novel object images during training, our P2C leverages cross-modality (image-text) association modules to ensure the above caption characteristics can be properly preserved. In the experiments, we not only show that our P2C achieves state-of-the-art performances on nocaps and COCO Caption datasets, we also verify the effectiveness and flexibility of our learning framework by replacing language and cross-modality association models for NOC. Implementation details and code are available in the supplementary materials.
Cheng-Fu Yang, Yao-Hung Tsai, Wan-Cyuan Fan, Ruslan Salakhutdinov, Louis-Philippe Morency, Frank Wang
NeurIPS1
2021 Multi-modal User Intent Classification Under the Scenario of Smart Factory (Student Abstract)
abstract
Question-answering systems are becoming increasingly popular in Natural Language Processing, especially when applied in smart factory settings. A common practice in designing those systems is through intent classification. However, in a multiple-stage task commonly seen in those settings, relying solely on intent classification may lead to erroneous answers, as questions rising from different work stages may share the same intent but have different contexts and therefore require different answers. To address this problem, we designed an interactive dialogue system that utilizes contextual information to assist intent classification in a multiple-stage task. Specifically, our system incorporates user’s utterances with real-time video feed to better situate users’ questions and analyze their intent.
Yu-Ching Chiu, Bo-Hao Chang, Tzu-Yu Chen, Cheng-Fu Yang, Nanyi Bi, Richard Tzong-Han Tsai, Hung-yi Lee, Yung-Jen Hsu 0001
AAAI4
2021 LayoutTransformer: Scene Layout Generation With Conceptual and Spatial Diversity
abstract
When translating text inputs into layouts or images, existing works typically require explicit descriptions of each object in a scene, including their spatial information or the associated relationships. To better exploit the text input, so that implicit objects or relationships can be properly inferred during layout generation, we propose a LayoutTransformer Network (LT-Net) in this paper. Given a scene-graph input, our LT-Net uniquely encodes the semantic features for exploiting their co-occurrences and implicit relationships. This allows one to manipulate conceptually diverse yet plausible layout outputs. Moreover, the decoder of our LT-Net translates the encoded contextual features into bounding boxes with self-supervised relation consistency preserved. By fitting their distributions to Gaussian mixture models, spatially-diverse layouts can be additionally produced by LT-Net. We conduct extensive experiments on the datasets of MS-COCO and Visual Genome, and confirm the effectiveness and plausibility of our LT-Net over recent layout generation models. Codes will be released at LayoutTransformer.
Cheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank Wang
CVPR1
2015 Light-emitting diodes for visible light communication
abstract
Conventional LEDs always pursue the high brightness for solid-state lighting, however, they always exhibit very low frequency bandwidth of tens MHz. As an application, visible light communication (VLC) is not only for illumination but also for data transmission communication, which is a kind of optical wireless communication that uses the visible white ray as the medium. In this study, we investigate the fabrication and characterization of high-modulation frequency GaN-based light-emitting diodes. The frequency response of LEDs is mainly limited by its diffusion capacitance and resistance, and the injected carriers in the active region of the device. By appropriate device design, gallium-doped ZnO (GZO) film deposited by atomic layer deposition (ALD) is used as the top contact layer with high lateral resistance to self-confine the current injection. In addition, a smaller bonding pad is used to reduce the RC time constant. By the ways, the blue GaN-based LEDs exhibit a 3-dB modulation bandwidth of 225.4 MHz and a light output power of 1.6 mW at the current of 35 mA. Also, the green GaN-based LEDs at 50 mA exhibit the highest 3-dB frequency bandwidth of 463 MHz and a relatively high light output power of 1.6 mW. The 3-dB modulation bandwidth increases with increasing drive current and decreases with increasing temperature. The highest bit rate of 4.4 Gbit/s is achieved at 1 m plastic fiber using nonreturn-to-zero modulation schemes. Such the LEDs can be applied to visible light communication in the future.
Chien-Lan Liao, Yung-Fu Chang, Chong-Lung Ho, Meng-Chyi Wu, Yuan-Tai Hsieh, Chien-Yu Li, Mau-Phon Houng, Cheng-Fu Yang
IWCMC8
2007 Design Small-Size and Wide-Band T-Shaped Patch Antenna on Ceramic Substrate
abstract
Due to the dielectric constant of traditional FR4 substrate is lower (epsivr=4.2); it is not suitable for the fabrication of small size microwave antennas. In order to reduce the size of the fabricated antennas, the ceramics with higher dielectric constant can be used as the substrates of the designed patch antenna. In this paper, the designed patch antenna is printed on aluminum oxide (Al2O3, epsivr=9.8). Due to larger dielectric constant of Al2O3ceramic, the size of the designed patch antenna can be reduced to about 25mm times 32mm. It is found that the patch size and the length of transmission line will influence the return loss (S11) and the resonant frequency of the fabricated patch antennas. The optimum bandwidth of the designed antenna can reach 400MHz/16.3%, the resonance frequency is near 2.45GHz, and the S11is -31dB.
Cheng-Fu Yang, Chang-Yi Hsieh, Chien-Min Cheng
WCNC1