Te-Lin Wu

dblp:166/3298 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
12since 2021 · last 2025
0009-0001-8761-4691ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Contrastive Visual Data Augmentation
abstract
Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-specific knowledge gaps in training also make them prone to confusing visually similar, commonly misrepresented, or low-resource concepts. To help LMMs better align nuanced visual features with language, improving their ability to recognize and reason about novel or rare concepts, we propose a Contrastive visual Data Augmentation (CoDA) strategy. CoDA extracts key contrastive textual and visual features of target concepts against the known concepts they are misrecognized as, and then uses multimodal generative models to produce targeted synthetic data. Automatic filtering of extracted features and augmented images is implemented to guarantee their quality, as verified by human annotators. We show the effectiveness and efficiency of CoDA on low-resource concept and diverse scene recognition datasets including INaturalist and SUN. We additionally collect NovelSpecies, a benchmark dataset consisting of newly discovered animal species that are guaranteed to be unseen by LMMs. LLaVA-1.6 1-shot updating results on these three datasets show CoDA significantly improves SOTA visual data augmentation strategies by 12.3% (NovelSpecies), 5.1% (SUN), and 6.0% (iNat) absolute gains in accuracy.
Yu Zhou 0030, Mohan Tang, Xiaomeng Jin, Te-Lin Wu, Kuan-Hao Huang, Heng Ji 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
ICML5
2024 InSpaceType: Dataset and Benchmark for Reconsidering Cross-Space Type Performance in Indoor Monocular Depth
Cho-Ying Wu, Quankai Gao, Chin-Cheng Hsu, Te-Lin Wu, Jing-Wen Chen, Ulrich Neumann
BMVC4
2024 MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning System
abstract
We present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems.
Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001
ACM Multimedia15
2024 LegalDiscourse: Interpreting When Laws Apply and To Whom
abstract
Alexander Spangher, Zihan Xue, Te-Lin Wu, Mark Hansen, Jonathan May. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Alexander Spangher, Zihan Xue, Te-Lin Wu, Mark Hansen, Jonathan May
NAACL-HLT3
2024 DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
abstract
Data analysis is a crucial analytical process essential for deriving insights from real-world databases. As shown in Figure 1, the need for data analysis typically arises from specific application scenarios, and requires diverse reasoning skills including mathematical reasoning, logical reasoning, and strategic reasoning. Existing work often focus on simple factual retrieval or arithmetic resolutions and thus are insufficient for addressing complex real-world queries. This work aims to propose new resources and benchmarks on this crucial yet challenging and under-explored task. Due to the prohibitively high cost of collecting expert annotations, we use large language models (LLMs) enhanced by code generation to automatically generate high-quality data analysis, which will later be refined by human annotators. We construct the DACO dataset, containing (1) 440 databases (of tabular data) collected from real-world scenarios, (2) ~2k automatically generated query-answer pairs that can serve as weak supervision for model training, and (3) a concentrated but high-quality test set with human refined annotations that serves as our main evaluation benchmark. Experiments show that while LLMs like GPT-4 exhibit promising data analysis capabilities, they are still evaluated as less helpful than human-written analysis on 58.1% cases. Leveraging our weak supervision data, we experiment with various fine-tuning methods, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). Our trained model outperforms existing baselines for table question answering, and RLHF further boosts the helpfulness of generated analysis on 58.5% cases.Data and code are released at https://github.com/shirley-wu/daco.
Xueqing Wu 0001, Jingzhen Sha, Te-Lin Wu, Hanyu Zhou, Mohan Tang, Kai-Wei Chang 0001, Nanyun Peng 0001, Haoran Huang
NeurIPS4
2023 SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR Streams
abstract
Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodriguez, Babak Damavandi, Nanyun Peng, Seungwhan Moon. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodríguez 0001, Babak Damavandi, Nanyun Peng 0001, Seungwhan Moon
ACL (1)1
2023 Learning Action Conditions from Instructional Manuals for Instruction Understanding
abstract
The ability to infer pre-and postconditions of an action is vital for comprehending complex instructions, and is essential for applications such as autonomous instruction-guided agents and assistive AI that supports humans to perform physical tasks.In this work, we propose a task dubbed action condition inference, which extracts mentions of preconditions and postconditions of actions in instructional manuals.We propose a weakly supervised approach utilizing automatically constructed large-scale training instances from online instructions, and curate a densely human-annotated and validated dataset to study how well the current NLP models do on the proposed task.We design two types of models differ by whether contextualized and global information is leveraged, as well as various combinations of heuristics to construct the weak supervisions.Our experiments show a >20% F1-score improvement with considering the entire instruction contexts and a > 6% F1-score benefit with the proposed heuristics.However, the best performing model is still well-behind human performance.1 standalone Heuristics Examples Descriptions Entity-Tracing & Coref.… Slice 500 grams of onions.… … Heat the pan with olive oil.… … Place them in the frying pan.… Precondition 1 Precondition 2The shared entities are pan and onions (linked via co-references to them).
Te-Lin Wu, Caiqi Zhang, Alexander Spangher, Nanyun Peng 0001
ACL (1)1
2023 ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos
abstract
Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Chandra, Marjorie Freedman, Ralph Weischedel, Nanyun Peng. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Te-Lin Wu, Zi-Yi Dou, Nischal Reddy Chandra, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng 0001
EMNLP1
2023 Localizing Active Objects from Egocentric Vision with Symbolic World Knowledge
abstract
The ability to actively ground task instructions from an egocentric view is crucial for AI agents to accomplish tasks or assist humans.One important step towards this goal is to localize and track key active objects that undergo major state change as a consequence of human actions/interactions in the environment (e.g., localizing and tracking the 'sponge' in video from the instruction "Dip the sponge into the bucket.")without being told exactly what/where to ground.While existing works approach this problem from a pure vision perspective, we investigate to which extent the language modality (i.e., task instructions) and their interaction with visual modality can be beneficial.Specifically, we propose to improve phrase grounding models' (Li* et al., 2022) ability in localizing the active objects by: (1) learning the role of objects undergoing change and accurately extracting them from the instructions, (2) leveraging pre-and post-conditions of the objects during actions, and (3) recognizing the objects more robustly with descriptional knowledge.We leverage large language models (LLMs) to extract the aforementioned actionobject knowledge, and design a per-object aggregation masking technique to effectively perform joint inference on object phrases with symbolic knowledge.We evaluate our framework on Ego4D ( Graumanet al., 2022) and Epic-Kitchens (Dunnhofer et al., 2022) datasets.Extensive experiments demonstrate the effectiveness of our proposed framework, which leads to > 54% improvements in all standard metrics on the TREK-150-OPE-Det localization + tracking task, > 7% improvements in all standard metrics on the TREK-150-OPE tracking task, and > 3% improvements in average precision (AP) on the Ego4D SCOD task.
Te-Lin Wu, Yu Zhou 0030, Nanyun Peng 0001
EMNLP1
2022 Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
abstract
Te-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph Weischedel, Nanyun Peng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Te-Lin Wu, Alexander Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng 0001
ACL (1)1
2022 Character-centric Story Visualization via Visual Planning and Token Alignment
abstract
Story visualization advances the traditional text-to-image generation by enabling multiple image generation based on a complete story.This task requires machines to 1) understand long text inputs and 2) produce a globally consistent image sequence that illustrates the contents of the story.A key challenge of consistent story visualization is to preserve characters that are essential in stories.To tackle the challenge, we propose to adapt a recent work that augments Vector-Quantized Variational Autoencoders (VQ-VAE) with a text-tovisual-token (transformer) architecture.Specifically, we modify the text-to-visual-token module with a two-stage framework: 1) character token planning model that predicts the visual tokens for characters only; 2) visual token completion model that generates the remaining visual token sequence, which is sent to VQ-VAE for finalizing image generations.To encourage characters to appear in the images, we further train the two-stage framework with a character-token alignment objective.Extensive experiments and evaluations demonstrate that the proposed method excels at preserving characters and can produce higher quality image sequences compared with the strong baselines.Code can be found in https: //github.com/PlusLabNLP/VP-CSV
Hong Chen 0017, Rujun Han, Te-Lin Wu, Hideki Nakayama, Nanyun Peng 0001
EMNLP3
2021 MELINDA: A Multimodal Dataset for Biomedical Experiment Method Classification
abstract
We introduce a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification. The dataset is collected in a fully automated distant supervision manner, where the labels are obtained from an existing curated database, and the actual contents are extracted from papers associated with each of the records in the database. We benchmark various state-of-the-art NLP and computer vision models, including unimodal models which only take either caption texts or images as inputs, and multimodal models. Extensive experiments and analysis show that multimodal models, despite outperforming unimodal ones, still need improvements especially on a less-supervised way of grounding visual concepts with languages, and better transferability to low resource domains. We release our dataset and the benchmarks to facilitate future research in multimodal learning, especially to motivate targeted improvements for applications in scientific domains.
Te-Lin Wu, Shikhar Singh, Sayan Paul, Gully A. P. C. Burns, Nanyun Peng 0001
AAAI1
2020 Program Guided Agent
Shao-Hua Sun, Te-Lin Wu, Joseph J. Lim
ICLR2
2018 Demo2Vec: Reasoning Object Affordances From Online Videos
abstract
Watching expert demonstrations is an important way for humans and robots to reason about affordances of unseen objects. In this paper, we consider the problem of reasoning object affordances through the feature embedding of demonstration videos. We design the Demo2Vec model which learns to extract embedded vectors of demonstration videos and predicts the interaction region and the action label on a target image of the same object. We introduce the Online Product Review dataset for Affordance (OPRA) by collecting and labeling diverse YouTube product review videos. Our Demo2Vec model outperforms various recurrent neural network baselines on the collected dataset.
Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese, Joseph J. Lim
CVPR2
2017 Feedback Networks
abstract
Urrently, the most successful learning models in computer vision are based on learning successive representations followed by a decision layer. This is usually actualized through feedforward multilayer neural networks, e.g. ConvNets, where each layer forms one of such successive representations. However, an alternative that can achieve the same goal is a feedback based approach in which the representation is formed in an iterative manner based on a feedback received from previous iterations output. We establish that a feedback based approach has several core advantages over feedforward: it enables making early predictions at the query time, its output naturally conforms to a hierarchical structure in the label space (e.g. a taxonomy), and it provides a new basis for Curriculum Learning. We observe that feedback develops a considerably different representation compared to feedforward counterparts, in line with the aforementioned advantages. We provide a general feedback based learning architecture, instantiated using existing RNNs, with the endpoint results on par or better than existing feedforward networks and the addition of the above advantages.
Amir Zamir, Te-Lin Wu, Lin Sun 0004, Bokui Shen, Bertram E. Shi, Jitendra Malik, Silvio Savarese
CVPR2
2016 Emergence of Euclidean geometrical intuitions in hierarchical generative models
Arianna Yuan, Te-Lin Wu, James L. McClelland
CogSci2
2015 A low phase-noise class-C VCO using novel 8-shaped transformer
abstract
A 9 GHz low phase noise class-C voltage controlled oscillator (VCOs) with an 8-shaped transformer configuration using 0.18-μm CMOS technology is presented. By utilizing the 8-shaped transformer, the proposed class-C VCO can be operated at reduced dc power consumption while maintaining circuit performance in terms of phase noise and limiting EMC (Electro-Magnetic Compatibility) issues when the symmetrical structures are considered. Consuming a dc current of 5.5 mA with the supply voltage of 1.8 V, the class-C VCO exhibits a frequency tuning range of 1.4 GHz, a phase noise of -117.4 dBc/Hz at 1 MHz offset frequency away from the 8.94 GHz carrier, and a figure of merit up to 186.4 dBc/Hz.
Ping-Yi Wang, Te-Lin Wu, Ming-Yu Chen 0005, Yun-Chun Shen, Yin-Cheng Chang, Da-Chiang Chang, Shawn S. H. Hsu
ISCAS2