Xin Wang 0061

dblp:10/5630-61 · also Xin Eric Wang · DBLP profile ↗
← Back
67ranked-venue papers
9as first author
44since 2021 · last 2026
0000-0003-2605-5504ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 9 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Interleaved Vision-and-Language Generation via Generative Voken
abstract
The effectiveness of Multimodal Large Language Models (MLLMs) demonstrates a profound capability in multimodal understanding. However, the simultaneous generation of images with coherent texts is still underdeveloped. Addressing this, we introduce a novel interleaved vision-and-language generation method, centered around the concept of "generative vokens". These vokens serve as pivotal elements contributing to coherent image-text outputs. Our method is marked by a unique two-stage training strategy for description-free multimodal generation, which does not necessitate extensive descriptions of images. We integrate classifier-free guidance to enhance the alignment of generated images and texts, ensuring more seamless and contextually relevant multimodal interactions. Our model, MiniGPT-5, exhibits substantial improvement over the baseline models on multimodal generation datasets, including MMDialog and VIST. The human evaluation shows MiniGPT-5 is better than the baseline model on more than 56% cases for multimodal generation, highlighting its efficacy across diverse benchmarks. Project page: https://eric-ai-lab.github.io/minigpt-5.github.io/.
Kaizhi Zheng, Xuehai He, Xin Wang 0061
WACV3
2025 GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
abstract
Graphical User Interface (GUI) action grounding, mapping language instructions to actionable elements on GUI screens, is important for assisting users in interactive tutorials, task automation, accessibility support, etc. Most recent works of GUI action grounding use large GUI datasets to fine-tune Multimodal Large Language Models (MLLMs). However, the fine-tuning data is inherently limited to specific GUI environments, leading to significant performance degradation in novel environments due to the generalization challenges in the GUI domain. Therefore, we argue that GUI action grounding models should be further aligned with novel environments before deployment to optimize their performance. To address this, we first propose GUI-Bee, an MLLM-based autonomous agent, to collect high-quality, environment-specific data through exploration and then continuously fine-tune GUI grounding models with the collected data. To ensure the GUI action grounding models generalize to various screens within the target novel environment after the continuous fine-tuning, we equip GUI-Bee with a novel Q-value-Incentive In-Context Reinforcement Learning (Q-ICRL) algorithm that optimizes exploration efficiency and exploration data quality. In the experiment, we introduce NovelScreenSpot to test how well the data can help align GUI action grounding models to novel environments. Furthermore, we conduct an ablation study to validate the Q-ICRL method in enhancing the efficiency of GUI-Bee.
Handong Zhao, Ruiyi Zhang 0002, Xin Wang 0061, Gang Wu 0013
EMNLP5
2025 Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
abstract
Multimodal large language models (MLLMs) are increasingly deployed in open-ended, realworld environments where inputs are messy, underspecified, and not always trustworthy.Unlike curated benchmarks, these settings frequently involve instructions that reference missing objects or contradictory facts, rely on ambiguous cues, or request infeasible actions.In such cases, success hinges not merely on task execution, but on the model's ability to detect when something is silently wrong.This paper presents a systematic analysis of how current MLLMs handle such underspecified and misspecified scenarios: cases where flaws must be inferred from context rather than explicitly stated.Using a curated diagnostic suite spanning four categories of real-world failure modes, we evaluate nine MLLMs, including o3 and GPT-4o, and find that models often fail to surface hidden issues, even when they possess the necessary perceptual and reasoning skills.Explicit prompting reveals that the underlying capabilities exist but are frequently suppressed in favor of user compliance.We further show that simple inference-time interventions, such as cautious persona prompting and, in particular, requiring a clarifying question, can substantially recover performance.Our findings highlight a persistent gap between reasoning competence and behavioral compliance in current MLLMs, and suggest practical strategies for making these systems more trustworthy in underconstrained environments.
Qianqi Yan, Hongquan Li, Xinze Guan, Ching-Chen Kuo, Xin Wang 0061
EMNLP7
2025 SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
abstract
Large Reasoning Models (LRMs) introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks.However, they pose great safety risks against harmful queries and adversarial attacks.While recent mainstream safety efforts on LRMs, supervised fine-tuning (SFT), improve safety performance, we find that SFT-aligned models struggle to generalize to unseen jailbreak prompts.After thorough investigation of LRMs' generation, we identify a safety aha moment that can activate safety reasoning and lead to a safe response.This aha moment typically appears in the 'key sentence', which follows models' query understanding process and can indicate whether the model will proceed safely.Based on these insights, we propose SafeKey, including two complementary objectives to better activate the safety aha moment in the key sentence: (1) a Dual-Path Safety Head to enhance the safety signal in the model's internal representations before the key sentence, and (2) a Query-Mask Modeling objective to improve the models' attention on its query understanding, which has important safety hints.Experiments across multiple safety benchmarks demonstrate that our methods significantly improve safety generalization to a wide range of jailbreak attacks and out-of-distribution harmful prompts, lowering the average harmfulness rate by 9.6%, while maintaining general abilities.Our analysis reveals how SafeKey enhances safety by reshaping internal attention and improving the quality of hidden representations.
Kaiwen Zhou 0002, Xuandong Zhao, Jayanth Srinivasa, Gaowen Liu, Aosong Feng, Dawn Song, Xin Wang 0061
EMNLP7
2025 VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
abstract
Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments.
Shijie Zhou 0003, Alexander Vilesov, Xuehai He, Ziyu Wan, Shuwang Zhang, Aditya Nagachandra, Di Chang, Xin Wang 0061, Achuta Kadambi
ICCV9
2025 Agent S: An Open Agentic Framework that Uses Computers Like a Human
abstract
We present Agent S, an open agentic framework that enables autonomous interaction with computers through Graphical User Interface (GUI), aimed at transforming human-computer interaction by automating complex, multi-step tasks. Agent S addresses three key challenges in automating computer tasks: acquiring domain-specific knowledge, planning over long task horizons, and handling dynamic, non-uniform interfaces. To this end, Agent S introduces experience-augmented hierarchical planning, which learns from external knowledge search and internal experience retrieval at multiple levels, facilitating efficient task planning and subtask execution. In addition, it employs an Agent-Computer Interface (ACI) to better elicit the reasoning and control capabilities of GUI agents based on Multimodal Large Language Models (MLLMs). Evaluation on the OSWorld benchmark shows that Agent S outperforms the baseline by 9.37\% on success rate (an 83.6\% relative improvement) and achieves a new state-of-the-art. Comprehensive analysis highlights the effectiveness of individual components and provides insights for future improvements. Furthermore, Agent S demonstrates broad generalizability to different operating systems on a newly-released WindowsAgentArena benchmark. Code available at https://github.com/simular-ai/Agent-S.
Saaket Agashe, Jiuzhou Han, Shuyu Gan, Ang Li 0006, Xin Wang 0061
ICLR6
2025 MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
abstract
Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models"---interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causalities. To this end, we introduce MMWorld, a new benchmark for multi-discipline, multi-faceted multimodal video understanding. MMWorld distinguishes itself from previous video understanding benchmarks with two unique advantages: (1) multi-discipline, covering various disciplines that often require domain expertise for comprehensive understanding; (2) multi-faceted reasoning, including explanation, counterfactual thinking, future prediction, etc. MMWorld consists of a human-annotated dataset to evaluate MLLMs with questions about the whole videos and a synthetic dataset to analyze MLLMs within a single modality of perception. Together, MMWorld encompasses 1,910 videos across seven broad disciplines and 69 subdisciplines, complete with 6,627 question-answer pairs and associated captions. The evaluation includes 4 proprietary and 11 open-source MLLMs, which struggle on MMWorld (e.g., GPT-4o performs the best with only 62.5% accuracy), showing large room for improvement. Further ablation studies reveal other interesting findings such as models' different skill sets from humans. We hope MMWorld can serve as an essential step towards world model evaluation in videos.
Xuehai He, Weixi Feng, Kaizhi Zheng, Wanrong Zhu, Zhengyuan Yang, William Yang Wang, Xin Wang 0061
ICLR14
2025 EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
abstract
Given the steep learning curve of professional 3D software and the time- consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augmented reality, and gaming. However, recent approaches to language-guided 3D scene editing either require manual interventions or focus only on appearance modifications without supporting comprehensive scene layout changes. In response, we propose EditRoom, a unified framework capable of executing a variety of layout edits through natural language commands, without requiring manual intervention. Specifically, EditRoom leverages Large Language Models (LLMs) for command planning and generates target scenes using a diffusion-based method, enabling six types of edits: rotate, translate, scale, replace, add, and remove. To address the lack of data for language-guided 3D scene editing, we have developed an automatic pipeline to augment existing 3D scene synthesis datasets and introduced EditRoom-DB, a large-scale dataset with 83k editing pairs, for training and evaluation. Our experiments demonstrate that our approach consistently outperforms other baselines across all metrics, indicating higher accuracy and coherence in language-guided scene layout editing.
Kaizhi Zheng, Xuehai He, Zhengyuan Yang, Xin Wang 0061
ICLR10
2025 Multimodal Situational Safety
abstract
Multimodal Large Language Models (MLLMs) are rapidly evolving, demonstrating impressive capabilities as multimodal assistants that interact with both humans and their environments. However, this increased sophistication introduces significant safety concerns. In this paper, we present the first evaluation and analysis of a novel safety challenge termed Multimodal Situational Safety, which explores how safety considerations vary based on the specific situation in which the user or agent is engaged. We argue that for an MLLM to respond safely—whether through language or action—it often needs to assess the safety implications of a language query within its corresponding visual context. To evaluate this capability, we develop the Multimodal Situational Safety benchmark (MSSBench) to assess the situational safety performance of current MLLMs. The dataset comprises 1,960 language query-image pairs, half of which the image context is safe, and the other half is unsafe. We also develop an evaluation framework that analyzes key safety aspects, including explicit safety reasoning, visual understanding, and, crucially, situational safety reasoning. Our findings reveal that current MLLMs struggle with this nuanced safety problem in the instruction-following setting and struggle to tackle these situational safety challenges all at once, highlighting a key area for future research. Furthermore, we develop multi-agent pipelines to coordinately solve safety challenges, which shows consistent improvement in safety over the original MLLM response.
Kaiwen Zhou 0002, Xuandong Zhao, Anderson Compalas, Dawn Song, Xin Wang 0061
ICLR6
2025 JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
abstract
Building a conversational embodied agent to execute real-life tasks has been a long-standing yet quite challenging research goal, as it requires effective human-agent communication, multi-modal understanding, long-range sequential decision making, etc. Traditional symbolic methods have scaling and generalization issues, while end-to-end deep learning models suffer from data scarcity and high task complexity, and are often hard to explain. To benefit from both worlds, we propose JARVIS, a neuro-symbolic commonsense reasoning framework for modular, generalizable, and interpretable conversational embodied agents. First, it acquires symbolic representations by prompting large language models (LLMs) for language understanding and sub-goal planning, and by constructing semantic maps from visual observations. Then the symbolic module reasons for sub-goal planning and action generation based on task- and action-level common sense. Extensive experiments on the TEACh dataset validate the efficacy and efficiency of our JARVIS framework, which achieves state-of-the-art (SOTA) results on all three dialog-based embodied tasks, including Execution from Dialog History (EDH), Trajectory from Dialog (TfD), and Two-Agent Task Completion (TATC) (e.g., our method boosts the unseen Success Rate on EDH from 6.1% to 15.8%). Moreover, we systematically analyze the essential factors that affect the task performance and also demonstrate the superiority of our method in few-shot settings.
Kaizhi Zheng, Kaiwen Zhou 0002, Zonglin Di, Xuehai He, Xin Wang 0061
NeSy8
2025 GRIT: Teaching MLLMs to Think with Images
abstract
Recent studies have demonstrated the efficacy of using Reinforcement Learning (RL) in building reasoning models that articulate chains of thoughts prior to producing final answers. However, despite ongoing advances that aim at enabling reasoning for vision-language tasks, existing open-source visual reasoning models typically generate reasoning content with pure natural language, lacking explicit integration of visual information. This limits their ability to produce clearly articulated and visually grounded reasoning chains. To this end, we propose Grounded Reasoning with Images and Texts (GRIT), a novel method for training MLLMs to think with images. GRIT introduces a grounded reasoning paradigm, in which models generate reasoning chains that interleave natural language and explicit bounding box coordinates. These coordinates point to regions of the input image that the model consults during its reasoning process. Additionally, GRIT is equipped with a reinforcement learning approach, GRPO-GR, built upon the GRPO algorithm. GRPO-GR employs robust rewards focused on the final answer accuracy and format of the grounded reasoning output, which eliminates the need for data with reasoning chain annotations or explicit bounding box labels. As a result, GRIT achieves exceptional data efficiency, requiring as few as 20 image-question-answer triplets from existing datasets. Comprehensive evaluations demonstrate that GRIT effectively trains MLLMs to produce coherent and visually grounded reasoning chains, showing a successful unification of reasoning and grounding abilities. All code, data, and checkpoints will be released.
Xuehai He, Diji Yang, Kaizhi Zheng, Ching-Chen Kuo, Xinze Guan, Xin Wang 0061
NeurIPS8
2025 More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
abstract
Test-time compute has empowered multimodal large language models to generate extended reasoning chains, yielding strong performance on tasks such as multimodal math reasoning. However, we observe that this improved reasoning ability often comes with increased hallucination: as generations become longer, models tend to drift away from image-grounded content and rely more on language priors. Attention analysis reveals that longer reasoning chains reduce focus on visual inputs, contributing to hallucination. To systematically study this phenomenon, we introduce RH-AUC, a metric that quantifies how a model's perception accuracy changes with reasoning length, enabling evaluation of whether the model preserves visual grounding while reasoning. We also release RH-Bench, a diagnostic benchmark covering diverse multimodal tasks, designed to jointly assess the balance of reasoning ability and hallucination. We find that (i) larger models generally exhibit a better balance between reasoning and perception; (ii) reasoning and perception balance depends more on the types and domains of the training data than its volume. Our findings highlight the need for evaluation frameworks that account for both reasoning quality and perceptual reliability.
Zhongxing Xu, Qingyue Wei, Juncheng Wu, James Zou 0001, Xin Wang 0061, Yuyin Zhou
NeurIPS6
2025 Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
abstract
Human cognition typically involves thinking through abstract, fluid concepts rather than strictly using discrete linguistic tokens. Current Large Language Models (LLMs), however, are constrained to reasoning within the boundaries of human language, processing discrete token embeddings that represent fixed points in semantic space. This discrete constraint restricts the expressive power and upper potential of such reasoning models, often causing incomplete exploration of reasoning paths, as standard Chain-of-Thought (CoT) methods rely on sampling one token per step. In this work, we introduce Soft Thinking, a training-free method that emulates human-like ``soft'' reasoning by generating abstract concept tokens in a continuous concept space. These concept tokens are created by the probability-weighted mixture of token embeddings, which span the continuous concept space, enabling smooth transitions and richer representations that transcend traditional discrete boundaries. In essence, each generated concept token encapsulates multiple meanings from related discrete tokens, implicitly exploring various reasoning paths to converge effectively toward the correct answer. Empirical evaluations on diverse mathematical and coding benchmarks consistently demonstrate the effectiveness and efficiency of Soft Thinking, improving pass@1 accuracy by up to 2.48 points while simultaneously reducing token usage by up to 22.4\% compared to standard CoT. Qualitative analysis further reveals that Soft Thinking outputs remain highly interpretable and readable, highlighting the potential of Soft Thinking to break the inherent limits of discrete language-based reasoning.
Zhen Zhang 0008, Xuehai He, Weixiang Yan, Xin Wang 0061
NeurIPS6
2024 Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
abstract
Yue Fan, Jing Gu, Kaiwen Zhou, Qianqi Yan, Shan Jiang, Ching-Chen Kuo, Yang Zhao, Xinze Guan, Xin Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Kaiwen Zhou 0002, Qianqi Yan, Ching-Chen Kuo, Xinze Guan, Xin Wang 0061
ACL (1)9
2024 Topology-aware Retrieval Augmentation for Text Generation
Yu Wang 0160, Nedim Lipka, Ruiyi Zhang 0002, Alexa F. Siu, Yuying Zhao, Bo Ni, Xin Wang 0061, Ryan Rossi, Tyler Derr
CIKM7
2024 SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing
Nanxuan Zhao, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Yilin Wang 0002, Xin Wang 0061
ECCV (32)10
2024 NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Gengze Zhou, Yicong Hong, Zun Wang 0001, Xin Wang 0061, Qi Wu 0001
ECCV (7)4
2024 ComCLIP: Training-Free Compositional Image and Text Matching
abstract
Kenan Jiang, Xuehai He, Ruize Xu, Xin Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kenan Jiang, Xuehai He, Ruize Xu, Xin Wang 0061
NAACL-HLT4
2024 WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking
abstract
While deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that without a sound model evaluation framework, the AI community's efforts cannot reach their full potential, thereby slowing the progress and transfer of innovation into real-world drug discovery.Thus, in this paper, we seek to establish a new gold standard for small molecule drug discovery benchmarking, WelQrate. Specifically, our contributions are threefold: WelQrate dataset collection - we introduce a meticulously curated collection of 9 datasets spanning 5 therapeutic target classes. Our hierarchical curation pipelines, designed by drug discovery experts, go beyond the primary high-throughput screen by leveraging additional confirmatory and counter screens along with rigorous domain-driven preprocessing, such as Pan-Assay Interference Compounds (PAINS) filtering, to ensure the high-quality data in the datasets; WelQrate Evaluation Framework - we propose a standardized model evaluation framework considering high-quality datasets, featurization, 3D conformation generation, evaluation metrics, and data splits, which provides a reliable benchmarking for drug discovery experts conducting real-world virtual screening; Benchmarking - we evaluate model performance through various research questions using the WelQrate dataset collection, exploring the effects of different models, dataset quality, featurization methods, and data splitting strategies on the results.In summary, we recommend adopting our proposed WelQrate as the gold standard in small molecule drug discovery benchmarking. The WelQrate dataset collection, along with the curation codes, and experimental scripts are all publicly available at www.WelQrate.org.
Yunchao Liu 0001, Ha Dong, Xin Wang 0061, Rocco Moretti, Yu Wang 0160, Zhaoqian Su, Jiawei Gu, Bobby Bodenheimer, Charles David Weaver, Jens Meiler, Tyler Derr
NeurIPS3
2023 Parameter-Efficient Model Adaptation for Vision Transformers
abstract
In computer vision, it has achieved great transfer learning performance via adapting large-scale pretrained vision models (e.g., vision transformers) to downstream tasks. Common approaches for model adaptation either update all model parameters or leverage linear probes. In this paper, we aim to study parameter-efficient model adaptation strategies for vision transformers on the image classification task. We formulate efficient model adaptation as a subspace training problem and perform a comprehensive benchmarking over different efficient adaptation methods. We conduct an empirical study on each efficient model adaptation method focusing on its performance alongside parameter cost. Furthermore, we propose a parameter-efficient model adaptation framework, which first selects submodules by measuring local intrinsic dimensions and then projects them into subspace for further decomposition via a novel Kronecker Adaptation method. We analyze and compare our method with a diverse set of baseline model adaptation methods (including state-of-the-art methods for pretrained language models). Our method performs the best in terms of the tradeoff between accuracy and parameter efficiency across 20 datasets under the few-shot setting and 7 image classification datasets under the full-shot setting.
Xuehai He, Chunyuan Li, Pengchuan Zhang, Xin Wang 0061
AAAI5
2023 Multimodal Graph Transformer for Multimodal Question Answering
abstract
Despite the success of Transformer models in vision and language tasks, they often learn knowledge from enormous data implicitly and cannot utilize structured input data directly.On the other hand, structured learning approaches such as graph neural networks (GNNs) that integrate prior information can barely compete with Transformer models.In this work, we aim to benefit from both worlds and propose a novel Multimodal Graph Transformer for question answering tasks that requires performing reasoning across multiple modalities.We introduce a graph-involved plug-and-play quasi-attention mechanism to incorporate multimodal graph information, acquired from text and visual data, to the vanilla self-attention as effective prior.In particular, we construct the text graph, dense region graph, and semantic graph to generate adjacency matrices, and then compose them with input vision and language features to perform downstream reasoning.Such a way of regularizing self-attention with graph information significantly improves the inferring ability and helps align features from different modalities.We validate the effectiveness of Multimodal Graph Transformer over its Transformer baselines on GQA, VQAv2, and MultiModalQA datasets.
Xuehai He, Xin Wang 0061
EACL2
2023 Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation
abstract
The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users.Despite the advancements of T2I models, a common issue encountered by users is the need for repetitive editing of input prompts in order to receive a satisfactory image, which is time-consuming and labor-intensive.Given the demonstrated text generation power of largescale language models, such as GPT-k, we investigate the potential of utilizing such models to improve the prompt editing process for T2I generation.We conduct a series of experiments to compare the common edits made by humans and GPT-k, evaluate the performance of GPT-k in prompting T2I, and examine factors that may influence this process.We found that GPT-k models focus more on inserting modifiers while humans tend to replace words and phrases, which includes changes to the subject matter.Experimental results show that GPT-k are more effective in adjusting modifiers rather than predicting spontaneous changes in the primary subject matters.Adopting the edit suggested by GPT-k models may reduce the percentage of remaining edits by 20-30%. 1 Our experiments are conducted upon StableDiffusion since it is a wide-adopted open-source large text-to-image generative model with SoTA performance.
Wanrong Zhu, Xinyi Wang 0003, Tsu-Jui Fu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
EMNLP5
2023 Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Pradyumna Narayana, Sugato Basu, Xin Wang 0061, William Yang Wang
ICLR8
2023 Neuro-Symbolic Procedural Planning with Commonsense Prompting
Weixi Feng, Wanrong Zhu, Wenda Xu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
ICLR5
2023 ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation
abstract
The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, which generalizes poorly to novel objects in unknown environments. In this work, we present a novel zero-shot object navigation method, Exploration with Soft Commonsense constraints (ESC), that transfers commonsense knowledge in pre-trained models to open-world object navigation without any navigation experience nor any other training on the visual environments. First, ESC leverages a pre-trained vision and language model for open-world prompt-based grounding and a pre-trained commonsense language model for room and object reasoning. Then ESC converts commonsense knowledge into navigation actions by modeling it as soft logic predicates for efficient exploration. Extensive experiments on MP3D, HM3D, and RoboTHOR benchmarks show that our ESC method improves significantly over baselines, and achieves new state-of-the-art results for zero-shot object navigation (e.g., 288% relative Success Rate improvement than CoW on MP3D).
Kaiwen Zhou 0002, Kaizhi Zheng, Connor Pryor, Yilin Shen, Hongxia Jin, Lise Getoor, Xin Wang 0061
ICML7
2023 LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
abstract
Attaining a high degree of user controllability in visual generation often requires intricate, fine-grained inputs like layouts. However, such inputs impose a substantial burden on users when compared to simple text inputs. To address the issue, we study how Large Language Models (LLMs) can serve as visual planners by generating layouts from text conditions, and thus collaborate with visual generative models. We propose LayoutGPT, a method to compose in-context visual demonstrations in style sheet language to enhance visual planning skills of LLMs. We show that LayoutGPT can generate plausible layouts in multiple domains, ranging from 2D images to 3D indoor scenes. LayoutGPT also shows superior performance in converting challenging language concepts like numerical and spatial relations to layout arrangements for faithful text-to-image generation. When combined with a downstream image generation model, LayoutGPT outperforms text-to-image models/systems by 20-40\% and achieves comparable performance as human users in designing visual layouts for numerical and spatial correctness. Lastly, LayoutGPT achieves comparable performance to supervised methods in 3D indoor scene synthesis, demonstrating its effectiveness and potential in multiple visual domains.
Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Xuehai He, Sugato Basu, Xin Wang 0061, William Yang Wang
NeurIPS8
2023 PHOTOSWAP: Personalized Subject Swapping in Images
abstract
In an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity. Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the original charm and composition of the image. We present \emph{Photoswap}, a novel approach that enables this immersive image editing experience through personalized subject swapping in existing images. \emph{Photoswap} first learns the visual concept of the subject from reference images and then swaps it into the target image using pre-trained diffusion models in a training-free manner. We establish that a well-conceptualized visual subject can be seamlessly transferred to any image with appropriate self-attention and cross-attention manipulation, maintaining the pose of the swapped subject and the overall coherence of the image. Comprehensive experiments underscore the efficacy and controllability of \emph{Photoswap} in personalized subject swapping. Furthermore, \emph{Photoswap} significantly outperforms baseline methods in human ratings across subject swapping, background preservation, and overall quality, revealing its vast application potential, from entertainment to professional editing.
Yilin Wang 0002, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Xin Wang 0061
NeurIPS11
2023 LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation
abstract
Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose LLMScore, a new framework that offers evaluation scores with multi-granularity compositionality. LLMScore leverages the large language models (LLMs) to evaluate text-to-image models. Initially, it transforms the image into image-level and object-level visual descriptions. Then an evaluation instruction is fed into the LLMs to measure the alignment between the synthesized image and the text, ultimately generating a score accompanied by a rationale. Our substantial analysis reveals the highest correlation of LLMScore with human judgments on a wide range of datasets (Attribute Binding Contrast, Concept Conjunction, MSCOCO, DrawBench, PaintSkills). Notably, our LLMScore achieves Kendall's tau correlation with human evaluations that is 58.8% and 31.2% higher than the commonly-used text-image matching metrics CLIP and BLIP, respectively.
Xianjun Yang, Xiujun Li, Xin Wang 0061, William Yang Wang
NeurIPS4
2023 CUDA-GHR: Controllable Unsupervised Domain Adaptation for Gaze and Head Redirection
abstract
The robustness of gaze and head pose estimation models is highly dependent on the amount of labeled data. Recently, generative modeling has shown excellent results in generating photo-realistic images, which can alleviate the need for annotations. However, adopting such generative models to new domains while maintaining their ability to provide fine-grained control over different image attributes, e.g., gaze and head pose directions, has been a challenging problem. This paper proposes CUDA-GHR, an unsupervised domain adaptation framework that enables fine-grained control over gaze and head pose directions while preserving the appearance-related factors of the person. Our framework simultaneously learns to adapt to new domains and disentangle visual attributes such as appearance, gaze direction, and head orientation by utilizing a label-rich source domain and an unlabeled target domain. Extensive experiments on the benchmarking datasets show that the proposed method can outperform state-of-the-art techniques on both quantitative and qualitative evaluations. Furthermore, we demonstrate the effectiveness of generated image-label pairs in the target domain for pretraining networks for the downstream task of gaze and head pose estimation. The source code and pre-trained models are available at https://github.com/jswati31/cuda-ghr.
Swati Jindal, Xin Wang 0061
WACV2
2022 Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
abstract
A long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks.Visionand-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives increasing attention from natural language processing, computer vision, robotics, and machine learning communities.In this paper, we review contemporary studies in the emerging field of VLN, covering tasks, evaluation metrics, methods, etc.Through structured analysis of current progress and challenges, we highlight the limitations of current VLN and opportunities for future work.This paper serves as a thorough reference for the VLN research community.1
Eliana Stefani, Qi Wu 0001, Jesse Thomason, Xin Wang 0061
ACL (1)5
2022 Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning
abstract
Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing attention in recent years. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, current temporal grounding datasets do not specifically test for the compositional generalizability. To systematically measure the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. Evaluating the state-of-the-art methods on our new dataset splits, we empirically find that they fail to generalize to queries with novel combinations of seen words. To tackle this challenge, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into multiple structured hierarchies and learns fine-grained semantic correspondence among them. Experiments illustrate the superior compositional generalizability of our approach. The repository of this work is at ht tps: / / gi thub. com/YYJMJC/ Composi tional- Temporal-Grounding.
Juncheng Li 0006, Junlin Xie, Linchao Zhu, Siliang Tang, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang, Xin Wang 0061
CVPR9
2022 M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers
abstract
Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language instructions to perform automatic editing would significantly improve accessibility. This paper introduces the language-based video editing (LBVE) task, which allows the model to edit, guided by text instruction, a source video into a target video. LBVE contains two features: 1) the scenario of the source video is preserved instead of generating a completely different video; 2) the semantic is presented differently in the target video, and all changes are controlled by the given instruction. We propose a Multi-Modal Multi-Level Transformer (M3L) to carry out LBVE. M3L dynamically learns the correspondence between video perception and language semantic at different levels, which benefits both the video understanding and video frame synthesis. We build three new datasets for evaluation, including two diagnostic and one from natural videos with human-labeled text. Extensive experimental results show that M3L is effective for video editing and that LBVE can lead to a new field toward vision-and-language research.
Tsu-Jui Fu, Xin Wang 0061, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
CVPR2
2022 Language-Driven Artistic Style Transfer
Tsu-Jui Fu, Xin Wang 0061, William Yang Wang
ECCV (36)2
2022 FedVLN: Privacy-Preserving Federated Vision-and-Language Navigation
Kaiwen Zhou 0002, Xin Wang 0061
ECCV (36)2
2022 CPL: Counterfactual Prompt Learning for Vision and Language Models
abstract
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun R. Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang 0061
EMNLP10
2022 Understanding Instance-Level Impact of Fairness Constraints
abstract
A variety of fairness constraints have been proposed in the literature to mitigate group-level statistical bias. Their impacts have been largely evaluated for different groups of populations corresponding to a set of sensitive attributes, such as race or gender. Nonetheless, the community has not observed sufficient explorations for how imposing fairness constraints fare at an instance level. Building on the concept of influence function, a measure that characterizes the impact of a training example on the target model and its predictive performance, this work studies the influence of training examples when fairness constraints are imposed. We find out that under certain assumptions, the influence function with respect to fairness constraints can be decomposed into a kernelized combination of training examples. One promising application of the proposed fairness influence function is to identify suspicious training examples that may cause model discrimination by ranking their influence scores. We demonstrate with extensive experiments that training on a subset of weighty data examples leads to lower fairness violations with a trade-off of accuracy.
Xin Wang 0061, Yang Liu 0018
ICML2
2022 Imagination-Augmented Natural Language Understanding
abstract
Yujie Lu, Wanrong Zhu, Xin Wang, Miguel Eckstein, William Yang Wang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Wanrong Zhu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
NAACL-HLT3
2022 Diagnosing Vision-and-Language Navigation: What Really Matters
abstract
Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Wang, Qi Wu, Miguel Eckstein, William Yang Wang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Wang 0061, Qi Wu 0001, Miguel P. Eckstein, William Yang Wang
NAACL-HLT6
2022 VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation
abstract
Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill the blank of the last mile of embodied agents---object manipulation by following human guidance, e.g., “move the red mug next to the box while keeping it upright.” To this end, we introduce an Automatic Manipulation Solver (AMSolver) system and build a Vision-and-Language Manipulation benchmark (VLMbench) based on it, containing various language instructions on categorized robotic manipulation tasks. Specifically, modular rule-based task templates are created to automatically generate robot demonstrations with language instructions, consisting of diverse object shapes and appearances, action types, and motion constraints. We also develop a keypoint-based model 6D-CLIPort to deal with multi-view observations and language input and output a sequence of 6 degrees of freedom (DoF) actions. We hope the new simulator and benchmark will facilitate future research on language-guided robotic manipulation.
Kaizhi Zheng, Odest Chadwicke Jenkins, Xin Wang 0061
NeurIPS4
2021 L2C: Describing Visual Differences Needs Semantic Understanding of Individuals
abstract
Recent advances in language and vision push forward the research of captioning a single image to describing visual differences between image pairs.Suppose there are two images, I 1 and I 2 , and the task is to generate a description W 1,2 comparing them, existing methods directly model ⟨I 1 , I 2 ⟩ → W 1,2 mapping without the semantic understanding of individuals.In this paper, we introduce a Learningto-Compare (L2C) model, which learns to understand the semantic structures of these two images and compare them while learning to describe each one.We demonstrate that L2C benefits from a comparison between explicit semantic representations and singleimage captions, and generalizes better on the new testing image pairs.It outperforms the baseline on both automatic evaluation and human evaluation for the Birds-to-Words dataset.
An Yan 0003, Xin Wang 0061, Tsu-Jui Fu, William Yang Wang
EACL2
2021 Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
abstract
Wanrong Zhu, Xin Wang, Tsu-Jui Fu, An Yan, Pradyumna Narayana, Kazoo Sone, Sugato Basu, William Yang Wang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Wanrong Zhu, Xin Wang 0061, Tsu-Jui Fu, An Yan 0003, Pradyumna Narayana, Kazoo Sone, Sugato Basu, William Yang Wang
EACL2
2021 Are Gender-Neutral Queries Really Gender-Neutral? Mitigating Gender Bias in Image Search
abstract
Internet search affects people's cognition of the world, so mitigating biases in search results and learning fair models is imperative for social good.We study a unique gender bias in image search in this work: the search images are often gender-imbalanced for genderneutral natural language queries.We diagnose two typical image search models, the specialized model trained on in-domain datasets and the generalized representation model pretrained on massive image and text data across the internet.Both models suffer from severe gender bias.Therefore, we introduce two novel debiasing approaches: an in-processing fair sampling method to address the gender imbalance issue for training models, and a postprocessing feature clipping method base on mutual information to debias multimodal representations of pre-trained models.Extensive experiments on MS-COCO (Lin et al., 2014) and Flickr30K (Young et al., 2014) benchmarks show that our methods significantly reduce the gender bias in image search models.
Yang Liu 0018, Xin Wang 0061
EMNLP (1)3
2021 Visual Question Rewriting for Increasing Response Rate
abstract
When a human asks questions online, or when a conversational virtual agent asks a human questions, questions triggering emotions or with details might more likely to get responses or answers. we explore how to automatically rewrite natural language questions to improve the response rate form people. In particular, a new task of Visual Question Rewriting (VQR) task is introduced to explore how visual information can be used to improve the new question(s). A data set containing -4K bland&attractive question-images triples is collected. We developed some baseline sequence to sequence models and more advanced transformer-based models, which take a bland question and a related image as input, and output a rewritten question that's expected to be more attractive. Offline experiments and mechanical Turk based evaluations show that it's possible to rewrite bland questions in a more detailed and attractive way to increase response rate, and images can be helpful.
Jiayi Wei, Xilian Li, Yi Zhang 0001, Xin Wang 0061
SIGIR4
2021 Vision-Language Navigation Policy Learning and Adaptation
abstract
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms baseline methods by 10 percent on Success Rate weighted by Path Length (SPL) and achieves the state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore and adapt to unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7 to 11.7 percent).
Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Generative Adversarial Zero-Shot Relational Learning for Knowledge Graphs
abstract
Large-scale knowledge graphs (KGs) are shown to become more important in current information systems. To expand the coverage of KGs, previous studies on knowledge graph completion need to collect adequate training instances for newly-added relations. In this paper, we consider a novel formulation, zero-shot learning, to free this cumbersome curation. For newly-added relations, we attempt to learn their semantic features from their text descriptions and hence recognize the facts of unseen relations with no examples being seen. For this purpose, we leverage Generative Adversarial Networks (GANs) to establish the connection between text and knowledge graph domain: The generator learns to generate the reasonable relation embeddings merely with noisy text descriptions. Under this setting, zero-shot learning is naturally converted to a traditional supervised classification task. Empirically, our method is model-agnostic that could be potentially applied to any version of KG embeddings, and consistently yields performance improvements on NELL and Wiki dataset.
Pengda Qin, Xin Wang 0061, Wenhu Chen, Chunyun Zhang, Weiran Xu, William Yang Wang
AAAI2
2020 Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation
abstract
Visual navigation is a task of training an embodied agent by intelligently navigating to a target object (e.g., television) using only visual observations. A key challenge for current deep reinforcement learning models lies in the requirements for a large amount of training data. It is exceedingly expensive to construct sufficient 3D synthetic environments annotated with the target object information. In this paper, we focus on visual navigation in the low-resource setting, where we have only a few training environments annotated with object information. We propose a novel unsupervised reinforcement learning approach to learn transferable meta-skills (e.g., bypass obstacles, go straight) from unannotated environments without any supervisory signals. The agent can then fast adapt to visual navigation through learning a high-level master policy to combine these meta-skills, when the visual-navigation-specified reward is provided. Experimental results show that our method significantly outperforms the baseline by 53.34% relatively on SPL, and further qualitative analysis demonstrates that our method learns transferable motor primitives for visual navigation.
Juncheng Li 0006, Xin Wang 0061, Siliang Tang, Haizhou Shi, Fei Wu 0001, Yueting Zhuang, William Yang Wang
CVPR2
2020 REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments
abstract
One of the long-term challenges of robotics is to enable robots to interact with humans in the visual world via natural language, as humans are visual animals that communicate through language. Overcoming this challenge requires the ability to perform a wide variety of complex tasks in response to multifarious instructions from humans. In the hope that it might drive progress towards more flexible and powerful human interactions with robots, we propose a dataset of varied and complex robot tasks, described in natural language, in terms of objects visible in a large set of real images. Given an instruction, success requires navigating through a previously-unseen environment to identify an object. This represents a practical challenge, but one that closely reflects one of the core visual problems in robotics. Several state-of-the-art vision-and-language navigation, and referring-expression models are tested to verify the difficulty of this new task, but none of them show promising results because there are many fundamental differences between our task and previous ones. A novel Interactive Navigator-Pointer model is also proposed that provides a strong baseline on the task. The proposed model especially achieves the best performance on the unseen test split, but still leaves substantial room for improvement compared to the human performance. Repository: https://github.com/YuankaiQi/REVERIE.
Yuankai Qi, Qi Wu 0001, Peter Anderson 0001, Xin Wang 0061, William Yang Wang, Chunhua Shen, Anton van den Hengel
CVPR4
2020 Counterfactual Vision-and-Language Navigation via Adversarial Path Sampler
Tsu-Jui Fu, Xin Wang 0061, Matthew F. Peterson, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
ECCV (6)2
2020 Environment-Agnostic Multitask Learning for Natural Language Grounded Navigation
Xin Wang 0061, Vihan Jain, Eugene Ie, William Yang Wang, Zornitsa Kozareva, Sujith Ravi
ECCV (24)1
2020 SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
abstract
Iterative Language-Based Image Editing (IL-BIE) tasks follow iterative instructions to edit images step by step.Data scarcity is a significant issue for ILBIE as it is challenging to collect large-scale examples of images before and after instruction-based changes.However, humans still accomplish these editing tasks even when presented with an unfamiliar image-instruction pair.Such ability results from counterfactual thinking and the ability to think about alternatives to events that have happened already.In this paper, we introduce a Self-Supervised Counterfactual Reasoning (SSCR) framework that incorporates counterfactual thinking to overcome data scarcity.SSCR allows the model to consider out-ofdistribution instructions paired with previous images.With the help of cross-task consistency (CTC), we train these counterfactual instructions in a self-supervised scenario.Extensive results show that SSCR improves the correctness of ILBIE in terms of both object identity and position, establishing a new state of the art (SOTA) on two IBLIE datasets (i-CLEVR and CoDraw).Even with only 50% of the training data, SSCR achieves a comparable result to using complete data.
Tsu-Jui Fu, Xin Wang 0061, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
EMNLP (1)2
2020 Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations
abstract
A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings.To do this, it is critical to ensure that our evaluation protocols are correct, and benchmarks are reliable.In this work, we set forth to design a set of experiments to understand an important but often ignored problem in visually grounded language generation: given that humans have different utilities and visual attention, how will the sample variance in multi-reference datasets affect the models' performance?Empirically, we study several multi-reference datasets and corresponding vision-and-language tasks.We show that it is of paramount importance to report variance in experiments; that humangenerated references could vary drastically in different datasets/tasks, revealing the nature of each task; that metric-wise, CIDEr has shown systematically larger variances than others.Our evaluations on reference-per-instance shed light on the design of reliable datasets in the future.
Wanrong Zhu, Xin Wang 0061, Pradyumna Narayana, Kazoo Sone, Sugato Basu, William Yang Wang
EMNLP (1)2
2020 Relational Graph Learning for Grounded Video Description Generation
abstract
Grounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the model from hallucinating object words in its description. However, such design mainly focuses on object word generation and thus may ignore fine-grained information and suffer from missing visual concepts. Moreover, relational words (e.g., 'jump left or right') are usual spatio-temporal inference results, i.e., these words cannot be grounded on certain spatial regions. To tackle the above limitations, we design a novel relational graph learning framework for GVD, in which a language-refined scene graph representation is designed to explore fine-grained visual concepts. Furthermore, the refined graph can be regarded as relational inductive knowledge to assist captioning models in selecting the relevant information it needs to generate correct words. We validate the effectiveness of our model through automatic metrics and human evaluation, and the results indicate that our approach can generate more fine-grained and accurate description, and it solves the problem of object hallucination to some extent.
Wenqiao Zhang, Xin Wang 0061, Siliang Tang, Haizhou Shi, Jun Xiao 0001, Yueting Zhuang, William Yang Wang
ACM Multimedia2
2019 Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
abstract
Although promising results have been achieved in video captioning, existing models are limited to the fixed inventory of activities in the training corpus, and do not generalize to open vocabulary scenarios. Here we introduce a novel task, zeroshot video captioning, that aims at describing out-of-domain videos of unseen activities. Videos of different activities usually require different captioning strategies in many aspects, i.e. word selection, semantic construction, and style expression etc, which poses a great challenge to depict novel activities without paired training data. But meanwhile, similar activities share some of those aspects in common. Therefore, we propose a principled Topic-Aware Mixture of Experts (TAMoE) model for zero-shot video captioning, which learns to compose different experts based on different topic embeddings, implicitly transferring the knowledge learned from seen activities to unseen ones. Besides, we leverage external topic-related text corpus to construct the topic embedding for each activity, which embodies the most relevant semantic vectors within the topic. Empirical results not only validate the effectiveness of our method in utilizing semantic knowledge for video captioning, but also show its strong generalization ability when describing novel activities.
Xin Wang 0061, Jiawei Wu 0003, Da Zhang 0001, Yu Su 0001, William Yang Wang
AAAI1
2019 Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models
abstract
Variational autoencoders (VAEs) have received much attention recently as an end-toend architecture for text generation with latent variables.However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the posterior collapse issue has been widely identified in text-VAEs.In this paper, we propose to leverage several multi-level structures to learn a VAE model for generating long, and coherent text.In particular, a hierarchy of stochastic layers between the encoder and decoder networks is employed to abstract more informative and semantic-rich latent codes.Besides, we utilize a multi-level decoder structure to capture the coherent long-term structure inherent in long-form texts, by generating intermediate sentence representations as highlevel plan vectors.Extensive experimental results demonstrate that the proposed multi-level VAE model produces more coherent and less repetitive long text compared to baselines as well as can mitigate the posterior-collapse issue.
Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang 0002, Liqun Chen 0001, Xin Wang 0061, Jianfeng Gao 0001, Lawrence Carin
ACL (1)5
2019 Self-Supervised Learning for Contextualized Extractive Summarization
abstract
Existing models for extractive summarization are usually trained from scratch with a crossentropy loss, which does not explicitly capture the global context at the document level.In this paper, we aim to improve this task by introducing three auxiliary pre-training tasks that learn to capture the document-level context in a self-supervised fashion.Experiments on the widely-used CNN/DM dataset validate the effectiveness of the proposed auxiliary tasks.Furthermore, we show that after pretraining, a clean model with simple building blocks is able to outperform previous state-ofthe-art that are carefully designed.1
Hong Wang 0023, Xin Wang 0061, Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang
ACL (1)2
2019 Self-Supervised Dialogue Learning
abstract
The sequential order of utterances is often meaningful in coherent dialogues, and the order changes of utterances could lead to lowquality and incoherent conversations.We consider the order information as a crucial supervised signal for dialogue learning, which, however, has been neglected by many previous dialogue systems.Therefore, in this paper, we introduce a self-supervised learning task, inconsistent order detection, to explicitly capture the flow of conversation in dialogues.Given a sampled utterance pair triple, the task is to predict whether it is ordered or misordered.Then we propose a samplingbased self-supervised network SSN to perform the prediction with sampled triple references from previous dialogue history.Furthermore, we design a joint learning framework where SSN can guide the dialogue systems towards more coherent and relevant dialogue learning through adversarial training.We demonstrate that the proposed methods can be applied to both open-domain and taskoriented dialogue scenarios, and achieve the new state-of-the-art performance on the Open-Subtitiles and Movie-Ticket Booking datasets.
Jiawei Wu 0003, Xin Wang 0061, William Yang Wang
ACL (1)2
2019 Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
abstract
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieves the new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%).
Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001
CVPR1
2019 MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment
abstract
This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex temporal dependencies, which often happens in real scenarios. We identify two crucial challenges: semantic misalignment and structural misalignment. However, existing approaches treat different moments separately and do not explicitly model complex moment-wise temporal relations. In this paper, we present Moment Alignment Network (MAN), a novel framework that unifies the candidate moment encoding and temporal structural reasoning in a single-shot feed-forward network. MAN naturally assigns candidate moment representations aligned with language semantics over different temporal locations and scales. Most importantly, we propose to explicitly model moment-wise temporal relations as a structured graph and devise an iterative graph adjustment network to jointly learn the best structure in an end-to-end manner. We evaluate the proposed approach on two challenging public benchmarks DiDeMo and Charades-STA, where our MAN significantly outperforms the state-of-the-art by a large margin.
Da Zhang 0001, Xiyang Dai, Xin Wang 0061, Yuan-Fang Wang, Larry Davis 0001
CVPR3
2019 TIGEr: Text-to-Image Grounding for Image Caption Evaluation
abstract
Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, Jianfeng Gao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ming Jiang 0018, Qiuyuan Huang, Lei Zhang 0001, Xin Wang 0061, Pengchuan Zhang, Zhe Gan, Jana Diesner, Jianfeng Gao 0001
EMNLP/IJCNLP (1)4
2019 VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
abstract
We present a new large-scale multilingual video description dataset, VATEX1, which contains over 41,250 videos and 825, 000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSRVTT dataset [64], VATEX is multilingual, larger, linguistically complex, and more diverse in terms of both video and natural language descriptions. We also introduce two tasks for video-and-language research based on VATEX: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2) Video-guided Machine Translation, to translate a source language description into the target language using the video information as additional spatiotemporal context. Extensive experiments on the VATEX dataset show that, first, the unified multilingual model can not only produce both English and Chinese descriptions for a video more efficiently, but also offer improved performance over the monolingual models. Furthermore, we demonstrate that the spatiotemporal video context can be effectively utilized to align source and target languages and thus assist machine translation. In the end, we discuss the potentials of using VATEXfor other video-and-language research.
Xin Wang 0061, Jiawei Wu 0003, Jun-Kun Chen, Lei Li 0005, Yuan-Fang Wang, William Yang Wang
ICCV1
2018 No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling
abstract
Though impressive results have been achieved in visual captioning, the task of generating abstract stories from photo streams is still a little-tapped problem.Different from captions, stories have more expressive language styles and contain many imaginary concepts that do not appear in the images.Thus it poses challenges to behavioral cloning algorithms.Furthermore, due to the limitations of automatic metrics on evaluating story quality, reinforcement learning methods with hand-crafted rewards also face difficulties in gaining an overall performance boost.Therefore, we propose an Adversarial REward Learning (AREL) framework to learn an implicit reward function from human demonstrations, and then optimize policy search with the learned reward function.Though automatic evaluation indicates slight performance boost over state-of-the-art (SOTA) methods in cloning expert behaviors, human evaluation shows that our approach achieves significant improvement in generating more human-like stories than SOTA systems.Code will be made available here 1 .
Xin Wang 0061, Wenhu Chen, Yuan-Fang Wang, William Yang Wang
ACL (1)1
2018 S3D: Single Shot multi-Span Detector via Fully 3D Convolutional Networks
Da Zhang 0001, Xiyang Dai, Xin Wang 0061, Yuan-Fang Wang
BMVC3
2018 Video Captioning via Hierarchical Reinforcement Learning
abstract
Video captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short video, it is still very challenging to caption a video containing multiple fine-grained actions with a detailed description. This paper aims to address the challenge by proposing a novel hierarchical reinforcement learning framework for video captioning, where a high-level Manager module learns to design sub-goals and a low-level Worker module recognizes the primitive actions to fulfill the sub-goal. With this compositional framework to reinforce video captioning at different levels, our approach significantly outperforms all the baseline methods on a newly introduced large-scale dataset for fine-grained video captioning. Furthermore, our non-ensemble model has already achieved the state-of-the-art results on the widely-used MSR-VTT dataset.
Xin Wang 0061, Wenhu Chen, Jiawei Wu 0003, Yuan-Fang Wang, William Yang Wang
CVPR1
2018 Look Before You Leap: Bridging Model-Free and Model-Based Reinforcement Learning for Planned-Ahead Vision-and-Language Navigation
Xin Wang 0061, Wenhan Xiong, Hongmin Wang, William Yang Wang
ECCV (16)1
2018 XL-NBT: A Cross-lingual Neural Belief Tracking Framework
abstract
Task-oriented dialog systems are becoming pervasive, and many companies heavily rely on them to complement human agents for customer service in call centers.With globalization, the need for providing cross-lingual customer support becomes more urgent than ever.However, cross-lingual support poses great challenges-it requires a large amount of additional annotated data from native speakers.In order to bypass the expensive human annotation and achieve the first step towards the ultimate goal of building a universal dialog system, we set out to build a cross-lingual state tracking framework.Specifically, we assume that there exists a source language with dialog belief tracking annotations while the target languages have no annotated dialog data of any form.Then, we pre-train a state tracker for the source language as a teacher, which is able to exploit easy-to-access parallel data.We then distill and transfer its own knowledge to the student state tracker in target languages.We specifically discuss two types of common parallel resources: bilingual corpus and bilingual dictionary, and design different transfer learning strategies accordingly.Experimentally, we successfully use English state tracker as the teacher to transfer its knowledge to both Italian and German trackers and achieve promising results.
Wenhu Chen, Jianshu Chen, Yu Su 0001, Xin Wang 0061, Dong Yu 0001, Xifeng Yan, William Yang Wang
EMNLP4
2018 Virtual dictionary based kernel sparse representation for face recognition
Zizhu Fan, Da Zhang 0001, Xin Wang 0061, Qi Zhu 0001, Yuan-Fang Wang
Pattern Recognit.3
2017 Multimodal Transfer: A Hierarchical Deep Convolutional Neural Network for Fast Artistic Style Transfer
abstract
Transferring artistic styles onto everyday photographs has become an extremely popular task in both academia and industry. Recently, offline training has replaced online iterative optimization, enabling nearly real-time stylization. When those stylization networks are applied directly to high-resolution images, however, the style of localized regions often appears less similar to the desired artistic style. This is because the transfer process fails to capture small, intricate textures and maintain correct texture scales of the artworks. Here we propose a multimodal convolutional neural network that takes into consideration faithful representations of both color and luminance channels, and performs stylization hierarchically with multiple losses of increasing scales. Compared to state-of-the-art networks, our network can also perform style transfer in nearly real-time by performing much more sophisticated training offline. By properly handling style and texture cues at multiple scales using several modalities, we can transfer not just large-scale, obvious style cues but also subtle, exquisite ones. That is, our scheme can generate results that are visually pleasing and more similar to multiple desired artistic styles with color and texture cues at multiple scales.
Xin Wang 0061, Geoffrey Oxholm, Da Zhang 0001, Yuan-Fang Wang
CVPR1