EDBT 2026 Demo / reviewers in the wild / expert
Wenhu Chen
dblp:136/0957
· DBLP profile ↗
74ranked-venue papers
14as first author
57since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 13 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search AgentsabstractZijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Sahel Sharifymoghaddam, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, Yanxi Li, Haoran Hong, Xinyu Shi, Xuye Liu, Hosna Oyarhoseini, Nandan Thakur, Crystina Zhang, Luyu Gao, Wenhu Chen, Jimmy Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xueguang Ma, Shengyao Zhuang, Ping Nie, Sahel Sharifymoghaddam, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, Haoran Hong, Xuye Liu, Hosna Oyarhoseini, Nandan Thakur, Xinyu Zhang 0018, Luyu Gao, Wenhu Chen, Jimmy Lin |
ACL (1) | 20 |
| 2025 | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at ScaleabstractJiawei Guo, Tianyu Zheng, Yizhi Li, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyu Zheng, Yuelin Bai, Bo Li 0080, Yubo Wang 0019, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue |
ACL (1) | 9 |
| 2025 | TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem UnderstandingabstractUnderstanding domain-specific theorems often requires more than just text-based reasoning; effective communication through structured visual explanations is crucial for deeper comprehension. While large language models (LLMs) demonstrate strong performance in text-based theorem reasoning, their ability to generate coherent and pedagogically meaningful visual explanations remains an open challenge. In this work, we introduce TheoremExplainAgent, an agentic approach for generating long-form theorem explanation videos (over 5 minutes) using Manim animations. To systematically evaluate multimodal theorem explanations, we propose TheoremExplainBench, a benchmark covering 240 theorems across multiple STEM disciplines, along with 5 automated evaluation metrics. Our results reveal that agentic planning is essential for generating detailed long-form videos, and the o3-mini agent achieves a success rate of 93.8% and an overall score of 0.77. However, our quantitative and qualitative studies show that most of the videos produced exhibit minor issues with visual element layout. Furthermore, multimodal explanations expose deeper reasoning flaws that text-based explanations fail to reveal, highlighting the importance of multimodal explanations. Max Ku, Cheuk Hei Chong, Jonathan Leung, Krish Shah, Alvin Yu, Wenhu Chen |
ACL (1) | 6 |
| 2025 | VISA: Retrieval Augmented Generation with Visual Source AttributionabstractGeneration with source attribution is important for enhancing the verifiability of retrievalaugmented generation (RAG) systems.However, existing approaches in RAG primarily link generated content to document-level references, making it challenging for users to locate evidence among multiple content-rich retrieved documents.To address this challenge, we propose Retrieval-Augmented Generation with Visual Source Attribution (VISA), a novel approach that combines answer generation with visual source attribution.Leveraging large vision-language models (VLMs), VISA identifies the evidence and highlights the exact regions that support the generated answers with bounding boxes in the retrieved document screenshots.To evaluate its effectiveness, we curated two datasets: Wiki-VISA, based on crawled Wikipedia webpage screenshots, and Paper-VISA, derived from PubLayNet and tailored to the medical domain.Experimental results demonstrate the effectiveness of VISA for visual source attribution on documents' original look, as well as highlighting the challenges for improvement.1 Xueguang Ma, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Wenhu Chen, Jimmy Lin |
ACL (1) | 5 |
| 2025 | MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkabstractXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 0019, Kai Zhang 0033, Shengbang Tong, Yuxuan Sun 0002, Botao Yu, Ge Zhang 0009, Huan Sun 0001, Yu Su 0001, Wenhu Chen, Graham Neubig |
ACL (1) | 12 |
| 2025 | ACECODER: Acing Coder RL via Automated Test-Case SynthesisabstractMost progress in recent coder models has been driven by supervised fine-tuning (SFT), while the potential of reinforcement learning (RL) remains largely unexplored, primarily due to the lack of reliable reward data/model in the code domain.In this paper, we address this challenge by leveraging automated large-scale testcase synthesis to enhance code model training.Specifically, we design a pipeline that generates extensive (question, test-cases) pairs from existing code data.Using these test cases, we construct preference pairs based on pass rates over sampled programs to train reward models with Bradley-Terry loss.It shows an average of 10-point improvement for Llama-3.1-8B-Ins and 5-point improvement for Qwen2.5-Coder-7B-Insthrough best-of-32 sampling, making the 7B model on par with 236B DeepSeek-V2.5.Furthermore, we conduct reinforcement learning with both reward models and testcase pass rewards, leading to consistent improvements across HumanEval, MBPP, Big-CodeBench, and LiveCodeBench (V4).Notably, we follow the R1-style training to start from Qwen2.5-Coder-base directly and show that our RL training can improve model on HumanEval-plus by over 25% and MBPP-plus by 6% for merely 80 optimization steps.We believe our results highlight the huge potential of reinforcement learning in coder models. Huaye Zeng, Dongfu Jiang, Haozhe Wang 0002, Ping Nie, Wenhu Chen |
ACL (1) | 6 |
| 2025 | VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal AugmentationabstractCurrent large multimodal models (LMMs) face significant challenges in processing and comprehending long-duration or high-resolution videos, which is mainly due to the lack of high-quality datasets. To address this issue from a Average Accuracy data-centric perspective, we propose VISTA, a simple yet effective VIdeo SpatioTemporal Augmentation framework that synthesizes long-duration and high-resolution video instruction-following pairs from existing video-caption datasets. VISTA spatially and temporally combines videos to create new synthetic videos with extended durations and enhanced resolutions, and subsequently produces question-answer pairs pertaining to these newly synthesized videos. Based on this paradigm, we develop seven video augmentation methods and curate VISTA-400K, a video instruction-following dataset aimed at enhancing long-duration and high-resolution video understanding. Finetuning various video LMMs on our data resulted in an average improvement of 3.3% across four challenging benchmarks for long-video understanding. Furthermore, we introduce the first comprehensive high-resolution video understanding benchmark HRVideoBench, on which our finetuned models achieve a 6.5% performance gain. These results highlight the effectiveness of our framework. Weiming Ren, Huan Yang 0005, Jie Min, Cong Wei 0001, Wenhu Chen |
CVPR | 5 |
| 2025 | VisualWebInstruct: Scaling up Multimodal Instruction Data through Web SearchabstractVision-Language Models have made significant progress on many perception-focused tasks.However, their progress on reasoningfocused tasks remains limited due to the lack of high-quality and diverse training data.In this work, we aim to address the scarcity of reasoning-focused multimodal datasets.We propose VisualWebInstruct, a novel approach that leverages search engines to create a diverse and high-quality dataset spanning multiple disciplines, including mathematics, physics, finance, and chemistry, etc. Starting with a meticulously selected set of 30,000 seed images, we employ Google Image Search to identify websites containing similar images.We collect and process HTML data from over 700K unique URLs.Through a pipeline of content extraction, filtering, and synthesis, we construct a dataset of approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs and the remaining comprising text-based QA pairs.Models finetuned on VisualWebInstruct demonstrate significant performance improvements: (1) finetuning on Llava-OV results in 10-20 absolute points improvement across benchmarks, and (2) fine-tuning from MAmmoTH-VL yields a 5 absolute points gain across benchmarks.Our best model, MAmmoTH-VL2, achieves the best known performance with SFT without RL within the 10B parameter class on MMMU-Pro (40.7),MathVerse (42.6), and DynaMath (55.7).These results highlight the effectiveness of our dataset in enhancing the reasoning capabilities of vision-language models for complex multimodal tasks. Yiming Jia, Xiang Yue, Bo Li 0080, Ping Nie, Wenhu Chen |
EMNLP | 7 |
| 2025 | Unleashing the Reasoning Potential of LLMs by Critique Fine-Tuning on One ProblemabstractWe have witnessed that strong LLMs like Qwen-Math, MiMo, and Phi-4 possess immense reasoning potential inherited from the pre-training stage.With reinforcement learning (RL), these models can improve dramatically on reasoning tasks.Recent studies have shown that even RL on a single problem (Wang et al., 2025a) can unleash these models' reasoning capabilities.However, RL is not only expensive but also unstable.Even one-shot RL requires hundreds of GPU hours.This raises a critical question: Is there a more efficient way to unleash the reasoning potential of these powerful base LLMs?In this work, we demonstrate that Critique Fine-Tuning (CFT) on only one problem can effectively unleash the reasoning potential of LLMs.Our method constructs critique data by collecting diverse model-generated solutions to a single problem and using teacher LLMs to provide detailed critiques.We finetune Qwen and Llama family models, ranging from 1.5B to 14B parameters, on the CFT data and observe significant performance gains across diverse reasoning tasks.For example, with just 5 GPU hours of training, Qwen-Math-7B-CFT show an average improvement of 15% on six math benchmarks and 16% on three logic reasoning benchmarks.These results are comparable to or even surpass the results from RL with 20x less compute.Ablation studies reveal the robustness of one-shot CFT across different prompt problems.These results highlight oneshot CFT as a simple, general, and computeefficient approach to unleashing the reasoning capabilities of modern LLMs. Ping Nie, Wenhu Chen |
EMNLP | 5 |
| 2025 | VAMBA: Understanding Hour-Long Videos with Hybrid Mamba-TransformersabstractState-of-the-art transformer-based large multimodal models (LMMs) struggle to handle hour-long video inputs due to the quadratic complexity of the causal self-attention operations, leading to high computational costs during training and inference. Existing token compression-based methods reduce the number of video tokens but often incur information loss and remain inefficient for extremely long sequences. In this paper, we explore an orthogonal direction to build a hybrid Mamba-Transformer model (VAMBA) that employs Mamba-2 blocks to encode video tokens with linear complexity. Without any token reduction, VAMBA can encode more than 1024 frames (640$\times$360) on a single GPU, while transformer-based models can only encode 256 frames. On long video input, VAMBA achieves at least 50% reduction in GPU memory usage during training and inference, and nearly doubles the speed per training step compared to transformer-based LMMs. Our experimental results demonstrate that VAMBA improves accuracy by 4.3% on the challenging hour-long video understanding benchmark LVBench over prior efficient video LMMs, and maintains strong performance on a broad spectrum of long and short video understanding tasks. Weiming Ren, Huan Yang 0005, Cong Wei 0001, Ge Zhang 0009, Wenhu Chen |
ICCV | 6 |
| 2025 | MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World TasksabstractWe present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users.
Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation.
In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MM-Bench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks.
Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions. Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang 0068, Yubo Wang 0019, Yuansheng Ni, Ziyan Jiang, Wang Zhu 0001, Bohan Lyu 0001, Dongfu Jiang, Hexiang Hu, Xiang Yue, Wenhu Chen |
ICLR | 16 |
| 2025 | VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding TasksabstractEmbedding models play a crucial role in a variety of downstream tasks, including semantic similarity, information retrieval, and clustering. While there has been a surge of interest in developing universal text embedding models that generalize across tasks (e.g., MTEB), progress in learning universal multimodal embedding models has been comparatively slow, despite their importance and practical applications. In this work, we explore the potential of building universal multimodal embeddings capable of handling a broad range of downstream tasks.
Our contributions are twofold: (1) we propose MMEB (Massive Multimodal Embedding Benchmark), which covers four meta-tasks (classification, visual question answering, multimodal retrieval, and visual grounding) and 36 datasets, including 20 training datasets and 16 evaluation datasets spanning both in-distribution and out-of-distribution tasks, and (2) VLM2Vec (Vision-Language Model → Vector), a contrastive training framework that transforms any vision-language model into an embedding model through contrastive training on MMEB.
Unlike previous models such as CLIP and BLIP, which encode text and images independently without task-specific guidance, VLM2Vec can process any combination of images and text while incorporating task instructions to generate a fixed-dimensional vector. We develop a series of VLM2Vec models based on state-of-the-art VLMs, including Phi-3.5-V, LLaVA-1.6, and Qwen2-VL, and evaluate them on MMEB’s benchmark. With LoRA tuning, VLM2Vec achieves a 10% to 20% improvement over existing multimodal embedding models on MMEB’s evaluation sets. Our findings reveal that VLMs are surprisingly strong embedding models. Ziyan Jiang, Xinyi Yang 0002, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen |
ICLR | 6 |
| 2025 | T2V-Turbo-v2: Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance DesignabstractIn this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a significant advancement by integrating various supervision signals, including high-quality training data, reward model feedback, and conditional guidance, into the consistency distillation process. Through comprehensive ablation studies, we highlight the crucial importance of tailoring datasets to specific learning objectives and the effectiveness of learning from diverse reward models for enhancing both the visual quality and text-video alignment. Additionally, we highlight the vast design space of conditional guidance strategies, which centers on designing an effective energy function to augment the teacher ODE solver. We demonstrate the potential of this approach by extracting motion guidance from the training datasets and incorporating it into the ODE solver, showcasing its effectiveness in improving the motion quality of the generated videos with the improved motion-related metrics from VBench and T2V-CompBench. Empirically, our T2V-Turbo-v2 establishes a new state-of-the-art result on VBench, **with a Total score of 85.13**, surpassing proprietary systems such as Gen-3 and Kling. Qian Long, Xiaofeng Gao 0002, Robinson Piramuthu, Wenhu Chen, William Yang Wang |
ICLR | 6 |
| 2025 | Harnessing Webpage UIs for Text-Rich Visual UnderstandingabstractText-rich visual understanding—the ability to interpret both textual content and visual elements within a scene—is crucial for multimodal large language models (MLLMs) to effectively interact with structured environments. We propose leveraging webpage UIs as a naturally structured and diverse data source to enhance MLLMs’ capabilities in this area. Existing approaches, such as rule-based extraction, multimodal model captioning, and rigid HTML parsing, are hindered by issues like noise, hallucinations, and limited generalization. To overcome these challenges, we introduce MultiUI, a dataset of 7.3 million samples spanning various UI types and tasks, structured using enhanced accessibility trees and task taxonomies. By scaling multimodal instructions from web UIs through LLMs, our dataset enhances generalization beyond web domains, significantly improving performance in document understanding, GUI comprehension, grounding, and advanced agent tasks. This demonstrates the potential of structured web data to elevate MLLMs’ proficiency in processing text-rich visual environments and generalizing across domains. Junpeng Liu 0001, Tianyue Ou, Yifan Song 0002, Yuxiao Qu, Wai Lam, Chenyan Xiong, Wenhu Chen, Graham Neubig, Xiang Yue |
ICLR | 7 |
| 2025 | OmniEdit: Building Image Editing Generalist Models Through Specialist SupervisionabstractInstruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from practical, real-life applications. We identify three primary challenges contributing to this gap. Firstly, existing models have limited editing skills due to the biased synthesis process. Secondly, these methods are trained with datasets with a high volume of noise and artifacts. This is due to the application of simple filtering methods like CLIP-score. Thirdly, all these datasets are restricted to a single low resolution and fixed aspect ratio, limiting the versatility to handle real-world use cases.
In this paper, we present OmniEdit, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) OmniEdit is trained by utilizing the supervision from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality. (3) we propose a new editing architecture called EditNet to greatly boost the editing success rate, (4) we provide images with different aspect ratios to ensure that our model can handle any image in the wild. We have curated a test set containing images of different aspect ratios, accompanied by diverse instructions to cover different tasks. Both automatic evaluation and human evaluations demonstrate that OmniEdit can significantly outperforms all the existing models. Cong Wei 0001, Zheyang Xiong, Weiming Ren, Xeron Du, Ge Zhang 0009, Wenhu Chen |
ICLR | 6 |
| 2025 | General-Reasoner: Advancing LLM Reasoning Across All DomainsabstractReinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs).
Particularly, the "Zero" reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate supervised fine-tuning stage.
Despite these advancements, current works for LLM reasoning mainly focus on mathematical and coding domains, largely due to data abundance and the ease of answer verification.
This limits the applicability and generalization of such models to broader domains, where questions often have diverse answer representations, and data is more scarce.
In this paper, we propose General-Reasoner, a novel training framework designed to enhance LLM reasoning capabilities across diverse domains. Our key contributions include: (1) constructing a large-scale, high-quality dataset of questions with verifiable answers curated by web crawling, covering a wide range of disciplines; and (2) developing a generative model-based answer verifier, which replaces traditional rule-based verification with the capability of chain-of-thought and context-awareness. We train a series of models and evaluate them on a wide range of datasets covering wide domains like physics, chemistry, finance, electronics etc. Our comprehensive evaluation across these 12 benchmarks (e.g. MMLU-Pro, GPQA, SuperGPQA, TheoremQA, BBEH and MATH AMC) demonstrates that General-Reasoner outperforms existing baseline methods, achieving robust and generalizable reasoning performance while maintaining superior effectiveness in mathematical reasoning tasks. Xueguang Ma, Qian Liu 0033, Dongfu Jiang, Ge Zhang 0009, Zejun Ma 0001, Wenhu Chen |
NeurIPS | 6 |
| 2025 | Pixel Reasoner: Incentivizing Pixel Space Reasoning via Curiosity-Driven Reinforcement LearningabstractChain-of-thought reasoning has significantly improved the performance of Large
Language Models (LLMs) across various domains. However, this reasoning process has been confined exclusively to textual space, limiting its effectiveness in visually intensive tasks. To address this limitation, we introduce the concept of pixel-space reasoning. Within this novel framework, Vision-Language Models (VLMs) are equipped with a suite of visual reasoning operations, such as zoom-in and select-frame. These operations enable VLMs to directly inspect, interrogate, and infer from visual evidences, thereby enhancing reasoning fidelity for visual tasks. Cultivating such pixel-space reasoning capabilities in VLMs presents notable challenges, including the model’s initially imbalanced competence and its reluctance to adopt the newly introduced pixel-space operations. We address these challenges through a two-phase training approach. The first phase employs instruction tuning on synthesized reasoning traces to familiarize the model with the novel visual operations. Following this, a reinforcement learning (RL) phase leverages a curiosity-driven reward scheme to balance exploration between pixel-space reasoning and textual reasoning. With these visual operations, VLMs can interact with complex visual inputs, such as information-rich images or videos to proactively gather necessary information. We demonstrate that this approach significantly improves VLM performance across diverse visual reasoning benchmarks. Our 7B model, Pixel-Reasoner, achieves 84% on V* bench, 74% on TallyQA-Complex, and 84% on InfographicsVQA, marking the highest accuracy achieved by any open-source model to date. These results highlight the importance of pixel-space reasoning and the effectiveness of our framework. Alex Su, Haozhe Wang 0002, Weiming Ren, Fangzhen Lin, Wenhu Chen |
NeurIPS | 5 |
| 2025 | Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningabstractContrastive learning (CL) is a prevalent technique for training embedding models, which pulls semantically similar examples (positives) closer in the representation space while pushing dissimilar ones (negatives) further apart. A key source of negatives are "in-batch" examples, i.e., positives from other examples in the batch. Effectiveness of such models is hence strongly influenced by the size and quality of training batches. In this work, we propose *Breaking the Batch Barrier* (B3), a novel batch construction strategy designed to curate high-quality batches for CL. Our approach begins by using a pretrained teacher embedding model to rank all examples in the dataset, from which a sparse similarity graph is constructed. A community detection algorithm is then applied to this graph to identify clusters of examples that serve as strong negatives for one another. The clusters are then used to construct batches that are rich in in-batch negatives. Empirical results on the MMEB multimodal embedding benchmark (36 tasks) demonstrate that our method sets a new state of the art, outperforming previous best methods by +1.3 and +2.9 points at the 7B and 2B model scales, respectively. Notably, models trained with B3 surpass existing state-of-the-art results even with a batch size as small as 64, which is 4–16× smaller than that required by other methods. Moreover, experiments show that B3 generalizes well across domains and tasks, maintaining strong performance even when trained with considerably weaker teachers. Raghuveer Thirukovalluru, Ye Liu 0006, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen, Bhuwan Dhingra |
NeurIPS | 9 |
| 2025 | VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningabstractRecently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection. They significantly outperform the best fast-thinking models, such as GPT-4o, on various math and science benchmarks. However, their multimodal reasoning capabilities remain on par with fast-thinking models. For instance, GPT-o1's performance on benchmarks like MathVista, MathVerse, and MathVision is similar to fast-thinking models. In this paper, we aim to enhance the slow-thinking capabilities of vision-language models using reinforcement learning (without relying on distillation) to advance the state of the art. First, we adapt the GRPO algorithm with a novel technique called Selective Sample Replay (SSR) to address the vanishing advantages problem. While this approach yields strong performance, the resulting RL-trained models exhibit limited self-reflection or self-verification. To further encourage slow-thinking, we introduce Forced Rethinking, which appends a rethinking trigger token to the end of rollouts in RL training, explicitly enforcing a self-reflection reasoning step. By combining these two techniques, our model, VL-Rethinker, advances state-of-the-art scores on MathVista, MathVerse to achieve 80.4%, 63.5% respectively. VL-Rethinker also achieves open-source SoTA on multi-disciplinary benchmarks such as MathVision, MMMU-Pro, EMMA, and MEGA-Bench, narrowing the gap with OpenAI-o1. We conduct comprehensive ablations and analysis to provide insights into the effectiveness of our approach. Haozhe Wang 0002, Chao Qu, Zuming Huang, Fangzhen Lin, Wenhu Chen |
NeurIPS | 6 |
| 2025 | MoCha: Towards Movie-Grade Talking Character GenerationabstractRecent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation.
We introduce Talking Characters, a more realistic task to generate talking character animations directly from speech and text. Unlike talking head tasks, Talking Characters aims at generating the full portrait of one or more characters beyond the facial region.
In this paper, we propose MoCha, the first of its kind to generate talking characters. To ensure precise synchronization between video and speech, we propose a localized audio attention mechanism that effectively aligns speech and video tokens.
To address the scarcity of large-scale speech-labelled video datasets, we introduce a joint training strategy that leverages both speech-labelled and text-labelled video data, significantly improving generalization across diverse character actions. We also design structured prompt templates with character tags, enabling, for the first time, multi-character conversation with turn-based dialogue—allowing AI-generated characters to engage in context-aware conversations with cinematic coherence.
Extensive qualitative and quantitative evaluations, including human evaluation studies and benchmark comparisons, demonstrate that MoCha sets a new standard for AI-generated cinematic storytelling, achieving superior realism, controllability and generalization. Cong Wei 0001, Ji Hou, Felix Juefei-Xu, Zecheng He, Xiaoliang Dai, Luxin Zhang, Tingbo Hou, Animesh Sinha, Peter Vajda, Wenhu Chen |
NeurIPS | 13 |
| 2024 | VIEScore: Towards Explainable Metrics for Conditional Image Synthesis EvaluationabstractIn the rapidly advancing field of conditional image generation research, challenges such as limited explainability lie in effectively evaluating the performance and capabilities of various models.This paper introduces VIESCORE, a Visual Instruction-guided Explainable metric for evaluating any conditional image generation tasks.VIESCORE leverages general knowledge from Multimodal Large Language Models (MLLMs) as the backbone and does not require training or fine-tuning.We evaluate VIESCORE on seven prominent tasks in conditional image tasks and found: (1) VIESCORE (GPT4-o) achieves a high Spearman correlation of 0.4 with human evaluations, while the human-to-human correlation is 0.45.(2) VI-ESCORE (with open-source MLLM) is significantly weaker than GPT-4o and GPT-4v in evaluating synthetic images.(3) VIESCORE achieves a correlation on par with human ratings in the generation tasks but struggles in editing tasks.With these results, we believe VIESCORE shows its great potential to replace human judges in evaluating image synthesis tasks. Max Ku, Dongfu Jiang, Cong Wei 0001, Xiang Yue, Wenhu Chen |
ACL (1) | 5 |
| 2024 | Instruct-Imagen: Image Generation with Multi-modal InstructionabstractThis paper presents Instruct-Imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce multi-modal in-struction for image generation, a task representation artic-ulating a range of generation intents with precision. It uses natural language to amalgamate disparate modalities (e.g., text, edge, style, subject, etc.), such that abundant generation intents can be standardized in a uniform format. We then build Instruct - Imagen by fine-tuning a pre-trained text-to-image diffusion model with two stages. First, we adapt the model using the retrieval-augmented training, to enhance model's capabilities to ground its generation on external multi-modal context. Subsequently, we fine-tune the adapted model on diverse image generation tasks that requires vision-language understanding (e.g., subject-driven generation, etc.), each paired with a multi-modal instruction encapsulating the task's essence. Human evaluation on various image generation datasets re-veals that Instruct-Imagen matches or surpasses prior task-specific models in-domain and demonstrates promising generalization to unseen and more complex tasks. Our evaluation suite will be made publicly available. Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen, Yandong Li, Kihyuk Sohn, Xue Ben, Boqing Gong, William W. Cohen, Ming-Wei Chang, Xuhui Jia |
CVPR | 4 |
| 2024 | MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIabstractWe introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence. Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
CVPR | 22 |
| 2024 | UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
Cong Wei 0001, Yang Chen 0065, Hexiang Hu, Ge Zhang 0009, Jie Fu 0001, Alan Ritter, Wenhu Chen |
ECCV (87) | 8 |
| 2024 | VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video GenerationabstractXuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, Wenhu Chen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Dongfu Jiang, Ge Zhang 0009, Max Ku, Achint Soni, Sherman Siu, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang 0068, Quy Duc Do, Yuansheng Ni, Bohan Lyu 0001, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Y. Lin, Wenhu Chen |
EMNLP | 19 |
| 2024 | Unifying Multimodal Retrieval via Document Screenshot EmbeddingabstractIn the real world, documents are organized in different formats and varied modalities.Traditional retrieval pipelines require tailored document parsing techniques and content extraction modules to prepare input for indexing.This process is tedious, prone to errors, and has information loss.To this end, we propose Document Screenshot Embedding (DSE), a novel retrieval paradigm that regards document screenshots as a unified input format, which does not require any content extraction preprocess and preserves all the information in a document (e.g., text, image and layout).DSE leverages a large vision-language model to directly encode document screenshots into dense representations for retrieval.To evaluate our method, we first craft the dataset of Wiki-SS, a 1.3M Wikipedia web page screenshots as the corpus to answer the questions from the Natural Questions dataset.In such a text-intensive document retrieval setting, DSE shows competitive effectiveness compared to other text retrieval methods relying on parsing.For example, DSE outperforms BM25 by 17 points in top-1 retrieval accuracy.Additionally, in a mixed-modality task of slide retrieval, DSE significantly outperforms OCR text retrieval methods by over 15 points in [email protected] experiments show that DSE is an effective document retrieval paradigm for diverse types of documents.Model checkpoints, code, and Wiki-SS collection are released at http://tevatron.ai. Xueguang Ma, Sheng-Chieh Lin, Minghan Li 0002, Wenhu Chen, Jimmy Lin |
EMNLP | 4 |
| 2024 | ImagenHub: Standardizing the evaluation of conditional image generation modelsabstractRecently, a myriad of conditional image generation and editing models have been developed to serve different downstream tasks, including text-to-image generation, text-guided image editing, subject-driven image generation, control-guided image generation, etc. However, we observe huge inconsistencies in experimental conditions: datasets, inference, and evaluation metrics -- render fair comparisons difficult.
This paper proposes ImagenHub, which is a one-stop library to standardize the inference and evaluation of all the conditional image generation models. Firstly, we define seven prominent tasks and curate high-quality evaluation datasets for them. Secondly, we built a unified inference pipeline to ensure fair comparison. Thirdly, we design two human evaluation scores, i.e. Semantic Consistency and Perceptual Quality, along with comprehensive guidelines to evaluate generated images. We train expert raters to evaluate the model outputs based on the proposed metrics. Our human evaluation achieves a high inter-worker agreement of Krippendorff’s alpha on 76\% models with a value higher than 0.4. We comprehensively evaluated a total of around 30 models and observed three key takeaways: (1) the existing models’ performance is generally unsatisfying except for Text-guided Image Generation and Subject-driven Image Generation, with 74\% models achieving an overall score lower than 0.5. (2) we examined the claims from published papers and found 83\% of them hold with a few exceptions. (3) None of the existing automatic metrics has a Spearman's correlation higher than 0.2 except subject-driven image generation. Moving forward, we will continue our efforts to evaluate newly published models and update our leaderboard to keep track of the progress in conditional image generation. Max Ku, Tianle Li, Kai Zhang 0033, Wenwen Zhuang, Wenhu Chen |
ICLR | 7 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 14 |
| 2024 | Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsabstractRecent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." Xichen Pan, Li Dong 0004, Shaohan Huang, Zhiliang Peng, Wenhu Chen, Furu Wei |
ICLR | 5 |
| 2024 | MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningabstractWe introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, our meticulously curated instruction tuning dataset. MathInstruct is compiled from 13 math datasets with intermediate rationales, six of which have rationales newly curated by us. It presents a unique hybrid of chain-of-thought (CoT) and program-of-thought (PoT) rationales, and also ensures extensive coverage of diverse fields in math. The hybrid of CoT and PoT not only unleashes the potential of tool use but also allows different thought processes for different math problems. As a result, the MAmmoTH series substantially outperform existing open-source models on nine mathematical reasoning datasets across all scales with an average accuracy gain between 16% and 32%. Remarkably, our MAmmoTH-7B model reaches 33% on MATH (a competition-level dataset), which exceeds the best open-source 7B model (WizardMath) by 23%, and the MAmmoTH-34B model achieves 44% accuracy on MATH, even surpassing GPT-4’s CoT result. Our work underscores the importance of diverse problem coverage and the use of hybrid rationales in developing superior math generalist models. Xiang Yue, Xingwei Qu, Ge Zhang 0009, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
ICLR | 8 |
| 2024 | Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths AggregationabstractPre-trained language models (LMs) are able to perform complex reasoning without explicit fine-tuning. To understand how pre-training with a next-token prediction objective contributes to the emergence of such reasoning capability, we propose that we can view an LM as deriving new conclusions by aggregating indirect reasoning paths seen at pre-training time. We found this perspective effective in two important cases of reasoning: logic reasoning with knowledge graphs (KGs) and chain-of-thought (CoT) reasoning. More specifically, we formalize the reasoning paths as random walk paths on the knowledge/reasoning graphs. Analyses of learned LM distributions suggest that a weighted sum of relevant random walk path probabilities is a reasonable way to explain how LMs reason. Experiments and analysis on multiple KG and CoT datasets reveal the effect of training on random walk paths and suggest that augmenting unlabeled random walk reasoning paths can improve real-world multi-step reasoning performance. Xinyi Wang 0003, Alfonso Amayuelas, Kexun Zhang, Liangming Pan, Wenhu Chen, William Yang Wang |
ICML | 5 |
| 2024 | MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsabstractImage retrieval, i.e., finding desired images given a reference image, inherently encompasses rich, multi-faceted search intents that are difficult to capture solely using image-based measures. Recent works leverage text instructions to allow users to more freely express their search intents. However, they primarily focus on image pairs that are visually similar and/or can be characterized by a small set of pre-defined relations. The core thesis of this paper is that text instructions can enable retrieving images with richer relations beyond visual similarity. To show this, we introduce MagicLens, a series of self-supervised image retrieval models that support open-ended instructions. MagicLens is built on a key novel insight: image pairs that naturally occur on the same web pages contain a wide range of implicit relations (e.g., inside view of), and we can bring those implicit relations explicit by synthesizing instructions via foundation models. Trained on 36.7M (query image, instruction, target image) triplets with rich semantic relations mined from the web, MagicLens achieves results comparable with or better than prior best on eight benchmarks of various image retrieval tasks, while maintaining high parameter efficiency with a significantly smaller model size. Additional human analyses on a 1.4M-image unseen corpus further demonstrate the diversity of search intents supported by MagicLens. Code and models are publicly available at the https://open-vision-language.github.io/MagicLens/. Kai Zhang 0033, Yi Luan, Hexiang Hu, Kenton Lee, Siyuan Qiao, Wenhu Chen, Yu Su 0001, Ming-Wei Chang |
ICML | 6 |
| 2024 | GenAI Arena: An Open Evaluation Platform for Generative ModelsabstractGenerative AI has made remarkable strides to revolutionize fields such as image and video generation. These advancements are driven by innovative algorithms, architecture, and data. However, the rapid proliferation of generative models has highlighted a critical gap: the absence of trustworthy evaluation metrics. Current automatic assessments such as FID, CLIP, FVD, etc often fail to capture the nuanced quality and user satisfaction associated with generative outputs. This paper proposes an open platform GenAI-Arena to evaluate different image and video generative models, where users can actively participate in evaluating these models. By leveraging collective user feedback and votes, GenAI-Arena aims to provide a more democratic and accurate measure of model performance. It covers three tasks of text-to-image generation, text-to-video generation, and image editing respectively. Currently, we cover a total of 35 open-source generative models. GenAI-Arena has been operating for seven months, amassing over 9000 votes from the community. We describe our platform, analyze the data, and explain the statistical methods for ranking the models. To further promote the research in building model-based evaluation metrics, we release a cleaned version of our preference data for the three tasks, namely GenAI-Bench. We prompt the existing multi-modal models like Gemini, and GPT-4o to mimic human voting. We compute the accuracy by comparing the model voting with the human voting to understand their judging abilities. Our results show existing multimodal models are still lagging in assessing the generated visual content, even the best model GPT-4o only achieves an average accuracy of $49.19\%$ across the three generative tasks. Open-source MLLMs perform even worse due to the lack of instruction-following and reasoning ability in complex vision scenarios. Dongfu Jiang, Max Ku, Tianle Li, Yuansheng Ni, Shizhuo Sun, Rongqi Fan, Wenhu Chen |
NeurIPS | 7 |
| 2024 | T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackabstractDiffusion-based text-to-video (T2V) models have achieved significant success but continue to be hampered by the slow sampling speed of their iterative sampling processes. To address the challenge, consistency models have been proposed to facilitate fast inference, albeit at the cost of sample quality. In this work, we aim to break the quality bottleneck of a video consistency model (VCM) to achieve **both fast and high-quality video generation**. We introduce T2V-Turbo, which integrates feedback from a mixture of differentiable reward models into the consistency distillation (CD) process of a pre-trained T2V model. Notably, we directly optimize rewards associated with single-step generations that arise naturally from computing the CD loss, effectively bypassing the memory constraints imposed by backpropagating gradients through an iterative sampling process. Remarkably, the 4-step generations from our T2V-Turbo achieve the highest total score on VBench, even surpassing Gen-2 and Pika. We further conduct human evaluations to corroborate the results, validating that the 4-step generations from our T2V-Turbo are preferred over the 50-step DDIM samples from their teacher models, representing more than a tenfold acceleration while improving video generation quality. Weixi Feng, Tsu-Jui Fu, Xinyi Wang 0003, Sugato Basu, Wenhu Chen, William Yang Wang |
NeurIPS | 6 |
| 2024 | WildVision: Evaluating Vision-Language Models in the Wild with Human PreferencesabstractRecent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that collects human preferences to evaluate VLMs. We curated WV-Bench by selecting 500 high-quality samples from 8,000 user submissions in WV-Arena. WV-Bench uses GPT-4 as the judge to compare each VLM with Claude-3-Sonnet, achieving a Spearman correlation of 0.94 with the WV-Arena Elo. This significantly outperforms other benchmarks like MMVet, MMMU, and MMStar.Our comprehensive analysis of 20K real-world interactions reveals important insights into the failure cases of top-performing VLMs. For example, we find that although GPT-4V surpasses many other models like Reka-Flash, Opus, and Yi-VL-Plus in simple visual recognition and reasoning tasks, it still faces challenges with subtle contextual cues, spatial reasoning, visual imagination, and expert domain knowledge. Additionally, current VLMs exhibit issues with hallucinations and safety when intentionally provoked. We are releasing our chat and feedback data to further advance research in the field of VLMs. Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi 0001, Bill Y. Lin |
NeurIPS | 3 |
| 2024 | MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkabstractIn the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates part of the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16\% to 33\% compared to MMLU, but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5\% in MMLU to just 2\% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is more discriminative benchmark to better track progress in the field. Yubo Wang 0019, Xueguang Ma, Ge Zhang 0009, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang 0068, Alex Zhuang, Rongqi Fan, Xiang Yue, Wenhu Chen |
NeurIPS | 17 |
| 2024 | MAmmoTH2: Scaling Instructions from the WebabstractInstruction tuning improves the reasoning abilities of large language models (LLMs), with data quality and scalability being the crucial factors. Most instruction tuning data come from human crowd-sourcing or GPT-4 distillation. We propose a paradigm to efficiently harvest 10 million naturally existing instruction data from the pre-training web corpus to enhance LLM reasoning. Our approach involves (1) recalling relevant documents, (2) extracting instruction-response pairs, and (3) refining the extracted pairs using open-source LLMs. Fine-tuning base LLMs on this dataset, we build MAmmoTH2 models, which significantly boost performance on reasoning benchmarks. Notably, MAmmoTH2-7B’s (Mistral) performance increases from 11% to 36.7% on MATH and from 36% to 68.4% on GSM8K without training on any in-domain data. Further training MAmmoTH2 on public instruction tuning datasets yields MAmmoTH2-Plus, achieving state-of-the-art performance on several reasoning and chatbot benchmarks. Our work demonstrates how to harvest large-scale, high-quality instruction data without costly human annotation or GPT-4 distillation, providing a new paradigm for building better instruction tuning data. Xiang Yue, Tianyu Zheng, Ge Zhang 0009, Wenhu Chen |
NeurIPS | 4 |
| 2024 | Synthesizing Coherent Story with Auto-Regressive Latent Diffusion ModelsabstractConditioned diffusion models have demonstrated state-of-the-art text-to-image synthesis capacity. Recently, most works focus on synthesizing independent images; While for real-world applications, it is common and necessary to generate a series of coherent images for story-stelling. In this work, we mainly focus on story visualization and continuation tasks and propose AR-LDM, a latent diffusion model auto-regressively conditioned on history captions and generated images. Moreover, AR-LDM can generalize to new characters through adaptation. To our best knowledge, this is the first work successfully leveraging diffusion models for coherent visual story synthesizing. It also extends the text-conditioned method to multimodal conditioning. Quantitative results show that AR-LDM achieves SoTA FID scores on PororoSV, FlintstonesSV, and the adopted challenging dataset VIST containing natural images. Large-scale human evaluations show that AR-LDM has superior performance in terms of quality, relevance, and consistency. Code available at this https URL Xichen Pan, Pengda Qin, Hui Xue 0001, Wenhu Chen |
WACV | 5 |
| 2023 | QA Is the New KR: Question-Answer Pairs as Knowledge BasesabstractWe propose a new knowledge representation (KR) based on knowledge bases (KBs) derived from text, based on question generation and entity linking. We argue that the proposed type of KB has many of the key advantages of a traditional symbolic KB: in particular, it consists of small modular components, which can be combined compositionally to answer complex queries, including relational queries and queries involving ``multi-hop'' inferences. However, unlike a traditional KB, this information store is well-aligned with common user information needs. We present one such KB, called a QEDB, and give qualitative evidence that the atomic components are high-quality and meaningful, and that atomic components can be combined in ways similar to the triples in a symbolic KB. We also show experimentally that questions reflective of typical user questions are more easily answered with a QEDB than a symbolic KB. William W. Cohen, Wenhu Chen, Michiel de Jong, Nitish Gupta, Alessandro Presta, Patrick Verga, John Wieting |
AAAI | 2 |
| 2023 | Few-shot In-context Learning on Knowledge Base Question AnsweringabstractQuestion answering over knowledge bases is considered a difficult problem due to the challenge of generalizing to a wide variety of possible natural language questions.Additionally, the heterogeneity of knowledge base schema items between different knowledge bases often necessitates specialized training for different knowledge base question-answering (KBQA) datasets.To handle questions over diverse KBQA datasets with a unified trainingfree framework, we propose KB-BINDER, which for the first time enables few-shot incontext learning over KBQA tasks.Firstly, KB-BINDER leverages large language models like Codex to generate logical forms as the draft for a specific question by imitating a few demonstrations.Secondly, KB-BINDER grounds on the knowledge base to bind the generated draft to an executable one with BM25 score matching.The experimental results on four public heterogeneous KBQA datasets show that KB-BINDER can achieve a strong performance with only a few in-context demonstrations.Especially on GraphQA and 3-hop MetaQA, KB-BINDER can even outperform the state-of-the-art trained models.On GrailQA and WebQSP, our model is also on par with other fully-trained models.We believe KB-BINDER can serve as an important baseline for future research.Our code is available at Tianle Li, Xueguang Ma, Alex Zhuang, Yu Gu 0016, Yu Su 0001, Wenhu Chen |
ACL (1) | 6 |
| 2023 | Augmenting Pre-trained Language Models with QA-Memory for Open-Domain Question AnsweringabstractExisting state-of-the-art methods for opendomain question-answering (ODQA) use anopen book approach in which information is first retrieved from a large text corpus or knowledge base (KB) and then reasoned over to produce an answer.A recent alternative is to retrieve from a collection of previouslygenerated question-answer pairs; this has several practical advantages including being more memory and compute-efficient.Questionanswer pairs are also appealing in that they can be viewed as an intermediate between text and KB triples: like KB triples, they often concisely express a single relationship, but like text, have much higher coverage than traditional KBs.In this work, we describe a new QA system that augments a text-to-text model with a large memory of question-answer pairs, and a new pre-training task for the latent step of question retrieval.The pre-training task substantially simplifies training and greatly improves performance on smaller QA benchmarks.Unlike prior systems of this sort, our QA system can also answer multi-hop questions that do not explicitly appear in the collection of stored question-answer pairs. Wenhu Chen, Patrick Verga, Michiel de Jong, John Wieting, William W. Cohen |
EACL | 1 |
| 2023 | TheoremQA: A Theorem-driven Question Answering DatasetabstractThe recent LLMs like GPT-4 and PaLM-2 have made tremendous progress in solving fundamental math problems like GSM8K by achieving over 90% accuracy.However, their capabilities to solve more challenging math problems which require domain-specific knowledge (i.e.theorem) have yet to be investigated.In this paper, we introduce TheoremQA, the first theorem-driven question-answering dataset designed to evaluate AI models' capabilities to apply theorems to solve challenging science problems.TheoremQA is curated by domain experts containing 800 high-quality questions covering 350 theorems 1 from Math, Physics, EE&CS, and Finance.We evaluate a wide spectrum of 16 large language and code models with different prompting strategies like Chain-of-Thoughts and Program-of-Thoughts.We found that GPT-4's capabilities to solve these problems are unparalleled, achieving an accuracy of 51% with Program-of-Thoughts Prompting.All the existing open-sourced models are below 15%, barely surpassing the random-guess baseline.Given the diversity and broad coverage of TheoremQA, we believe it can be used as a better benchmark to evaluate LLMs' capabilities to solve challenging science problems. Wenhu Chen, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang 0003, Tony Xia |
EMNLP | 1 |
| 2023 | EDIS: Entity-Driven Image Search over Multimodal Web ContentabstractMaking image retrieval methods practical for real-world search applications requires significant progress in dataset scales, entity comprehension, and multimodal information fusion.In this work, we introduce Entity-Driven Image Search (EDIS), a challenging dataset for cross-modal image search in the news domain.EDIS consists of 1 million web images from actual search engine results and curated datasets, with each image paired with a textual description.Unlike datasets that assume a small set of single-modality candidates, EDIS reflects realworld web image search scenarios by including a million multimodal image-text pairs as candidates.EDIS encourages the development of retrieval models that simultaneously address cross-modal information fusion and matching.To achieve accurate ranking results, a model must: 1) understand named entities and events from text queries, 2) ground entities onto images or text descriptions, and 3) effectively fuse textual and visual representations.Our experimental results show that EDIS challenges stateof-the-art methods with dense entities and the large-scale candidate set.The ablation study also proves that fusing textual features with visual features is critical in improving retrieval results. Weixi Feng, Tsu-Jui Fu, Wenhu Chen, William Yang Wang |
EMNLP | 4 |
| 2023 | Re-Imagen: Retrieval-Augmented Text-to-Image Generator
Wenhu Chen, Hexiang Hu, Chitwan Saharia, William W. Cohen |
ICLR | 1 |
| 2023 | Attacking Open-domain Question Answering by Injecting MisinformationabstractLiangming Pan, Wenhu Chen, Min-Yen Kan, William Yang Wang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liangming Pan, Wenhu Chen, Min-Yen Kan, William Yang Wang |
IJCNLP (1) | 2 |
| 2023 | Subject-driven Text-to-Image Generation via Apprenticeship LearningabstractRecent text-to-image generation models like DreamBooth have made remarkable progress in generating highly customized images of a target subject, by fine-tuning an ``expert model'' for a given subject from a few examples.
However, this process is expensive, since a new expert model must be learned for each subject.
In this paper, we present SuTI, a Subject-driven Text-to-Image generator that replaces subject-specific fine tuning with {in-context} learning.
Given a few demonstrations of a new subject, SuTI can instantly generate novel renditions of the subject in different scenes, without any subject-specific optimization.
SuTI is powered by {apprenticeship learning}, where a single apprentice model is learned from data generated by a massive number of subject-specific expert models.
Specifically, we mine millions of image clusters from the Internet, each centered around a specific visual subject. We adopt these clusters to train a massive number of expert models, each specializing in a different subject. The apprentice model SuTI then learns to imitate the behavior of these fine-tuned experts.
SuTI can generate high-quality and customized subject-specific images 20x faster than optimization-based SoTA methods. On the challenging DreamBench and DreamBench-v2, our human evaluation shows that SuTI significantly outperforms existing models like InstructPix2Pix, Textual Inversion, Imagic, Prompt2Prompt, Re-Imagen and DreamBooth. Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, William W. Cohen |
NeurIPS | 1 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 18 |
| 2023 | MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image EditingabstractText-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop.However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise.Thus, they still require lots of manual tuning to produce desirable outcomes in practice.To address this issue, we introduce MagicBrush, the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing.MagicBrush comprises over 10K manually annotated triplets (source image, instruction, target image), which supports trainining large-scale text-guided image editing models.We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation.We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations.The results reveal the challenging nature of our dataset and the gap between current baselines and real-world editing needs. Kai Zhang 0033, Lingbo Mo, Wenhu Chen, Huan Sun 0001, Yu Su 0001 |
NeurIPS | 3 |
| 2023 | Knowledge Discovery from Unstructured Data in Financial Services (KDF) WorkshopabstractKnowledge discovery from unstructured data, including business documents, web content, and news articles, has been a key AI challenge for the financial services industry. Comprehending these corpora and discovering knowledge from them, which could be textual, tabular, or graphic, are the cornerstone of supporting business decisions in the financial services domain, where information retrieval and content analysis techniques are of fundamental importance. We propose a workshop on knowledge discovery from unstructured data in financial services at SIGIR 2023 to highlight the current and emerging opportunities, invite original research, and prompt success sharing between researchers. Sameena Shah, Xiaodan Zhu 0001, Wenhu Chen, Manling Li, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Yulong Pei, Akshat Gupta |
SIGIR | 3 |
| 2022 | MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextabstractWhile language Models store a massive amount of world knowledge implicitly in their parameters, even very large models often fail to encode information about rare entities and events, while incurring huge computational costs.Recently, retrieval-augmented models, such as REALM, RAG, and RETRO, have incorporated world knowledge into language generation by leveraging an external nonparametric index and have demonstrated impressive performance with constrained model sizes.However, these methods are restricted to retrieving only textual knowledge, neglecting the ubiquitous amount of knowledge in other modalities like images -much of which contains information not covered by any text.To address this limitation, we propose the first Multimodal Retrieval-Augmented Transformer (MuRAG), which accesses an external non-parametric multimodal memory to augment language generation.MuRAG is pretrained with a mixture of large-scale imagetext and text-only corpora using a joint contrastive and generative loss.We perform experiments on two different datasets that require retrieving and reasoning over both images and text to answer a given query: We-bQA, and MultimodalQA.Our results show that MuRAG achieves state-of-the-art accuracy, outperforming existing models by 10-20% absolute on both datasets and under both distractor and full-wiki settings. Wenhu Chen, Hexiang Hu, Xi Chen 0071, Patrick Verga, William W. Cohen |
EMNLP | 1 |
| 2021 | A Systematic Investigation of KB-Text Embedding Alignment at ScaleabstractVardaan Pahuja, Yu Gu, Wenhu Chen, Mehdi Bahrami, Lei Liu, Wei-Peng Chen, Yu Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Vardaan Pahuja, Yu Gu 0016, Wenhu Chen, Mehdi Bahrami, Wei-Peng Chen, Yu Su 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | FinQA: A Dataset of Numerical Reasoning over Financial DataabstractZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhiyu Chen 0002, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao 'Kenneth' Huang, Bryan R. Routledge, William Yang Wang |
EMNLP (1) | 2 |
| 2021 | Open Question Answering over Tables and Text
Wenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang, William W. Cohen |
ICLR | 1 |
| 2021 | Unsupervised Multi-hop Question Answering by Question GenerationabstractLiangming Pan, Wenhu Chen, Wenhan Xiong, Min-Yen Kan, William Yang Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Liangming Pan, Wenhu Chen, Wenhan Xiong, Min-Yen Kan, William Yang Wang |
NAACL-HLT | 2 |
| 2021 | Local Explanation of Dialogue Response GenerationabstractIn comparison to the interpretation of classification models, the explanation of sequence generation models is also an important problem, however it has seen little attention. In this work, we study model-agnostic explanations of a representative text generation task -- dialogue response generation. Dialog response generation is challenging with its open-ended sentences and multiple acceptable responses. To gain insights into the reasoning process of a generation model, we propose a new method, local explanation of response generation (LERG) that regards the explanations as the mutual interaction of segments in input and output sentences. LERG views the sequence prediction as uncertainty estimation of a human response and then creates explanations by perturbing the input and calculating the certainty change over the human response. We show that LERG adheres to desired properties of explanations for text generation including unbiased approximation, consistency and cause identification. Empirically, our results show that our method consistently improves other widely used methods on proposed automatic- and human- evaluation metrics for this new task by $4.4$-$12.8$\%. Our analysis demonstrates that LERG can extract both explicit and implicit relations between input and output segments. Yi-Lin Tuan, Connor Pryor, Wenhu Chen, Lise Getoor, William Yang Wang |
NeurIPS | 3 |
| 2021 | Counterfactual Maximum Likelihood Estimation for Training Deep NetworksabstractAlthough deep learning models have driven state-of-the-art performance on a wide array of tasks, they are prone to spurious correlations that should not be learned as predictive clues. To mitigate this problem, we propose a causality-based training framework to reduce the spurious correlations caused by observed confounders. We give theoretical analysis on the underlying general Structural Causal Model (SCM) and propose to perform Maximum Likelihood Estimation (MLE) on the interventional distribution instead of the observational distribution, namely Counterfactual Maximum Likelihood Estimation (CMLE). As the interventional distribution, in general, is hidden from the observational data, we then derive two different upper bounds of the expected negative log-likelihood and propose two general algorithms, Implicit CMLE and Explicit CMLE, for causal predictions of deep learning models using observational data. We conduct experiments on both simulated data and two real-world tasks: Natural Language Inference (NLI) and Image Captioning. The results show that CMLE methods outperform the regular MLE method in terms of out-of-domain generalization performance and reducing spurious correlations, while maintaining comparable performance on the regular evaluations. Xinyi Wang 0003, Wenhu Chen, Michael Saxon, William Yang Wang |
NeurIPS | 2 |
| 2021 | Meta Module Network for Compositional Visual ReasoningabstractNeural Module Network (NMN) exhibits strong interpretability and compositionality thanks to its handcrafted neural modules with explicit multi-hop reasoning capability. However, most NMNs suffer from two critical draw-backs: 1) scalability: customized module for specific function renders it impractical when scaling up to a larger set of functions in complex tasks; 2) generalizability: rigid pre-defined module inventory makes it difficult to generalize to unseen functions in new tasks/domains. To design a more powerful NMN architecture for practical use, we propose Meta Module Network (MMN) centered on a novel meta module, which can take in function recipes and morph into diverse instance modules dynamically. The instance modules are then woven into an execution graph for complex visual reasoning, inheriting the strong explainability and compositionality of NMN. With such a flexible instantiation mechanism, the parameters of instance modules are inherited from the central meta module, retaining the same model complexity as the function set grows, which promises better scalability. Meanwhile, as functions are encoded into the embedding space, unseen functions can be readily represented based on its structural similarity with previously observed ones, which ensures better generalizability. Experiments on GQA and CLEVR datasets validate the superiority of MMN over state-of-the-art NMN designs. Synthetic experiments on held-out unseen functions from GQA dataset also demonstrate the strong generalizability of MMN. Our code and model are released in Github1. Wenhu Chen, Zhe Gan, Yu Cheng 0001, William Yang Wang, Jingjing Liu 0001 |
WACV | 1 |
| 2020 | Generative Adversarial Zero-Shot Relational Learning for Knowledge GraphsabstractLarge-scale knowledge graphs (KGs) are shown to become more important in current information systems. To expand the coverage of KGs, previous studies on knowledge graph completion need to collect adequate training instances for newly-added relations. In this paper, we consider a novel formulation, zero-shot learning, to free this cumbersome curation. For newly-added relations, we attempt to learn their semantic features from their text descriptions and hence recognize the facts of unseen relations with no examples being seen. For this purpose, we leverage Generative Adversarial Networks (GANs) to establish the connection between text and knowledge graph domain: The generator learns to generate the reasonable relation embeddings merely with noisy text descriptions. Under this setting, zero-shot learning is naturally converted to a traditional supervised classification task. Empirically, our method is model-agnostic that could be potentially applied to any version of KG embeddings, and consistently yields performance improvements on NELL and Wiki dataset. Pengda Qin, Xin Wang 0061, Wenhu Chen, Chunyun Zhang, Weiran Xu, William Yang Wang |
AAAI | 3 |
| 2020 | Logical Natural Language Generation from Open-Domain TablesabstractNeural natural language generation (NLG) models have recently shown remarkable progress in fluency and coherence.However, existing studies on neural NLG are primarily focused on surface-level realizations with limited emphasis on logical inference, an important aspect of human thinking and language.In this paper, we suggest a new NLG task where a model is tasked with generating natural language statements that can be logically entailed by the facts in an open-domain semi-structured table.To facilitate the study of the proposed logical NLG problem, we use the existing Tab-Fact dataset (Chen et al., 2019) featured with a wide range of logical/symbolic inferences as our testbed, and propose new automatic metrics to evaluate the fidelity of generation models w.r.t.logical inference.The new task poses challenges to the existing monotonic generation frameworks due to the mismatch between sequence order and logical order.In our experiments, we comprehensively survey different generation architectures (LSTM, Transformer, Pre-Trained LM) trained with different algorithms (RL, Adversarial Training, Coarse-to-Fine) on the dataset and made following observations: 1) Pre-Trained LM can significantly boost both the fluency and logical fidelity metrics, 2) RL and Adversarial Training are trading fluency for fidelity, 3) Coarse-to-Fine generation can help partially alleviate the fidelity issue while maintaining high language fluency. Wenhu Chen, Jianshu Chen, Yu Su 0001, Zhiyu Chen 0002, William Yang Wang |
ACL | 1 |
| 2020 | Few-Shot NLG with Pre-Trained Language ModelabstractNeural-based end-to-end approaches to natural language generation (NLG) from structured data or knowledge are data-hungry, making their adoption for real-world applications difficult with limited data. In this work, we propose the new task of few-shot natural language generation. Motivated by how humans tend to summarize tabular data, we propose a simple yet effective approach and show that it not only demonstrates strong performance but also provides good generalization across domains. The design of the model architecture is based on two aspects: content selection from input data and language modeling to compose coherent sentences, which can be acquired from prior knowledge. With just 200 training examples, across multiple domains, we show that our approach achieves very reasonable performances and outperforms the strongest baseline by an average of over 8.0 BLEU points improvement. Our code and data can be found at https://github.com/czyssrs/Few-Shot-NLG Zhiyu Chen 0002, Harini Eavani, Wenhu Chen, Yinyin Liu, William Yang Wang |
ACL | 3 |
| 2020 | Violin: A Large-Scale Dataset for Video-and-Language InferenceabstractWe introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the video content, a model needs to infer whether the hypothesis is entailed or contradicted by the given video clip. A new large-scale dataset, named Violin (VIdeO-and-Language INference), is introduced for this task, which consists of 95,322 video-hypothesis pairs from 15,887 video clips, spanning over 582 hours of video. These video clips contain rich content with diverse temporal dynamics, event shifts, and people interactions, collected from two sources: (i) popular TV shows, and (ii) movie clips from YouTube channels. In order to address our new multimodal inference task, a model is required to possess sophisticated reasoning skills, from surface-level grounding (e.g., identifying objects and characters in the video) to in-depth commonsense reasoning (e.g., inferring causal relations of events in the video). We present a detailed analysis of the dataset and an extensive evaluation over many strong baselines, providing valuable insights on the challenges of this new task. Jingzhou Liu, Wenhu Chen, Yu Cheng 0001, Zhe Gan, Licheng Yu, Yiming Yang 0002, Jingjing Liu 0001 |
CVPR | 2 |
| 2020 | KGPT: Knowledge-Grounded Pre-Training for Data-to-Text GenerationabstractData-to-text generation has recently attracted substantial interests due to its wide applications. Existing methods have shown impressive performance on an array of tasks. However, they rely on a significant amount of labeled data for each task, which is costly to acquire and thus limits their application to new tasks and domains. In this paper, we propose to leverage pre-training and transfer learning to address this issue. We propose a knowledge-grounded pre-training (KGPT), which consists of two parts, 1) a general knowledge-grounded generation model to generate knowledge-enriched text. 2) a pre-training paradigm on a massive knowledge-grounded text corpus crawled from the web. The pre-trained model can be fine-tuned on various data-to-text generation tasks to generate task-specific text. We adopt three settings, namely fully-supervised, zero-shot, few-shot to evaluate its effectiveness. Under the fully-supervised setting, our model can achieve remarkable gains over the known baselines. Under zero-shot setting, our model without seeing any examples achieves over 30 ROUGE-L on WebNLG while all other baselines fail. Under the few-shot setting, our model only needs about one-fifteenth as many labeled examples to achieve the same level of performance as baseline models. These experiments consistently prove the strong generalization ability of our proposed framework. Wenhu Chen, Yu Su 0001, Xifeng Yan, William Yang Wang |
EMNLP (1) | 1 |
| 2020 | TabFact: A Large-scale Dataset for Table-based Fact Verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 0002, Hong Wang 0023, Xiyou Zhou, William Yang Wang |
ICLR | 1 |
| 2019 | Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-AttentionabstractSemantically controlled neural response generation on limited-domain has achieved great performance.However, moving towards multi-domain large-scale scenarios are shown to be difficult because the possible combinations of semantic inputs grow exponentially with the number of domains.To alleviate such scalability issue, we exploit the structure of dialog acts to build a multi-layer hierarchical graph, where each act is represented as a rootto-leaf route on the graph.Then, we incorporate such graph structure prior as an inductive bias to build a hierarchical disentangled self-attention network, where we disentangle attention heads to model designated nodes on the dialog act graph.By activating different (disentangled) heads at each layer, combinatorially many dialog act semantics can be modeled to control the neural response generation.On the large-scale Multi-Domain-WOZ dataset, our model can yield a significant improvement over the baselines on various automatic and human evaluation metrics. Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan, William Yang Wang |
ACL (1) | 1 |
| 2019 | Global Textual Relation Embedding for Relational UnderstandingabstractPre-trained embeddings such as word embeddings and sentence embeddings are fundamental tools facilitating a wide range of downstream NLP tasks.In this work, we investigate how to learn a general-purpose embedding of textual relations, defined as the shortest dependency path between entities.Textual relation embedding provides a level of knowledge between word/phrase level and sentence level, and we show that it can facilitate downstream tasks requiring relational understanding of the text.To learn such an embedding, we create the largest distant supervision dataset by linking the entire English ClueWeb09 corpus to Freebase.We use global co-occurrence statistics between textual and knowledge base relations as the supervision signal to train the embedding.Evaluation on two relational understanding tasks demonstrates the usefulness of the learned textual relation embedding. Zhiyu Chen 0002, Hanwen Zha, Honglei Liu 0001, Wenhu Chen, Xifeng Yan, Yu Su 0001 |
ACL (1) | 4 |
| 2019 | Interpreting and Improving Deep Neural SLU Models via Vocabulary Importance
Yilin Shen, Wenhu Chen, Hongxia Jin |
INTERSPEECH | 2 |
| 2019 | Mining Algorithm Roadmap in Scientific PublicationsabstractThe number of scientific publications is ever increasing. The long time to digest a scientific paper posts great challenges on the number of papers people can read, which impedes a quick grasp of major activities in new research areas especially for intelligence analysts and novice researchers. To accelerate such a process, we first define a new problem called mining algorithm roadmap in scientific publications, and then propose a new weakly supervised method to build the roadmap. The algorithm roadmap describes evolutionary relation between different algorithms, and sketches the undergoing research and the dynamics of the area. It is a tool for analysts and researchers to locate the successors and families of algorithms when analyzing and surveying a research field. We first propose abbreviated words as candidates for algorithms and then use tables as weak supervision to extract these candidates and labels. Next we propose a new method called Cross-sentence Attention NeTwork for cOmparative Relation (CANTOR) to extract comparative algorithms from text. Finally, we derive order for individual algorithm pairs with time and frequency to construct the algorithm roadmap. Through comprehensive experiments, our proposed algorithm shows its superiority over the baseline methods on the proposed task. Hanwen Zha, Wenhu Chen, Keqian Li, Xifeng Yan |
KDD | 2 |
| 2019 | Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series ForecastingabstractTime series forecasting is an important problem across many domains, including predictions of solar plant energy output, electricity consumption, and traffic jam situation. In this paper, we propose to tackle such forecasting problem with Transformer. Although impressed by its performance in our preliminary study, we found its two major weaknesses: (1) locality-agnostics: the point-wise dot- product self-attention in canonical Transformer architecture is insensitive to local context, which can make the model prone to anomalies in time series; (2) memory bottleneck: space complexity of canonical Transformer grows quadratically with sequence length L, making directly modeling long time series infeasible. In order to solve these two issues, we first propose convolutional self-attention by producing queries and keys with causal convolution so that local context can be better incorporated into attention mechanism. Then, we propose LogSparse Transformer with only O(L(log L)^2) memory cost, improving forecasting accuracy for time series with fine granularity and strong long-term dependencies under constrained memory budget. Our experiments on both synthetic data and real- world datasets show that it compares favorably to the state-of-the-art. Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang 0003, Xifeng Yan |
NeurIPS | 5 |
| 2018 | No Metrics Are Perfect: Adversarial Reward Learning for Visual StorytellingabstractThough impressive results have been achieved in visual captioning, the task of generating abstract stories from photo streams is still a little-tapped problem.Different from captions, stories have more expressive language styles and contain many imaginary concepts that do not appear in the images.Thus it poses challenges to behavioral cloning algorithms.Furthermore, due to the limitations of automatic metrics on evaluating story quality, reinforcement learning methods with hand-crafted rewards also face difficulties in gaining an overall performance boost.Therefore, we propose an Adversarial REward Learning (AREL) framework to learn an implicit reward function from human demonstrations, and then optimize policy search with the learned reward function.Though automatic evaluation indicates slight performance boost over state-of-the-art (SOTA) methods in cloning expert behaviors, human evaluation shows that our approach achieves significant improvement in generating more human-like stories than SOTA systems.Code will be made available here 1 . Xin Wang 0061, Wenhu Chen, Yuan-Fang Wang, William Yang Wang |
ACL (1) | 2 |
| 2018 | Triangular Architecture for Rare Language TranslationabstractNeural Machine Translation (NMT) performs poor on the low-resource language pair (X, Z), especially when Z is a rare language.By introducing another rich language Y , we propose a novel triangular training architecture (TA-NMT) to leverage bilingual data (Y, Z) (may be small) and (X, Y ) (can be rich) to improve the translation performance of lowresource pairs.In this triangular architecture, Z is taken as the intermediate latent variable, and translation models of Z are jointly optimized with a unified bidirectional EM algorithm under the goal of maximizing the translation likelihood of (X, Y ).Empirical results demonstrate that our method significantly improves the translation quality of rare languages on MultiUN and IWSLT2012 datasets, and achieves even better performance combining back-translation methods. Shuo Ren 0002, Wenhu Chen, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL (1) | 2 |
| 2018 | Video Captioning via Hierarchical Reinforcement LearningabstractVideo captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short video, it is still very challenging to caption a video containing multiple fine-grained actions with a detailed description. This paper aims to address the challenge by proposing a novel hierarchical reinforcement learning framework for video captioning, where a high-level Manager module learns to design sub-goals and a low-level Worker module recognizes the primitive actions to fulfill the sub-goal. With this compositional framework to reinforce video captioning at different levels, our approach significantly outperforms all the baseline methods on a newly introduced large-scale dataset for fine-grained video captioning. Furthermore, our non-ensemble model has already achieved the state-of-the-art results on the widely-used MSR-VTT dataset. Xin Wang 0061, Wenhu Chen, Jiawei Wu 0003, Yuan-Fang Wang, William Yang Wang |
CVPR | 2 |
| 2018 | XL-NBT: A Cross-lingual Neural Belief Tracking FrameworkabstractTask-oriented dialog systems are becoming pervasive, and many companies heavily rely on them to complement human agents for customer service in call centers.With globalization, the need for providing cross-lingual customer support becomes more urgent than ever.However, cross-lingual support poses great challenges-it requires a large amount of additional annotated data from native speakers.In order to bypass the expensive human annotation and achieve the first step towards the ultimate goal of building a universal dialog system, we set out to build a cross-lingual state tracking framework.Specifically, we assume that there exists a source language with dialog belief tracking annotations while the target languages have no annotated dialog data of any form.Then, we pre-train a state tracker for the source language as a teacher, which is able to exploit easy-to-access parallel data.We then distill and transfer its own knowledge to the student state tracker in target languages.We specifically discuss two types of common parallel resources: bilingual corpus and bilingual dictionary, and design different transfer learning strategies accordingly.Experimentally, we successfully use English state tracker as the teacher to transfer its knowledge to both Italian and German trackers and achieve promising results. Wenhu Chen, Jianshu Chen, Yu Su 0001, Xin Wang 0061, Dong Yu 0001, Xifeng Yan, William Yang Wang |
EMNLP | 1 |
| 2018 | Generative Bridging Network for Neural Sequence PredictionabstractWenhu Chen, Guanlin Li, Shuo Ren, Shujie Liu, Zhirui Zhang, Mu Li, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Wenhu Chen, Shuo Ren 0002, Shujie Liu 0001, Zhirui Zhang, Mu Li 0001, Ming Zhou 0001 |
NAACL-HLT | 1 |
| 2018 | Variational Knowledge Graph ReasoningabstractWenhu Chen, Wenhan Xiong, Xifeng Yan, William Yang Wang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Wenhu Chen, Wenhan Xiong, Xifeng Yan, William Yang Wang |
NAACL-HLT | 1 |