Yuwei Fang

dblp:227/2871 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-subject Open-set Personalization in Video Generation
abstract
Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, require time-consuming optimization per subject, or support only a single subject. We present Video Alchemist—a video model with built-in multi-subject, openset personalization capabilities for both foreground objects and background, eliminating the need for time-consuming test-time optimization. Our model is built on a new Diffusion Transformer module that fuses each conditional reference image and its corresponding subject-level text prompt with cross-attention layers. Developing such a large model presents two main challenges: dataset and evaluation. First, as paired datasets of reference images and videos are extremely hard to collect, we sample selected video frames as reference images and synthesize a clip of the target video. However, while models can easily denoise training videos given reference frames, they fail to generalize to new contexts. To mitigate this issue, we design a new automatic data construction pipeline with extensive image augmentations. Second, evaluating open-set video personalization is a challenge in itself. To address this, we introduce a personalization benchmark that focuses on accurate subject fidelity and supports diverse personalization scenarios. Finally, our extensive experiments show that our method significantly outperforms existing personalization methods in both quantitative and qualitative evaluations.
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang, Kwot Sin Lee, Ivan Skorokhodov, Kfir Aberman, Jun-Yan Zhu, Ming-Hsuan Yang 0001, Sergey Tulyakov
CVPR4
2025 Mind the Time: Temporally-Controlled Multi-Event Video Generation
abstract
Real-world videos consist of sequences of events. Generating such sequences with precise temporal control is infeasible with existing video generators that rely on a single paragraph of text as input. When tasked with generating multiple events described using a single prompt, such methods often ignore some of the events or fail to arrange them in the correct order. To address this limitation, we present MinT, a multi-event video generator with temporal control. Our key insight is to bind each event to a specific period in the generated video, which allows the model to focus on one event at a time. To enable time-aware interactions between event captions and video tokens, we design a time-based positional encoding method, dubbed ReRoPE. This encoding helps to guide the cross-attention operation. By fine-tuning a pre-trained video diffusion transformer on temporally grounded data, our approach produces coherent videos with smoothly connected events. For the first time in the literature, our model offers control over the timing of events in generated videos. Extensive experiments demonstrate that MinT outperforms existing commercial and open-source models by a large margin. Additional results and details are available at our project page.
Ziyi Wu 0002, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Yuwei Fang, Varnith Chordia, Igor Gilitschenski, Sergey Tulyakov
CVPR5
2024 Evaluating Very Long-Term Conversational Memory of LLM Agents
abstract
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Adyasha Maharana, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang
ACL (1)6
2024 PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning
abstract
Zhihan Zhang, Dong-Ho Lee, Yuwei Fang, Wenhao Yu, Mengzhao Jia, Meng Jiang, Francesco Barbieri. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhihan Zhang 0001, Yuwei Fang, Wenhao Yu 0002, Mengzhao Jia, Meng Jiang 0001, Francesco Barbieri
ACL (1)3
2024 Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
abstract
The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more time-consuming, as it requires an annotator to watch an entire video. Second, videos have a temporal dimension, consisting of several scenes stacked together, and showing multiple actions. Accordingly, to establish a video dataset with high-quality captions, we propose an automatic approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames. Specifically, we curate 3.8M high-resolution videos from the publicly available HD-VILA-100M dataset. We then split them into semantically consistent video clips, and apply multiple cross-modality teacher models to obtain captions for each video. Next, we finetune a retrieval model on a small subset where the best caption of each video is manually selected and then employ the model in the whole dataset to select the best caption as the annotation. In this way, we get 70M videos paired with high-quality text captions. We dub the dataset as Panda-70M. We show the value of the proposed dataset on three downstream tasks: video captioning, video and text retrieval, and text-driven video generation. The models trained on the proposed data score substantially better on the majority of metrics across all the tasks.
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee 0001, Jian Ren 0005, Ming-Hsuan Yang 0001, Sergey Tulyakov
CVPR7
2024 Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
abstract
Contemporary models for generating images show remarkable quality and versatility. Swayed by these advantages, the research community repurposes them to generate videos. Since video content is highly redundant, we argue that naively bringing advances of image models to the video generation domain reduces motion fidelity, visual quality and impairs scalability. In this work, we build Snap Video, a video-first model that systematically addresses these challenges. To do that, we first extend the EDM framework to take into account spatially and temporally redundant pixels and naturally support video generation. Second, we show that a U-Net—a workhorse behind image generation—scales poorly when generating videos, requiring significant computational overhead. Hence, we propose a new transformer-based architecture that trains 3.31 times faster than U-Nets (and is ∼4.5 faster at inference). This allows us to efficiently train a text-to-video model with billions of parameters for the first time, reach state-of-the-art results on a number of benchmarks, and generate videos with substantially higher quality, temporal consistency, and motion complexity. The user studies showed that our model was favored by a large margin over the most recent methods.
Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Ekaterina Deyneka, Tsai-Shien Chen, Anil Kag, Yuwei Fang, Aleksei Stoliar, Elisa Ricci 0001, Jian Ren 0005, Sergey Tulyakov
CVPR7
2024 VIMI: Grounding Video Generation through Multi-modal Instruction
abstract
Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov
EMNLP1
2024 MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
Kuan-Chieh Wang, Daniil Ostashev, Yuwei Fang, Sergey Tulyakov, Kfir Aberman
SIGGRAPH Asia3
2023 i-Code: An Integrative and Composable Multimodal Learning Framework
abstract
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.
Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001
AAAI2
2023 Unifying Vision, Text, and Layout for Universal Document Processing
abstract
We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to model image, text, and layout modalities with one uniform representation. With a novel Vision-Text-Layout Transformer, UDOP unifies pretraining and multi-domain downstream tasks into a prompt-based sequence generation scheme. UDOP is pretrained on both large-scale unlabeled document corpora using innovative self-supervised objectives and diverse labeled data. UDOP also learns to generate document images from text and layout modalities via masked image reconstruction. To the best of our knowledge, this is the first time in the field of document AI that one model simultaneously achieves high-quality neural document editing and content customization. Our method sets the state-of-the-art on 8 Document AI tasks, e.g., document understanding and QA, across diverse data domains like finance reports, academic papers, and web-sites. UDOP ranks first on the leaderboard of the Document Understanding Benchmark.11Code and models: https://github.com/microsoft/i-Code/tree/main/i-Code-Doc
Zineng Tang, Ziyi Yang 0011, Yuwei Fang, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001, Cha Zhang, Mohit Bansal
CVPR4
2023 MACSum: Controllable Summarization with Mixed Attributes
abstract
Abstract Controllable summarization allows users to generate customized summaries with specified attributes. However, due to the lack of designated annotations of controlled summaries, existing work has to craft pseudo datasets by adapting generic summarization benchmarks. Furthermore, most research focuses on controlling single attributes individually (e.g., a short summary or a highly abstractive summary) rather than controlling a mix of attributes together (e.g., a short and highly abstractive summary). In this paper, we propose MACSum, the first human-annotated summarization dataset for controlling mixed attributes. It contains source texts from two domains, news articles and dialogues, with human-annotated summaries controlled by five designed attributes (Length, Extractiveness, Specificity, Topic, and Speaker). We propose two simple and effective parameter-efficient approaches for the new task of mixed controllable summarization based on hard prompt tuning and soft prefix tuning. Results and analysis demonstrate that hard prompt models yield the best performance on most metrics and human evaluations. However, mixed-attribute control is still challenging for summarization tasks. Our dataset and code are available at https://github.com/psunlpgroup/MACSum.
Yusen Zhang 0001, Yang Liu 0124, Ziyi Yang 0011, Yuwei Fang, Yulong Chen 0001, Dragomir R. Radev, Chenguang Zhu 0001, Michael Zeng 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics4
2022 RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling
abstract
Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. We propose a framework that alleviates this data constraint by jointly training a grounded generator and document retriever on the language model signal. The model learns to reward retrieval of the documents with the highest utility in generation, and attentively combines them using a Mixture-of-Experts (MoE) ensemble to generate follow-on text. We demonstrate that both generator and retriever can take advantage of this joint training and work synergistically to produce more informative and relevant text in both prose and dialogue generation.
Yizhe Zhang 0002, Xiang Gao 0011, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao 0001, William B. Dolan
AAAI4
2022 Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data
abstract
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu 0124, Ruochen Xu, Chenguang Zhu 0001, Michael Zeng 0001
ACL (1)3
2022 KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question Answering
abstract
Donghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Donghan Yu, Chenguang Zhu 0001, Yuwei Fang, Wenhao Yu 0002, Shuohang Wang, Yichong Xu, Xiang Ren 0001, Yiming Yang 0002, Michael Zeng 0001
ACL (1)3
2022 Retrieval Augmentation for Commonsense Reasoning: A Unified Approach
abstract
A common thread of retrieval-augmented methods in the existing literature focuses on retrieving encyclopedic knowledge, such as Wikipedia, which facilitates well-defined entity and relation spaces that can be modeled.However, applying such methods to commonsense reasoning tasks faces two unique challenges, i.e., the lack of a general large-scale corpus for retrieval and a corresponding effective commonsense retriever.In this paper, we systematically investigate how to leverage commonsense knowledge retrieval to improve commonsense reasoning tasks.We proposed a unified framework of Retrieval-Augmented Commonsense reasoning (called RACO), including a newly constructed commonsense corpus with over 20 million documents and novel strategies for training a commonsense retriever.We conducted experiments on four different commonsense reasoning tasks.Extensive evaluation results showed that our proposed RACO can significantly outperform other knowledgeenhanced method counterparts, achieving new SoTA performance on the CommonGen 1 and CREAK 2 leaderboards.Our code is available at https://github.com/wyu97/RACo.
Wenhao Yu 0002, Chenguang Zhu 0001, Zhihan Zhang 0001, Shuohang Wang, Zhuosheng Zhang 0001, Yuwei Fang, Meng Jiang 0001
EMNLP6
2022 GAIM: Graph-aware Feature Interactional Model for Spam Movie Review Detection
abstract
Nowadays, more and more people decide whether to watch a certain movie by reading online movie reviews. Driven by large commercial interest, a growing number of spammers try to manipulate the online word-of-mouth of movies by publishing spam reviews. This has severely destroyed the credibility of movie review platforms and affected the healthy development of the lm industry. However, there is little research on the detection of spam movie reviews. Meanwhile, there are still great challenges for spam movie review detection, such as the lack of publicly available datasets, insufficient features, and low-performance detection models. In this paper, we firstly construct a dataset for spam movie reviews with our proposed labeling strategy and release it publicly. Secondly, we design 28 features to detect spam movie reviews, 13 of which are completely new features. Thirdly, in order to mine the characteristics of collusive attack behavior deeply, we propose a Graph-aware Feature Interactional Model (GAIM), which combines TextCNN, MLP (Multilayer Perceptron), and GAT (Graph Attention Network). After performance evaluation, the experimental results show that GAIM is more effective than the state-of-the-art baselines with an F1-score of 91.88% for spam movie review.
Lei Zhang 0103, Xueqiang Song, Yuwei Fang, Dong Li 0051, Haizhou Wang 0001
ICPR4
2022 A Novel Chinese Sarcasm Detection Model Based on Retrospective Reader
Lei Zhang 0103, Xueqiang Song, Yuwei Fang, Dong Li 0051, Haizhou Wang 0001
MMM (2)4
2021 FILTER: An Enhanced Fusion Method for Cross-lingual Language Understanding
abstract
Large-scale cross-lingual language models (LM), such as mBERT, Unicoder and XLM, have achieved great success in cross-lingual representation learning. However, when applied to zero-shot cross-lingual transfer tasks, most existing methods use only single-language input for LM finetuning, without leveraging the intrinsic cross-lingual alignment between different languages that proves essential for multilingual tasks. In this paper, we propose FILTER, an enhanced fusion method that takes cross-lingual data as input for XLM finetuning. Specifically, FILTER first encodes text input in the source language and its translation in the target language independently in the shallow layers, then performs cross-language fusion to extract multilingual knowledge in the intermediate layers, and finally performs further language-specific encoding. During inference, the model makes predictions based on the text input in the target language and its translation in the source language. For simple tasks such as classification, translated text in the target language shares the same label as the source language. However, this shared label becomes less accurate or even unavailable for more complex tasks such as question answering, NER and POS tagging. To tackle this issue, we further propose an additional KL-divergence self-teaching loss for model training, based on auto-generated soft pseudo-labels for translated text in the target language. Extensive experiments demonstrate that FILTER achieves new state of the art on two challenging multilingual multi-task benchmarks, XTREME and XGLUE.
Yuwei Fang, Shuohang Wang, Zhe Gan, Jingjing Liu 0001
AAAI1
2021 LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval
abstract
Siqi Sun, Yen-Chun Chen, Linjie Li, Shuohang Wang, Yuwei Fang, Jingjing Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yen-Chun Chen 0001, Shuohang Wang, Yuwei Fang, Jingjing Liu 0001
NAACL-HLT5
2020 Hierarchical Graph Network for Multi-hop Question Answering
abstract
In this paper, we present Hierarchical Graph Network (HGN) for multi-hop question answering.To aggregate clues from scattered texts across multiple paragraphs, a hierarchical graph is created by constructing nodes on different levels of granularity (questions, paragraphs, sentences, entities), the representations of which are initialized with pre-trained contextual encoders.Given this hierarchical graph, the initial node representations are updated through graph propagation, and multihop reasoning is performed via traversing through the graph edges for each subsequent sub-task (e.g., paragraph selection, supporting facts extraction, answer prediction).By weaving heterogeneous nodes into an integral unified graph, this hierarchical differentiation of node granularity enables HGN to support different question answering sub-tasks simultaneously.Experiments on the HotpotQA benchmark demonstrate that the proposed model achieves new state of the art, outperforming existing multi-hop QA approaches. 1
Yuwei Fang, Zhe Gan, Rohit Pillai, Shuohang Wang, Jingjing Liu 0001
EMNLP (1)1
2020 Contrastive Distillation on Intermediate Representations for Language Model Compression
abstract
Existing language model compression methods mostly use a simple L 2 loss to distill knowledge in the intermediate representations of a large BERT model to a smaller one.Although widely used, this objective by design assumes that all the dimensions of hidden representations are independent, failing to capture important structural knowledge in the intermediate layers of the teacher network.To achieve better distillation efficacy, we propose Contrastive Distillation on Intermediate Representations (CODIR), a principled knowledge distillation framework where the student is trained to distill knowledge through intermediate layers of the teacher via a contrastive objective.By learning to distinguish positive sample from a large set of negative samples, CoDIR facilitates the student's exploitation of rich information in teacher's hidden layers.CoDIR can be readily applied to compress large-scale language models in both pretraining and finetuning stages, and achieves superb performance on the GLUE benchmark, outperforming state-of-the-art compression methods. 1
Zhe Gan, Yuwei Fang, Yu Cheng 0001, Shuohang Wang, Jingjing Liu 0001
EMNLP (1)3
2020 Cross-Thought for Sentence Encoder Pre-training
abstract
In this paper, we propose Cross-Thought, a novel approach to pre-training sequence encoder, which is instrumental in building reusable sequence embeddings for large-scale NLP tasks such as question answering.Instead of using the original signals of full sentences, we train a Transformer-based sequence encoder over a large set of short sequences, which allows the model to automatically select the most useful information for predicting masked words.Experiments on question answering and textual entailment tasks demonstrate that our pre-trained encoder can outperform state-of-the-art encoders trained with continuous sentence signals as well as traditional masked language modeling baselines.Our proposed approach also achieves new state of the art on HotpotQA (full-wiki setting) by improving intermediate information retrieval performance.1
Shuohang Wang, Yuwei Fang, Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001, Jing Jiang 0001
EMNLP (1)2