EDBT 2026 Demo / reviewers in the wild / expert
Ruiyi Zhang 0002
dblp:08/7975-2
· DBLP profile ↗
56ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0002-8157-6364ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 8 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents ArchiveabstractThe opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and public policy. Analyzing how these interconnected systems simultaneously failed to protect public health requires innovative analytic approaches for exploring the vast amounts of data and documents disclosed in the UCSF-JHU Opioid Industry Documents Archive (OIDA). The complexity, multimodal nature, and specialized characteristics of these healthcare-related legal and corporate documents necessitate more advanced methods and models tailored to specific data types and detailed annotations, ensuring the precision and professionalism in the analysis. In this paper, we tackle this challenge by organizing the original dataset according to document attributes and constructing a benchmark with 400k training documents and 10k for testing. From each document, we extract rich multimodal information—including textual content, visual elements, and layout structures—to capture a comprehensive range of features. Using multiple AI models, we then generate a large-scale dataset comprising 360k training QA pairs and 10k testing QA pairs. Building on this foundation, we develop domain-specific multimodal Large Language Models (LLMs) and explore the impact of multimodal inputs on task performance. To further enhance response accuracy, we incorporate historical QA pairs as contextual grounding for answering current queries. Additionally, we incorporate page references within the answers and introduce an importance-based page classifier, further improving the precision and relevance of the information provided. Preliminary results indicate the improvements with our AI assistant in document information extraction and question-answering tasks. Xuan Shen, Brian Wingenroth, Zichao Wang 0001, Jason Kuen, Wanrong Zhu, Ruiyi Zhang 0002, Lichun Ma, Anqi Liu 0001, Tong Sun 0005, Kevin S. Hawkins, Kate Tasker, G. Caleb Alexander, Jiuxiang Gu |
AAAI | 6 |
| 2026 | VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-useabstractWhile vision-language models (VLMs) have demonstrated remarkable performance across various tasks combining textual and visual information, they continue to struggle with fine-grained visual perception tasks that require detailed pixel-level analysis. Effectively eliciting comprehensive reasoning from VLMs on such intricate visual elements remains an open challenge. In this paper, we present VipAct, an agent framework that enhances VLMs by integrating multi-agent collaboration and vision expert models, enabling more precise visual understanding and comprehensive reasoning. VipAct consists of an orchestrator agent, which manages task requirement analysis, planning, and coordination, along with specialized agents that handle specific tasks such as image captioning and vision expert models that provide high-precision perceptual information. This multi-agent approach allows VLMs to better perform fine-grained visual perception tasks by synergizing planning, reasoning, and tool use. We evaluate VipAct on benchmarks featuring a diverse set of visual perception tasks, with experimental results demonstrating significant performance improvements over state-of-the-art baselines across all tasks. Furthermore, comprehensive ablation studies reveal the critical role of multi-agent collaboration in eliciting more detailed System-2 reasoning and highlight the importance of image input for task planning. Additionally, our error analysis identifies patterns of VLMs' inherent limitations in visual perception, providing insights into potential future improvements. VipAct offers a flexible and extensible framework, paving the way for more advanced visual perception systems across various real-world applications. Zhehao Zhang 0001, Ryan Rossi, Tong Yu 0001, Franck Dernoncourt, Ruiyi Zhang 0002, Jiuxiang Gu, Sungchul Kim, Xiang Chen 0010, Zichao Wang 0001, Nedim Lipka |
AAAI | 5 |
| 2026 | Lizard: An Efficient Linearization Framework for Large Language ModelsabstractChien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang 0002, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Jayakumar Subramanian, Ryan Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen |
ACL (1) | 3 |
| 2026 | CachePrune: Teaching LLMs What Not to Follow via KV-Cache EditingabstractRui Wang, Junda Wu, Yu Xia, Tong Yu, Ruiyi Zhang, Ryan A. Rossi, Subrata Mitra, Lina Yao, Julian McAuley. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Wang 0088, Junda Wu, Yu Xia 0007, Tong Yu 0001, Ruiyi Zhang 0002, Ryan Rossi, Subrata Mitra, Lina Yao 0001, Julian J. McAuley |
ACL (1) | 5 |
| 2026 | Federated Large Language Models: Current Progress and Future Directions
Yuhang Yao 0003, Junda Wu, Chengkai Huang, Yu Xia 0007, Tong Yu 0001, Ruiyi Zhang 0002, Sungchul Kim, Ryan Rossi, Ang Li 0005, Lina Yao 0001, Julian J. McAuley, Yiran Chen 0001, Carlee Joe-Wong |
PAKDD (4) | 7 |
| 2025 | Numerical Pruning for Efficient Autoregressive ModelsabstractTransformers have emerged as the leading architecture in deep learning, proving to be versatile and highly effective across diverse domains beyond language and image processing. However, their impressive performance often incurs high computational costs due to their substantial model size. This paper focuses on compressing decoder-only transformer-based autoregressive models through structural weight pruning to improve the model efficiency while preserving performance for both language and image generation tasks. Specifically, we propose a training-free pruning method that calculates a numerical score with Newton's method for the Attention and MLP modules, respectively. Besides, we further propose another compensation algorithm to recover the pruned model for better performance. To verify the effectiveness of our method, we provide both theoretical support and extensive experiments. Our experiments show that our method achieves state-of-the-art performance with reduced memory usage and faster generation speeds on GPUs. Xuan Shen, Zhao Song 0002, Yufa Zhou 0001, Bo Chen 0029, Jing Liu 0001, Ruiyi Zhang 0002, Ryan Rossi, Hao Tan 0005, Tong Yu 0001, Xiang Chen 0010, Yufan Zhou 0001, Tong Sun 0005, Pu Zhao 0001, Yanzhi Wang 0001, Jiuxiang Gu |
AAAI | 6 |
| 2025 | From Selection to Generation: A Survey of LLM-based Active LearningabstractYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao, Nedim Lipka, Seunghyun Yoon, Ting-Hao Kenneth Huang, Zichao Wang, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee, Zhehao Zhang, Namyong Park, Thien Huu Nguyen, Jiebo Luo, Ryan A. Rossi, Julian McAuley. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Xia 0007, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li 0001, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen 0003, Franck Dernoncourt, Branislav Kveton, Tong Yu 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang 0160, Xiang Chen 0010, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao 0016, Nedim Lipka, Seunghyun Yoon 0002, Ting-Hao 'Kenneth' Huang, Zichao Wang 0001, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee 0001, Zhehao Zhang 0001, Namyong Park 0001, Thien Huu Nguyen, Jiebo Luo 0001, Ryan Rossi, Julian J. McAuley |
ACL (1) | 13 |
| 2025 | A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction GenerationabstractLarge multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but its quality is poor, as multimodal alignment remains a hurdle even for the largest models. In this work, we propose LLaVAR-2, to enhance multimodal alignment for text-rich images through hybrid instruction generation between human annotators and large language models. Specifically, it involves detailed image captions from human annotators, followed by the use of these annotations in tailored text prompts for GPT-4o to curate a dataset. It also implements several mechanisms to filter out low-quality data, and the resulting dataset comprises 424k high-quality pairs of instructions. Empirical results show that models fine-tuned on this dataset exhibit impressive enhancements over those trained with self-instruct data. Shijie Zhou 0008, Ruiyi Zhang 0002, Yufan Zhou 0001, Changyou Chen |
COLING | 2 |
| 2025 | GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous ExplorationabstractGraphical User Interface (GUI) action grounding, mapping language instructions to actionable elements on GUI screens, is important for assisting users in interactive tutorials, task automation, accessibility support, etc. Most recent works of GUI action grounding use large GUI datasets to fine-tune Multimodal Large Language Models (MLLMs). However, the fine-tuning data is inherently limited to specific GUI environments, leading to significant performance degradation in novel environments due to the generalization challenges in the GUI domain. Therefore, we argue that GUI action grounding models should be further aligned with novel environments before deployment to optimize their performance. To address this, we first propose GUI-Bee, an MLLM-based autonomous agent, to collect high-quality, environment-specific data through exploration and then continuously fine-tune GUI grounding models with the collected data. To ensure the GUI action grounding models generalize to various screens within the target novel environment after the continuous fine-tuning, we equip GUI-Bee with a novel Q-value-Incentive In-Context Reinforcement Learning (Q-ICRL) algorithm that optimizes exploration efficiency and exploration data quality. In the experiment, we introduce NovelScreenSpot to test how well the data can help align GUI action grounding models to novel environments. Furthermore, we conduct an ablation study to validate the Q-ICRL method in enhancing the efficiency of GUI-Bee. Handong Zhao, Ruiyi Zhang 0002, Xin Wang 0061, Gang Wu 0013 |
EMNLP | 3 |
| 2025 | Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
Shijie Zhou 0008, Ruiyi Zhang 0002, Huaisheng Zhu, Branislav Kveton, Yufan Zhou 0001, Jiuxiang Gu, Jian Chen 0043, Changyou Chen |
ICCV | 2 |
| 2025 | SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document UnderstandingabstractMultimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named **S**elf-**V**isual **R**etrieval-**A**ugmented **G**eneration (SV-RAG), which can broaden horizons of *any* MLLM to support long-document understanding. We demonstrate that **MLLMs themselves can be an effective multimodal retriever** to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG. Jian Chen 0043, Ruiyi Zhang 0002, Yufan Zhou 0001, Tong Yu 0001, Franck Dernoncourt, Jiuxiang Gu, Ryan Rossi, Changyou Chen, Tong Sun 0005 |
ICLR | 2 |
| 2025 | ARTIST: Improving the Generation of Text-Rich Images with Disentangled Diffusion Models and Large Language ModelsabstractDiffusion models have demonstrated exceptional capabilities in generating a broad spectrum of visual content, yet their proficiency in rendering text is still limited: they often generate inaccurate characters or words that fail to blend well with the underlying image. To address these shortcomings, we introduce a novel framework named ARTIST, which incorporates a dedicated textual diffusion model to focus on the learning of text structures specifically. Initially, we pretrain this textual model to capture the intricacies of text representation. Subsequently, we finetune a visual diffusion model, enabling it to assimilate textual structure information from the pretrained textual model. This disentangled architecture design and training strategy significantly enhance the text rendering ability of the diffusion models for text-rich image generation. Additionally, we leverage the capabilities of pretrained large language models to interpret user intentions better, contributing to improved generation quality. Empirical results on the MARIO-Eval benchmark underscore the effectiveness of the proposed method, showing an improvement of up to 15% in various metries. Yufan Zhou 0001, Jiuxiang Gu, Curtis Wigington, Tong Yu 0001, Yiran Chen 0001, Tong Sun 0005, Ruiyi Zhang 0002 |
WACV | 8 |
| 2024 | Knowledge Graph Prompting for Multi-Document Question AnsweringabstractThe `pre-train, prompt, predict' paradigm of large language models (LLMs) has achieved remarkable success in open-domain question answering (OD-QA). However, few works explore this paradigm in multi-document question answering (MD-QA), a task demanding a thorough understanding of the logical associations among the contents and structures of documents. To fill this crucial gap, we propose a Knowledge Graph Prompting (KGP) method to formulate the right context in prompting LLMs for MD-QA, which consists of a graph construction module and a graph traversal module. For graph construction, we create a knowledge graph (KG) over multiple documents with nodes symbolizing passages or document structures (e.g., pages/tables), and edges denoting the semantic/lexical similarity between passages or document structural relations. For graph traversal, we design an LLM-based graph traversal agent that navigates across nodes and gathers supporting passages assisting LLMs in MD-QA. The constructed graph serves as the global ruler that regulates the transitional space among passages and reduces retrieval latency. Concurrently, the graph traversal agent acts as a local navigator that gathers pertinent context to progressively approach the question and guarantee retrieval quality. Extensive experiments underscore the efficacy of KGP for MD-QA, signifying the potential of leveraging graphs in enhancing the prompt design and retrieval augmented generation for LLMs. Our code: https://github.com/YuWVandy/KG-LLM-MDQA. Yu Wang 0160, Nedim Lipka, Ryan Rossi, Alexa F. Siu, Ruiyi Zhang 0002, Tyler Derr |
AAAI | 5 |
| 2024 | Topology-aware Retrieval Augmentation for Text Generation
Yu Wang 0160, Nedim Lipka, Ruiyi Zhang 0002, Alexa F. Siu, Yuying Zhao, Bo Ni, Xin Wang 0061, Ryan Rossi, Tyler Derr |
CIKM | 3 |
| 2024 | TRINS: Towards Multimodal Language Models that Can ReadabstractLarge multimodal language models have shown remarkable proficiency in understanding and editing images. However, a majority of these visually-tuned models struggle to comprehend the textual content embedded in images, primar-ily due to the limitation of training data. In this work, we introduce TRINS: a Text-Rich image11In this work, we use the phrase “text-rich images” to describe images with rich textual information, such as posters and book covers. INStruction dataset, with the objective of enhancing the reading ability of the multimodal large language model. TRINS is built upon LAION22Work done during Q3 2023. using hybrid data annotation strategies that include machine-assisted and human-assisted annotation process. It contains 39,153 text-rich images, captions, and 102,437 questions. Specifically, we show that the number of words per annotation in TRINS is significantly longer than that of related datasets, providing new challenges. Furthermore, we introduce a simple and effective architecture, called a Language-Vision Reading Assistant (LaRA), which is good at understanding textual content within images. LaRA outperforms existing state-of-the-art multimodal large language models on the TRINS dataset as well as other classical benchmarks. Lastly, we conducted a comprehensive evaluation with TRINS on various text-rich image understanding and generation tasks, demonstrating its effectiveness. Ruiyi Zhang 0002, Jian Chen 0043, Yufan Zhou 0001, Jiuxiang Gu, Changyou Chen, Tong Sun 0005 |
CVPR | 1 |
| 2024 | Customization Assistant for Text-to-image GenerationabstractCustomizing pre-trained text-to-image generation model has attracted massive research interest recently, due to its huge potential in real-world applications. Although existing methods are able to generate creative content for a novel concept contained in single user-input image, their capability are still far from perfection. Specifically, most existing methods require fine-tuning the generative model on testing images. Some existing methods do not require fine-tuning, while their performance are unsatisfactory. Furthermore, the interaction between users and models are still limited to directive and descriptive prompts such as instructions and captions. In this work, we build a customization assistant based on pre-trained large language model and diffusion model, which can not only perform customized generation in a tuning-free manner, but also enable more user-friendly interactions: users can chat with the assistant and input ei-ther ambiguous text or clear instruction. Specifically, we propose a new framework consists of a new model design and a novel training strategy. The resulting assistant can perform customized generation in 2–5 seconds without any test time fine-tuning. Extensive experiments are conducted, competitive results have been obtained across different domains, illustrating the effectiveness of the proposed method. Yufan Zhou 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Tong Sun 0005 |
CVPR | 2 |
| 2024 | Few-Shot Dialogue Summarization via Skeleton-Assisted Prompt Transfer in Prompt TuningabstractKaige Xie, Tong Yu, Haoliang Wang, Junda Wu, Handong Zhao, Ruiyi Zhang, Kanak Mahadik, Ani Nenkova, Mark Riedl. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kaige Xie, Tong Yu 0001, Junda Wu, Handong Zhao, Ruiyi Zhang 0002, Kanak Mahadik, Ani Nenkova, Mark O. Riedl |
EACL (1) | 6 |
| 2024 | Towards Building The Federatedgpt: Federated Instruction TuningabstractWhile "instruction-tuned" generative large language models (LLMs) have demonstrated an impressive ability to generalize to new tasks, the training phases heavily rely on large amounts of diverse and high-quality instruction data (such as ChatGPT and GPT-4). Unfortunately, acquiring high-quality data, especially when it comes to human-written data, can pose significant challenges both in terms of cost and accessibility. Moreover, concerns related to privacy can further limit access to such data, making the process of obtaining it a complex and nuanced undertaking. To tackle this issue, our study introduces a new approach called Federated Instruction Tuning (FedIT), which leverages federated learning (FL) as the learning framework for the instruction tuning of LLMs. This marks the first exploration of FL-based instruction tuning for LLMs. This is especially important since text data is predominantly generated by end users. For example, collecting extensive amounts of everyday user conversations can be a useful approach to improving the generalizability of LLMs, allowing them to generate authentic and natural responses. Therefore, it is imperative to design and adapt FL approaches to effectively leverage these users’ diverse instructions stored on local devices while mitigating concerns related to the data sensitivity and the cost of data transmission. In this study, we leverage extensive qualitative analysis, including the prevalent GPT-4 auto-evaluation to illustrate how our FedIT framework enhances the performance of LLMs. Utilizing diverse instruction sets on the client side, FedIT outperforms centralized training with only limited local instructions. Saeed Vahidian, Martin Kuo, Chunyuan Li, Ruiyi Zhang 0002, Tong Yu 0001, Guoyin Wang 0002, Yiran Chen 0001 |
ICASSP | 5 |
| 2024 | SOHES: Self-supervised Open-world Hierarchical Entity SegmentationabstractOpen-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive generalization capabilities on unseen images and concepts. Despite its promise, existing entity segmentation methods like Segment Anything Model (SAM) rely heavily on costly expert annotators. This work presents Self-supervised Open-world Hierarchical Entity Segmentation (SOHES), a novel approach that eliminates the need for human annotations. SOHES operates in three phases: self-exploration, self-instruction, and self-correction. Given a pre-trained self-supervised representation, we produce abundant high-quality pseudo-labels through visual feature clustering. Then, we train a segmentation model on the pseudo-labels, and rectify the noises in pseudo-labels via a teacher-student mutual-learning procedure. Beyond segmenting entities, SOHES also captures their constituent parts, providing a hierarchical understanding of visual entities. Using raw images as the sole training data, our method achieves unprecedented performance in self-supervised open-world segmentation, marking a significant milestone towards high-quality open-world entity segmentation in the absence of human-annotated masks. Project page: https://SOHES.github.io. Shengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan 0002, Ruiyi Zhang 0002, Handong Zhao, Ani Nenkova, Liangyan Gui, Tong Sun 0005, Yu-Xiong Wang |
ICLR | 5 |
| 2024 | Towards Aligned Layout Generation via Diffusion Model with Aesthetic ConstraintsabstractControllable layout generation refers to the process of creating a plausible visual arrangement of elements within a graphic design (*e.g.*, document and web designs) with constraints representing design intentions. Although recent diffusion-based models have achieved state-of-the-art FID scores, they tend to exhibit more pronounced misalignment compared to earlier transformer-based models. In this work, we propose the **LA**yout **C**onstraint diffusion mod**E**l (LACE), a unified model to handle a broad range of layout generation tasks, such as arranging elements with specified attributes and refining or completing a coarse layout design. The model is based on continuous diffusion models. Compared with existing methods that use discrete diffusion models, continuous state-space design can enable the incorporation of continuous aesthetic constraint functions in training more naturally. For conditional generation, we propose injecting layout conditions in the form of masks or gradient guidance during inference. Empirical results show that LACE produces high-quality layouts and outperforms existing state-of-the-art baselines. We will release our source code and model checkpoints. Jian Chen 0043, Ruiyi Zhang 0002, Yufan Zhou 0001, Changyou Chen |
ICLR | 2 |
| 2024 | ADOPD: A Large-Scale Document Page Decomposition DatasetabstractResearch in document image understanding is hindered by limited high-quality document data. To address this, we introduce ADOPD, a comprehensive dataset for document page decomposition. ADOPD stands out with its data-driven approach for document taxonomy discovery during data collection, complemented by dense annotations. Our approach integrates large-scale pretrained models with a human-in-the-loop process to guarantee diversity and balance in the resulting data collection. Leveraging our data-driven document taxonomy, we collect and densely annotate document images, addressing four document image understanding tasks: Doc2Mask, Doc2Box, Doc2Tag, and Doc2Seq. Specifically, for each image, the annotations include human-labeled entity masks, text bounding boxes, as well as automatically generated tags and captions that have been manually cleaned. We conduct comprehensive experimental analyses to validate our data and assess the four tasks using various models. We envision ADOPD as a foundational dataset with the potential to drive future research in document understanding. Jiuxiang Gu, Xiangxi Shi, Jason Kuen, Ruiyi Zhang 0002, Anqi Liu 0001, Ani Nenkova, Tong Sun 0005 |
ICLR | 5 |
| 2024 | Bias and Fairness in Large Language Models: A SurveyabstractAbstract Rapid advancements of large language models (LLMs) have enabled the processing, understanding, and generation of human-like text, with increasing integration into systems that touch our social sphere. Despite this success, these models can learn, perpetuate, and amplify harmful social biases. In this article, we present a comprehensive survey of bias evaluation and mitigation techniques for LLMs. We first consolidate, formalize, and expand notions of social bias and fairness in natural language processing, defining distinct facets of harm and introducing several desiderata to operationalize fairness for LLMs. We then unify the literature by proposing three intuitive taxonomies, two for bias evaluation, namely, metrics and datasets, and one for mitigation. Our first taxonomy of metrics for bias evaluation disambiguates the relationship between metrics and evaluation datasets, and organizes metrics by the different levels at which they operate in a model: embeddings, probabilities, and generated text. Our second taxonomy of datasets for bias evaluation categorizes datasets by their structure as counterfactual inputs or prompts, and identifies the targeted harms and social groups; we also release a consolidation of publicly available datasets for improved access. Our third taxonomy of techniques for bias mitigation classifies methods by their intervention during pre-processing, in-training, intra-processing, and post-processing, with granular subcategories that elucidate research trends. Finally, we identify open problems and challenges for future work. Synthesizing a wide range of recent research, we aim to provide a clear guide of the existing literature that empowers researchers and practitioners to better understand and prevent the propagation of bias in LLMs. Isabel O. Gallegos, Ryan Rossi, Joe Barrow, Md. Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu 0001, Ruiyi Zhang 0002, Nesreen K. Ahmed |
Comput. Linguistics | 8 |
| 2023 | Few-Shot Composition Learning for Image Retrieval with Prompt TuningabstractWe study the problem of composition learning for image retrieval, for which we learn to retrieve target images with search queries in the form of a composition of a reference image and a modification text that describes desired modifications of the image. Existing models of composition learning for image retrieval are generally built with large-scale datasets, demanding extensive training samples, i.e., query-target pairs, as supervision, which restricts their application for the scenario of few-shot learning with only few query-target pairs available. Recently, prompt tuning with frozen pretrained language models has shown remarkable performance when the amount of training data is limited. Inspired by this, we propose a prompt tuning mechanism with the pretrained CLIP model for the task of few-shot composition learning for image retrieval. Specifically, we regard the representation of the reference image as a trainable visual prompt, prefixed to the embedding of the text sequence. One challenge is to efficiently train visual prompt with few-shot samples. To deal with this issue, we further propose a self-upervised auxiliary task via ensuring that the reference image can retrieve itself when no modification information is given from the text, which facilitates training for the visual prompt, while not requiring additional annotations for query-target pairs. Experiments on multiple benchmarks show that our proposed model can yield superior performance when trained with only few query-target pairs. Junda Wu, Rui Wang 0088, Handong Zhao, Ruiyi Zhang 0002, Chaochao Lu, Shuai Li 0010, Ricardo Henao |
AAAI | 4 |
| 2023 | Learning Navigational Visual Representations with Semantic Map SupervisionabstractBeing able to perceive the semantics and the spatial structure of the environment is essential for visual navigation of a household robot. However, most existing works only employ visual backbones pre-trained either with independent images for classification or with self-supervised learning methods to adapt to the indoor navigation domain, neglecting the spatial relationships that are essential to the learning of navigation. Inspired by the behavior that humans naturally build semantically and spatially meaningful cognitive maps in their brains during navigation, in this paper, we propose a novel navigational-specific visual representation learning method by contrasting the agent’s egocentric views and semantic maps (Ego2-Map). We apply the visual transformer as the backbone encoder and train the model with data collected from the large-scale HabitatMatterport3D environments. Ego2-Map learning transfers the compact and rich information from a map, such as objects, structure and transition, to the agent’s egocentric representations for navigation. Experiments show that agents using our learned representations on object-goal navigation outperform recent visual pre-training methods. Moreover, our representations significantly improve vision-and-language navigation in continuous environments for both high-level and low-level action spaces, achieving new state-of-the-art results of 47% SR and 41% SPL on the test server. Yicong Hong, Yang Zhou 0009, Ruiyi Zhang 0002, Franck Dernoncourt, Trung Bui, Stephen Gould, Hao Tan 0002 |
ICCV | 3 |
| 2023 | Label-Retrieval-Augmented Diffusion Models for Learning from Noisy LabelsabstractLearning from noisy labels is an important and long-standing problem in machine learning for real applications. One of the main research lines focuses on learning a label corrector to purify potential noisy labels. However, these methods typically rely on strict assumptions and are limited to certain types of label noise. In this paper, we reformulate the label-noise problem from a generative-model perspective, *i.e.*, labels are generated by gradually refining an initial random guess. This new perspective immediately enables existing powerful diffusion models to seamlessly learn the stochastic generative process. Once the generative uncertainty is modeled, we can perform classification inference using maximum likelihood estimation of labels. To mitigate the impact of noisy labels, we propose the **L**abel-**R**etrieval-**A**ugmented (LRA) diffusion model, which leverages neighbor consistency to effectively construct pseudo-clean labels for diffusion training. Our model is flexible and general, allowing easy incorporation of different types of conditional information, *e.g.*, use of pre-trained models, to further boost model performance. Extensive experiments are conducted for evaluation. Our model achieves new state-of-the-art (SOTA) results on all the standard real-world benchmark datasets. Remarkably, by incorporating conditional information from the powerful CLIP model, our method can boost the current SOTA accuracy by 10-20 absolute points in many cases. Code is available: https://anonymous.4open.science/r/LRA-diffusion-5F2F Jian Chen 0043, Ruiyi Zhang 0002, Tong Yu 0001, Rohan Sharma, Tong Sun 0005, Changyou Chen |
NeurIPS | 2 |
| 2023 | InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language UnderstandingabstractSoft prompt tuning achieves superior performances across a wide range of few-shot tasks. However, the performances of prompt tuning can be highly sensitive to the initialization of the prompts. We have also empirically observed that conventional prompt tuning methods cannot encode and learn sufficient task-relevant information from prompt tokens. In this work, we develop an information-theoretic framework that formulates soft prompt tuning as maximizing the mutual information between prompts and other model parameters (or encoded representations). This novel view helps us to develop a more efficient, accurate and robust soft prompt tuning method, InfoPrompt. With this framework, we develop two novel mutual information based loss functions, to (i) explore proper prompt initialization for the downstream tasks and learn sufficient task-relevant information from prompt tokens and (ii) encourage the output representation from the pretrained language model to be more aware of the task-relevant information captured in the learnt prompts. Extensive experiments validate that InfoPrompt can significantly accelerate the convergence of the prompt tuning and outperform traditional prompt tuning methods. Finally, we provide a formal theoretical result to show that a gradient descent type algorithm can be used to train our mutual information loss. Junda Wu, Tong Yu 0001, Rui Wang 0088, Zhao Song 0002, Ruiyi Zhang 0002, Handong Zhao, Chaochao Lu, Shuai Li 0010, Ricardo Henao |
NeurIPS | 5 |
| 2022 | Text-Based Interactive Recommendation via Offline Reinforcement LearningabstractInteractive recommendation with natural-language feedback can provide richer user feedback and has demonstrated advantages over traditional recommender systems. However, the classical online paradigm involves iteratively collecting experience via interaction with users, which is expensive and risky. We consider an offline interactive recommendation to exploit arbitrary experience collected by multiple unknown policies. A direct application of policy learning with such fixed experience suffers from the distribution shift. To tackle this issue, we develop a behavior-agnostic off-policy correction framework to make offline interactive recommendation possible. Specifically, we leverage the conservative Q-function to perform off-policy evaluation, which enables learning effective policies from fixed datasets without further interactions. Empirical results on the simulator derived from real-world datasets demonstrate the effectiveness of our proposed offline training framework. Ruiyi Zhang 0002, Tong Yu 0001, Yilin Shen, Hongxia Jin |
AAAI | 1 |
| 2022 | TiGAN: Text-Based Interactive Image Generation and ManipulationabstractUsing natural-language feedback to guide image generation and manipulation can greatly lower the required efforts and skills. This topic has received increased attention in recent years through refinement of Generative Adversarial Networks (GANs); however, most existing works are limited to single-round interaction, which is not reflective of real world interactive image editing workflows. Furthermore, previous works dealing with multi-round scenarios are limited to predefined feedback sequences, which is also impractical. In this paper, we propose a novel framework for Text-based Interactive image generation and manipulation (TiGAN) that responds to users' natural-language feedback. TiGAN utilizes the powerful pre-trained CLIP model to understand users' natural-language feedback and exploits contrastive learning for a better text-to-image mapping. To maintain the image consistency during interactions, TiGAN generates intermediate feature vectors aligned with the feedback and selectively feeds these vectors to our proposed generative model. Empirical results on several datasets show that TiGAN improves both interaction efficiency and image quality while better avoids undesirable image manipulation during interactions. Yufan Zhou 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Chris Tensmeyer, Tong Yu 0001, Changyou Chen, Jinhui Xu 0001, Tong Sun 0005 |
AAAI | 2 |
| 2022 | Few-Shot Class-Incremental Learning for Named Entity RecognitionabstractRui Wang, Tong Yu, Handong Zhao, Sungchul Kim, Subrata Mitra, Ruiyi Zhang, Ricardo Henao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Rui Wang 0088, Tong Yu 0001, Handong Zhao, Sungchul Kim, Subrata Mitra, Ruiyi Zhang 0002, Ricardo Henao |
ACL (1) | 6 |
| 2022 | Towards Language-Free Training for Text-to-Image GenerationabstractOne of the major challenges in training text-to-image generation models is the need of a large number of highquality image-text pairs. While image samples are often easily accessible, the associated text descriptions typically require careful human captioning, which is particularly time- and cost-consuming. In this paper, we propose the first work to train text-to-image generation models without any text data. Our method leverages the well-aligned multi-modal semantic space of the powerful pre-trained CLIP model: the requirement of text-conditioning is seamlessly alleviated via generating text features from image features. Extensive experiments are conducted to illustrate the effectiveness of the proposed method. We obtain state-of-the-art results in the standard text-to-image generation tasks. Importantly, the proposed language-free model outperforms most existing models trained with full image-text pairs. Furthermore, our method can be applied in fine-tuning pretrained models, which saves both training time and cost in training text-to-image generation models. Our pre-trained model obtains competitive results in zero-shot text-to-image generation on the MS-COCO dataset, yet with around only 1% of the model size and training data size relative to the recently proposed large DALL-E model. Yufan Zhou 0001, Ruiyi Zhang 0002, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu 0001, Jiuxiang Gu, Jinhui Xu 0001, Tong Sun 0005 |
CVPR | 2 |
| 2022 | Federated Non-negative Matrix Factorization for Short Texts Topic Modeling with Mutual InformationabstractNon-negative matrix factorization (NMF) based topic modeling is widely used in natural language processing (NLP) to uncover hidden topics of short text documents. Usually, training a high-quality topic model requires large amount of textual data. In many real-world scenarios, customer textual data should be private and sensitive, precluding uploading to data centers. This paper proposes a Federated NMF (FedNMF) framework, which allows multiple clients to collaboratively train a high-quality NMF based topic model with locally stored data. However, standard federated learning will significantly undermine the performance of topic models in downstream tasks (e.g., text classification) when the data distribution over clients is heterogeneous. To alleviate this issue, we further propose FedNMF+MI, which simultaneously maximizes the mutual information (MI) between the count features of local texts and their topic weight vectors to mitigate the performance degradation. Experimental results show that our FedNMF+MI methods outperform Federated Latent Dirichlet Allocation (FedLDA) and the FedNMF without MI methods for short texts by a significant margin on both coherence score and classification F1 score. Shijing Si, Jianzong Wang, Ruiyi Zhang 0002, Qinliang Su, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | CIPhy: Causal Intervention with Physical Confounder from IoT Sensor Data for Robust Occupant Information InferenceabstractOccupant information inference with IoT sensor data enables many smart applications, such as patients'/older adults' in-home monitoring. The difficulty of collecting labeled real-world IoT sensor data often leads to reliability and scalability issues for those systems. Extensive prior works (e.g., domain adaptation) focus on the domain shift issues, i.e., the inconsistent data feature and label relationship, and dataset bias is often neglected. Dataset bias is commonly caused by limited and varied accessibility to labeled data for each class, and it is inevitable for real-world datasets. The model trained with a biased dataset fits into the bias, hence cannot further generalize to the testing data for accurate inference. Zhizhang Hu, Tong Yu 0001, Ruiyi Zhang 0002, Shijia Pan |
SenSys | 3 |
| 2022 | Dynamics-Aware Adaptation for Reinforcement Learning Based Cross-Domain Interactive RecommendationabstractInteractive recommender systems (IRS) have received wide attention in recent years. To capture users' dynamic preferences and maximize their long-term engagement, IRS are usually formulated as reinforcement learning (RL) problems. Despite the promise to solve complex decision-making problems, RL-based methods generally require a large amount of online interaction, restricting their applications due to economic considerations. One possible direction to alleviate this issue is cross-domain recommendation that aims to leverage abundant logged interaction data from a source domain (e.g., adventure genre in movie recommendation) to improve the recommendation quality in the target domain (e.g., crime genre). Nevertheless, prior studies mostly focus on adapting the static representations of users/items. Few have explored how the temporally dynamic user-item interaction patterns transform across domains. Junda Wu, Zhihui Xie 0002, Tong Yu 0001, Handong Zhao, Ruiyi Zhang 0002, Shuai Li 0010 |
SIGIR | 5 |
| 2021 | Unsupervised Abstractive Dialogue Summarization for Tete-a-TetesabstractHigh-quality dialogue-summary paired data is expensive to produce and domain-sensitive, making abstractive dialogue summarization a challenging task. In this work, we propose the first unsupervised abstractive dialogue summarization model for tete-a-tetes (SuTaT). Unlike standard text summarization, a dialogue summarization method should consider the multi-speaker scenario where the speakers have different roles, goals, and language styles. In a tete-a-tete, such as a customer-agent conversation, SuTaT aims to summarize for each speaker by modeling the customer utterances and the agent utterances separately while retaining their correlations. SuTaT consists of a conditional generative module and two unsupervised summarization modules. The conditional generative module contains two encoders and two decoders in a variational autoencoder framework where the dependencies between two latent spaces are captured. With the same encoders and decoders, two unsupervised summarization modules equipped with sentence-level self-attention mechanisms generate summaries without using any annotations. Experimental results show that SuTaT is superior on unsupervised dialogue summarization for both automatic and human evaluations, and is capable of dialogue classification and single-turn conversation generation. Xinyuan Zhang 0001, Ruiyi Zhang 0002, Manzil Zaheer, Amr Ahmed 0001 |
AAAI | 2 |
| 2021 | Improving Zero-Shot Voice Style Transfer via Disentangled Representation Learning
Siyang Yuan, Pengyu Cheng, Ruiyi Zhang 0002, Weituo Hao, Zhe Gan, Lawrence Carin |
ICLR | 3 |
| 2020 | Learning Diverse Stochastic Human-Action Generators by Learning Smooth Latent TransitionsabstractHuman-motion generation is a long-standing challenging task due to the requirement of accurately modeling complex and diverse dynamic patterns. Most existing methods adopt sequence models such as RNN to directly model transitions in the original action space. Due to high dimensionality and potential noise, such modeling of action transitions is particularly challenging. In this paper, we focus on skeleton-based action generation and propose to model smooth and diverse transitions on a latent space of action sequences with much lower dimensionality. Conditioned on a latent sequence, actions are generated by a frame-wise decoder shared by all latent action-poses. Specifically, an implicit RNN is defined to model smooth latent sequences, whose randomness (diversity) is controlled by noise from the input. Different from standard action-prediction methods, our model can generate action sequences from pure noise without any conditional action poses. Remarkably, it can also generate unseen actions from mixed classes during training. Our model is learned with a bi-directional generative-adversarial-net framework, which can not only generate diverse action sequences of a particular class or mix classes, but also learns to classify action sequences within the same model. Experimental results show the superiority of our method in both diverse action-sequence generation and classification, relative to existing methods. Zhenyi Wang 0001, Ruiyi Zhang 0002, Yufan Zhou 0001, Junsong Yuan 0001, Changyou Chen |
AAAI | 4 |
| 2020 | Improving Adversarial Text Generation by Modeling the Distant FutureabstractAuto-regressive text generation models usually focus on local fluency, and may cause inconsistent semantic meaning in long text generation.Further, automatically generating words with similar semantics is challenging, and hand-crafted linguistic rules are difficult to apply.We consider a text planning scheme and present a model-based imitation-learning approach to alleviate the aforementioned issues.Specifically, we propose a novel guider network to focus on the generative process over a longer horizon, which can assist next-word prediction and provide intermediate rewards for generator optimization.Extensive experiments demonstrate that the proposed method leads to improved performance. Ruiyi Zhang 0002, Changyou Chen, Zhe Gan, Wenlin Wang, Dinghan Shen, Guoyin Wang 0002, Lawrence Carin |
ACL | 1 |
| 2020 | Nested-Wasserstein Self-Imitation Learning for Sequence GenerationabstractReinforcement learning (RL) has been widely studied for improving sequence-generation models. However, the conventional rewards used for RL training typically cannot capture sufficient semantic information and therefore render model bias. Further, the sparse and delayed rewards make RL exploration inefficient. To alleviate these issues, we propose the concept of nested-Wasserstein distance for distributional semantic matching. To further exploit it, a novel nested-Wasserstein self-imitation learning framework is developed, encouraging the model to exploit historical high-rewarded sequences for enhanced exploration and better semantic matching. Our solution can be understood as approximately executing proximal policy optimization with Wasserstein trust-regions. Experiments on a variety of unconditional and conditional sequence-generation tasks demonstrate the proposed approach consistently leads to improved performance. Ruiyi Zhang 0002, Changyou Chen, Zhe Gan, Wenlin Wang, Lawrence Carin |
AISTATS | 1 |
| 2020 | Stochastic Particle-Optimization Sampling and the Non-Asymptotic Convergence TheoryabstractParticle-optimization-based sampling (POS) is a recently developed effective sampling technique that interactively updates a set of particles. A representative algorithm is the Stein variational gradient descent (SVGD). We prove, under certain conditions, SVGD experiences a theoretical pitfall, {\it i.e.}, particles tend to collapse. As a remedy, we generalize POS to a stochastic setting by injecting random noise into particle updates, thus termed stochastic particle-optimization sampling (SPOS). Notably, for the first time, we develop non-asymptotic convergence theory for the SPOS framework (related to SVGD), characterizing algorithm convergence in terms of the 1-Wasserstein distance w.r.t. the numbers of particles and iterations. Somewhat surprisingly, with the same number of updates (not too large) for each particle, our theory suggests adopting more particles does not necessarily lead to a better approximation of a target distribution, due to limited computational budget and numerical errors. This phenomenon is also observed in SVGD and verified via a synthetic experiment. Extensive experimental results verify our theory and demonstrate the effectiveness of our proposed framework. Ruiyi Zhang 0002, Lawrence Carin, Changyou Chen |
AISTATS | 2 |
| 2020 | Repulsive Attention: Rethinking Multi-head Attention as Bayesian InferenceabstractBang An, Jie Lyu, Zhenyi Wang, Chunyuan Li, Changwei Hu, Fei Tan, Ruiyi Zhang, Yifan Hu, Changyou Chen. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Bang An 0001, Jie Lyu 0004, Zhenyi Wang 0001, Chunyuan Li, Changwei Hu, Fei Tan 0002, Ruiyi Zhang 0002, Yifan Hu 0001, Changyou Chen |
EMNLP (1) | 7 |
| 2020 | Improving Text Generation with Student-Forcing Optimal TransportabstractJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuhchen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jianqiao Li, Chunyuan Li, Guoyin Wang 0002, Hao Fu 0002, Yuh-Chen Lin, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Ruiyi Zhang 0002, Wenlin Wang, Dinghan Shen, Qian Yang 0003, Lawrence Carin |
EMNLP (1) | 9 |
| 2020 | Bayesian Meta Sampling for Fast Uncertainty Adaptation
Zhenyi Wang 0001, Ruiyi Zhang 0002, Changyou Chen |
ICLR | 4 |
| 2020 | Graphical Models Meet Bandits: A Variational Thompson Sampling ApproachabstractWe propose a novel framework for structured bandits, which we call an influence diagram bandit. Our framework uses a graphical model to capture complex statistical dependencies between actions, latent variables, and observations; and thus unifies and extends many existing models, such as combinatorial semi-bandits, cascading bandits, and low-rank bandits. We develop novel online learning algorithms that learn to act efficiently in our models. The key idea is to track a structured posterior distribution of model parameters, either exactly or approximately. To act, we sample model parameters from their posterior and then use the structure of the influence diagram to find the most optimistic action under the sampled parameters. We empirically evaluate our algorithms in three structured bandit problems, and show that they perform as well as or better than problem-specific state-of-the-art baselines. Tong Yu 0001, Branislav Kveton, Ruiyi Zhang 0002, Ole J. Mengshoel |
ICML | 4 |
| 2020 | Figure Captioning with Relation Maps for ReasoningabstractFigures, such as line plots, pie charts, bar charts, are widely used to convey important information in a concise format. In this work, we investigate the problem of figure caption generation where the goal is to automatically generate a natural language description for a given figure. While natural image captioning has been studied extensively, figure captioning has received relatively little attention and remains a challenging problem. A successful solution to this task has many potential applications, such as: 1) automatic parsing large amount of figures in PDF document; 2) improving user experience by allowing figure content to be accessible to those with visual impairment. To solve this problem, we introduce a dataset FigCAP and propose novel attention mechanism. In order to solve the exposure bias issue, we further train the captioning model with sequence-level policy based on reinforcement learning, which directly optimizes evaluation metrics. Extensive experiments show that the proposed method outperforms the baselines, thus demonstrating a significant potential for automatic generating captions for figures. Ruiyi Zhang 0002, Eunyee Koh, Sungchul Kim, Scott Cohen, Ryan Rossi |
WACV | 2 |
| 2019 | Scalable Thompson Sampling via Optimal TransportabstractThompson sampling (TS) is a class of algorithms for sequential decision-making, which requires maintaining a posterior distribution over a reward model. However, calculating exact posterior distributions is intractable for all but the simplest models. Consequently, how to computationally-efficiently approximate a posterior distribution is a crucial problem for scalable TS with complex models, such as neural networks. In this paper, we use distribution optimization techniques to approximate the posterior distribution, solved via Wasserstein gradient flows. Based on the framework, a principled particle-optimization algorithm is developed for TS to approximate the posterior efficiently. Our approach is scalable and does not make explicit distribution assumptions on posterior approximations. Extensive experiments on both synthetic data and large-scale real data demonstrate the superior performance of the proposed methods. Ruiyi Zhang 0002, Changyou Chen, Tong Yu 0001, Lawrence Carin |
AISTATS | 1 |
| 2019 | Improving Sequence-to-Sequence Learning via Optimal Transport
Liqun Chen 0001, Yizhe Zhang 0002, Ruiyi Zhang 0002, Chenyang Tao, Zhe Gan, Bai Li 0001, Dinghan Shen, Changyou Chen, Lawrence Carin |
ICLR (Poster) | 3 |
| 2019 | Understanding and Accelerating Particle-Based Variational InferenceabstractParticle-based variational inference methods (ParVIs) have gained attention in the Bayesian inference literature, for their capacity to yield flexible and accurate approximations. We explore ParVIs from the perspective of Wasserstein gradient flows, and make both theoretical and practical contributions. We unify various finite-particle approximations that existing ParVIs use, and recognize that the approximation is essentially a compulsory smoothing treatment, in either of two equivalent forms. This novel understanding reveals the assumptions and relations of existing ParVIs, and also inspires new ParVIs. We propose an acceleration framework and a principled bandwidth-selection method for general ParVIs; these are based on the developed theory and leverage the geometry of the Wasserstein space. Experimental results show the improved convergence by the acceleration framework and enhanced sample accuracy by the bandwidth-selection method. Chang Liu 0030, Jingwei Zhuo, Pengyu Cheng, Ruiyi Zhang 0002, Jun Zhu 0001 |
ICML | 4 |
| 2019 | Variational Annealing of GANs: A Langevin PerspectiveabstractThe generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to improve GAN training. To enrich the understanding of this fast-growing yet almost exclusively heuristic-driven subject, we elucidate the theoretical roots of some of the empirical attempts to stabilize and improve GAN training with the introduction of likelihoods. We highlight new insights from variational theory of diffusion processes to derive a likelihood-based regularizing scheme for GAN training, and present a novel approach to train GANs with an unnormalized distribution instead of empirical samples. To substantiate our claims, we provide experimental evidence on how our theoretically-inspired new algorithms improve upon current practice. Chenyang Tao, Shuyang Dai, Liqun Chen 0001, Ke Bai 0001, Junya Chen, Chang Liu 0030, Ruiyi Zhang 0002, Georgiy V. Bobashev, Lawrence Carin |
ICML | 7 |
| 2019 | Vision-Language Recommendation via Attribute Augmented Multimodal Reinforcement LearningabstractInteractive recommenders have demonstrated the advantage over traditional recommenders with dynamic change of items. However, the traditional user feedback in the format of clicks or ratings, provides limited user preference information and limited history tracking capabilities. As a result, it takes a user many interactions to find a desired item. Data of other modalities, such as item visual appearance and user comments in natural language, may enable richer user feedback. However, there are several critical challenges to be addressed when utilizing these multimodal data: multimodal matching, user preference tracking, and adaptation to dynamic unseen items. Without properly handling these challenges, the recommendations can easily violate the users' preference from their past natural language feedback. In this paper, we introduce a novel approach, called vision-language recommendation, that enables users to provide natural language feedback on visual products to have more natural and effective interactions. To model more explicit and accurate multimodal matching, we propose a novel visual attribute augmented reinforcement learning approach that enhances the grounding of natural language to visual items. Furthermore, to effectively track the users' preference and overcome the performance deficiency on dynamic unseen items after deployment, we propose a novel history multimodal matching reward to continuously adapt the model on-the-fly. Empirical results show that, our system augmented by visual attribute and history multimodal matching can significantly increase the success rate, reduce the number of recommendations that violate the user's previous feedback, and need less number of user interactions to find the desired items. Tong Yu 0001, Yilin Shen, Ruiyi Zhang 0002, Xiangyu Zeng 0003, Hongxia Jin |
ACM Multimedia | 3 |
| 2019 | Improving Textual Network Learning with Variational Homophilic EmbeddingsabstractThe performance of many network learning applications crucially hinges on the success of network embedding algorithms, which aim to encode rich network information into low-dimensional vertex-based vector representations. This paper considers a novel variational formulation of network embeddings, with special focus on textual networks. Different from most existing methods that optimize a discriminative objective, we introduce Variational Homophilic Embedding (VHE), a fully generative model that learns network embeddings by modeling the semantic (textual) information with a variational autoencoder, while accounting for the structural (topology) information through a novel homophilic prior design. Homophilic vertex embeddings encourage similar embedding vectors for related (connected) vertices. The VHE encourages better generalization for downstream tasks, robustness to incomplete observations, and the ability to generalize to unseen vertices. Extensive experiments on real-world networks, for multiple tasks, demonstrate that the proposed method achieves consistently superior performance relative to competing state-of-the-art approaches. Wenlin Wang, Chenyang Tao, Zhe Gan, Guoyin Wang 0002, Liqun Chen 0001, Xinyuan Zhang 0001, Ruiyi Zhang 0002, Qian Yang 0003, Ricardo Henao, Lawrence Carin |
NeurIPS | 7 |
| 2019 | Text-Based Interactive Recommendation via Constraint-Augmented Reinforcement LearningabstractText-based interactive recommendation provides richer user preferences and has demonstrated advantages over traditional interactive recommender systems. However, recommendations can easily violate preferences of users from their past natural-language feedback, since the recommender needs to explore new items for further improvement. To alleviate this issue, we propose a novel constraint-augmented reinforcement learning (RL) framework to efficiently incorporate user preferences over time. Specifically, we leverage a discriminator to detect recommendations violating user historical preference, which is incorporated into the standard RL objective of maximizing expected cumulative future rewards. Our proposed framework is general and is further extended to the task of constrained text generation. Empirical results show that the proposed method yields consistent improvement relative to standard RL methods. Ruiyi Zhang 0002, Tong Yu 0001, Yilin Shen, Hongxia Jin, Changyou Chen |
NeurIPS | 1 |
| 2018 | Learning Structural Weight Uncertainty for Sequential Decision-MakingabstractLearning probability distributions on the weights of neural networks (NNs) has recently proven beneficial in many applications. Bayesian methods, such as Stein variational gradient descent (SVGD), offer an elegant framework to reason about NN model uncertainty. However, by assuming independent Gaussian priors for the individual NN weights (as often applied), SVGD does not impose prior knowledge that there is often structural information (dependence) among weights. We propose efficient posterior learning of structural weight uncertainty, within an SVGD framework, by employing matrix variate Gaussian priors on NN parameters. We further investigate the learned structural uncertainty in sequential decision-making problems, including contextual bandits and reinforcement learning. Experiments on several synthetic and real datasets indicate the superiority of our model, compared with state-of-the-art methods. Ruiyi Zhang 0002, Chunyuan Li, Changyou Chen, Lawrence Carin |
AISTATS | 1 |
| 2018 | Variational Inference and Model Selection with Generalized Evidence BoundsabstractRecent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the conventional variational bound. It naturally subsumes the importance-weighted and Renyi bounds as special cases, and it is provably sharper than these counterparts. We also present an improved estimator for variational learning, and advocate a novel high signal-to-variance ratio update rule for the variational parameters. We discuss model-selection issues associated with existing evidence-lower-bound-based variational inference procedures, and show how to leverage the flexibility of our new formulation to address them. Empirical evidence is provided to validate our claims. Liqun Chen 0001, Chenyang Tao, Ruiyi Zhang 0002, Ricardo Henao, Lawrence Carin |
ICML | 3 |
| 2018 | Policy Optimization as Wasserstein Gradient FlowsabstractPolicy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical success, its correspondence to policy-distribution optimization has been unclear mathematically. We place policy optimization into the space of probability measures, and interpret it as Wasserstein gradient flows. On the probability-measure space, under specified circumstances, policy optimization becomes convex in terms of distribution optimization. To make optimization feasible, we develop efficient algorithms by numerically solving the corresponding discrete gradient flows. Our technique is applicable to several RL settings, and is related to many state-of-the-art policy-optimization algorithms. Specifically, we define gradient flows on both the parameter-distribution space and policy-distribution space, leading to what we term indirect-policy and direct-policy learning frameworks, respectively. Extensive experiments verify the effectiveness of our framework, often obtaining better performance compared to related algorithms. Ruiyi Zhang 0002, Changyou Chen, Chunyuan Li, Lawrence Carin |
ICML | 1 |
| 2018 | Adversarial Text Generation via Feature-Mover's DistanceabstractGenerative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose to improve text-generation GAN via a novel approach inspired by optimal transport. Specifically, we consider matching the latent feature distributions of real and synthetic sentences using a novel metric, termed the feature-mover's distance (FMD). This formulation leads to a highly discriminative critic and easy-to-optimize objective, overcoming the mode-collapsing and brittle-training problems in existing methods. Extensive experiments are conducted on a variety of tasks to evaluate the proposed model empirically, including unconditional text generation, style transfer from non-parallel text, and unsupervised cipher cracking. The proposed model yields superior performance, demonstrating wide applicability and effectiveness. Liqun Chen 0001, Shuyang Dai, Chenyang Tao, Zhe Gan, Dinghan Shen, Yizhe Zhang 0002, Guoyin Wang 0002, Ruiyi Zhang 0002, Lawrence Carin |
NeurIPS | 9 |
| 2018 | A Unified Particle-Optimization Framework for Scalable Bayesian Sampling
Changyou Chen, Ruiyi Zhang 0002, Wenlin Wang, Bai Li 0001, Liqun Chen 0001 |
UAI | 2 |