Paul Pu Liang

dblp:207/9749 · DBLP profile ↗
← Back
67ranked-venue papers
14as first author
47since 2021 · last 2026
0000-0001-7768-3610ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 12 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 6 since 2021Computer networks · 2
YearPublicationVenuePosition
2026 Breaking Negative Cycles: A Reflection-to-Action System for Adaptive Change
abstract
Breaking negative mental health cycles, including rumination and recurring regrets, requires reflection that translates awareness into behavioral change. Grounded in the Transtheoretical Model (TTM) and Gross’s Emotion Regulation (ER) Process Model, we examine how Technologies Supporting Self-Reflection (TSR) bridge reflection and action. In a 15-day in-the-wild study (N = 20), participants used a voice-based journaling system to capture regrets and wishes and engaged in WhatIf-Planning, a novel structured reflection module integrating counterfactual thinking with if–then planning. Participants were randomized to either a free-form condition or a Gross-guided condition, which maps the five processes of Gross’s ER model into explicit journaling prompts. We contribute: (1) a unified reflection-to-action TSR system that operationalizes the Preparation stage of TTM to bridge Contemplation and Action, and (2) triangulated empirical evidence from an in-the-wild journaling study that first operationalizes Gross’s Process Model, revealing effects on coping flexibility and emotion regulation in daily life. Results show significant pre–post improvements in coping flexibility, indicating adaptive self-regulation across conditions, with the Gross-guided group generating more counterfactual alternatives, articulating concrete if–then action plans, and implementing more plans for self-driven change.
Minsol Michelle Kim, Daniel Low, David Lafond, Eugene Shim, Michelle Han, Mohanad Kandil, Chenyu Zhang 0009, Theo Kitsberg, Chelsea Boccagno, Paul Pu Liang, Pattie Maes
CHI10
2026 Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
abstract
Given the growing prevalence of fake information, including increasingly realistic AI-generated news, there is an urgent need to train people to better evaluate and detect misinformation. While interactions with AI have been shown to durably reduce people’s beliefs in false information, it is unclear whether these interactions also teach people the skills to discern false information themselves. We conducted a month-long study where 67 participants classified news headline-image pairs as real or fake, discussed their assessments with an AI system, followed by an unassisted evaluation of unseen news items to measure accuracy before, during, and after AI assistance. While AI assistance produced immediate improvements during AI-assisted sessions (+21% average), participants’ unassisted performance on new items declined significantly by 15.3% in week 4 compared to week 0. These results indicate that while AI may help immediately, it may ultimately degrade long-term misinformation detection abilities.
Anku Rani, Valdemar Danry, Paul Pu Liang, Andy Lippman, Pattie Maes
CHI3
2026 Mind Mapper: Modeling and Predicting Behavioral Patterns from Everyday Conversations with Wearable AI Systems and LLMs
abstract
Everyday conversations are more than exchanges of words—they reveal how people think, react, and adapt across situations. Through our speech we reveal the cognitive patterns that shape our behavior over time. Yet much of these patterns remains implicit: people operate through recurring heuristics and habits—some helpful, some limiting, many unnoticed. However, identifying such patterns requires self awareness that humans struggle with—and that current user modeling systems, often fragmented and narrowly tailored to specific tasks, fail to capture. Recognizing these patterns can enable user support systems to go beyond reactive assistance towards anticipatory support. We present Mind Mapper, an always-on wearable AI system that mines behavioral patterns from everyday real-life conversations. Mind Mapper employs a multi-stage LLM pipeline to generate, refine, and evaluate human-readable behavioral patterns. In a field study with 12 participants capturing over 700 hours of real-life conversational data, Mind Mapper generated behavioral patterns that participants consistently rated as accurate, unique, and helpful for reflection and behavior change. We further illustrate how such behavioral pattern modeling might enable new forms of human-AI interaction—anticipating user behavior through proactive interventions, adaptive content delivery, cognitive reframing, and behavioral simulation. Our results show the potential of always-on wearable systems with LLM-driven user models to support new forms of cognitively scaffolding, context-aware human-AI interactions.
Valdemar Danry, Jean Ghislain Billa, Yasith Samaradivakara, Paul Pu Liang, Pattie Maes
IUI4
2026 WiReSens Toolkit: An Open-source Platform towards Accessible Wireless Tactile Sensing
abstract
Past research has widely explored the design and fabrication of resistive matrix-based tactile sensors for creating touch-sensitive devices. However, real-world deployment of resistive tactile sensing systems remains difficult for individuals with limited prior experience in embedded sensing due to challenges of portability, adaptivity, and efficiency. We introduce the WiReSens Toolkit, an accessible, open-source platform to bridge this gap. Central to our approach is adaptive hardware for interfacing with resistive sensors and a web-based GUI that streamlines access to advanced features for building scalable tactile sensing systems, including multi-device programming and wireless visualization across three communication protocols, autocalibration for adaptive sensitivity, and intermittent data transmission for low-power use. We validated the toolkit’s usability through a user study with 11 novice participants, who, on average, configured a tactile sensor with over 95% accuracy in under five minutes, calibrated sensors 10× faster than baseline methods, and showed improved sense-making of tactile data.
Devin Murphy, Junyi Zhu 0001, Akshay Gadre, Antonio Torralba 0001, Paul Pu Liang, Wojciech Matusik, Yiyue Luo
TEI5
2025 VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
abstract
Visually linking matching cues is a crucial ability in daily life, such as identifying the same person in multiple photos based on their cues, even without knowing who they are. Despite the extensive knowledge that vision-language models (VLMs) possess, it remains largely unexplored whether they are capable of performing this fundamental task. To address this, we introduce VLM2-Bench, a benchmark designed to assess whether VLMs can Visually Link Matching cues, with 9 subtasks and over 3,000 test cases. Comprehensive evaluation across twelve VLMs, along with further analysis of various language-side and vision-side prompting methods, leads to a total of eight key findings. We identify critical challenges in models’ ability to link visual cues, highlighting a significant performance gap. Based on these insights, we advocate for (i) enhancing core visual capabilities to improve adaptability and reduce reliance on prior knowledge, (ii) establishing clearer principles for integrating language-based reasoning in vision-centric tasks to prevent unnecessary biases, and (iii) shifting vision-text training paradigms toward fostering models’ ability to independently structure and infer relationships among visual cues.
Jianshu Zhang 0003, Dongyu Yao, Renjie Pi, Paul Pu Liang, Yi R. Fung 0001
ACL (1)4
2025 Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
abstract
Social reasoning abilities are crucial for AI systems to effectively interpret and respond to multimodal human communication and interaction within social contexts.We introduce SOCIAL GENOME, the first benchmark for fine-grained, grounded social reasoning abilities of multimodal models.SOCIAL GENOME contains 272 videos of interactions and 1,486 humanannotated reasoning traces related to inferences about these interactions.These traces contain 5,777 reasoning steps that reference evidence from visual cues, verbal cues, vocal cues, and external knowledge (contextual knowledge external to videos).SOCIAL GENOME is also the first modeling challenge to study external knowledge in social reasoning.SOCIAL GENOME computes metrics to holistically evaluate semantic and structural qualities of modelgenerated social reasoning traces.We demonstrate the utility of SOCIAL GENOME through experiments with state-of-the-art models, identifying performance gaps and opportunities for future research to improve the grounded social reasoning abilities of multimodal models.
Leena Mathur, Marian Qian, Paul Pu Liang, Louis-Philippe Morency
EMNLP3
2025 Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation
abstract
Counterfactual reasoning typically involves considering alternatives to actual events.While often applied to understand past events, a distinct form-forward counterfactual reasoningfocuses on anticipating plausible future developments.This type of reasoning is invaluable in dynamic financial markets, where anticipating market developments can powerfully unveil potential risks and opportunities for stakeholders, guiding their decision-making.However, performing this at scale is challenging due to the cognitive demands involved, underscoring the need for automated solutions.LLMs offer promise, but remain unexplored for this application.To address this gap, we introduce a novel benchmark, FIN-FORCE-FINancial FORward Counterfactual Evaluation.By curating financial news headlines and providing structured evaluation, FIN-FORCE supports LLM based forward counterfactual generation.This paves the way for scalable and automated solutions for exploring and anticipating future market developments, thereby providing structured insights for decision-making.Through experiments on FIN-FORCE, we evaluate state-of-theart LLMs and counterfactual generation methods, analyzing their limitations and proposing insights for future research.We release the benchmark, supplementary data and all experimental codes at the following link
Keane Ong, Rui Mao 0010, Deeksha Varshney, Paul Pu Liang, Erik Cambria, Gianmarco Mengaldo
EMNLP4
2025 OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
abstract
In recent years, there has been increasing interest in automatic facial behavior analysis systems from computing communities such as vision, multimodal interaction, robotics, and affective computing. Building upon the widespread utility of prior open-source facial analysis systems, we introduce OpenFace 3.0, an open-source toolkit capable of facial landmark detection, facial action unit detection, eye-gaze estimation, and facial emotion recognition. OpenFace 3.0 contributes a lightweight unified model for facial analysis, trained with a multi-task architecture across diverse populations, head poses, lighting conditions, video resolutions, and facial analysis tasks. By leveraging the benefits of parameter sharing through a unified model and training paradigm, OpenFACE 3.0 exhibits improvements in prediction performance, inference speed, and memory efficiency over similar toolkits and rivals state-of-theart models. Openface 3.0 can be installed and run with a single line of code and operate in real-time without specialized hardware. OpenFace 3.0 code for training models and running the system is freely available for research purposes and supports contributions from the community.
Jiewen Hu, Leena Mathur, Paul Pu Liang, Louis-Philippe Morency
FG3
2025 OS-ATLAS: Foundation Action Model for Generalist GUI Agents
abstract
Existing efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision. Practitioners are often reluctant to use open-source VLMs due to their significant performance lag compared to their closed-source counterparts, particularly in GUI grounding and Out-Of-Distribution (OOD) scenarios. To facilitate future research in this area, we developed OS-Atlas—a foundational GUI action model that excels at GUI grounding and OOD agentic tasks through innovations in both data and modeling. We have invested significant engineering effort in developing an open-source toolkit for synthesizing GUI grounding data across multiple platforms, including Windows, Linux, MacOS, Android, and the web. Leveraging this toolkit, we are releasing the largest open-source cross-platform GUI grounding corpus to date, which contains over 13 million GUI elements. This dataset, combined with innovations in model training, provides a solid foundation for OS-Atlas to understand GUI screenshots and generalize to unseen interfaces. Through extensive evaluation across six benchmarks spanning three different platforms (mobile, desktop, and web), OS-Atlas demonstrates significant performance improvements over previous state-of-the-art models. Our evaluation also uncovers valuable insights into continuously improving and scaling the agentic capabilities of open-source VLMs.
Zhiyong Wu 0003, Fangzhi Xu, Yian Wang 0003, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding 0002, Paul Pu Liang, Yu Qiao 0001
ICLR10
2025 Progressive Compositionality in Text-to-Image Generative Models
abstract
Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes, especially in complex settings. Existing approaches through building compositional architectures or generating difficult negative captions often assume a fixed prespecified compositional structure, which limits generalization to new distributions. In this paper, we argue that curriculum training is crucial to equipping generative models with a fundamental understanding of compositionality. To achieve this, we leverage large-language models (LLMs) to automatically compose complex scenarios and harness Visual-Question Answering (VQA) checkers to automatically curate a contrastive dataset, ConPair, consisting of 15k pairs of high-quality contrastive images. These pairs feature minimal visual discrepancies and cover a wide range of attribute categories, especially complex and natural scenarios. To learn effectively from these error cases (i.e., hard negative images), we propose EvoGen, a new multi-stage curriculum for contrastive learning of diffusion models. Through extensive experiments across a wide range of compositional scenarios, we showcase the effectiveness of our proposed framework on compositional T2I benchmarks.
Linghao Jin, Paul Pu Liang
ICLR4
2025 VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
abstract
Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text or static imagery can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focus- ing on text or static image inputs. To bridge this gap, we introduce VideoWebArena (VideoWA), a benchmark for evaluating the capabilities of long-context multimodal agents for video understanding. VideoWA consists of 2,021 web agent tasks based on manually crafted video tutorials, which total almost four hours of content. For our benchmark, we define a taxonomy of long-context video-based agent tasks with two main areas of focus: skill retention and factual retention. While skill retention tasks evaluate whether an agent can use a given human demonstration to complete a task efficiently, the factual retention task evaluates whether an agent can retrieve instruction-relevant information from a video to complete a task. We find that the best model achieves a 13.3% success rate on factual retention tasks and 45.8% on factual retention QA pairs—far below human success rates of 73.9% and 79.3%, respectively. On skill retention tasks, long-context models perform worse with tutorials than without, exhibiting a 5% performance decrease in WebArena tasks and a 10.3% decrease in VisualWebArena tasks. Our work highlights performance gaps in the agentic abilities of long-context multimodal models and provides as a testbed for the future development of long-context video agents.
Lawrence Jang, Yinheng Li, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, Kazuhito Koishida
ICLR6
2025 TeaserGen: Generating Teasers for Long Documentaries
abstract
Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling capability for the input videos, while necessitating maintaining audiovisual alignments, managing scene transitions and preserving factual accuracy for the output teasers. Due to the lack of a publicly-available dataset, progress along this research direction has been hindered. In this work, we present DocumentaryNet, a collection of 1,269 documentaries paired with their teasers, featuring multimodal data streams of video, speech, music, sound effects and narrations. With DocumentaryNet, we propose a new two-stage system for generating teasers from long documentaries. The proposed TeaserGen system first generates the teaser narration from the transcribed narration from the documentary using a pretrained large language model, and then selects the most relevant visual content to accompany the generated narration through language-vision models. For narration-video matching, we explore two approaches: a pretraining-based model using pretrained contrastive language-vision models and a deep sequential model that learns the mapping between the narrations and visuals. Our experimental results show that the pretraining-based approach is more effective at identifying relevant visual content than directly trained deep autoregressive models.
Weihan Xu, Paul Pu Liang, Haven Kim, Julian J. McAuley, Taylor Berg-Kirkpatrick, Hao-Wen Dong
ICLR2
2025 CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models
abstract
Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the development of large-scale multimodal methods that can make holistic assessments of patient health and well-being. To bridge this gap, we introduce Clinical Large-scale Integrative Multimodal Benchmark (CLIMB), a comprehensive clinical benchmark unifying diverse clinical data across imaging, language, temporal, and graph modalities. CLIMB comprises 4.51 million patient samples totaling 19.01 terabytes distributed across 2D imaging, 3D video, time series, graphs, and multimodal data. Through extensive empirical evaluation, we demonstrate that multitask pretraining significantly improves performance on understudied domains, achieving up to 29% improvement in ultrasound and 23% in ECG analysis over single-task learning. Pretraining on CLIMB also effectively improves models’ generalization capability to new tasks, and strong unimodal encoder performance translates well to multimodal performance when paired with task-appropriate fusion strategies. Our findings provide a foundation for new architecture designs and pretraining strategies to adavance clinical AI research. Code is released at https://github.com/DDVD233/climb.
Wei Dai 0013, Malinda Lu, Haowen Wei, Hejie Cui, Paul Pu Liang
ICML7
2025 Understanding the Emergence of Multimodal Representation Alignment
abstract
Multimodal representation learning is fundamentally about transforming incomparable modalities into comparable representations. While prior research has primarily focused on explicitly aligning these representations through targeted learning objectives and model architectures, a recent line of work has found that independently trained unimodal models of increasing scale and performance can become implicitly aligned with each other. These findings raise fundamental questions regarding the emergence of aligned representations in multimodal learning. Specifically: (1) when and why does alignment emerge implicitly? and (2) is alignment a reliable indicator of performance? Through a comprehensive empirical investigation, we demonstrate that both the emergence of alignment and its relationship with task performance depend on several critical data characteristics. These include, but are not necessarily limited to, the degree of similarity between the modalities and the balance between redundant and unique information they provide for the task. Our findings suggest that alignment may not be universally beneficial; rather, its impact on performance varies depending on the dataset and task. These insights can help practitioners determine whether increasing alignment between modalities is advantageous or, in some cases, detrimental to achieving optimal performance.
Megan Tjandrasuwita, Chanakya Ajit Ekbote, Liu Ziyin 0001, Paul Pu Liang
ICML4
2025 QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
abstract
Clinical decision‑making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision‑centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time‑series signals, and text reports. QoQ-Med is trained with Domain‑aware Relative Policy Optimization (DRPO), a novel reinforcement‑learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro‑F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and downstream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces.
David Dai, Chanakya Ajit Ekbote, Paul Pu Liang
NeurIPS4
2025 What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
abstract
In-context learning (ICL) is a hallmark capability of transformers, through which trained models learn to adapt to new tasks by leveraging information from the input context. Prior work has shown that ICL emerges in transformers due to the presence of special circuits called induction heads. Given the equivalence between induction heads and conditional $k$-grams, a recent line of work modeling sequential inputs as Markov processes has revealed the fundamental impact of model depth on its ICL capabilities: while a two-layer transformer can efficiently represent a conditional $1$-gram model, its single-layer counterpart cannot solve the task unless it is exponentially large. However, for higher order Markov sources, the best known constructions require at least three layers (each with a single attention head) - leaving open the question: *can a two-layer single-head transformer represent any $k^{\text{th}}$-order Markov process?* In this paper, we precisely address this and theoretically show that a two-layer transformer with one head per layer can indeed represent any conditional $k$-gram. Thus, our result provides the tightest known characterization of the interplay between transformer depth and Markov order for ICL. Building on this, we further analyze the learning dynamics of our two-layer construction, focusing on a simplified variant for first-order Markov chains, illustrating how effective in-context representations emerge during training. Together, these results deepen our current understanding of transformer-based ICL and illustrate how even shallow architectures can surprisingly exhibit strong ICL capabilities on structured sequence modeling tasks.
Chanakya Ajit Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman, Michael Gastpar, Jason D. Lee, Paul Pu Liang
NeurIPS7
2025 Balancing Multimodal Training Through Game-Theoretic Regularization
abstract
Multimodal learning holds the promise for richer information extraction by capturing dependencies across data sources. Yet, current training methods often underperform due to modality competition, a phenomenon where modalities contend for training resources, leaving some underoptimized. This raises a pivotal question: how can we address training imbalances, ensure adequate optimization across all modalities, and achieve consistent performance improvements as we transition from unimodal to multimodal data? This paper proposes the Multimodal Competition Regularizer (MCR), inspired by a mutual information (MI) decomposition designed to prevent the adverse effects of competition in multimodal training. Our key contributions are: 1) A game-theoretic framework that adaptively balances modality contributions by encouraging each to maximize its informative role in the final prediction. 2) Refining lower and upper bounds for each MI term to enhance the extraction of both task-relevant unique and shared information across modalities. 3) Proposing latent space permutations for conditional MI estimation, significantly improving computational efficiency. MCR outperforms all previously suggested training strategies and simple baselines, demonstrating that training modalities jointly lead to important performance gains on synthetic and large real-world datasets. We release our code and models at https://github.com/kkontras/MCR.
Konstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang, Matthew B. Blaschko, Maarten De Vos
NeurIPS4
2025 MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
abstract
As AI becomes more closely integrated with peoples' daily activities, socially intelligent AI that can understand and interact seamlessly with humans in daily lives is increasingly important. However, current works in AI social reasoning all rely on language-only or language-dominant approaches to benchmark and training models, resulting in systems that are improving in verbal communication but struggle with nonverbal social understanding. To address this limitation, we tap into a novel data source rich in nonverbal social interactions -- mime videos. Mimes refer to the art of expression through gesture and movement without spoken words, which presents unique challenges and opportunities in interpreting nonverbal social communication. We contribute a new dataset called MimeQA, obtained by sourcing ~8 hours of videos clips from YouTube and developing a comprehensive video question-answering benchmark comprising 806 carefully annotated and verified question-answer pairs, designed to probe nonverbal social reasoning capabilities. Using MimeQA, we evaluate state-of-the-art video large language models (VideoLLMs) and find that they achieve low accuracy, generally ranging from 20-30%, while humans score 86\%. Our analysis reveals that VideoLLMs often fail to ground imagined objects and over-rely on the text prompt while ignoring subtle nonverbal interactions. We hope to inspire future work in AI models that embody true social intelligence capable of interpreting non-verbal human interactions.
Hengzhi Li, Megan Tjandrasuwita, Yi R. Fung 0001, Armando Solar-Lezama, Paul Pu Liang
NeurIPS5
2025 REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
abstract
Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing models for generating shorts that feature a coherent narrative with embedded video insertions extracted from a long input video. We propose a novel retrieval-embedded generation framework that allows a large language model to quote multimodal resources while maintaining a coherent narrative. Our proposed REGen system first generates the output story script with quote placeholders using a finetuned large language model, and then uses a novel retrieval model to replace the quote placeholders by selecting a video clip that best supports the narrative from a pool of candidate quotable video clips. We examine the proposed method on the task of documentary teaser generation, where short interview insertions are commonly used to support the narrative of a documentary. Our objective evaluations show that the proposed method can effectively insert short video clips while maintaining a coherent narrative. In a subjective survey, we show that our proposed method outperforms existing abstractive and extractive approaches in terms of coherence, alignment, and realism in teaser generation.
Weihan Xu, Yimeng Ma, Jingyue Huang, Wenye Ma, Taylor Berg-Kirkpatrick, Julian J. McAuley, Paul Pu Liang, Hao-Wen Dong
NeurIPS8
2025 Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions
abstract
The study of multimodality has garnered significant interest in fields where analyzing interactions among multiple information sources can enhance predictive modeling, data fusion, and interpretability. Partial information decomposition (PID) has emerged as a useful information-theoretic framework to quantify the degree to which individual modalities independently, redundantly, or synergistically convey information about a target variable. However, existing PID methods depend on optimizing over a joint distribution constrained by estimated pairwise probability distributions, which are costly and inaccurate for continuous and high-dimensional modalities. Our first key insight is that the problem can be solved efficiently when the pairwise distributions are multivariate Gaussians, and we refer to this problem as Gaussian PID (GPID). We propose a new gradient-based algorithm that substantially enhances computational efficiency for GPID based on an alternative formulation of the underlying optimization problem. To generalize the applicability to non-Gaussian data, we learn information-preserving encoders to transform random variables of arbitrary input distributions into pairwise Gaussian random variables. Along the way, we resolved an open problem regarding the optimality of joint Gaussian solutions for GPID. Empirical validation on diverse synthetic examples demonstrates that our proposed method provides more accurate and efficient PID estimates than existing baselines. We further evaluate on a series of large-scale multimodal benchmarks to show its utility in real-world applications of quantifying PID in multimodal datasets and selecting high-performing models.
Wenyuan Zhao, Adithya Balachandran, Chao Tian 0002, Paul Pu Liang
NeurIPS4
2025 POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
Evans Xu Han, Alice Qian Zhang, Haiyi Zhu, Hong Shen 0004, Paul Pu Liang, Jane Hsieh
UIST5
2025 Comparative Knowledge Distillation
abstract
In the era of large-scale pretrained models, Knowledge Distillation (KD) serves an important role in transferring the wisdom of computationally-heavy teacher models to lightweight, efficient student models while preserving performance. Yet KD settings often assume readily available access to teacher models capable of performing many in-ferences-a notion increasingly at odds with the realities of costly large-scale models. Addressing this gap, we study an important question: how KD algorithms fare as the number of teacher inferences decreases, a setting we term Reduced-Teacher-Inference Knowledge Distillation (RTI-KD). We observe that the performance of prevalent KD techniques and state-of-the-art data augmentation strategies suffers considerably as the number of teacher inferences is reduced. One class of approaches, termed “relational” knowledge distillation underperforms the rest, yet we hypothesize that they hold promise for reduced dependency on teacher models because they can augment the effective dataset size without additional teacher calls. We find that a simple change - performing high-dimensional comparisons instead of low-dimensional relations, which we term Comparative Knowledge Distillation - vaults performance well over existing KD approaches. We perform empirical evaluation across varied experimental settings and rigorous analysis to understand the learning outcomes of our method. All code is made publicly available.
Alex Tianyi Xu, Alex Wilf, Paul Pu Liang, Alexander Obolenskiv, Daniel Fried, Louis-Philippe Morency
WACV3
2024 Think Twice: Perspective-Taking Improves Large Language Models' Theory-of-Mind Capabilities
abstract
Human interactions are deeply rooted in the interplay of thoughts, beliefs, and desires made possible by Theory of Mind (ToM): our cognitive ability to understand the mental states of ourselves and others.Although ToM may come naturally to us, emulating it presents a challenge to even the most advanced Large Language Models (LLMs).Recent improvements to LLMs' reasoning capabilities from simple yet effective prompting techniques such as Chain-of-Thought (CoT) (Wei et al., 2022) have seen limited applicability to ToM (Gandhi et al., 2023).In this paper, we turn to the prominent cognitive science theory "Simulation Theory" to bridge this gap.We introduce SIMTOM, a novel two-stage prompting framework inspired by Simulation Theory's notion of perspective-taking.To implement this idea on current ToM benchmarks, SIMTOM first filters context based on what the character in question knows before answering a question about their mental state.Our approach, which requires no additional training and minimal prompt-tuning, shows substantial improvement over existing methods, and our analysis reveals the importance of perspective-taking to Theory-of-Mind capabilities.Our findings suggest perspectivetaking as a promising direction for future research into improving LLMs' ToM capabilities.Our code is publicly available.
Alex Wilf, Sihyun Shawn Lee, Paul Pu Liang, Louis-Philippe Morency
ACL (1)3
2024 Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival Prediction
abstract
Integrating whole-slide images (WSIs) and bulk tran-scriptomics for predicting patient survival can improve our understanding of patient prognosis. However, this multi-modal task is particularly challenging due to the different nature of these data: WSIs represent a very high-dimensional spatial description of a tumor, while bulk tran-scriptomics represent a global description of gene expression levels within that tumor. In this context, our work aims to address two key challenges: (1) how can we tokenize transcriptomics in a semantically meaningful and interpretable way?, and (2) how can we capture dense multi-modal interactions between these two modalities? Here, we propose to learn biological pathway tokens from transcriptomics that can encode specific cellular functions. Together with histology patch tokens that encode the slide morphology, we argue that they form appropriate reasoning units for interpretability. We fuse both modalities using a memoryefficient multimodal Transformer that can model interactions between pathway and histology patch tokens. Our model, Survpath, achieves state-of-the-art performance when evaluated against unimodal and multimodal baselines on five datasets from The Cancer Genome Atlas. Our interpretability framework identifies key multimodal prognostic factors, and, as such, can provide valuable insights into the interaction between genotype and phenotype. Code available at https://github.com/mahmoodlab/SurvPath.
Guillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson, Paul Pu Liang, Faisal Mahmood 0001
CVPR5
2024 FLHetBench: Benchmarking Device and State Heterogeneity in Federated Learning
abstract
Federated learning (FL) is a powerful technology that enables collaborative training of machine learning models without sharing private data among clients. The fundamental challenge in FL lies in learning over extremely hetero-geneous data distributions, device capacities, and device state availabilities, all of which adversely impact performance and communication efficiency. While data hetero-geneity has been well-studied in the literature, this paper introduces FLHetBench, the first FL benchmark targeted toward understanding device and state heterogeneity. FL-HetBench comprises two new sampling methods to generate real-world device and state databases with varying het-erogeneity and new metrics for quantifying the success of FL methods under these real-world constraints. Using FL-HetBench, we conduct a comprehensive evaluation of existing methods and find that they struggle under these settings, which inspires us to propose BiasPrompt+, a new method employing staleness-aware aggregation and fast weights to tackle these new heterogeneity challenges. Experiments on various FL tasks and datasets validate the effectiveness of our BiasPrompt+ method and highlight the value of FLHet-Bench in fostering the development of more efficient and robust FL solutions under real-world device and state constraints.
Junyuan Zhang, Shuang Zeng, Miao Zhang 0030, Runxi Wang, Yuyin Zhou, Paul Pu Liang, Liangqiong Qu
CVPR7
2024 Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
abstract
Building socially-intelligent AI agents (Social-AI) is a multidisciplinary, multimodal research goal that involves creating agents that can sense, perceive, reason about, learn from, and respond to affect, behavior, and cognition of other agents (human or artificial).Progress towards Social-AI has accelerated in the past decade across several computing communities, including natural language processing, machine learning, robotics, human-machine interaction, computer vision, and speech.Natural language processing, in particular, has been prominent in Social-AI research, as language plays a key role in constructing the social world.In this position paper, we identify a set of underlying technical challenges and open questions for researchers across computing communities to advance Social-AI.We anchor our discussion in the context of social intelligence concepts and prior progress in Social-AI research.
Leena Mathur, Paul Pu Liang, Louis-Philippe Morency
EMNLP2
2024 MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
abstract
Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled.Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like imagetext matching.However, this covers only a subset of real-world interactions.Novel interactions, such as sarcasm expressed through opposing spoken words and gestures or humor expressed through utterances and tone of voice, remain challenging.In this paper, we introduce an approach to enhance multimodal models, which we call Multimodal Mixtures of Experts (MMOE).The key idea in MMOE is to train separate expert models for each type of multimodal interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both modalities are fused.On a sarcasm detection task (MUStARD) and a humor detection task (URFunny), we obtain new state-of-the-art results.MMOE is also able to be applied to various types of models to gain improvement.
Haofei Yu, Zhengyang Qi, Lawrence Jang, Ruslan Salakhutdinov, Louis-Philippe Morency, Paul Pu Liang
EMNLP6
2024 Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
abstract
In many machine learning systems that jointly learn from multiple modalities, a core research question is to understand the nature of multimodal interactions: how modalities combine to provide new task-relevant information that was not present in either alone. We study this challenge of interaction quantification in a semi-supervised setting with only labeled unimodal data and naturally co-occurring multimodal data (e.g., unlabeled images and captions, video and corresponding audio) but when labeling them is time-consuming. Using a precise information-theoretic definition of interactions, our key contribution is the derivation of lower and upper bounds to quantify the amount of multimodal interactions in this semi-supervised setting. We propose two lower bounds: one based on the shared information between modalities and the other based on disagreement between separately trained unimodal classifiers, and derive an upper bound through connections to approximate algorithms for min-entropy couplings. We validate these estimated bounds and show how they accurately track true interactions. Finally, we show how these theoretical results can be used to estimate multimodal model performance, guide data collection, and select appropriate multimodal models for various tasks.
Paul Pu Liang, Chun Kai Ling, Alex Obolenskiy, Rohan Pandey, Alex Wilf, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR1
2024 HEMM: Holistic Evaluation of Multimodal Foundation Models
abstract
Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study progress in multimodal foundation models, given the range of possible modeling decisions, tasks, and domains. In this paper, we introduce Holistic Evaluation of Multimodal Models (HEMM) to systematically evaluate the capabilities of multimodal foundation models across a set of 3 dimensions: basic skills, information flow, and real-world use cases. Basic multimodal skills are internal abilities required to solve problems, such as learning interactions across modalities, fine-grained alignment, multi-step reasoning, and the ability to handle external knowledge. Information flow studies how multimodal content changes during a task through querying, translation, editing, and fusion. Use cases span domain-specific challenges introduced in real-world multimedia, affective computing, natural sciences, healthcare, and human-computer interaction applications. Through comprehensive experiments across the 30 tasks in HEMM, we (1) identify key dataset dimensions (e.g., basic skills, information flows, and use cases) that pose challenges to today’s models, and (2) distill performance trends regarding how different modeling dimensions (e.g., scale, pre-training data, multimodal alignment, pre-training, and instruction tuning objectives) influence performance. Our conclusions regarding challenging multimodal interactions, use cases, and tasks requiring reasoning and external knowledge, the benefits of data and model scale, and the impacts of instruction-tuning yield actionable insights for future work in multimodal foundation models.
Paul Pu Liang, Akshay Goindani, Talha Chafekar, Leena Mathur, Haofei Yu, Ruslan Salakhutdinov, Louis-Philippe Morency
NeurIPS1
2023 Demystify the Gravity Well in the Optimization Landscape (Student Abstract)
abstract
We provide both empirical and theoretical insights to demystify the gravity well phenomenon in the optimization landscape. We start from describe the problem setup and theoretical results (an escape time lower bound) of the Softmax Gravity Well (SGW) in the literature. Then we move toward the understanding of a recent observation called ASR gravity well. We provide an explanation of why normal distribution with high variance can lead to suboptimal plateaus from an energy function point of view. We also contribute to the empirical insights of curriculum learning by comparison of policy initialization by different normal distributions. Furthermore, we provide the ASR escape time lower bound to understand the ASR gravity well theoretically. Future work includes more specific modeling of the reward as a function of time and quantitative evaluation of normal distribution’s influence on policy initialization.
Jason Xiaotian Dou, Runxue Bao, Susan Song, Shuran Yang, Yanfu Zhang, Paul Pu Liang, Haiyi Harry Mao
AAAI6
2023 Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
abstract
Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL (1)3
2023 Face-to-Face Contrastive Learning for Social Intelligence Question-Answering
abstract
Creating artificial social intelligence – algorithms that can understand the nuances of multi-person interactions – is an exciting and emerging challenge in processing facial expressions and gestures from multimodal videos. Recent multimodal methods have set the state of the art on many tasks, but have difficulty modeling the complex face-to-face conversational dynamics across speaking turns in social interaction, particularly in a self-supervised setup. In this paper, we propose Face-to-Face Contrastive Learning (F2F-CL), a graph neural network designed to model social interactions using factorization nodes to contextualize the multimodal face-to-face interaction along the boundaries of the speaking turn. With the F2F-CL model, we propose to perform contrastive learning between the factorization nodes of different speaking turns within the same video. We experimentally evaluate our method on the challenging Social-IQ dataset and show state-of-the-art results.
Alex Wilf, Martin Q. Ma, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
FG3
2023 Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational Videos
abstract
Many educational videos use slide presentations, a sequence of visual pages that contain text and figures accompanied by spoken language, which are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and psychology attribute the effectiveness of lecture presentations to their multimodal nature. As a step toward developing vision-language models to aid in student learning as intelligent teacher assistants, we introduce the Lecture Presentations Multimodal (LPM) Dataset as a large-scale benchmark testing the capabilities of vision-and-language models in multimodal understanding of educational videos. Our dataset contains aligned slides and spoken language, for 180+ hours of video and 9000+ slides, with 10 lecturers from various subjects (e.g., computer science, dentistry, biology). We introduce three research tasks, (1) figure-to-text retrieval, (2) text-to-figure retrieval, and (3) generation of slide explanations, which are grounded in multimedia learning and psychology principles to test a vision-language model’s understanding of multimodal content. We provide manual annotations to help implement these tasks and establish baselines on them. Comparing baselines and human student performances, we find that state-of-the-art vision-language models (zero-shot and fine-tuned) struggle in (1) weak crossmodal alignment between slides and spoken text, (2) learning novel visual mediums, (3) technical language, and (4) long-range sequences. We introduce PolyViLT, a novel multimodal transformer trained with a multi-instance learning loss that is more effective than current approaches for retrieval. We conclude by shedding light on the challenges and opportunities in multimodal understanding of educational presentation videos.
Dong Won Lee 0007, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu, Louis-Philippe Morency
ICCV3
2023 MultiViz: Towards Visualizing and Understanding Multimodal Models
Paul Pu Liang, Yiwei Lyu 0001, Gunjan Chhablani, Nihal Jain, Xingbo Wang 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR1
2023 Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework
abstract
The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interactions that are necessary to solve a multimodal task? Subsequently, what are the most suitable multimodal models to capture these interactions? To answer these questions, we propose an information-theoretic approach to quantify the degree of redundancy, uniqueness, and synergy relating input modalities with an output task. We term these three measures as the PID statistics of a multimodal distribution (or PID for short), and introduce two new estimators for these PID statistics that scale to high-dimensional distributions. To validate PID estimation, we conduct extensive experiments on both synthetic datasets where the PID is known and on large-scale multimodal benchmarks where PID estimations are compared with human annotations. Finally, we demonstrate their usefulness in (1) quantifying interactions within multimodal datasets, (2) quantifying interactions captured by multimodal models, (3) principled approaches for model selection, and (4) three real-world case studies engaging with domain experts in pathology, mood prediction, and robotic perception where our framework helps to recommend strong multimodal models for each application.
Paul Pu Liang, Chun Kai Ling, Suzanne Nie, Richard J. Chen, Nicholas B. Allen, Randy Auerbach, Faisal Mahmood 0001, Ruslan Salakhutdinov, Louis-Philippe Morency
NeurIPS1
2023 Factorized Contrastive Learning: Going Beyond Multi-view Redundancy
abstract
In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FactorCL, a new multimodal representation learning method to go beyond multi-view redundancy. FactorCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FactorCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.
Paul Pu Liang, Martin Q. Ma, James Zou 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
NeurIPS1
2023 Localized Symbolic Knowledge Distillation for Visual Commonsense Models
abstract
Instruction following vision-language (VL) models offer a flexible interface that supports a broad range of multimodal tasks in a zero-shot fashion. However, interfaces that operate on full images do not directly enable the user to “point to" and access specific regions within images. This capability is important not only to support reference-grounded VL benchmarks, but also, for practical applications that require precise within-image reasoning. We build Localized Visual Commonsense model which allows users to specify (multiple) regions- as-input. We train our model by sampling localized commonsense knowledge from a large language model (LLM): specifically, we prompt a LLM to collect commonsense knowledge given a global literal image description and a local literal region description automatically generated by a set of VL models. This pipeline is scalable and fully automatic, as no aligned or human-authored image and text pairs are required. With a separately trained critic model that selects high quality examples, we find that training on the localized commonsense corpus expanded solely from images can successfully distill existing VL models to support a reference-as-input interface. Empirical results and human evaluations in zero-shot settings demonstrate that our distillation method results in more precise VL models of reasoning compared to a baseline of passing a generated referring expression.
Jack Hessel, Khyathi Raghavi Chandu, Paul Pu Liang, Ximing Lu, Peter West, Youngjae Yu, Qiuyuan Huang, Jianfeng Gao 0001, Ali Farhadi, Yejin Choi 0001
NeurIPS4
2023 Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
abstract
High sample complexity has long been a challenge for RL. On the other hand, humans learn to perform tasks not only from interaction or demonstrations, but also by reading unstructured text documents, e.g., instruction manuals. Instruction manuals and wiki pages are among the most abundant data that could inform agents of valuable features and policies or task-specific environmental dynamics and reward structures. Therefore, we hypothesize that the ability to utilize human-written instruction manuals to assist learning policies for specific tasks should lead to a more efficient and better-performing agent. We propose the Read and Reward framework. Read and Reward speeds up RL algorithms on Atari games by reading manuals released by the Atari game developers. Our framework consists of a QA Extraction module that extracts and summarizes relevant information from the manual and a Reasoning module that evaluates object-agent interactions based on information from the manual. An auxiliary reward is then provided to a standard A2C RL agent, when interaction is detected. Experimentally, various RL algorithms obtain significant improvement in performance and training speed when assisted by our design. Code at github.com/Holmeswww/RnR
Yue Wu 0001, Yewen Fan, Paul Pu Liang, Amos Azaria, Yuanzhi Li, Tom M. Mitchell
NeurIPS3
2023 MultiZoo and MultiBench: A Standardized Toolkit for Multimodal Deep Learning
abstract
Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robustness, we release MultiZoo, a public toolkit consisting of standardized implementations of >20 core multimodal algorithms and MultiBench, a large-scale benchmark spanning 15 datasets, 10 modalities, 20 prediction tasks, and 6 research areas. Together, these provide an automated end-to-end machine learning pipeline that simplifies and standardizes data loading, experimental setup, and model evaluation. To enable holistic evaluation, we offer a comprehensive methodology to assess (1) generalization, (2) time and space complexity, and (3) modality robustness. MultiBench paves the way towards a better understanding of the capabilities and limitations of multimodal models, while ensuring ease of use, accessibility, and reproducibility. Our toolkits are publicly available, will be regularly updated, and welcome inputs from the community.
Paul Pu Liang, Yiwei Lyu 0001, Arav Agarwal, Louis-Philippe Morency, Ruslan Salakhutdinov
J. Mach. Learn. Res.1
2022 DIME: Fine-grained Interpretations of Multimodal Models via Disentangled Local Explanations
abstract
The ability for a human to understand an Artificial Intelligence (AI) model's decision-making process is critical in enabling stakeholders to visualize model behavior, perform model debugging, promote trust in AI models, and assist in collaborative human-AI decision-making. As a result, the research fields of interpretable and explainable AI have gained traction within AI communities as well as interdisciplinary scientists seeking to apply AI in their subject areas. In this paper, we focus on advancing the state-of-the-art in interpreting multimodal models - a class of machine learning methods that tackle core challenges in representing and capturing interactions between heterogeneous data sources such as images, text, audio, and time-series data. Multimodal models have proliferated numerous real-world applications across healthcare, robotics, multimedia, affective computing, and human-computer interaction. By performing model disentanglement into unimodal contributions (UC) and multimodal interactions (MI), our proposed approach, DIME, enables accurate and fine-grained analysis of multimodal models while maintaining generality across arbitrary modalities, model architectures, and tasks. Through a comprehensive suite of experiments on both synthetic and real-world multimodal tasks, we show that DIME generates accurate disentangled explanations, helps users of multimodal models gain a deeper understanding of model behavior, and presents a step towards debugging and improving these models for real-world deployment.
Yiwei Lyu 0001, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
AIES2
2022 Rethinking Architecture Design for Tackling Data Heterogeneity in Federated Learning
abstract
Federated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential for catastrophic forgetting across real-world heterogeneous devices. In this paper, we demonstrate that self-attention-based architectures (e.g., Transformers) are more robust to distribution shifts and hence improve federated learning over heterogeneous data. Concretely, we conduct the first rigorous empirical investigation of different neural architectures across a range of federated algorithms, real-world benchmarks, and heterogeneous data splits. Our experiments show that simply replacing convolutional networks with Transformers can greatly reduce catastrophic forgetting of previous devices, accelerate convergence, and reach a better global model, especially when dealing with heterogeneous data. We release our code and pretrained models to encourage future exploration in robust architectures as an alternative to current research efforts on the optimization front.
Liangqiong Qu, Yuyin Zhou, Paul Pu Liang, Yingda Xia, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Daniel L. Rubin
CVPR3
2022 PACS: A Dataset for Physical Audiovisual CommonSense Reasoning
Samuel Yu, Peter Wu, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
ECCV (37)3
2021 Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile Data
abstract
Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nick Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nicholas B. Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL/IJCNLP (1)1
2021 Anchor & Transform: Learning Sparse Embeddings for Large Vocabularies
Paul Pu Liang, Manzil Zaheer, Amr Ahmed 0001
ICLR1
2021 Towards Understanding and Mitigating Social Biases in Language Models
abstract
As machine learning methods are deployed in real-world settings such as healthcare, legal systems, and social science, it is crucial to recognize how they shape social biases and stereotypes in these sensitive decision-making processes. Among such real-world deployments are large-scale pretrained language models (LMs) that can be potentially dangerous in manifesting undesirable representational biases - harmful biases resulting from stereotyping that propagate negative generalizations involving gender, race, religion, and other social constructs. As a step towards improving the fairness of LMs, we carefully define several sources of representational biases before proposing new benchmarks and metrics to measure them. With these tools, we propose steps towards mitigating social biases during text generation. Our empirical results and human evaluation demonstrate effectiveness in mitigating bias while retaining crucial contextual information for high-fidelity text generation, thereby pushing forward the performance-fairness Pareto frontier.
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan Salakhutdinov
ICML1
2021 Cross-Modal Generalization: Learning in Low Resource Modalities via Meta-Alignment
abstract
How can we generalize to a new prediction task at test time when it also uses a new modality as input? More importantly, how can we do this with as little annotated data as possible? This problem of cross-modal generalization is a new research milestone with concrete impact on real-world applications. For example, can an AI system start understanding spoken language from mostly written text? Or can it learn the visual steps of a new recipe from only text descriptions? In this work, we formalize cross-modal generalization as a learning paradigm to train a model that can (1) quickly perform new tasks (from new domains) while (2) being originally trained on a different input modality. Such a learning paradigm is crucial for generalization to low-resource modalities such as spoken speech in rare languages while utilizing a different high-resource modality such as text. One key technical challenge that makes it different from other learning paradigms such as meta-learning and domain adaptation is the presence of different source and target modalities which will require different encoders. We propose an effective solution based on meta-alignment, a novel method to align representation spaces using strongly and weakly paired cross-modal data while ensuring quick generalization to new tasks across different modalities. This approach uses key ideas from cross-modal learning and meta-learning, and presents strong results on the cross-modal generalization problem. We benchmark several approaches on 3 real-world classification tasks: few-shot recipe classification from text to images of recipes, object classification from images to audio of objects, and language classification from text to spoken speech across 100 languages spanning many rare languages. Our results demonstrate strong performance even when the new target modality has only a few (1-10) labeled samples and in the presence of noisy labels, a scenario particularly prevalent in low-resource modalities.
Paul Pu Liang, Peter Wu, Liu Ziyin 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ACM Multimedia1
2021 StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer
abstract
Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yiwei Lyu 0001, Paul Pu Liang, Hai Pham, Eduard H. Hovy, Barnabás Póczos, Ruslan Salakhutdinov, Louis-Philippe Morency
NAACL-HLT2
2020 Towards Debiasing Sentence Representations
abstract
As natural language processing methods are increasingly deployed in real-world scenarios such as healthcare, legal systems, and social science, it becomes necessary to recognize the role they potentially play in shaping social biases and stereotypes.Previous work has revealed the presence of social biases in widely used word embeddings involving gender, race, religion, and other social constructs.While some methods were proposed to debias these word-level embeddings, there is a need to perform debiasing at the sentence-level given the recent shift towards new contextualized sentence representations such as ELMo and BERT.In this paper, we investigate the presence of social biases in sentence-level representations and propose a new method, SENT-DEBIAS, to reduce these biases.We show that SENT-DEBIAS is effective in removing biases, and at the same time, preserves performance on sentence-level downstream tasks such as sentiment analysis, linguistic acceptability, and natural language understanding.We hope that our work will inspire future research on characterizing and removing social biases from widely adopted sentence representations for fairer NLP.
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL1
2020 Diverse and Admissible Trajectory Forecasting Through Multimodal Context Understanding
Seong Hyeon Park, Gyubok Lee, Jimin Seo, Manoj Bhat, Jonathan Francis, Ashwin R. Jadhav, Paul Pu Liang, Louis-Philippe Morency
ECCV (11)8
2020 CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French
abstract
Modeling multimodal language is a core research area in natural language processing. While languages such as English have relatively large multimodal language resources, other widely spoken languages across the globe have few or no large-scale datasets in this area. This disproportionately affects native speakers of languages other than English. As a step towards building more equitable and inclusive multimodal systems, we introduce the first large-scale multimodal language dataset for Spanish, Portuguese, German and French. The proposed dataset, called CMU-MOSEAS (CMU Multimodal Opinion Sentiment, Emotions and Attributes), is the largest of its kind with 40, 000 total labelled sentences. It covers a diverse set topics and speakers, and carries supervision of 20 labels including sentiment (and subjectivity), emotions, and attributes. Our evaluations on a state-of-the-art multimodal model demonstrates that CMU-MOSEAS enables further research for multilingual studies in multimodal language.
Amir Zadeh 0001, Yansheng Cao, Smon Hessner, Paul Pu Liang, Soujanya Poria, Louis-Philippe Morency
EMNLP (1)4
2019 Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities
abstract
Multimodal sentiment analysis is a core research area that studies speaker sentiment expressed from the language, visual, and acoustic modalities. The central challenge in multimodal learning involves inferring joint representations that can process and relate information from these modalities. However, existing work learns joint representations by requiring all modalities as input and as a result, the learned representations may be sensitive to noisy or missing modalities at test time. With the recent success of sequence to sequence (Seq2Seq) models in machine translation, there is an opportunity to explore new ways of learning joint representations that may not require all input modalities at test time. In this paper, we propose a method to learn robust joint representations by translating between modalities. Our method is based on the key insight that translation from a source to a target modality provides a method of learning joint representations using only the source modality as input. We augment modality translations with a cycle consistency loss to ensure that our joint representations retain maximal information from all modalities. Once our translation model is trained with paired multimodal data, we only need data from the source modality at test time for final sentiment prediction. This ensures that our model remains robust from perturbations or missing information in the other modalities. We train our model with a coupled translationprediction objective and it achieves new state-of-the-art results on multimodal sentiment analysis datasets: CMU-MOSI, ICTMMMO, and YouTube. Additional experiments show that our model learns increasingly discriminative joint representations with more input modalities while maintaining robustness to missing or perturbed modalities.
Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, Barnabás Póczos
AAAI2
2019 Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors
abstract
Humans convey their intentions through the usage of both verbal and nonverbal behaviors during face-to-face communication. Speaker intentions often vary dynamically depending on different nonverbal contexts, such as vocal patterns and facial expressions. As a result, when modeling human language, it is essential to not only consider the literal meaning of the words but also the nonverbal contexts in which these words appear. To better model human language, we first model expressive nonverbal representations by analyzing the fine-grained visual and acoustic patterns that occur during word segments. In addition, we seek to capture the dynamic nature of nonverbal intents by shifting word representations based on the accompanying nonverbal behaviors. To this end, we propose the Recurrent Attended Variation Embedding Network (RAVEN) that models the fine-grained structure of nonverbal subword sequences and dynamically shifts word representations based on nonverbal cues. Our proposed model achieves competitive performance on two publicly available datasets for multimodal sentiment analysis and emotion recognition. We also visualize the shifted word representations in different nonverbal contexts and summarize common patterns regarding multimodal variations of word representations.
Yansen Wang, Ying Shen 0006, Zhun Liu, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
AAAI4
2019 Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization
abstract
There has been an increased interest in multimodal language processing including multimodal dialog, question answering, sentiment analysis, and speech recognition.However, naturally occurring multimodal data is often imperfect as a result of imperfect modalities, missing entries or noise corruption.To address these concerns, we present a regularization method based on tensor rank minimization.Our method is based on the observation that high-dimensional multimodal time series data often exhibit correlations across time and modalities which leads to low-rank tensor representations.However, the presence of noise or incomplete values breaks these correlations and results in tensor representations of higher rank.We design a model to learn such tensor representations and effectively regularize their rank.Experiments on multimodal language data show that our model achieves good results across various levels of imperfection.
Paul Pu Liang, Zhun Liu, Yao-Hung Tsai, Qibin Zhao, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL (1)1
2019 Multimodal Transformer for Unaligned Multimodal Language Sequences
abstract
Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data non-alignment due to variable sampling rates for the sequences from each modality; and 2) long-range dependencies between elements across modalities. In this paper, we introduce the Multimodal Transformer (MulT) to generically address the above issues in an end-to-end manner without explicitly aligning the data. At the heart of our model is the directional pairwise cross-modal attention, which attends to interactions between multimodal sequences across distinct time steps and latently adapt streams from one modality to another. Comprehensive experiments on both aligned and non-aligned multimodal time-series show that our model outperforms state-of-the-art methods by a large margin. In addition, empirical analysis suggests that correlated crossmodal signals are able to be captured by the proposed crossmodal attention mechanism in MulT.
Yao-Hung Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov
ACL (1)3
2019 Social-IQ: A Question Answering Benchmark for Artificial Social Intelligence
abstract
As intelligent systems increasingly blend into our everyday life, artificial social intelligence becomes a prominent area of research. Intelligent systems must be socially intelligent in order to comprehend human intents and maintain a rich level of interaction with humans. Human language offers a unique unconstrained approach to probe through questions and reason through answers about social situations. This unconstrained approach extends previous attempts to model social intelligence through numeric supervision (e.g. sentiment and emotions labels). In this paper, we introduce Social-IQ, a unconstrained benchmark specifically designed to train and evaluate socially intelligent technologies. By providing a rich source of open-ended questions and answers, Social-IQ opens the door to explainable social intelligence. The dataset contains rigorously annotated and validated videos, questions and answers, as well as annotations for the complexity level of each question and answer. Social-IQ contains 1,250 natural in-the-wild social situations, 7,500 questions and 52,500 correct and incorrect answers. Although humans can reason about social situations with very high accuracy (95.08%), existing state-of-the-art computational models struggle on this task. As a result, Social-IQ brings novel challenges that will spark future research in social intelligence modeling, visual reasoning, and multimodal question answering (QA).
Amir Zadeh 0001, Paul Pu Liang, Edmund Tong, Louis-Philippe Morency
CVPR3
2019 Learning Factorized Multimodal Representations
Yao-Hung Tsai, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR (Poster)2
2019 Deep Gamblers: Learning to Abstain with Portfolio Theory
abstract
We deal with the selective classification problem (supervised-learning problem with a rejection option), where we want to achieve the best performance at a certain level of coverage of the data. We transform the original $m$-class classification problem to (m+1)-class where the (m+1)-th class represents the model abstaining from making a prediction due to disconfidence. Inspired by portfolio theory, we propose a loss function for the selective classification problem based on the doubling rate of gambling. Minimizing this loss function corresponds naturally to maximizing the return of a horse race, where a player aims to balance between betting on an outcome (making a prediction) when confident and reserving one's winnings (abstaining) when not confident. This loss function allows us to train neural networks and characterize the disconfidence of prediction in an end-to-end fashion. In comparison with previous methods, our method requires almost no modification to the model inference algorithm or model architecture. Experiments show that our method can identify uncertainty in data points, and achieves strong results on SVHN and CIFAR10 at various coverages of the data.
Liu Ziyin 0001, Zhikang Wang, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency, Masahito Ueda
NeurIPS3
2018 Memory Fusion Network for Multi-view Sequential Learning
abstract
Multi-view sequential learning is a fundamental problem in machine learning dealing with multi-view sequences. In a multi-view sequence, there exists two forms of interactions between different views: view-specific interactions and cross-view interactions. In this paper, we present a new neural architecture for multi-view sequential learning called the Memory Fusion Network (MFN) that explicitly accounts for both interactions in a neural architecture and continuously models them through time. The first component of the MFN is called the System of LSTMs, where view-specific interactions are learned in isolation through assigning an LSTM function to each view. The cross-view interactions are then identified using a special attention mechanism called the Delta-memory Attention Network (DMAN) and summarized through time with a Multi-view Gated Memory. Through extensive experimentation, MFN is compared to various proposed approaches for multi-view sequential learning on multiple publicly available benchmark datasets. MFN outperforms all the multi-view approaches. Furthermore, MFN outperforms all current state-of-the-art models, setting new state-of-the-art results for all three multi-view datasets.
Amir Zadeh 0001, Paul Pu Liang, Navonil Majumder, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
AAAI2
2018 Multi-attention Recurrent Network for Human Communication Comprehension
abstract
Human face-to-face communication is a complex multimodal signal. We use words (language modality), gestures (vision modality) and changes in tone (acoustic modality) to convey our intentions. Humans easily process and understand face-to-face communication, however, comprehending this form of communication remains a significant challenge for Artificial Intelligence (AI). AI must understand each modality and the interactions between them that shape the communication. In this paper, we present a novel neural architecture for understanding human communication called the Multi-attention Recurrent Network (MARN). The main strength of our model comes from discovering interactions between modalities through time using a neural component called the Multi-attention Block (MAB) and storing them in the hybrid memory of a recurrent component called the Long-short Term Hybrid Memory (LSTHM). We perform extensive comparisons on six publicly available datasets for multimodal sentiment analysis, speaker trait recognition and emotion recognition. MARN shows state-of-the-art results performance in all the datasets.
Amir Zadeh 0001, Paul Pu Liang, Soujanya Poria, Prateek Vij, Erik Cambria, Louis-Philippe Morency
AAAI2
2018 Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
abstract
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Amir Zadeh 0001, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
ACL (1)2
2018 Efficient Low-rank Multimodal Fusion With Modality-Specific Factors
abstract
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, Louis-Philippe Morency. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Zhun Liu, Ying Shen 0006, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
ACL (1)4
2018 An Empirical Evaluation of Sketched SVD and its Application to Leverage Score Ordering
abstract
The power of randomized algorithms in numerical methods have led to fast solutions which use the Singular Value Decomposition (SVD) as a core routine. However, given the large data size of modern and the modest runtime of SVD, most practical algorithms would require some form of approximation, such as sketching, when running SVD. While these approximation methods satisfy many theoretical guarantees, we provide the first algorithmic implementations for sketch-and-solve SVD problems on real-world, large-scale datasets. We provide a comprehensive empirical evaluation of these algorithms and provide guidelines on how to ensure accurate deployment to real-world data. As an application of sketched SVD, we present Sketched Leverage Score Ordering, a technique for determining the ordering of data in the training of neural networks. Our technique is based on the distributed computation of leverage scores using random projections. These computed leverage scores provide a flexible and efficient method to determine the optimal ordering of training data without manual intervention or annotations. We present empirical results on an extensive set of experiments across image classification, language sentiment analysis, and multi-modal sentiment analysis. Our method is faster compared to standard randomized projection algorithms and shows improvements in convergence and results.
Hui Han Chin, Paul Pu Liang
ACML2
2018 Multimodal Language Analysis with Recurrent Multistage Fusion
abstract
Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities.Comprehending multimodal language requires modeling not only the interactions within each modality (intra-modal interactions) but more importantly the interactions between modalities (cross-modal interactions).In this paper, we propose the Recurrent Multistage Fusion Network (RMFN) which decomposes the fusion problem into multiple stages, each of them focused on a subset of multimodal signals for specialized, effective fusion.Crossmodal interactions are modeled using this multistage fusion approach which builds upon intermediate representations of previous stages.Temporal and intra-modal interactions are modeled by integrating our proposed fusion approach with a system of recurrent neural networks.The RMFN displays state-of-the-art performance in modeling human multimodal language across three public datasets relating to multimodal sentiment analysis, emotion recognition, and speaker traits recognition.We provide visualizations to show that each stage of fusion focuses on a different subset of multimodal signals, learning increasingly discriminative multimodal representations.
Paul Pu Liang, Liu Ziyin 0001, Amir Zadeh 0001, Louis-Philippe Morency
EMNLP1
2018 Robust Modulation Classification under Uncertain Noise Condition Using Recurrent Neural Network
abstract
Modulation classification using deep neural networks has recently received increasing attention due to its capability in learning rich features of data. In this paper, we propose a low- complexity blind data-driven modulation classifier. Our classifier operates robustly over Rayleigh fading channels under uncertain noise conditions modeled using a mixture of three types of noise, namely, white Gaussian noise, white non- Gaussian noise and correlated non-Gaussian noise. The proposed classifier consists of several layers of recurrent neural networks (RNN) which is well-suited for learning representations from time-correlated data. The classifier is trained using the labeled raw signal samples generated under different noise conditions. Simulation results show that the performance of our proposed classifier approaches that of maximum likelihood classifiers with perfect channel knowledge and outperforms existing expectation maximum (EM) and expectation conditional maximum (ECM) classifiers which iteratively estimate channel and noise parameters.
Shisheng Hu, Yiyang Pei, Paul Pu Liang, Ying-Chang Liang
GLOBECOM3
2018 A Machine Learning Approach to MIMO Communications
abstract
Inspired by the phenomenon that the received signals naturally form clusters, we propose a novel machine learning framework to design multi-input multi-output (MIMO) communication systems. In the proposed framework, the MIMO detection problem is converted into a clustering problem, and known labels are transmitted to assist the receiver for labeling the clusters. A modulation-constrained Gaussian mixture model (MC-GMM) and the associated optimization algorithm are developed to reduce the number of parameters to be learnt in the clustering algorithm. Furthermore, we propose a method called label reconstruction to minimize the overhead of label transmission, and the design of the optimal labels is studied. Simulation results are presented to verify the effectiveness of the proposed label-assisted clustering (LAC) receiver in approaching the optimal maximum likelihood detection (MLD) with perfectly known channel knowledge for typical MIMO systems.
Yudi Huang, Paul Pu Liang, Qianqian Zhang 0001, Ying-Chang Liang
ICC2
2018 Multimodal Local-Global Ranking Fusion for Emotion Recognition
abstract
Emotion recognition is a core research area at the intersection of artificial intelligence and human communication analysis. It is a significant technical challenge since humans display their emotions through complex idiosyncratic combinations of the language, visual and acoustic modalities. In contrast to traditional multimodal fusion techniques, we approach emotion recognition from both direct person-independent and relative person-dependent perspectives. The direct person-independent perspective follows the conventional emotion recognition approach which directly infers absolute emotion labels from observed multimodal features. The relative person-dependent perspective approaches emotion recognition in a relative manner by comparing partial video segments to determine if there was an increase or decrease in emotional intensity. Our proposed model integrates these direct and relative prediction perspectives by dividing the emotion recognition task into three easier subtasks. The first subtask involves a multimodal local ranking of relative emotion intensities between two short segments of a video. The second subtask uses local rankings to infer global relative emotion ranks with a Bayesian ranking algorithm. The third subtask incorporates both direct predictions from observed multimodal behaviors and relative emotion ranks from local-global rankings for final emotion prediction. Our approach displays excellent performance on an audio-visual emotion recognition benchmark and improves over other algorithms for multimodal fusion.
Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
ICMI1
2017 Multimodal sentiment analysis with word-level fusion and reinforcement learning
abstract
With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment analysis which focus on holistic information in speech segments such as bag of words representations and average facial expression intensity, we propose a novel deep architecture for multimodal sentiment analysis that is able to perform modality fusion at the word level. In this paper, we propose the Gated Multimodal Embedding LSTM with Temporal Attention (GME-LSTM(A)) model that is composed of 2 modules. The Gated Multimodal Embedding allows us to alleviate the difficulties of fusion when there are noisy modalities. The LSTM with Temporal Attention can perform word level fusion at a finer fusion resolution between the input modalities and attends to the most important time steps. As a result, the GME-LSTM(A) is able to better model the multimodal structure of speech through time and perform better sentiment comprehension. We demonstrate the effectiveness of this approach on the publicly-available Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis (CMU-MOSI) dataset by achieving state-of-the-art sentiment classification and regression results. Qualitative analysis on our model emphasizes the importance of the Temporal Attention Layer in sentiment prediction because the additional acoustic and visual modalities are noisy. We also demonstrate the effectiveness of the Gated Multimodal Embedding in selectively filtering these noisy modalities out. These results and analysis open new areas in the study of sentiment analysis in human communication and provide new models for multimodal fusion.
Minghai Chen, Paul Pu Liang, Tadas Baltrusaitis, Amir Zadeh 0001, Louis-Philippe Morency
ICMI3