VLDB 2026 Research / reviewers in the wild / expert
Boyang Li 0001
dblp:70/1211-1 · also Boyang "Albert" Li
· DBLP profile ↗
63ranked-venue papers
9as first author
39since 2021 · last 2026
0000-0002-6230-2376ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 5 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Animate Images from A Few Videos to Portray Delicate Human ActionsabstractDespite recent progress, video generative models still struggle to generate delicate human actions (e.g., gymnastics), particularly when they are required to start from a user-provided reference image. In this paper, we explore the task of learning to animate images into videos that portray delicate human actions using a small number of videos — 16 or fewer — which reduces the need for extensive data collection and enhances practicality for real-world applications. Learning generalizable motion patterns that smoothly transition from user-provided reference images in such a few-shot setting is highly challenging. We propose FLASH (Few-shot Learning to Animate and Steer Humans), which enhances generalization of motion by training the model to reconstruct a video using the motion features and cross-frame correspondences extracted from another video with the same motion but different appearance. This encourages the learning of transferable motion and mitigates overfitting to the appearance in limited training data. Additionally, FLASH extends the decoder with additional layers to propagate details from the reference image to generated frames, improving transition smoothness. Human judges significantly favor FLASH, with 65.78% of 488 responses prefer FLASH over baselines. We strongly recommend watching the videos on the webpage1, as motion artifacts are hard to notice from images. Haoxin Li, Yingchen Yu, Hanwang Zhang, Song Bai 0001, Boyang Li 0001 |
WACV | 6 |
| 2026 | GELD: A unified neural model for efficiently solving traveling salesman problems across different scales
Yubin Xiao, Di Wang 0004, Xuan Wu 0004, Boyang Li 0001, You Zhou 0008 |
Pattern Recognit. | 5 |
| 2026 | Copycat vs. Original: Multi-Modal Pretraining and Variable Importance in Box-Office PredictionabstractMovie production and investment are associated with a high level of risk, motivating machine learning research to predict box-office revenue. Furthermore, identifying variables that have a significant influence on box-office revenue may aid in human decision-making. In this study, we collect a large movie dataset, including user-generated keywords and movie posters, and integrate these modalities to better predict box-office revenue. We utilize visual information from movie posters to visually ground the movie keywords, thereby acquiring more semantically precise text representations, resulting in a substantial 14.5% enhancement in box-office prediction accuracy. Also, we develop metrics to quantify content similarity based on the keywords, facilitating the identification of “copycat movies,” a term that can be extended beyond traditional sequels and franchise movies. Subsequently, we analyze the importance of copycat features in box-office revenue prediction using two explanatory methods: Attention Rollout and LIME. Our analyses show the importance of copycat features in box-office prediction and reveal a positive relationship between copycat movies and box-office revenues. However, this effect diminishes with an increase in the number of similar movies and the similarity of their content. Overall, our work establishes a comprehensive process of predicting movie box-office revenue by utilizing multi-modal data and providing valuable business insights. Qin Chao, Eunsoo Kim, Boyang Li 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical EvaluationabstractCurrent vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessary for human-like understanding and real-world applications. To address this gap, we develop SPHERE (Spatial Perception and Hierarchical Evaluation of REasoning), a hierarchical evaluation framework supported by a new human-annotated dataset. SPHERE systematically probes models across increasing levels of complexity, from fundamental skills to multi-skill integration and high-level reasoning that combines spatial, visual, and logical understanding. Benchmark evaluation of state-of-the-art models reveals significant deficiencies, especially in reasoning about distance and proximity, understanding both egocentric and allocentric perspectives, and applying spatial logic in physical contexts. These findings expose critical blind spots in existing models and underscore the need for more advanced spatial reasoning techniques, driving the development of vision-language models that align more closely with human spatial cognition. Wei En Ng, Lixin Ma, Junqi Zhao, Allison Koenecke, Boyang Li 0001, Lu Wang 0003 |
ACL (1) | 7 |
| 2025 | Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold LabelsabstractLarge Language Models (LLMs) have demonstrated remarkable performance through supervised fine-tuning or in-context learning using gold labels. However, this paradigm is limited by the availability of gold labels, while in certain scenarios, LLMs may need to perform tasks that are too complex for humans to provide such labels. To tackle this challenge, this study explores whether solely utilizing unlabeled data can elicit strong model capabilities. We propose a new paradigm termed zero-to-strong generalization. We iteratively prompt LLMs to annotate unlabeled data and retain high-quality labels by filtering. Surprisingly, we obverse that this iterative process gradually unlocks LLMs’ potential on downstream tasks. Our experiments on extensive classification and reasoning tasks confirm the effectiveness of our proposed framework. Our analysis indicates that this paradigm is effective for both in-context learning and fine-tuning, and for various model sizes. Chaoqun Liu, Qin Chao, Wenxuan Zhang 0001, Xiaobao Wu, Boyang Li 0001, Anh Tuan Luu, Lidong Bing |
COLING | 5 |
| 2025 | Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable EventsabstractThe commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios [1], [23], [42], making it difficult to discern whether model performance stems from keen perception and reasoning skills, or reliance on pure statistical recall. We argue that by focusing on atypical events in videos, clearer insights can be gained on the core capabilities of VLMs. Explaining and understanding such out-of-distribution events requires models to extend beyond basic pattern recognition and regurgitation of their prior knowledge. To this end, we introduce Black-SwanSuite, a benchmark for evaluating VLMs’ ability to reason about unexpected events through abductive and defeasible tasks. Our tasks artificially limit the amount of visual information provided to models while questioning them about hidden unexpected events, or provide new visual information that could change an existing hypothesis about the event. We curate a comprehensive benchmark suite comprising over 3,800 MCQ, 4,900 generative and 6,700 yes/no questions, spanning 1,655 videos. After extensively evaluating various state-of-the-art VLMs, including GPT-4o and Gemini 1.5 Pro, as well as open-source VLMs such as LLaVA-Video, we find significant performance gaps of up to 32% from humans on these tasks. Our findings reveal key limitations in current VLMs, emphasizing the need for enhanced model architectures and training strategies. Our data and leaderboard is available at https://blackswan.cs.ubc.ca. Aditya Chinchure, Sahithya Ravi, Raymond T. Ng, Vered Shwartz, Boyang Li 0001, Leonid Sigal |
CVPR | 5 |
| 2025 | Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic DataabstractPaired image-text data with subtle variations in-between (e.g., people holding surfboards vs. people holding shovels) hold the promise of producing Vision-Language Models with proper compositional understanding. Synthesizing such training data from generative models is a highly coveted prize due to the reduced cost of data collection. However, synthesizing training images for compositional learning presents three challenges: (1) efficiency in generating large quantities of images, (2) text alignment between the generated image and the caption in the exact place of the subtle change, and (3) image fidelity in ensuring sufficient similarity with the original real images in all other places. We propose SPARCL (Synthetic Perturbations for Advancing Robust Compositional Learning), which integrates image feature injection into a fast text-to-image generative model, followed by an image style transfer step, to meet the three challenges. Further, to cope with any residual issues of text alignment, we propose an adaptive margin loss to filter out potentially incorrect synthetic samples and focus the learning on informative hard samples. Evaluation on four compositional understanding benchmarks demonstrates that SPARCL significantly improves the compositionality of CLIP, boosting the average accuracy of the CLIP base model by over 8% across all benchmarks and outperforming state-of-the-art methods by 2% on three benchmarks. Haoxin Li, Boyang Li 0001 |
CVPR | 2 |
| 2025 | Synopses of Movie Narratives: a Video-Language Dataset for Story UnderstandingabstractComputational story understanding is a crucial but under-explored area of AI, hampered by a lack of suitable datasets. To address this, we collect, preprocess and publicly release SYMON (Synopses of Movie Narratives), a new video-language dataset containing 5,193 human-narrated, short movie summary videos sourced from YouTube. SYMON features naturalistic storytelling videos for human audiences made by human creators. Compared to existing movie story datasets, the videos in SYMON are shorter yet provide higher coverage of key story events, making it ideal for computational story understanding. We establish benchmarks on story video-text alignment and story video narration generation, demonstrating significant performance improvements when models are trained on SYMON. These results underscore the value of SYMON for advancing research in vision-language story understanding and generation. Qin Chao, Yangfeng Ji, Boyang Li 0001 |
ICME | 4 |
| 2025 | CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingabstractMulti-task model merging offers a promising paradigm for integrating multiple expert models into a unified system without additional training. Existing state-of-the-art techniques, such as Task Arithmetic and its variants, merge models by accumulating task vectors—defined as the parameter differences between pre-trained and fine-tuned models. However, task vector accumulation is often hindered by knowledge conflicts, where conflicting components across different task vectors can lead to performance degradation during the merging process. To address this challenge, we propose Conflict-Aware Task Merging (CAT Merging), a novel training-free framework that selectively trims conflict-prone components from the task vectors. CAT Merging introduces several parameter-specific strategies, including projection for linear weights and masking for scaling and shifting parameters in normalization layers. Extensive experiments on vision and vision-language tasks demonstrate that CAT Merging effectively suppresses knowledge conflicts, achieving average accuracy improvements of up to 4.7% (ViT-B/32) and 2.0% (ViT-L/14) over state-of-the-art methods. Wenju Sun, Qingyong Li, Boyang Li 0001 |
ICML | 4 |
| 2025 | Conversational Explanations: Discussing Explainable AI with Non-AI ExpertsabstractExplainable AI (XAI) aims to provide insights into the decisions made by AI models. To date, most XAI approaches provide only one-time, static explanations, which cannot cater to users' diverse knowledge levels and information needs. Conversational explanations have been proposed as an effective method to customize XAI explanations. However, building conversational explanation systems is hindered by the scarcity of training data. Training with synthetic data faces two main challenges: lack of data diversity and hallucination in the generated data. To alleviate these issues, we introduce a repetition penalty to promote data diversity and exploit a hallucination detector to filter out untruthful synthetic conversation turns. We conducted both automatic and human evaluations on the proposed system, fEw-shot Multi-round ConvErsational Explanation (EMCEE). For automatic evaluation, EMCEE achieves relative improvements of 81.6% in BLEU and 80.5% in ROUGE compared to the baselines. EMCEE also mitigates the degeneration of data quality caused by training on synthetic data. In human evaluations (N = 60), EMCEE outperforms baseline models and the control group in improving users' comprehension, acceptance, trust, and collaboration with static explanations by large margins. Through a fine-grained analysis of model responses, we further demonstrate that training on self-generated synthetic data improves the model's ability to generate more truthful and understandable answers, leading to better user interactions. To the best of our knowledge, this is the first conversational explanation method that can answer free-form user questions following static explanations. Mengao Zhang, Wei Yan Low, Xi Jessie Yang, Boyang Li 0001 |
IUI | 5 |
| 2025 | Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge ConflictsabstractMulti-task model merging offers an efficient solution for integrating knowledge from multiple fine-tuned models, mitigating the significant computational and storage demands associated with multi-task training. As a key technique in this field, Task Arithmetic (TA) defines task vectors by subtracting the pre-trained model (0 pre) from the fine-tuned task models in parameter space, then adjusting the weight between these task vectors and 0 pre to balance task-generalized and task-specific knowledge. Despite the promising performance of TA, conflicts can arise among the task vectors, particularly when different tasks require distinct model adaptations. In this paper, we formally define this issue as knowledge conflicts, characterized by the performance degradation of one task after merging with a model fine-tuned for another task. Through in-depth analysis, we show that these conflicts stem primarily from the components of task vectors that align with the gradient of task-specific losses at 0 pre. To address this, we propose Task Arithmetic in Trust Region (TATR), which defines the trust region as dimensions in the model parameter space that cause only small changes (corresponding to the task vector components with gradient orthogonal direction) in the task-specific losses. Restricting parameter merging within this trust region, TATR can effectively alleviate knowledge conflicts. Moreover, TATR serves as a plug-and-play module compatible with a wide range of TA-based methods. Extensive empirical evaluations on visual and visual-language tasks robustly demonstrate that TATR improves the multi-task performance of several TA-based model merging methods. Wenju Sun, Qingyong Li, Wen Wang 0019, Boyang Li 0001 |
ACM Multimedia | 5 |
| 2025 | Two Causally Related Needles in a Video HaystackabstractProperly evaluating the ability of Video-Language Models (VLMs) to understand long videos remains a challenge. We propose a long-context video understanding benchmark, Causal2Needles, that assesses two crucial abilities insufficiently addressed by existing benchmarks: (1) extracting information from two separate locations (two needles) in a long video and understanding them jointly, and (2) modeling the world in terms of cause and effect in human behaviors. Causal2Needles evaluates these abilities using noncausal one-needle, causal one-needle, and causal two-needle questions. The most complex question type, causal two-needle questions, require extracting information from both the cause and effect events from a long video and the associated narration text. To prevent textual bias, we introduce two complementary question formats: locating the video clip containing the answer, and verbal description of a visual detail from that video clip. Our experiments reveal that models excelling on existing benchmarks struggle with causal 2-needle questions, and the model performance is negatively correlated with the distance between the two needles. These findings highlight critical limitations in current VLMs. Miaoyu Li, Qin Chao, Boyang Li 0001 |
NeurIPS | 3 |
| 2025 | Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationabstractMulti-task model merging aims to consolidate knowledge from multiple fine-tuned task-specific experts into a unified model while minimizing performance degradation. Existing methods primarily approach this by minimizing differences between task-specific experts and the unified model, either from a parameter-level or a task-loss perspective. However, parameter-level methods exhibit a significant performance gap compared to the upper bound, while task-loss approaches entail costly secondary training procedures. In contrast, we observe that performance degradation closely correlates with feature drift, i.e., differences in feature representations of the same sample caused by model merging. Motivated by this observation, we propose Layer-wise Optimal Task Vector Merging (LOT Merging), a technique that explicitly minimizes feature drift between task-specific experts and the unified model in a layer-by-layer manner. LOT Merging can be formulated as a convex quadratic optimization problem, enabling us to analytically derive closed-form solutions for the parameters of linear and normalization layers. Consequently, LOT Merging achieves efficient model consolidation through basic matrix operations. Extensive experiments across vision and vision-language benchmarks demonstrate that LOT Merging significantly outperforms baseline methods, achieving improvements of up to 4.4% (ViT-B/32) over state-of-the-art approaches. The source code is available at https://github.com/SunWenJu123/model-merging. Wenju Sun, Qingyong Li, Wen Wang 0019, Yang Liu 0352, Boyang Li 0001 |
NeurIPS | 6 |
| 2025 | Local Masked Reconstruction for Efficient Self-Supervised Learning on High-Resolution ImagesabstractSelf-supervised learning for computer vision has progressed tremendously and improved many downstream vision tasks, such as image classification, semantic segmentation, and object detection. Among these, generative self-supervised vision learning approaches, such as MAE and BEiT, show promising performance. However, their global reconstruction mechanism is computationally demanding, especially for high-resolution images. The computational cost increases extensively when scaled to a large-scale dataset. To address this issue, we propose local masked reconstruction (LoMaR), a simple yet effective approach that reconstructs image patches from small neighboring regions. The strategy can be easily integrated into any generative self-supervised learning techniques and improves the trade-off between efficiency and accuracy compared to reconstruction over the entire image. LoMaR is$2.5\times faster$than MAE and 5.0x faster than BEiT on$384\times 384$ImageNet pretraining and surpasses them by 0.2% and 0.8% in accuracy, respectively. It is$2.1\times faster$than MAE on iNaturalist pretraining and gains 0.2% in accuracy. On MS COCO, LoMaR outperforms MAE by 0.5$AP^{box}$on object detection and 0.5$AP^{mask}$on instance segmentation. It also outperforms$MAE$by 0.2% on semantic segmentation. Our code and pretrained models are available at: https://github.com/junchen14/LoMaR. Jun Chen 0021, Faizan Farooq Khan, Ammar Sherif, ZongYuan Ge, Boyang Li 0001, Mohamed Elhoseiny 0001 |
WACV | 6 |
| 2025 | May I Ask a Follow-up Question? Understanding the Benefits of Conversations in Neural Network ExplainabilityabstractResearch in explainable AI (XAI) aims to provide insights into the decision-making process of opaque AI models. To date, most XAI methods offer one-off and static explanations, which cannot cater to the diverse backgrounds and understanding levels of users. With this paper, we investigate if free-form conversations can enhance users’ comprehension of static explanations in image classification, improve acceptance and trust in the explanation methods, and facilitate human-AI collaboration. We conduct a human-subject experiment with 120 participants. Half serve as the experimental group and engage in a conversation with a human expert regarding the static explanations, while the other half are in the control group and read the materials regarding static explanations independently. We measure the participants’ objective and self-reported comprehension, acceptance, and trust of static explanations. Results show that conversations significantly improve participants’ comprehension, acceptance , trust, and collaboration with static explanations, while reading the explanations independently does not have these effects and even decreases users’ acceptance of explanations. Our findings highlight the importance of customized model explanations in the format of free-form conversations and provide insights for the future design of conversational explanations. Xi Jessie Yang, Boyang Li 0001 |
Int. J. Hum. Comput. Interact. | 3 |
| 2025 | Improving generalization of neural Vehicle Routing Problem solvers through the lens of model architecture
Yubin Xiao, Di Wang 0004, Xuan Wu 0004, Yuesong Wu, Boyang Li 0001, Wei Du 0002, Liupu Wang, You Zhou 0008 |
Neural Networks | 5 |
| 2025 | Reinforcement Learning-Based Nonautoregressive Solver for Traveling Salesman ProblemsabstractThe traveling salesman problem (TSP) is a well-known combinatorial optimization problem (COP) with broad real-world applications. Recently, neural networks (NNs) have gained popularity in this research area because as shown in the literature, they provide strong heuristic solutions to TSPs. Compared to autoregressive neural approaches, nonautoregressive (NAR) networks exploit the inference parallelism to elevate inference speed but suffer from comparatively low solution quality. In this article, we propose a novel NAR model named NAR4TSP, which incorporates a specially designed architecture and an enhanced reinforcement learning (RL) strategy. To the best of our knowledge, NAR4TSP is the first TSP solver that successfully combines RL and NAR networks. The key lies in the incorporation of NAR network output decoding into the training process. NAR4TSP efficiently represents TSP-encoded information as rewards and seamlessly integrates it into RL strategies, while maintaining consistent TSP sequence constraints during both training and testing phases. Experimental results on both synthetic and real-world TSPs demonstrate that NAR4TSP outperforms five state-of-the-art (SOTA) models in terms of solution quality, inference speed, and generalization to unseen scenarios. Yubin Xiao, Di Wang 0004, Boyang Li 0001, Huanhuan Chen 0001, Wei Pang 0001, Xuan Wu 0004, Dong Xu 0002, Yanchun Liang 0001, You Zhou 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Distilling Autoregressive Models to Obtain High-Performance Non-autoregressive Solvers for Vehicle Routing Problems with Faster Inference SpeedabstractNeural construction models have shown promising performance for Vehicle Routing Problems (VRPs) by adopting either the Autoregressive (AR) or Non-Autoregressive (NAR) learning approach. While AR models produce high-quality solutions, they generally have a high inference latency due to their sequential generation nature. Conversely, NAR models generate solutions in parallel with a low inference latency but generally exhibit inferior performance. In this paper, we propose a generic Guided Non-Autoregressive Knowledge Distillation (GNARKD) method to obtain high-performance NAR models having a low inference latency. GNARKD removes the constraint of sequential generation in AR models while preserving the learned pivotal components in the network architecture to obtain the corresponding NAR models through knowledge distillation. We evaluate GNARKD by applying it to three widely adopted AR models to obtain NAR VRP solvers for both synthesized and real-world instances. The experimental results demonstrate that GNARKD significantly reduces the inference time (4-5 times faster) with acceptable performance drop (2-3%). To the best of our knowledge, this study is first-of-its-kind to obtain NAR VRP solvers from AR ones through knowledge distillation. Yubin Xiao, Di Wang 0004, Boyang Li 0001, Mingzhao Wang, Xuan Wu 0004, Changliang Zhou, You Zhou 0008 |
AAAI | 3 |
| 2024 | Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsabstractFrom image-text pairs, large-scale vision-language models (VLMs) learn to implicitly associate image regions with words, which prove effective for tasks like visual question answering. However, leveraging the learned association for open-vocabulary semantic segmentation remains a challenge. In this paper, we propose a simple, yet extremely effective, training-free technique, Plug-and-Play Open- Vocabulary Semantic Segmentation (PnP-OVSS) for this task. PnP-OVSS leverages a VLM with direct text-to-image cross-attention and an image-text matching loss. To balance between over-segmentation and under-segmentation, we introduce Salience Dropout; by iteratively dropping patches that the model is most attentive to, we are able to better resolve the entire extent of the segmentation mask. PnP-OVSS does not require any neural net-work training and performs hyperparameter tuning without the need for any segmentation annotations, even for a validation set. PnP-OVSS demonstrates substantial improvements over comparable baselines (+29.4% mIoU on Pascal VOC, +13.2% mIoU on Pascal Context, +14.0% mIoU on MS COCO, +2.4% mIoU on COCO Stuff) and even outper-forms most baselines that conduct additional network training on top of pretrained VLMs. Our codebase is at https://github.com/letitiabanana/PnP-OVSS. Jiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang Li 0001 |
CVPR | 4 |
| 2024 | Concept-skill Transferability-based Data Selection for Large Vision-Language ModelsabstractInstruction tuning, or supervised finetuning on extensive task-specific data, is necessary for Large Vision-Language Models (LVLMs) to generalize well across a broad range of visionlanguage (VL) tasks.However, training on large VL datasets can become prohibitively expensive.In this work, we introduce COIN-CIDE, an effective and scalable data selection technique that uses a small model as a reference model to select visual instruction tuning data for efficient finetuning of a target LVLM, focusing on diversity and transferability.Specifically, we cluster the training data using internal activations from a small model, which identifies VL concept-skill compositions needed by a target LVLM.We then sample data from these diverse clusters by considering their density and transferability, or the ability to transfer well to other concept-skill compositions.This approach ensures the diversity of these compositions, which is vital for LVLM generalization.Extensive experiments demonstrate that COINCIDE achieves superior performance and data selection efficiency against 8 strong baselines on two distinct datasets: LLaVA-1.5 and Vision-Flan.Using only 20% of the LLaVA-1.5 dataset, COINCIDE achieves performance comparable to the LVLM finetuned on the whole dataset, with 70% reduction of the wall-clock running time.On the Vision-Flan dataset, our method achieves superior results with only 16.7% of the training data. Jaewoo Lee 0001, Boyang Li 0001, Sung Ju Hwang |
EMNLP | 2 |
| 2024 | Event Causality Is Key to Computational Story UnderstandingabstractYidan Sun, Qin Chao, Boyang Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Qin Chao, Boyang Li 0001 |
NAACL-HLT | 3 |
| 2024 | What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and BiasesabstractAnthony Tiong, Junqi Zhao, Boyang Li, Junnan Li, Steven Hoi, Caiming Xiong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Anthony Meng Huat Tiong, Junqi Zhao, Boyang Li 0001, Junnan Li 0001, Steven C. H. Hoi, Caiming Xiong |
NAACL-HLT | 3 |
| 2023 | Is GPT-3 a Good Data Annotator?abstractBosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia, Boyang Li, Shafiq Joty, Lidong Bing. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Bosheng Ding, Chengwei Qin, Yew Ken Chia, Boyang Li 0001, Shafiq R. Joty, Lidong Bing |
ACL (1) | 5 |
| 2023 | From Images to Textual Prompts: Zero-shot Visual Question Answering with Frozen Large Language ModelsabstractLarge language models (LLMs) have demonstrated excellent zero-shot generalization to new language tasks. However, effective utilization of LLMs for zero-shot visual question-answering (VQA) remains challenging, primarily due to the modality disconnect and task disconnect between the LLM and VQA tasks. End-to-end training on multimodal data may bridge the disconnects, but is inflexible and computationally expensive. To address this issue, we propose Img2LLM, a plug-and-play module that provides LLM prompts to enable LLMs to perform zeroshot VQA tasks without end-to-end training. We develop LLM-agnostic models describe image content as exemplar question-answer pairs, which prove to be effective LLM prompts. Img2LLM offers the following benefits: 1) It achieves comparable or better performance than methods relying on end-to-end training. For example, we outperform Flamingo [3] by 5.6% on VQAv2. On the challenging A-OKVQA dataset, our method outperforms few-shot methods by as much as 20%. 2) It flexibly interfaces with a wide range of LLMs to perform VQA. 3) It eliminates the need to specialize LLMs using end-to-end finetuning and serve highly specialized LLMs to end users, thereby reducing cost. Code is available via the LAVIS [28] framework at https://github.com/salesforce/LAVIS/tree/main/projects/img2llm-vqa. Jiaxian Guo, Junnan Li 0001, Dongxu Li 0003, Anthony Meng Huat Tiong, Boyang Li 0001, Dacheng Tao, Steven C. H. Hoi |
CVPR | 5 |
| 2023 | Mitigating and Evaluating Static Bias of Action Representations in the Background and the ForegroundabstractIn video action recognition, shortcut static features can interfere with the learning of motion features, resulting in poor out-of-distribution (OOD) generalization. The video background is clearly a source of static bias, but the video foreground, such as the clothing of the actor, can also provide static bias. In this paper, we empirically verify the existence of foreground static bias by creating test videos with conflicting signals from the static and moving portions of the video. To tackle this issue, we propose a simple yet effective technique, StillMix, to learn robust action representations. Specifically, StillMix identifies bias-inducing video frames using a 2D reference network and mixes them with videos for training, serving as effective bias suppression even when we cannot explicitly extract the source of bias within each video frame or enumerate types of bias. Finally, to precisely evaluate static bias, we synthesize two new benchmarks, SCUBA for static cues in the background, and SCUFO for static cues in the foreground. With extensive experiments, we demonstrate that StillMix mitigates both types of static bias and improves video representations for downstream applications. Code is available at https://github.com/lihaoxin05/StillMix. Haoxin Li, Yuan Liu 0002, Hanwang Zhang, Boyang Li 0001 |
ICCV | 4 |
| 2023 | Training Multimedia Event Extraction With Generated Images and CaptionsabstractContemporary news reporting increasingly features multimedia content, motivating research on multimedia event extraction. However, the task lacks annotated multimodal training data and artificially generated training data suffer from the distribution shift from the real-world data. In this paper, we propose Cross-modality Augmented Multimedia Event Learning (CAMEL), which successfully utilizes artificially generated multimodal training data and achieves state-of-the-art performance. Conditioned on unimodal training data, we generate multimodal training data using off-the-shelf image generators like Stable Diffusion [45] and image captioners like BLIP [24]. After that, we train the network on the resultant multimodal datasets. In order to learn robust features that are effective across domains, we devise an iterative and gradual training strategy. Substantial experiments show that CAMEL surpasses state-of-the-art (SOTA) baselines on the M2E2 benchmark. On multimedia events in particular, we outperform the prior SOTA by 4.2% F1 on event mention identification and by 9.8% F1 on argument identification, which demonstrates that CAMEL learns synergistic representations from the two modalities. Our work demonstrates a recipe to unleash the power of synthetic training data in structured prediction. Zilin Du, Yunxin Li, Xu Guo 0002, Boyang Li 0001 |
ACM Multimedia | 5 |
| 2023 | InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningabstractLarge-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-language models is challenging due to the rich input distributions and task diversity resulting from the additional visual input. Although vision-language pretraining has been widely studied, vision-language instruction tuning remains under-explored. In this paper, we conduct a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models. We gather 26 publicly available datasets, covering a wide variety of tasks and capabilities, and transform them into instruction tuning format. Additionally, we introduce an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction. Trained on 13 held-in datasets, InstructBLIP attains state-of-the-art zero-shot performance across all 13 held-out datasets, substantially outperforming BLIP-2 and larger Flamingo models. Our models also lead to state-of-the-art performance when finetuned on individual downstream tasks (e.g., 90.7% accuracy on ScienceQA questions with image contexts). Furthermore, we qualitatively demonstrate the advantages of InstructBLIP over concurrent multimodal models. All InstructBLIP models are open-source. Wenliang Dai, Junnan Li 0001, Dongxu Li 0003, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li 0001, Pascale Fung, Steven C. H. Hoi |
NeurIPS | 7 |
| 2023 | Improving Tail-Class Representation with Centroid Contrastive Learning
Anthony Meng Huat Tiong, Junnan Li 0001, Guosheng Lin, Boyang Li 0001, Caiming Xiong, Steven C. H. Hoi |
Pattern Recognit. Lett. | 4 |
| 2022 | VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningabstractThe limited availability of annotated data often hinders real-world applications of machine learning. To efficiently learn from small quantities of multimodal data, we leverage the linguistic knowledge from a large pre-trained language model (PLM) and quickly adapt it to new domains of image captioning. To effectively utilize a pretrained model, it is critical to balance the visual input and prior linguistic knowledge from pretraining. We propose VisualGPT, which employs a novel self-resurrecting encoder-decoder attention mechanism to quickly adapt the PLM with a small amount of in-domain image-text data. The proposed self-resurrecting activation unit produces sparse activations that prevent accidental overwriting of linguistic knowledge. When trained on 0.1%, 0.5% and 1% of the respective training sets, VisualGPT surpasses the best baseline by up to 10.0% CIDEr on MS COCO [43] and 17.9% CIDEr on Conceptual Captions [63]. Furthermore, VisualGPT achieves the state-of-the-art result on IU X-ray [15], a medical report generation dataset. Our code is available at https://github.com/Vision-CAIR/VisualGPT. Jun Chen 0021, Kai Yi, Boyang Li 0001, Mohamed Elhoseiny 0001 |
CVPR | 4 |
| 2022 | Semi-Supervised Federated Heterogeneous Transfer Learning
Siwei Feng, Boyang Li 0001, Han Yu 0001, Yang Liu 0165, Qiang Yang 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Federated Learning for Personalized Humor RecognitionabstractComputational understanding of humor is an important topic under creative language understanding and modeling. It can play a key role in complex human-AI interactions. The challenge here is that human perception of humorous content is highly subjective. The same joke may receive different funniness ratings from different readers. This makes it highly challenging for humor recognition models to achieve personalization in practical scenarios. Existing approaches are generally designed based on the assumption that users have a consensus on whether a given text is humorous or not. Thus, they cannot handle diverse humor preferences well. In this article, we propose the FedHumor approach for the recognition of humorous content in a personalized manner through Federated Learning (FL). Extending a pre-trained language model, FedHumor guides the fine-tuning process by considering diverse distributions of humor preferences from individuals. It incorporates a diversity adaptation strategy into the FL paradigm to train a personalized humor recognition model. To the best of our knowledge, FedHumor is the first text-based personalized humor recognition model through federated learning. Extensive experiments demonstrate the advantage of FedHumor in recognizing humorous texts compared to nine state-of-the-art humor recognition approaches with superior capability for handling the diversity in humor labels produced by users with diverse preferences. Xu Guo 0002, Han Yu 0001, Boyang Li 0001, Hao Wang 0005, Pengwei Xing, Siwei Feng, Zaiqing Nie, Chunyan Miao |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | HyDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural NetworksabstractThe behaviors of deep neural networks (DNNs) are notoriously resistant to human interpretations. In this paper, we propose Hypergradient Data Relevance Analysis, or HyDRA, which interprets the predictions made by DNNs as effects of their training data. Existing approaches generally estimate data contributions around the final model parameters and ignore how the training data shape the optimization trajectory. By unrolling the hypergradient of test loss w.r.t. the weights of training data, HyDRA assesses the contribution of training data toward test data points throughout the training trajectory. In order to accelerate computation, we remove the Hessian from the calculation and prove that, under moderate conditions, the approximation error is bounded. Corroborating this theoretical claim, empirical results indicate the error is indeed small. In addition, we quantitatively demonstrate that HyDRA outperforms influence functions in accurately estimating data contribution and detecting noisy data labels. The source code is available at https://github.com/cyyever/aaai_hydra. Yuanyuan Chen 0012, Boyang Li 0001, Han Yu 0001, Chunyan Miao |
AAAI | 2 |
| 2021 | Proof of Learning (PoLe): Empowering Machine Learning with Consensus Building on Blockchains (Demo)abstractThe consensus algorithm is the core component of a blockchain system, which determines the efficiency, security, and scalability of the blockchain network. The representative consensus algorithm is the proof of work (PoW) proposed in Bitcoin, where the consensus process consumes large amount of compute in solving meaningless Hash puzzel. Meanwhile, the deep learning (DL) has brought unprecedented performance gains at heavy computate cost. In this demo, we channels the otherwise wasted computational power to the practical purpose of training neural network models, through the proposed proof of learning (PoL) consensus algorithm. In PoLe, the training/testing data are released to the entire blockchain network (BCN) and the consensus nodes train NN models on the data, which serves as the proof of learning. When the consensus on the BCN considers a NN model to be valid, a new block is appended to the blockchain. Through our system, we investigate the potential of enpowering machine learning with consensus building on blockchains. Yixiao Lan, Yuan Liu 0002, Boyang Li 0001, Chunyan Miao |
AAAI | 3 |
| 2021 | Noise-Resistant Deep Metric Learning With Ranking-Based Instance SelectionabstractThe existence of noisy labels in real-world data negatively impacts the performance of deep learning models. Although much research effort has been devoted to improving robustness to noisy labels in classification tasks, the problem of noisy labels in deep metric learning (DML) remains open. In this paper, we propose a noise-resistant training technique for DML, which we name Probabilistic Ranking-based Instance Selection with Memory (PRISM). PRISM identifies noisy data in a minibatch using average similarity against image features extracted by several previous versions of the neural network. These features are stored in and retrieved from a memory bank. To alleviate the high computational cost brought by the memory bank, we introduce an acceleration method that replaces individual data points with the class centers. In extensive comparisons with 12 existing approaches under both synthetic and real-world label noise, PRISM demonstrates superior performance of up to 6.06% in Precision@1. Chang Liu 0040, Han Yu 0001, Boyang Li 0001, Zhiqi Shen 0001, Zhanning Gao, Peiran Ren, Xuansong Xie, Li-Zhen Cui 0001, Chunyan Miao |
CVPR | 3 |
| 2021 | Exploring Long Tail Visual Relationship Recognition with Large VocabularyabstractSeveral approaches have been proposed in recent literature to alleviate the long-tail problem, mainly in object classification tasks. In this paper, we make the first largescale study concerning the task of Long-Tail Visual Relationship Recognition (LTVRR). LTVRR aims at improving the learning of structured visual relationships that come from the long-tail (e.g., "rabbit grazing on grass"). In this setup, the subject, relation, and object classes each follow a long-tail distribution. To begin our study and make a future benchmark for the community, we introduce two LTVRR-related benchmarks, dubbed VG8K-LT and GQA-LT, built upon the widely used Visual Genome and GQA datasets. We use these benchmarks to study the performance of several state-of-the-art long-tail models on the LTVRR setup. Lastly, we propose a visiolinguistic hubless (VilHub) loss and a Mixup augmentation technique adapted to LTVRR setup, dubbed as RelMix. Both VilHub and RelMix can be easily integrated on top of existing models and despite being simple, our results show that they can remarkably improve the performance, especially on tail classes. Benchmarks, code, and models have been made available at: https://github.com/Vision-CAIR/LTVRR. Sherif Abdelkarim, Aniket Agarwal, Panos Achlioptas, Jun Chen 0021, Jiaji Huang, Boyang Li 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
ICCV | 6 |
| 2021 | Initialization Matters: Regularizing Manifold-informed Initialization for Neural Recommendation SystemsabstractProper initialization is crucial to the optimization and the generalization of neural networks. However, most existing neural recommendation systems initialize the user and item embeddings randomly. In this work, we propose a new initialization scheme for user and item embeddings called Laplacian Eigenmaps with Popularity-based Regularization for Isolated Data (LEPORID). LEPORID endows the embeddings with information regarding multi-scale neighborhood structures on the data manifold and performs adaptive regularization to compensate for high embedding variance on the tail of the data distribution. Exploiting matrix sparsity, LEPORID embeddings can be computed efficiently. We evaluate LEPORID in a wide range of neural recommendation models. In contrast to the recent surprising finding that the simple K-nearest-neighbor (KNN) method often outperforms neural recommendation systems, we show that existing neural systems initialized with LEPORID often perform on par or better than KNN. To maximize the effects of the initialization, we propose the Dual-Loss Residual Recommendation (DLR^2) network, which, when initialized with LEPORID, substantially outperforms both traditional and state-of-the-art neural recommender systems. Yinan Zhang 0002, Boyang Li 0001, Yong Liu 0020, Hao Wang 0005, Chunyan Miao |
KDD | 2 |
| 2021 | Latent-Optimized Adversarial Neural Transfer for Sarcasm DetectionabstractThe existence of multiple datasets for sarcasm detection prompts us to apply transfer learning to exploit their commonality.The adversarial neural transfer (ANT) framework utilizes multiple loss terms that encourage the source-domain and the target-domain feature distributions to be similar while optimizing for domain-specific performance.However, these objectives may be in conflict, which can lead to optimization difficulties and sometimes diminished transfer.We propose a generalized latent optimization strategy that allows different losses to accommodate each other and improves training dynamics.The proposed method outperforms transfer learning and meta-learning baselines.In particular, we achieve 10.02% absolute performance gain over the previous state of the art on the iSarcasm dataset. Xu Guo 0002, Boyang Li 0001, Han Yu 0001, Chunyan Miao |
NAACL-HLT | 2 |
| 2021 | Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature DistributionsabstractThe task of video and text sequence alignment is a pre-requisite step toward joint understanding of movie videos and screenplays. However, supervised methods face the obstacle of limited realistic training data. With this pa-per, we attempt to enhance data efficiency of the end-to-end alignment network NeuMATCH [15]. Recent research [56] suggests that network components dealing with different modalities may overfit and generalize at different speeds, creating difficulties for training. We propose to employ (1) layer-wise adaptive rate scaling (LARS) to align the magnitudes of gradient updates in different layers and balance the pace of learning and (2) sequence-wise batch normalization (SBN) to align the internal feature distributions from different modalities. Finally, we leverage random projection to reduce the dimensionality of input features. On the YouTube Movie Summary dataset, the combined use of these technique closes the performance gap when the pretraining on the LSMDC dataset is omitted and achieves the state-of-the-art result. Extensive empirical comparisons and analysis reveal that these techniques improve optimization and regularize the network more effectively than two different setups of layer normalization. Boyang Li 0001, Yanwei Fu 0001 |
WACV | 2 |
| 2021 | Proof of Learning (PoLe): Empowering neural network training with consensus building on blockchainsabstractThe advent of neural network (NN) based deep learning, especially the recent development of the automatic design of networks, has brought unprecedented performance gains at heavy computational cost. On the other hand, in order to generate a new consensus block, Proof of Work (PoW) based blockchain systems routinely perform a huge amount of computation that does not achieve practical purposes but to solving a difficult cryptographic hash puzzle problem.In this study, we propose a new consensus mechanism, Proof of Learning (PoLe), which directs the computation spent for block consensus toward optimization of neural networks. In our design, the training and testing data are released to the entire blockchain network and the consensus nodes train NN models on the data, which serves as the proof of learning. As a core component of PoLe, we design a secure mapping layer (SML) to prevent consensus nodes from cheating, which can be straightforwardly implemented as a linear NN layer. When the consensus on the blockchain network is achieved, a new block is appended to the blockchain. We experimentally compare the PoLe protocol with Proof of Work (PoW) and show that PoLe can achieve a more stable block generation rate, which leads to more efficient transaction processing. Experimental evaluation also shows the PoLe can achieve a stable block generation rate without significantly sacrificing training performance. Yuan Liu 0002, Yixiao Lan, Boyang Li 0001, Chunyan Miao, Zhihong Tian 0001 |
Comput. Networks | 3 |
| 2020 | Predicting Personality from Book Preferences with User-Generated Content LabelsabstractPsychological studies have shown that personality traits are associated with book preferences. However, past findings based on questionnaires are limited to conventional book genres and do not capture niche content (e.g., family drama) and reading behaviors (e.g., backburners). For a more comprehensive measure of book content, this study harnesses a massive archive of content labels, also known as `tags', created by users of a book review website, Goodreads.com. Combined with data on preferences and personality scores collected from Facebook users, the tag labels achieve high accuracy in personality prediction by psychological standards. Additionally, we group tags into broader genres to check their validity against past findings. Our results are robust across both tag-level and genre-level analyses and are consistent with existing literature. Moreover, user-generated tag labels reveal unexpected insights, such as cultural differences, book reading behaviors, and other non-content factors affecting preferences. To our knowledge, this is currently the largest study that explores the relationship between personality and book content preferences. Ng Annalyn, Maarten W. Bos, Leonid Sigal, Boyang Li 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2020 | A Multi-Task Neural Approach for Emotion Attribution, Classification, and SummarizationabstractEmotional content is a crucial ingredient in user-generated videos. However, the sparsity of emotional expressions in the videos poses an obstacle to visual emotion analysis. In this paper, we propose a new neural approach, Bi-stream Emotion Attribution-Classification Network (BEAC-Net), to solve three related emotion analysis tasks: emotion recognition, emotion attribution, and emotion-oriented summarization, in a single integrated framework. BEAC-Net has two major constituents, an attribution network and a classification network. The attribution network extracts the main emotional segment that classification should focus on in order to mitigate the sparsity issue. The classification network utilizes both the extracted segment and the original video in a bi-stream architecture. We contribute a new dataset for the emotion attribution task with human-annotated ground-truth labels for emotion segments. Experiments on two video datasets demonstrate superior performance of the proposed framework and the complementary nature of the dual classification streams. Guoyun Tu, Yanwei Fu 0001, Boyang Li 0001, Jiarui Gao, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Understanding Actors and Evaluating Personae with Gaussian Embeddings
Hannah Kim 0001, Denys Katerenchuk, Daniel Billet, Jun Huan, Haesun Park, Boyang Li 0001 |
AAAI | 6 |
| 2019 | Joint Event Detection and Description in Continuous Video StreamsabstractDense video captioning is a fine-grained video understanding task that involves two sub-problems: localizing distinct events in a long video stream, and generating captions for the localized events. We propose the Joint Event Detection and Description Network (JEDDi-Net), which solves the dense video captioning task in an end-to-end fashion. Our model continuously encodes the input video stream with three-dimensional convolutional layers, proposes variable-length temporal events based on pooled features, and generates their captions. Proposal features are extracted within each proposal segment through 3D Segment-of-Interest pooling from shared video feature encoding. In order to explicitly model temporal relationships between visual events and their captions in a single video, we also propose a two-level hierarchical captioning module that keeps track of context. On the large-scale ActivityNet Captions dataset, JEDDi-Net demonstrates improved results as measured by standard metrics. We also present the first dense captioning results on the TACoS-MultiLevel dataset. Huijuan Xu 0001, Boyang Li 0001, Vasili Ramanishka, Leonid Sigal, Kate Saenko |
WACV | 2 |
| 2018 | A Neural Multi-Sequence Alignment TeCHnique (NeuMATCH)abstractThe alignment of heterogeneous sequential data (video to text) is an important and challenging problem. Standard techniques for this task, including Dynamic Time Warping (DTW) and Conditional Random Fields (CRFs), suffer from inherent drawbacks. Mainly, the Markov assumption implies that, given the immediate past, future alignment decisions are independent of further history. The separation between similarity computation and alignment decision also prevents end-to-end training. In this paper, we propose an end-to-end neural architecture where alignment actions are implemented as moving data between stacks of Long Short-term Memory (LSTM) blocks. This flexible architecture supports a large variety of alignment tasks, including one-to-one, one-to-many, skipping unmatched elements, and (with extensions) non-monotonic alignment. Extensive experiments on semi-synthetic and real datasets show that our algorithm outperforms state-of-the-art baselines. Pelin Dogan-Schönberger, Boyang Li 0001, Leonid Sigal, Markus Gross 0001 |
CVPR | 2 |
| 2018 | Annotating High-Level Structures of Short Stories and Personal Anecdotes
Boyang Li 0001, Beth Cardier, Tong Wang 0007, Florian Metze |
LREC | 1 |
| 2018 | Heterogeneous Knowledge Transfer in Video Emotion Recognition, Attribution and SummarizationabstractEmotion is a key element in user-generated video. However, it is difficult to understand emotions conveyed in such videos due to the complex and unstructured nature of user-generated content and the sparsity of video frames expressing emotion. In this paper, for the first time, we propose a technique for transferring knowledge from heterogeneous external sources, including image and textual data, to facilitate three related tasks in understanding video emotion: emotion recognition, emotion attribution and emotion-oriented summarization. Specifically, our framework (1) learns a video encoding from an auxiliary emotional image dataset in order to improve supervised video emotion recognition, and (2) transfers knowledge from an auxiliary textual corpora for zero-shot recognition of emotion classes unseen during training. The proposed technique for knowledge transfer facilitates novel applications of emotion attribution and emotion-oriented summarization. A comprehensive set of experiments on multiple datasets demonstrate the effectiveness of our framework. Baohan Xu, Yanwei Fu 0001, Yu-Gang Jiang 0001, Boyang Li 0001, Leonid Sigal |
IEEE Trans. Affect. Comput. | 4 |
| 2017 | Collaborative Storytelling between Robot and Child: A Feasibility StudyabstractJoint storytelling is a common parent-child activity and brings multiple benefits such as improved language learning for children. Most existing storytelling robots offer rigid interaction with children and do not contribute to children's stories. In this paper, we envision a robot that collaborates with a child to create oral stories in a highly interactive manner. We performed a Wizard-of-Oz feasibility study, which involved 78 children between 4 and 10 years old, to compare two collaboration strategies: (1) inserting new story content and relating it to the existing story and (2) inserting content without relating it to the existing story. We hypothesize the first strategy can foster true collaboration and create rapport, whereas the second is a safe strategy when the robot cannot understand the story. We observed that, although the first strategy creates a heavier cognitive load, it was as enjoyable as the second. We also observed some indications that the first strategy may mitigate the difficulties in story creation for young children under the age of 7 and encourage children to speak more. This study suggests that a mixture strategy is feasible for robots in collaborative storytelling, providing sufficient cognitive challenge while concealing its shortcomings on natural language understanding. Iolanda Leite, Jill Fain Lehman, Boyang Li 0001 |
IDC | 4 |
| 2017 | Game Engine Learning from VideoabstractIntelligent agents need to be able to make predictions about their environment. In this work we present a novel approach to learn a forward simulation model via simple search over pixel input. We make use of a video game, Super Mario Bros., as an initial test of our approach as it represents a physics system that is significantly less complex than reality. We demonstrate the significant improvement of our approach in predicting future states compared with a baseline CNN and apply the learned model to train a game playing agent. Thus we evaluate the algorithm in terms of the accuracy and value of its output model. Matthew Guzdial, Boyang Li 0001, Mark O. Riedl |
IJCAI | 2 |
| 2017 | Predicting the Quality of Short Narratives from Social MediaabstractAn important and difficult challenge in building computational models for narratives is the automatic evaluation of narrative quality. Quality evaluation connects narrative understanding and generation as generation systems need to evaluate their own products. To circumvent difficulties in acquiring annotations, we employ upvotes in social media as an approximate measure for story quality. We collected 54,484 answers from a crowd-powered question-and-answer website, Quora, and then used active learning to build a classifier that labeled 28,320 answers as stories. To predict the number of upvotes without the use of social network features, we create neural networks that model textual regions and the interdependence among regions, which serve as strong benchmarks for future research. To our best knowledge, this is the first large-scale study for automatic evaluation of narrative quality. Tong Wang 0007, Ping Chen 0001, Boyang Li 0001 |
IJCAI | 3 |
| 2017 | Learning and Reusing Dialog for Repeated Interactions with a Situated Social Agent
James Kennedy 0001, Iolanda Leite, André Pereira 0001, Boyang Li 0001, Rishub Jain, Ricson Cheng, Eli Pincus, Elizabeth J. Carter, Jill Fain Lehman |
IVA | 5 |
| 2016 | Semi-situated learning of verbal and nonverbal content for repeated human-robot interactionabstractContent authoring of verbal and nonverbal behavior is a limiting factor when developing agents for repeated social interactions with the same user. We present PIP, an agent that crowdsources its own multimodal language behavior using a method we call semi-situated learning. PIP renders segments of its goal graph into brief stories that describe future situations, sends the stories to crowd workers who author and edit a single line of character dialog and its manner of expression, integrates the results into its goal state representation, and then uses the authored lines at similar moments in conversation. We present an initial case study in which the language needed to host a trivia game interaction is learned pre-deployment and tested in an autonomous system with 200 users "in the wild." The interaction data suggests that the method generates both meaningful content and variety of expression. Iolanda Leite, André Pereira 0001, Allison Funkhouser, Boyang Li 0001, Jill Fain Lehman |
ICMI | 4 |
| 2016 | Video Emotion Recognition with Transferred Deep Feature EncodingsabstractDespite growing research interest, emotion understanding for user-generated videos remains a challenging problem. Major obstacles include the diversity and complexity of video content, as well as the sparsity of expressed emotions. For the first time, we systematically study large-scale video emotion recognition by transferring deep feature encodings. In addition to the traditional, supervised recognition, we study the problem of zero-shot emotion recognition, where emotions in the test set are unseen during training. To cope with this task, we utilize knowledge transferred from auxiliary image and text corpora. A novel auxiliary Image Transfer Encoding (ITE) process is proposed to efficiently encode and generate video representation. We also thoroughly investigate different configurations of convolutional neural networks. Comprehensive experiments on multiple datasets demonstrate the effectiveness of our framework. Baohan Xu, Yanwei Fu 0001, Yu-Gang Jiang 0001, Boyang Li 0001, Leonid Sigal |
ICMR | 4 |
| 2015 | Scheherazade: Crowd-Powered Interactive Narrative GenerationabstractInteractive narrative is a form of storytelling in which users affect a dramatic storyline through actions by assuming the role of characters in a virtual world.This extended abstract outlines the Scheherazade-IF system, which uses crowdsourcing and artificial intelligence to automatically construct text-based interactive narrative experiences. Boyang Li 0001, Mark O. Riedl |
AAAI | 1 |
| 2015 | Crowdsourcing Open Interactive Narrative
Matthew Guzdial, Brent E. Harrison, Boyang Li 0001, Mark O. Riedl |
FDG | 3 |
| 2014 | Storytelling with Adjustable Narrator Styles and Sentiments
Boyang Li 0001, Mohini Thakkar, Mark O. Riedl |
ICIDS | 1 |
| 2014 | From Data to Storytelling Agents
Boyang Li 0001, Mohini Thakkar, Mark O. Riedl |
IVA | 1 |
| 2013 | Story Generation with Crowdsourced Plot GraphsabstractStory generation is the problem of automatically selecting a sequence of events that meet a set of criteria and can be told as a story. Story generation is knowledge-intensive; traditional story generators rely on a priori defined domain models about fictional worlds, including characters, places, and actions that can be performed. Manually authoring the domain models is costly and thus not scalable. We present a novel class of story generation system that can generate stories in an unknown domain. Our system (a) automatically learns a domain model by crowdsourcing a corpus of narrative examples and (b) generates stories by sampling from the space defined by the domain model. A large-scale evaluation shows that stories generated by our system for a previously unknown topic are comparable in quality to simple stories authored by untrained humans Boyang Li 0001, Stephen Lee-Urban, George Johnston, Mark O. Riedl |
AAAI | 1 |
| 2013 | Crowdsourcing interactive fiction games
Boyang Li 0001, Stephen Lee-Urban, Mark O. Riedl |
FDG | 1 |
| 2012 | Goal-Driven Conceptual Blending: A Computational Approach for Creativity
Boyang Li 0001, Alexander Zook, Nicholas Davis 0001, Mark O. Riedl |
ICCC | 1 |
| 2011 | Distributed creative cognition in digital filmmakingabstractThis paper reports on an empirical study that uses a Grounded Theory approach to investigate the creative practices of Machinima filmmakers. Machinima is a new digital film production technique that uses the 3D graphics and real time rendering capability of video game engines to create films. In contrast to practices used in traditional film production, we've found that Machinima filmmakers explore and evaluate ideas in real time. These filmmakers generate vague and underspecified mental images, which are then explored and refined using the real time rendering capabilities of game engines. The game engine assists the filmmaker to fill in indeterminate details, which allows creative exploration of scenes through playfully experimenting with parameters such as camera angle and position, lighting, and character position. Creative exploration distributes the cognitive task of evaluation between the human user and the Machinima tool to enable evaluation through exploring possible scene configurations. Nicholas Davis 0001, Boyang Li 0001, Mark O. Riedl, Michael Nitsche |
Creativity & Cognition | 2 |
| 2011 | Creative gadget design in fictions: generalized planning in analogical spacesabstractScience-fiction and fantasy stories often contain objects never envisioned previously. Inventing gadgets like lightsabers or mythical creatures like griffins is a creative task. Traditional computational storytelling systems are limited in their expressivity because they cannot create new types of objects or gadgets. The Japanese manga series Doraemon exemplifies the role of new and creative gadgets in creating fun and successful stories. We surveyed five volumes of Doraemon and identified 9 cognitive strategies of gadget creation, unified in a 5-step process. We present an algorithm to create new types of gadgets in the context of story generation. The algorithm is a combination of partial-order planning and analogical reasoning. Although Doraemon is our motivating example, we can also generate gadgets commonly seen in other science fictions and fairy tales. Boyang Li 0001, Mark O. Riedl |
Creativity & Cognition | 1 |
| 2008 | Memetic Gradient SearchabstractThis paper reviews the different gradient-based schemes and the sources of gradient, their availability, precision and computational complexity, and explores the benefits of using gradient information within a memetic framework in the context of continuous parameter optimization, which is labeled here as memetic gradient search. In particular, we considered a quasi-Newton method with analytical gradient and finite differencing, as well as simultaneous perturbation stochastic approximation, used as the local searches. Empirical study on the impact of using gradient information showed that memetic gradient search outperformed the traditional GA and analytical, precise gradient brings considerable benefit to gradient-based local search (LS) schemes. Though gradient-based searches can sometimes get trapped in local optima, memetic gradient searches were still able to converge faster than the conventional GA. Boyang Li 0001, Yew-Soon Ong, Minh Nghia Le, Chi Keong Goh |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | The national weather sensor gridabstractWith the rapid advances in technologies such as MEMS sensors, low-power embedded processing and wireless networking, sensor networks are becoming more powerful in terms of data acquisition and processing capabilities. Sensor networks can now be deployed in the physical world for various important applications such as environmental monitoring, weather monitoring and modeling, military surveillance, healthcare monitoring, tracking of goods and manufacturing processes, smart homes and offices, etc. Hock-Beng Lim, Keck Voon Ling, Yuxia Yao, Mudasser Iqbal, Boyang Li 0001, Xiaonan Yin |
SenSys | 6 |