Feiqi Cao

dblp:318/3086 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-4910-5925ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 3M-Game: Multi-Modal Multi-Task Multi-Teacher Learning for Game Event Detection (Student Abstract)
abstract
Esports has rapidly emerged as a global phenomenon with an ever-expanding audience on livestream platforms. However, due to the complex nature of the game, it becomes challenging for newcomers to comprehend the gaming situation. This research introduces a 3M-Game that integrates multi-modal (MM) information from the livestream platform, including chat and livestream, to uncover the event. While conventional MM models typically prioritise aligning MM data through concurrent training towards a unified objective, our framework leverages multiple independent teachers trained on different tasks to accomplish game event detection. The results show the effectiveness of the proposed framework. The code and appendix are in https://github.com/adlnlp/3m_game.
Thye Shan Ng, Feiqi Cao, Soyeon Caren Han
AAAI2
2024 The Language Model Can Have the Personality: Joint Learning for Personality Enhanced Language Model (Student Abstract)
abstract
With the introduction of large language models, chatbots are becoming more conversational to communicate effectively and capable of handling increasingly complex tasks. To make a chatbot more relatable and engaging, we propose a new language model idea that maps the human-like personality. In this paper, we propose a systematic Personality-Enhanced Language Model (PELM) approach by using a joint learning mechanism of personality classification and language generation tasks. The proposed PELM leverages a dataset of defined personality typology, Myers-Briggs Type Indicator, and produces a Personality-Enhanced Language Model by using a joint learning and cross-teaching structure consisting of a classification and language modelling to incorporate personalities via both distinctive types and textual information. The results show that PELM can generate better personality-based outputs than baseline models.
Feiqi Cao, Yihao Ding, Soyeon Caren Han
AAAI2
2024 PEACH: Pretrained-Embedding Explanation across Contextual and Hierarchical Structure
Feiqi Cao, Soyeon Caren Han, Hyunsuk Chung
IJCAI1
2024 Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
abstract
This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the foundational concepts of multimodality, the evolution of multimodal research, and the key technical challenges addressed by these models. We will cover the latest multimodal datasets and pretrained models, including those beyond vision and language. Additionally, the tutorial will delve into the intricacies of multimodal large models and instruction tuning strategies to optimise performance for specific tasks. Hands-on laboratories will offer practical experience with state-of-the-art multimodal models, demonstrating real-world applications like visual storytelling and visual question answering. This tutorial aims to equip researchers, practitioners, and newcomers with the knowledge and skills to leverage multimodal AI. ACM Multimedia 2024 is the ideal venue for this tutorial, aligning perfectly with our goal of understanding multimodal pretrained and large language models, and their tuning mechanisms.
Soyeon Caren Han, Feiqi Cao, Josiah Poon, Roberto Navigli
ACM Multimedia2
2023 In-Game Toxic Language Detection: Shared Task and Attention Residuals (Student Abstract)
abstract
In-game toxic language becomes the hot potato in the gaming industry and community. There have been several online game toxicity analysis frameworks and models proposed. However, it is still challenging to detect toxicity due to the nature of in-game chat, which has extremely short length. In this paper, we describe how the in-game toxic language shared task has been established using the real-world in-game chat data. In addition, we propose and introduce the model/framework for toxic language token tagging (slot filling) from the in-game chat. The data and code will be released.
Yuanzhe Jia, Weixuan Wu, Feiqi Cao, Soyeon Caren Han
AAAI3
2022 Understanding Attention for Vision-and-Language Tasks
abstract
Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While attention has been widely used in VL tasks, it has not been examined the capability of different attention alignment calculation in bridging the semantic gap between visual and textual clues. In this research, we conduct a comprehensive analysis on understanding the role of attention alignment by looking into the attention score calculation methods and check how it actually represents the visual region’s and textual token’s significance for the global assessment. We also analyse the conditions which attention score calculation mechanism would be more (or less) interpretable, and which may impact the model performance on three different VL tasks, including visual question answering, text-to-image generation, text-and-image matching (both sentence and image retrieval). Our analysis is the first of its kind and provides useful insights of the importance of each attention alignment score calculation when applied at the training phase of VL tasks, commonly ignored in attention-based cross modal models, and/or pretrained models. Our code is available at: https://github.com/adlnlp/Attention_VL
Feiqi Cao, Soyeon Caren Han, Siqu Long, Changwei Xu, Josiah Poon
COLING1
2022 Vision-and-Language Pretrained Models: A Survey
abstract
Pretrained models have produced great success in both Computer Vision (CV) and Natural Language Processing (NLP). This progress leads to learning joint representations of vision and language pretraining by feeding visual and linguistic contents into a multi-layer transformer, Visual-Language Pretrained Models (VLPMs). In this paper, we present an overview of the major advances achieved in VLPMs for producing joint representations of vision and language. As the preliminaries, we briefly describe the general task definition and genetic architecture of VLPMs. We first discuss the language and vision data encoding methods and then present the mainstream VLPM structure as the core content. We further summarise several essential pretraining and fine-tuning strategies. Finally, we highlight three future directions for both CV and NLP researchers to provide insightful guidance.
Siqu Long, Feiqi Cao, Soyeon Caren Han, Haiqin Yang
IJCAI2