Bolin Lai

dblp:250/6136 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-7578-7336ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
abstract
Spurious bias, a tendency to exploit spurious correlations between superficial input attributes and prediction targets, has revealed a severe robustness pitfall in classical machine learning problems. Multimodal Large Language Models (MLLMs), which leverage pretrained vision and language models, have recently demonstrated strong capability in joint vision-language understanding. However, both the presence and severity of spurious biases in MLLMs remain poorly understood. In this work, we address this gap by analyzing the spurious biases in the multimodal setting and uncovering the specific inference-time data patterns that can manifest this problem. To support this analysis, we introduce MM-SpuBench, a comprehensive, human-verified benchmark dataset consisting of image-class pairs annotated with core and spurious attributes, grounded in our taxonomy of nine distinct types of spurious correlations. The benchmark is constructed using human-interpretable attribute information to capture a wide range of spurious patterns reflective of real-world knowledge. Leveraging this benchmark, we conduct a comprehensive evaluation of the state-of-the-art open-source and proprietary MLLMs with both standard accuracy and the proposed Conditional Generation Likelihood Advantage (CGLA). Our findings highlight the persistence of reliance on spurious correlations and the difficulty of mitigation on our benchmark. We hope this work can inspire new technical strides to mitigate these biases. Our benchmark is publicly available at https://huggingface.co/datasets/mmbench/MM-SpuBench.
Wenqian Ye, Bohan Liu 0008, Guangtao Zheng, Di Wang 0053, Yunsheng Ma, Bolin Lai, James M. Rehg, Aidong Zhang 0001
KDD (1)7
2025 SocialGesture: Delving into Multi-person Gesture Understanding
abstract
Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalities like language and speech. To address this issue, we introduce SocialGesture, the first large-scale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social interactions. Furthermore, we propose a novel visual question answering (VQA) task to benchmark vision language models’ (VLMs) performance on social gesture understanding. Our findings highlight several limitations of current gesture recognition models, offering insights into future directions for improvement in this field. SocialGesture is available at hugging-face.co/datasets/IrohXu/SocialGesture.
Pranav Virupaksha, Wenqi Jia 0001, Bolin Lai, Fiona Ryan, Sangmin Lee 0001, James M. Rehg
CVPR4
2025 Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs
abstract
Long-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited context windows. In this work, we introduce VideoMindPalace, a new framework inspired by the "Mind Palace", which organizes critical video moments into a topologically structured semantic graph. VideoMindPalace organizes key information through (i) hand-object tracking and interaction, (ii) clustered activity zones representing specific areas of recurring activities, and (iii) environment layout mapping, allowing natural language parsing by LLMs to provide grounded insights on spatio-temporal and 3D context. In addition, we propose the Video Mind-Palace Benchmark (VMB), to assess human-like reasoning, including spatial localization, temporal reasoning, and layout-aware sequential understanding. Evaluated on VMB and established video QA datasets, including EgoSchema, NExT-QA, IntentQA, and the Active Memories Benchmark, VideoMindPalace demonstrates notable gains in spatiotemporal coherence and human-aligned reasoning, advancing long-form video analysis capabilities in VLMs. Project website: https://oodbag.github.io/VMP/.
Zeyi Huang, Yuyang Ji, Nikhil Mehta 0002, Tong Xiao 0003, Donghyun Lee 0004, Sigmund Vanvalkenburgh, Shengxin Zha, Bolin Lai, Licheng Yu, Yong Jae Lee, Miao Liu 0007
CVPR9
2025 Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
abstract
Text-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been applied for instructions that are underrepresented in the training set, or difficult to describe purely in language. However, learning from visual prompts requires strong reasoning capability, which diffusion models are struggling with. To address this issue, we introduce a novel multi-modal autoregressive model, dubbed InstaManip, that can instantly learn a new image manipulation operation from textual and visual guidance via in-context learning, and apply it to new query images. Specifically, we propose an innovative group self-attention mechanism to break down the in-context learning process into two separate stages – learning and applying, which simplifies the complex problem into two easier tasks. We also introduce a relation regularization method to further disentangle image transformation features from irrelevant contents in exemplar images. Extensive experiments suggest that our method surpasses previous few-shot image manipulation models by a notable margin (≥19% in human evaluation). We also find our model can be further boosted by increasing the number or diversity of exemplar images. Please check out our project page (https://bolinlai.github.io/projects/InstaManip/).
Bolin Lai, Felix Juefei-Xu, Miao Liu 0007, Xiaoliang Dai, Nikhil Mehta 0002, Zeyi Huang, James M. Rehg, Sangmin Lee 0001, Tong Xiao 0003
CVPR1
2025 Toward Human Deictic Gesture Target Estimation
abstract
Humans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central role in human communication. Understanding the intended targets of another individual's deictic gestures enables inference of their intentions, comprehension of their current actions, and prediction of upcoming behaviors. Despite its significance, gesture target estimation remains an underexplored task within the computer vision community. In this paper, we introduce GestureTarget, a novel task designed specifically for comprehensive evaluation of social deictic gesture semantic target estimation. To address this task, we propose TransGesture, a set of Transformer-based gesture target prediction models. Given an input image and the spatial location of a person, our models predict the intended target of their gesture within the scene. Critically, our gaze-aware joint cross attention fusion model demonstrates how incorporating gaze-following cues significantly improves gesture target mask prediction IoU by 6% and gesture existence prediction accuracy by 10%. Our results underscore the complexity and importance of integrating gaze cues into deictic gesture intention understanding, advocating for increased research attention to this emerging area. All data, code will be made publicly available upon acceptance. Code of TransGesture is available at GitHub.com/IrohXu/TransGesture.
Pranav Virupaksha, Sangmin Lee 0001, Bolin Lai, Wenqi Jia 0001, Jintai Chen, James M. Rehg
NeurIPS4
2024 Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
abstract
Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behav-iors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Conse-quently, they are limited in modeling the intricate dynam-ics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dy-namics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to cu-rate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Exper-iments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.iolprojectslMMSI.
Sangmin Lee 0001, Bolin Lai, Fiona Ryan, Bikram Boote, James M. Rehg
CVPR2
2024 LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
Bolin Lai, Xiaoliang Dai, Lawrence Chen 0002, Guan Pang, James M. Rehg, Miao Liu 0007
ECCV (9)1
2024 Listen to Look Into the Future: Audio-Visual Egocentric Gaze Anticipation
Bolin Lai, Fiona Ryan, Wenqi Jia 0001, Miao Liu 0007, James M. Rehg
ECCV (9)1
2024 In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation and Beyond
abstract
Predicting human's gaze from egocentric videos serves as a critical role for human intention understanding in daily activities. In this paper, we present the first transformer-based model to address the challenging problem of egocentric gaze estimation. We observe that the connection between the global scene context and local visual information is vital for localizing the gaze fixation from egocentric video frames. To this end, we design the transformer encoder to embed the global context as one additional visual token and further propose a novel global-local correlation module to explicitly model the correlation of the global token and each local token. We validate our model on two egocentric video datasets - EGTEA Gaze + and Ego4D. Our detailed ablation studies demonstrate the benefits of our method. In addition, our approach exceeds the previous state-of-the-art model by a large margin. We also apply our model to a novel gaze saccade/fixation prediction task and the traditional action recognition problem. The consistent gains suggest the strong generalization capability of our model. We also provide additional visualizations to support our claim that global-local correlation serves a key representation for predicting gaze fixation from egocentric videos. More details can be found in our website (https://bolinlai.github.io/GLC-EgoGazeEst).
Bolin Lai, Miao Liu 0007, Fiona Ryan, James M. Rehg
Int. J. Comput. Vis.1
2022 In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
Bolin Lai, Miao Liu 0007, Fiona Ryan, James M. Rehg
BMVC1
2021 Semi-supervised Vein Segmentation of Ultrasound Images for Autonomous Venipuncture
abstract
Venipuncture is an indispensable procedure for both diagnosis and treatment. In this paper, unlike existing solutions that fully or partially rely on professional assistance, a compact robotic system integrating both novel hardware and software developments is introduced. The hardware consists of a set of units to facilitate the supporting, positioning, puncturing, and imaging functionalities. To achieve full automation, a novel deep learning framework — semi-ResNeXt-Unet for semi-supervised vein segmentation from ultrasound images is proposed. The depth information of vein is calculated and enables the automated navigation for the puncturing unit. The algorithm is validated on 40 volunteers, and the proposed semi-ResNeXt-Unet improves the dice similarity coefficient (DSC) by 5.36%, decreases the centroid error by 1.38 pixels and decreases the failure rate by 5.60%, compared to fully-supervised ResNeXt-Unet.
Bolin Lai, Nanyang Ye 0001, Zhongyuan Ren, Xiaoyun Zhou 0001, Peng Qi 0001
IROS3