Binzhu Xie

dblp:327/5879 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights
abstract
Traffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints to anomaly scenarios such as crashes and honking. Our contributions are twofold. First, we compile AV-TAU, the first large-scale audio-visual dataset for TAU, providing 29,865 traffic anomaly videos and 149,325 Q&A pairs, while supporting five essential TAU tasks. Second, we develop EchoTraffic, a multimodal LLM that integrates audio and visual data for TAU, through our audio-insight frame selector and dynamic connector to effectively extract crucial audio cues for anomaly understanding with a two-phase training framework. Experimental results on AV-TAU manifest that EchoTraffic sets a new SOTA performance in TAU, outperforming the existing multimodal LLMs. Our contributions, including AV-TAU and EchoTraffic, pave a new direction for multimodal TAU.
Zhenghao Xing, Hao Chen 0193, Binzhu Xie, Xuemiao Xu, Jianye Hao, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng
CVPR3
2025 EgoLife: Towards Egocentric Life Assistant
abstract
We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities—including discussions, shopping, cooking, social-izing, and entertainment—using AI glasses for multimodal person-view video references. This effort resulted in EgoLife Dataset, a comprehensive 300-hour egocentric, terpersonal, multiview, and multimodal daily life with intensive annotation. Leveraging this dataset, we troduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful sistance in daily life by addressing practical questions as recalling past relevant events, monitoring health and offering personalized recommendations.To address the key technical challenges of 1) developing robust visual-audio models for egocentric data, 2) enabling identity recognition, and 3) facilitating long-context question answering over extensive temporal information, we introduce EgoBulter, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants.
Shuai Liu 0002, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li 0080, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Jörg Widmer, Francesco Gringoli, Lei Yang 0059, Ziwei Liu 0002
CVPR9
2025 Trade-Offs in Image Generation: How Do Different Dimensions Interact?
Binzhu Xie, Zhonghao Yan, Shi Qiu 0001, Guoyang Xie, Zhichao Lu
ICCV2
2025 Generative Multi-Sensory Meditation: Exploring Immersive Depth and Activation in Virtual Reality
abstract
This work introduces MindfulVerse, an AI-Generated Content (AIGC)-driven application for personalized mindfulness experiences in VR. Using fNIRS to observe brain activation, the study demonstrates that generative meditation improves neural activation in self-regulation regions and positively impacts emotional regulation and user participation compared to static content.
Yuyang Jiang 0002, Binzhu Xie, Xiaokang Lei, Shi Qiu 0001, Luwen Yu, Pan Hui 0001
ACM Multimedia2
2024 DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm Understanding
abstract
Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks and approaches usually focus on sentence-level MSU. In document-level news, sarcasm clues are sparse or small and are often concealed in long text. Moreover, compared to sentence-level comments like tweets, which mainly focus on only a few trends or hot topics (e.g., sports events), content in the news is considerably diverse. Models created for sentence-level MSU may fail to capture sarcasm clues in document-level news. To fill this gap, we present a comprehensive benchmark for Document-level Multimodal Sarcasm Understanding (DocMSU). Our dataset contains 102,588 pieces of news with text-image pairs, covering 9 diverse topics such as health, business, etc. The proposed large-scale and diverse DocMSU significantly facilitates the research of document-level MSU in real-world scenarios. To take on the new challenges posed by DocMSU, we introduce a fine-grained sarcasm comprehension method to properly align the pixel-level image features with word-level textual features in documents. Experiments demonstrate the effectiveness of our method, showing that it can serve as a baseline approach to the challenging DocMSU.
Guoshun Nan, Binzhu Xie, Junrui Xu, Hehe Fan, Qimei Cui, Xiaofeng Tao 0001
AAAI4
2024 Uncovering what, why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
abstract
Video anomaly understanding (VAU) aims to automat-ically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial manufacturing. While existing VAU benchmarks primarily concentrate on anomaly detection and localization, our focus is on more practicality, prompting us to raise the following crucial questions: “what anomaly occurred?”,”why did it happen?”, and “how severe is this abnormal event?”. In pursuit of these answers, we present a comprehensive benchmark for Causation Understanding of Video Anomaly (CUVA). Specifically, each instance of the proposed benchmark involves three sets of human annotations to indicate the”what”, “why” and “how” of an anomaly, including 1) anomaly type, start and end times, and event descriptions, 2) natural language explanations for the cause of an anomaly, and 3) free text reflecting the effect of the abnormality. In addition, we also introduce MMEval, a novel evaluation metric designed to better align with human preferences for CUVA, facilitating the measurement of existing LLMs in comprehending the underlying cause and corresponding effect of video anoma-lies. Finally, we propose a novel prompt-based method that can serve as a baseline approach for the challenging CUVA. We conduct extensive experiments to show the superiority of our evaluation metric and the prompt-based approach. Our code and dataset are available at https://github.com/fesvhtr/CUVA.
Binzhu Xie, Guoshun Nan, Junrui Xu, Hangyu Liu 0001, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Linli Chen, Xuhuan Li, Jianhang Chen, Qimei Cui, Xiaofeng Tao 0001
CVPR3
2024 🐱 FunQA: Towards Surprising Video Comprehension
Binzhu Xie, Zitang Zhou, Bo Li 0080, Yuanhan Zhang, Jack Hessel, Ziwei Liu 0002
ECCV (1)1
2023 Change-Aware Network for Damaged Roads Recognition and Assessment Based on Multi-temporal Remote Sensing Imageries
Ming Wu 0001, Binzhu Xie
PRCV (4)4