Boshen Xu

dblp:293/8958 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 39% Language models and text generation · 14% 3D vision · 12%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language pretraining
egocentric video-language pretraining
1.722025
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining · NeurIPS 2025
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions? · ICLR 2025
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis
0.912025
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions? · ICLR 2025
Computer vision › Image recognition and object detection
hand-object interaction understanding
0.912025
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions? · ICLR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Computer vision › Vision and language
temporal grounding
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Computer vision › Vision and language › vision-language pretraining
video-language pre-training
0.912025
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining · NeurIPS 2025
Computer vision › 3D vision
hand-object interaction
0.712023
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-view World · ACM Multimedia 2023
Computer vision › 3D vision › 3d scene understanding › object relation reasoning
human-object interaction
0.712023
Open-Category Human-Object Interaction Pre-training via Language Modeling Framework · CVPR 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
Open-Category Human-Object Interaction Pre-training via Language Modeling Framework · CVPR 2023
Machine learning › Representation and self-supervised learning › pre-training
weakly supervised pre-training
0.712023
Open-Category Human-Object Interaction Pre-training via Language Modeling Framework · CVPR 2023
Computer vision › 3D vision
depth estimation
0.312025
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining · NeurIPS 2025
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
0.312025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.7verifiable reward · 0.9supervised fine-tuning · 0.9reinforcement learning · 0.9foundation model · 0.9depth estimation · 0.9vision-language pretraining · 0.7sequence generation · 0.7language modeling · 0.7interactive masking · 0.7
YearPublicationVenuePosition
2026 An Integrated Wearable Electromagnetic Sensing System with Wireless Vector Readout for Noninvasive Glucose Monitoring
Shiquan Wang, Boshen Xu, Yange Wang, Yuanjin Zheng
ISCAS2
2025 SPAFormer: Sequential 3D Part Assembly with Transformers
abstract
We introduce SPAFormer, an innovative model designed to overcome the combinatorial explosion challenge in the 3D Part Assembly (3D-PA) task. This task requires accu-rate prediction of each part's poses in sequential steps. As the number of parts increases, the possible assembly com-binations increase exponentially, leading to a combinato-rial explosion that severely hinders the efficacy of 3D-PA. SPAFormer addresses this problem by leveraging weak con-straints from assembly sequences, effectively reducing the solution space's complexity. Since the sequence of parts conveys construction rules similar to sentences structured through words, our model explores both parallel and au-toregressive generation. We further strengthen SPAFormer through knowledge enhancement strategies that utilize the attributes of parts and their sequence information, enabling it to capture the inherent assembly pattern and relationships among sequentially ordered parts. We also construct a more challenging benchmark named PartNet-Assembly cov-ering 21 varied categories to more comprehensively val-idate the effectiveness of SPAFormer. Extensive experi-ments demonstrate the superior generalization capabilities of SPAFormer, particularly with multi-tasking and in sce-narios requiring long-horizon assembly. Code is available at https://github.com/xuboshen/SPAFormer.
Boshen Xu, Sipeng Zheng, Qin Jin
3DV1
2025 Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
abstract
Egocentric video-language pretraining is a crucial step in advancing the understanding of hand-object interactions in first-person scenarios. Despite successes on existing testbeds, we find that current EgoVLMs can be easily misled by simple modifications, such as changing the verbs or nouns in interaction descriptions, with models struggling to distinguish between these changes. This raises the question: "Do EgoVLMs truly understand hand-object interactions?'' To address this question, we introduce a benchmark called $\textbf{EgoHOIBench}$, revealing the performance limitation of current egocentric models when confronted with such challenges. We attribute this performance gap to insufficient fine-grained supervision and the greater difficulty EgoVLMs experience in recognizing verbs compared to nouns. To tackle these issues, we propose a novel asymmetric contrastive objective named $\textbf{EgoNCE++}$. For the video-to-text objective, we enhance text supervision by generating negative captions using large language models or leveraging pretrained vocabulary for HOI-related word substitutions. For the text-to-video objective, we focus on preserving an object-centric feature space that clusters video representations based on shared nouns. Extensive experiments demonstrate that EgoNCE++ significantly enhances EgoHOI understanding, leading to improved performance across various EgoVLMs in tasks such as multi-instance retrieval, action recognition, and temporal understanding. Our code is available at https://github.com/xuboshen/EgoNCEpp.
Boshen Xu, Yang Du 0011, Zhinan Song, Sipeng Zheng, Qin Jin
ICLR1
2025 Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
abstract
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities. Project Page: https://xuboshen.github.io/Time-R1/.
Boshen Xu, Yang Du 0011, Kejun Lin, Zihan Xiao 0001, Zihao Yue, Jianzhong Ju, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang 0001, Junqi Lin, Jian Luan 0001, Qin Jin
NeurIPS3
2025 EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
abstract
Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most previous works learn from 1D text or 2D visual cues, such as bounding boxes, which inherently lack 3D understanding. To bridge this gap, we introduce EgoDTM, an Egocentric Depth- and Text-aware Model, jointly trained through large-scale 3D-aware video pretraining and video-text contrastive learning. EgoDTM incorporates a lightweight 3D-aware decoder to efficiently learn 3D-awareness from pseudo depth maps generated by depth estimation models. To further facilitate 3D-aware video pretraining, we enrich the original brief captions with hand-object visual cues by organically combining several foundation models. Extensive experiments demonstrate EgoDTM's superior performance across diverse downstream tasks, highlighting its superior 3D-aware visual understanding. Code: \url{https://github.com/xuboshen/EgoDTM}.
Boshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng, Qin Jin
NeurIPS1
2023 Open-Category Human-Object Interaction Pre-training via Language Modeling Framework
abstract
Human-object interaction (HOI) has long been plagued by the conflict between limited supervised data and a vast number of possible interaction combinations in real life. Current methods trained from closed-set data predict HOIs as fixed-dimension log its, which restricts their scalability to open-set categories. To address this issue, we introduce Open Cat, a language modeling framework that reformulates HOI prediction as sequence generation. By converting HOI triplets into a token sequence through a serialization scheme, our model is able to exploit the open-set vocabulary of the language modeling framework to predict novel interaction classes with a high degree of freedom. In addition, inspired by the great success of vision-language pre-training, we collect a large amount of weakly-supervised data related to HOI from image-caption pairs, and devise several auxiliary proxy tasks, including soft relational matching and human-object relation prediction, to pre-train our model. Extensive experiments show that our OpenCat significantly boosts HOI performance, particularly on a broad range of rare and unseen categories.
Sipeng Zheng, Boshen Xu, Qin Jin
CVPR2
2023 POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-view World
abstract
We humans are good at translating third-person observations of hand-object interactions (HOI) into an egocentric view. However, current methods struggle to replicate this ability of view adaptation from third-person to first-person. Although some approaches attempt to learn view-agnostic representation from large-scale video datasets, they ignore the relationships among multiple third-person views. To this end, we propose a Prompt-Oriented View-agnostic learning (POV) framework in this paper, which enables this view adaptation with few egocentric videos. Specifically, We introduce interactive masking prompts at the frame level to capture fine-grained action information, and view-aware prompts at the token level to learn view-agnostic representation. To verify our method, we establish two benchmarks for transferring from multiple third-person views to the egocentric view. Our extensive experiments on these benchmarks demonstrate the efficiency and effectiveness of our POV framework and prompt tuning techniques in terms of view adaptation and view generalization.
Boshen Xu, Sipeng Zheng, Qin Jin
ACM Multimedia1