VLDB 2026 Research / reviewers in the wild / expert
Wenxuan Song
dblp:372/3054
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot PerceiverabstractRecent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispersed. To guide the visual attention grounding on the correct target, we propose ReconVLA, a reconstructive VLA model with an implicit grounding paradigm. Conditioned on the model's visual outputs, a diffusion transformer aims to reconstruct the gaze region of the image, which corresponds to the target manipulated objects. This process prompts the VLA model to learn fine-grained representations and accurately allocate visual attention, thus effectively leveraging task-specific visual information and conducting precise manipulation. Moreover, we curate a large-scale pretraining dataset comprising over 100k trajectories and 2 million data samples from open-source robotic datasets, further boosting the model’s generalization in visual reconstruction. Extensive experiments in simulation and the real world demonstrate the superiority of our implicit grounding method, showcasing its capabilities of precise manipulation and generalization. Wenxuan Song, Han Zhao 0008, Pengxiang Ding, Haodong Yan, Haoang Li |
AAAI | 1 |
| 2026 | VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelabstractVision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how to effectively bridge vision-language (VL) representations to action (A). We introduce VLA-Adapter, a novel paradigm designed to reduce the reliance of VLA models on large-scale VLMs and extensive pre-training. To this end, we first systematically analyze the effectiveness of various VL conditions and present key findings on which conditions are essential for bridging perception and action spaces. Based on these insights, we propose a lightweight Policy module with Bridge Attention, which autonomously injects the optimal condition into the action space. In this way, our method achieves high performance using only a 0.5B-parameter backbone, without any robotic data pre-training. Extensive experiments on both simulated and real-world robotic benchmarks show that VLA-Adapter not only achieves state-of-the-art level performance, but also offers the fast inference speed reported to date. Furthermore, thanks to the proposed advanced bridging paradigm, VLA-Adapter enables the training of a powerful VLA model on a single consumer-grade GPU, greatly lowering the barrier to deploying VLA model. Yihao Wang 0006, Pengxiang Ding, Can Cui 0008, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao 0008, Pengxu Hou, Siteng Huang, Ru Zhang 0002 |
AAAI | 7 |
| 2026 | "In my defense, only three hours on Instagram": Designing Toward Digital Self-Awareness and WellbeingabstractScreen use pervades daily life, shaping work, leisure, and social connections while raising concerns for digital wellbeing. Yet, reducing screen time alone risks oversimplifying technology’s role and neglecting its potential for meaningful engagement. We posit self-awareness—reflecting on one’s digital behavior—as a critical pathway to digital wellbeing. We developed WellScreen, a lightweight probe that scaffolds daily reflection by asking people to estimate and report smartphone use. In a two-week deployment with college students (\(\mathtt {N}\)=25) focused on generating formative insights, we examined how discrepancies between estimated and actual usage shaped digital awareness and wellbeing. Participants often underestimated productivity and social media while overestimating entertainment app use. They showed a 10% improvement in positive affect, rating WellScreen as moderately useful. Interviews revealed that structured reflection supported recognition of patterns, adjustment of expectations, and more intentional engagement with technology. Our findings highlight the promise of lightweight reflective interventions for supporting self-awareness and intentional digital engagement, offering implications for designing digital wellbeing tools. Karthik S. Bhat, Jiayue Melissa Shi, Wenxuan Song, Dong Whi Yoo, Koustuv Saha |
CHI | 3 |
| 2026 | Does Sequencing Matter? Evaluating AI and Human Simulations for High-Stakes Communication Training in Law EnforcementabstractTraining professionals in high-stakes, trauma-informed communication is critical across domains such as law enforcement, healthcare, and counseling. While live role-play with trained actors remains the gold standard, it is resource-intensive and emotionally demanding. Generative AI offers scalable alternatives, but what is gained or lost when training shifts to AI? We developed an AI-powered sexual assault victim interview training system and conducted a mixed-methods study with 35 police recruits, each completing both an AI-based and a live, actor-based training session. By varying the sequence (AI-first vs. human-first), we examined differences in self-efficacy, perceptions of the AI system, and perceived learning experience. Although both modalities supported learning, the order in which they were experienced significantly shaped learners’ emotional engagement, sense of preparedness, and interpretation of each simulation’s role. Building on these insights, we introduce a conceptual design framework that identifies social–emotional, temporal, and embodied distance as key pedagogical dimensions, and we offer implications for sequencing hybrid simulations to scaffold preparation, performance, and reflection. Our findings position AI not as a replacement for human realism, but as a complementary modality that expands opportunities for safe, scalable practice in sensitive communication training. Kyrian Liang, Qingxiao Zheng 0001, Wenxuan Song, Jen Whiting, Mike Yao 0001, Caroline G. L. Cao |
CHI | 4 |
| 2026 | Should the AI Speak First? Evaluating Proactive vs. Reactive Facilitation in Mixed-Reality Medical TrainingabstractAs AI support tools become more common in immersive medical training, designers face a critical interaction-design question: When should an AI facilitator take initiative, and when should it wait for the learner? To investigate this design tension, we compared two versions of an AI facilitator in a mixed-reality (XR) lumbar puncture simulator training conditions: one in which the AI proactively initiated guidance and encouragement, and another in which the AI responded only when prompted. Using a mixed-methods approach, we examined how medical students (n=22) engaged with, interpreted, and reacted to these two facilitation styles. We found no significant differences in learning outcomes, interaction frequency, or overall experience ratings. However, interviews and behavioral analyses revealed nuanced differences in how learners perceived AI interventions across distinct task phases. AI-initiated support was seen as helpful in some moments and disruptive in others, depending on task phase, cognitive load, and personal preferences. Based on these findings, we contribute a boundary framework which offers actionable design guidance for calibrating AI proactivity in immersive training systems, and extends HCI research on proactive agents and human–AI collaboration within high-cognitive-load environments. Wenxuan Song, Jianwei Ni, Qingxiao Zheng 0001, Kyrian Liang, Mike Yao 0001, Caroline G. L. Cao |
CHI | 2 |
| 2025 | WaterSplatting: Fast Underwater 3D Scene Reconstruction Using Gaussian SplattingabstractThe underwater 3D scene reconstruction is a challenging, yet interesting problem with applications ranging from naval robots to VR experiences. The problem was successfully tackled by fully volumetric NeRF-based methods which can model both the geometry and the medium (water). Unfortunately, these methods are slow to train and do not offer real-time rendering. More recently, 3D Gaussian Splatting (3DGS) method offered a fast alternative to NeRFs. However, because it is an explicit method that renders only the geometry, it cannot render the medium and is therefore unsuited for underwater reconstruction. Therefore, we propose a novel approach that fuses volumetric rendering with 3DGS to handle underwater data effectively. Our method employs 3DGS for explicit geometry representation and a separate volumetric field (queried once per pixel) for capturing the scattering medium. This dual representation further allows the restoration of the scenes by removing the scattering medium. Our method outperforms state-of-the-art NeRF-based methods in rendering quality on the underwater SeaThru-NeRF dataset. Furthermore, it does so while offering real-time rendering performance, addressing the efficiency limitations of existing methods. Huapeng Li, Wenxuan Song, Tianao Xu, Alexandre Elsig, Jonas Kulhanek |
3DV | 2 |
| 2025 | Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingabstractRecent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness. Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge |
CVPR | 13 |
| 2025 | GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection
Wenxue Li 0003, Tian Ye 0001, Xinyu Xiong, Jinbin Bai, Wenxuan Song, Zhaohu Xing, Lie Ju, Guanbin Li, Lei Zhu 0003 |
ICCV | 6 |
| 2025 | MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action ModelsabstractDeveloping versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinforcement learning (RL) for fine-tuning large-scale VLA models with a large amount of mixed-quality data. MoRE integrates multiple low-rank adaptation modules as distinct experts within a dense multi-modal large language model (MLLM), forming a sparse-activated mixture-of-experts model. This design enables the model to effectively adapt to a wide array of downstream tasks. Moreover, we employ a reinforcement learning-based training objective to train our model as a Q-function after deeply exploring the structural properties of our tasks. Effective learning from automatically collected mixed-quality data enhances data efficiency and model performance. Extensive experiments demonstrate that MoRE outperforms all baselines across six different skills and exhibits superior generalization capabilities in out-of-distribution scenarios. We further validate our method in real-world scenarios, confirming the practicality of our approach and laying a solid foundation for future research on multi-task learning in quadruped robots. Han Zhao 0008, Wenxuan Song, Xinyang Tong, Pengxiang Ding, Xuelian Cheng, ZongYuan Ge |
ICRA | 2 |
| 2025 | PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel DecodingabstractVision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in VLA models with increased chunking sizes. This reduces the inference efficiency. Therefore, accelerating VLA integrated with action chunking is an urgent need. To tackle this problem, we propose PD-VLA, the first parallel decoding framework for VLA models integrated with action chunking. Our framework reformulates autoregressive decoding as a nonlinear system solved by parallel fixed-point iterations. This approach preserves model performance with mathematical guarantees while significantly improving decoding speed. In addition, it enables training-free acceleration without architectural changes, as well as seamless synergy with existing acceleration techniques. Extensive simulations validate that our PD-VLA maintains competitive success rates while achieving 2.52× execution frequency on manipulators (with 7 degrees of freedom) compared with the fundamental VLA model. Furthermore, we experimentally identify the most effective settings for acceleration. Finally, real-world experiments validate its high applicability across different tasks. Wenxuan Song, Pengxiang Ding, Han Zhao 0008, Zhide Zhong, ZongYuan Ge, Jun Ma 0008, Haoang Li |
IROS | 1 |
| 2024 | QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Pengxiang Ding, Han Zhao 0008, Wenxuan Song, Min Zhang 0068, Siteng Huang, Ningxi Yang |
ECCV (5) | 4 |
| 2024 | GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped RobotabstractMulti-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in collecting training datasets. In this paper, we propose GeRM (Generalist Robotic Model). We utilize offline reinforcement learning to optimize data utilization strategies to learn from both demonstrations and sub-optimal data, thus surpassing the limitations of human demonstrations. Thereafter, we employ a transformer-based VLA network to process multi-modal inputs and output actions. By introducing the Mixture-of-Experts structure, GeRM allows faster inference speed with higher whole model capacity, and thus resolves the issue of limited RL parameters, enhancing model performance in multi-task learning while controlling computational costs. Through a series of experiments, we demonstrate that GeRM outperforms other methods across all tasks, while also validating its efficiency in both training and inference processes. Additionally, we uncover its potential to acquire emergent skills. Additionally, we contribute the QUARD-Auto dataset, collected automatically to support our training approach and foster advancements in multi-task quadruped robot learning. This work presents a new paradigm for reducing the cost of collecting robot data and driving progress in the multi-task learning community.You can reach our project and video through the link: https://songwxuan.github.io/GeRM/. Wenxuan Song, Han Zhao 0008, Pengxiang Ding, Can Cui 0008, Shangke Lyu, Yaning Fan |
IROS | 1 |
| 2024 | ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-IdentificationabstractTo address the occlusion issues in person Re-Identification (ReID) tasks, many methods have been proposed to extract part features by introducing external spatial information. However, due to missing part appearance information caused by occlusion and noisy spatial information from external model, these purely vision-based approaches fail to correctly learn the features of human body parts from limited training data and struggle in accurately locating body parts, ultimately leading to misaligned part features. To tackle these challenges, we propose a Prompt-guided Feature Disentangling method (ProFD), which leverages the rich pre-trained knowledge in the textual modality facilitate model to generate well-aligned part features. ProFD first designs part-specific prompts and utilizes noisy segmentation mask to preliminarily align visual and textual embedding, enabling the textual prompts to have spatial awareness. Furthermore, to alleviate the noise from external masks, ProFD adopts a hybrid-attention decoder, ensuring spatial and semantic consistency during the decoding process to minimize noise impact. Additionally, to avoid catastrophic forgetting, we employ a self-distillation strategy, retaining pre-trained knowledge of CLIP to mitigate over-fitting. Evaluation results on the Market1501, DukeMTMC-ReID, Occluded-Duke, Occluded-ReID, and P-DukeMTMC datasets demonstrate that ProFD achieves state-of-the-art results. Can Cui 0008, Siteng Huang, Wenxuan Song, Pengxiang Ding, Min Zhang 0068 |
ACM Multimedia | 3 |