Xinrun Xu

dblp:341/8034 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-4765-3952ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Foundational Skills Influence VLM-based Embodied Agents: A Native Perspective
abstract
Recent advances in vision–language models (VLMs) have shed light on human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents still rely on high-level commands or discretised action spaces—``non-native'' settings that diverge markedly from the real world. Moreover, current benchmarks focus exclusively on high-level tasks, while lacking joint evaluation and analysis on both low- and high-level. To bridge these gaps, we present \textbf{NativeEmbodied}, a challenging benchmark for VLM-driven embodied agents that adopts a unified, native low-level action space. Built upon diverse simulated scenes, NativeEmbodied first designs three representative high-level tasks in complex scenarios to evaluate overall performance. For more detailed and comprehensive performance analysis, we further decouple the entangled skills behind complex tasks and construct four types of low-level tasks, each corresponding to a key fundamental embodied skill. This joint evaluation across task and skill granularities enables a fine-grained assessment of embodied agent. Comprehensive experiments on the best VLMs reveal pronounced deficiencies in certain fundamental embodied skills. Further analysis shows that these bottlenecks severely constrain performance on high-level tasks. Our NativeEmbodied not only pinpoints the key challenges faced by current VLM-driven embodied agents, but also provides valuable insight for future development of this field.
Pi Bu, Keyu Pan, Xinrun Xu, Yingxiu Zhao, Tong Xu 0001
AAAI4
2026 DeepPhy: Benchmarking Agentic VLMs on Physical Reasoning
abstract
Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs' understanding and reasoning about fundamental physical principles through a series of challenging simulated environments. DeepPHY integrates multiple physical reasoning environments of varying difficulty levels and incorporates fine-grained evaluation metrics. Our evaluation finds that even state-of-the-art VLMs struggle to translate descriptive physical knowledge into precise, predictive control.
Xinrun Xu, Pi Bu, Börje Karlsson 0001, Tengtao Song, Zhiming Ding, Bo Zheng 0007
AAAI1
2025 Catastrophic Forgetting Mitigation via Discrepancy-Weighted Experience Replay
Xinrun Xu, Jianwen Yang, Qiuhong Zhang, Zhanbiao Lian, Zhiming Ding
ICANN (1)1
2025 DISEncoder: A Dual-Branch Query Encoder Using Graph Models for Distributed Databases
Jianwen Yang, Qiuhong Zhang, Zhiming Ding, Meiling Zhu, Xinrun Xu
ICANN (4)7
2025 Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
abstract
Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains unexplored. The hybrid text often appears in the form of hybrid long documents (HLDs), which far exceed the token limit of LLMs. Consequently, we apply an Automated Information Extraction framework (AIE) to enable LLMs to process the HLDs and carry out experiments to analyse four important aspects of information extraction from HLDs. Given the findings: 1) The effective way to select and summarize the useful part of a HLD. 2) An easy table serialization way is enough for LLMs to understand tables. 3) The naive AIE has adaptability in many complex scenarios. 4) The useful prompt engineering to enhance LLMs on HLDs. To address the issue of dataset scarcity in HLDs and support future work, we also propose the Financial Reports Numerical Extraction (FINE) dataset. The dataset and code are publicly available in the attachments.
Chongjian Yue, Xinrun Xu, Xiaojun Ma 0001, Lun Du, Zhiming Ding, Shi Han, Dongmei Zhang 0001, Qi Zhang 0066
ICASSP2
2025 KnobTuneX:LLM-Enhanced Automatic Database Tuning via Structured Reasoning
abstract
Cross-database knob tuning has long been recognized as a critical but complex task. Modern database systems expose hundreds of configuration knobs that play key roles in memory management, concurrency control, and query optimization. Proper tuning can significantly improve performance, while improper settings can cause severe degradation. Despite progress with black-box methods like reinforcement learning and bayesian optimization, as well as LLM-based tuning guides, challenges remain, such as modeling knob dependencies, high cold-start trial costs, and weak handling of dynamic workloads. We propose KnobTuneX, a structure-aware, LLM-enhanced framework for automatic database tuning that integrates domain knowledge, historical behaviors, and reasoning capabilities to adapt to diverse workloads. The approach features an offline learning stage to capture knob-performance relationships and build a historical RAG store, and an online inference stage that dynamically recommends knobs based on structured reasoning and historical insights. By explicitly modeling dependencies among knobs and leveraging LLMs for informed decision-making, the framework achieves both interpretability and adaptability. Finally, we evaluate KnobTuneX on PostgreSQL and show that the method outperforms mainstream approaches in efficiency, scalability, and overall tuning quality across OLTP and OLAP workloads. The implementation of our work can be found at https://github.com/vjwww/KnobTuneX.
Jianwen Yang, Qiuhong Zhang, Xinrun Xu, Yurong Wu, Zhiming Ding
ICDM3
2025 MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
abstract
MLLM agents demonstrate potential for complex embodied tasks by retrieving multimodal task-relevant trajectory data. However, current retrieval methods primarily focus on surface-level similarities of textual or visual cues in trajectories, neglecting their effectiveness for the specific task at hand. To address this issue, we propose a novel method, MART, which enhances the performance of embodied agents by utilizing interaction data to fine-tune an MLLM retriever based on preference learning, such that the retriever fully considers the effectiveness of trajectories and prioritize them for unseen tasks. We also introduce Trajectory Abstraction, a mechanism that leverages MLLMs' summarization capabilities to represent trajectories with fewer tokens while preserving key information, enabling agents to better comprehend milestones in the trajectory. Experimental results across various environments demonstrate our method significantly improves task success rates in unseen scenes compared to baseline methods. This work presents a new paradigm for multimodal retrieval in embodied agents, by fine-tuning a general-purpose MLLM as the retriever to assess trajectory effectiveness. All the code for benchmark tasks, simulator modifications and the MLLM retriever is available at https://github.com/PKU-RL/MART.
Junpeng Yue, Xinrun Xu, Börje Karlsson 0001, Zongqing Lu 0002
ICLR2
2025 Cradle: Empowering Foundation Agents towards General Computer Control
abstract
Despite their success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the General Computer Control (GCC) setting to restrict foundation agents to interact with software through the most unified and standardized interface, i.e., using screenshots as input and keyboard and mouse actions as output. We introduce Cradle, a modular and flexible LMM-powered framework, as a preliminary attempt towards GCC. Enhanced by six key modules, Information Gathering, Self-Reflection, Task Inference, Skill Curation, Action Planning, and Memory, Cradle is able to understand input screenshots and output executable code for low-level keyboard and mouse control after high-level planning and information retrieval, so that Cradle can interact with any software and complete long-horizon complex tasks without relying on any built-in APIs. Experimental results show that Cradle exhibits remarkable generalizability and impressive performance across four previously unexplored commercial video games (Red Dead Redemption 2, Cities:Skylines, Stardew Valley and Dealer’s Life 2), five software applications (Chrome, Outlook, Feishu, Meitu and CapCut), and a comprehensive benchmark, OSWorld. With a unified interface to interact with any software, Cradle greatly extends the reach of foundation agents thus paving the way for generalist agents.
Weihao Tan, Wentao Zhang 0007, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li 0003, Junpeng Yue, Jiechuan Jiang, Yewen Li, Ruyi An, Molei Qin, Chuqiao Zong, Longtao Zheng, Xiaoqiang Chai, Yifei Bi, Tianbao Xie, Pengjie Gu, Xiyun Li, Ceyao Zhang, Chaojie Wang 0001, Xinrun Wang, Börje Karlsson 0001, Bo An 0001, Shuicheng Yan, Zongqing Lu 0002
ICML3
2025 High-Quality Pseudo-Label Generation Based on Visual Prompt Assisted Cloud Model Update
abstract
Generating high-quality pseudo-labels on the cloud side is crucial for cloud-edge collaborative object detection, especially in dynamic traffic monitoring scenarios where the target data distribution continuously evolves. Existing methods often assume a perfectly reliable cloud model, neglecting the potential for errors in the cloud’s predictions, or employ simple adaptation techniques that struggle to handle complex distribution shifts. This paper proposes a novel Cloud-Adaptive High-Quality Pseudo-label generation algorithm (CA-HQP) that addresses these limitations by incorporating a learnable Visual Prompt Generator (VPG) and a dual feature alignment strategy into the cloud model updating process. The VPG enables parameter-efficient adaptation of the large pre-trained cloud model by injecting task-specific visual prompts into the model’s input, enhancing its flexibility without extensive fine-tuning. To mitigate domain discrepancies, CA-HQP introduces two complementary feature alignment techniques: a global Domain Query Feature Alignment (DQFA) that captures scene-level distribution shifts and a fine-grained Temporal Instance-Aware Feature Embedding Alignment (TIAFA) that addresses instance-level variations. Extensive experiments on the Bellevue traffic dataset, a challenging real-world traffic monitoring dataset, demonstrate that CA-HQP significantly improves the quality of pseudo-labels compared to existing state-of-the-art cloud-edge collaborative object detection methods. This translates to notable performance gains for the edge model, showcasing the effectiveness of CA-HQP in adapting to dynamic environments. Further ablation studies validate the contribution of each individual component (DQFA, TIAFA, VPG) and confirm the synergistic effect of combining global and instance-level feature alignment strategies. The results highlight the importance of adaptive cloud model updates and sophisticated domain adaptation techniques for achieving robust and accurate object detection in continuously evolving scenarios. The proposed CA-HQP algorithm provides a promising solution for enhancing the performance and reliability of cloud-edge collaborative object detection systems in real-world applications.
Xinrun Xu, Qiuhong Zhang, Jianwen Yang, Zhanbiao Lian, Zhiming Ding
IJCNN1
2025 From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
abstract
Achieving general agile whole-body control on humanoid robots remains a major challenge due to diverse motion demands and data conflicts. While existing frameworks excel in training single motion-specific policies, they struggle to generalize across highly varied behaviors due to conflicting control requirements and mismatched data distributions. In this work, we propose BumbleBee (BB), an expert-generalist learning framework that combines motion clustering and sim-to-real adaptation to overcome these challenges. BB first leverages an autoencoder-based clustering method to group behaviorally similar motions using motion features and motion descriptions. Expert policies are then trained within each cluster and refined with real-world data through iterative delta action modeling to bridge the sim-to-real gap. Finally, these experts are distilled into a unified generalist controller that preserves agility and robustness across all motion types. Experiments on two simulations and a real humanoid robot demonstrate that BB achieves state-of-the-art general whole-body control, setting a new benchmark for agile, robust, and generalizable humanoid performance in the real world.
Gang Ding, Weishuai Zeng, Xinrun Xu, Haobin Jiang, Zongqing Lu 0002
NeurIPS6
2024 Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear Queries
abstract
Tabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the Text2Analysis benchmark, incorporating advanced analysis tasks that go beyond the SQL-compatible operations and require more in-depth analysis. We also develop five innovative and effective annotation methods, harnessing the capabilities of large language models to enhance data quality and quantity. Additionally, we include unclear queries that resemble real-world user questions to test how well models can understand and tackle such challenges. Finally, we collect 2249 query-result pairs with 347 tables. We evaluate five state-of-the-art models using three different metrics and the results show that our benchmark presents introduces considerable challenge in the field of tabular data analysis, paving the way for more advanced research opportunities.
Mengyu Zhou, Xinrun Xu, Xiaojun Ma 0001, Rui Ding 0001, Lun Du, Yan Gao 0002, Ran Jia, Xu Chen 0022, Shi Han, Zejian Yuan, Dongmei Zhang 0001
AAAI3
2024 A Clustering Method with Graph Maximum Decoding Information
abstract
The clustering method based on graph models has garnered increased attention for its widespread applicability across various knowledge domains. Its adaptability to integrate seamlessly with other relevant applications endows the graph model-based clustering analysis with the ability to robustly extract "natural associations" or "graph structures" within datasets, facilitating the modelling of relationships between data points. Despite its efficacy, the current clustering method utilizing the graph-based model overlooks the uncertainty associated with random walk access between nodes and the embedded structural information in the data. To address this gap, we present a novel Clustering method for Maximizing Decoding Information within graph-based models, named CMDI. CMDI innovatively incorporates two-dimensional structural information theory into the clustering process, consisting of two phases: graph structure extraction and graph vertex partitioning. Within CMDI, graph partitioning is reformulated as an abstract clustering problem, leveraging maximum decoding information to minimize uncertainty associated with random visits to vertices. Empirical evaluations on three real-world datasets demonstrate that CMDI outperforms classical baseline methods, exhibiting a superior decoding information ratio (DI-R). Furthermore, CMDI showcases heightened efficiency, particularly when considering prior knowledge (PK). These findings underscore the effectiveness of CMDI in enhancing decoding information quality and computational efficiency, positioning it as a valuable tool in graph-based clustering analyses.
Xinrun Xu, Manying Lv, Zhanbiao Lian, Yurong Wu, Zhiming Ding
IJCNN1
2024 A Multi-constraint and Multi-objective Allocation Model for Emergency Rescue in IoT Environment
abstract
Emergency relief operations are essential in disaster aftermaths, necessitating effective resource allocation to minimize negative impacts and maximize benefits. In prolonged crises or extensive disasters, a systematic, multi-cycle approach is key for timely and informed decision-making. Leveraging advancements in IoT and spatio-temporal data analytics, we’ve developed the Multi-Objective Shuffled Gray-Wolf Frog Leaping Model (MSGW-FLM). This multi-constraint, multi-objective resource allocation model has been rigorously tested against 28 diverse challenges, showing superior performance in comparison to established models such as NSGA-II, IBEA, and MOEA/D. MSGW-FLM’s effectiveness is particularly notable in complex, multi-cycle emergency rescue scenarios, which involve numerous constraints and objectives. This model represents a significant step forward in optimizing resource distribution in emergency response situations.
Xinrun Xu, Zhanbiao Lian, Yurong Wu, Manying Lv, Zhiming Ding
ISCAS1
2023 Optimal Node Embedding Dimension Selection Using Overall Entropy
Xinrun Xu, Zhiming Ding, Yurong Wu, Qinglong Cui
ICANN (9)1