Xuhong He

dblp:80/11379 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0002-6360-4488ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Multi-agent systems · 55% Language models and text generation · 14% Vision and language · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-language information retrieval
1.012026
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation · SIGIR 2026
Information retrieval › evaluation › benchmark
multilingual benchmark
1.012026
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation · SIGIR 2026
Knowledge, reasoning and agents › Multi-agent systems › multi-robot coordination
cooperative search
0.912025
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
0.912025
SimWorld: An Open-ended Simulator for Agents in Physical and Social Worlds · NeurIPS 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
SimWorld: An Open-ended Simulator for Agents in Physical and Social Worlds · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
multi-robot coordination
0.912025
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration · NeurIPS 2025
Robotics › Robot navigation and mapping › mobile robot navigation › outdoor navigation
urban navigation
0.912025
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration · NeurIPS 2025
Computer vision › Vision and language
vision-and-language navigation
0.912025
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration · NeurIPS 2025
Information retrieval
evaluation
0.312026
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation · SIGIR 2026
Information retrieval › evaluation › offline evaluation
simulation-based evaluation
0.312026
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation · SIGIR 2026
Robotics › Motion planning and robot control › motion planning
replanning
0.312025
SimWorld: An Open-ended Simulator for Agents in Physical and Social Worlds · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 2.6unreal engine 5 · 1.7procedural scene generation · 1.7large language model query simulation · 1.0large language model · 0.9
YearPublicationVenuePosition
2026 Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation
abstract
Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collections for Chinese, Japanese, Korean, and English, using an LLM-based query simulation framework. We systematically study how prompt language and source document language affect the fidelity of simulated ToT queries, validating synthetic queries through system rank correlation against real user queries. Our results show that effective ToT simulation requires language-aware design choices: non-English language sources are generally important, while English Wikipedia can be beneficial when non-English sources provide insufficient information for query generation. Based on these findings, we release four ToT test collections with 5,000 queries per language across multiple domains. This work provides the first large-scale multilingual ToT benchmark and offers practical guidance for constructing realistic ToT datasets beyond English.
Xuhong He, To Eun Kim, Maik Fröbe, Jaime Arguello, Bhaskar Mitra 0001, Fernando Diaz 0001
SIGIR1
2025 SimWorld: An Open-ended Simulator for Agents in Physical and Social Worlds
abstract
While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building agents that can survive and thrive in the real world (e.g., by autonomously earning income) requires massive-scale interaction, reasoning, training, and evaluation across diverse scenarios. However, existing world simulators for such development fall short: they often rely on limited hand-crafted environments, simulate simplified game-like physics and social rules, and lack native support for LLM/VLM agents. We introduce SimWorld, a new simulator built on Unreal Engine 5, designed for developing and evaluating LLM/VLM agents in rich, real-world-like settings. SimWorld offers three core capabilities: (1) realistic, open-ended world simulation, including accurate physical and social dynamics and language-driven procedural environment generation; (2) rich interface for LLM/VLM agents, with multi-modal world inputs/feedback and open-vocabulary action outputs at varying levels of abstraction; and (3) diverse physical and social reasoning scenarios that are easily customizable by users. We demonstrate SimWorld by deploying frontier LLM agents (e.g., Gemini-2.5-Flash, Claude-3.5, GPT-4o, and DeepSeek-Prover-V2) on both short-horizon navigation tasks requiring grounded re-planning, and long-horizon multi-agent food delivery tasks involving strategic cooperation and competition. The results reveal distinct reasoning patterns and limitations across models. We open-source SimWorld and hope it becomes a foundational platform for advancing real-world agent intelligence across disciplines. Please refer to the project website for the most up-to-date information: http://simworld.org/.
Xiaokang Ye, Xuhong He, Yiming Liang, Yiqing Yang, Mrinaal Dogra, Xianrui Zhong, Eric Liu 0006, Kevin Benavente, Rajiv Mandya Nagaraju, Dhruv Vivek Sharma, Ziqiao Ma 0001, Tianmin Shu, Zhiting Hu, Lianhui Qin
NeurIPS4
2025 SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration
abstract
Recent advances in foundation models have shown promising results in developing generalist robotics that can perform diverse tasks in open-ended scenarios given multimodal inputs. However, current work has been mainly focused on indoor, household scenarios. In this work, we present SimWorld-Robotics (SWR), a simulation platform for embodied AI in large-scale, photorealistic urban environments. Built on Unreal Engine 5, SWR procedurally generates unlimited photorealistic urban scenes populated with dynamic elements such as pedestrians and traffic systems, surpassing prior urban simulations in realism, complexity, and scalability. It also supports multi-robot control and communication. With these key features, we build two challenging robot benchmarks: (1) a multimodal instruction-following task, where a robot must follow vision-language navigation instructions to reach a destination in the presence of pedestrians and traffic; and (2) a multi-agent search task, where two robots must communicate to cooperatively locate and meet each other. Unlike existing benchmarks, these two new benchmarks comprehensively evaluate a wide range of critical robot capacities in realistic scenarios, including (1) multimodal instructions grounding, (2) 3D spatial reasoning in large environments, (3) safe, long-range navigation with people and traffic, (4) multi-robot collaboration, and (5) grounded communication. Our experimental results demonstrate that state-of-the-art models, including vision-language models (VLMs), struggle with our tasks, lacking robust perception, reasoning, and planning abilities necessary for urban environments.
Xiaokang Ye, Jianzhi Shen, Tianai Yue, Muhammad Faayez, Xuhong He, Xiyan Zhang, Ziqiao Ma 0001, Lianhui Qin, Zhiting Hu, Tianmin Shu
NeurIPS8