Haojun Shi

dblp:386/3000 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Knowledge representation and reasoning · 48% Language models and text generation · 28% Vision and language · 24%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
web agents
1.012026
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.912025
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind · AAAI 2025
Computer vision › Vision and language
multimodal reasoning
0.912025
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind · AAAI 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind
0.912025
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind · AAAI 2025
Human-AI interaction › GUI agent
GUI grounding
0.312026
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users · AAAI 2026

Methods — techniques the papers use, named apart from their topics

large language model · 2.0benchmark evaluation · 2.0large multimodal model · 0.9inverse planning · 0.9
YearPublicationVenuePosition
2026 RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
abstract
To achieve successful assistance with long-horizon web-based tasks, AI agents must be able to sequentially follow real-world user instructions over a long period. Unlike existing web-based agent benchmarks, sequential instruction following in the real world poses significant challenges beyond performing a single, clearly defined task. For instance, real-world human instructions can be ambiguous, require different levels of AI assistance, and may evolve over time, reflecting changes in the user's mental state. To address this gap, we introduce RealWebAssist, a novel benchmark designed to evaluate sequential instruction-following in realistic scenarios involving long-horizon interactions with the web, visual GUI grounding, and understanding ambiguous real-world user instructions. RealWebAssist includes a dataset of sequential instructions collected from real-world human users. Each user instructs a web-based assistant to perform a series of tasks on multiple websites. A successful agent must reason about the true intent behind each instruction, keep track of the mental state of the user, understand user-specific routines, and ground the intended tasks to actions on the correct GUI elements. Our experimental results show that state-of-the-art models struggle to understand and ground user instructions, posing critical challenges in following real-world user instructions for long-horizon web assistance.
Suyu Ye, Haojun Shi, Darren Shih, Hyokun Yun, Tanya G. Roosta, Tianmin Shu
AAAI2
2025 MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
abstract
Understanding people's social interactions in complex real-world scenarios often relies on intricate mental reasoning. To truly understand how and why people interact with one another, we must infer the underlying mental states that give rise to the social interactions, i.e., Theory of Mind reasoning in multi-agent interactions. Additionally, social interactions are often multi-modal -- we can watch people's actions, hear their conversations, and/or read about their past behaviors. For AI systems to successfully and safely interact with people in real-world environments, they also need to understand people's mental states as well as their inferences about each other's mental states based on multi-modal information about their interactions. For this, we introduce MuMA-ToM, a Multi-modal Multi-Agent Theory of Mind benchmark. MuMA-ToM is the first multi-modal Theory of Mind benchmark that evaluates mental reasoning in embodied multi-agent interactions. In MuMA-ToM, we provide video and text descriptions of people's multi-modal behavior in realistic household environments. Based on the context, we then ask questions about people's goals, beliefs, and beliefs about others' goals. We validated MuMA-ToM in a human experiment and provided a human baseline. We also proposed a novel multi-modal, multi-agent ToM model, LIMP (Language model-based Inverse Multi-agent Planning). Our experimental results show that LIMP significantly outperforms state-of-the-art methods, including large multi-modal models (e.g., GPT-4o, Gemini-1.5 Pro) and a recent multi-modal ToM model, BIP-ALM.
Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Leyla Isik, Yen-Ling Kuo, Tianmin Shu
AAAI1