EDBT 2026 Demo / reviewers in the wild / expert
Zhixuan Shen
dblp:372/2812
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Robot navigation and mapping · 33% Vision and language · 33% Language models and text generation · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › chain-of-thought reasoning
multimodal chain-of-thought |
0.9 | 1 | 2025 | Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration · AAAI 2025 |
Robotics › Robot navigation and mapping › visual navigation
semantic navigation |
0.9 | 1 | 2025 | Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration · AAAI 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9chain-of-thought · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score CollaborationabstractUnderstanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot systems. Previous methods primarily focused on single-robot centralized planning strategies, which severely limited exploration efficiency. Recent research has considered decentralized planning strategies for multiple robots, assigning separate planning models to each robot, but these approaches often overlook communication costs. In this work, we propose Multimodal Chain-of-Thought Co-Navigation (MCoCoNav), a modular approach that utilizes multimodal Chain-of-Thought to plan collaborative semantic navigation for multiple robots. MCoCoNav combines visual perception with Vision Language Models (VLMs) to evaluate exploration value through probabilistic scoring, thus reducing time costs and achieving stable outputs. Additionally, a global semantic map is used as a communication bridge, minimizing communication overhead while integrating observational results. Guided by scores that reflect exploration trends, robots utilize this map to assess whether to explore new frontier points or revisit history nodes. Experiments on HM3D_v0.2 and MP3D demonstrate the effectiveness of our approach. Zhixuan Shen, Haonan Luo 0002, Kexun Chen, Fengmao Lv, Tianrui Li 0001 |
AAAI | 1 |
| 2025 | Role-Specific Reward Design with Large Language Model for StarCraft IIabstractReward acts as a signal to guide the agent’s learning process in Reinforcement Learning (RL), evaluating and assigning rewards to the agent’s actions based on theiralignment with goals. Designing reward is challenging in multiagent environment such as StarCraft II benchmark since agents face credit allocation and role adaptation problems. Recent studies have successfully exploited the language understanding and reasoning capabilities of large language models (LLMs) to learn manipulation tasks. Impressed by the remarkable power of LLMs, this paper employs LLMs as role-specific reward designer for playing StarCraft II, making rewards more flexible and task-oriented. Firstly, we develop an interactive text and multiagent RL environment to study real-time strategy generation in StarCraft II. Secondly, we use LLMs to interpret the game situation and understand agent roles from user instructions. Then, by assigning appropriate subtasks, LLMs quantify the completion of these subtasks to generate role-specific rewards. Further, credit assignment problem is addressed by introducing dynamic reward weights in value decomposition method. In StarCraft II maps, experiments show that role-aligned RL agents trained with our framework achieve superior policy performance, and win rate results demonstrates the effectiveness of our approach in decision-making for micromanagement and long-term planning. Haonan Lou, Zhixuan Shen, Tianrui Li 0001 |
ICASSP | 5 |
| 2025 | A Continual Learning Approach for Embodied Question Answering with Generative Adversarial Imitation LearningabstractEmbodied Question Answering (EQA) is a task in artificial intelligence where an intelligent agent is required to answer questions about its environment. For example, to answer a question such as "Is the TV on or off?", the agent must navigate to the room with the TV and answer with either "On." or "Off." after recognizing the status. Unlike traditional question-answering systems that rely solely on text or static images, EQA involves agents that can move through a physical or simulated space, interact with the environment, and gather information to respond accurately. The agent must interpret both visual and linguistic inputs, navigate the environment, and complete tasks or locate objects based on the user’s questions. However, in the real world, the agent always faces unseen environments (i.e. different people’s houses), which makes the pre-trained model fail. Meanwhile, re-training in an unseen environment can cause high costs. Therefore, it is significant for the agent to learn continually by itself to cope with the challenges of unseen environments. In this work, we proposed a continual learning method based on generative adversarial imitation learning and self-supervision to support the agent when facing unseen environments. Besides, we designed a policy generator and policy quality discriminator to generate action policy sequences and evaluate the quality of the policy, respectively. Extensive experiments on the MP3D-EQA dataset demonstrate that our method reaches state-of-the-art performance. Haonan Luo 0002, Zihang Wang 0002, Zhixuan Shen, Tianrui Li 0001 |
ICASSP | 5 |
| 2025 | MoPE: Mixture of Policy Experts and Verification with Multimodal Information for Instance ImageGoal NavigationabstractInstance ImageGoal Navigation (IIN) entails an agent autonomously seeking out a specific object instance depicted by a goal image in an unknown environment. While Large Language Models (LLMs) have shown promise in navigation tasks similar to IIN, their application to IIN remains unexplored. Furthermore, existing LLM-based exploration faces challenges such as inaccurate reasoning due to informative environmental information available to the agent, especially in the early episode stages, and the inability of reinforcement learning(RL) exploration to fully leverage gathered information. Moreover, previous IIN methods did not productively verify potentially distant goal objects discovered during exploration. This work proposes MoPE–Mixture of Policy Experts for exploration and potential goal verification with multimodal information when exploring. Specifically, the hybrid exploration policy comprises an LLM and an RL-based Policy Network (RLPN) to generate an exploration goal to explore efficiently. Our MoPE model surpasses prior approaches on the HM3D datasets significantly. Yijie Zeng, Kexun Chen, Zhixuan Shen, Haonan Luo 0002, Tianrui Li 0001 |
ICME | 4 |
| 2024 | Adversarial Training with OCR modality Perturbation for Scene-Text Visual Question AnsweringabstractScene-Text Visual Question Answering (ST-VQA) aims to understand scene text in images and answer questions related to the text content. Most existing methods heavily rely on the accuracy of Optical Character Recognition (OCR) systems, and aggressive fine-tuning based on limited spatial location information and erroneous OCR text information often leads to inevitable overfitting. In this paper, we propose a multimodal adversarial training architecture with spatial awareness capabilities. Specifically, we introduce an Adversarial OCR Enhancement (AOE) module, which leverages adversarial training in the embedding space of OCR modality to enhance fault-tolerant representation of OCR texts, thereby reducing noise caused by OCR errors. Simultaneously, We add a Spatial-Aware Self-Attention (SASA) mechanism to help the model better capture the spatial relationships among OCR tokens. Various experiments demonstrate that our method achieves significant performance improvements on both the ST-VQA and TextVQA datasets and provides a novel paradigm for multimodal adversarial training. Zhixuan Shen, Haonan Luo 0002, Tianrui Li 0001 |
ICME | 1 |
| 2024 | VLAI: Exploration and Exploitation based on Visual-Language Aligned Information for Robotic Object Goal Navigation
Haonan Luo 0002, Yijie Zeng, Kexun Chen, Zhixuan Shen, Fengmao Lv |
Image Vis. Comput. | 5 |