Bingchen Miao

dblp:389/6313 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0009-9028-9058ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 20% Vision and language · 19% Reinforcement learning · 18%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
virtual agents
1.722025
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark · ICML 2025
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities · ICML 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.222026
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities · ICML 2025
Evolving Generalist Virtual Agents with Generative and Associative Memory · AAAI 2026
Natural language and speech › Language models and text generation › LLM agents
agent memory
1.012026
Evolving Generalist Virtual Agents with Generative and Associative Memory · AAAI 2026
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning
1.012026
Evolving Generalist Virtual Agents with Generative and Associative Memory · AAAI 2026
Machine learning › Time series and sequential data
anomaly detection
0.912025
Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark · ACM Multimedia 2025
Computer vision › Vision and language › multimodal representation
missing modality learning
0.912025
Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark · ACM Multimedia 2025
Natural language and speech › Language models and text generation › LLM agents
multimodal large language model agent
0.912025
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities · ICML 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark · ICML 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark · ICML 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark · ACM Multimedia 2025
Machine learning › Reinforcement learning
action selection
0.312025
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark · ICML 2025
Computer vision › Image recognition and object detection › industrial visual inspection
industrial defect detection
0.312025
Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark · ACM Multimedia 2025
Natural language and speech › Language models and text generation
test-time scaling
0.312025
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark · ICML 2025

Methods — techniques the papers use, named apart from their topics

spreading activation · 1.0memory graph · 1.0generative recombination · 1.0triple-m strategy · 0.9pseudo-modality · 0.9multimodal transformer · 0.9knowledge distillation · 0.9graph-based task synthesis · 0.9automated benchmark generation · 0.9MCTS-P · 0.9
YearPublicationVenuePosition
2026 Evolving Generalist Virtual Agents with Generative and Associative Memory
abstract
Generalist Virtual Agents (GVAs) powered by Multimodal Large Language Models (MLLMs) exhibit impressive capabilities. However, their long-term learning is hampered by a core limitation: a failure to evolve beyond existing trajectories. This stems from memory systems that treat experiences as isolated fragments and rely on brittle semantic retrieval, preventing the synthesis of novel solutions from disparate knowledge. To address this, we introduce CA3Mem, a framework inspired by the human hippocampus that organizes experiences into a structured memory graph. Leveraging this graph, CA3Mem features two key innovations: 1) a generative memory recombination mechanism that synthesizes novel solutions to drive agent evolution, and 2) an associative retrieval algorithm that employs spreading activation to recall a comprehensive and contextually-aware set of experiences. Experiments on OSWorld and WebArena demonstrate that CA3Mem significantly enhances agent capabilities, leading to marked improvements in long-horizon planning, compositional generalization for novel tasks, and continuous adaptation from experience.
Zhenkui Zhang, Wendong Bu, Kaihang Pan, Bingchen Miao, Wenqiao Zhang, Guoming Wang, Wei Ji 0008, Juncheng Li 0006, Siliang Tang
AAAI4
2025 What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
abstract
As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate. Training on our graph-structured data shows that it improves generalization across environments. We conduct multidimensional evaluations for virtual agents, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io.
Wendong Bu, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Liyunfei, Mengze Li 0001, Wei Ji 0008, Juncheng Li 0006, Siliang Tang, Yueting Zhuang
ICML5
2025 Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
abstract
The development of Generalist Virtual Agents (GVAs) has shown significant promise in autonomous task execution. However, current training paradigms face critical limitations, including reliance on outcome supervision and labor-intensive human annotations. To address these challenges, we propose Similar, a step-wise multi-dimensional generalist reward model, which offers fine-grained signals for agent training and can choose better actions for inference-time scaling. Specifically, we begin by systematically defining five dimensions for evaluating agent actions. Building on this framework, we design an MCTS-P algorithm to automatically collect and annotate step-wise, five-dimensional agent execution data. Using this data, we train Similar with our crafted Triple-M strategy. Furthermore, we introduce the first benchmark in the virtual agent domain for step-wise, multi-dimensional reward model training and evaluation, named SRM. This benchmark consists of two components: SRMTrain, which serves as the training set for Similar, and SRMEval, a manually selected test set for evaluating the reward model. Experimental results demonstrate that Similar, through its step-wise, multi-dimensional assessment and synergistic gain, provides GVAs with effective intermediate signals during both training and inference-time scaling. The code is available at https://github.com/antgroup/Similar.
Bingchen Miao, Minghe Gao, Wendong Bu, Wenqiao Zhang, Siliang Tang, Tat-Seng Chua, Juncheng Li 0006
ICML1
2025 Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark
abstract
Multimodal Industrial Anomaly Detection (MIAD)-fusing 3D point clouds and 2D RGB for product defect detection-is critical to quality inspection. However, existing MIAD methods assume all modalities are available and paired, overlooking real-scenario modality-missing and risking overfitting to incomplete data. To address these, we conduct the first comprehensive study on Modality-Incomplete Industrial Anomaly Detection (MIIAD) and establish MIIAD Bench , a benchmark covering diverse missing settings. Meanwhile, we propose RADAR, a robust two-stage Robust modAlity-instructive fusing & Detecting frAmewoRk. RADAR integrates i) a Modality-Incomplete Instruction mechanism-guiding the multimodal Transformer to focus more on available modal info, and ii) a Double-Pseudo Hybrid Module to highlight unique modality combinations and reduce overfitting. Our results show RADAR outperforms prior methods markedly on MIIAD Bench.
Bingchen Miao, Wenqiao Zhang, Juncheng Li 0006, Wangyu Wu, Siliang Tang, Zhaocheng Li, Jun Xiao 0001, Yueting Zhuang
ACM Multimedia1