EDBT 2026 Demo / reviewers in the wild / expert
Mengkang Hu
dblp:321/0644
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2025
0009-0009-3779-3378ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AnalogCoder: Analog Circuit Design via Training-Free Code GenerationabstractAnalog circuit design is a significant task in modern chip technology, focusing on the selection of component types, connectivity, and parameters to ensure proper circuit functionality. Despite advances made by Large Language Models (LLMs) in digital circuit design, the complexity and scarcity of data in analog circuitry pose significant challenges. To mitigate these issues, we introduce AnalogCoder, the first training-free LLM agent for designing analog circuits through Python code generation. Firstly, AnalogCoder incorporates a feedback-enhanced flow with tailored domain-specific prompts, enabling the automated and self-correcting design of analog circuits with a high success rate. Secondly, it proposes a circuit tool library to archive successful designs as reusable modular sub-circuits, simplifying composite circuit creation. Thirdly, extensive experiments on a benchmark designed to cover a wide range of analog circuit tasks show that AnalogCoder outperforms other LLM-based methods. It has successfully designed 20 circuits, 5 more than standard GPT-4o. We believe AnalogCoder can significantly improve the labor-intensive chip design process, enabling non-experts to design analog circuits efficiently. Yao Lai, Sungyoung Lee 0004, Guojin Chen, Souradip Poddar, Mengkang Hu, David Z. Pan, Ping Luo 0002 |
AAAI | 5 |
| 2025 | Cross-Lingual Text-Rich Visual Comprehension: An Information Theory PerspectiveabstractRecent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. However, the abundant text within such images may increase the model's sensitivity to language. This raises the need to evaluate LVLM performance on cross-lingual text-rich visual inputs, where the language in the image differs from the language of the instructions. To address this, we introduce XT-VQA (Cross-Lingual Text-Rich Visual Question Answering), a benchmark designed to assess how LVLMs handle language inconsistency between image text and questions. XT-VQA integrates five existing text-rich VQA datasets and a newly collected dataset, XPaperQA, covering diverse scenarios that require faithful recognition and comprehension of visual information despite language inconsistency. Our evaluation of prominent LVLMs on XT-VQA reveals a significant drop in performance for cross-lingual scenarios, even for models with multilingual capabilities. A mutual information analysis suggests that this performance gap stems from cross-lingual questions failing to adequately activate relevant visual information. To mitigate this issue, we propose MVCL-MI (Maximization of Vision-Language Cross-Lingual Mutual Information), where a visual-text cross-lingual alignment is built by maximizing mutual information between the model's outputs and visual information. This is achieved by distilling knowledge from monolingual to cross-lingual settings through KL divergence minimization, where monolingual output logits serve as a teacher. Experimental results on the XT-VQA demonstrate that MVCL-MI effectively reduces the visual-text cross-lingual performance disparity while preserving the inherent capabilities of LVLMs, shedding new light on the potential practice for improving LVLMs. Xinmiao Yu, Minghui Liao, Ya-Qi Yu, Xiachong Feng, Weihong Zhong, Ruihan Chen 0001, Mengkang Hu, Jihao Wu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
AAAI | 9 |
| 2025 | HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language ModelabstractLarge Language Model (LLM)-based agents exhibit significant potential across various domains, operating as interactive systems that process environmental observations to generate executable actions for target tasks.The effectiveness of these agents is significantly influenced by their memory mechanism, which records historical experiences as sequences of actionobservation pairs.We categorize memory into two types: cross-trial memory, accumulated across multiple attempts, and in-trial memory (working memory), accumulated within a single attempt.While considerable research has optimized performance through cross-trial memory, the enhancement of agent performance through improved working memory utilization remains underexplored.Instead, existing approaches often involve directly inputting entire historical action-observation pairs into LLMs, leading to redundancy in long-horizon tasks.Inspired by human problem-solving strategies, this paper introduces HIAGENT, a framework that leverages subgoals as memory chunks to manage the working memory of LLM-based agents hierarchically.Specifically, HIAGENT prompts LLMs to formulate subgoals before generating executable actions and enables LLMs to decide proactively to replace previous subgoals with summarized observations, retaining only the action-observation pairs relevant to the current subgoal.Experimental results across five long-horizon tasks demonstrate that HIAGENT achieves a twofold increase in success rate and reduces the average number of steps required by 3.8.Additionally, our analysis shows that HIAGENT consistently improves performance across various steps, highlighting its robustness and generalizability. Mengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu 0001, Wenqi Shao, Ping Luo 0002 |
ACL (1) | 1 |
| 2025 | EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM AgentsabstractHeterogeneous multi-robot systems (HMRS) have emerged as a powerful ap-
proach for tackling complex tasks that single robots cannot manage alone. Current
large-language-model-based multi-agent systems (LLM-based MAS) have shown
success in areas like software development and operating systems, but applying
these systems to robot control presents unique challenges. In particular, the ca-
pabilities of each agent in a multi-robot system are inherently tied to the physical
composition of the robots, rather than predefined roles. To address this issue,
we introduce a novel multi-agent framework designed to enable effective collab-
oration among heterogeneous robots with varying embodiments and capabilities,
along with a new benchmark named Habitat-MAS. One of our key designs is
Robot Resume: Instead of adopting human-designed role play, we propose a self-
prompted approach, where agents comprehend robot URDF files and call robot
kinematics tools to generate descriptions of their physics capabilities to guide
their behavior in task planning and action execution. The Habitat-MAS bench-
mark is designed to assess how a multi-agent framework handles tasks that require
embodiment-aware reasoning, which includes 1) manipulation, 2) perception, 3)
navigation, and 4) comprehensive multi-floor object rearrangement. The experi-
mental results indicate that the robot’s resume and the hierarchical design of our
multi-agent system are essential for the effective operation of the heterogeneous
multi-robot system within this intricate problem context. Checheng Yu, Xunzhe Zhou, Yao Mu 0001, Mengkang Hu, Wenqi Shao, Guohao Li 0013, Lin Shao 0002 |
ICLR | 6 |
| 2025 | AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
Mengkang Hu, Pu Zhao 0004, Can Xu 0002, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo 0002, Saravan Rajmohan |
KDD (1) | 1 |
| 2025 | OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data SynthesisabstractThe rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks.
However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level decision making with low-level robot control based on both global scene understanding and current agent state. To address this complexity, we propose a novel multi-modal agent architecture that maintains multi-view scene frames and agent states for decision-making and controls the robot by function calling.
A second challenge is the hallucination from domain shift. To enhance the agent performance, we further introduce an agentic data synthesis pipeline for the OWMM task to adapt the VLM model to our task domain with instruction fine-tuning. We highlight our fine-tuned OWMM-VLM as the first dedicated foundation model for mobile manipulators with global scene understanding, robot state tracking, and multi-modal action generation in a unified model. Through experiments, we demonstrate that our model achieves SOTA performance compared to other foundation models including GPT-4o and strong zero-shot generalization in real world.
The project page is at https://hhyhrhy.github.io/owmm-agent-project. Haotian Liang, Lingxiao Du, Weiyun Wang, Mengkang Hu, Yao Mu 0001, Wenhai Wang, Jifeng Dai, Ping Luo 0002, Wenqi Shao, Lin Shao 0002 |
NeurIPS | 5 |
| 2025 | OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationabstractLarge Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature.
Current approaches face two critical shortcomings: they require complete architectural redesign and full retraining of all components when applied to new domains.
We introduce **Workforce**, a hierarchical multi-agent framework that decouples strategic planning from specialized execution through a modular architecture comprising:
*(i)* a *domain-agnostic* **Planner** for task decomposition,
*(ii)* a **Coordinator** for subtask management, and
*(iii)* specialized **Workers** with *domain-specific* tool-calling capabilities.
This decoupling enables cross-domain transferability during both inference and training phases:
During inference, Workforce seamlessly adapts to new domains by adding or modifying worker agents;
For training, we introduce **Optimized Workforce Learning (OWL)**, which improves generalization across domains by optimizing a domain-agnostic planner with reinforcement learning from real-world feedback.
To validate our approach, we evaluate Workforce on the GAIA benchmark, covering various realistic, multi-domain agentic tasks.
Experimental results demonstrate Workforce achieves open-source state-of-the-art performance (**69.70%**), outperforming commercial systems like OpenAI's Deep Research by **2.34%**.
More notably, our OWL-trained 32B model achieves **52.73%** accuracy (**+16.37%**) and demonstrates performance comparable to GPT-4o on challenging tasks.
To summarize, by enabling scalable generalization and modular domain transfer, our work establishes a foundation for the next generation of general-purpose AI assistants.
*Our code is available at [Anonymous URL](https://anonymous.4open.science/r/annonymous-owl/), and our data is available at [Anonymous URL](https://huggingface.co/anonymous21016).* Mengkang Hu, Wendong Fan, Yuzhou Nie, Ziyu Ye, Bowei Xia, Zhaoxuan Jin, Yingru Li, Qianshuo Ye, Bernard Ghanem, Ping Luo 0002, Guohao Li 0001 |
NeurIPS | 1 |
| 2024 | KET-QA: A Dataset for Knowledge Enhanced Table Question AnsweringabstractDue to the concise and structured nature of tables, the knowledge contained therein may be incomplete or missing, posing a significant challenge for table question answering (TableQA) systems. However, most existing datasets either overlook the challenge of missing knowledge in TableQA or only utilize unstructured text as supplementary information for tables. In this paper, we propose to use a knowledge base (KB) as the external knowledge source for TableQA and construct a dataset KET-QA with fine-grained gold evidence annotation. Each table in the dataset corresponds to a sub-graph of the entire KB, and every question requires the integration of information from both the table and the sub-graph to be answered. To extract pertinent information from the vast knowledge sub-graph and apply it to TableQA, we design a retriever-reasoner structured pipeline model. Experimental results demonstrate that our model consistently achieves remarkable relative performance improvements ranging from 1.9 to 6.5 times on EM scores across three distinct settings (fine-tuning, zero-shot, and few-shot), in comparison with solely relying on table information. However, even the best model achieves a 60.23% EM score, which still lags behind the human-level performance, highlighting the challenging nature of KET-QA for the question-answering community. Mengkang Hu, Haoyu Dong 0001, Ping Luo 0002, Shi Han, Dongmei Zhang 0001 |
LREC/COLING | 1 |
| 2024 | OpenTE: Open-Structure Table Extraction From TextabstractThis paper presents an Open-Structure Table Extraction (OpenTE) task, which aims to extract a table with intrinsic semantic, calculational, and hierarchical structure from unstructured text. We devise a novel Identification-Extraction-Grounding (IEG) framework for language models (LMs) comprising three chaining steps: (1) identifying semantic and calculational relationships among columns, (2) extracting structured data from unstructured text, and (3) aligning extracted data with the source text and the table structure with a separate discrete grounding model. Experiment results suggest that OpenTE presents a significant challenge for state-of-the-art LMs and demonstrate that the IEG framework achieves superior performance on both datasets, with over 9% F1 improvements in the few-shot setting for GPT-3.5&4 and other large language models (LLMs) and over 4.9% F1 enhancements in the fine-tuning setting for open-source BART. We’ll release the dataset to facilitate future research. Haoyu Dong 0001, Mengkang Hu, Qinyu Xu, Yue Hu 0002 |
ICASSP | 2 |
| 2024 | Tree-Planner: Efficient Close-loop Task Planning with Large Language ModelsabstractThis paper studies close-loop task planning, which refers to the process of generating a sequence of skills (a plan) to accomplish a specific goal while adapting the plan based on real-time observations.
Recently, prompting Large Language Models (LLMs) to generate actions iteratively has become a prevalent paradigm due to its superior performance and user-friendliness.
However, this paradigm is plagued by two inefficiencies: high token consumption and redundant error correction, both of which hinder its scalability for large-scale testing and applications.
To address these issues, we propose Tree-Planner, which reframes task planning with LLMs into three distinct phases:
plan sampling, action tree construction, and grounded deciding.
Tree-Planner starts by using an LLM to sample a set of potential plans before execution, followed by the aggregation of them to form an action tree.
Finally, the LLM performs a top-down decision-making process on the tree, taking into account real-time environmental information.
Experiments show that Tree-Planner achieves state-of-the-art performance while maintaining high efficiency.
By decomposing LLM queries into a single plan-sampling call and multiple grounded-deciding calls,
a considerable part
of the prompt are less likely to be repeatedly consumed.
As a result, token consumption is reduced by 92.2\% compared to the previously best-performing model.
Additionally, by enabling backtracking on the action tree as needed, the correction process becomes more flexible, leading to a 40.5\% decrease in error corrections. Mengkang Hu, Yao Mu 0001, Xinmiao Yu, Mingyu Ding, Shiguang Wu 0004, Wenqi Shao, Qiguang Chen, Bin Wang 0034, Yu Qiao 0001, Ping Luo 0002 |
ICLR | 1 |
| 2024 | RoboCodeX: Multimodal Code Generation for Robotic Behavior SynthesisabstractRobotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these conceptual understandings into detailed robotic actions while achieving generalization across various scenarios. In this paper, we propose a tree-structured multimodal code generation framework for generalized robotic behavior synthesis, termed RoboCodeX. RoboCodeX decomposes high-level human instructions into multiple object-centric manipulation units consisting of physical preferences such as affordance and safety constraints, and applies code generation to introduce generalization ability across various robotics platforms. To further enhance the capability to map conceptual and perceptual understanding into control commands, a specialized multimodal reasoning dataset is collected for pre-training and an iterative self-updating methodology is introduced for supervised fine-tuning. Extensive experiments demonstrate that RoboCodeX achieves state-of-the-art performance in both simulators and real robots on four different kinds of manipulation tasks and one embodied navigation task. Yao Mu 0001, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang 0026, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao 0001, Mingyu Ding, Ping Luo 0002 |
ICML | 9 |
| 2024 | Needle In A Multimodal HaystackabstractWith the rapid advancement of multimodal large language models (MLLMs), their evaluation has become increasingly comprehensive. However, understanding long multimodal content, as a foundational ability for real-world applications, remains underexplored. In this work, we present Needle In A Multimodal Haystack (MM-NIAH), the first benchmark specifically designed to systematically evaluate the capability of existing MLLMs to comprehend long multimodal documents. Our benchmark includes three types of evaluation tasks: multimodal retrieval, counting, and reasoning. In each task, the model is required to answer the questions according to different key information scattered throughout the given multimodal document. Evaluating the leading MLLMs on MM-NIAH, we observe that existing models still have significant room for improvement on these tasks, especially on vision-centric evaluation. We hope this work can provide a platform for further research on long multimodal document comprehension and contribute to the advancement of MLLMs. Code and benchmark are released at https://github.com/OpenGVLab/MM-NIAH. Weiyun Wang, Shuibo Zhang, Yiming Ren 0001, Yuchen Duan, Tiantong Li, Mengkang Hu, Zhe Chen 0017, Kaipeng Zhang, Lewei Lu, Xizhou Zhu, Ping Luo 0002, Yu Qiao 0001, Jifeng Dai, Wenqi Shao, Wenhai Wang |
NeurIPS | 7 |
| 2023 | EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtabstractEmbodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments.
In this work, we introduce EmbodiedGPT, an end-to-end multi-modal foundation model for embodied AI, empowering embodied agents with multi-modal understanding and execution capabilities. To achieve this, we have made the following efforts: (i) We craft a large-scale embodied planning dataset, termed EgoCOT. The dataset consists of carefully selected videos from the Ego4D dataset, along with corresponding high-quality language instructions. Specifically, we generate a sequence of sub-goals with the "Chain of Thoughts" mode for effective embodied planning.
(ii) We introduce an efficient training approach to EmbodiedGPT for high-quality plan generation, by adapting a 7B large language model (LLM) to the EgoCOT dataset via prefix tuning. (iii) We introduce a paradigm for extracting task-related features from LLM-generated planning queries to form a closed loop between high-level planning and low-level control.
Extensive experiments show the effectiveness of EmbodiedGPT on embodied tasks, including embodied planning, embodied control, visual captioning, and visual question answering.
Notably, EmbodiedGPT significantly enhances the success rate of the embodied control task by extracting more effective features. It has achieved a remarkable 1.6 times increase in success rate on the Franka Kitchen benchmark and a 1.3 times increase on the Meta-World benchmark, compared to the BLIP-2 baseline fine-tuned with the Ego4D dataset. Yao Mu 0001, Mengkang Hu, Wenhai Wang, Mingyu Ding, Bin Wang 0034, Jifeng Dai, Yu Qiao 0001, Ping Luo 0002 |
NeurIPS | 3 |
| 2022 | TaCube: Pre-computing Data Cubes for Answering Numerical-Reasoning Questions over Tabular DataabstractExisting auto-regressive pre-trained language models (PLMs) like T5 and BART, have been well applied to table question answering by UNIFIEDSKG and TAPEX, respectively, and demonstrated state-of-the-art results on multiple benchmarks.However, auto-regressive PLMs are challenged by recent emerging numerical reasoning datasets, such as TAT-QA, due to the error-prone implicit calculation.In this paper, we present TACUBE, to precompute aggregation/arithmetic results for the table in advance, so that they are handy and readily available for PLMs to answer numerical reasoning questions.TACUBE systematically and comprehensively covers a collection of computational operations over table segments.By simply concatenating TACUBE to the input sequence of PLMs, it shows significant experimental effectiveness.TACUBE promotes the F1 score from 49.6% to 66.2% on TAT-QA and achieves new state-of-the-art results on WikiTQ (59.6% denotation accuracy).TACUBE 's improvements on numerical reasoning cases are even more notable: on TAT-QA, TACUBE promotes the exact match accuracy of BART-large by 39.6% on sum, 52.5% on average, 36.6% on subtraction and 22.2% on division.We believe that TACUBE is a general and portable pre-computation solution that can be potentially integrated into various numerical reasoning frameworks.Data and code will be available at https://github.com/ microsoft/TaCube. Mengkang Hu, Haoyu Dong 0001, Zhoujun Cheng, Fan Cheng 0002, Shi Han, Dongmei Zhang 0001 |
EMNLP | 2 |