EDBT 2026 Demo / reviewers in the wild / expert
Mingchen Zhuge
dblp:283/5310
· DBLP profile ↗
18ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0003-2561-7712ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language ModelsabstractLong-form writing agents require flexible integration and interaction across information retrieval, reasoning, and composition.Current approaches rely on predefined workflows and rigid thinking patterns to generate outlines before writing, resulting in constrained adaptability during writing.In this paper we propose WriteHERE, a general agent framework that achieves human-like adaptive writing through recursive task decomposition and dynamic integration of three fundamental task types: retrieval, reasoning, and composition.Our methodology features: 1) a planning mechanism that interleaves recursive task decomposition and execution, eliminating artificial restrictions on writing workflow; and 2) integration of task types that facilitates heterogeneous task decomposition.Evaluations on both fiction writing and technical report generation show that our method consistently outperforms state-of-the-art approaches across all automatic evaluation metrics, demonstrating the effectiveness and broad applicability of our proposed framework.We have publicly released our code and prompts to facilitate further research. Ruibin Xiong, Dmitrii Khizbullin, Mingchen Zhuge, Jürgen Schmidhuber |
EMNLP | 4 |
| 2025 | OpenHands: An Open Platform for AI Software Developers as Generalist AgentsabstractSoftware is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and effect change in their surrounding environments. In this paper, we introduce OpenHands, a platform for the development of powerful and flexible AI agents that interact with the world in similar ways to a human developer: by writing code, interacting with a command line, and browsing the web. We describe how the platform allows for the implementation of new agents, utilization of various LLMs, safe interaction with sandboxed environments for code execution, and incorporation of evaluation benchmarks. Based on our currently incorporated benchmarks, we perform an evaluation of agents over 13 challenging tasks, including software engineering (e.g., SWE-Bench) and web browsing (e.g., WebArena), amongst others. Released under the permissive MIT license, OpenHands is a community project spanning academia and industry with more than 2K contributions from over 186 contributors in less than six months of development, and will improve going forward. Xingyao Wang 0002, Boxuan Li, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Yueqi Song, Bowen Li 0002, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang 0002, Binyuan Hui, Junyang Lin |
ICLR | 6 |
| 2025 | AFlow: Automating Agentic Workflow GenerationabstractLarge language models (LLMs) have demonstrated remarkable potential in solving complex tasks across diverse domains, typically by employing agentic workflows that follow detailed instructions and operational sequences. However, constructing these workflows requires significant human effort, limiting scalability and generalizability. Recent research has sought to automate the generation and optimization of these workflows, but existing methods still rely on initial manual setup and fall short of achieving fully automated and effective workflow generation. To address this challenge, we reformulate workflow optimization as a search problem over code-represented workflows, where LLM-invoking nodes are connected by edges. We introduce AFLOW, an automated framework that efficiently explores this space using Monte Carlo Tree Search, iteratively refining workflows through code modification, tree-structured experience, and execution feedback. Empirical evaluations across six benchmark datasets demonstrate AFLOW's efficacy, yielding a 5.7% average improvement over state-of-the-art baselines. Furthermore, AFLOW enables smaller models to outperform GPT-4o on specific tasks at 4.55% of its inference cost in dollars. The code is available at https://github.com/FoundationAgents/AFlow. Jiayi Zhang 0017, Jinyu Xiang, Zhaoyang Yu 0004, Fengwei Teng, Xionghui Chen, Mingchen Zhuge, Sirui Hong, Bingnan Zheng, Bang Liu 0003, Yuyu Luo, Chenglin Wu 0001 |
ICLR | 7 |
| 2025 | Agent-as-a-Judge: Evaluate Agents with AgentsabstractContemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes—ignoring the step-by-step nature of the thinking done by agentic systems—or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is a natural extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving processes for more precise evaluations. We apply the Agent-as-a-Judge framework to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic AI code generation tasks. DevAI includes rich manual annotations, like a total of 365 hierarchical solution requirements, which make it particularly suitable for an agentic evaluator. We benchmark three of the top code-generating agentic systems using Agent-as-a-Judge and find that our framework dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that this work represents a concrete step towards enabling vastly more sophisticated agentic systems. To help that, our dataset and the full implementation of Agent-as-a-Judge will be publically available at https://github.com/metauto-ai/agent-as-a-judge Mingchen Zhuge, Changsheng Zhao 0002, Dylan R. Ashley, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jürgen Schmidhuber |
ICML | 1 |
| 2025 | Mindstorms in Natural Language-Based Societies of MindabstractInspired by Minsky's Society of Mind, Schmidhuber's Learning to Think, and other more recent works, this paper proposes and advocates for the concept of natural language-based societies of mind (NLSOMs). We imagine these societies as consisting of a collection of multimodal neural networks, including large language models, which engage in a “mindstorm” to solve problems using a shared natural language interface. Here, we work to identify and discuss key questions about the social structure, governance, and economic principles for NLSOMs, emphasizing their impact on the future of AI. Our demonstrations with NLSOMs-which feature up to 129 agents-show their effectiveness in various tasks, including visual question answering, image captioning, and prompt generation for text-to-image synthesis. Mingchen Zhuge, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li 0024, Guohao Li 0001, Shuming Liu 0001, Jinjie Mai, Piotr Piekos, Aditya A. Ramesh, Imanol Schlag, Aleksandar Stanic, Yuhui Wang 0004, Mengmeng Xu 0006, Deng-Ping Fan, Bernard Ghanem, Jürgen Schmidhuber |
Comput. Vis. Media | 1 |
| 2024 | Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Mingchen Zhuge, Jian Ding 0001, Deyao Zhu, Jürgen Schmidhuber, Mohamed Elhoseiny 0001 |
ECCV (29) | 5 |
| 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkabstractRecently, remarkable progress has been made on automated problem solving through societies of agents based on large language models (LLMs). Previous LLM-based multi-agent systems can already solve simple dialogue tasks. More complex tasks, however, face challenges through logic inconsistencies due to cascading hallucinations caused by naively chaining LLMs. Here we introduce MetaGPT, an innovative meta-programming framework incorporating efficient human workflows into LLM-based multi-agent collaborations. MetaGPT encodes Standardized Operating Procedures (SOPs) into prompt sequences for more streamlined workflows, thus allowing agents with human-like domain expertise to verify intermediate results and reduce errors. MetaGPT utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together. On collaborative software engineering benchmarks, MetaGPT generates more coherent solutions than previous chat-based multi-agent systems. Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu 0001, Jürgen Schmidhuber |
ICLR | 2 |
| 2024 | GPTSwarm: Language Agents as Optimizable GraphsabstractVarious human-designed prompt engineering techniques have been proposed to improve problem solvers based on Large Language Models (LLMs), yielding many disparate code bases. We unify these approaches by describing LLM-based agents as computational graphs. The nodes implement functions to process multimodal data or query LLMs, and the edges describe the information flow between operations. Graphs can be recursively combined into larger composite graphs representing hierarchies of inter-agent collaboration (where edges connect operations of different agents). Our novel automatic graph optimizers (1) refine node-level LLM prompts (node optimization) and (2) improve agent orchestration by changing graph connectivity (edge optimization). Experiments demonstrate that our framework can be used to efficiently develop, integrate, and automatically improve various LLM agents. Our code is public. Mingchen Zhuge, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, Jürgen Schmidhuber |
ICML | 1 |
| 2024 | Multimodal Inplace Prompt Tuning for Open-set Object DetectionabstractThe integration of large language models into open-world detection frameworks significantly improves versatility in new environments. Prompt representations derived from these models help establish classification boundaries for both base and novel categories within open-world detectors. However, we are the first to discover that directly fine-tuning language models in detection systems results in redundant attention patterns and leads to suboptimal prompt representations. In order to fully leverage the capabilities of large language models and augment prompt encoding for detection, this study introduces a redundancy assessment metric to identify uniform attention patterns. Furthermore, in areas with high redundancy, we incorporate multimodal inplace prompt tuning (MIPT) to enrich the text prompt with visual clues. Experimental results validate the efficacy of our MIPT framework, achieving a notable increase across benchmarks, e.g. elevating GLIP-L from 22.6% to 25.0% on ODinW-35, and 9.0% improvement on LVIS. Mengdan Zhang, Xiawu Zheng, Peixian Chen, Yunhang Shen, Mingchen Zhuge, Chenglin Wu 0001, Fei Chao 0001, Ke Li 0015, Xing Sun 0001, Rongrong Ji |
ACM Multimedia | 7 |
| 2023 | Skating-Mixer: Long-Term Sport Audio-Visual Modeling with MLPsabstractFigure skating scoring is challenging because it requires judging players’ technical moves as well as coordination with the background music. Most learning-based methods struggle for two reasons: 1) each move in figure skating changes quickly, hence simply applying traditional frame sampling will lose a lot of valuable information, especially in 3 to 5 minutes lasting videos; 2) prior methods rarely considered the critical audio-visual relationship in their models. Due to these reasons, we introduce a novel architecture, named Skating-Mixer. It extends the MLP framework into a multimodal fashion and effectively learns long-term representations through our designed memory recurrent unit (MRU). Aside from the model, we collected a high-quality audio-visual FS1000 dataset, which contains over 1000 videos on 8 types of programs with 7 different rating metrics, overtaking other datasets in both quantity and diversity. Experiments show the proposed method achieves SOTAs over all major metrics on the public Fis-V and our FS1000 dataset. In addition, we include an analysis applying our method to the recent competitions in Beijing 2022 Winter Olympic Games, proving our method has strong applicability. Jingfei Xia, Mingchen Zhuge, Tiantian Geng, Shun Fan, Yuantai Wei, Zhenyu He 0001, Feng Zheng 0001 |
AAAI | 2 |
| 2023 | NewsNet: A Novel Dataset for Hierarchical Temporal SegmentationabstractTemporal video segmentation is the get-to- go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations can help compare the semantics in the corresponding scales, but lack a wider view of larger temporal spans, especially when the video is complex and structured. Therefore, we present two abstractive levels of temporal segmentations and study their hierarchy to the existing fine-grained levels. Accordingly, we collect NewsNet, the largest news video dataset consisting of 1,000 videos in over 900 hours, associated with several tasks for hierarchical temporal video segmentation. Each news video is a collection of stories on different topics, represented as aligned audio, visual, and textual data, along with extensive frame-wise annotations in four granularities. We assert that the study on NewsNet can advance the understanding of complex structured video and benefit more areas such as short-video creation, personalized advertisement, digital instruction, and education. Our dataset and code is publicly available at https://github.com/NewsNet-Benchmark/NewsNet. Haoqian Wu, Mingchen Zhuge, Bing Li 0024, Ruizhi Qiao, Xiujun Shu, Bei Gan, Liangsheng Xu, Bo Ren 0002, Mengmeng Xu 0006, Wentian Zhang, Ramachandra Raghavendra, Chia-Wen Lin, Bernard Ghanem |
CVPR | 4 |
| 2023 | Learning to Identify Critical States for Reinforcement Learning from VideosabstractRecent work on deep reinforcement learning (DRL) has pointed out that algorithmic information about good policies can be extracted from offline data which lack explicit information about executed actions [45], [46], [30]. For example, videos of humans or robots may convey a lot of implicit information about rewarding action sequences, but a DRL machine that wants to profit from watching such videos must first learn by itself to identify and recognize relevant states/actions/rewards. Without relying on ground-truth annotations, our new method called Deep State Identifier learns to predict returns from episodes encoded as videos. Then it uses a kind of mask-based sensitivity analysis to extract/identify important critical states. Extensive experiments showcase our method’s potential for understanding and improving agent behavior. The source code and the generated datasets are available at https://github.com/AI-Initiative-KAUST/VideoRLCS. Mingchen Zhuge, Bing Li 0024, Yuhui Wang 0004, Francesco Faccio, Bernard Ghanem, Jürgen Schmidhuber |
ICCV | 2 |
| 2023 | Salient Object Detection via Integrity LearningabstractAlthough current salient object detection (SOD) works have achieved significant progress, they are limited when it comes to the integrity of the predicted salient regions. We define the concept of integrity at both a micro and macro level. Specifically, at the micro level, the model should highlight all parts that belong to a certain salient object. Meanwhile, at the macro level, the model needs to discover all salient objects in a given image. To facilitate integrity learning for SOD, we design a novel Integrity Cognition Network (ICON), which explores three important components for learning strong integrity features. 1) Unlike existing models, which focus more on feature discriminability, we introduce a diverse feature aggregation (DFA) component to aggregate features with various receptive fields (i.e., kernel shape and context) and increase feature diversity. Such diversity is the foundation for mining the integral salient objects. 2) Based on the DFA features, we introduce an integrity channel enhancement (ICE) component with the goal of enhancing feature channels that highlight the integral salient objects, while suppressing the other distracting ones. 3) After extracting the enhanced features, the part-whole verification (PWV) method is employed to determine whether the part and whole object features have strong agreement. Such part-whole agreements can further improve the micro-level integrity for each salient object. To demonstrate the effectiveness of our ICON, comprehensive experiments are conducted on seven challenging benchmarks. Our ICON outperforms the baseline methods in terms of a wide range of metrics. Notably, our ICON achieves ∼ 10% relative improvement over the previous best model in terms of average false negative ratio (FNR), on six datasets. Codes and results are available at: https://github.com/mczhuge/ICON. Mingchen Zhuge, Deng-Ping Fan, Nian Liu 0002, Dingwen Zhang, Dong Xu 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
Ge-Peng Ji, Lei Zhu 0012, Mingchen Zhuge, Keren Fu |
Pattern Recognit. | 3 |
| 2022 | CubeNet: X-shape connection for camouflaged object detection
Mingchen Zhuge, Xiankai Lu, Yiyou Guo, Zhihua Cai, Shuhan Chen |
Pattern Recognit. | 1 |
| 2021 | Kaleido-BERT: Vision-Language Pre-Training on Fashion DomainabstractWe present a new vision-language (VL) pre-training model dubbed Kaleido-BERT , which introduces a novel kaleido strategy for fashion cross-modality representations from transformers. In contrast to random masking strategy of recent VL models, we design alignment guided masking to jointly focus more on image-text semantic relations. To this end, we carry out five novel tasks, i.e., rotation, jigsaw, camouflage, grey-to-color, and blank-to-color for self-supervised VL pre-training at patches of different scale. Kaleido-BERT is conceptually simple and easy to extend to the existing BERT framework, it attains state-of-the-art results by large margins on four downstream tasks, including text retrieval (R@1: 4.03% absolute improvement), image retrieval (R@1: 7.13% abs imv.), category recognition (ACC: 3.28% abs imv.), and fashion captioning (Bleu4: 1.2 abs imv.). We validate the efficiency of Kaleido-BERT on a wide range of e-commerical websites, demonstrating its broader potential in real-world applications. Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Haoming Zhou, Minghui Qiu, Ling Shao 0001 |
CVPR | 1 |
| 2021 | Object Decoupling with Graph Correlation for Fine-Grained Image ClassificationabstractFine-grained image classification has drawn increasing attention as it is much closer to practical applications than generic image classification. The majority of current fine-grained approaches locate the discriminative regions and leverage the features of these regions for classification as their magic weapons. However, these approaches simply ignore the internal semantic region correlation. As is well known, the correlation reveals the salient information of images, which can further boost the performance of fine-grained image classification. To this end, we propose an Object Decoupling with Graph Correlation network (ODGC) to explore the informative potentials of region correlation. A Responsive Object Location Module (ROLM) is first introduced to obtain the fine-grained object within a bounding box automatically. A Semantic Decoupling Module (SDM) then segments the object into different parts. ODGC learns the representations of these parts by transferring these part features into a Graph Correlation Module (GCM). Consists of these three main modules, ODGC is trained for fine-grained image classification in an end-to-end way. Extensive experiments conducted on CUB-200-2011 demonstrate that the aforementioned modules significantly improve the ODGC, and it achieves a new state-of-the-art performance to 88.2% top-1 accuracy. Besides, we collect a practical business e-commercial dataset, named Ecom-15K. The evaluation on it further validates the applicability of our method in practical scenarios. Qiushi Guo, Mingchen Zhuge, Dehong Gao, Huiling Zhou, Xiaonan Meng |
ICME | 2 |
| 2021 | Cooperative Spectral-Spatial Attention Dense Network for Hyperspectral Image ClassificationabstractRecently, deep learning-based methods have made great progress in hyperspectral image (HSI) classification (HSIC). Different from ordinary images, the intrinsic complexity of HSIs data still limits the performance of many common convolutional neural network (CNN) models. Thus, the network architecture becomes more and more complex to extract discriminative spectral-spatial features. For instance, 3-D CNN usually has a large number of trainable parameters, thus increasing the computational complexity of the HSIC. In this letter, we designed a cooperative spectral-spatial attention dense network (CS2ADN) that takes raw 3-D HSI data as input data. Specifically, the attention module consists of spectral and spatial axes, by which the salient spectral-spatial features will be emphasized. Furthermore, we combined these attention modules with the dense connection, which is termed as the lightweight dense block; it has a lower computation cost and achieves better classification performance. At the same time, we introduced the center loss, by jointly using the supervision of the center loss and the softmax loss, where the discriminative features could be clearly observed, particularly for small data sets. Experimental results on the biased and unbiased HSI data show that our method outperforms several state-of-the-art methods in HSIC with small training samples. Zhimin Dong, Yaoming Cai, Zhihua Cai, Xiaobo Liu 0001, Zhaoyu Yang, Mingchen Zhuge |
IEEE Geosci. Remote. Sens. Lett. | 6 |