EDBT 2026 Demo / reviewers in the wild / expert
Muzhi Zhu
dblp:157/1679
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Denoising Pre-training Framework for Accelerating Novel Material DiscoveryabstractCrystal materials play an important role in the development of society. The discovery of new materials is critical to achieving sustainable development goals (SDGs), such as climate change mitigation, affordable and clean energy, and fostering innovation in industry and infrastructure. Recent advances in deep learning for crystal property prediction have accelerated material discovery, but these methods typically rely on labeled data, which is often limited and varies across different properties. This limitation hinders the full utilization of the vast amount of unlabeled data in materials science. To overcome this challenge, we introduce an unsupervised Denoising Pre-training Framework (DPF) tailored for crystal structures. DPF trains a model to reconstruct the original crystal structure by recovering the masked atom types, perturbed atom positions, and perturbed crystal lattices. Through pre-training, models learn the intrinsic features of crystal structures and capture the key features influencing crystal properties. We pre-train models on a dataset of 380,743 unlabeled crystal structures and fine-tune them on downstream property prediction tasks. Extensive experiments demonstrate the effectiveness of our framework, showing its potential to significantly advance material science and contribute to the development of society by accelerating the discovery of materials crucial for sustainable technologies. Shuaike Shen, Ke Liu 0012, Muzhi Zhu, Hao Chen 0041 |
AAAI | 3 |
| 2025 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesabstractWhile MLLMs have demonstrated impressive image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks such as VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Although segmentation is foundational for pixel-level understanding, existing methods often require MLLMs to generate implicit tokens, decoded through external pixel decoders. This approach disrupts the MLLM’s text output space, potentially compromising language capabilities and reducing flexibility and extensibility while failing to reflect the model’s intrinsic pixel-level understanding. Thus, we introduce the Human-Like Mask Annotation Task (HLMAT), a new paradigm where MLLMs mimic human annotators using interactive segmentation tools. Modelling segmentation as a multi-step Markov Decision Process, HLMAT enables MLLMs to iteratively generate text-based click points, achieving high-quality masks without architectural changes or implicit tokens. Through this setup, we develop SegAgent, a model fine-tuned on human-like annotation trajectories, which achieves performance comparable to SoTA methods and supports additional tasks like mask refinement and annotation filtering. HLMAT provides a protocol for assessing fine-grained pixel understanding in MLLMs and introduces a vision-centric, multi-step decision-making task that facilitates the exploration of MLLMs’ visual reasoning abilities. Our adaptations of policy improvement method StaR and PRM guided tree search further enhance model robustness in complex segmentation tasks, laying a foundation for future advancements in fine-grained visual perception and multi-step decision-making for MLLMs. Code can be found at https://github.com/aim-uofa/SegAgent. Muzhi Zhu, Yuzhuo Tian, Hao Chen 0041, Chunluan Zhou, Qingpei Guo, Yang Liu 0357, Ming Yang 0007, Chunhua Shen |
CVPR | 1 |
| 2025 | Unified Open-World Segmentation with Multi-Modal Prompts
Yang Liu 0357, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen 0041, Yuling Xi, Hao Wang 0052, Chunhua Shen |
ICCV | 4 |
| 2025 | DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksabstractThis paper's primary objective is to develop a robust generalist perception model capable of addressing multiple tasks under constraints of computational resources and limited training data. We leverage text-to-image diffusion models pre-trained on billions of images and successfully introduce our DICEPTION, a visual generalist model. Exhaustive evaluations demonstrate that DICEPTION effectively tackles diverse perception tasks, even achieving performance comparable to SOTA single-task specialist models. Specifically, we achieve results on par with SAM-vit-h using only 0.06% of their data (e.g., 600K vs.\ 1B pixel-level annotated images). We designed comprehensive experiments on architectures and input paradigms, demonstrating that the key to successfully re-purposing a single diffusion model for multiple perception tasks lies in maximizing the preservation of the pre-trained model's prior knowledge. Consequently, DICEPTION can be trained with substantially lower computational costs than conventional models requiring training from scratch. Furthermore, adapting DICEPTION to novel tasks is highly efficient, necessitating fine-tuning on as few as 50 images and approximately 1% of its parameters. Finally, we demonstrate that a subtle application of classifier-free guidance can improve the model's performance on depth and normal estimation. We also show that pixel-aligned training, as is characteristic of perception tasks, significantly enhances the model's ability to preserve fine details. DICEPTION offers valuable insights and presents a promising direction for the development of advanced diffusion-based visual generalist models. Canyu Zhao, Yanlong Sun, Huanyi Zheng, Muzhi Zhu, Zhiyue Zhao, Hao Chen 0041, Tong He 0001, Chunhua Shen |
NeurIPS | 5 |
| 2025 | Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System CollaborationabstractLong-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution inputs. We tackle this trade-off with a two-system architecture: a Global Reasoning System selects informative keyframes and rewrites the task at low spatial cost, while a Detail Understanding System performs pixel-level grounding on the selected high-resolution snippets.
Because "optimal" keyframe selection and reformulation are ambiguous and hard to supervise, we formulate them as a reinforcement-learning (RL) problem and present Omni-R1, an end-to-end RL framework built on Group Relative Policy Optimization.
Omni-R1 trains the Global Reasoning System through hierarchical rewards obtained via online collaboration with the Detail Understanding System, requiring only one epoch of RL on small task splits.
Experiments on two challenging benchmarks, Referring Audio-Visual Segmentation (RefAVS) and Reasoning Video Object Segmentation (REVOS), show that Omni-R1 not only surpasses strong supervised baselines but also outperforms specialized state-of-the-art models, while substantially improving out-of-domain generalization and mitigating multimodal hallucination.
Our results demonstrate the first successful application of RL to large-scale omnimodal reasoning and highlight a scalable path toward universally foundation models. Muzhi Zhu, Zongze Du, Canyu Zhao, Wen Wang 0015, Hao Chen 0041, Chunhua Shen |
NeurIPS | 2 |
| 2025 | Masked Channel Modeling for Bootstrapping Visual Pre-training
Yang Liu 0357, Muzhi Zhu, Yue Cao 0001, Tiejun Huang 0001, Chunhua Shen |
Int. J. Comput. Vis. | 3 |
| 2025 | Segment Anything in Context with Vision Foundation Models
Yang Liu 0357, Muzhi Zhu, Hao Chen 0041, Hao Wang 0052, Raviteja Vemulapalli, Chunhua Shen |
Int. J. Comput. Vis. | 2 |
| 2024 | DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative DataabstractInstance segmentation is data-hungry, and as model capacity increases, data scale becomes crucial for improving the accuracy. Most instance segmentation datasets today require costly manual annotation, limiting their data scale. Models trained on such data are prone to overfitting on the training set, especially for those rare categories. While recent works have delved into exploiting generative models to create synthetic datasets for data augmentation, these approaches do not efficiently harness the full potential of generative models. To address these issues, we introduce a more efficient strategy to construct generative datasets for data augmentation, termed DiverGen. Firstly, we provide an explanation of the role of generative data from the perspective of distribution discrepancy. We investigate the impact of different data on the distribution learned by the model. We argue that generative data can expand the data distribution that the model can learn, thus mitigating overfitting. Additionally, we find that the diversity of generative data is crucial for improving model performance and enhance it through various strategies, including category diversity, prompt diversity, and generative model diversity. With these strategies, we can scale the data to millions while maintaining the trend of model performance improvement. On the LVIS dataset, DiverGen significantly outperforms the strong model X-Paste, achieving +1.1 box AP and +1.1 mask AP across all categories, and +1.9 box AP and +2.5 mask AP for rare categories. Our codes are available at https://github.com/aim-uofa/DiverGen. Chengxiang Fan, Muzhi Zhu, Hao Chen 0041, Yang Liu 0357, Weijia Wu 0001, Huaqi Zhang, Chunhua Shen |
CVPR | 2 |
| 2024 | Matcher: Segment Anything with One Shot Using All-Purpose Feature MatchingabstractPowered by large-scale pre-training, vision foundation models exhibit significant potential in open-world image understanding. However, unlike large language models that excel at directly tackling various language tasks, vision foundation models require a task-specific model structure followed by fine-tuning on specific tasks. In this work, we present $\textbf{Matcher}$, a novel perception paradigm that utilizes off-the-shelf vision foundation models to address various perception tasks. Matcher can segment anything by using an in-context example without training. Additionally, we design three effective components within the Matcher framework to collaborate with these foundation models and unleash their full potential in diverse perception tasks. Matcher demonstrates impressive generalization performance across various segmentation tasks, all without training. For example, it achieves 52.7% mIoU on COCO-20$^i$ with one example, surpassing the state-of-the-art specialist model by 1.6%. In addition, Matcher achieves 33.0% mIoU on the proposed LVIS-92$^i$ for one-shot semantic segmentation, outperforming the state-of-the-art generalist model by 14.4%. Our visualization results further showcase the open-world generality and flexibility of Matcher when applied to images in the wild. Yang Liu 0357, Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chunhua Shen |
ICLR | 2 |
| 2024 | De novo Protein Design Using Geometric Vector Field NetworksabstractAdvances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Only a few basic encoders, like IPA, have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we introduce the Vector Field Network (VFN), that enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential $\textit{universal encoder}$. In protein diffusion (frame modeling), VFN exhibits a impressive performance advantage over IPA, excelling in terms of both designability ($\textbf{67.04}$\% vs. 53.58\%) and diversity ($\textbf{66.54}$\% vs. 51.98\%). In inverse folding(frame and atom modeling), VFN outperforms the previous SoTA model, PiFold ($\textbf{54.7}$\% vs. 51.66\%), on sequence recovery rate; we also propose a method of equipping VFN with the ESM model, which significantly surpasses the previous ESM-based SoTA ($\textbf{62.67}$\% vs. 55.65\%), LM-Design, by a substantial margin. Code is available at https://github.com/aim-uofa/VFN Weian Mao, Muzhi Zhu, Shuaike Shen, Lin Wu 0001, Hao Chen 0041, Chunhua Shen |
ICLR | 2 |
| 2024 | Generative Active Learning for Long-tailed Instance SegmentationabstractRecently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do not thoroughly explore how to better select and utilize generated data. On the other hand, there is still a lack of research oriented towards active learning on generated data. In this paper, we explore how to perform active learning specifically for generated data in the long-tailed instance segmentation task. Subsequently, we propose BSGAL, a new algorithm that estimates the contribution of the current batch-generated data based on gradient cache. BSGAL is meticulously designed to cater for unlimited generated data and complex downstream segmentation tasks. BSGAL outperforms the baseline approach and effectually improves the performance of long-tailed segmentation. Muzhi Zhu, Chengxiang Fan, Hao Chen 0041, Yang Liu 0357, Weian Mao, Xiaogang Xu 0002, Chunhua Shen |
ICML | 1 |
| 2024 | Unveiling Universal Forensics of Diffusion Models with Adversarial PerturbationsabstractWith state-of-the-art performance in image synthesis especially in text-to-image generation, diffusion models (DMs) have received unprecedented attention. Despite promising application prospects, fake images generated from diffusion models are causing potential security concerns. To this end, in this work, we aim to investigate whether the generated images of DMs are different from other generated models e.g. generative adversarial networks (GANs) and whether a universal forensic classifier exists. To perform this work, we first collected a dataset consisting of 409k fake images generated from different types of DMs. Through a comprehensive analysis on this benchmark, we showcased that common forensic artifacts are shared among DMs and a forensic classifier trained for one model can generalize well for other agnostic generative models. Specifically, we first demonstrated that despite photorealism of the images generated by DMs, they still contain artifacts namely non-robust visual features which are hard for human but easy for machine to recognize. Then we studied the characteristics of the artifacts from the view of adversarial attack and unexpectedly found there exists a universal adversarial perturbation to fool the classifier. Furthermore, we devised visualization and analysis tools focusing on the spectral properties of the generated samples and adversarial features which demonstrates augmentations in the frequency domain greatly affect the performance of the detectors. Kangyang Xie, Jiaan Liu, Muzhi Zhu, Ganggui Ding, Zide Liu, Hao Chen 0041, Hangyue Chen |
IJCNN | 3 |
| 2024 | A Simple Image Segmentation Framework via In-Context ExamplesabstractRecently, there have been explorations of generalist segmentation models that can effectively tackle a variety of image segmentation tasks within a unified in-context learning framework. However, these methods still struggle with task ambiguity in in-context segmentation, as not all in-context examples can accurately convey the task information. In order to address this issue, we present SINE, a simple image $\textbf{S}$egmentation framework utilizing $\textbf{in}$-context $\textbf{e}$xamples. Our approach leverages a Transformer encoder-decoder structure, where the encoder provides high-quality image representations, and the decoder is designed to yield multiple task-specific output masks to eliminate task ambiguity effectively. Specifically, we introduce an In-context Interaction module to complement in-context information and produce correlations between the target image and the in-context example and a Matching Transformer that uses fixed matching and a Hungarian algorithm to eliminate differences between different tasks. In addition, we have further perfected the current evaluation system for in-context image segmentation, aiming to facilitate a holistic appraisal of these models. Experiments on various segmentation tasks show the effectiveness of the proposed method. Yang Liu 0357, Chenchen Jing, Hengtao Li, Muzhi Zhu, Hao Chen 0041, Chunhua Shen |
NeurIPS | 4 |
| 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic SegmentationabstractThe Diffusion Model has not only garnered noteworthy achievements in the realm of image generation
but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data.
Drawing from the extensive potential unveiled by the Diffusion Model in both semantic correspondence and open vocabulary segmentation, our work initiates an investigation into employing the Latent Diffusion Model for Few-shot Semantic Segmentation.
Recently, inspired by the in-context learning ability of large language models, Few-shot Semantic Segmentation has evolved into In-context Segmentation tasks, morphing into a crucial element in assessing generalist segmentation models.
In this context, we concentrate
on Few-shot Semantic Segmentation,
establishing a solid foundation for the future development of a Diffusion-based generalist model for segmentation. Our initial focus lies in understanding how to facilitate interaction between the query image and the support image, resulting in the proposal of a KV fusion method within the self-attention framework.
Subsequently, we delve deeper into optimizing the infusion of information from the support mask and simultaneously re-evaluating how to provide reasonable supervision from the query mask.
Based on our analysis, we establish a simple and effective framework named DiffewS, maximally retaining the original Latent Diffusion Model's generative framework and effectively utilizing the pre-training prior. Experimental results demonstrate that our method significantly outperforms the previous SOTA models in multiple settings. Muzhi Zhu, Yang Liu 0357, Zekai Luo, Chenchen Jing, Hao Chen 0041, Guangkai Xu, Chunhua Shen |
NeurIPS | 1 |
| 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt LearningabstractCurrent closed-set instance segmentation models rely on pre-defined class labels for each mask during training and evaluation, largely limiting their ability to detect novel objects. Open-world instance segmentation (OWIS) models address this challenge by detecting unknown objects in a class-agnostic manner. However, previous OWIS approaches completely erase category information during training to keep the model’s ability to generalize to unknown objects. In this work, we propose a novel training mechanism termed SegPrompt that uses category information to improve the model’s class-agnostic segmentation ability for both known and unknown categories. In addition, the previous OWIS training setting exposes the unknown classes to the training set and brings information leakage, which is unreasonable in the real world. Therefore, we provide a new open-world benchmark closer to a real-world scenario by dividing the dataset classes into known-seen-unseen parts. For the first time, we focus on the model’s ability to discover objects that never appear in the training set images.Experiments show that SegPrompt can improve the overall and unseen detection performance by 5.6% and 6.1% in AR on our new benchmark without affecting the inference efficiency. We further demonstrate the effectiveness of our method on existing cross-dataset transfer and strongly supervised settings, leading to 5.5% and 12.3% relative improvement. Code and data are released at: https://github.com/aim-uofa/SegPrompt Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chengxiang Fan, Weian Mao, Chenchen Jing, Yifan Liu 0001, Chunhua Shen |
ICCV | 1 |