EDBT 2026 Demo / reviewers in the wild / expert
Mude Hui
dblp:342/2735
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 47% Vision and language · 26% 3D vision · 15% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 76% Image and video coding · 24% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.1 | 3 | 2025 | Unifying Layout Generation with a Decoupled Diffusion Model · CVPR 2023 What If We Recaption Billions of Web Images with LLaMA-3? · ICML 2025 MicroDiffusion: Implicit Representation-Guided Diffusion for 3D Reconstruction from Limited 2D Microscopy Projections · CVPR 2024 |
Machine learning › Generative modeling › diffusion model
image editing |
0.9 | 1 | 2025 | HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.9 | 1 | 2025 | HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025 |
Computer vision › Vision and language
vision-language pretraining |
0.9 | 1 | 2025 | What If We Recaption Billions of Web Images with LLaMA-3? · ICML 2025 |
Visual content generation and editing
image editing |
0.9 | 1 | 2025 | HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025 |
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing |
0.9 | 1 | 2025 | HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025 |
Computer vision › 3D vision
3d reconstruction |
0.8 | 1 | 2024 | MicroDiffusion: Implicit Representation-Guided Diffusion for 3D Reconstruction from Limited 2D Microscopy Projections · CVPR 2024 |
Digital forensics and information hiding › watermarking
image watermarking |
0.8 | 1 | 2024 | FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent Space · NeurIPS 2024 |
Digital forensics and information hiding › watermarking › multimedia watermarking
invisible watermarking |
0.8 | 1 | 2024 | FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent Space · NeurIPS 2024 |
Machine learning › Generative modeling › scene generation
layout generation |
0.7 | 1 | 2023 | Unifying Layout Generation with a Decoupled Diffusion Model · CVPR 2023 |
Visual content generation and editing › layout generation
graphic layout generation |
0.7 | 1 | 2023 | Unifying Layout Generation with a Decoupled Diffusion Model · CVPR 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.3 | 1 | 2025 | What If We Recaption Billions of Web Images with LLaMA-3? · ICML 2025 |
Computer vision › 3D vision
implicit neural representation |
0.2 | 1 | 2024 | MicroDiffusion: Implicit Representation-Guided Diffusion for 3D Reconstruction from Limited 2D Microscopy Projections · CVPR 2024 |
Machine learning › Generative modeling
variational autoencoder |
0.2 | 1 | 2024 | FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent Space · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
latent frequency optimization · 2.3VAE encoder · 2.3GPT-4V · 1.7DALL-E 3 · 1.7decoupled diffusion · 1.3multimodal large language model · 0.9large language model · 0.9linear interpolation · 0.8implicit neural representation · 0.8denoising diffusion probabilistic model · 0.8layout diffusion generative model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingabstractThis study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data. Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Peng Wang 0001, Cihang Xie, Yuyin Zhou |
ICLR | 1 |
| 2025 | What If We Recaption Billions of Web Images with LLaMA-3?abstractWeb-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investigations in this area remain predominantly closed-source. Our paper aims to bridge this community effort, leveraging the powerful and $\textit{open-sourced}$ LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered LLaVA-1.5 and then employ it to recaption ~1.3 billion images from the DataComp-1B dataset. Our empirical results confirm that this enhanced dataset, Recap-DataComp-1B, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe an average of 3.1% enhanced zero-shot performance cross four cross-modal retrieval tasks using a mixed set of the original and our captions. For generative models like text-to-image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Our project page is https://www.haqtu.me/Recap-Datacomp-1B/. Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang 0008, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu 0017, Huangjie Zheng, Yuyin Zhou, Cihang Xie |
ICML | 3 |
| 2024 | MicroDiffusion: Implicit Representation-Guided Diffusion for 3D Reconstruction from Limited 2D Microscopy ProjectionsabstractVolumetric optical microscopy using non-diffracting beams enables rapid imaging of 3D volumes by projecting them axially to 2D images but lacks crucial depth information. Addressing this, we introduce MicroDiffusion, a pi-oneering tool facilitating high-quality, depth-resolved 3D volume reconstruction from limited 2D projections. While existing Implicit Neural Representation (INR) models often yield incomplete outputs and Denoising Diffusion Prob-abilistic Models (DDPM) excel at capturing details, our method integrates INR's structural coherence with DDPM's fine-detail enhancement capabilities. We pretrain an INR model to transform 2D axially-projected images into a pre-liminary 3D volume. This pretrained INR acts as a global prior guiding DDPM's generative process through a linear interpolation between INR outputs and noise inputs. This strategy enriches the diffusion process with structured 3D information, enhancing detail and reducing noise in localized 2D images. By conditioning the diffusion model on the closest 2D projection, MicroDiffusion substantially enhances fidelity in resulting 3D reconstructions, surpassing INR and standard DDPM outputs with unparalleled image quality and structural fidelity. Our code and dataset are available at https://github.com/UCSC-VLAA/MicroDiffusion. Mude Hui, Zihao Wei, Hongru Zhu, Yuyin Zhou |
CVPR | 1 |
| 2024 | FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent SpaceabstractInvisible watermarking is essential for safeguarding digital content, enabling copyright protection and content authentication.
However, existing watermarking methods fall short in robustness against regeneration attacks.
In this paper, we propose a novel method called FreqMark that involves unconstrained optimization of the image latent frequency space obtained after VAE encoding. Specifically, FreqMark embeds the watermark by optimizing the latent frequency space of the images and then extracts the watermark through a pre-trained image encoder. This optimization allows a flexible trade-off between image quality with watermark robustness and effectively resists regeneration attacks.
Experimental results demonstrate that FreqMark offers significant advantages in image quality and robustness, permits flexible selection of the encoding bit number, and achieves a bit accuracy exceeding 90\% when encoding a 48-bit hidden message under various attack scenarios. Yiyang Guo, Mude Hui, Hanzhong Guo, Chuangjian Cai, Le Wan, Shangfei Wang |
NeurIPS | 3 |
| 2023 | Unifying Layout Generation with a Decoupled Diffusion ModelabstractLayout generation aims to synthesize realistic graphic scenes consisting of elements with different attributes in-cluding category, size, position, and between-element relation. It is a crucial task for reducing the burden on heavyduty graphic design works for formatted scenes, e.g., publications, documents, and user interfaces (UIs). Diverse application scenarios impose a big challenge in unifying various layout generation subtasks, including conditional and unconditional generation. In this paper, we propose a Layout Diffusion Generative Model (LDGM) to achieve such unification with a single decoupled diffusion model. LDGM views a layout of arbitrary missing or coarse element attributes as an intermediate diffusion status from a completed layout. Since different attributes have their individual semantics and characteristics, we propose to decouple the diffusion processes for them to improve the diversity of training samples and learn the reverse process jointly to exploit global-scope contexts for facilitating generation. As a result, our LDGM can generate layouts either from scratch or conditional on arbitrary available attributes. Extensive qualitative and quantitative experiments demonstrate our proposed LDGM outperforms existing layout generation models in both functionality and performance. Mude Hui, Zhizheng Zhang 0004, Wenxuan Xie, Yuwang Wang, Yan Lu 0001 |
CVPR | 1 |