VLDB 2026 Research / reviewers in the wild / expert
Hu Ye
dblp:201/8427
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-0848-9301ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 84% Representation and self-supervised learning · 10% Vision and language · 6% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models · AAAI 2025 IMAGDressing-v1: Customizable Virtual Dressing · AAAI 2025 Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models · ICLR 2024 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
1.6 | 2 | 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models · AAAI 2025 Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models · ICLR 2024 |
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
0.9 | 1 | 2025 | TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable image generation |
0.9 | 1 | 2025 | IMAGDressing-v1: Customizable Virtual Dressing · AAAI 2025 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.9 | 1 | 2025 | TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation · CVPR 2025 |
Machine learning › Generative modeling › cross-modal generation
story visualization |
0.9 | 1 | 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models · AAAI 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models · AAAI 2025 |
Visual content generation and editing
virtual try-on |
0.9 | 1 | 2025 | IMAGDressing-v1: Customizable Virtual Dressing · AAAI 2025 |
Visual content generation and editing › image generation
person image generation |
0.8 | 1 | 2024 | Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models · ICLR 2024 |
Visual content generation and editing › image generation › person image generation
pose-guided person image synthesis |
0.8 | 1 | 2024 | Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models · ICLR 2024 |
Computer vision › Vision and language
multimodal understanding |
0.3 | 1 | 2025 | TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation · CVPR 2025 |
Computer vision › Vision and language › multimodal representation
visual-semantic representation |
0.3 | 1 | 2025 | TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
u-net · 1.7hybrid attention · 1.7controlnet · 1.7VAE · 1.7IP-Adapter · 1.7CLIP · 1.7transformer diffusion · 0.9frame-prior transformer · 0.9dual-codebook architecture · 0.9conditional diffusion model · 0.9progressive conditional diffusion · 0.8inpainting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IMAGDressing-v1: Customizable Virtual DressingabstractExisting virtual try-on (VTON) methods provide only limited user control over garment attributes and generally overlook essential factors such as face, pose, and scene context. To address these limitations, we introduce the virtual dressing (VD) task, which aims to synthesize freely editable human images conditioned on fixed garments and optional user-defined inputs. We further propose a comprehensive affinity metric index (CAMI) to quantify the consistency between generated outputs and reference garments. We present IMAGDressing-v1, which leverages a garment-specific U-Net to integrate semantic features from CLIP and texture features from a VAE. To incorporate these garment features into a frozen denoising U-Net for flexible text-driven scene control, we employ a hybrid attention mechanism composed of frozen self-attention and trainable cross-attention layers. IMAGDressing-v1 seamlessly integrates with extension modules, such as ControlNet and IP-Adapter, enabling enhanced diversity and controllability. To alleviate data constraints, we introduce the Interactive Garment Pairing (IGPair) dataset, comprising over 300,000 garment–image pairs and a standardized data assembly pipeline. Extensive experiments demonstrate that IMAGDressing-v1 achieves state-of-the-art performance in controlled human image synthesis. The code and model will be available at https://github.com/muzishen/IMAGDressing. Fei Shen 0004, Xin Jiang 0010, Hu Ye, Cong Wang 0034, Xiaoyu Du 0002, Zechao Li, Jinhui Tang 0001 |
AAAI | 4 |
| 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsabstractRecent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which primarily generate stories in a caption-dependent manner, often overlook the importance of contextual consistency and the relevance of frames during sequential generation. To address this, we propose a novel Rich-contextual Conditional Diffusion Models (RCDMs), a two-stage approach designed to enhance story generation's semantic consistency and temporal consistency. Specifically, in the first stage, the frame-prior transformer diffusion model is presented to predict the frame semantic embedding of the unknown clip by aligning the semantic correlations between the captions and frames of the known clip. The second stage establishes a robust model with rich contextual conditions, including reference images of the known clip, the predicted frame semantic embedding of the unknown clip, and text embeddings of all captions. By jointly injecting these rich contextual conditions at the image and feature levels, RCDMs can generate semantic and temporal consistency stories. Moreover, RCDMs can generate consistent stories with a single forward inference compared to autoregressive models. Our qualitative and quantitative results demonstrate that our proposed RCDMs outperform in challenging scenarios. Fei Shen 0004, Hu Ye, Sibo Liu, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011 |
AAAI | 2 |
| 2025 | TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationabstractWe present TokenFlow, a novel unified image tokenizer that bridges the long-standing gap between multimodal understanding and generation. Prior research attempt to employ a single reconstruction-targeted Vector Quantization (VQ) encoder for unifying these two tasks. We observe that understanding and generation require fundamentally different granularities of visual information. This leads to a critical trade-off, particularly compromising performance in multimodal understanding tasks. TokenFlow addresses this challenge through an innovative dual-codebook architecture that decouples semantic and pixel-level feature learning while maintaining their alignment via a shared mapping mechanism. This design enables direct access to both high-level semantic representations crucial for understanding tasks and fine-grained visual features essential for generation through shared indices. Our extensive experiments demonstrate TokenFlow’s superiority across multiple dimensions. Leveraging TokenFlow, we demonstrate for the first time that discrete visual input can surpass LLaVA-1.5 13B in understanding performance, achieving a 7.2% average improvement. For image reconstruction, we achieve a strong FID score of 0.63 at 384×384 resolution. Moreover, TokenFlow establishes state-of-the-art performance in autoregressive image generation with a GenEval score of 0.55 at 256×256 resolution, achieving comparable results to SDXL. Our code and models are released at https://github.com/ByteFlow-Ai/TokenFlow. Liao Qu, Huichao Zhang, Yi Jiang 0009, Hu Ye, Daniel K. Du, Zehuan Yuan |
CVPR | 7 |
| 2025 | Active Security Control for Networked Jumping Systems Under Asynchronous Dual-Channel DoS Attacks: An HMM-Based ApproachabstractThis paper aims to address the design problem of asynchronous controllers for networked jumping systems (NJSs), in which two communication channels are vulnerable to asynchronous denial-of-service (DoS) attacks. To this end, a hidden Markov model (HMM)-based active security control strategy is proposed to counter asynchronous dual-channel DoS attacks. Firstly, two maximum consecutive numbers are introduced to describe DoS attacks randomly initiated by adversaries. Subsequently, a mode-dependent predictor is designed to generate predictive states, which can be employed in the switched controller to stabilize the NJSs. Furthermore, sufficient conditions are derived by constructing Lyapunov functions to ensure the asymptotic mean-square stability of the NJSs under asynchronous dual-channel DoS attacks. Finally, a simulation using a space robot manipulator model is presented to validate the effectiveness and practicality of the proposed active security control strategy. Peng Cheng 0010, Hu Ye, Di Wu 0058, Weidong Zhang 0004 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion ModelsabstractRecent work has showcased the significant potential of diffusion models in pose-guided person image synthesis.
However, owing to the inconsistency in pose between the source and target images, synthesizing an image with a distinct pose, relying exclusively on the source image and target pose information, remains a formidable challenge.
This paper presents Progressive Conditional Diffusion Models (PCDMs) that incrementally bridge the gap between person images under the target and source poses through three stages.
Specifically, in the first stage, we design a simple prior conditional diffusion model that predicts the global features of the target image by mining the global alignment relationship between pose coordinates and image appearance.
Then, the second stage establishes a dense correspondence between the source and target images using the global features from the previous stage, and an inpainting conditional diffusion model is proposed to further align and enhance the contextual features, generating a coarse-grained person image.
In the third stage, we propose a refining conditional diffusion model to utilize the coarsely generated image from the previous stage as a condition, achieving texture restoration and enhancing fine-detail consistency.
The three-stage PCDMs work progressively to generate the final high-quality and high-fidelity synthesized image.
Both qualitative and quantitative results demonstrate the consistency and photorealism of our proposed PCDMs under challenging scenarios.
The code and model will be available at https://github.com/tencent-ailab/PCDMs. Fei Shen 0004, Hu Ye, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011 |
ICLR | 2 |