VLDB 2026 Research / reviewers in the wild / expert
Zili Yi
dblp:191/1594
· DBLP profile ↗
26ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FSA-GS: Fine-Structure-Aware Gaussian Splatting for sparse-view novel view synthesis
Yangeng Li, Hongchen Tan, Zili Yi, Jingchun Zhou |
Comput. Vis. Image Underst. | 6 |
| 2026 | DreamBarbie: Text to Barbie-Style 3D AvatarsabstractTo integrate digital humans into everyday life, there is a strong demand for generating high-quality, fine-grained disentangled 3D avatars that support expressive animation and simulation capabilities, ideally from low-cost textual inputs. Although text-driven 3D avatar generation has made significant progress by leveraging 2D generative priors, existing methods still struggle to fulfill all these requirements simultaneously. To address this challenge, we propose DreamBarbie, a novel text-driven framework for generating animatable 3D avatars with separable shoes, accessories, and simulation-ready garments, truly capturing the iconic "Barbie doll" aesthetic. The core of our framework lies in an expressive 3D representation combined with appropriate modeling constraints. Unlike prior methods, we use G-Shell to uniformly model watertight components (e.g., bodies, shoes) and non-watertight garments. By reformulating boundaries as euclidean field intersections instead of manifold geodesics, we propose an SDF-based initialization and a hole regularization loss that together achieve a $100\times$100× speedup and stable open topology without image input. These disentangled 3D representations are then optimized by specialized expert diffusion models tailored to each domain, ensuring high-fidelity outputs. To mitigate geometric artifacts and texture conflicts when combining different expert models, we further propose several effective geometric losses and strategies. Extensive experiments demonstrate that DreamBarbie outperforms existing methods in both dressed human and outfit generation. Our framework further enables diverse applications, including apparel combination, editing, expressive animation, and physical simulation. Xiaokun Sun, Zhenyu Zhang 0005, Ying Tai, Hao Tang 0005, Zili Yi, Jian Yang 0003 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | SigStyle: Signature Style Transfer via Personalized Text-to-Image ModelsabstractStyle transfer enables the seamless integration of artistic styles from a style image into a content image, resulting in visually striking and aesthetically enriched outputs. Despite numerous advances in this field, existing methods did not explicitly focus on the signature style, which represents the distinct and recognizable visual traits of the image such as geometric and structural patterns, color palettes and brush strokes etc. In this paper, we introduce SigStyle, a framework that leverages the semantic priors that embedded in a personalized text-to-image diffusion model to capture the signature style representation. This style capture process is powered by a hypernetwork that efficiently fine-tunes the diffusion model for any given single style image. Style transfer then is conceptualized as the reconstruction process of content image through learned style tokens from the personalized diffusion model. Additionally, to ensure the content consistency throughout the style transfer process, we introduce a time-aware attention swapping technique that incorporates content information from the original image into the early denoising steps of target image generation. Beyond enabling high-quality signature style transfer across a wide range of styles, SigStyle supports multiple interesting applications, such as local style transfer, texture transfer, style fusion and style-guided text-to-image generation. Quantitative and qualitative evaluations demonstrate our approach outperforms existing style transfer methods for recognizing and transferring the signature styles. Tongyuan Bai, Xuping Xie, Zili Yi, Rui Ma 0011 |
AAAI | 4 |
| 2025 | Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image GenerationabstractRecent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduced control flexibility. These challenges arise from current end-to-end inpainting models, which suffer from inaccurate training masks, limited foreground semantic understanding, data distribution biases, and inherent interference between visual and textual prompts. To overcome these limitations, we present Anywhere, a multi-agent framework that departs from the traditional end-to-end approach. In this framework, each agent is specialized in a distinct aspect, such as foreground understanding, diversity enhancement, object integrity protection, and textual prompt consistency. Our framework is further enhanced with the ability to incorporate optional user textual inputs, perform automated quality assessments, and initiate re-generation as needed. Comprehensive experiments demonstrate that this modular design effectively overcomes the limitations of existing end-to-end models, resulting in higher fidelity, quality, diversity and controllability in foreground-conditioned image generation. Additionally, the Anywhere framework is extensible, allowing it to benefit from future advancements in each individual agent. Tianyidan Xie, Rui Ma 0011, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang 0005, Lanjun Wang, Zili Yi |
AAAI | 9 |
| 2025 | SingleDream: Attribute-Driven T2I Customization from a Single Reference Image
Zili Yi, Tieru Wu, Rui Ma 0011 |
CVM (2) | 3 |
| 2025 | OmniStyle: Filtering High Quality Style Transfer Data at ScaleabstractIn this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual descriptions and instruction prompts. We show that OmniStyle-1M can not only enable efficient and scalable of style transfer models through supervised training but also facilitate precise control over target stylization. Especially, to ensure the quality of the dataset, we introduce OmniFilter, a comprehensive style transfer quality assessment framework, which filters high-quality triplets based on content preservation, style consistency, and aesthetic appeal. Building upon this foundation, we propose OmniStyle, a framework based on the Diffusion Transformer (DiT) architecture designed for high-quality and efficient style transfer. This framework supports both instruction-guided and imageguided style transfer, generating high resolution outputs with exceptional detail. Extensive qualitative and quantitative evaluations demonstrate OmniStyle’s superior performance compared to existing approaches, highlighting its efficiency and versatility. OmniStyle-1M and its accompanying methodologies provide a significant contribution to advancing high-quality style transfer, offering a valuable resource for the research community. Jiang Lin, Zili Yi, Rui Ma 0011 |
CVPR | 5 |
| 2025 | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral PerspectiveabstractUltra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restoration problem into three progressive stages: zero-frequency enhancement, low-frequency restoration, and high-frequency refinement. Building on this insight, we propose a novel framework, ERR, which comprises three collaborative sub-networks: the zero-frequency enhancer (ZFE), the low-frequency restorer (LFR), and the high-frequency refiner (HFR). Specifically, the ZFE integrates global priors to learn global mapping, while the LFR restores low-frequency information, emphasizing reconstruction of coarse-grained content. Finally, the HFR employs our designed frequency-windowed kolmogorov-arnold networks (FW-KAN) to refine textures and details, producing high-quality image restoration. Our approach significantly outperforms previous UHD methods across various tasks, with extensive ablation studies validating the effectiveness of each component. The code is available at here. Zhizhou Chen, Yunzhe Xu, Enxuan Gu, Zili Yi, Ying Tai |
CVPR | 6 |
| 2025 | One-Shot Learning for Pose-Guided Person Image Synthesis in the WildabstractCurrent Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU. Dongqi Fan, Rui Ma 0011, Qiang Tang 0002, Zili Yi |
ICASSP | 6 |
| 2025 | FreeControl: Efficient, Training-Free Structural Control via One-Step Attention ExtractionabstractControlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising.
We present \textbf{FreeControl}, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs \textit{one-step attention extraction} from a single, optimally chosen timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining.
To further improve quality and stability, we introduce \textit{Latent-Condition Decoupling (LCD)}: a principled separation of the timestep condition and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources, enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control—enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at ~5\% additional cost. Jiang Lin, Zhiqiu Zhang, Jizhi Zhang, Qiang Tang 0002, Zili Yi |
NeurIPS | 10 |
| 2025 | DP-Adapter: Dual-pathway adapter for boosting fidelity and text consistency in customizable human image generationabstractWith the growing popularity of personalized human content creation and sharing, there is a rising demand for advanced techniques in customized human image generation. However, current methods struggle to simultaneously maintain the fidelity of human identity and ensure the consistency of textual prompts, often resulting in suboptimal outcomes. This shortcoming is primarily due to the lack of effective constraints during the simultaneous integration of visual and textual prompts, leading to unhealthy mutual interference that compromises the full expression of both types of input. Building on prior research that suggests visual and textual conditions influence different regions of an image in distinct ways, we introduce a novel Dual-Pathway Adapter (DP-Adapter) to enhance both high-fidelity identity preservation and textual consistency in personalized human image generation. Our approach begins by decoupling the target human image into visually sensitive and text-sensitive regions. For visually sensitive regions, DP-Adapter employs an Identity-Enhancing Adapter (IEA) to preserve detailed identity features. For text-sensitive regions, we introduce a Textual-Consistency Adapter (TCA) to minimize visual interference and ensure the consistency of textual semantics. To seamlessly integrate these pathways, we develop a Fine-Grained Feature-Level Blending (FFB) module that efficiently combines hierarchical semantic features from both pathways, resulting in more natural and coherent synthesis outcomes. Additionally, DP-Adapter supports various innovative applications, including controllable headshot-to-full-body portrait generation, age editing, old-photo to reality, and expression editing. Extensive experiments demonstrate that DP-Adapter outperforms state-of-the-art methods in both visual fidelity and text consistency, highlighting its effectiveness and versatility in the field of human image generation. Xuping Xie, Lanjun Wang, Zili Yi, Rui Ma 0011 |
Graph. Model. | 5 |
| 2025 | Real-time dual-eye collaborative eyeblink detection with contrastive learning
Hanli Zhao, Yu Wang 0219, Wanglong Lu, Zili Yi, Minglun Gong |
Pattern Recognit. | 4 |
| 2025 | Draw What You Hear: High-Fidelity Image Generation and Manipulation via SoundAdapterabstractCurrently, the text-to-image (T2I) generation has established itself as a cornerstone within the realm of AI-generated content (AIGC), due its remarkable success to the availability of extensive datasets comprising paired text-vision samples. Nevertheless, the absence of audio-visual pairs hinders the growth of audio-to-image (A2I). Although prior approaches have pioneered the A2I task, the tight entanglement between initial audio and image encoders imposes the challenge of gathering audio-visual samples, resulting in degraded performance and limited sound flexibility. Therefore, this article proposes a novel SoundAdapter to draw what you hear. Specifically, the SoundAdapter's structure is meticulously designed around transformer blocks, which are critical for capturing overarching patterns and dependencies within the data. In addition, it integrates a sophisticated multigranularity approach coupled with a hybrid supervisory signal, ensuring both fine-grained semantic alignment and seamless optimization across various levels of representation. Extensive tests demonstrate that the SoundAdapter excels in training, setting new benchmarks in zero-shot audio classification, as well as in creating and modifying images across a variety of datasets. The implementation code and several demos supporting this study are openly accessible at https://github.com/CV-MM-Lab/SoundAdapter, facilitating reproducibility and further research. Mingjie Wang 0002, Song Yuan, Xian-Feng Han, Zili Yi |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
Zili Yi, Rui Ma 0011 |
ECCV (30) | 3 |
| 2023 | ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image EditingabstractThe StyleGAN family succeed in high-fidelity image generation and allow for flexible and plausible editing of generated images by manipulating the semantic-rich latent style space. However, projecting a real image into its latent space encounters an inherent trade-off between inversion quality and editability. Existing encoder-based or optimization-based StyleGAN inversion methods attempt to mitigate the trade-off but see limited performance. To fundamentally resolve this problem, we propose a novel two-phase framework by designating two separate networks to tackle editing and reconstruction respectively, instead of balancing the two. Specifically, in Phase I, a W-space-oriented StyleGAN inversion network is trained and used to perform image inversion and edit- ing, which assures the editability but sacrifices reconstruction quality. In Phase II, a carefully designed rectifying network is utilized to rectify the inversion errors and perform ideal reconstruction. Experimental results show that our approach yields near-perfect reconstructions without sacrificing the editability, thus allowing accurate manipulation of real images. Further, we evaluate the performance of our rectifying net- work, and see great generalizability towards unseen manipulation types and out-of-domain images. Bingchuan Li, Tianxiang Ma, Miao Hua, Zili Yi |
AAAI | 7 |
| 2023 | DyStyle: Dynamic Neural Network for Multi-Attribute-Conditioned Style EditingsabstractThe semantic controllability of StyleGAN is enhanced by unremitting research. Although the existing weak supervision methods work well in manipulating the style codes along one attribute, the accuracy of manipulating multiple attributes is neglected. Multi-attribute representations are prone to entanglement in the StyleGAN latent space, while sequential editing leads to error accumulation. To address these limitations, we design a Dynamic Style Manipulation Network (DyStyle) whose structure and parameters vary by input samples, to perform nonlinear and adaptive manipulation of latent codes for flexible and precise attribute control. In order to efficient and stable optimization of the DyStyle network, we propose a Dynamic Multi-Attribute Contrastive Learning (DmaCL) method: including dynamic multi-attribute contrastor and dynamic multi-attribute contrastive loss, which simultaneously disentangle a variety of attributes from the generative image and latent space of model. As a result, our approach demonstrates fine-grained disentangled edits along multiple numeric and binary attributes. Qualitative and quantitative comparisons with existing style manipulation methods verify the superiority of our method in terms of the multi-attribute control accuracy and identity preservation without compromising photorealism. Bingchuan Li, Shaofei Cai, Miao Hua, Zili Yi |
WACV | 7 |
| 2022 | XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font GenerationabstractGenerating a new font library is a very labor-intensive and time-consuming job for glyph-rich scripts. Few-shot font generation is thus required, as it requires only a few glyph references without fine-tuning during test. Existing methods follow the style-content disentanglement paradigm and expect novel fonts to be produced by combining the style codes of the reference glyphs and the content representations of the source. However, these few-shot font generation methods either fail to capture content-independent style representations, or employ localized component-wise style representations, which is insufficient to model many Chinese font styles that involve hyper-component features such as inter-component spacing and “connected-stroke”. To resolve these drawbacks and make the style representations more reliable, we propose a self-supervised cross-modality pre-training strategy and a cross-modality transformer-based encoder that is conditioned jointly on the glyph image and the corresponding stroke labels. The cross-modality encoder is pre-trained in a self-supervised manner to allow effective capture of cross- and intra-modality correlations, which facilitates the content-style disentanglement and modeling style representations of all scales (strokelevel, component-level and character-level). The pretrained encoder is then applied to the downstream font generation task without fine-tuning. Experimental comparisons of our method with state-of-the-art methods demonstrate our method successfully transfers styles of all scales. In addition, it only requires one reference glyph and achieves the lowest rate of bad cases in the few-shot font generation task (28% lower than the second best). Fangyue Liu, Zili Yi |
CVPR | 5 |
| 2022 | Region-Aware Face SwappingabstractThis paper presents a novel Region-Aware Face Swapping (RAFSwap) network to achieve identity-consistent harmonious high-resolution face generation in a local-global manner: 1) Local Facial Region-Aware (FRA) branch augments local identity-relevant features by introducing the Transformer to effectively model misaligned crossscale semantic interaction. 2) Global Source Feature-Adaptive (SFA) branch further complements global identity-relevant cues for generating identity-consistent swapped faces. Besides, we propose a Face Mask Predictor (FMP) module incorporated with StyleGAN2 to predict identity-relevant soft facial masks in an unsupervised manner that is more practical for generating harmonious high-resolution faces. Abundant experiments qualitatively and quantitatively demonstrate the superiority of our method for generating more identity-consistent high-resolution swapped faces over SOTA methods, e.g., obtaining 96.70 ID retrieval that outperforms SOTA MegaFS by$5.87\uparrow$. Chao Xu 0023, Jiangning Zhang, Miao Hua, Zili Yi, Yong Liu 0007 |
CVPR | 5 |
| 2021 | Spatial-Temporal Residual Aggregation for High Resolution Video Inpainting
Vishnu Sanjay Ramiya Srinivasan, Rui Ma 0011, Qiang Tang 0002, Zili Yi |
BMVC | 4 |
| 2020 | Contextual Residual Aggregation for Ultra High-Resolution Image InpaintingabstractRecently data-driven image inpainting methods have made inspiring progress, impacting fundamental image editing tasks such as object removal and damaged image repairing. These methods are more effective than classic approaches, however, due to memory limitations they can only handle low-resolution inputs, typically smaller than 1K. Meanwhile, the resolution of photos captured with mobile devices increases up to 8K. Naive up-sampling of the low-resolution inpainted result can merely yield a large yet blurry result. Whereas, adding a high-frequency residual image onto the large blurry image can generate a sharp result, rich in details and textures. Motivated by this, we propose a Contextual Residual Aggregation (CRA) mechanism that can produce high-frequency residuals for missing contents by weighted aggregating residuals from contextual patches, thus only requiring a low-resolution prediction from the network. Since convolutional layers of the neural network only need to operate on low-resolution inputs and outputs, the cost of memory and computing power is thus well suppressed. Moreover, the need for high-resolution training datasets is alleviated. In our experiments, we train the proposed model on small images with resolutions 512 × 512 and perform inference on high-resolution images, achieving compelling inpainting quality. Our model can inpaint images as large as 8K with considerable hole sizes, which is intractable with previous learning-based approaches. We further elaborate on the light-weight design of the network architecture, achieving real-time performance on 2K images on a GTX 1080 Ti GPU. Codes are available at: https://github. com/Ascend-Huawei/Ascend-Canada/tree/ master/Models/Research_HiFIll_Model. Zili Yi, Qiang Tang 0002, Shekoofeh Azizi, Daesik Jang |
CVPR | 1 |
| 2020 | Animating Through Warping: An Efficient Method for High-Quality Facial Expression AnimationabstractAdvances in deep neural networks have considerably improved the art of animating a still image without operating in 3D domain. Whereas, prior arts can only animate small images (typically no larger than 512x512) due to memory limitations, difficulty of training and lack of high-resolution (HD) training datasets, which significantly reduce their potential for applications in movie production and interactive systems. Motivated by the idea that HD images can be generated by adding high-frequency residuals to low-resolution results produced by a neural network, we propose a novel framework known as Animating Through Warping (ATW) to enable efficient animation of HD images. Zili Yi, Qiang Tang 0002, Vishnu Sanjay Ramiya Srinivasan |
ACM Multimedia | 1 |
| 2020 | BSD-GAN: Branched Generative Adversarial Network for Scale-Disentangled Representation Learning and Image SynthesisabstractWe introduce BSD-GAN, a novel multi-branch and scale-disentangled training method which enables unconditional Generative Adversarial Networks (GANs) to learn image representations at multiple scales, benefiting a wide range of generation and editing tasks. The key feature of BSD-GAN is that it is trained in multiple branches, progressively covering both the breadth and depth of the network, as resolutions of the training images increase to reveal finer-scale features. Specifically, each noise vector, as input to the generator network of BSD-GAN, is deliberately split into several sub-vectors, each corresponding to, and is trained to learn, image representations at a particular scale. During training, we progressively "de-freeze" the sub-vectors, one at a time, as a new set of higher-resolution images is employed for training and more network layers are added. A consequence of such an explicit sub-vector designation is that we can directly manipulate and even combine latent (sub-vector) codes which model different feature scales. Extensive experiments demonstrate the effectiveness of our training method in scale-disentangled learning of image representations and synthesis of novel image contents, without any extra labels and without compromising quality of the synthesized high-resolution images. We further demonstrate several image generation and manipulation applications enabled or improved by BSD-GAN. Zili Yi, Hao Cai 0004, Wendong Mao, Minglun Gong, Hao (Richard) Zhang |
IEEE Trans. Image Process. | 1 |
| 2019 | A Global-Matching Framework for Multi-View Stereopsis
Wendong Mao, Minglun Gong, Xin Huang 0030, Hao Cai 0004, Zili Yi |
CAIP (1) | 5 |
| 2019 | Blind quality assessment of gamut-mapped images via local and global statistical analysis
Hao Cai 0004, Leida Li, Zili Yi, Minglun Gong |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Towards a blind image quality evaluator using multi-scale second-order statistics
Hao Cai 0004, Leida Li, Zili Yi, Minglun Gong |
Signal Process. Image Commun. | 3 |
| 2017 | DualGAN: Unsupervised Dual Learning for Image-to-Image TranslationabstractConditional Generative Adversarial Networks (GANs) for cross-domain image-to-image translation have made much progress recently [7, 8, 21, 12, 4, 18]. Depending on the task complexity, thousands to millions of labeled image pairs are needed to train a conditional GAN. However, human labeling is expensive, even impractical, and large quantities of data may not always be available. Inspired by dual learning from natural language translation [23], we develop a novel dual-GAN mechanism, which enables image translators to be trained from two sets of unlabeled images from two domains. In our architecture, the primal GAN learns to translate images from domain U to those in domain V, while the dual GAN learns to invert the task. The closed loop made by the primal and dual tasks allows images from either domain to be translated and then reconstructed. Hence a loss function that accounts for the reconstruction error of images can be used to train the translators. Experiments on multiple image translation tasks with unlabeled data show considerable performance gain of DualGAN over a single GAN. For some tasks, DualGAN can even achieve comparable or slightly better results than conditional GAN trained on fully labeled data. Zili Yi, Hao (Richard) Zhang, Ping Tan 0002, Minglun Gong |
ICCV | 1 |
| 2017 | Artistic stylization of face photos based on a single exemplar
Zili Yi, Songyuan Ji, Minglun Gong |
Vis. Comput. | 1 |