VLDB 2026 Research / reviewers in the wild / expert
Ziyang Yuan
dblp:192/1411
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PHIDE: A Parallel Hybrid Direct-Iterative Eigensolver for Hermitian Eigenvalue ProblemsabstractIn this paper, we propose a Parallel Hybrid Direct-Iterative Eigensolver for Hermitian Eigenvalue Problems without tridiagonalization, denoted byPHIDE, which combines direct and iterative methods.PHIDEfirst reduces a Hermitian matrix to banded form, then applies a spectrum slicing algorithm to the banded matrix, and finally computes the eigenvectors of the original matrix via backtransformation. Compared with conventional direct eigensolvers,PHIDEavoids tridiagonalization, which involves many memory-bound operations. InPHIDE, the banded eigenvalue problem is solved using the contour integral method implemented in FEAST, which may yield slightly lower accuracy than tridiagonalization-based approaches. For sequences of correlated Hermitian eigenvalue problems arising in density functional theory (DFT),PHIDEachieves an average speedup of$1.22\times$over the state-of-the-art direct solver in ELPA when using 1024 processes. Numerical experiments are conducted on dense Hermitian matrices from real applications as well as large sparse matrices from the SuiteSparse and ELSES collections. Shengguo Li, Xinzhe Wu, José E. Román, Ziyang Yuan, Ruibo Wang, Xuguang Chen |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Image Conductor: Precision Control for Interactive Video SynthesisabstractFilmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video creation, achieving precise control over motion for interactive video asset generation remains challenging. To this end, we propose Image Conductor, a method for precise control of camera transitions and object movements to generate video assets from a single image. An well-cultivated training strategy is proposed to separate distinct camera and object motion by camera LoRA weights and object LoRA weights. To further eliminate motion ambiguity from ill-posed trajectories, we introduce a camera-free guidance technique during inference process, enhancing object movements while eliminating camera transitions. Additionally, we develop a trajectory-oriented video motion data curation pipeline for training. Quantitative and qualitative experiments demonstrate our method's precision and fine-grained control in generating motion-controllable videos from images, advancing the practical application of interactive video synthesis. Yaowei Li 0001, Xintao Wang 0002, Zhaoyang Zhang 0004, Zhouxia Wang, Ziyang Yuan, Liangbin Xie, Ying Shan, Yuexian Zou |
AAAI | 5 |
| 2025 | SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse ViewpointsabstractRecent advancements in video diffusion models demonstrate remarkable capabilities in simulating real-world dynamics and 3D consistency. This progress motivates us to explore the potential of these models to maintain dynamic consistency across diverse viewpoints, a feature highly sought after in applications like virtual filming. Unlike existing methods focused on multi-view generation of single objects for 4D reconstruction, our interest lies in generating open-world videos from arbitrary viewpoints, incorporating six degrees of freedom (6 DoF) camera poses.
To achieve this, we propose a plug-and-play module that enhances a pre-trained text-to-video model for multi-camera video generation, ensuring consistent content across different viewpoints. Specifically, we introduce a multi-view synchronization module designed to maintain appearance and geometry consistency across these viewpoints. Given the scarcity of high-quality training data, we also propose a progressive training scheme that leverages multi-camera images and monocular videos as a supplement to Unreal Engine-rendered multi-camera videos. This comprehensive approach significantly benefits our model.
Experimental results demonstrate the superiority of our proposed method over existing competitors and several baselines. Furthermore, our method enables intriguing extensions, such as re-rendering a video from multiple novel viewpoints. Project webpage: https://jianhongbai.github.io/SynCamMaster/ Jianhong Bai, Menghan Xia, Xintao Wang 0002, Ziyang Yuan, Zuozhu Liu, Haoji Hu, Pengfei Wan 0001, Di Zhang 0026 |
ICLR | 4 |
| 2025 | 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video GenerationabstractThis paper aims to manipulate multi-entity 3D motions in video generation. Previous methods on controllable video generation primarily leverage 2D control signals to manipulate object motions and have achieved remarkable synthesis results. However, 2D control signals are inherently limited in expressing the 3D nature of object motions. To overcome this problem, we introduce 3DTrajMaster, a robust controller that regulates multi-entity dynamics in 3D space, given user-desired 6DoF pose (location and rotation) sequences of entities. At the core of our approach is a plug-and-play 3D-motion grounded object injector that fuses multiple input entities with their respective 3D trajectories through a gated self-attention mechanism. In addition, we exploit an injector architecture to preserve the video diffusion prior, which is crucial for generalization ability. To mitigate video quality degradation, we introduce a domain adaptor during training and employ an annealed sampling strategy during inference. To address the lack of suitable training data, we construct a 360-Motion Dataset, which first correlates collected 3D human and animal assets with GPT-generated trajectory and then captures their motion with 12 evenly-surround cameras on diverse 3D UE platforms. Extensive experiments show that 3DTrajMaster sets a new state-of-the-art in both accuracy and generalization for controlling multi-entity 3D motions. Project page: http://fuxiao0719.github.io/projects/3dtrajmaster Xintao Wang 0002, Sida Peng, Menghan Xia, Xiaoyu Shi 0002, Ziyang Yuan, Pengfei Wan 0001, Di Zhang 0026, Dahua Lin |
ICLR | 7 |
| 2025 | Improving Video Generation with Human FeedbackabstractVideo generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs. Jie Liu 0047, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Menghan Xia, Xintao Wang 0002, Xiaohong Liu 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Yujiu Yang 0001, Wanli Ouyang |
NeurIPS | 4 |
| 2025 | AVCPNet: An AAV-Vehicle Collaborative Perception Network for 3-D Object DetectionabstractWith the advancement of collaborative perception, the role of autonomous aerial vehicle (AAV)–vehicle collaborative perception has become increasingly significant. The demand for collaborative perception from various perspectives to construct comprehensive perceptual information is rising. However, challenges emerge due to differences in the field of view (FOV) between cross-domain agents and their varying sensitivities to image information. Furthermore, accurate depth information is essential for collaboration to transform image features into bird’s eye view (BEV) features. To address these challenges, we propose a framework specifically designed for aerial-ground collaboration. First, to address the deficiency of datasets for aerial-ground collaboration, we have developed a virtual dataset named V2U-COO for our research. Second, we design a cross-domain cross-adaptation (CDCA) module to align the target information obtained from different domains, thereby achieving more accurate perception results. Finally, we introduce a collaborative depth optimization (CDO) module to obtain more precise depth estimation results, leading to more accurate perception results. We conduct extensive experiments on both our virtual dataset and a public dataset to validate the effectiveness of our framework. Our method resolves the feature fusion issue under significant height differences, a challenge that previous BEV generation methods struggled to address effectively. Our experiments on the V2U-COO and DAIR-V2X datasets demonstrate improvements in detection accuracy of 6.1% and 2.7%, respectively. Our code will be released athttps://github.com/wyccoo/uvcp. Zhirui Wang 0003, Peirui Cheng, Pengju Tian, Ziyang Yuan, Liangjin Zhao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Blind Face Restoration under Extreme Conditions: Leveraging 3D-2D Prior Fusion for Superior Structural and Texture RecoveryabstractBlind face restoration under extreme conditions involves reconstructing high-quality face images from severely degraded inputs. These input images are often in poor quality and have extreme facial poses, leading to errors in facial structure and unnatural artifacts within the restored images. In this paper, we show that utilizing 3D priors effectively compensates for structure knowledge deficiencies in 2D priors while preserving the texture details. Based on this, we introduce FREx (Face Restoration under Extreme conditions) that combines structure-accurate 3D priors and texture-rich 2D priors in pretrained generative networks for blind face restoration under extreme conditions. To fuse the different information in 3D and 2D priors, we introduce an adaptive weight module that adjusts the importance of features based on the input image's condition. With this approach, our model can restore structure-accurate and natural-looking faces even when the images have lost a lot of information due to degradation and extreme pose. Extensive experimental results on synthetic and real-world datasets validate the effectiveness of our methods. Zhengrui Chen, Liying Lu, Ziyang Yuan, Yu Li 0003, Chun Yuan 0003, Weihong Deng |
AAAI | 3 |
| 2024 | SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language ModelsabstractCurrent instruction-based image editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper introduces SmartEdit, a novel approach of instruction-based image editing that leverages Multimodal Large Language Models (MLLMs) to enhance its understanding and reasoning capabilities. However, direct integration of these elements still faces challenges in situations requiring complex reasoning. To mitigate this, we propose a Bidirectional Interaction Module (BIM) that enables comprehensive bidirectional information interactions between the input image and the MLLM output. During training, we initially incorporate perception data to boost the perception and understanding capabilities of diffusion models. Subsequently, we demonstrate that a small amount of complex instruction editing data can effectively stimulate SmartEdit’ s editing capabilities for more complex instructions. We further construct a new evaluation dataset, Reason-Edit, specifically tailored for complex instruction-based image editing. Both quantitative and qualitative results on this evaluation dataset indicate that our SmartEdit surpasses previous methods, paving the way for the practical application of complex instruction-based image editing. Yuzhou Huang, Liangbin Xie, Xintao Wang 0002, Ziyang Yuan, Xiaodong Cun, Yixiao Ge, Jiantao Zhou 0001, Chao Dong 0005, Ruimao Zhang, Ying Shan |
CVPR | 4 |
| 2024 | CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion ModelsabstractIncorporating a customized object into image generation presents an attractive feature in text-to-image (T2I) generation. Some methods finetune T2I models for each object individually at test-time, which tend to be overfitted and time-consuming. Others train an extra encoder to extract object visual information for customization efficiently but struggle to preserve the object's identity. To address these limitations, we present CustomNet, a unified encoder-based object customization framework that explicitly incorporates 3D novel view synthesis capabilities into the customization process. This integration facilitates the adjustment of spatial positions and viewpoints, producing diverse outputs while effectively preserving the object's identity. To train our model effectively, we propose a dataset construction pipeline to better handle real-world objects and complex backgrounds. Additionally, we introduce delicate designs that enable location control and flexible background control through textual descriptions or user-defined backgrounds. Our method allows for object customization without the need of test-time optimization, providing simultaneous control over viewpoints, location, and text. Experimental results show that our method outperforms other customization methods regarding identity preservation, diversity, and harmony. Codes are available at https://github.com/TencentARC/CustomNet. Ziyang Yuan, Mingdeng Cao, Xintao Wang 0002, Zhongang Qi, Chun Yuan 0003, Ying Shan |
ACM Multimedia | 1 |
| 2024 | MiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsabstractSora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for generating Sora-like videos, as they mainly contain short videos with low motion intensity and brief captions. To address these issues, we propose MiraData, a high-quality video dataset that surpasses previous ones in video duration, caption detail, motion strength, and visual quality. We curate MiraData from diverse, manually selected sources and meticulously process the data to obtain semantically consistent clips. GPT-4V is employed to annotate structured captions, providing detailed descriptions from four different perspectives along with a summarized dense caption. To better assess temporal consistency and motion intensity in video generation, we introduce MiraBench, which enhances existing benchmarks by adding 3D consistency and tracking-based motion strength metrics. MiraBench includes 150 evaluation prompts and 17 metrics covering temporal consistency, motion strength, 3D consistency, visual quality, text-video alignment, and distribution similarity. To demonstrate the utility and effectiveness of MiraData, we conduct experiments using our DiT-based video generation model, MiraDiT. The experimental results on MiraBench demonstrate the superiority of MiraData, especially in motion strength. Xuan Ju, Yiming Gao 0007, Zhaoyang Zhang 0004, Ziyang Yuan, Xintao Wang 0002, Ailing Zeng, Qiang Xu 0001, Ying Shan |
NeurIPS | 4 |
| 2024 | Untrained neural network embedded Fourier phase retrieval from few measurements
Ningyi Leng, Ziyang Yuan |
Signal Process. | 4 |
| 2024 | Phase Retrieval With Background Information: Decreased References and Efficient MethodsabstractFourier phase retrieval (PR) is a severely ill-posed inverse problem that arises in various applications. To guarantee a unique solution and relieve the dependence on the initialization, background information can be exploited as a structural prior. However, the requirement for the background information may be challenging when moving to high-resolution imaging. At the same time, the previously proposed projected gradient descent (PGD) method also demands much background information. In this paper, we present an improved theoretical result about the demand for the background information, along with two Douglas Rachford (DR) based methods. Analytically, we demonstrate that the background information required to ensure a unique solution can be decreased by nearly$1/2$for the 2-D signals compared to the 1-D signals. By generalizing the results into d-dimension, we show that the length of the background information more than$\left ({{2^{\frac {d+1}{d}}-1}}\right)$folds of the signal is sufficient to ensure uniqueness. At the same time, we also analyze the stability and robustness of the model when the measurements and background information are corrupted by noise. Furthermore, two methods called Background Douglas Rachford (BDR) and Convex Background Douglas Rachford (CBDR) are proposed. BDR, which is a kind of non-convex method, is proven to have the local R-linear convergence rate under mild assumptions. Instead, the CBDR method uses the techniques of convexification and can be proven to have a global convergence guarantee as long as the background information is sufficient. To support this, a new property called F-RIP is established. We test the performance of the proposed methods through simulations as well as real experimental measurements, and demonstrate that they achieve a higher recovery rate with less background information compared to the PGD method. Ziyang Yuan, Haoxing Yang, Ningyi Leng |
IEEE Trans. Inf. Theory | 1 |
| 2023 | Make Encoder Great Again in 3D GAN Inversion through Geometry and Occlusion-Aware Encodingabstract3D GAN inversion aims to achieve high reconstruction fidelity and reasonable 3D geometry simultaneously from a single image input. However, existing 3D GAN inversion methods rely on time-consuming optimization for each individual case. In this work, we introduce a novel encoder-based inversion framework based on EG3D, one of the most widely-used 3D GAN models. We leverage the inherent properties of EG3D’s latent space to design a discriminator and a background depth regularization. This enables us to train a geometry-aware encoder capable of converting the input image into corresponding latent code. Additionally, we explore the feature space of EG3D and develop an adaptive refinement stage that improves the representation ability of features in EG3D to enhance the recovery of fine-grained textural details. Finally, we propose an occlusion-aware fusion operation to prevent distortion in unobserved regions. Our method achieves impressive results comparable to optimization-based methods while operating up to 500 times faster. Our framework is well-suited for applications such as semantic editing. Ziyang Yuan, Yu Li 0003, Chun Yuan 0003 |
ICCV | 1 |
| 2022 | One Model to Edit Them All: Free-Form Text-Driven Image Manipulation with Semantic ModulationsabstractFree-form text prompts allow users to describe their intentions during image manipulation conveniently. Based on the visual latent space of StyleGAN[21] and text embedding space of CLIP[34], studies focus on how to map these two latent spaces for text-driven attribute manipulations. Currently, the latent mapping between these two spaces is empirically designed and confines that each manipulation model can only handle one fixed text prompt. In this paper, we propose a method named Free-Form CLIP (FFCLIP), aiming to establish an automatic latent mapping so that one manipulation model handles free-form text prompts. Our FFCLIP has a cross-modality semantic modulation module containing semantic alignment and injection. The semantic alignment performs the automatic latent mapping via linear transformations with a cross attention mechanism. After alignment, we inject semantics from text prompt embeddings to the StyleGAN latent space. For one type of image (e.g., human portrait'), one FFCLIP model can be learned to handle free-form text prompts. Meanwhile, we observe that although each training text prompt only contains a single semantic meaning, FFCLIP can leverage text prompts with multiple semantic meanings for image manipulation. In the experiments, we evaluate FFCLIP on three types of images (i.e.,human portraits', cars', andchurches'). Both visual and numerical results show that FFCLIP effectively produces semantically accurate and visually realistic images. Project page: https://github.com/KumapowerLIU/FFCLIP. Yibing Song, Ziyang Yuan, Xintong Han, Chun Yuan 0003, Qifeng Chen 0001, Jue Wang 0001 |
NeurIPS | 4 |