VLDB 2026 Research / reviewers in the wild / expert
Ruoyu Feng 0001
dblp:251/9999-1
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0001-5226-1905ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HomoGen: Enhanced Video Inpainting via Homography Propagation and DiffusionabstractIn this paper, we present HomoGen, an enhanced video inpainting method based on homography propagation and diffusion models. HomoGen leverages homography registration to propagate contextual pixels as priors for generating missing content in corrupted videos. Unlike previous flow-based propagation methods, which introduce local distortions due to point-to-point optical flows, homography-induced artifacts are typically global structural distortions that preserve semantic integrity. To effectively utilize these priors for generation, we employ a video diffusion model that inherently prioritizes semantic information within the priors over pixel-level details. A content-adaptive control mechanism is proposed to scale and inject the priors into intermediate video latents during iterative denoising. In contrast to existing transformer-based networks that often suffer from artifacts within priors, leading to error accumulation and unrealistic results, our denoising diffusion network can smooth out artifacts and ensure natural outputs. Extensive experiments demonstrate the effectiveness of the proposed method qualitatively and quantitatively. Ding Ding 0004, Yueming Pan, Ruoyu Feng 0001, Qi Dai 0001, Jianmin Bao, Chong Luo 0001, Zhenzhong Chen 0001 |
CVPR | 3 |
| 2025 | Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative PriorabstractImage compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligent tasks and aids human understanding. Concurrently, enhanced perceptual quality not only improves visual appeal but also, by ensuring realistic image distributions, benefits semantic feature extraction for machine tasks.
Based on this insight, we propose Diff-ICMH, a generative image compression framework aiming for harmonizing machine and human vision in image compression. It ensures perceptual realism by leveraging generative priors and simultaneously guarantees semantic fidelity through the incorporation of Semantic Consistency loss (SC loss) during training.
Additionally, we introduce the Tag Guidance Module (TGM) that leverages highly semantic image-level tags to stimulate the pre-trained diffusion model's generative capabilities, requiring minimal additional bit rates. Consequently, Diff-ICMH supports multiple intelligent tasks through a single codec and bitstream without any task-specific adaptation, while preserving high-quality visual experience for human perception. Extensive experimental results demonstrate Diff-ICMH's superiority and generalizability across diverse tasks, while maintaining visual appeal for human perception. Ruoyu Feng 0001, Yunpeng Qi, Jinming Liu 0001, Xin Li 0082, Xin Jin 0014, Zhibo Chen 0001 |
NeurIPS | 1 |
| 2024 | CCEdit: Creative and Controllable Video Editing via Diffusion ModelsabstractIn this paper, we present CCEdit, a versatile generative video editing framework based on diffusion models. Our approach employs a novel trident network structure that separates structure and appearance control, ensuring precise and creative editing capabilities. Utilizing the foundational ControlNet architecture, we maintain the structural integrity of the video during editing. The incorporation of an additional appearance branch enables users to exert fine-grained control over the edited key frame. These two side branches seamlessly integrate into the main branch, which is constructed upon existing text-to-image (T2I) generation models, through learnable temporal layers. The versatility of our framework is demonstrated through a diverse range of choices in both structure representations and personalized T2I models, as well as the option to provide the edited key frame. To facilitate comprehensive evaluation, we introduce the BalanceCC benchmark dataset, comprising 100 videos and 4 target prompts for each video. Our extensive user studies compare CCEdit with eight state-of-the-art video editing methods. The outcomes demonstrate CCEdit's substantial superiority over all other methods. Ruoyu Feng 0001, Wenming Weng, Yuhui Yuan, Jianmin Bao, Chong Luo 0001, Zhibo Chen 0001, Baining Guo |
CVPR | 1 |
| 2024 | SeD: Semantic-Aware Discriminator for Image Super-ResolutionabstractGenerative Adversarial Networks (GANs) have been widely used to recover vivid textures in image super-resolution (SR) tasks. In particular, one discriminator is utilized to enable the SR network to learn the distribution of real-world high-quality images in an adversarial training manner. However, the distribution learning is overly coarse-grained, which is susceptible to virtual textures and causes counter-intuitive generation results. To mitigate this, we propose the simple and effective Semantic-aware Discriminator (denoted as SeD), which encourages the SR network to learn the fine-grained distributions by introducing the semantics of images as a condition. Concretely, we aim to excavate the semantics of images from a well-trained semantic extractor. Under different semantics, the discriminator is able to distinguish the real-fake images individually and adaptively, which guides the SR network to learn the more fine-grained semantic-aware textures. To obtain accurate and abundant semantics, we take full advantage of recently popular pretrained vision models (PVMs) with extensive datasets, and then incorporate its semantic features into the discriminator through a well-designed spatial cross-attention module. In this way, our proposed semantic-aware discriminator empowered the SR network to produce more photo-realistic and pleasing images. Extensive experiments on two typical tasks, i.e., SR and Real SR have demonstrated the effectiveness of our proposed methods. The code will be available at https://github.com/1bc12345/SeD. Bingchen Li 0001, Xin Li 0082, Hanxin Zhu, Yeying Jin, Ruoyu Feng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001 |
CVPR | 5 |
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 4 |
| 2024 | Rate-Distortion-Cognition Controllable Versatile Neural Image Compression
Jinming Liu 0001, Ruoyu Feng 0001, Yunpeng Qi, Qiuyu Chen, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
ECCV (56) | 2 |
| 2024 | Rethinking Domain Adaptation and Generalization in the ERA Of ClipabstractIn recent studies on domain adaptation, significant emphasis has been placed on the advancement of learning shared knowledge from a source domain to a target domain. Recently, the large vision-language pre-trained model (i.e., CLIP) has shown strong ability on zero-shot recognition, and parameter efficient tuning can further improve its performance on specific tasks. This work demonstrates that a simple domain prior boosts CLIP’s zero-shot recognition in a specific domain. Besides, CLIP’s adaptation relies less on source domain data due to its diverse pre-training dataset. Furthermore, we create a benchmark for zero-shot adaptation and pseudo-labeling based self-training with CLIP. Last but not least, we propose to improve the task generalization ability of CLIP from multiple unlabeled domains, which is a more practical and unique scenario. We believe our findings motivate a rethinking of domain adaptation benchmarks and the associated role of related algorithms in the era of CLIP. Ruoyu Feng 0001, Tao Yu 0012, Xin Jin 0014, Xiaoyuan Yu, Zhibo Chen 0001 |
ICIP | 1 |
| 2024 | Local Patch AutoAugment With Multi-Agent CollaborationabstractData augmentation (DA) plays a critical role in improving the generalization of deep learning models. Recent works on automatically searching for DA policies from data have achieved great success. However, existing automated DA methods generally perform the search at the image level, which limits the exploration of diversity in local regions. In this paper, we propose a more fine-grained automated DA approach, dubbed Patch AutoAugment, to divide an image into a grid of patches and search for the joint optimal augmentation policies for the patches. We formulate it as a multi-agent reinforcement learning (MARL) problem, where each agent learns an augmentation policy for each patch based on its content together with the semantics of the whole image. The agents cooperate with each other to achieve the optimal augmentation effect of the entire image by sharing a team reward. We show the effectiveness of our method on multiple benchmark datasets of image classification, fine-grained image recognition and object detection (e.g., CIFAR-10, CIFAR-100, ImageNet, CUB-200-2011, Stanford Cars, FGVC-Aircraft and Pascal VOC 2007). Extensive experiments demonstrate that our method outperforms the state-of-the-art DA methods while requiring fewer computational resources. Shiqi Lin, Tao Yu 0012, Ruoyu Feng 0001, Xin Li 0082, Xiaoyuan Yu, Zhibo Chen 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Semantically Structured Image Compression via Irregular Group-Based DecouplingabstractImage compression techniques typically focus on compressing rectangular images for human consumption, however, resulting in transmitting redundant content for downstream applications. To overcome this limitation, some previous works propose to semantically structure the bitstream, which can meet specific application requirements by selective transmission and reconstruction. Nevertheless, they divide the input image into multiple rectangular regions according to semantics and ignore avoiding information interaction among them, causing waste of bitrate and distorted reconstruction of region boundaries. In this paper, we propose to decouple an image into multiple groups with irregular shapes based on a customized group mask and compress them independently. Our group mask describes the image at a finer granularity, enabling significant bitrate saving by reducing the transmission of redundant content. Moreover, to ensure the fidelity of selective reconstruction, this paper proposes the concept of group-independent transform that maintain the independence among distinct groups. And we instantiate it by the proposed Group-Independent Swin-Block (GI Swin-Block). Experimental results demonstrate that our framework structures the bitstream with negligible cost, and exhibits superior performance on both visual quality and intelligent task supporting. Ruoyu Feng 0001, Xin Jin 0014, Runsen Feng, Zhibo Chen 0001 |
ICCV | 1 |
| 2023 | Composable Image Coding for Machine via Task-oriented Internal Adaptor and External PriorabstractTraditional image coding standards are typically optimized with a focus on human perception, which conflicts with the fact that most of the images are now analyzed by machines. To enable a variety of downstream intelligent tasks, contemporary approaches either utilize traditional codecs for image compression which are then used for task analysis, or develop a unified feature compression paradigm with deep learning techniques. However, they might suffer from accumulative errors and poor compatibility/generalization due to the conflict between standardized codecs and diverse machine tasks. We argue that a favorable image coding for machine (ICM) framework should have highly efficient adaptation capability, and take the ultimate task goals into account. Oriented at this, we propose a composable ICM solution dubbed Com-ICM, which develops plug-and-play lightweight internal adaptors injected into the codec architecture for efficient task transfer, and leverages off-the-shelf (large) models to provide external prior information for further task-oriented semantics learning. The internal adaptors (from the architectural aspect) and external priors (from the precondition aspect) complement each other, resulting in a mutually beneficial effect. We evaluate Com-ICM on diverse vision benchmarks, including image classification, object detection, and semantic segmentation, demonstrating its effectiveness and superiority. We are also actively submitting Com-ICM as a technical proposal to the international organization for standardization. Jinming Liu 0001, Xin Jin 0014, Ruoyu Feng 0001, Zhibo Chen 0001, Wenjun Zeng 0001 |
VCIP | 3 |
| 2023 | Image Coding for Machines based on Non-Uniform Importance AllocationabstractIn the Internet era, the explosive growth of media data processing poses significant challenges for the research of Image Coding for Machines (ICM) in improving the efficiency of AI models while reducing the burdens of data storage and transmission. Existing ICM methods face challenges in achieving sufficient generalization ability when developing a single codec to handle diverse downstream tasks. To address these issues, we propose a unified ICM framework that facilitates diverse downstream tasks with a novel importance allocation mechanism. Equipped with a spatially variable-rate image compression codec, we introduce two options: online updating and offline predicting the non-uniform quality map, which governs the quality distribution of reconstructed images based on specific downstream tasks. Our proposed method is rigorously evaluated through extensive experiments on diverse and comprehensive fine-grained image classification datasets. The experiment results conclusively demonstrate the effectiveness of the proposed method in achieving a superior rate-distortion trade-off for ICM. Yunpeng Qi, Ruoyu Feng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001 |
VCIP | 2 |
| 2023 | Semantical video coding: Instill static-dynamic clues into structured bitstream for AI tasks
Xin Jin 0014, Ruoyu Feng 0001, Simeng Sun, Runsen Feng, Tianyu He, Zhibo Chen 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Cloth-Changing Person Re-identification from A Single Image with Gait Prediction and RegularizationabstractCloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting, i.e., just from a single image, which enables an efficient and latency-free person identity matching for surveillance. Specifically, we introduce Gait recognition as an auxiliary task to drive the Image ReID model to learn cloth-agnostic representations by leveraging personal unique and cloth-independent gait information, we name this framework as GI-ReID. GI-ReID adopts a two-stream architecture that consists of an image ReID-Stream and an auxiliary gait recognition stream (Gait-Stream). The Gait-Stream, that is discarded in the inference for high efficiency, acts as a regulator to encourage the ReID-Stream to capture cloth-invariant biometric motion features during the training. To get temporal continuous motion cues from a single image, we design a Gait Sequence Prediction (GSP) module for Gait-Stream to enrich gait information. Finally, a semantics consistency constraint over two streams is enforced for effective knowledge regularization. Extensive experiments on multiple image-based Cloth-Changing ReID benchmarks, e.g., LTCC, PRCC, Real28, and VC-Clothes, demonstrate that GI-ReID performs favorably against the state-of-the-art methods. Xin Jin 0014, Tianyu He, Kecheng Zheng, Zhiheng Yin, Xu Shen 0001, Zhen Huang 0007, Ruoyu Feng 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001 |
CVPR | 7 |
| 2022 | Image Coding for Machines with Omnipotent Feature Learning
Ruoyu Feng 0001, Xin Jin 0014, Zongyu Guo, Runsen Feng, Tianyu He, Zhizheng Zhang 0004, Simeng Sun, Zhibo Chen 0001 |
ECCV (37) | 1 |