EDBT 2026 Demo / reviewers in the wild / expert
Yurui Ren
dblp:204/2612
· DBLP profile ↗
15ranked-venue papers
7as first author
7since 2021 · last 2023
0000-0003-0178-4460ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Flow-Guided Attention Deformation for Person Image GenerationabstractPose-guided person image generation aims to transfer reference images to target poses while preserving the source appearance. Recent approaches achieve considerable improvement by using spatial transformation modules such as attention operation. However, the commonly used vanilla attention tends to generate a dense correlation matrix which means that the value of a target position is the weighted sum of many source positions, resulting in blurry appearance. In this paper, we propose a novel model named Flow-guided Attention Deformation (FAD) to perform the spatial transformation. Our model first establishes the correlation between sources and targets with a flow-guided attention operation. Then, with the obtained correlation matrix, we perform an accurate deformation for source features to generate the predicted image. Extensive results demonstrate the superiority of the proposed method, outperforming state-of-the-art methods quantitatively and qualitatively. Ablation studies clarify the efficiency of the proposed modules and verify our hypothesis. Yubo Wu, Yurui Ren, Yuanqi Chen |
ICME | 2 |
| 2022 | Neural Texture Extraction and Distribution for Controllable Person Image SynthesisabstractWe deal with the controllable person image synthesis task which aims to re-render a human from a reference image with explicit control over body pose and appearance. Observing that person images are highly structured, we propose to generate desired images by extracting and distributing semantic entities of reference images. To achieve this goal, a neural texture extraction and distribution operation based on double attention is described. This operation first extracts semantic neural textures from reference feature maps. Then, it distributes the extracted neural textures according to the spatial distributions learned from target poses. Our model is trained to predict human images in arbitrary poses, which encourages it to extract disentangled and expressive neural textures representing the appearance of different semantic entities. The disentangled representation further enables explicit appearance control. Neural textures of different reference images can be fused to control the appearance of the interested areas. Experimental comparisons show the superiority of the proposed model. Code is available at https://github.com/RenYurui/Neural-Texture-Extraction-Distribution. Yurui Ren, Ge Li 0002, Shan Liu 0001, Thomas H. Li |
CVPR | 1 |
| 2022 | Flow-Based Point Cloud Completion Network with Adversarial RefinementabstractPoint cloud completion is the task of estimating the complete point cloud from the partial observation. Most of the existing methods tend to recover global shapes of 3D objects and usually lack local details. These methods rely merely on distance metrics between point sets as loss functions, which have the insufficient capability of supervising fine structures. In this work, we propose a coarse-to-fine approach to complete the partial point cloud with two stages: 1) Flow-based Completion Network, a principled probabilistic model that built on continuous normalizing flow to generate coarse completions conditioned on partial inputs. 2) Adversarial Refinement Network, a hierarchical refinement network constrained by the proposed patch discriminator to refine local details based on coarse completions. Experimental results show that our method can progressively complete 3D point clouds with fine details. Compared with other competitive methods, our method achieves better results on both quantitative and qualitative evaluations. Rong Bao, Yurui Ren, Ge Li 0002, Wei Gao 0003, Shan Liu 0001 |
ICASSP | 2 |
| 2022 | Context-Aware Hierarchical Transformer for Fine-Grained Video-Text RetrievalabstractVideo-Text Retrieval aims to perform accurate retrieval process that adopts texts to retrieve the corresponding videos, and vice versa. Typically, mainstream methods solve this problem by learning a common joint embedding space, and then measure the similarities between videos and texts. However, these methods lack the ability to represent detailed semantic information. Therefore, we first utilize three pre-trained models to construct the video embeddings of different semantic levels, and then propose a Context-aware Hierarchical Transformer (CHT) model to encode the context information between these levels. More specifically, our model builds fine-grained hierarchical video embeddings of three semantic levels: global, objects, and actions. Attention-based contextual transformers are utilized to establish the context interactions between different semantic levels. Experimental results on two benchmark video-text retrieval datasets demonstrate the superiority of our CHT model. Ablation studies also prove the effectiveness of our proposed model. Yurui Ren, Ge Li 0002 |
ICIP | 3 |
| 2022 | Deep Geometry Post-Processing for Decompressed Point CloudsabstractPoint cloud compression plays a crucial role in reducing the huge cost of data storage and transmission. However, distortions can be introduced into the decompressed point clouds due to quantization. In this paper, we propose a novel learning-based post-processing method to enhance the decompressed point clouds. Specifically, a voxelized point cloud is first divided into small cubes. Then, a 3D convolutional network is proposed to predict the occupancy probability for each location of a cube. We leverage both local and global contexts by generating multi-scale probabilities. These probabilities are progressively summed to predict the results in a coarse-to-fine manner. Finally, we obtain the geometry-refined point clouds based on the predicted probabilities. Different from previous methods, we deal with decompressed point clouds with huge variety of distortions using a single model. Experimental results show that the proposed method can significantly improve the quality of the decompressed point clouds, achieving 9.30dB BDPSNR gain on three representative datasets on average. Ge Li 0002, Dingquan Li, Yurui Ren, Wei Gao 0003, Thomas H. Li |
ICME | 4 |
| 2021 | PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingabstractGenerating portrait images by controlling the motions of existing faces is an important task of great consequence to social media industries. For easy use and intuitive control, semantically meaningful and fully disentangled parameters should be used as modifications. However, many existing techniques do not provide such fine-grained controls or use indirect editing methods i.e. mimic motions of other individuals. In this paper, a Portrait Image Neural Renderer (PIRenderer) is proposed to control the face motions with the parameters of three-dimensional morphable face models (3DMMs). The proposed model can generate photo-realistic portrait images with accurate movements according to intuitive modifications. Experiments on both direct and indirect editing tasks demonstrate the superiority of this model. Meanwhile, we further extend this model to tackle the audio-driven facial reenactment task by extracting sequential motions from audio inputs. We show that our model can generate coherent videos with convincing movements from only a single reference image and a driving audio stream. Our source code is available at https://github.com/RenYurui/PIRender. Yurui Ren, Ge Li 0002, Yuanqi Chen, Thomas H. Li, Shan Liu 0001 |
ICCV | 1 |
| 2021 | Combining Attention with Flow for Person Image SynthesisabstractPose-guided person image synthesis aims to synthesize person images by transforming reference images into target poses. In this paper, we observe that the commonly used spatial transformation blocks have complementary advantages. We propose a novel model by combining the attention operation with the flow-based operation. Our model not only takes the advantage of the attention operation to generate accurate target structures but also uses the flow-based operation to sample realistic source textures. Both objective and subjective experiments demonstrate the superiority of our model. Meanwhile, comprehensive ablation studies verify our hypotheses and show the efficacy of the proposed modules. Besides, additional experiments on the portrait image editing task demonstrate the versatility of the proposed combination. Yurui Ren, Yubo Wu, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ACM Multimedia | 1 |
| 2020 | Over-Exposure Correction via Exposure and Scene Information Disentanglement
Yuhui Cao, Yurui Ren, Thomas H. Li, Ge Li 0002 |
ACCV (4) | 2 |
| 2020 | Deep Image Spatial Transformation for Person Image GenerationabstractPose-guided person image generation is to transform a source person image to a target pose. This task requires spatial manipulations of source data. However, Convolutional Neural Networks are limited by the lack of ability to spatially transform the inputs. In this paper, we propose a differentiable global-flow local-attention framework to reassemble the inputs at the feature level. Specifically, our model first calculates the global correlations between sources and targets to predict flow fields. Then, the flowed local patch pairs are extracted from the feature maps to calculate the local attention coefficients. Finally, we warp the source features using a content-aware sampling method with the obtained local attention coefficients. The results of both subjective and objective experiments demonstrate the superiority of our model. Besides, additional results in video animation and view synthesis show that our model is applicable to other tasks requiring spatial transformation. Our source code is available at https://github.com/RenYurui/Global-Flow-Local-Attention. Yurui Ren, Xiaoming Yu, Thomas H. Li, Ge Li 0002 |
CVPR | 1 |
| 2020 | Context-aware Attention Network for Predicting Image Aesthetic SubjectivityabstractImage aesthetic assessment involves both fine-grained details and the holistic layout of images. However, most of current approaches learn the local and the holistic information separately, which has a potential loss of contextual information. Additionally, learning-based methods mainly cast image aesthetic assessment as a binary classification or a regression problem, which cannot sufficiently delineate the potential diversity of human aesthetic experience. To address these limitations, we attempt to render the contextual information and model the varieties of aesthetic experience. Specifically, we explore a context-aware attention module in two dimensions: hierarchical and spatial. The hierarchical context is introduced to present the concern of multi-level aesthetic details while the spatial context is served to yield the long-range perception of images. Based on the attention model, we predict the distribution of human aesthetic ratings of images, which reflects the diversity and similarity of human subjective opinions. We conduct extensive experiments on the prevailing AVA dataset to validate the effectiveness of our approach. Experimental results demonstrate that our approach achieves state-of-the-art results. Munan Xu, Jia-Xing Zhong, Yurui Ren, Shan Liu 0001, Ge Li 0002 |
ACM Multimedia | 3 |
| 2020 | Deep Spatial Transformation for Pose-Guided Person Image Generation and AnimationabstractPose-guided person image generation and animation aim to transform a source person image to target poses. These tasks require spatial manipulation of source data. However, Convolutional Neural Networks are limited by the lack of ability to spatially transform the inputs. In this paper, we propose a differentiable global-flow local-attention framework to reassemble the inputs at the feature level. This framework first estimates global flow fields between sources and targets. Then, corresponding local source feature patches are sampled with content-aware local attention coefficients. We show that our framework can spatially transform the inputs in an efficient manner. Meanwhile, we further model the temporal consistency for the person image animation task to generate coherent videos. The experiment results of both image generation and animation tasks demonstrate the superiority of our model. Besides, additional results of novel view synthesis and face image animation show that our model is applicable to other tasks requiring spatial transformation. The source code of our project is available at https://github.com/RenYurui/Global-Flow-Local-Attention. Yurui Ren, Ge Li 0002, Shan Liu 0001, Thomas H. Li |
IEEE Trans. Image Process. | 1 |
| 2019 | Base-detail image inpainting
Ruonan Zhang 0002, Yurui Ren, Jingfei Qiu, Ge Li 0002 |
BMVC | 2 |
| 2019 | StructureFlow: Image Inpainting via Structure-Aware Appearance FlowabstractImage inpainting techniques have shown significant improvements by using deep neural networks recently. However, most of them may either fail to reconstruct reasonable structures or restore fine-grained textures. In order to solve this problem, in this paper, we propose a two-stage model which splits the inpainting task into two parts: structure reconstruction and texture generation. In the first stage, edge-preserved smooth images are employed to train a structure reconstructor which completes the missing structures of the inputs. In the second stage, based on the reconstructed structures, a texture generator using appearance flow is designed to yield image details. Experiments on multiple publicly available datasets show the superior performance of the proposed network. Yurui Ren, Xiaoming Yu, Ruonan Zhang 0002, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ICCV | 1 |
| 2019 | LECARM: Low-Light Image Enhancement Using the Camera Response ModelabstractLow-light image enhancement algorithms can improve the visual quality of low-light images and support the extraction of valuable information for some computer vision techniques. However, existing techniques inevitably introduce color and lightness distortions when enhancing the images. To lower the distortions, we propose a novel enhancement framework using the response characteristics of cameras. First, we discuss how to determine a reasonable camera response model and its parameters. Then, we use the illumination estimation techniques to estimate the exposure ratio for each pixel. Finally, the selected camera response model is used to adjust each pixel to the desired exposure according to the estimated exposure ratio map. Experiments show that our method can obtain enhancement results with fewer color and lightness distortions compared with the several state-of-the-art methods. Yurui Ren, Zhenqiang Ying, Thomas H. Li, Ge Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | A New Image Contrast Enhancement Algorithm Using Exposure Fusion Framework
Zhenqiang Ying, Ge Li 0002, Yurui Ren, Ronggang Wang, Wenmin Wang 0001 |
CAIP (2) | 3 |