VLDB 2026 Research / reviewers in the wild / expert
Yuren Cong
dblp:256/4899 · also Cong Yuren
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-7505-8563ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Generative modeling · 43% Vision and language · 15% Segmentation and scene understanding · 14% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.8 | 3 | 2025 | FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing · ICLR 2024 GenTron: Diffusion Transformers for Image and Video Generation · CVPR 2024 Learning Flow Fields in Attention for Controllable Person Image Generation · CVPR 2025 |
Computer vision › Segmentation and scene understanding
scene graph generation |
1.7 | 2 | 2026 | SPAN: Learning Similarity Between Scene Graphs and Images With Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2026 RelTR: Relation Transformer for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
1.6 | 2 | 2025 | Attribute-Centric Compositional Text-to-Image Generation · Int. J. Comput. Vis. 2025 GenTron: Diffusion Transformers for Image and Video Generation · CVPR 2024 |
Computer vision › Vision and language
cross-modal retrieval |
1.0 | 1 | 2026 | SPAN: Learning Similarity Between Scene Graphs and Images With Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › Image recognition and object detection
image retrieval |
1.0 | 1 | 2026 | SPAN: Learning Similarity Between Scene Graphs and Images With Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Generative modeling › diffusion model › controllable generation
attribute-controlled generation |
0.9 | 1 | 2025 | Attribute-Centric Compositional Text-to-Image Generation · Int. J. Comput. Vis. 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
compositional text-to-image generation |
0.9 | 1 | 2025 | Attribute-Centric Compositional Text-to-Image Generation · Int. J. Comput. Vis. 2025 |
Visual content generation and editing › image generation
person image generation |
0.9 | 1 | 2025 | Learning Flow Fields in Attention for Controllable Person Image Generation · CVPR 2025 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.8 | 1 | 2024 | GenTron: Diffusion Transformers for Image and Video Generation · CVPR 2024 |
Visual content generation and editing › video editing
temporally consistent editing |
0.8 | 1 | 2024 | FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing · ICLR 2024 |
Visual content generation and editing › video editing
text-to-video editing |
0.8 | 1 | 2024 | FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing · ICLR 2024 |
Machine learning › Deep learning architectures and training › transformer
relation-enhanced transformer |
0.7 | 1 | 2023 | RelTR: Relation Transformer for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision › geometric deep learning › set learning
set prediction |
0.7 | 1 | 2023 | RelTR: Relation Transformer for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
video scene graph generation |
0.5 | 1 | 2021 | Spatial-Temporal Transformer for Dynamic Scene Graph Generation · ICCV 2021 |
Computer vision › Video understanding and tracking › dynamic scene analysis
video scene understanding |
0.5 | 1 | 2021 | Spatial-Temporal Transformer for Dynamic Scene Graph Generation · ICCV 2021 |
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations |
0.4 | 1 | 2020 | NODIS: Neural Ordinary Differential Scene Understanding · ECCV (20) 2020 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.4 | 1 | 2020 | NODIS: Neural Ordinary Differential Scene Understanding · ECCV (20) 2020 |
Machine learning › Generative modeling › diffusion model
controllable generation |
0.3 | 1 | 2025 | Learning Flow Fields in Attention for Controllable Person Image Generation · CVPR 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.3 | 1 | 2025 | Attribute-Centric Compositional Text-to-Image Generation · Int. J. Comput. Vis. 2025 |
Computer vision › Vision and language
visual relationship detection |
0.1 | 1 | 2021 | Spatial-Temporal Transformer for Dynamic Scene Graph Generation · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.6attention mechanism · 2.2contrastive learning · 1.9flow field · 1.7attention regularization · 1.7graph transformer · 1.0graph serialization · 1.0feature augmentation · 0.9optical flow · 0.8motion-free guidance · 0.8diffusion transformer · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPAN: Learning Similarity Between Scene Graphs and Images With TransformersabstractLearning similarity between scene graphs and images aims to estimate a similarity score given a scene graph and an image. There is currently no research dedicated to this task, although it is critical for scene graph generation and downstream applications. Scene graph generation is conventionally evaluated by Recall$@K$@K and mean Recall$@K$@K, which measure the ratio of predicted triplets that appear in the human-labeled triplet set. However, such triplet-oriented metrics fail to demonstrate the overall semantic difference between a scene graph and an image and are sensitive to annotation bias and noise. Using generated scene graphs in the downstream applications is therefore limited. To address this issue, for the first time, we propose a Scene graPh-imAge coNtrastive learning framework, SPAN, that can measure the similarity between scene graphs and images. Our novel framework consists of a graph Transformer and an image Transformer to align scene graphs and their corresponding images in the shared latent space. We introduce a novel graph serialization technique that transforms a scene graph into a sequence with structural encodings. Based on our framework, we propose R-Precision measuring image retrieval accuracy as a new evaluation metric for scene graph generation. We establish new benchmarks on the Visual Genome and Open Images datasets. Extensive experiments are conducted to verify the effectiveness of SPAN, which shows great potential as a scene graph encoder. Yuren Cong, Wentong Liao, Bodo Rosenhahn, Michael Ying Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Learning Flow Fields in Attention for Controllable Person Image GenerationabstractControllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person’s appearance or pose. However, prior methods often distort fine-grained details from the reference image, despite achieving high overall image quality. We attribute these distortions to inadequate attention to corresponding regions in the reference image. To address this, we thereby propose learning flow fields in attention (Leffa), which explicitly guides the target query to attend to the correct reference key in the attention layer during training. Specifically, it is realized via a regularization loss on top of the attention map within a diffusionbased baseline. Our extensive experiments show that Leffa achieves state-of-the-art performance in controlling appearance and pose, significantly reducing fine-grained detail distortion while maintaining high image quality. Additionally, we show that our loss is model-agnostic and can be used to improve the performance of other diffusion models. Zijian Zhou 0002, Shikun Liu, Kam Woh Ng, Tian Xie 0003, Yuren Cong, Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Aditya Patel, Tao Xiang 0002, Miaojing Shi, Sen He 0001 |
CVPR | 7 |
| 2025 | Attribute-Centric Compositional Text-to-Image GenerationabstractAbstract Despite the recent impressive breakthroughs in text-to-image generation, generative models have difficulty in capturing the data distribution of underrepresented attribute compositions while over-memorizing overrepresented attribute compositions, which raises public concerns about their robustness and fairness. To tackle this challenge, we propose ACTIG, an attribute-centric compositional text-to-image generation framework. We present an attribute-centric feature augmentation and a novel image-free training scheme, which greatly improves model’s ability to generate images with underrepresented attributes. We further propose an attribute-centric contrastive loss to avoid overfitting to overrepresented attribute compositions. We validate our framework on the CelebA-HQ and CUB datasets. Extensive experiments show that the compositional generalization of ACTIG is outstanding, and our framework outperforms previous works in terms of image quality and text-image consistency. The source code and trained models are publicly available at https://github.com/yrcong/ACTIG . Yuren Cong, Martin Renqiang Min, Li Erran Li, Bodo Rosenhahn, Michael Ying Yang |
Int. J. Comput. Vis. | 1 |
| 2024 | GenTron: Diffusion Transformers for Image and Video GenerationabstractIn this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexibility and scalability, the visual generative domain primarily utilizes CNN-based U-Net architectures, particularly in diffusion-based models. We introduce GenTron, a family of Generative models employing Transformer-based diffusion, to address this gap. Our initial step was to adapt Diffusion Transformers (DiTs) from class to text conditioning, a process involving thorough empirical exploration of the conditioning mechanism. We then scale GenTron from approximately 900M to over 3B parameters, observing improvements in visual quality. Furthermore, we extend GenTron to text-to-video generation, incorporating novel motion-free guidance to enhance video quality. In human evaluations against SDXL, GenTron achieves a 51.1% win rate in visual quality (with a 19.8% draw rate), and a 42.3% win rate in text alignment (with a 42.9% draw rate). GenTron notably performs well in T2I-CompBench, highlighting its compositional generation ability. We hope GenTron could provide meaningful insights and serve as a valuable reference for future research. Please refer to the website11https://www.shoufachen.com/gentron_website/ and the arXiv version for the most up-to-date results: https://arxiv.org/abs/2312.04557. Shoufa Chen, Mengmeng Xu 0006, Jiawei Ren 0001, Yuren Cong, Sen He 0001, Yanping Xie, Animesh Sinha, Ping Luo 0002, Tao Xiang 0002, Juan-Manuel Pérez-Rúa |
CVPR | 4 |
| 2024 | Segment Any Object Model (SAOM): Real-To-Simulation Fine-Tuning Strategy For Multi-Class Multi-Instance SegmentationabstractMulti-class multi-instance segmentation is the task of identifying masks for multiple object classes and multiple instances of the same class within an image. The foundational Segment Anything Model (SAM) is designed for promptable multi-class multi-instance segmentation but tends to output part or sub-part masks in the “everything” mode for various real-world applications. Whole object segmentation masks play a crucial role for indoor scene understanding, especially in robotics applications. We propose a new domain invariant Real-to-Simulation (Real-Sim) fine-tuning strategy for SAM. We use object images and ground truth data collected from Ai2Thor simulator during fine-tuning (real-to-sim). To allow our Segment Any Object Model (SAOM) to work in the “everything” mode, we propose the novel nearest neighbour assignment method, updating point embeddings for each ground-truth mask. SAOM is evaluated on our own dataset collected from Ai2Thor simulator. SAOM significantly improves on SAM, with a $28 \%$ increase in mIoU and a $25 \%$ increase in mAcc for 54 frequently-seen indoor object classes. Moreover, our Real-to-Simulation fine-tuning strategy demonstrates promising generalization performance in real environments without being trained on the real-world data (sim-to-real). The dataset and the code are available here. Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, Jumana M. Abu-Khalaf, David Suter |
ICIP | 3 |
| 2024 | FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingabstractText-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts.
A major challenge in this task is to ensure that all frames in the edited video are visually consistent.
Most recent works apply advanced text-to-image diffusion models to this task by inflating 2D spatial attention in the U-Net into spatio-temporal attention.
Although temporal context can be added through spatio-temporal attention, it may introduce some irrelevant information for each patch and therefore cause inconsistency in the edited video.
In this paper, for the first time, we introduce optical flow into the attention module in diffusion model's U-Net to address the inconsistency issue for text-to-video editing.
Our method, FLATTEN, enforces the patches on the same flow path across different frames to attend to each other in the attention module, thus improving the visual consistency in the edited videos.
Additionally, our method is training-free and can be seamlessly integrated into any diffusion based text-to-video editing methods and improve their visual consistency.
Experiment results on existing text-to-video editing benchmarks show that our proposed method achieves the new state-of-the-art performance. In particular, our method excels in maintaining the visual consistency in the edited videos. Yuren Cong, Mengmeng Xu 0006, Christian Simon, Shoufa Chen, Jiawei Ren 0001, Yanping Xie, Juan-Manuel Pérez-Rúa, Bodo Rosenhahn, Tao Xiang 0002, Sen He 0001 |
ICLR | 1 |
| 2024 | Worldafford: Affordance Grounding Based on Natural Language InstructionsabstractAffordance grounding aims to localize the interaction regions for the manipulated objects in the scene image according to given instructions, which is essential for Embodied AI and manipulation tasks. A key challenge in affordance grounding is enabling the agent to understand human instructions, identify usable tools in the environment, and determine how to use them to complete the task. Most recent works primarily support simple action labels as input instructions for localizing affordance regions, failing to capture complex human objectives. Moreover, these approaches typically identify affordance regions of only a single object in object-centric images, ignoring the object context and struggling to localize affordance regions of multiple objects in complex scenes for practical applications. To address this concern, for the first time, we introduce a new task of affordance grounding based on natural language instructions, extending it from previously using simple labels for complex human instructions. For this new task, we propose a new framework, WorldAfford. We design a novel Affordance Reasoning Chain-of-Thought Prompting to reason about affordance knowledge from LLMs more precisely and logically. Subsequently, we use SAM and CLIP to localize the objects related to the affordance knowledge in the image. We identify the affordance regions of the objects through an affordance region localization module. To benchmark this new task and validate our framework, an affordance grounding dataset, LLMaFF, is constructed. We conduct extensive experiments to verify that WorldAfford performs state-of-the-art on the previous AGD20K and the new LLMaFF dataset. In particular, WorldAfford can localize the affordance regions of multiple objects and provide an alternative when objects in the environment cannot fully match the given instruction. Our Project page: https://worldafford.github.io/. Changmao Chen, Yuren Cong, Zhen Kan |
ICTAI | 2 |
| 2024 | Indoor Scene Change Understanding (SCU): Segment, Describe, and Revert Any ChangeabstractUnderstanding of scene changes is crucial for embodied AI applications, such as visual room rearrangement, where the agent must revert changes by restoring the objects to their original locations or states. Visual changes between two scenes, pre- and post-rearrangement, encompass two tasks: scene change detection (locating changes) and image difference captioning (describing changes). While previous methods, focused on sequential 2D images, have addressed these tasks separately, it is essential to emphasize the significance of their combination. Therefore, we propose a new Scene Change Understanding (SCU) task for simultaneous change detection and description. Moreover, we go beyond change language description generation and aim to generate rearrangement instructions for the robotic agent to revert changes. To solve this task, we propose a novel method - EmbSCU, which allows to compare instance-level change object masks (for 53 frequently-seen indoor object classes) before and after changes and generate rearrangement language instructions for the agent. EmbSCU is built on our Segment Any Object Model (SAOMv2) - a fine-tuned version of Segment Anything Model (SAM), adapted to obtain instance-level object masks for both foreground and background objects in indoor embodied environments. EmbSCU is evaluated on our own dataset of sequential 2D image pairs before and after changes, collected from the Ai2Thor simulator. The proposed framework achieves promising results in both change detection and change description. Moreover, EmbSCU demonstrates positive generalization results on real-world scenes without using any real-life data during training. The dataset and the code are available here. Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, David Suter, Jumana M. Abu-Khalaf |
IROS | 3 |
| 2023 | RelTR: Relation Transformer for Scene Graph GenerationabstractDifferent objects in the same scene are more or less related to each other, but only a limited number of these relationships are noteworthy. Inspired by Detection Transformer, which excels in object detection, we view scene graph generation as a set prediction problem. In this article, we propose an end-to-end scene graph generation model Relation Transformer (RelTR), which has an encoder-decoder architecture. The encoder reasons about the visual feature context while the decoder infers a fixed-size set of triplets subject-predicate-object using different types of attention mechanisms with coupled subject and object queries. We design a set prediction loss performing the matching between the ground truth and predicted triplets for the end-to-end training. In contrast to most existing scene graph generation methods, RelTR is a one-stage method that predicts sparse scene graphs directly only using visual appearance without combining entities and labeling all possible predicates. Extensive experiments on the Visual Genome, Open Images V6, and VRD datasets demonstrate the superior performance and fast inference of our model. Yuren Cong, Michael Ying Yang, Bodo Rosenhahn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Spatial-Temporal Transformer for Dynamic Scene Graph GenerationabstractDynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic interpretation. In this paper, we propose Spatial-temporal Transformer (STTran), a neural network that consists of two core modules: (1) a spatial encoder that takes an input frame to extract spatial context and reason about the visual relationships within a frame, and (2) a temporal decoder which takes the output of the spatial encoder as input in order to capture the temporal dependencies between frames and infer the dynamic relationships. Furthermore, STTran is flexible to take varying lengths of videos as input without clipping, which is especially important for long videos. Our method is validated on the benchmark dataset Action Genome (AG). The experimental results demonstrate the superior performance of our method in terms of dynamic scene graphs. Moreover, a set of ablative studies is conducted and the effect of each proposed module is justified. Code available at: https://github.com/yrcong/STTran. Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, Michael Ying Yang |
ICCV | 1 |
| 2020 | NODIS: Neural Ordinary Differential Scene Understanding
Yuren Cong, Hanno Ackermann, Wentong Liao, Michael Ying Yang, Bodo Rosenhahn |
ECCV (20) | 1 |