VLDB 2026 Research / reviewers in the wild / expert
Yuyu Guo 0001
dblp:205/3190-1
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-4376-6922ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Informative Scene Graph Generation via Debiasing
Lianli Gao, Xinyu Lyu, Yuyu Guo 0001, Yuan-Fang Li, Xu Lu 0004, Heng Tao Shen, Jingkuan Song |
Int. J. Comput. Vis. | 3 |
| 2023 | AMANet: Adaptive Multi-Path Aggregation for Learning Human 2D-3D CorrespondencesabstractLearning human 2D-3D correspondences aims to map all human 2D pixels to a 3D human template, namely human densepose estimation, involving surface patch recognition (i.e., Index-to-Patch (I)) and regression of patch-specific UV coordinates. Despite recent progress, it remains challenging especially under the condition of “in the wild”, where RGB images capture real-world scenes with backgrounds, occlusions, scale variations, and postural diversity. In this paper, we address three vital problems in this task: 1) how to perceive multi-scale visual information for instances “in the wild”; 2) how to design learning objectives to address the precise instance representation harassed by “multiple instances in one bounding box” phenomenon; and 3) how to boost the performance of index-to-patch prediction faced by limited supervision. To tackle problems above, we propose an end-to-end deep Adaptive Multi-path Aggregation network (AMA-net) for Human DensePose Estimation. First, we introduce an adaptive multi-path aggregation algorithm to extract varying-sized instance-level features, which capture multi-scale information of a bounding-box and are then utilized for parsing different instances. Second, we adopt an instance augmentation learning objective to further distinguish the target instance from other interference instances. Third, taking advantage of 2D human parsers that are trained from sufficient annotations, we introduce a task transformer that bridges the “gap” between 2D human parsing and densepose estimation, thus benefiting the performance of densepose estimator. Experimental results on the challenging DensePose-COCO dataset demonstrate that our approach sets a new record, and it significantly outperforms the state-of-the-art methods. Codes and models are publicly available. Xuanhan Wang, Yuyu Guo 0001, Jingkuan Song, Lianli Gao, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2022 | Fine-Grained Predicates Learning for Scene Graph GenerationabstractThe performance of current Scene Graph Generation models is severely hampered by some hard-to-distinguish predicates, e.g., “woman-on/standing on/walking on-beach” or “woman-near/looking at/in front of-child”. While general SGG models are prone to predict head predicates and existing re-balancing strategies prefer tail categories, none of them can appropriately handle these hard-to-distinguish predicates. To tackle this issue, inspired by fine-grained image classification, which focuses on differentiating among hard-to-distinguish object classes, we propose a method named Fine-Grained Predicates Learning (FGPL) which aims at differentiating among hard-to-distinguish predicates for Scene Graph Generation task. Specifically, we first introduce a Predicate Lattice that helps SGG models to figure out fine-grained predicate pairs. Then, utilizing the Predicate Lattice, we propose a Category Discriminating Loss and an Entity Discriminating Loss, which both contribute to distinguishing fine-grained predicates while maintaining learned discriminatory power over recognizable ones. The proposed model-agnostic strategy significantly boosts the performances of three benchmark models (Transformer, VCTree, and Motif) by 22.8%, 24.1% and 21.7% of Mean Recall (mR@100) on the Predicate Classification sub-task, respectively. Our model also outperforms state-of-the-art methods by a large margin (i.e., 6.1%, 4.6%, and 3.2% of Mean Recall (mR@100)) on the Visual Genome dataset. Codes are publicly available11https://github.com/XinyuLyu/FGPL. Xinyu Lyu, Lianli Gao, Yuyu Guo 0001, Zhou Zhao 0001, Heng Tao Shen, Jingkuan Song |
CVPR | 3 |
| 2022 | Multi-Scale Graph Attention Network for Scene Graph GenerationabstractScene graph provides a high-level scene understanding of the image, which has a wide range of applications in computer vision. Previous methods elaborately design many message passing strategies and uniformly treat instances in the image to capture contextual information. These methods, however, fail to grasp the salient objects and their relations, which are the basis of understanding the content of images. To capture the interaction among salient instances, we propose a novel Multi-Scale Graph Attention Network (MSGAT) that gradually shrinks the graph scale to retain salient instances, and then expands it to encode the multi-scale context. Our proposed MSGAT contains two sub-modules: Multi-Scale Message Passing (MSMP) and Relationship Filtering Module (RFM), which are designed to enhance features of salient instances and filter redundant relationships, respectively. Extensive experiments demonstrate that MSGAT outperforms previous methods and achieves state-of-the-art performances on Visual Genome. Xinyu Lyu, Yuyu Guo 0001, Lianli Gao, Jingkuan Song |
ICME | 3 |
| 2022 | Learning to Generate Scene Graph from Head to TailabstractScene Graph Generation (SGG) represents objects and their interactions with a graph structure. Recently, many works are devoted to solving the imbalanced problem in SGG. However, underestimating the head predicates in the whole training process, they wreck the features of head predicates that provide general features for tail ones. Besides, assigning excessive attention to the tail predicates leads to semantic deviation. Based on this, we propose a novel SGG framework, learning to generate scene graphs from Head to Tail (SGG-HT), containing Curriculum Re-weight Mechanism (CRM) and Semantic Context Module (SCM). CRM learns head/easy samples firstly for robust features of head predicates and then gradually focuses on tail/hard ones. SCM is proposed to relieve semantic deviation by ensuring the semantic consistency between the generated scene graph and the ground truth in global and local representations. Experiments show that SGG-HT significantly alleviates the biased problem and achieves state-of-the-art performances on Visual Genome. Chaofan Zheng, Xinyu Lyu, Yuyu Guo 0001, Pengpeng Zeng, Jingkuan Song, Lianli Gao |
ICME | 3 |
| 2022 | Dynamic Scene Graph Generation via Temporal Prior InferenceabstractReal-world videos are composed of complex actions with inherent temporal continuity (eg "person-touching-bottle" is usually followed by "person-holding-bottle"). In this work, we propose a novel method to mine such temporal continuity for dynamic scene graph generation (DSGG), namely Temporal Prior Inference (TPI). As opposed to current DSGG methods, which individually capture the temporal dependence of each video by refining representations, we make the first attempt to explore the temporal continuity by extracting the entire co-occurrence patterns of action categories from a variety of videos in Action Genome (AG) dataset. Then, these inherent patterns are organized as Temporal Prior Knowledge (TPK) which serves as prior knowledge for models' learning and inference. Furthermore, given the prior knowledge, human-object relationships in current frames can be effectively inferred from adjacent frames via the robust Temporal Prior Inference algorithm with tiny computation cost. Specifically, to efficiently guide the generating of temporal-consistent dynamic scene graphs, we incorporate the temporal prior inference into a DSGG framework by introducing frame enhancement, continuity loss, and fast inference. The proposed model-agnostic strategies significantly boost the performances of existing state-of-the-art models on the Action Genome dataset, achieving 69.7 and 72.6 for [email protected] and [email protected] on PredCLS. In addition, the inference speed can be significantly reduced by 41% with an acceptable drop on [email protected] (69.7 to 66.8) by utilizing fast inference. Lianli Gao, Xinyu Lyu, Yuyu Guo 0001, Pengpeng Zeng, Jingkuan Song |
ACM Multimedia | 4 |
| 2022 | Relation Regularized Scene Graph GenerationabstractScene graph generation (SGG) is built on top of detected objects to predict object pairwise visual relations for describing the image content abstraction. Existing works have revealed that if the links between objects are given as prior knowledge, the performance of SGG is significantly improved. Inspired by this observation, in this article, we propose a relation regularized network (R2-Net), which can predict whether there is a relationship between two objects and encode this relation into object feature refinement and better SGG. Specifically, we first construct an affinity matrix among detected objects to represent the probability of a relationship between two objects. Graph convolution networks (GCNs) over this relation affinity matrix are then used as object encoders, producing relation-regularized representations of objects. With these relation-regularized features, our R2-Net can effectively refine object labels and generate scene graphs. Extensive experiments are conducted on the visual genome dataset for three SGG tasks (i.e., predicate classification, scene graph classification, and scene graph detection), demonstrating the effectiveness of our proposed method. Ablation studies also verify the key roles of our proposed components in performance improvement. Yuyu Guo 0001, Lianli Gao, Jingkuan Song, Peng Wang 0023, Nicu Sebe, Heng Tao Shen, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | From General to Specific: Informative Scene Graph Generation via Balance AdjustmentabstractThe scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene understanding. However, current models are stuck in common predicates, e.g., "on" and "at", rather than informative ones, e.g., "standing on" and "looking at", resulting in the loss of precise information and overall performance. If a model only uses "stone on road" rather than "blocking" to describe an image, it is easy to misunderstand the scene. We argue that this phenomenon is caused by two key imbalances between informative predicates and common ones, i.e., semantic space level imbalance and training sample level imbalance. To tackle this problem, we propose BA-SGG, a simple yet effective SGG framework based on balance adjustment but not the conventional distribution fitting. It integrates two components: Semantic Adjustment (SA) and Balanced Predicate Learning (BPL), respectively for adjusting these imbalances. Benefited from the model-agnostic process, our method is easily applied to the state-of-the-art SGG models and significantly improves the SGG performance. Our method achieves 14.3%, 8.0%, and 6.1% higher Mean Recall (mR) than that of the Transformer model at three scene graph generation sub-tasks on Visual Genome, respectively. Codes are publicly available1. Yuyu Guo 0001, Lianli Gao, Xuanhan Wang, Xing Xu 0001, Xu Lu 0004, Heng Tao Shen, Jingkuan Song |
ICCV | 1 |
| 2021 | SKANet: Structured Knowledge-Aware Network for Visual DialogabstractVisual dialog aims to generate an answer to each question based on an image and dialog history. Despite recent progress, existing methods still undergo degradation on the condition of complex scenarios. Handling these scenarios depends on logical reasoning that requires common sense priors. In this paper, we propose a novel visual dialog pipeline, named Structured Knowledge-Aware Network (SKANet), consisting of a Multi-Modality Fusion Module, an Image Knowledge-Aware Module, and a Caption Knowledge-Aware Module. The Multi-Modality Fusion Module explores the textual context about the dialog history and visual content. To deal with the complex scenarios, the Image and Caption Knowledge-Aware Modules construct common sense knowledge graphs from ConceptNet. Experimental results on the VisDial v1.0 dataset show that our proposed method effectively outperforms comparative methods. Lei Zhao 0017, Lianli Gao, Yuyu Guo 0001, Jingkuan Song, Heng Tao Shen |
ICME | 3 |
| 2020 | One-shot Scene Graph GenerationabstractAs a structured representation of the image content, the visual scene graph (visual relationship) acts as a bridge between computer vision and natural language processing. Existing models on the scene graph generation task notoriously require tens or hundreds of labeled samples. By contrast, human beings can learn visual relationships from a few or even one example. Inspired by this, we design a task named One-Shot Scene Graph Generation, where each relationship triplet (e.g., "dog-has-head'') comes from only one labeled example. The key insight is that rather than learning from scratch, one can utilize rich prior knowledge. In this paper, we propose Multiple Structured Knowledge (Relational Knowledge and Commonsense Knowledge) for the one-shot scene graph generation task. Specifically, the Relational Knowledge represents the prior knowledge of relationships between entities extracted from the visual content, e.g., the visual relationships "standing in'', "sitting in'', and "lying in'' may exist between "dog'' and "yard'', while the Commonsense Knowledge encodes "sense-making'' knowledge like "dog can guard yard''. By organizing these two kinds of knowledge in a graph structure, Graph Convolution Networks (GCNs) are used to extract knowledge-embedded semantic features of the entities. Besides, instead of extracting isolated visual features from each entity generated by Faster R-CNN, we utilize an Instance Relation Transformer encoder to fully explore their context information. Based on a constructed one-shot dataset, the experimental results show that our method significantly outperforms existing state-of-the-art methods by a large margin. Ablation studies also verify the effectiveness of the Instance Relation Transformer encoder and the Multiple Structured Knowledge. Yuyu Guo 0001, Jingkuan Song, Lianli Gao, Heng Tao Shen |
ACM Multimedia | 1 |
| 2019 | Adaptive Multi-Path Aggregation for Human DensePose Estimation in the WildabstractDense human pose "in the wild'' task aims to map all 2D pixels of the detected human body to a 3D surface by establishing surface correspondences, i.e., surface patch index and part-specific UV coordinates. It remains challenging especially under the condition of "in the wild'', where RGB images capture complex, real-world scenes with background, occlusions, scale variations, and postural diversity. In this paper, we propose an end-to-end deep Adaptive Multi-path Aggregation network (AMA-net) for Dense Human Pose Estimation. In the proposed framework, we address two main problems: 1) how to design a simple yet effective pipeline for supporting distinct sub-tasks (e.g., instance segmentation, body part segmentation, and UV estimation); and 2) how to equip this pipeline with the ability of handling "in the wild''. To solve these problems, we first extend FPN by adding a branch for mapping 2D pixels to a 3D surface in parallel with the existing branch for bounding box detection. Then, in AMA-net, we extract variable-sized object-level feature maps (e.g., 7×7, 14×14, and 28×28), named multi-path, from multi-layer feature maps, which capture rich information of objects and are then adaptively utilized in different tasks. AMA-net is simple to train and adds only a small overhead to FPN. We discover that aside from the deep feature map, Adaptive Multi-path Aggregation is of particular importance for improving the accuracy of dense human pose estimation "in the wild''. The experimental results on the challenging Dense-COCO dataset demonstrate that our approach sets a new record for Dense Human Pose Estimation task, and it significantly outperforms the state-of-the-art methods. Our code: \urlhttps://github.com/nobody-g/AMA-net. Yuyu Guo 0001, Lianli Gao, Jingkuan Song, Peng Wang 0023, Wuyuan Xie, Heng Tao Shen |
ACM Multimedia | 1 |
| 2019 | From Deterministic to Generative: Multimodal Stochastic RNNs for Video CaptioningabstractVideo captioning, in essential, is a complex natural process, which is affected by various uncertainties stemming from video content, subjective judgment, and so on. In this paper, we build on the recent progress in using encoder-decoder framework for video captioning and address what we find to be a critical deficiency of the existing methods that most of the decoders propagate deterministic hidden states. Such complex uncertainty cannot be modeled efficiently by the deterministic models. In this paper, we propose a generative approach, referred to as multimodal stochastic recurrent neural networks (MS-RNNs), which models the uncertainty observed in the data using latent stochastic variables. Therefore, MS-RNN can improve the performance of video captioning and generate multiple sentences to describe a video considering different random factors. Specifically, a multimodal long short-term memory (LSTM) is first proposed to interact with both visual and textual features to capture a high-level representation. Then, a backward stochastic LSTM is proposed to support uncertainty propagation by introducing latent variables. Experimental results on the challenging data sets, microsoft video description and microsoft research video-to-text, show that our proposed MS-RNN approach outperforms the state-of-the-art video captioning benchmarks. Jingkuan Song, Yuyu Guo 0001, Lianli Gao, Xuelong Li 0001, Alan Hanjalic, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Exploiting long-term temporal dynamics for video captioning
Yuyu Guo 0001, Jingqiu Zhang, Lianli Gao |
World Wide Web | 1 |