Gaoyang Zhang

dblp:345/2903 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 50% 3D vision · 25% Robot navigation and mapping · 25%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Robotics › Robot navigation and mapping › spatial cognition › spatial knowledge
spatial semantics
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Visual content generation and editing
image generation
0.312025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025

Methods — techniques the papers use, named apart from their topics

token encoding ordering · 1.7spatial constraints · 1.7data engine · 1.7
YearPublicationVenuePosition
2025 CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
abstract
Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS.
Gaoyang Zhang, Bingtao Fu, Qingnan Fan, Qi Zhang 0029, Runxing Liu, Huaqi Zhang, Xinguo Liu
ICCV1
2025 Efficient RGB-D scene understanding via multi-task adaptive learning and cross-dimensional feature guidance
Guodong Sun 0002, Gaoyang Zhang, Yang Zhang 0053
Knowl. Based Syst.3
2024 Point Cloud Segmentation with Guided Sampling and Continuous Interpolation
Gaoyang Zhang, Xinguo Liu
CVM (1)1
2023 Knowledge Graph based Explainable Question Retrieval for Programming Tasks
abstract
Developers often seek solutions for their programming problems by retrieving existing questions on technical Q&A sites such as Stack Overflow. In many cases, they fail to find relevant questions due to the knowledge gap between the questions and the queries or feel it hard to choose the desired questions from the returned results due to the lack of explanations about the relevance. In this paper, we propose KGXQR, a knowledge graph based explainable question retrieval approach for programming tasks. It uses BERT-based sentence similarity to retrieve candidate Stack Overflow questions that are relevant to a given query. To bridge the knowledge gap and enhance the performance of question retrieval, it constructs a software development related concept knowledge graph and trains a question relevance prediction model to re-rank the candidate questions. The model is trained based on a combined sentence representation of BERT-based sentence embedding and graph-based concept embedding. To help understand the relevance of the returned Stack Overflow questions, KGXQR further generates explanations based on the association paths between the concepts involved in the query and the Stack Overflow questions. The evaluation shows that KGXQR outperforms the baselines in terms of accuracy, recall, MRR, and MAP and the generated explanations help the users to find the desired questions faster and more accurately.
Mingwei Liu 0002, Simin Yu, Xin Peng 0001, Xueying Du, Tianyong Yang, Huanjun Xu, Gaoyang Zhang
ICSME7