Runxing Liu

dblp:320/1946 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 50% 3D vision · 25% Robot navigation and mapping · 25%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Robotics › Robot navigation and mapping › spatial cognition › spatial knowledge
spatial semantics
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025
Visual content generation and editing
image generation
0.312025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models · ICCV 2025

Methods — techniques the papers use, named apart from their topics

token encoding ordering · 1.7spatial constraints · 1.7data engine · 1.7
YearPublicationVenuePosition
2025 CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
abstract
Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS.
Gaoyang Zhang, Bingtao Fu, Qingnan Fan, Qi Zhang 0029, Runxing Liu, Huaqi Zhang, Xinguo Liu
ICCV5