Qiushan Guo

dblp:231/1814 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Generative modeling · 28% Vision and language · 20% Efficient and distributed learning · 18%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 28 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.232025
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception · NeurIPS 2025
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths · NeurIPS 2023
EGC: Image Generation and Classification via a Diffusion Energy-Based Model · ICCV 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.632022
Rethinking Resolution in the Context of Efficient Video Recognition · NeurIPS 2022
Scale-Equivalent Distillation for Semi-Supervised Object Detection · CVPR 2022
Online Knowledge Distillation via Collaborative Learning · CVPR 2020
Computer vision › Vision and language
vision-language model
1.522024
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models · NeurIPS 2024
RegionGPT: Towards Region Understanding Vision Language Model · CVPR 2024
Machine learning › Generative modeling
autoregressive model
0.912025
OmniGen-AR: AutoRegressive Any-to-Image Generation · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception · NeurIPS 2025
Visual content generation and editing › image generation
conditional image synthesis
0.912025
OmniGen-AR: AutoRegressive Any-to-Image Generation · NeurIPS 2025
Visual content generation and editing
video generation
0.912025
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception · NeurIPS 2025
Computer vision › 3D vision
3d scene understanding
0.812024
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models · NeurIPS 2024
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.812024
RegionGPT: Towards Region Understanding Vision Language Model · CVPR 2024
Computer vision › Vision and language › image captioning › grounded image captioning
region captioning
0.812024
RegionGPT: Towards Region Understanding Vision Language Model · CVPR 2024
Computer vision › Vision and language
region-level understanding
0.812024
RegionGPT: Towards Region Understanding Vision Language Model · CVPR 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.812024
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models · NeurIPS 2024
Machine learning › Generative modeling
energy-based model
0.712023
EGC: Image Generation and Classification via a Diffusion Energy-Based Model · ICCV 2023
Computer vision › Image recognition and object detection
image classification
0.712023
EGC: Image Generation and Classification via a Diffusion Energy-Based Model · ICCV 2023
Machine learning › Deep learning architectures and training
mixture of experts
0.712023
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.712023
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths · NeurIPS 2023
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization
0.612022
Scale-Equivalent Distillation for Semi-Supervised Object Detection · CVPR 2022
Computer vision › Video understanding and tracking › video classification
efficient video recognition
0.612022
Rethinking Resolution in the Context of Efficient Video Recognition · NeurIPS 2022
Computer vision › Image recognition and object detection
object detection
0.612022
Scale-Equivalent Distillation for Semi-Supervised Object Detection · CVPR 2022
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection
0.612022
Scale-Equivalent Distillation for Semi-Supervised Object Detection · CVPR 2022
Machine learning › Efficient and distributed learning
collaborative learning
0.412020
Online Knowledge Distillation via Collaborative Learning · CVPR 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation › online knowledge distillation
mutual learning
0.412020
Online Knowledge Distillation via Collaborative Learning · CVPR 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
online knowledge distillation
0.412020
Online Knowledge Distillation via Collaborative Learning · CVPR 2020
Machine learning › Deep learning architectures and training › normalization
batch normalization
0.412019
Dynamic Recursive Neural Network · CVPR 2019
Machine learning › Efficient and distributed learning
dynamic neural network
0.412019
Dynamic Recursive Neural Network · CVPR 2019
Machine learning › Deep learning architectures and training
normalization
0.412019
Dynamic Recursive Neural Network · CVPR 2019
Computer vision › 3D vision
depth estimation
0.312025
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception · NeurIPS 2025
Computer vision › Video understanding and tracking
video classification
0.212022
Rethinking Resolution in the Context of Efficient Video Recognition · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.4visual tokenizer · 1.7segmented noise scheduling · 1.7rectified flow · 1.7next-token prediction · 1.7memory bank · 1.7disentangled causal attention · 1.7knowledge distillation · 1.1region caption data generation · 0.8instruction prompting · 0.8
YearPublicationVenuePosition
2025 TEVLA: Text-oriented Enhancement for Vision-Language Alignment in Relation Extraction
abstract
With the explosive growth of multimedia data storage, multimodal learning is an inevitable trend for Information Extraction (IE). However, the noise and irrelevance of web- crawled samples cause adverse degradation in each modality. Additionally, previous researches inadequately address the above issue, and the potential of cross-modal fusion remains underexplored. We propose a strengthened alignment module, using a generative text augmentation submodule to reduce noise contamination and emphasize incorporating visual features into texts. We further propose a fusion adapter utilizing a soft-prompt structure for profound fusion. To activate logical capabilities, we apply prompts with a multi-turn dialogue. For the Multimodal Relation Extraction task over the MNRE dataset, our method exceeds the previous SOTA model with a 7% increase in F1-score. It has superior generalization for other multimodal IE tasks, achieving SOTA on Named Entity Recognition over Twitter-15/17 datasets and on Event Extraction over M2E2dataset.
Junlin Chen, Qiushan Guo, Ka Chun Cheung, Mingrui Liang, Dezhi Chen
ICME2
2025 WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
abstract
Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these issues, we introduce WorldWeaver, a robust framework for long video generation that jointly models RGB frames and perceptual conditions within a unified long-horizon modeling scheme. Our training framework offers three key advantages. First, by jointly predicting perceptual conditions and color information from a unified representation, it significantly enhances temporal consistency and motion dynamics. Second, by leveraging depth cues, which we observe to be more resistant to drift than RGB, we construct a memory bank that preserves clearer contextual information, improving quality in long-horizon video generation. Third, we employ segmented noise scheduling for training prediction groups, which further mitigates drift and reduces computational cost. Extensive experiments on both diffusion and rectified flow-based models demonstrate the effectiveness of WorldWeaver in reducing temporal drift and improving the fidelity of generated videos.
Xueqing Deng, Shoufa Chen, Angtian Wang, Qiushan Guo, Mingfei Han 0003, Zeyue Xue, Mengzhao Chen, Ping Luo 0002
NeurIPS5
2025 OmniGen-AR: AutoRegressive Any-to-Image Generation
abstract
Autoregressive (AR) models have demonstrated strong potential in visual generation, offering competitive performance with simple architectures and optimization objectives. However, existing methods are typically limited to single-modality conditions, \eg, text or category labels, restricting their applicability in real-world scenarios that demand image synthesis from diverse forms of controls. In this work, we present \system, the first unified autoregressive framework for Any-to-Image generation. By discretizing various visual conditions through a shared visual tokenizer and text prompts with a text tokenizer, \system supports a broad spectrum of conditional inputs within a single model, including text (text-to-image generation), spatial signals (segmentation-to-image and depth-to-image), and visual context (image editing, frame prediction, and text-to-video generation). To mitigate the risk of information leakage from condition tokens to content tokens, we introduce Disentangled Causal Attention (DCA), which separates the full-sequence causal mask into condition causal attention and content causal attention. It serves as a training-time regularizer without affecting the standard next-token prediction during inference. With this design, \system achieves new state-of-the-art results across a range of benchmark, \eg, 0.63 on GenEval and 80.02 on VBench, demonstrating its effectiveness in flexible and high-fidelity visual generation.
Qiushan Guo, Peize Sun, Zuxuan Wu, Yu-Gang Jiang 0001
NeurIPS3
2024 RegionGPT: Towards Region Understanding Vision Language Model
abstract
Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder, and the use of coarse-grained training data that lacks detailed, region-specific captions. To address this, we introduce RegionGPT (short as RGPT), a novel framework designed for complex region-level captioning and understanding. RGPT enhances the spatial awareness of regional representation with simple yet effective modifications to existing visual encoders in VLMs. We further improve performance on tasks requiring a specific output scope by integrating task-guided instruction prompts during both training and inference phases, while maintaining the model's versatility for general-purpose tasks. Additionally, we develop an automated region caption data generation pipeline, enriching the training set with detailed region-level captions. We demonstrate that a universal RGPT model can be effectively applied and significantly enhancing performance across a range of region-level tasks, including but not limited to complex region descriptions, reasoning, object classification, and referring expressions comprehension. Code will be released at the project page.
Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon, Ka Chun Cheung, Yizhou Yu, Ping Luo 0002, Sifei Liu
CVPR1
2024 SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
abstract
Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT) to enhance VLMs’ spatial perception and reasoning capabilities. SpatialRGPT advances VLMs’ spatial understanding through two key innovations: (i) a data curation pipeline that enables effective learning of regional representation from 3D scene graphs, and (ii) a flexible ``plugin'' module for integrating depth information into the visual encoder of existing VLMs. During inference, when provided with user-specified region proposals, SpatialRGPT can accurately perceive their relative directions and distances. Additionally, we propose SpatialRGBT-Bench, a benchmark with ground-truth 3D annotations encompassing indoor, outdoor, and simulated environments, for evaluating 3D spatial cognition in Vision-Language Models (VLMs). Our results demonstrate that SpatialRGPT significantly enhances performance in spatial reasoning tasks, both with and without local region prompts. The model also exhibits strong generalization capabilities, effectively reasoning about complex spatial relations and functioning as a region-aware dense reward annotator for robotic tasks. Code, dataset, and benchmark are released at https://www.anjiecheng.me/SpatialRGPT.
An-Chieh Cheng, Hongxu Yin, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang 0004, Sifei Liu
NeurIPS4
2023 EGC: Image Generation and Classification via a Diffusion Energy-Based Model
abstract
Learning image classification and image generation using the same set of network parameters presents a formidable challenge. Recent advanced approaches perform well in one task often exhibit poor performance in the other. This work introduces an energy-based classifier and generator, namely EGC, which can achieve superior performance in both tasks using a single neural network. Unlike conventional classifiers that produce a label given an image (i.e., a conditional distribution p(y|x)), the forward pass in EGC is a classification model that yields a joint distribution p(x,y), enabling a diffusion model in its backward pass by marginalizing out the label y to estimate the score function. Furthermore, EGC can be adapted for unsupervised learning by considering the label as latent variables. EGC achieves competitive generation results compared with state-of-the-art approaches on ImageNet-1k, CelebA-HQ and LSUN Church, while achieving superior classification accuracy and robustness against adversarial attacks on CIFAR-10. This work marks the inaugural success in mastering both domains using a unified network parameter set. We believe that EGC bridges the gap between discriminative and generative learning. Code will be released at https://github.com/GuoQiushan/EGC.
Qiushan Guo, Chuofan Ma, Yi Jiang 0009, Zehuan Yuan, Yizhou Yu, Ping Luo 0002
ICCV1
2023 RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
abstract
Text-to-image generation has recently witnessed remarkable achievements. We introduce a text-conditional image diffusion model, termed RAPHAEL, to generate highly artistic images, which accurately portray the text prompts, encompassing multiple nouns, adjectives, and verbs. This is achieved by stacking tens of mixture-of-experts (MoEs) layers, i.e., space-MoE and time-MoE layers, enabling billions of diffusion paths (routes) from the network input to the output. Each path intuitively functions as a "painter" for depicting a particular textual concept onto a specified image region at a diffusion timestep. Comprehensive experiments reveal that RAPHAEL outperforms recent cutting-edge models, such as Stable Diffusion, ERNIE-ViLG 2.0, DeepFloyd, and DALL-E 2, in terms of both image quality and aesthetic appeal. Firstly, RAPHAEL exhibits superior performance in switching images across diverse styles, such as Japanese comics, realism, cyberpunk, and ink illustration. Secondly, a single model with three billion parameters, trained on 1,000 A100 GPUs for two months, achieves a state-of-the-art zero-shot FID score of 6.61 on the COCO dataset. Furthermore, RAPHAEL significantly surpasses its counterparts in human evaluation on the ViLG-300 benchmark. We believe that RAPHAEL holds the potential to propel the frontiers of image generation research in both academia and industry, paving the way for future breakthroughs in this rapidly evolving field. More details can be found on a webpage: https://raphael-painter.github.io/.
Zeyue Xue, Guanglu Song, Qiushan Guo, Boxiao Liu, Zhuofan Zong, Yu Liu 0015, Ping Luo 0002
NeurIPS3
2022 Scale-Equivalent Distillation for Semi-Supervised Object Detection
abstract
Recent Semi-Supervised Object Detection (SS-OD) methods are mainly based on self-training, i.e., generating hard pseudo-labels by a teacher model on unlabeled data as supervisory signals. Although they achieved certain success, the limited labeled data in semi-supervised learning scales up the challenges of object detection. We analyze the challenges these methods meet with the empirical experiment results. We find that the massive False Negative samples and inferior localization precision lack consideration. Besides, the large variance of object sizes and class imbalance (i.e., the extreme ratio between back-ground and object) hinder the performance of prior arts. Further, we overcome these challenges by introducing a novel approach, Scale-Equivalent Distillation (SED), which is a simple yet effective end-to-end knowledge distillation framework robust to large object size variance and class imbalance. SED has several appealing benefits compared to the previous works. (1) SED imposes a consistency regularization to handle the large scale variance problem. (2) SED alleviates the noise problem from the False Negative samples and inferior localization precision. (3) A re-weighting strategy can implicitly screen the potential foreground regions of the unlabeled data to reduce the effect of class imbalance. Extensive experiments show that SED consistently outperforms the recent state-of-the-art methods on different datasets with significant margins. For example, it surpasses the supervised counterpart by more than 10 mAP when using 5% and 10% labeled data on MS-COCO.
Qiushan Guo, Yao Mu 0001, Jianyu Chen 0002, Yizhou Yu, Ping Luo 0002
CVPR1
2022 Rethinking Resolution in the Context of Efficient Video Recognition
abstract
In this paper, we empirically study how to make the most of low-resolution frames for efficient video recognition. Existing methods mainly focus on developing compact networks or alleviating temporal redundancy of video inputs to increase efficiency, whereas compressing frame resolution has rarely been considered a promising solution. A major concern is the poor recognition accuracy on low-resolution frames. We thus start by analyzing the underlying causes of performance degradation on low-resolution frames. Our key finding is that the major cause of degradation is not information loss in the down-sampling process, but rather the mismatch between network architecture and input scale. Motivated by the success of knowledge distillation (KD), we propose to bridge the gap between network and input size via cross-resolution KD (ResKD). Our work shows that ResKD is a simple but effective method to boost recognition accuracy on low-resolution frames. Without bells and whistles, ResKD considerably surpasses all competitive methods in terms of efficiency and accuracy on four large-scale benchmark datasets, i.e., ActivityNet, FCVID, Mini-Kinetics, Something-Something V2. In addition, we extensively demonstrate its effectiveness over state-of-the-art architectures, i.e., 3D-CNNs and Video Transformers, and scalability towards super low-resolution frames. The results suggest ResKD can serve as a general inference acceleration method for state-of-the-art video recognition. Our code will be available at https://github.com/CVMI-Lab/ResKD.
Chuofan Ma, Qiushan Guo, Yi Jiang 0009, Ping Luo 0002, Zehuan Yuan, Xiaojuan Qi 0001
NeurIPS2
2020 Online Knowledge Distillation via Collaborative Learning
abstract
This work presents an efficient yet effective online Knowledge Distillation method via Collaborative Learning, termed KDCL, which is able to consistently improve the generalization ability of deep neural networks (DNNs) that have different learning capacities. Unlike existing two-stage knowledge distillation approaches that pre-train a DNN with large capacity as the ''teacher'' and then transfer the teacher's knowledge to another ''student'' DNN unidirectionally (i.e. one-way), KDCL treats all DNNs as ''students'' and collaboratively trains them in a single stage (knowledge is transferred among arbitrary students during collaborative training), enabling parallel computing, fast computations, and appealing generalization ability. Specifically, we carefully design multiple methods to generate soft target as supervisions by effectively ensembling predictions of students and distorting the input images. Extensive experiments show that KDCL consistently improves all the ''students'' on different datasets, including CIFAR-100 and ImageNet. For example, when trained together by using KDCL, ResNet-50 and MobileNetV2 achieve 78.2% and 74.0% top-1 accuracy on ImageNet, outperforming the original results by 1.4% and 2.0% respectively. We also verify that models pre-trained with KDCL transfer well to object detection and semantic segmentation on MS COCO dataset. For instance, the FPN detector is improved by 0.9% mAP.
Qiushan Guo, Xinjiang Wang, Yichao Wu, Ding Liang, Xiaolin Hu 0001, Ping Luo 0002
CVPR1
2020 Companion Guided Soft Margin for Face Recognition
Yingcheng Su, Yichao Wu, Zhenmao Li, Qiushan Guo, Ding Liang, Xiaolin Hu 0001
ECML/PKDD (3)4
2019 Dynamic Recursive Neural Network
abstract
This paper proposes the dynamic recursive neural network (DRNN), which simplifies the duplicated building blocks in deep neural network. Different from forwarding through different blocks sequentially in previous networks, we demonstrate that the DRNN can achieve better performance with fewer blocks by employing block recursively. We further add a gate structure to each block, which can adaptively decide the loop times of recursive blocks to reduce the computational cost. Since the recursive networks are hard to train, we propose the Loopy Variable Batch Normalization (LVBN) to stabilize the volatile gradient. Further, we improve the LVBN to correct statistical bias caused by the gate structure. Experiments show that the DRNN reduces the parameters and computational cost and while outperforms the original model in term of the accuracy consistently on CIFAR-10 and ImageNet-1k. Lastly we visualize and discuss the relation between image saliency and the number of loop time.
Qiushan Guo, Yichao Wu, Ding Liang, Haoyu Qin
CVPR1
2018 MSFD: Multi-Scale Receptive Field Face Detector
abstract
We aim to study the multi-scale receptive fields of a single convolutional neural network to detect faces of varied scales. This paper presents our Multi-Scale Receptive Field Face Detector (MSFD), which has superior performance on detecting faces at different scales and enjoys real-time inference speed. MSFD agglomerates context and texture by hierarchical structure. More additional information and rich receptive field bring significant improvement but generate marginal time consumption. We simultaneously propose an anchor assignment strategy which can cover faces with a wide range of scales to improve the recall rate of small faces and rotated faces. To reduce the false positive rate, we train our detector with focal loss which keeps the easy samples from overwhelming. As a result, MSFD reaches superior results on the FDDB, Pascal-Faces and WIDER FACE datasets, and can run at 31 FPS on GPU for VGA-resolution images.
Qiushan Guo, Hongliang Bai
ICPR1