Longtian Qiu

dblp:308/0937 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Vision and language · 33% Generative modeling · 22% Representation and self-supervised learning · 11%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.832025
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models · ICML 2024
SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models · ECCV (62) 2024
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation · NeurIPS 2025
Machine learning › Reinforcement learning › value function estimation
advantage estimation
0.912025
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation · ICLR 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation · ICLR 2025
Machine learning › Generative modeling
flow matching
0.912025
Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation · ICLR 2025
Computer vision › Vision and language › multimodal reasoning
multimodal chain-of-thought reasoning
0.912025
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation · ICLR 2025
Computer vision › Vision and language
cross-modal alignment
0.812024
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training · AAAI 2024
Natural language and speech › Language models and text generation
large language model training
0.812024
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models · ICML 2024
Machine learning › Deep learning architectures and training
mixture of experts
0.812024
SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models · ECCV (62) 2024
Computer vision › Vision and language › image captioning › low-shot image captioning
zero-shot image captioning
0.812024
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training · AAAI 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked autoencoder
cross-modal masked autoencoding
0.712023
Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training · IJCAI 2023
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.712023
Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training · IJCAI 2023
Computer vision › Image recognition and object detection
human-object interaction detection
0.712023
HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models · CVPR 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.712023
Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training · IJCAI 2023
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud pre-training
0.712023
Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training · IJCAI 2023
Computer vision › Vision and language
vision-language knowledge transfer
0.712023
HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models · CVPR 2023
Computer vision › Vision and language
vision-language model
0.712023
CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free Attention · AAAI 2023
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.712023
CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free Attention · AAAI 2023
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.712023
CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free Attention · AAAI 2023
Machine learning › Generative modeling › video generation
text-to-video generation
0.312025
Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation · ICLR 2025
Computer vision › Image recognition and object detection
visual recognition
0.212024
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models · ICML 2024
Computer vision › Vision and language › visual question answering
zero-shot visual question answering
0.212024
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training · AAAI 2024

Methods — techniques the papers use, named apart from their topics

contrastive language-image pretraining · 2.1noise injection · 1.6zero-initialized attention · 0.9flow matching · 0.9bayesian inference · 0.9RoPE · 0.9KQ-Norm · 0.9GRPO · 0.9text-only training · 0.8reranking · 0.8
YearPublicationVenuePosition
2025 Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
abstract
Sora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based Large Diffusion Transformers (Flag-DiT) equipped with zero-initialized attention, as a simple and scalable generative framework that can be adapted to various modalities, e.g., transforming noise into images, videos, multi-view 3D objects, or audio clips conditioned on text instructions. By tokenizing the latent spatial-temporal space and incorporating learnable placeholders such as |[nextline]| and |[nextframe]| tokens, Lumina-T2X seamlessly unifies the representations of different modalities across various spatial-temporal resolutions. Advanced techniques like RoPE, KQ-Norm, and flow matching enhance the stability, flexibility, and scalability of Flag-DiT, enabling models of Lumina-T2X to scale up to 7 billion parameters and extend the context window to 128K tokens. This is particularly beneficial for creating ultra-high-definition images with our Lumina-T2I model and long 720p videos with our Lumina-T2V model. Remarkably, Lumina-T2I, powered by a 5-billion-parameter Flag-DiT, requires only 35% of the training computational costs of a 600-million-parameter naive DiT (PixArt-alpha), indicating that increasing the number of parameters significantly accelerates convergence of generative models without compromising visual quality. Our further comprehensive analysis underscores Lumina-T2X's preliminary capability in resolution extrapolation, high-resolution editing, generating consistent 3D views, and synthesizing videos with seamless transitions. All code and checkpoints of Lumina-T2X are released at https://github.com/Alpha-VLLM/Lumina-T2X to further foster creativity, transparency, and diversity in the generative AI community.
Peng Gao 0007, Le Zhuo, Ruoyi Du, Longtian Qiu, Rongjie Huang 0001, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang 0001, Tianshuo Yang, Weicai Ye, Tong He 0001, Jingwen He, Junjun He, Yu Qiao 0001, Hongsheng Li 0001
ICLR6
2025 NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
abstract
Reinforcement learning (RL) has shown promise in enhancing the general Chain-of-Thought (CoT) reasoning capabilities of multimodal large language models (MLLMs). However, when applied to improve general CoT reasoning, existing RL frameworks often struggle to generalize beyond the training distribution. To address this, we propose NoisyGRPO, a systematic multimodal RL framework that introduces controllable noise into visual inputs for enhanced exploration and explicitly models the advantage estimation process via a Bayesian framework. Specifically, NoisyGRPO improves RL training by: (1) \textbf{Noise-Injected Exploration Policy}: Perturbing visual inputs with Gaussian noise to encourage exploration across a wider range of visual scenarios; and (2) \textbf{Bayesian Advantage Estimation}: Formulating advantage estimation as a principled Bayesian inference problem, where the injected noise level serves as a prior and the observed trajectory reward as the likelihood. This Bayesian modeling fuses both sources of information to compute a robust posterior estimate of trajectory advantage, effectively guiding MLLMs to prefer visually grounded trajectories over noisy ones. Experiments on standard CoT quality, general capability, and hallucination benchmarks demonstrate that NoisyGRPO substantially improves generalization and robustness, especially in RL settings with small-scale MLLMs such as Qwen2.5-VL 3B.
Longtian Qiu, Shan Ning, Jiaxuan Sun
NeurIPS1
2024 Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
abstract
Image captioning aims at generating descriptive and meaningful textual descriptions of images, enabling a broad range of vision-language applications. Prior works have demonstrated that harnessing the power of Contrastive Image Language Pre-training (CLIP) offers a promising approach to achieving zero-shot captioning, eliminating the need for expensive caption annotations. However, the widely observed modality gap in the latent space of CLIP harms the performance of zero-shot captioning by breaking the alignment between paired image-text features. To address this issue, we conduct an analysis on the CLIP latent space which leads to two findings. Firstly, we observe that the CLIP's visual feature of image subregions can achieve closer proximity to the paired caption due to the inherent information loss in text descriptions. In addition, we show that the modality gap between a paired image-text can be empirically modeled as a zero-mean Gaussian distribution. Motivated by the findings, we propose a novel zero-shot image captioning framework with text-only training to reduce the modality gap. In particular, we introduce a subregion feature aggregation to leverage local region information, which produces a compact visual representation for matching text representation. Moreover, we incorporate a noise injection and CLIP reranking strategy to boost captioning performance. We also extend our framework to build a zero-shot VQA pipeline, demonstrating its generality. Through extensive experiments on common captioning and VQA datasets such as MSCOCO, Flickr30k and VQAV2, we show that our method achieves remarkable performance improvements. Code is available at https://github.com/Artanic30/MacCap.
Longtian Qiu, Shan Ning, Xuming He 0001
AAAI1
2024 SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
Renrui Zhang, Peng Gao 0007, Longtian Qiu, Han Xiao 0010, Han Qiu 0010, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang 0004, Xuming He 0001, Yu Qiao 0001, Hongsheng Li 0001
ECCV (62)5
2024 SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
abstract
We propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8$\times$7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory.
Renrui Zhang, Longtian Qiu, Siyuan Huang 0004, Weifeng Lin, Shitian Zhao, Shijie Geng, Kaipeng Zhang, Wenqi Shao, Conghui He, Junjun He, Hao Shao, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007
ICML3
2023 CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free Attention
abstract
Contrastive Language-Image Pre-training (CLIP) has been shown to learn visual representations with promising zero-shot performance. To further improve its downstream accuracy, existing works propose additional learnable modules upon CLIP and fine-tune them by few-shot training sets. However, the resulting extra training cost and data requirement severely hinder the efficiency for model deployment and knowledge transfer. In this paper, we introduce a free-lunch enhancement method, CALIP, to boost CLIP's zero-shot performance via a parameter-free attention module. Specifically, we guide visual and textual representations to interact with each other and explore cross-modal informative features via attention. As the pre-training has largely reduced the embedding distances between two modalities, we discard all learnable parameters in the attention and bidirectionally update the multi-modal features, enabling the whole process to be parameter-free and training-free. In this way, the images are blended with textual-aware signals and the text representations become visual-guided for better adaptive zero-shot alignment. We evaluate CALIP on various benchmarks of 14 datasets for both 2D image and 3D point cloud few-shot classification, showing consistent zero-shot performance improvement over CLIP. Based on that, we further insert a small number of linear layers in CALIP's attention module and verify our robustness under the few-shot settings, which also achieves leading performance compared to existing methods. Those extensive experiments demonstrate the superiority of our approach for efficient enhancement of CLIP. Code is available at https://github.com/ZiyuGuo99/CALIP.
Renrui Zhang, Longtian Qiu, Xianzheng Ma, Xupeng Miao, Xuming He 0001, Bin Cui 0001
AAAI3
2023 HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models
abstract
Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Recently, Contrastive Language-Image Pre-training (CLIP) has shown great potential in providing interaction prior for HOI detectors via knowledge distillation. However, such approaches often rely on large-scale training data and suffer from inferior performance under few/zero-shot scenarios. In this paper, we propose a novel HOI detection framework that efficiently extracts prior knowledge from CLIP and achieves better generalization. In detail, we first introduce a novel interaction decoder to extract informative regions in the visual feature map of CLIP via a cross-attention mechanism, which is then fused with the detection backbone by a knowledge integration block for more accurate human- object pair detection. In addition, prior knowledge in CLIP text encoder is leveraged to generate a classifier by embedding HOI descriptions. To distinguish fine-grained interactions, we build a verb classifier from training data via visual semantic arithmetic and a lightweight verb representation adapter. Furthermore, we propose a training-free enhancement to exploit global HOI predictions from CLIP. Extensive experiments demonstrate that our method outperforms the state of the art by a large margin on various settings, e.g. +4.04 mAP on HICO-Det. The source code is available in https://github.com/Artanic30/HOICLIP.
Shan Ning, Longtian Qiu, Yongfei Liu, Xuming He 0001
CVPR2
2023 Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training
abstract
Masked Autoencoders (MAE) have shown promising performance in self-supervised learning for both 2D and 3D computer vision. However, existing MAE-style methods can only learn from the data of a single modality, i.e., either images or point clouds, which neglect the implicit semantic and geometric correlation between 2D and 3D. In this paper, we explore how the 2D modality can benefit 3D masked autoencoding, and propose Joint-MAE, a 2D-3D joint MAE framework for self-supervised 3D point cloud pre-training. Joint-MAE randomly masks an input 3D point cloud and its projected 2D images, and then reconstructs the masked information of the two modalities. For better cross-modal interaction, we construct our JointMAE by two hierarchical 2D-3D embedding modules, a joint encoder, and a joint decoder with modal-shared and model-specific decoders. On top of this, we further introduce two cross-modal strategies to boost the 3D representation learning, which are local-aligned attention mechanisms for 2D-3D semantic cues, and a cross-reconstruction loss for 2D-3D geometric constraints. By our pre-training paradigm, Joint-MAE achieves superior performance on multiple downstream tasks, e.g., 92.4% accuracy for linear SVM on ModelNet40 and 86.07% accuracy on the hardest split of ScanObjectNN.
Renrui Zhang, Longtian Qiu, Xianzhi Li 0001, Pheng-Ann Heng
IJCAI3